Technology

OpenAI's Two New AI Models Cost Less and Make Fewer Mistakes

Martin HollowayPublished 2w ago3 min readBased on 7 sources
Reading level
OpenAI's Two New AI Models Cost Less and Make Fewer Mistakes
source:openai.com

OpenAI introduced two new AI models, GPT-6 Sol and GPT-6 Luna. In a Sept. 16 post, the company said both belong to the GPT-6 model family. OpenAI

OpenAI said it trained Sol and Luna in a similar way as GPT-6 Astra. The price developers pay to use Sol and Luna through the API, the paid service that lets apps connect to the models, is 50% lower than the promotional price for GPT-5.6.

For Sol, OpenAI claimed fewer factual mistakes. The company said Sol makes about half as many mistakes as the prior model on an internal test based on real chats where users pointed out errors.

For work with AI agents, software that can do multi-step jobs with tools, OpenAI pointed to a test called AutomationBench. The company said GPT-6 Sol on its xhigh effort setting beats Claude Opus 5 on its max setting on that test, at 9% of the cost per task.

That price comparison goes back to GPT-5.6. That family had three versions: Sol as the main workhorse, Terra as the middle option, and Luna as the low-cost option. TechCrunch At first, Sol cost $5 for every million input tokens and $30 for every million output tokens, while Luna cost $1 for input and $6 for output. Tokens are small pieces of words and text. TechCrunch

OpenAI later cut those prices. It lowered the price for GPT-5.6 Sol by over 20% for three months and cut GPT-5.6 Luna by 80%, making Luna its cheapest model. OpenAI After the cut, Luna cost $0.20 per million input tokens and $1.20 per million output tokens. Yahoo Finance

That earlier launch was not open to all at first. OpenAI gave GPT-5.6 only to approved partners before wider release, and Reuters reported the company said Sol, Terra and Luna would launch on Thursday. Reuters Early descriptions presented Sol as a next-generation model for coding, science and cybersecurity, and Luna as a fast and low-cost model. OpenAI

The broader context here is that price shapes how developers build. When rates fall by half from an already lowered price, teams can give the model more text to read at once, like handing an assistant a thicker file, and ask it to do more steps and checks. That factuality test counts mistakes everyday users spotted in chats, not exam questions. The AutomationBench claim links effort level and cost for each finished job.

In my view, buyers should check two claims first: fewer mistakes and lower cost per job. Half as many user-reported mistakes, if true in many uses, means less time fixing output. That fixing time often blocks wider use. And a win at 9% of the cost changes the choice between building in-house and buying automation. Teams will want to test both with their own data, their own lists of failure types and retry limits. Worth flagging, company tests and company cost math are a starting point, not a replacement for that checking.

Looking at what this means for deployment, the practical result is simple. More thinking work per dollar, with fewer fixes needed later. That mix is what helps a new model stay in everyday use.