OpenAI Just Made Its Most Powerful AI Model 20% Cheaper to Use

OpenAI has lowered the price of using its GPT-5.6 Sol model by more than 20% for a three-month period, with the new rates reflected in its developer pricing documentation as of August 24, 2026. Reuters reported the cut on August 21, 2026, and OpenAI's own GPT-5.6 product page carries an update of the same date announcing the reduction. The lower baseline pricing now sits at $4.00 per 1 million input tokens and $20.00 per 1 million output tokens, down from prior rates of $5.00 and $30.00 respectively, per OpenAI's model documentation.
A "token" is a small chunk of text, roughly four characters. Think of it as the unit of currency for AI: you pay per million tokens that go into the model (input) and per million tokens that come out (output). If you send the model a page of text to summarize, that page is measured in input tokens. The summary it writes back is measured in output tokens.
The current pricing structure on OpenAI's developer pricing page, dated August 24, 2026, organizes costs along two things: how fast you need the response, and how much text the model has to process at once. Three speed options are available. Standard mode is the baseline. Fast mode, renamed from "Priority processing" on July 30, 2026, offers quicker responses at a higher price. Batch and Flex modes provide the lowest rates for work that does not need an immediate answer. API requests can still use either the priority or fast value for the service_tier parameter. Each mode is further split by short and long amounts of text.
In Standard mode with a shorter amount of text, GPT-5.6 Sol is priced at $4.00 per 1M input tokens, $0.40 per 1M cached input tokens, $5.00 per 1M cache write tokens, and $20.00 per 1M output tokens. "Cached input" means the model has already seen part of your prompt and saved it, so it charges you less to process that part again. "Cache write" is the cost of saving your prompt for future reuse. Long-text Standard pricing doubles input costs to $8.00 per 1M tokens, raises cached input to $0.80 per 1M, drops cache write to $10.00 per 1M, and sets output at $30.00 per 1M.
Fast mode costs twice as much as Standard across the board. Short-text Fast pricing is $8.00 per 1M input, $0.80 per 1M cached input, $10.00 per 1M cache write, and $40.00 per 1M output. Long-text Fast pricing reaches $16.00 per 1M input, $1.60 per 1M cached input, $20.00 per 1M cache write, and $60.00 per 1M output.
Batch and Flex modes cut the Standard rates in half. Short-text pricing is $2.00 per 1M input, $0.20 per 1M cached input, $2.50 per 1M cache write, and $10.00 per 1M output. Long-text pricing is $4.00 per 1M input, $0.40 per 1M cached input, $5.00 per 1M cache write, and $15.00 per 1M output. All figures are sourced from OpenAI's API pricing documentation.
The price reduction narrows the gap between GPT-5.6 Sol's launch pricing and its current rates. OpenAI's launch announcement, "Previewing GPT-5.6 Sol," initially priced the model family across three sizes: Sol at $5 input / $30 output, Terra at $2.50 input / $15 output, and Luna at $1 input / $6 output, per OpenAI's launch page. The August 21 update also reduced GPT-5.6 Terra pricing by 20%.
The Sol and Cyber variants occupy different tiers. GPT-5.6 Cyber, exposed through the daybreak-red-latest API alias, is priced substantially higher in Standard mode with short text: $12.50 per 1M input tokens, $1.25 per 1M cached input, $15.625 per 1M cache write, and $75.00 per 1M output. GPT-5.6 Sol is accessible via the daybreak-blue-latest alias.
Prompt caching for GPT-5.6 and later models requires a strict minimum of 1,024 tokens, down from the 1,024-to-2,048 range that applied to earlier models, per OpenAI's prompt caching guide. The 10x discount on cached input reads relative to fresh input tokens holds across all modes and text lengths.
One structural cost factor worth noting: regional processing endpoints, which provide data residency, carry a 10% uplift for OpenAI models released on or after March 5, 2026. GPT-5.6 Sol and Cyber both fall within that window, so any workload routing through regional endpoints will see the listed rates plus that surcharge.
The broader context here is that OpenAI's pricing now spans a 3x range between the cheapest and fastest modes for the same model, with the fastest long-text option reaching $60 per 1M output tokens. The cache write pricing is notably inconsistent across tiers. In Standard mode, cache writes are cheaper at long text ($10.00) than at short text ($5.00 per 1M for short, effectively $10.00 for long), while in Fast and Batch/Flex modes the long-text cache write is exactly double the short-text rate. This suggests the write cost is being tied to the input-token rate rather than treated as a fixed overhead, which has implications for workloads that frequently rebuild large prompt caches.
For teams running high-volume work, the arithmetic favors routing anything that does not need an immediate answer through Batch or Flex whenever possible. A workload doing 1M input and 1M output tokens in short text costs $24 in Standard, $48 in Fast, and $12 in Batch/Flex. The 4x spread between Batch and Fast for identical volumes is the widest gap in the current lineup, and it reflects the degree to which OpenAI is pricing speed as a premium feature rather than a baseline expectation.
The three-month duration of the Sol reduction, announced August 21, positions this as a promotional adjustment rather than a permanent repricing. Whether the current rates persist beyond that window is not yet indicated in the published documentation.


