OpenAI Cuts GPT-5.6 Sol API Pricing by 20% in Three-Month Promotion

OpenAI has lowered API and credit pricing for its GPT-5.6 Sol model by more than 20% for a three-month period, with the new rates reflected in its developer pricing documentation as of August 24, 2026. Reuters reported the cut on August 21, 2026, and OpenAI's own GPT-5.6 product page carries an update of the same date announcing the reduction. The lower Standard-mode short-context pricing now sits at $4.00 per 1M input tokens and $20.00 per 1M output tokens, down from prior rates of $5.00 and $30.00 respectively, per OpenAI's model documentation.
For those less familiar with the terminology: a "token" is roughly a piece of a word, about four characters. API pricing is quoted per million tokens processed. "Input tokens" are what you send to the model; "output tokens" are what it generates back.
The current pricing structure on OpenAI's developer pricing page, dated August 24, 2026, organizes costs along two axes: processing mode and context length. Three modes are available. Standard mode serves as the baseline. Fast mode, renamed from "Priority processing" on July 30, 2026, offers lower latency at a premium. Batch and Flex modes provide the lowest rates for asynchronous or flexible-scheduling workloads. API requests can still use either the priority or fast value for the service_tier parameter. Each mode is further split by short-context and long-context rates. "Context length" refers to how many tokens the model considers at once; longer contexts cost more because they require more computation.
In Standard mode with short context, GPT-5.6 Sol is priced at $4.00 per 1M input tokens, $0.40 per 1M cached input tokens, $5.00 per 1M cache write tokens, and $20.00 per 1M output tokens. "Cached input" means the model has seen part of your prompt before and stored it, so it charges less to process it again. "Cache write" is the cost of storing that prompt for future reuse. Long-context Standard pricing doubles input costs to $8.00 per 1M tokens, raises cached input to $0.80 per 1M, drops cache write to $10.00 per 1M, and sets output at $30.00 per 1M.
Fast mode carries a consistent 2x premium over Standard across the board. Short-context Fast pricing is $8.00 per 1M input, $0.80 per 1M cached input, $10.00 per 1M cache write, and $40.00 per 1M output. Long-context Fast pricing reaches $16.00 per 1M input, $1.60 per 1M cached input, $20.00 per 1M cache write, and $60.00 per 1M output.
Batch and Flex modes halve the Standard rates. Short-context pricing is $2.00 per 1M input, $0.20 per 1M cached input, $2.50 per 1M cache write, and $10.00 per 1M output. Long-context pricing is $4.00 per 1M input, $0.40 per 1M cached input, $5.00 per 1M cache write, and $15.00 per 1M output. All figures are sourced from OpenAI's API pricing documentation.
The price reduction narrows the gap between GPT-5.6 Sol's launch pricing and its current rates. OpenAI's launch announcement, "Previewing GPT-5.6 Sol," initially priced the model family across three sizes: Sol at $5 input / $30 output, Terra at $2.50 input / $15 output, and Luna at $1 input / $6 output, per OpenAI's launch page. The August 21 update also reduced GPT-5.6 Terra pricing by 20%.
The Sol and Cyber variants occupy different tiers. GPT-5.6 Cyber, exposed through the daybreak-red-latest API alias, is priced substantially higher in Standard mode with short context: $12.50 per 1M input tokens, $1.25 per 1M cached input, $15.625 per 1M cache write, and $75.00 per 1M output. GPT-5.6 Sol is accessible via the daybreak-blue-latest alias.
Prompt caching for GPT-5.6 and later models requires a strict minimum of 1,024 tokens, down from the 1,024-to-2,048 range that applied to earlier models, per OpenAI's prompt caching guide. The 10x discount on cached input reads relative to fresh input tokens holds across all modes and context lengths.
One structural cost factor worth noting: regional processing endpoints, which provide data residency, carry a 10% uplift for OpenAI models released on or after March 5, 2026. GPT-5.6 Sol and Cyber both fall within that window, so any workload routing through regional endpoints will see the listed rates plus that surcharge.
The broader context here is that OpenAI's pricing now spans a 3x range between Batch/Flex and Fast modes for the same model, with long-context Fast reaching $60 per 1M output tokens. The cache write token pricing is notably inconsistent across tiers. In Standard mode, cache writes are cheaper at long context ($10.00) than at short context ($5.00 per 1M for short, effectively $10.00 for long), while in Fast and Batch/Flex modes the long-context cache write is exactly double the short-context rate. This suggests the write cost is being pegged to the input-token rate rather than treated as a fixed overhead, which has implications for workloads that frequently rebuild large prompt caches.
For teams running high-volume inference, the arithmetic favors routing non-latency-sensitive workloads through Batch or Flex wherever possible. A workload doing 1M input and 1M output tokens in short context costs $24 in Standard, $48 in Fast, and $12 in Batch/Flex. The 4x spread between Batch and Fast for identical token volumes is the widest gap in the current lineup, and it reflects the degree to which OpenAI is pricing latency as a premium product feature rather than a baseline expectation.
The three-month duration of the Sol reduction, announced August 21, positions this as a promotional adjustment rather than a permanent repricing. Whether the current rates persist beyond that window is not yet indicated in the published documentation.


