· via dev.to (home feed)
OpenAI's GPT-6.1 Sol halves cached input pricing and drops none reasoning effort
OpenAI released GPT-6.1 Sol at DevDay 2026 with cached input cut from $0.20 to $0.10 per million tokens. Migration from gpt-6-sol is mostly a model ID swap, except reasoning.effort no longer accepts none or minimal.

What OpenAI shipped
OpenAI introduced GPT-6.1 Sol at its DevDay event on September 29, 2026, according to developer guides published on dev.to. The release is an incremental update to GPT-6 Sol rather than a new model family: standard pricing, context window, endpoints, and rate limits all carry over unchanged. The two substantive changes are a halved cached-input price and a tightened reasoning-effort API that removes two previously supported values.
Pricing: one number moves
The dev.to guides, consolidating OpenAI's model and pricing pages, report the following per-million-token rates. Standard input stays at $2 and output at $10, identical to GPT-6 Sol. Cached input drops from $0.20 to $0.10, while cache writes remain $2.50. The context window, maximum input, and maximum output are unchanged at 1,050,000, 922,000, and 128,000 tokens respectively.
The tiered pricing keeps the same shape, per the dev.to breakdown: Batch and Flex both run $1 input, $0.05 cached input, $1.25 cache writes, and $5 output; Fast runs $4 input, $0.20 cached, $5 writes, and $20 output. One surcharge applies to very long prompts: requests exceeding 272,000 input tokens are billed at double the input and cache rates and 1.5 times the output rate for the entire request.
Rate limits are untouched, with Tier 1 at 500 requests per minute and 500,000 tokens per minute, and Tier 5 at 15,000 RPM and 40 million TPM.
The breaking change: reasoning.effort
The main compatibility issue for existing code is that GPT-6.1 Sol no longer accepts none or minimal as values for reasoning.effort. OpenAI's guidance, as relayed by the dev.to guides, is to map both to low and re-evaluate any pipeline that depended on none. The accepted values are now low, medium (the default), high, xhigh, and max.
This interacts with function calling. In GPT-6 Sol, Chat Completions supported tool calls only when reasoning_effort was set to none. GPT-6.1 Sol has no equivalent combination, so any workflow using tools must move to the Responses API; Chat Completions remains available only for requests without tools.
Developers should also strip sampling parameters when effort is not none: temperature, top_p, top_logprobs, and logprobs in Chat Completions. Code that previously combined temperature with reasoning_effort none needs that pairing removed.
One more behavioural shift: the knowledge cutoff moves from April 20 to April 30, 2026, so date-sensitive evaluations should be rerun after migrating.
Migration in four steps
The dev.to guides reduce the move from gpt-6-sol to four code changes:
- Swap the model ID to gpt-6.1-sol, keeping it in configuration or an environment variable so rollback is a single edit.
- Remap none and minimal to low, then run representative evaluations before shifting traffic.
- Remove the incompatible sampling parameters listed above.
- Move tool-calling flows from Chat Completions to the Responses API.
The guides also flag operational details worth knowing: requests go to the /v1/responses endpoint with the model string gpt-6.1-sol; OpenAI's reasoning documentation suggests reserving at least 25,000 tokens for reasoning and output while experimenting; and responses can return as incomplete with reason max_output_tokens before producing visible text. To change effort mid-conversation without invalidating the prompt cache, the guides recommend a configuration_update input item rather than changing reasoning.effort at the request level. Usage reporting exposes cached_tokens, cache_write_tokens, and reasoning_tokens, which makes the cheaper cache directly measurable.
Reported benchmark results
The guides cite OpenAI's launch figures for GPT-6.1 Sol. At low effort, the rate of responses containing a factual error on user-flagged conversations falls from 11.4% with GPT-6 Sol to 7.7%. At medium, AutomationBench 1.0.6 shows a 2.2 percentage-point gain over Claude Opus 5.5 at roughly a third of the cost, and a 4.8-point gain over GPT-6 Sol in the same configuration. At max, OSWorld 2.0 improves 7 points over GPT-6 Sol for less than half the cost, and Terminal-Bench Science 0.1 costs $5.47 per task versus $23.21 for Opus 5.5 and $23.80 for GPT-6 Astra.
The guides add two caveats: the factuality set consists of previously flagged conversations rather than typical traffic, and GPT-6 Astra still scores highest on Terminal-Bench Science at 68.1%, with OpenAI recommending Astra for the hardest scientific work.
Why it matters
For teams already on GPT-6 Sol, this is a low-friction release with a real cost lever. Cached input at $0.10 per million tokens rewards prompt caching aggressively, which matters most for agentic and long-context workloads that resend large system prompts. The friction concentrates in one place: anything built on reasoning_effort none, especially Chat Completions function calling, requires rework rather than a string swap. Teams that re-evaluate at low effort and move tools to the Responses API get the price cut without surprises, while the unchanged rate limits and context window mean no capacity replanning is needed.
- #openai
- #llm
- #api
- #pricing
- #developer-tools