· via dev.to (home feed)
GPT-6.1 Sol keeps token prices flat, halves cached input, drops none effort
OpenAI's GPT-6.1 Sol API keeps standard token prices unchanged but halves cached-input cost to $0.10 per million tokens, and it drops the none reasoning setting many gpt-6-sol callers rely on.

OpenAI released GPT-6.1 Sol at its DevDay event on 29 September 2026, and the model is already reachable through the Responses API under the identifier gpt-6.1-sol. According to two dev.to guides published the following day, standard pricing carries over from gpt-6-sol unchanged — $2 per million input tokens and $10 per million output tokens — while cached input falls from $0.20 to $0.10 per million tokens. That cache reduction is the headline change for workloads that lean heavily on prompt caching.
Pricing and limits mostly carry over
Beyond the cache cut, dev.to reports that nearly every other specification is identical to the predecessor. Cache writes still cost $2.50 per million tokens, the context window remains 1,050,000 tokens with a 922,000-token input ceiling and a 128,000-token output ceiling, and rate limits are unchanged: tier 1 at 500 requests and 500,000 tokens per minute, tier 5 at 15,000 requests and 40 million tokens per minute. Chat Completions, Responses and Batch endpoints all remain available.
Two line items do move. The knowledge cutoff advances from 20 April 2026 to 30 April 2026, so evaluations sensitive to recency should be rerun after migration. And per the model page cited by dev.to, any request containing more than 272,000 input tokens is billed at twice the standard input and cache rates plus 1.5 times the output rate for the entire request.
Migration is a model-ID swap with real exceptions
For most codebases, upgrading amounts to replacing one string, and the guides recommend keeping the model ID in configuration or an environment variable so rollback is a single change. The exceptions need attention first:
- The reasoning.effort parameter no longer accepts none or minimal; both must be mapped to low. On GPT-6 Astra, sending none produces an HTTP 400, so the guides warn to fix this configuration before shifting production traffic.
- Sampling parameters — temperature, top_p and top_logprobs, plus logprobs on Chat Completions — must be removed whenever effort is not none.
- Function calling in Chat Completions is gone entirely. gpt-6-sol permitted it only alongside reasoning_effort none, a combination with no equivalent in the new model, so tool-based workflows have to move to the Responses API.
- Date-sensitive checks should be rerun given the knowledge-cutoff shift.
The GPT-6 Sol model page now points readers to GPT-6.1 Sol as the newer Sol model, a signal that the old ID is being phased down.
Reasoning effort remains the main dial
Effort, which defaults to medium, is described by dev.to as the primary control over cost, latency and quality. Benchmarks OpenAI reports for the new model, as summarised in the guides: at low effort, responses containing a factual error in previously flagged conversations dropped from 11.4 percent on GPT-6 Sol to 7.7 percent; at medium, AutomationBench 1.0.6 improved 2.2 percentage points over Claude Opus 5.5 at roughly a third of the cost, and 4.8 points over GPT-6 Sol at the same setting; at max, OSWorld 2.0 gained 7 points over GPT-6 Sol, and Terminal-Bench Science cost $5.47 per task versus $23.21 for Opus 5.5 and $23.80 for GPT-6 Astra.
The guides attach caveats: the factuality pool consists of conversations users had already marked as faulty rather than typical traffic, and GPT-6 Astra still leads Terminal-Bench Science at 68.1 percent, so OpenAI continues to recommend Astra for the hardest scientific work. For latency-sensitive callers that previously used none, the suggested path is to start at low and measure latency, cost and quality on representative prompts before committing.
Response fields worth watching
When handling responses, dev.to highlights several details. A status of completed signals success, while incomplete with incomplete_details.reason set to max_output_tokens means the output budget ran out, possibly before any visible text. The output field is an array, so code should locate the item whose type is message rather than rely on index positions. Billing-wise, usage.output_tokens includes reasoning tokens charged at the output rate, with the count broken out in output_tokens_details, while input_tokens_details exposes cached_tokens and cache_write_tokens — the place to verify the cheaper cache is actually being applied. OpenAI's reasoning guidance suggests reserving at least 25,000 tokens for reasoning and output while experimenting, and effort can be changed mid-conversation via a configuration_update input item instead of the request-level parameter, which preserves the prompt cache.
Why it matters
Halving cached-input pricing meaningfully changes the economics of large, repeated-context workloads in a model with a roughly one-million-token window, rewarding architectures that keep stable prefixes warm. At the same time, a migration that looks like a string swap hides two breaking changes — the removal of the none effort level and the loss of tool calling on Chat Completions — either of which can take a production pipeline down with client errors. With effort now spanning only low through max, and reasoning tokens billed as output, the effort setting and cache behaviour together dominate total cost, making measurement before and after the switch the real work of this upgrade.
- #openai
- #api
- #llm
- #pricing
- #migration