· via dev.to (home feed)
OpenAI ships GPT-6.1 Sol Ultrafast tier across API, Codex and ChatGPT Work
OpenAI is rolling out an Ultrafast service tier for GPT-6.1 Sol across its API, Codex and ChatGPT Work, with API pricing at 6x Standard rates and eligibility limited to Pro plans and select workspaces elsewhere.

What OpenAI shipped
OpenAI has started rolling out an "Ultrafast" service tier for GPT-6.1 Sol, bringing the speed-focused option to its API, the Codex coding product and ChatGPT Work. According to dev.to, the company positions the tier as its fastest serving option for the model, but access is gated by plan and workspace eligibility, and usage costs considerably more than the standard tier.
How developers enable it
In the API, the tier is exposed through the Responses API: developers select the gpt-6.1-sol model and set the service_tier parameter to ultrafast. OpenAI prices the API tier at six times the Standard rate, per dev.to's reading of the documentation.
There is no equivalent parameter in ChatGPT Work or Codex. There, Ultrafast simply becomes available for GPT-6.1 Sol on eligible accounts. Launch access is limited to Pro-level plans, listed at $500, and qualifying Enterprise and Edu workspaces. The rollout is described as ongoing rather than universal, so not every plan tier can use it yet.
Pricing differs by product
The cost structure is not uniform across platforms, which is the detail most likely to trip up teams budgeting for the feature:
- API: 6x Standard pricing, applied directly via the service-tier setting.
- Codex: initial usage listed at 8x the Standard rate, moving to 6x for credit-based usage under the documented conditions.
- ChatGPT Work: usage draws down plan credits under documented rates.
dev.to cautions that the API multiplier should not be treated as a single universal price across OpenAI's products, and that teams should verify the rules attached to their own accounts.
GPT-6.1 Sol itself is available on eligible paid plans in Work and Codex alongside GPT-6 Sol and GPT-6 Luna. Ultrafast is narrower at launch: OpenAI describes it as a feature for GPT-6 Astra and GPT-6.1 Sol in those products.
No latency numbers published
Notably, OpenAI has not published a response-time figure for the tier in the documentation dev.to cites. The intent is clear — maximum speed — but the absence of a stated latency improvement means teams will have to measure the difference inside their own applications rather than rely on a vendor benchmark before committing spend.
The source points to workflows where response time has a visible effect as the strongest candidates: interactive support drafting while a customer conversation is live, code assistance where shorter waits preserve a developer's flow, triage and routing of short incoming requests, and human-reviewed content operations that need quick first drafts or summaries. These are examples of suitable use cases, not claims that the tier improves output quality or removes the need for review.
Why it matters
A cost multiplier of 6x to 8x makes blanket adoption of Ultrafast hard to justify, and dev.to argues the sensible pattern is selective routing. Identify the points in a process where a person is genuinely blocked waiting on a result, reserve the premium tier for those moments, and let background jobs, batch work and other latency-tolerant tasks continue on Standard.
For API users that means keeping the service-tier choice explicit in every request and measuring the business effect in context — whether a support queue moves faster, whether a code-review handoff becomes smoother. For Work and Codex users, administrators should confirm plan or workspace eligibility first and understand how Ultrafast draws down credits, since the consumption rate can differ from the headline API price.
More broadly, the release extends a model of tiered model-serving in which speed, not capability, is the variable you pay for. Ultrafast for GPT-6.1 Sol is a premium, eligibility-gated capability rather than a default upgrade, and the value case depends entirely on whether faster responses translate into a measurably better workflow.
- #openai
- #api
- #latency
- #llm
- #pricing