TL;DR: OpenAI released GPT-6.1 Sol on September 29 with unchanged standard input/output prices of $2/$10 per million tokens, while cached input falls from GPT-6 Sol's $0.20 to $0.10.[2][4] Tool calling requires Responses, and neither none nor minimal reasoning is supported, so the upgrade needs an integration test as well as a cost comparison.[6]
OpenAI's September 29 release puts a new question in front of teams that adopted Sol only a week earlier: how much should they change before the first evaluation has settled? The company positions GPT-6.1 Sol close to Astra for complex work. Its launch report claims a 4.8-percentage-point improvement over GPT-6 Sol on AutomationBench at medium effort.[1] That is a useful reason to test the model, not a forecast for your own queue.
The real story isn't another decimal in a model name. OpenAI is trying to move more consequential work into an existing price bracket. That makes the relevant purchasing unit a completed task that passes review. A cheap response that triggers another run, breaks an integration, or requires manual repair belongs on the same invoice as the successful response.
Our original Sol and Luna analysis covered the September 22 price ladder. This October 3 analysis addresses the new migration decision: what changes with 6.1, what stays constant, and how to prove that a nominal upgrade improves your operating economics.
Why This Matters Now
Cover: Generated editorial artwork about completed work and accounting. It does not depict a benchmark or OpenAI infrastructure.
The Price Ledger: A Cache Cut, Not a General Discount
The headline input and output rates have not fallen against GPT-6 Sol. OpenAI still lists $2 per million ordinary input tokens and $10 per million output tokens. The direct price change is cached input: half the previous Sol rate.[3][4] A workload that never reuses context gets no automatic saving from that cut.
GPT-6.1 Sol's Standard Token Ledger
Per million tokens, according to OpenAI.
Per million tokens, according to OpenAI.
Per million tokens, according to OpenAI.
Per million tokens, according to OpenAI.
The long-context boundary matters more than the model's large advertised window. Above 272,000 input tokens, OpenAI applies $4 input, $0.20 cached input, $5 cache-write and $15 output rates per million tokens to the whole request.[5] The first 272,000 tokens do not keep their cheaper rate. Before attaching another repository dump, measure which information the task actually needs.
The Migration Contract: Tools Move to Responses
GPT-6 Sol allowed Chat Completions function calling with reasoning_effort: "none".[4] That route does not carry forward. GPT-6.1 Sol needs Responses for tools, and its lowest supported effort is low. The supported sequence is low, medium, high, xhigh, max; medium remains the default.[3]
OpenAI's migration guide also tells reasoning users to remove unsupported sampling and log-probability parameters, including temperature and top_p.[6] A wrapper that automatically sends the old options needs attention before anyone debates model quality.
Test a complete tool cycle: request, tool selection, tool result and final answer. Include a deliberately failed tool call so the application proves it can recover or report failure. Keep the old route available during evaluation. Otherwise an integration regression can masquerade as weak intelligence, and an apparently cheaper run can simply be a run that stopped early.
The Cache Example: Fifty Percent Becomes Nine Cents
Here is an illustrative accounting exercise, not a measured model result. Assume ten requests, each containing the same 100,000-token prefix, a changing 10,000-token suffix and 2,000 billed output tokens including reasoning. Assume the prefix is written once and fully read nine times; the suffix is never written to cache. Every request stays below the long-context boundary. Exclude tools, retries and regional surcharges.
OpenAI charges cache writes at 1.25 times ordinary input, with later reads at 0.05 times ordinary input for GPT-6.1 Sol versus 0.1 times for GPT-6 Sol. A persistent session does not itself guarantee those hits.[7]
| Assumed billable work across ten requests | GPT-6 Sol | GPT-6.1 Sol |
|---|---|---|
| One 100,000-token prefix write | $0.25 | $0.25 |
| Nine 100,000-token prefix reads | $0.18 | $0.09 |
| Ten 10,000-token uncached suffixes | $0.20 | $0.20 |
| Ten 2,000-token billed outputs | $0.20 | $0.20 |
| Total under these assumptions | $0.83 | $0.74 |
LLM Rumors arithmetic using OpenAI's published Standard rates [5, 7]. Token counts and cache hits are assumed equal; this table measures neither capability nor actual consumption.
The cache price halves, but the total falls by $0.09 because writes, ordinary input and output still cost money. Now impose a hypothetical acceptance test. If the old model completes eight acceptable tasks, $0.83 divided by eight is $0.10375 per accepted task. If the new model completes only seven, $0.74 divided by seven is $0.10571, rounded to five decimals. It is more expensive per accepted task despite its lower bill. With nine accepted tasks, the same new-model bill becomes $0.08222 per accepted task.
Those acceptance counts are invented scenarios, not estimates of either model. They show why the denominator belongs in the purchase decision. In production, different reasoning consumption and retries also change the numerator.
The Evidence Test: Vendor Progress Needs Local Replication
OpenAI's launch comparisons justify attention, but their measurement conditions matter. The company says its research or API evaluations can differ from production ChatGPT because prompts, tools and effort differ. Its difficult factuality prompts were selected from user-flagged errors, rather than typical traffic.[1] A launch chart cannot tell a support team its future correction rate.
The system-card addendum supplies another easily missed warning: comparison results for older models can reflect versions newer than their original launch evaluations.[8] Freeze the exact configurations used in your trial. A spreadsheet mixing old scores and current prices can produce a precise answer to the wrong question.
Use a held-out queue of actual tasks and the same acceptance rules. OpenAI's evaluation guidance supports reproducible datasets and trace grading.[10] Our recommendation is to record accepted tasks, every billable token category, tool fees, retries, escalation costs, elapsed time and human repair minutes. Inspect failures manually before treating an automated grader as the final judge.
Acceptance must mean more than plausible prose. For a repository change, it may require tests and a reviewer-approved diff. For document work, it may require correct numbers and traceable citations. Count escalation to Astra against the original task, rather than hiding the failed Sol attempt in a different budget.
The Rollout Decision: Buy More Useful Work
Availability also needs a precise label. OpenAI's current support page places GPT-6.1 Sol in ChatGPT Work and Codex, with access dependent on plan, rollout and workspace settings. It is unavailable in ordinary Chat conversations.[9] API dollars should not be translated into a guaranteed number of subscription tasks.
Start with work that has a clear finish line, then expand only after the migration passes its compatibility and acceptance checks. Keep reporting the original model, effort, tool environment and pricing tier. That record lets the next update face the same test instead of restarting the argument from marketing copy.
The Discount Ends at the Acceptance Test
The uncomfortable truth is that teams can spend less and accomplish less at the same time. GPT-6.1 Sol deserves a serious trial because it changes capability claims and cache economics within Sol's existing input/output price bracket. The winning migration is the one that lowers the verified cost of finished work.
Sources & References
Key sources and references used in this article
| # | Source | Outlet | Date | Key Takeaway |
|---|---|---|---|---|
| 1 | OpenAI | 2026-09-29 | Vendor capability claims and evaluation limitations. | |
| 2 | OpenAI | 2026-09-29 | Release date, launch rates and tool endpoint. | |
| 3 | OpenAI | Accessed 2026-10-03 | Current supported efforts, tools and Standard token rates. | |
| 4 | OpenAI | Accessed 2026-10-03 | Previous Sol's pricing and Chat Completions behavior. | |
| 5 | OpenAI | Accessed 2026-10-03 | Short- and long-context pricing for both models. | |
| 6 | OpenAI | Accessed 2026-10-03 | Reasoning and parameter changes needed for migration. | |
| 7 | OpenAI | Accessed 2026-10-03 | Read/write billing and the limits of session reuse. | |
| 8 | OpenAI | 2026-09-29 | Comparison versions can differ from original launch evaluations. | |
| 9 | OpenAI | Accessed 2026-10-03 | Work/Codex availability and exclusion from regular Chat. | |
| 10 | OpenAI | Accessed 2026-10-03 | Datasets and trace grading for repeatable evaluations. |
Last updated: October 3, 2026




