Cover: generated editorial illustration of two drafting desks connected through an inspection gate. It depicts delegation and review, not measured model performance.
TL;DR: Claude Sonnet 5.5 launched on September 28 at $2 per million input tokens and $10 per million output tokens.[1] Its five effort levels create a practical choice: begin well-defined agent work at medium, measure completed outcomes, and escalate difficult judgment to Opus 5.5.[2] The workflows below are proposed operating recipes, not results from an LLM Rumors benchmark.
The real story isn't another model that can write code and presentations. It is whether a team can turn an instruction into a checked deliverable without supervising every intermediate step. Sonnet 5.5 gives that question a fresh commercial incentive: cheaper execution becomes valuable when the result actually survives review.
Anthropic positions Sonnet as the faster complement to Opus, while reserving the strongest endorsement for Opus on complex, open-ended judgment. Its reported generation-speed improvement of 30%+ over Sonnet 5 is a vendor claim, not a prediction of your application's latency. We have not independently tested these workflows.[1] For the broader commercial change, see our Opus 5.5 pricing and migration analysis.
Why This Matters Now
Effort Settings: Make the Starting Point Explicit
Anthropic's Claude apps and Claude Code default to Medium for Sonnet 5.5; the Claude Platform defaults to High.[1] Set effort deliberately when comparing environments. Otherwise, the same task can become a comparison of different operating choices disguised as a model comparison.
The Documented Operating Envelope
low, medium, high, xhigh and max in Anthropic's effort documentation.
Anthropic's documented capacity, not a guarantee of complete recall.
Anthropic's ordinary model output limit; a ceiling rather than a recommended spend.
Anthropic recommends medium for well-specified agent tasks, high for harder ones, and reserving xhigh or max for demonstrated quality gains.[2] The model page documents a 1M-token context window and 128K maximum output.[3] Large limits give a workflow room; they do not tell it when to stop.
Our proposed selection rule is simple: choose the lowest setting that passes your acceptance checks. Preserve the failed examples. They reveal whether the bottleneck is missing context, an unclear instruction, a broken tool, or insufficient reasoning. Buying more effort before diagnosing the failure turns an integration problem into a recurring bill.
Bug Repair: Give Sonnet a Reproduction and a Finish Line
Begin with one failing behavior, a reproduction command, and the relevant repository instructions. Ask for a bounded repair rather than a general improvement. Anthropic's prompting guide warns that low effort can skip meaningful verification; your completion contract should therefore require observable checks.[4]
A reusable task prompt:
Fix the checkout bug described in the attached issue.
Reproduce it using the supplied command and inspect the relevant code.
Keep the change within the affected behavior and existing conventions.
Run the check that exercises the failure, then the relevant project checks.
Report the cause, changed files, commands and outcomes, and any blocker.
Finish when the requested behavior works and the checks pass.
Start this proposed workflow at medium. Escalate to high when the repair requires tracing several interacting modules. Ask Opus for a focused diagnosis when the unresolved question concerns competing designs or a persistent failure whose cause is still unclear. A second model should receive the reproduction, attempted patch, and actual logs, not an unsupported claim that the first model was confused.
The business gain is reduced reviewer reconstruction. A repair with its evidence attached can enter review immediately. A confident completion message with no executed check sends the next person back to the beginning.
Source to Brief: Separate Evidence From Recommendation
For a product or operating brief, prepare a source packet before requesting prose. Include publication dates, direct URLs, the reporting period, and the decision the document must support. Give the workflow a search tool if it needs current information. Sonnet's knowledge cutoff is June 2026, so September facts require supplied material or retrieval.[3]
Use this proposed prompt:
Create a two-page operating brief from the supplied source packet.
Attach a source URL and reporting period to every material factual claim.
Separate observed results, management statements and your analysis.
Reconcile conflicting figures or mark the conflict unresolved.
Finish with three options, their assumptions, and the evidence missing.
Return the claim ledger alongside the brief.
Start at high for synthesis with conflicting evidence; use medium for formatting an already reviewed ledger. Ask a reviewer to follow each consequential statement back to its source. Send the unresolved decision to Opus with the same packet and an explicit question. Do not pay for a wholesale rewrite when the actual uncertainty is one assumption.
What's often overlooked is the information handoff. Sonnet 5.5 and Opus 5.5 do not share readable thinking blocks in either direction. Keep the complete API history as documented, but provide an explicit summary of facts, decisions and unresolved issues for the receiving model.[5] A handoff should transfer evidence that a person can inspect.
Spreadsheet Extraction: A Valid Shape Is Only the First Check
A third workflow turns invoices, tables, or operating data into a structured record. Anthropic's structured outputs use output_config.format; strict tool inputs use strict: true. Supported object schemas require additionalProperties: false.[6] Those contracts constrain format. They do not establish that an extracted amount belongs to the right period.
Define fields for source location, currency, period, value, and unresolved ambiguity. Preserve a missing value as missing. Recalculate totals in code and compare them with the source document. For a dense visual table, give the agent a way to inspect the relevant region instead of expecting a single screenshot to carry every detail.
A proposed extraction prompt:
Extract the attached table into the provided schema.
Keep currency and reporting period with each value.
Use null for absent data; never infer an unreported number.
Include the page and row that support each record.
Flag conflicting totals for review before writing the final workbook.
Start the reasoning step with adaptive thinking at high; evaluate whether a cheaper setting passes the same reconciliation checks. Treat max_tokens, refusal, and model_context_window_exceeded as distinct outcomes in the application, not interchangeable empty answers.[7]
API Setup: Use a Small Request Before Connecting the Workflow
The native Claude API model ID is claude-sonnet-5-5. Adaptive thinking is the default; the lowest thinking mode is between_tools. The old disabled mode is rejected, and between_tools accepts only low, medium, or high effort. Parse responses by block type and retain thinking blocks unchanged in tool loops.[8]
Install the Python SDK with pip install anthropic in your project environment and set ANTHROPIC_API_KEY to your own key as an environment variable. This request is adapted from the documented API shape. It has not been executed against a paid endpoint:
import anthropic
client = anthropic.Anthropic()
reply = client.messages.create(
model="claude-sonnet-5-5",
max_tokens=4096,
output_config={"effort": "medium"},
messages=[{"role": "user", "content": "List acceptance checks for this bug report."}],
)
print(reply.stop_reason)
for block in reply.content:
if block.type == "text":
print(block.text)
Keep repeated instructions stable when using prompt caching. Sonnet 5.5's minimum cacheable prefix is 512 tokens.[9] Measure cache hits rather than assuming that repeated text qualifies. For a more integrated Sonnet-to-Opus workflow, Anthropic's advisor tool is a beta option; Sonnet 5.5 with Opus 5.5 returns encrypted advice rather than client-readable advisor text.[10] An ordinary explicit handoff is easier to inspect while you establish a baseline.
Acceptance First: Buy Escalation With Evidence
Run a small evaluation on representative work before changing the default model. Record accepted deliverables, retries, manual corrections, total spend, and time until acceptance. Hold the task packet, tools, and reviewer rubric constant. Compare effort within Sonnet first, then test Opus on the failures that survive better instructions and working tools.
The Cost That Matters
The uncomfortable truth is that delegating work demands a clear definition of finished. Sonnet 5.5 makes execution attractive; Opus provides a route for harder judgment. The durable advantage belongs to the team that knows when the job is done.
Sources & References
Key sources and references used in this article
| # | Source | Outlet | Date | Key Takeaway |
|---|---|---|---|---|
| 1 | Anthropic | 2026-09-28 | Launch, price, product defaults and vendor-reported speed claim. | |
| 2 | Anthropic | Accessed 2026-09-29 | Five effort levels and task-specific starting settings. | |
| 3 | Anthropic | Accessed 2026-09-29 | Context, output limits, model ID and knowledge cutoff. | |
| 4 | Anthropic | Accessed 2026-09-29 | Scope, verification and tools for visual inputs. | |
| 5 | Anthropic | Accessed 2026-09-29 | Model switches can drop unreadable reasoning blocks. | |
| 6 | Anthropic | Accessed 2026-09-29 | JSON contracts and supported schema limitations. | |
| 7 | Anthropic | Accessed 2026-09-29 | Different completion and truncation states require different handling. | |
| 8 | Anthropic | Accessed 2026-09-29 | Thinking modes and response block handling. | |
| 9 | Anthropic | Accessed 2026-09-29 | Sonnet 5.5 minimum cacheable prefix is 512 tokens. | |
| 10 | Anthropic | Accessed 2026-09-29 | Beta advisor support and encrypted Sonnet 5.5 advisor results. |
Last updated: September 29, 2026




