Sonnet 5.5 vs Opus 5.5: which coding tasks need Opus?
Choose Sonnet 5.5 or Opus 5.5 for bug fixes, code review and complex changes. Compare task scope, effort, cache costs and the evidence behind the decision.
Start by evaluating Sonnet 5.5 on well-defined coding tasks and Opus 5.5 on work where judgment, ambiguity or repeated failure justifies the extra trial. That is a workload-selection approach, not proof that Sonnet always handles small tasks or that Opus always wins on large ones. Your repository tests and review standards remain the acceptance criteria.
Anthropic positions Sonnet as a faster, lower-cost complement to Opus, while describing Opus as suited to complex work requiring careful judgment. The two models’ billing and effort settings make the choice less simple than “Sonnet is half the price.” This guide uses documentation checked September 29, 2026; no original head-to-head model trial is claimed.
Rates do not settle the choice
| Official Claude API category | Sonnet 5.5 | Opus 5.5 |
|---|---|---|
| Input, per million tokens | $2 | $4 |
| Output, per million tokens | $10 | $20 |
| Cache reads, per million tokens | $0.20 | $0.20 |
| API default effort | high | medium |
| Context window | 1M tokens | 1M tokens |
Sources: Sonnet specification and Opus specification. These are vendor rates, not subscription quotas or an Ofox quote. The cache-read row is especially relevant to repeated repository context: a large cached prefix does not have the same two-to-one difference as uncached input and output.
Imagine two synthetic requests with the same 100,000 cache-read tokens and 2,000 billable output tokens, with all other billable categories excluded. Sonnet would cost $0.04 and Opus $0.06. The difference is not twofold because cache reading costs the same in this example. If one model takes additional turns, the comparison changes again.
Match the task to a testable outcome
| Task | Useful starting comparison | Evidence to retain |
|---|---|---|
| Bug with a deterministic reproduction | Sonnet at a modest effort setting, then Opus if needed | Failing test, patch, full regression run |
| Small feature with explicit requirements | Sonnet versus your current working baseline | Acceptance checklist, scope of edits |
| Ambiguous cross-module failure | Include Opus from the outset | Competing explanations, inspected files, verified cause |
| Repository review | Give both the same scope | Confirmed findings versus false positives |
| Risky migration | Compare planning and validation separately | Migration checklist, rollback path, integration tests |
These are trial designs. They are not measured pass-rate claims. A short patch can require difficult reasoning, and a long but mechanical change can be easy. File count or lines changed alone is a poor measure of task difficulty.
Read benchmark claims with their conditions
Artificial Analysis’s Sonnet launch report shows strong results on several tasks but much higher output consumption at max effort. It also notes a pre-release structured-output issue affecting the tested deployment and planned reruns. Those results can justify testing both models; they cannot establish your cheapest configuration.
Do not compare an Opus medium run to a Sonnet max run and call the result a pure model comparison. It is a configuration comparison. That can still be valuable, but publish both settings and the actual budgets. Equal effort names also do not guarantee equal computation.
Anthropic’s own release announcement describes different model strengths and evaluation conditions. Keep vendor claims separate from independent measurements and from your own observations. A disagreement between charts may reflect tasks or settings rather than a mysterious contradiction.
A simple escalation policy
Before a run, write a stopping condition: for example, one proposed fix plus a regression check. If the check fails, inspect the failure before rerunning. If the model misunderstood the task, clarify the input rather than blindly increasing effort. If it found the right area but cannot produce a valid fix, trying Opus becomes a useful controlled escalation.
Carry the issue description, relevant files and test results into the new run. Do not assume signed thinking blocks are portable across models. Sonnet 5.5 has model- and conversation-specific rules described in the migration guide. Retain visible evidence rather than relying on hidden reasoning continuity.
Record the cost of the initial attempt and the escalation together. Otherwise a routing system can appear cheaper by attributing the first failure to one model and counting only the final successful run for another. Also record review time: a patch that compiles but needs substantial cleanup has not finished the job.
Calculate three different price relationships
The same two models can have very different cost ratios depending on the workload. The following are synthetic vendor-list calculations with equal token counts, standard service and no other charges. They isolate billing differences before any claim about model behavior.
| Token pattern per request | Sonnet 5.5 | Opus 5.5 | Interpretation |
|---|---|---|---|
| 20K uncached input + 2K output | $0.06 | $0.12 | Uncached input and output are both priced at 2× on Opus |
| 100K cache read + 2K output | $0.04 | $0.06 | Equal cache-read rates narrow the ratio to 1.5× |
| 100K five-minute cache write + 2K output | $0.27 | $0.54 | Creation costs matter on the first request |
Opus’s documented five-minute and one-hour cache creation rates are $5 and $8 per million tokens; Sonnet’s are $2.50 and $4. At one 100K five-minute write followed by nine reads, plus 2K output on every call, the ten-call subtotal is $0.63 for Sonnet and $1.08 for Opus. Sonnet: $0.25 creation + $0.18 reads + $0.20 output. Opus: $0.50 + $0.18 + $0.40. This is about 1.71×, not a fixed 2× multiplier.
Do not assume that changing models preserves a previous cache. The example gives each model its own initial write. Actual token counts and reuse can differ. A model that needs twice as many attempts may erase a lower rate, while a more expensive model can still be the wrong choice if it creates the same failed patch with a longer explanation.
What “difficult enough for Opus” means in practice
Difficulty is often uncertainty about the correct behavior rather than the size of the edit. A change to one authorization condition can require more care than replacing an API name across 30 files. Anthropic’s positioning makes Opus a reasonable candidate for long-running, judgment-heavy work; it does not provide a guarantee that any particular one-line bug needs Opus.
Use the following four questions to decide whether to include Opus early. Is the root cause still unknown? Are there conflicting requirements? Can the result create costly side effects? Does acceptance require evaluating alternatives rather than checking a deterministic answer? Several “yes” answers justify spending time on a second configuration before committing to a patch. They do not justify granting broader tool permissions.
For a bounded formatting bug with a failing snapshot test, start with Sonnet and require the specific test plus nearby regression tests. For a payment retry bug that may duplicate a charge, the acceptance plan must cover idempotency, partial failure and recovery. Include Opus in the investigation if useful, but use fixtures and human review for the sensitive change. A more capable model is not a substitute for isolating external side effects.
For an architecture task, ask both candidates to identify constraints, propose alternatives and describe a reversible migration. Judge whether each proposal fits the actual code and deployment limits. A confident diagram is not sufficient evidence. Require references to relevant modules and an implementation sequence whose intermediate states remain valid. If those inputs are absent, supply them before treating any disagreement as a model weakness.
Set an escalation budget before the first attempt
A simple policy is one bounded Sonnet attempt, one diagnosis of any failure, then a deliberate choice between clarifying the task and trying Opus. Do not send the same unchanged prompt through an unlimited retry loop. If the failure is a missing dependency or a broken fixture, changing models will not repair the environment. Fix the test setup first and retain that incident separately from a reasoning failure.
Consider a hypothetical population in which each Sonnet attempt costs $0.06 and each Opus escalation costs $0.12, with at most one escalation. If a fraction e of tasks escalates, the average token cost per submitted task is $0.06 + $0.12e. At e = 25%, that is $0.09; at e = 50%, it is $0.12, equal to one direct Opus attempt under these assumptions. Above 50%, this routing policy spends more than that direct baseline.
This break-even calculation is not a pass-rate prediction. It excludes differences in success, context preparation, cache state, review and delay. If Opus also fails or takes additional turns, include those costs too. Report both accepted-task cost and the fraction left unresolved. For urgent work, a cheaper average may still be a bad policy if escalation repeatedly adds delay to the tasks that matter most.
Give the second model an evidence packet
When escalating, create a short visible handoff: intended behavior; current commit; reproduction command; actual failure; attempted patch; and the reason it was rejected. Include only relevant files or references the next run can access. Ask it to explain the root cause before revising the patch when the cause remains uncertain. This prevents the second attempt from repeating an already disproved assumption.
Do not turn that packet into “the previous model said X, so X is true.” Preserve actual test output separately from the model’s interpretation. If the first attempt changed files, either reset to the same baseline or explicitly record the changed starting state. Without that distinction, an apparent Opus success may actually reflect useful work Sonnet already performed, and an apparent failure may inherit a damaged workspace.
For review tasks, independently verify each reported issue. Record the file and line, triggering condition, observed or reproducible consequence, and whether the issue existed before the patch. Deduplicate overlapping findings before counting them. A second model can help challenge the first explanation, but two models agreeing is still not the same as a reproduced defect.
A practical routing choice
Keep Sonnet as a candidate for frequent, well-specified work where acceptance is cheap to check. Include Opus for the uncertain or costly cases and retain a direct-Opus baseline so you can tell when escalation itself is wasteful. If your sample shows that a task category consistently needs escalation, route that category directly only after confirming the pattern is not caused by a poor prompt or broken integration.
Review the policy per task category. Repository search, implementation and final review may have different needs; splitting them can also add handoff overhead. Measure the complete workflow before claiming a saving. For a subscription, use the same quality and timing framework but do not substitute the API arithmetic for account-specific allowances. The final choice should explain which tasks go where, why, and which evidence would change that decision.
What to do on a subscription
Do not translate the API table into an exact number of Claude Code prompts. Subscription allowances and usage limits depend on the account and service conditions. Check the model actually selected, especially after a client update or provider change, and keep billing mode separate from the model name.
The Claude Code setup guide explains version checks and explicit selection. The effort guide separates API defaults from Claude Code defaults. If you are considering another provider, the Sonnet versus Sol guide gives a controlled comparison method.
Frequently Asked Questions
- Is Sonnet always half the cost of Opus?
- No. Its listed uncached input/output rates are half, but cache rates, token consumption, retries and tools affect completed-task cost. The worked example above deliberately isolates those differences.
- Should code review always use Opus?
- No universal rule follows from the documentation. Compare confirmed findings and false positives on representative reviews, including the time a human spends checking each finding.
- Can I move the same conversation between models?
- Visible messages can be part of a supported workflow, but thinking blocks have compatibility rules. Check the migration documentation rather than assuming the entire hidden state transfers.


