Sonnet 5.5 effort: when to use medium, high or max

Choose Sonnet 5.5 effort using acceptance tests, latency and task cost. Learn why API and Claude Code defaults differ and when higher effort is worth testing.

Art-line illustration of compass with the title Sonnet 5.5 Effort.

Use Sonnet 5.5 effort as a parameter to evaluate, not a quality guarantee. Anthropic recommends starting at medium for well-specified agentic coding and multistep tool tasks, moving to high for harder or longer work. For latency-sensitive chat it suggests medium or low. The native API’s default remains high; Claude Code has its own default.

Those recommendations come from the Sonnet 5.5 behavior documentation, checked September 29, 2026. This article explains how to test the trade-off. It does not claim an Ofox benchmark across all effort levels.

Start with the work, then choose a setting

WorkloadDocumented starting directionWhat to inspect
Short, latency-sensitive chatlow or mediumResponse time and missing required details
Well-specified agentic codingmediumTests, scope and tool turns
Harder or longer tool taskshighAccepted result and repeated failure modes
A difficult task still failingTest higher effort against a baselineWhether extra cost changes the outcome

The last row is an evaluation suggestion, not an official promise that xhigh or max will fix the failure. A task with missing requirements, unavailable tools or contradictory instructions may remain unsolved at any effort.

The model page specifies high as the API default. Claude Code’s configuration documentation describes medium for Sonnet 5.5 in that client. Record the entry point before comparing results. Otherwise two runs described as “default Sonnet” may use different settings.

Why max is not a universal recommendation

Artificial Analysis’s launch evaluation found high output-token consumption at max effort and an unfavorable cost trade-off against some alternatives. The same report shows substantial benchmark capability. Both can be true: a model can reach a strong result by spending more tokens.

That report tested a pre-release deployment with a structured-output bug and states that relevant evaluations will be rerun. Its benchmark cost is not a price quote for your bug fix, nor does it establish that every max run is wasteful. Use it as a reason to collect cost and quality together.

Higher effort can also change the kind of behavior you see. More exploration is useful when it tests relevant explanations; it is harmful when the model expands beyond the requested change or spends time on irrelevant work. Your acceptance criteria should include scope, not just whether one test passes.

Design a small effort sweep

Prepare a set of representative tasks with expected outputs or acceptance checks. Keep the model version, tools, inputs and starting repository state fixed. Run the baseline setting, then test a higher or lower level on the same tasks. Use independent sessions when you want to avoid the first run teaching the next run the answer.

Record at least:

Task | Effort | Applied setting | Accepted | Attempts
Elapsed seconds | Input tokens | Cache categories | Output tokens
Tool charges | Total cost | Out-of-scope edits | Review notes

Separate first-attempt results from results after retries. If you cap time or tokens, mark a capped run explicitly rather than interpreting it as an ordinary completed answer. Preserve failures in the dataset. A table containing only successful examples cannot establish the most reliable configuration.

One useful decision rule is to retain a higher setting only when it improves an outcome that matters enough to justify its added cost or delay. Define that threshold before examining the results. For example, a team might value fewer incorrect patches more than slightly lower latency; a chat product might have the opposite constraint. Do not borrow a universal threshold from someone else’s benchmark.

Set effort without creating an invalid request

For the native API, an adaptive-thinking request can include:

{
  "model": "claude-sonnet-5-5",
  "max_tokens": 2048,
  "thinking": {"type": "adaptive"},
  "output_config": {"effort": "high"},
  "messages": [{"role": "user", "content": "List the acceptance checks for a CSV parser fix."}]
}

This is a documentation-based request body, not a live API test. max_tokens caps thinking plus response text; thinking tokens are billed as output even when their text is omitted. It is not a requested amount of thinking or an all-inclusive dollar budget. Authentication, version headers and your response handling are separate requirements.

If you use between_tools to disable up-front thinking, keep effort at high or below. That mode does not support xhigh or max. Changing its effort mid-conversation also has constraints. Use the migration checklist before copying older disabled or manual-budget configurations.

In Claude Code, launch with --effort medium or select a supported level with /effort. Managed settings can cap what actually runs. Requested and applied effort must not be treated as equivalent without checking the client and account behavior.

Separate effort, output limit and tool permissions

Three controls answer different questions. Effort influences how much reasoning the model applies. max_tokens limits the response’s token budget, including thinking. Tool permissions determine which actions the surrounding application can perform. Raising effort cannot create an unavailable file tool or grant access to a private repository; raising the output limit cannot fix a contradictory task specification.

For a practical example, consider “repair a CSV total and run the tests.” If the model has the source and an executable test, effort is a meaningful variable to compare. If it has only a screenshot of the error and cannot inspect the source, the first intervention is to supply the missing input. If the correct patch is present but the test command is denied, the next intervention is an approved verification route. Treating all three situations as “needs max” hides the actual cause.

Use this decision table before changing a setting:

Observed resultFirst interventionWhen an effort comparison is useful
Required input is absentAdd the missing source or clarify the requirementAfter both trials receive the same complete input
Tool or account access is deniedResolve approved access or record the blockAfter the task can actually execute
Response stops at its token capInspect truncation and choose an appropriate limitWith the same adequate cap in both trials
Patch is plausible but misses an edge caseAdd the edge case to acceptance criteriaCompare fresh runs on the same corrected task
Multiple valid approaches require careful trade-offsState the decision criteriaCompare quality and cost at medium and high

This is an editorial diagnostic framework. It is not a measurement that any particular effort level will pass a task. Keep the original failure in your records when you improve the task specification; otherwise the new prompt’s benefit can be falsely credited to the new effort setting.

A worked accepted-task cost example

Imagine two configurations tested on the same ten small tasks. The following figures are synthetic teaching data, not Sonnet measurements. Medium spends $0.40 across all attempts and finishes eight tasks acceptably. High spends $0.60 and finishes nine. The cost of an accepted result is $0.40 / 8 = $0.05 for medium and $0.60 / 9 ≈ $0.0667 for high.

Synthetic batchTasks attemptedAccepted tasksTotal spend, including failuresSpend per accepted task
medium108$0.40$0.0500
high109$0.60$0.0667

High costs one third more per accepted task in this example, but it delivers one additional accepted result. Whether that is worthwhile depends on the value of completion and the handling of the failed cases. It is not enough to say that medium is “25% less accurate” or that high is “always better.” The sample is small, the tasks are synthetic and neither number establishes production reliability.

Now consider a medium-first policy: run medium on all ten tasks, then try high only on its two failures. If those two escalations together cost $0.12 and both succeed, the policy’s total is $0.52 for ten accepted tasks, or $0.052 each. That arithmetic illustrates why escalation can be useful; it does not predict that high will fix every medium failure. If neither escalation succeeds, spend rises to $0.52 while the accepted count stays eight, making the result $0.065 per accepted task.

Include every attempt in spend, including abandoned runs and failures. Keep tool charges and human review time separate if they use different accounting units. If no task is accepted, report the ratio as undefined, not $0.00. For subscriptions without trustworthy per-task billing, report the observed usage and latency separately rather than inventing a dollar figure from an API price table.

Run a comparison that another person can inspect

Choose tasks before examining model outputs. For a coding set, include the task description, starting commit, permitted files, test command and an explicit success rule. For document work, keep the same source packet and claim-to-source requirements. Avoid making one setting solve the original problem while another sees the first setting’s answer, unless your intended product is specifically a repair pipeline.

Run each setting in a fresh session from the same state. Alternate the order across tasks when practical, because transient service conditions or local cache warmth can favor whichever configuration always runs second. Record cache state instead of assuming it stayed constant. If you repeat a task, count the repeats as trials of the same task, not additional independent tasks.

A useful result row includes the task ID, exact model and endpoint, requested and applied effort when observable, completion reason, elapsed time, usage categories, retries, test result and scope violations. Preserve the response or diff needed to audit the acceptance decision. If the API does not expose an applied setting, record “not independently observable”; do not fill the field by copying the requested value and label it verified.

Compare paired task outcomes, not just averages. If medium and high pass the same tasks, paying more for high has not demonstrated a completion benefit on that set. If high fixes an important failure but slows every easy task, route that failure class to high rather than changing all requests automatically. If the difference is one uncertain subjective score, enlarge or clarify the evaluation before claiming a winner.

Define escalation and stopping before the run

A practical policy might start bounded coding work at medium, retry a verified reasoning failure once at high with the same complete task, and stop for review when the second attempt remains invalid. That is an example policy, not an Anthropic default. Decide which errors qualify before running it: missing access, unsupported parameters and absent input should not consume an effort-escalation retry.

Reserve xhigh or max trials for cases where additional reasoning could plausibly address the observed failure and where your latency and spend limits allow the experiment. Use adaptive thinking for those levels. With between_tools, high is the highest documented level, and changing effort within the same conversation is constrained. Start a fresh controlled test when comparing modes rather than constructing an invalid mixed conversation.

Set a task-level stop condition that includes repeated attempts. A per-response token limit does not cap an agent that can make many calls. Stop when the allotted attempts, elapsed time or application budget is reached, retain partial artifacts and label the outcome capped. Never mark a capped run as a normal successful completion solely because its final text sounds conclusive.

The resulting recommendation should name its scope: for example, “medium is our baseline for these validated CSV fixes; high is evaluated for the listed failures.” That statement is more useful than a universal best-effort label, and it remains auditable when a model version or task mix changes.

Decide whether to change model instead

When a task remains difficult, compare increasing Sonnet effort with trying another model. Sonnet versus Opus addresses the within-Claude choice; Sonnet versus Sol covers a cross-provider trial. Keep the task and acceptance test fixed when changing the configuration.

Once you choose a baseline, record why you chose it and which failures justify escalation. This makes future model updates easier to evaluate. Sonnet 5.5’s effort levels were recalibrated relative to Sonnet 5, so reusing the old label without testing is not evidence of equivalent behavior.

Frequently Asked Questions

Is high the default everywhere?
No. The native API and Claude Code have different documented defaults. Account controls and explicit settings can change what is applied.
Can I use max with between_tools?
No. The documented mode supports low, medium and high. Use adaptive thinking for higher effort levels.
Will lower effort always reduce the cost of a completed task?
Not necessarily. It may reduce tokens per attempt but require more attempts or fail more often. Measure accepted-task cost rather than only one response.