Sonnet 5.5 API pricing: calculate cache and task costs

Calculate Claude Sonnet 5.5 API costs for input, output, caching and Batch jobs, with worked examples that separate token prices from completed-task costs.

Art-line illustration of price tags with the title Sonnet 5.5 API.

Claude Sonnet 5.5 costs $2 per million input tokens and $10 per million output tokens on the standard Claude API. A cache read costs $0.20 per million tokens. Those are vendor list prices checked on September 29, 2026, not an Ofox quote or the price of a Claude subscription. The important budgeting question is how many billable tokens your application uses to finish an accepted task.

For developers moving from Sonnet 5, the rate card itself is unchanged. Anthropic’s claim of lower costs on much of its work concerns model behavior and token consumption; it is not a blanket discount on every request. Start with the official model specification, then estimate the workload you actually intend to run.

The Sonnet 5.5 rate card

Token categoryUSD per million tokensWhat to count
Uncached input$2.00Input not charged as cache creation or reading
Output$10.00Billable output, including billed thinking usage
Five-minute cache creation$2.50Tokens written to that cache tier
One-hour cache creation$4.00Tokens written to that cache tier
Cache reading$0.20Tokens actually served from the prompt cache

The model page also lists a 50% Batch API discount on input and output and a 512-token minimum cacheable prompt. Reaching that length does not automatically create a cache: configure caching and check the usage fields. Do not count cached tokens twice by adding them to ordinary input again.

Tool charges, data-residency options and other service-specific charges need separate checking against the full pricing documentation. A provider’s supported features, currency and billing rules can differ. Check its current terms before applying this vendor calculation to another endpoint.

A small request: $0.04, not a monthly budget

For a hypothetical request using 10,000 uncached input tokens and 2,000 billable output tokens:

Input:  10,000 / 1,000,000 × $2  = $0.02
Output:  2,000 / 1,000,000 × $10 = $0.02
Total:                               $0.04

At exactly that usage, 1,000 requests would cost $40 before other charges. This is arithmetic, not a measurement of Sonnet’s typical response length. A longer reasoning trace, another tool turn or a retry changes the total. Multiplying the visible answer’s word count by a guessed conversion factor is not a reliable billing method.

When caching changes the calculation

Suppose ten requests share a 100,000-token prefix. Assume one five-minute cache write, nine actual cache hits within the relevant lifetime, and no other input or output. The prefix costs $0.25 to write plus 9 × $0.02 to read: $0.43, versus $2.00 for ten fully uncached copies. That is a saving of $1.57, or 78.5%, for this prefix alone.

It is not a 78.5% discount on the whole application. Changing the prefix, missing the cache or generating substantial output can reduce the overall saving. Record cache creation and read usage for each request. A cache-shaped prompt is not evidence that the provider actually billed a cache hit.

For eligible non-interactive work, evaluate Batch separately. Applying its advertised 50% input/output discount to the uncached $0.04 example gives $0.02 under the stated assumptions. This example does not combine Batch with caching or assert how every combination is billed.

Measure accepted-task cost

Use this worksheet for a bug fix, document or extraction job:

Task ID | Model | Effort | Input | Cache write | Cache read
Output | Tool charges | Attempts | Accepted? | Total USD

Include failed attempts and retries in the cost of the task. Divide the sum by accepted tasks only after defining acceptance: tests pass, required fields are correct, or the deliverable meets a review checklist. Keep subscription usage in a separate column; a CLI’s estimated API-equivalent cost is not automatically a charge on your subscription invoice.

There is a reason to record effort. Artificial Analysis’s launch evaluation found very high output consumption at max effort. Its benchmark workload does not predict your bill. It also used a pre-release deployment affected by a structured-output issue, so its own rerun caveat matters. This is a warning to measure task usage, not a reason to call Sonnet universally expensive.

Build a bill from usage, not from the chat transcript

For the standard rates above, a complete token calculation is:

USD = (2 × uncached_input
     + 2.5 × cache_write_5m
     + 4 × cache_write_1h
     + 0.2 × cache_read
     + 10 × billable_output) / 1,000,000

These variables are mutually exclusive billing categories. In a Claude usage response, distinguish input_tokens, cache_creation_input_tokens and cache_read_input_tokens; where different cache lifetimes are used, retain their detailed creation breakdown. Do not multiply both the aggregate creation count and its five-minute/one-hour components. Use the selected provider’s documented usage semantics if a gateway changes the response shape. An absent field means the record is incomplete, not that its cost was zero.

Consider an invented request with 8,000 ordinary input tokens, a 40,000-token five-minute cache write, 60,000 cache-read tokens and 3,000 output tokens. Its cost is $0.016 + $0.100 + $0.012 + $0.030 = $0.158. The total input context can contain all three input categories, but they do not all receive the ordinary input rate. Save the raw usage response next to this calculation so a later invoice discrepancy can be traced to a category instead of guessed from prompt length.

Thinking adds another source of confusion. A short visible answer can have substantial billed thinking usage; hiding thinking text does not remove that usage. Likewise, max_tokens limits thinking and text together. It is an output ceiling, not a promise that the model will finish correctly within it. If a run ends at the ceiling, log the incomplete attempt before continuing. Discarding it would understate the task’s real spend.

Five-minute or one-hour cache: calculate the break-even point

Take the same 100,000-token prefix and ignore all other usage. With N calls, including the first write, uncached input costs $0.20N. One five-minute write followed by N−1 actual hits costs $0.25 + $0.02(N−1). That becomes cheaper at two calls: $0.27 versus $0.40. A one-hour write costs $0.40 + $0.02(N−1), so two calls cost $0.42, slightly more than uncached input; three calls cost $0.44 versus $0.60.

Those thresholds assume a single write and successful reads. They are not a guarantee based on wall-clock spacing alone. Use the actual usage records to count hits and writes. If ten requests each recreate the five-minute cache instead of reading it, the prefix costs $2.50, more than the $2.00 uncached baseline. A longer-lived cache is useful only when avoided rewrites justify its higher creation rate.

Synthetic prefix-only patternCostWhat it demonstrates
10 uncached copies$2.00Baseline for 100K input tokens per call
1 five-minute write + 9 hits$0.43Reuse with one creation
1 one-hour write + 9 hits$0.58Higher creation cost, same read rate
10 separate five-minute writes$2.50A cache configuration can cost more when it misses

Put genuinely stable material before changing task input, but never alter required context merely to improve a hit-rate metric. A cached irrelevant repository dump can still waste context, while a smaller uncached excerpt may be the better request. Also distinguish changing the model from reusing the same model: never budget cross-model cache reuse without evidence that the service supports and bills it that way.

Three workload budgets you can reproduce

The following scenarios use standard synchronous Claude API rates, no caching, no Batch discount and no separate tool charges. They are planning examples, not usage forecasts or reported model results.

Workload assumptionCalculation per taskTasks per monthToken subtotal
Short extraction: 4K input, 500 output$0.008 + $0.005 = $0.01310,000$130
Drafting: 12K input, 4K output$0.024 + $0.040 = $0.0641,000$64
Coding loop: four turns, each 30K input and 2K output4 × ($0.060 + $0.020) = $0.32500$160

The coding row counts input on every turn. Sending the previous conversation again can increase later input, so four identical turns are an explicit simplifying assumption. Replace each turn with its measured usage before treating the result as a forecast. Tool results can become input on a later turn; the model’s output containing a tool request and any separately billed tool execution are different costs.

If 100 of those 500 coding tasks require one additional $0.08 turn, add $8, giving $168. If only 480 tasks ultimately pass acceptance, the token subtotal per accepted task is $168 / 480 = $0.35. Retain the 20 failures in the report: that is a 96% acceptance rate, not 500 successful jobs. Human review, hosting and any external tool fees still sit outside the token subtotal.

Diagnose an unexpectedly high invoice

Work through the largest component first. If output dominates, inspect effort, actual thinking usage, truncation and unnecessary verbosity. If input dominates, inspect repeated history and whether retrieval sends more material than the task requires. If cache creation dominates, compare write and read events rather than assuming the cache is working. If retries dominate, classify invalid requests, tool failures and rejected results separately: a schema bug needs a client repair, while a correct but inadequate answer needs task or model investigation.

For a spending guardrail, enforce a task-level limit in the application as well as a per-request output limit. Stop and surface an incomplete task when its budget is exhausted. Never represent an application guardrail as an official rate limit or silently substitute a cheaper model without retaining the actual model in the task record. Before choosing Batch for savings, confirm the workflow tolerates asynchronous completion and that it is eligible for the documented service. A lower token rate does not make a delayed response appropriate for an interactive user.

Choose the next check

If the rate card is clear but the bill is not, compare actual output tokens, cache hits and retry counts before switching providers. For tuning, use the Sonnet 5.5 effort guide. For model selection, see Sonnet versus Opus for coding tasks.

For Ofox integration, open the Claude Sonnet 5.5 model page to check the model ID, current provider, pricing and supported protocols. The calculations above use Anthropic’s published rates; they are not a guarantee that every Anthropic feature is available through each provider.

Frequently Asked Questions

Did Sonnet 5.5 lower the token price from Sonnet 5?
No. The current official documentation says the prices are unchanged. A task can still use fewer or more tokens, which changes its total cost without changing the per-token rate.
Does a Claude Pro or Team subscription include API credits?
Do not treat a subscription allowance as API credit. Confirm the billing route used by your client and review the relevant plan separately from the Claude API rate card.
Is max effort the cheapest way to finish a task?
There is no general guarantee. Compare accepted results at a few effort settings, including retries and elapsed time. A higher setting can consume much more output without improving your acceptance rate.