Claude Sonnet 5.5: Specs, Price, and Modelflare Access
A source-backed account of claude-sonnet-5-5: published specs, official token prices, how to read the 28 September 2026 benchmark table, and the Modelflare groups and user prices returned on 30 September 2026.
Anthropic released Claude Sonnet 5.5 on 28 September 2026. The API model ID is claude-sonnet-5-5. It is the second model in the Claude 5.5 family. Anthropic places it beside Opus 5.5: well-scoped everyday tasks, bug fixes, and documents, slides, and spreadsheets go here; complex, open-ended work that needs sustained judgment stays on Opus 5.5. The list price matches Sonnet 5, at $2 per million input tokens, $10 per million output tokens, and $0.20 per million cache-read tokens.
This article keeps three layers apart. The specification and the breaking changes come from the model page, the what's-new page, and the migration guide. The benchmark table is copied from the launch page. Modelflare did not rerun those evaluations. Groups and user prices come from GET https://modelflare.dev/api/pricing read on 30 September 2026. Vendor statements such as "up to 30% less" and "30%+ faster" describe Anthropic's own harness.
What the sources support
On the published list price, Sonnet 5.5 is the same card as Sonnet 5. Input and output are half of Opus 5.5. On the coding and knowledge-work rows printed on the launch page, it is well ahead of Sonnet 5. Terminal-Bench 4.0 at 70.6% is above the 66.4% printed there for Opus 5.5. CursorBench 4.0 at 55.5% and GDPval-AA v2.1 at 1844 sit next to Opus 5.5 and below it. Anthropic also says that, in their own use and in external testing, Opus 5.5 remains stronger at complex, open-ended work that needs sustained judgment.
On 30 September 2026 the Modelflare pricing response already contained claude-sonnet-5-5. The unit prices match the official input, output, cache-read, and 5-minute cache-write rates. The public groups are claude-award, claude-stable, claude-premium, and claude-tmp-cheap. The main-site notice of 29 September 2026 named the groups open that day, claude-award and claude-stable. By this reading, the public list also included the latter two.
Published specification
| Item | Claude Sonnet 5.5 |
|---|---|
| Claude API ID | claude-sonnet-5-5 |
| Amazon Bedrock ID | anthropic.claude-sonnet-5-5 |
| Google Cloud, Microsoft Foundry, Claude Platform on AWS | claude-sonnet-5-5 |
| Release date | 28 September 2026 |
| Retirement commitment | Not sooner than 28 September 2027 |
| Context / synchronous max output | 1M / 128K tokens |
| Batch API max output | 300K tokens, beta header output-300k-2026-03-24 |
| Input to output | Text and images to text |
| Thinking | Adaptive, on by default |
| Default effort on the Claude API | high |
| Default effort in the Claude apps and Claude Code | medium |
| Effort values | low, medium, high, xhigh, max |
| Reliable knowledge cutoff | June 2026 |
| Minimum cacheable prompt | 512 tokens |
| Tokenizer | Same as Sonnet 5, so the same text has the same token count |
The model page labels latency as Fast. That is a comparison inside the current Claude lineup, not a measured tokens-per-second figure. The lowest setting that turns off up-front thinking is thinking: {"type": "between_tools"}, accepted only at low, medium, and high. xhigh and max go back to adaptive thinking. Setting temperature, top_p, or top_k away from the default returns 400.
The launch page says Haiku 5.5 will follow in the coming weeks. This article does not price an ID that is not available yet.
Official prices
Amounts are US dollars per million tokens. The Sonnet 5.5 column is the model page. The Opus 5.5 column is the comparison table on the same launch page. The what's-new page says Sonnet 5.5 has the same prices as Sonnet 5, including prompt caching and batch processing. The batch discount is 50% on input and output.
| Per 1M tokens | Sonnet 5.5 | Opus 5.5 on the launch page |
|---|---|---|
| Input | $2 | $4 |
| Output | $10 | $20 |
| Cache read | $0.20 | $0.20 |
| Cache write | 5-minute $2.50; 1-hour $4 | $5 |
| Batch input / output | 50% of list | Not printed in that table |
The launch page prints one Opus 5.5 cache-write figure, $5, without a 5-minute and 1-hour split. The Sonnet 5.5 model page splits the two. Against Opus 5.5 input and output, this card is half. Cache reads are $0.20 on both.
Anthropic also publishes two workload statements: the same work usually takes fewer tokens, and most tasks cost up to 30% less; output generation is 30%+ faster than Sonnet 5. The cost charts add per-task comparisons at about one tenth, one fifth, and one fifteenth. Those results belong to their harness. A real bill still depends on effort, cache hits, tool rounds, and whether the call used batch.
How to read the benchmarks
The table below is the one printed on the launch page. A dash is an empty cell on that page. The Terminal-Bench 4.0 and CursorBench 4.0 cost charts use GPT-5.6 Sol because OpenAI had not published GPT-6 Sol on those two. This article does not copy a GPT-5.6 Sol chart point into the GPT-6 Sol column. Modelflare did not rerun the tests.
| Benchmark | Sonnet 5.5 | Sonnet 5 | Opus 5.5 | GPT-6 Sol |
|---|---|---|---|---|
| Terminal-Bench 4.0 | 70.6% | 10.3% | 66.4% | — |
| FrontierCode v1.1, Main | 46.2% (max); 52.1% (xhigh) | 42.4% | 54.4% | 49.3% |
| CursorBench 4.0 | 55.5% | 34.1% | 57.8% | — |
| GDPval-AA v2.1 | 1844 | 1449 | 1846 | 1487 |
| AA-Briefcase v1.1 | 1811 | 1359 | 1822 | 1483 |
| Humanity's Last Exam, with tools | 64.5% | 54.9% | 67.7% | — |
| OSWorld 2.1, partial | 80.1% | 57.0% | 81.8% | — |
| Chartography, no tools | 61.6% | 15.6% | 64.4% | 53.6% |
The footnotes belong next to the scores. Opus 5.5 on Terminal-Bench 4.0 is at xhigh, which the launch page calls that model's highest score. On FrontierCode, Sonnet 5.5 at max scores below xhigh. The launch page explains that max more often runs a code review split across subagents, and that this produced a timeout or edits beyond the task, which the benchmark penalizes. Artificial Analysis ran GDPval-AA and AA-Briefcase on a pre-release deployment that had a bug which could degrade structured outputs. The launch page says the bug was later fixed and that any effect would understate Sonnet 5.5. The GPT-6 Sol cells for GDPval-AA, AA-Briefcase, and Chartography carry a further note: OpenAI had fixed a bug that hurt image understanding, and those outside scores may not yet have been updated.
The cost charts on the same page are a different reading. Medium effort is the default in the Claude apps. On several benchmarks, Sonnet 5.5 at Low or Medium beats Sonnet 5's best score at about a tenth of the cost per task. High effort is the default on the Claude Platform. The launch page says that on FrontierCode at High it scores about 10 points above Sonnet 5 at the same setting, at about one fifteenth of the cost per task, and that it matches GPT-6 Sol's best score at about one fifth of the cost. On CursorBench at Low it exceeds Sonnet 5's best score at less than a tenth of the cost. On AA-Briefcase at Medium it beats Sonnet 5's best score at about one ninth of the cost. The printed 55.5% and 1844 are the points in the table. The cost sentences are another harness. Selected customer quotes on the launch page can suggest what to replay. They are not part of this table.
The launch page also says this is the first Sonnet to beat Pokémon Red from screenshots alone. That is the vendor's product narrative, not a row in the table above.
Modelflare prices and groups
In the pricing response of 30 September 2026, claude-sonnet-5-5 has input_price 2 and completion_ratio 5, so output is $10. cache_ratio is 0.1, so a cache read is $0.20. create_cache_ratio is 1.25, so the catalog cache write is $2.50. The row has no billing_mode and no separate 1-hour write price. Do not multiply the official 1-hour $4, or the batch half-price, by a group ratio and treat the product as the posted charge. The usage record is the charge.
The endpoint types are anthropic and openai. The group is chosen on the API key. The four public groups call the same model ID.
Price table
| Group | Ratio | Input / 1M | Cache read / 1M | Catalog cache write / 1M | Output / 1M | Catalog description that day |
|---|---|---|---|---|---|---|
claude-award |
0.25 | $0.50 | $0.05 | $0.625 | $2.50 | Price-priority routing for budget-sensitive, retryable development, testing, and flexible workloads |
claude-stable |
0.315 | $0.63 | $0.063 | $0.7875 | $3.15 | Everyday development, longer projects, and production use where stability matters |
claude-premium |
0.5 | $1.00 | $0.10 | $1.25 | $5.00 | Continuity and request success, billed at 50% of official list |
claude-tmp-cheap |
0.06 | $0.12 | $0.012 | $0.15 | $0.60 | Limited-time promotional route |
The arithmetic is the catalog unit price times the group ratio. On claude-stable, $0.63 is $2 × 0.315, and the cache write of $0.7875 is $2.50 × 0.315. The description is the text in the group definition. It is not a measured availability agreement.
The same response also lists non-public claude-first-topup at ratio 0.03, described as unlocked when a single top-up reaches $50. It is not a public selectable group, so it has no row above. When a key also has fallback groups, a request that leaves the campaign group is billed on the group that serves it.
Three examples
Each shape is 100,000 tokens of the named input class plus 5,000 output tokens. The cache rows assume the usage record actually classifies those tokens as a read or a write. At small amounts, quota rounding can move the posted figure.
| Request | List price | claude-award |
claude-stable |
claude-premium |
claude-tmp-cheap |
|---|---|---|---|---|---|
| 100,000 uncached input, 5,000 output | $0.2500 | $0.0625 | $0.07875 | $0.1250 | $0.0150 |
| 100,000 cache read, 5,000 output | $0.0700 | $0.0175 | $0.02205 | $0.0350 | $0.0042 |
| 100,000 catalog cache write, 5,000 output | $0.3000 | $0.0750 | $0.0945 | $0.1500 | $0.0180 |
The list price on the first row is 0.1 × $2 + 0.005 × $10. Each group cell is that $0.25 times the ratio. The third row uses the catalog write price of $2.50, the official 5-minute tier, rather than the 1-hour $4. The live page is the claude-sonnet-5-5 pricing page. The full catalog is Models and pricing.
When one failure is expensive, the catalog describes claude-premium as the continuity route. For everyday production and longer projects, the catalog describes claude-stable as that kind of default. Retryable development and testing use claude-award. claude-tmp-cheap is the limited-time promotional route at 0.06. Choose the group by the cost of failure, and read the ratio together with that description.
Calling it on Modelflare
Create a key on one of the four groups in the table. Send the model ID claude-sonnet-5-5 exactly. The canonical base URL is https://modelflare.dev/v1. The same API is also published at https://cf.modelflare.dev/v1 and https://origin.modelflare.dev/v1. The root domain is the canonical site.
The Messages API is the entrance that sets effort. The request below leaves thinking adaptive and sets effort to medium. max_tokens has to leave room for thinking blocks.
export MODELFLARE_API_KEY="YOUR_MODELFLARE_API_KEY"
curl -sS https://modelflare.dev/v1/messages \
-H "x-api-key: $MODELFLARE_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "Content-Type: application/json" \
-d '{
"model": "claude-sonnet-5-5",
"max_tokens": 1024,
"thinking": {"type": "adaptive"},
"output_config": {"effort": "medium"},
"messages": [
{
"role": "user",
"content": "Review this migration plan and list the three assumptions that most need a test."
}
]
}'
Chat Completions is also listed on this ID. This example sends text only and does not set reasoning_effort.
curl -sS https://modelflare.dev/v1/chat/completions \
-H "Authorization: Bearer $MODELFLARE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "claude-sonnet-5-5",
"messages": [
{
"role": "user",
"content": "Review this migration plan and list the three assumptions that most need a test."
}
]
}'
To choose an effort, use the Messages request above. The gateway currently turns Chat Completions reasoning_effort into a manual thinking budget with budget_tokens, and Sonnet 5.5 rejects a manual budget. When effort is omitted, the Claude API default is high. The difference between the two entrances is covered in Responses API versus Chat Completions. How usage meets the bill is covered in AI API cost tracking.
Moving from Sonnet 5
The what's-new page lists five changes that make code already running on Sonnet 5 return 400, plus one change that leaves the request successful and changes the response shape.
Requests that return 400
thinking: {"type": "disabled"}
thinking: {"type": "enabled", "budget_tokens": N}
thinking: {"type": "between_tools"} at effort xhigh or max
tool_choice: {"type": "any"}
tool_choice: {"type": "tool", "name": "..."}
temperature, top_p, or top_k set away from the default
tool type computer_20251124 on the Claude API and Google Cloud
advisor tool using Opus 4.8, Opus 4.7, or Sonnet 5 as the advisor
To turn off up-front thinking, send between_tools and keep effort at high or below. Keep tool_choice at auto or none. For schema-valid tool input, turn on strict, or move the schema to structured outputs, and say in the prompt when the tool applies. On the Claude API and Google Cloud, computer use moves to computer_toolset_20260801. Amazon Bedrock still accepts computer_20251124. The advisor tool accepts Mythos 5.1, Fable 5.1, Mythos 5, Fable 5, Opus 5.5, Opus 5, or Sonnet 5.5 itself. Advisor content comes back encrypted, and the client cannot read the text.
Accounts created on or after 31 August 2026, 00:00 UTC, have one more 400 on the Claude API, Amazon Bedrock, and Google Cloud: replaying a Sonnet 5.5 thinking block after a change to the system prompt, the tool list, or an earlier message fails. Keep the conversation append-only. When instructions must change, use a mid-conversation system message.
Behavior that changes on a successful response
When effort is omitted, the API runs at high. The same word does not buy the same amount of thinking it did on Sonnet 5. The what's-new page says to sweep effort again rather than copy the old setting. A well-scoped tool loop can start at medium, and a longer or harder one can move to high. Latency-sensitive chat can start at medium or low.
Progress text longer than a sentence or two between tool calls comes back inside thinking blocks. At the default display of omitted, the text of those blocks is empty, so a UI goes quiet between tools while the request still succeeds. Set display when the product needs those lines. With between_tools, that text comes back.
A thinking block records the model that produced it. Sonnet 5.5 reads thinking blocks from Sonnet 5, Opus 4.8, Haiku 4.5, and earlier models. It does not read blocks from Opus 5, Opus 5.5, Fable, or Mythos. No other model reads Sonnet 5.5 thinking blocks. A block the target model cannot read is dropped before the model sees it, and a dropped block is not billed. A block Sonnet 5.5 produces also stays with the account that produced it. Replaying it from an unlinked account drops the block, and the request still succeeds.
Refusal and fallback
A declined request returns HTTP 200 with stop_reason set to "refusal". The categories named on the what's-new page are cyber, bio, frontier_llm, reasoning_extraction, and general_harms. The launch page says its cybersecurity capability is close to Opus 5, so this is the first Sonnet to ship with that class of cyber safeguards. Higher-risk cybersecurity tasks fall back to Sonnet 5. Biology safeguards match Sonnet 5. Routine software development and most life-sciences work sit outside those two narrow safeguards.
The what's-new page says the Claude API beta default fallback retries cyber and frontier_llm on Sonnet 5, and does not retry bio, reasoning_extraction, or general_harms. That is Anthropic's documented behavior. Whether this Modelflare route performs the same fallback is answered by the model name on the response. Count refusals apart from transport failures. Whether a refusal is billed depends on its category. The launch page also says Sonnet 5.5 is available with zero data retention.
Checks before shifting traffic
- Pin
claude-sonnet-5-5. Also check the model name on the completed result. A safeguard fallback can answer as Sonnet 5. - Confirm the key group is one of the four public groups in the table, and choose it by the cost of failure.
- Replay one non-sensitive set at
low,medium,high,xhigh, andmax. Record success, latency, input, cache reads, cache writes, output, and charge. - Put effort in
output_config.efforton the Messages request. Do not use Chat Completionsreasoning_effortto set the thinking level for this model. - Change clients that still send
disabled, forced tool choice, a manual thinking budget, orcomputer_20251124before raising volume. - Keep conversations that replay thinking blocks append-only.
- Count HTTP 200 refusals apart from transport failures.
- Price one short prompt and one prompt that should hit cache, then read the posted amount. The official 1-hour write price and the batch half-price follow the usage record.
- Keep complex, open-ended work that needs sustained judgment on
claude-opus-5-5.claude-sonnet-5stays available.
Sources
Checked on 30 September 2026. Modelflare did not independently reproduce the benchmarks.
- Launch page, 28 September 2026: https://www.anthropic.com/claude-sonnet-5-5
- Model page: https://platform.claude.com/docs/en/models/sonnet-5-5/overview
- What's new: https://platform.claude.com/docs/en/models/sonnet-5-5/whats-new-sonnet-5-5
- Migration guide: https://platform.claude.com/docs/en/models/sonnet-5-5/migration-guide
- Modelflare pricing API: https://modelflare.dev/api/pricing
- Pricing page: https://modelflare.dev/pricing/claude-sonnet-5-5