by Anthropic · Claude 4 family · best for best value production workhorse
Claude Sonnet 4.6 is Anthropic's balanced workhorse, released February 17, 2026, and the best value in the lineup: SWE-bench Verified 79.6% sits within ~1.2 points of Opus 4.6 at 60% of the input cost, with a 1M-token context at standard pricing. For a buyer, the single sentence is this: route everything that is not frontier-hard here — it is the default Claude Code model and the cost anchor for any Anthropic-centric production stack.
| Benchmark | Score | Source |
|---|---|---|
| MMLU | 89.1% | benchlm.ai 2026-02-17T00:00:00.000Z |
| MMMU | 83.6% | morphllm.com 2026-02-17T00:00:00.000Z |
| MATH-500 | 89% | benchlm.ai 2026-02-17T00:00:00.000Z |
| MMLU-Pro | 87.3% | benchlm.ai 2026-02-17T00:00:00.000Z |
| AIME 2025 | 94% | benchlm.ai 2026-02-17T00:00:00.000Z |
| HumanEval | 98% | nxcode.io 2026-02-17T00:00:00.000Z |
| LMArena Elo | 1460 | openlm.ai 2026-05-28T00:00:00.000Z |
| GPQA Diamond | 74.1% | morphllm.com 2026-02-17T00:00:00.000Z |
| LiveCodeBench | 79.7% | rootly.com 2026-02-17T00:00:00.000Z |
| LMArena Coding Elo | 1500 | openlm.ai 2026-05-28T00:00:00.000Z |
| SWE-bench Verified | 79.6% | anthropic.com 2026-02-17T00:00:00.000Z |
| Artificial Analysis Index | 51 | artificialanalysis.ai 2026-02-17T00:00:00.000Z |
Six personas, six verdicts — the same panel that reviews every product on TopReviewed.
“Sonnet 4.6 is the model I standardize on: frontier-enough for code and agents at a price I can run on every request.”
For most workloads this is the right default. It clears the capability bar for coding and agent work while keeping per-call cost low enough to put on every user request, and multi-cloud availability (Bedrock global/regional, Vertex three-tier, first-party API) makes failover real. The 1M context at flat pricing simplifies capacity planning. The strategic risk is the gap to Opus 4.7 once the hardest decile of jobs arrives — plan to route those upstream. Lock-in is mitigated by the multi-region story and stable tokenizer.
“Opus-class coding at Sonnet price is the wedge that makes Anthropic the safe default for the whole production stack.”
Sonnet 4.6's positioning is the strongest in the lineup commercially: it neutralizes the "Opus is too expensive for production" objection and keeps teams from defecting to cheaper rivals. Leading GDPval-AA and Terminal-Bench (edging Opus 4.6) while priced at $3/$15 is a differentiation story that compounds Anthropic's ecosystem gravity in Claude Code, Cursor, and Copilot. Market timing is excellent: it landed as agentic coding became the dominant enterprise use case, capturing the high-volume middle of the market.
“At $3/$15 with deep cache and batch discounts, Sonnet 4.6 is the cost anchor of any Anthropic budget.”
This is the sweet spot of the lineup. Cached input ($0.30) and batch ($1.50/$7.50) cut bills 50-90% on stable workloads, and the 1M context at flat pricing removes the long-context premium entirely. Anthropic's own worked example (~$37 per 10k support conversations on Haiku) scales to a competitive ~3x for Sonnet, still well inside GPT-class peer ranges. The one budget watch-item is extended-thinking and max-effort spend — Sonnet 4.6 can use ~3x the output tokens of 4.5 if left at max effort by default. Set effort deliberately.
“This is the model I reach for when building — close to Opus on routine code, fast, with a clean thinking-budget knob.”
For hands-on builders, Sonnet 4.6 is the everyday driver. Coding feels close to Opus on routine work, latency is meaningfully better, and the explicit extended-thinking budget tunes cost without rewriting prompts. The full Anthropic tool surface works cleanly with the SDK; structured output and prompt caching behave as expected, and cache hits at 10% of input make repeated agent loops genuinely cheap. The OSWorld jump matters in practice: computer-use agents that flaked on Sonnet 4.5 are now usable. Where it falls short is the hardest debugging and refactor cases, which still want Opus 4.7.
“Fast, sharp, and rarely refuses — Sonnet 4.6 is the everyday Claude that just keeps up with you.”
For a heavy daily user, Sonnet 4.6 is a strong default chat surface. Latency is in the fast bucket, conversation quality is high, refusals are well calibrated, and it is more concise than Opus when asked. It handles pasted screenshots well enough to act on them. Some users notice a more direct, less effusive personality versus Sonnet 4.5. The only soft spot is the hardest scientific or math questions, where Opus 4.7 materially outperforms — but for the 95% case, Sonnet 4.6 is faster and feels better day to day.
“Great value, but 'within 1.2 points of Opus' leans on SWE-bench while GPQA Diamond sits 17 points lower.”
The value story is real and the coding parity on SWE-bench is genuine — but the "near-Opus" framing is selective. GPQA Diamond at 74.1% is roughly 17 points below Opus 4.6's 91.3%, so on hard science Sonnet 4.6 is not close. The headline GPQA number is also methodology-dependent (74.1% no-tools vs ~89.9% with extended thinking in some reports), which makes cross-model comparison fragile. The ~3x output-token inflation in max-effort mode quietly erodes the cost advantage if teams leave effort high. And it still trails Opus 4.7 meaningfully on SWE-bench Pro. It is the best value Claude; just do not mistake it for an Opus substitute on reasoning.
The full research notes behind this review — verified against primary sources.
Anthropic discloses no parameter count, layer count, or attention mechanism, so those fields are null/unknown. Disclosed: a 1M-token context window served at standard pricing; 64k synchronous max output (300k via batch beta); and support for both an explicit extended-thinking control (with token budgets) and adaptive thinking. Sonnet 4.6 retains the prior Claude tokenizer (Opus 4.7's new tokenizer is the exception), so token budgets and cost models from Sonnet 4.5 carry over cleanly.
Coding (9.0): SWE-bench Verified 79.6%, LiveCodeBench 79.7%, HumanEval ~98%, LMArena coding Elo 1500 — close to Opus on routine work at far lower cost. Reasoning (8.3): GPQA Diamond 74.1% (well below Opus tier), ARC-AGI-2 58.3%, AA Index 51 (second only to Opus 4.6 at the time of release). Math (8.5): AIME 2025 ~94% with tools, MATH ~89%. Agentic/tool use (8.8): OSWorld 72.5% (tied with Opus 4.6), full first-party tool suite, leads Terminal-Bench in AA testing. Long-context (9.0): 1M tokens at standard pricing, strong retrieval across the window. Multilingual (9.0): broad coverage, Arena multilingual 91.3. Vision (8.5) and document/OCR (8.3): solid for charts, screenshots, and documents though below Opus 4.7's high-res pipeline. Instruction-following (8.8): improved and more consistent than Sonnet 4.5. Function-calling (9.0): robust structured output and parallel calls. Safety calibration (9.0): ASL-3, balanced refusals. Realtime-data (7.0): no native post-August-2025 knowledge, but web search and web fetch close the gap.
| Benchmark | Score | vs Predecessor | vs Top Competitor | Source |
|---|---|---|---|---|
| SWE-bench Verified | 79.6% | +2.4 vs Sonnet 4.5 (77.2%) | within 1.2 pts of Opus 4.6 (80.8%) | Anthropic |
| GPQA Diamond | 74.1% | improved | behind Opus 4.6 (91.3%) | Morph |
| AIME 2025 | 94.0% | improved | frontier (with tools) | BenchLM |
| MMLU-Pro | 87.3% | improved | near-frontier | BenchLM |
| LiveCodeBench | 79.7% | improved | competitive frontier coder | Rootly |
| HumanEval | 98% | improved | near-saturated | NxCode |
| MMMU | 83.6% | improved | frontier vision tier | Morph |
| OSWorld-Verified | 72.5% | +11.1 vs Sonnet 4.5 (61.4%) | ~tied Opus 4.6 (72.7%) | Morph |
| ARC-AGI-2 | 58.3% | +44.7 vs Sonnet 4.5 (13.6%) | behind Opus 4.6 (68.8%) | Morph |
| LMArena Elo | 1460 | +40 vs Sonnet 4.5 (1420) | #6 tier | OpenLM |
| LMArena Coding Elo | 1500 | +36 vs Sonnet 4.5 (1464) | top-tier | OpenLM |
| Artificial Analysis Index | 51 | +8 vs Sonnet 4.5 (43) | behind Opus 4.6 (53) | AA |
(GPQA Diamond varies by methodology: the 74.1% no-tools figure is used here; some aggregators report ~89.9% with extended thinking. Aider Polyglot, BBH, Tau-bench, Terminal-Bench, MRCR, SimpleQA, HLE carry no clean published Sonnet-4.6 figure and are null.)
Sonnet 4.6 sits firmly in the fast latency tier — Anthropic labels it "fast," with output around 75 tokens/sec and time-to-first-token under ~2 seconds in typical use. This is the reason it is the default Claude Code model: it feels responsive in interactive IDE loops and chat while still clearing the frontier-enough bar for coding. Extended thinking adds latency only when explicitly enabled with a budget, giving developers a clean speed/quality knob.
| Surface | Cost | Notes |
|---|---|---|
| API input | $3 / 1M tok | Standard rate |
| API output | $15 / 1M tok | Standard rate |
| Cached input (read/hit) | $0.30 / 1M tok | 0.1x base |
| Cache write (5m / 1h) | $3.75 / $6 per 1M tok | 1.25x / 2x base |
| Batch (in/out) | $1.50 / $7.50 per 1M tok | 50% off both |
| Web search tool | $10 / 1,000 searches | plus token costs |
| Direct UI | $20/mo Pro · $100/mo Max 5x · $200/mo Max 20x | claude.ai; also on Free plan |
| Free tier | claude.ai Free plan | daily message caps |
| Rate limits | Tiered (Tier 1–4 + Enterprise) | Priority Tier supported |
Proprietary, no open weights or self-hosting. Available first-party via the Claude API and Claude Platform on AWS, plus Amazon Bedrock (global and regional endpoints), Google Vertex AI (global, multi-region, regional), and Microsoft Foundry. Regional/multi-region endpoints carry a 10% premium; first-party US-only routing via inference_geo: "us" adds 1.1x. Data residency options include US and global.
Governed by Anthropic's RSP v3.0 and deployed under ASL-3 protections. No training on API inputs by default; opt-out and zero-retention available. Compliance: SOC 2 Type II, ISO 27001:2022, ISO/IEC 42001:2023, HIPAA (BAA available), GDPR. No forced content-moderation classifier on API output; refusal calibration is mature and slightly more direct than Sonnet 4.5.
SDKs in Python, TypeScript, Java, Go, Ruby, and C#. First-class in the Claude Agent SDK and Claude Code (its default model), plus LangChain, LlamaIndex, Vercel AI SDK, and Pydantic AI. Selectable in Cursor, GitHub Copilot, Windsurf, Replit, and Sourcegraph. Popularity is dominant — it is the highest-volume model in most Anthropic-centric production stacks.
Cost and speed. You get ~90% of the coding capability at 60% of input price and noticeably faster latency; reserve Opus for frontier-hard jobs.
Yes — served at standard per-token pricing with no long-context premium; caching and batch apply across the full window.
Use the explicit extended-thinking budget or set effort deliberately; max-effort can use ~3x the output tokens.
Yes — no training on inputs, SOC 2 Type II, ISO 27001/42001, HIPAA BAA, GDPR, plus data-residency options.
First-party Claude API plus Bedrock, Vertex AI, and Microsoft Foundry with regional endpoints.
Now for new builds — 4.6 is the same price with 5x context and large OSWorld/ARC gains.
Comparable workhorse tier; OpenAI is generally cheaper on input, Anthropic stronger on agentic coding and computer use.
Cheaper at the low end with large context, but weaker on SWE-bench Verified and OSWorld.
Primary references used to verify this review.
Does not train on API inputs by default
Last verified 2026-05-27