by Anthropic · Claude 4 family · best for prior frontier, stable production target
Claude Opus 4.6 was Anthropic's flagship from February 5, 2026 until Opus 4.7 superseded it in April. It remains fully supported and widely deployed: a frontier all-rounder with SWE-bench Verified 80.8%, GPQA Diamond 91.3%, a 1M-token context at standard pricing, and the distinction of a stable tokenizer that Opus 4.7's new tokenizer broke. For a buyer, the single sentence is this: a still-excellent frontier model whose main remaining advantage over 4.7 is prompt/tokenizer stability and the fact that identical text can cost less than on 4.7.
| Benchmark | Score | Source |
|---|---|---|
| Humanity's Last Exam | 53.1% | vellum.ai 2026-02-05T00:00:00.000Z |
| MMMU | 77.3% | vellum.ai 2026-02-05T00:00:00.000Z |
| MMLU-Pro | 88.3% | vellum.ai 2026-02-05T00:00:00.000Z |
| AIME 2025 | 85% | datacamp.com 2026-02-05T00:00:00.000Z |
| HumanEval | 95% | morphllm.com 2026-02-05T00:00:00.000Z |
| TAU-bench | 91.9% | vellum.ai 2026-02-05T00:00:00.000Z |
| LMArena Elo | 1490 | openlm.ai 2026-05-28T00:00:00.000Z |
| GPQA Diamond | 91.3% | vellum.ai 2026-02-05T00:00:00.000Z |
| Terminal-Bench | 65.4% | vellum.ai 2026-02-05T00:00:00.000Z |
| MRCR Long Context | 76% | morphllm.com 2026-02-05T00:00:00.000Z |
| LMArena Coding Elo | 1535 | openlm.ai 2026-05-28T00:00:00.000Z |
| SWE-bench Verified | 80.8% | vellum.ai 2026-02-05T00:00:00.000Z |
| Artificial Analysis Index | 53 | artificialanalysis.ai 2026-02-05T00:00:00.000Z |
Six personas, six verdicts — the same panel that reviews every product on TopReviewed.
“Opus 4.6 is a solid production target, but the strategic question is now migration timing, not continued use.”
Opus 4.6 remains a strong production model, and Anthropic's own docs recommend migrating to 4.7. Pricing is identical, infra fit is unchanged, and the capability lift on 4.7 is real on agentic coding. The argument for staying is stability: Opus 4.7's new tokenizer and more-literal instruction following can shift behavior on tuned prompt suites, and 4.6's tokenizer keeps cost models intact. For most buyers the right call is a controlled migration over a quarter rather than treating 4.6 as a long-term destination.
“Opus 4.6 was the model that briefly gave Anthropic the #1 intelligence-index slot — its legacy is the price-parity ladder.”
Opus 4.6's market role was to take the top of the Artificial Analysis Index for the first time while holding the $5/$25 price Opus 4.5 had set, proving Anthropic could lead on raw intelligence without raising prices. That positioning compounded ecosystem trust. Today its differentiation is mostly historical: 4.7 owns the coding narrative and GPT-5.5 leads the index. Its remaining strategic value is as the stable rung on a price-parity ladder that lets teams adopt Opus capability without repeated cost renegotiation.
“Identical rate card to 4.7 — and because 4.6 keeps the old tokenizer, the same work can actually bill less here.”
On headline rates Opus 4.6 is a wash with 4.7 at $5/$25, with the same cache and batch discounts and the same 1M-context flat pricing. The wrinkle that favors 4.6 is the tokenizer: Opus 4.7's new tokenizer can use up to 35% more tokens for identical text, so for workloads where 4.6's output quality is sufficient, staying can be cheaper per equivalent task for a quarter or two while planning migration. There is no cost penalty to remaining on 4.6 short-term, and a real (if modest) cost reason to do so.
“Opus 4.6 was the first 'throw the whole repo at it' model and it still is — but new code I write on 4.7.”
For builders, Opus 4.6 was a great model and still is, but 4.7 has materially better coding numbers at the same price, so new code goes to 4.7. Where 4.6 stays useful is integrations tuned tightly to its instruction-following quirks, which do not always port cleanly. Adaptive thinking and four effort levels work well, tool use is identical to 4.7, and the 1M context made this generation the first practical whole-repo model. The honest framing: 4.6 is the previous best, 4.7 is the current best.
“On casual chat most users can't tell 4.6 from 4.7 — and some prefer 4.6's slightly warmer voice.”
For a consumer chat product, Opus 4.6 still delivers a polished experience: high conversation quality, calibrated refusals, working vision, and moderate latency with a lower time-to-first-token than 4.7's max-effort mode. The voice has a slight Anthropic warmth that some users prefer to 4.7's more direct, literal style. The May 2025 reliable knowledge cutoff is starting to feel dated for current events, but web search fills the gap. Most end users will not perceive a difference between 4.6 and 4.7 in casual use.
“A genuinely strong model now mostly notable for being cheaper-per-task than its own successor.”
Opus 4.6's benchmarks are real and were briefly best-in-class, so there is no hype problem with the model itself. The skeptical point is positioning: it has been superseded across nearly every eval by 4.7 at the same list price, and its main live advantage is the accident that its older tokenizer bills fewer tokens than 4.7's. The ARC-AGI-2 jump to 68.8% was impressive but is a single benchmark, and like all Opus models the latency is poor for interactive use. As a data point it is honest; as a purchase decision for new work it loses cleanly to 4.7.
The full research notes behind this review — verified against primary sources.
Anthropic discloses no parameter count, layer count, or attention mechanism — null/unknown. Disclosed: a 1M-token context window at standard pricing, 128k synchronous max output (300k via batch beta), adaptive thinking with four effort levels, and context compaction for long-running agents. It uses the standard pre-4.7 Claude tokenizer, so cost models and prompt suites built for Opus 4.5/4.6 remain stable — a meaningful operational advantage over migrating to Opus 4.7's new tokenizer.
Coding (9.3): SWE-bench Verified 80.8%, HumanEval 95%, LMArena coding Elo 1535 — frontier, just behind Opus 4.7's agentic-coding lead. Reasoning (9.3): GPQA Diamond 91.3%, MMLU-Pro 88.3%, ARC-AGI-2 68.8%, HLE with tools 53.1%, AA Index 53 (top of the index at release). Math (8.8): AIME 2025 85.0%. Agentic/tool use (9.3): Terminal-Bench 2.0 65.4%, OSWorld 72.7%, Tau2-bench retail 91.9% / telecom 99.3%, plus context compaction and agent teams. Long-context (9.2): 1M tokens at standard pricing, MRCR v2 76.0%. Multilingual (9.0): MMMLU 91.1%. Vision (8.5) and document/OCR (8.3): MMMU-Pro 77.3% with tools; solid but below Opus 4.7's high-res pipeline. Instruction-following (9.0): strong, with a slightly less literal style than 4.7. Function-calling (9.3): robust. Safety calibration (9.3): ASL-3. Realtime-data (7.0): May 2025 cutoff plus web search/fetch.
| Benchmark | Score | vs Predecessor | vs Successor | Source |
|---|---|---|---|---|
| SWE-bench Verified | 80.8% | ~flat vs Opus 4.5 (80.9%) | behind Opus 4.7 (87.6%) | Vellum |
| SWE-bench Pro | 53.4% | improved | behind Opus 4.7 (64.3%) | Vellum |
| GPQA Diamond | 91.3% | +4.3 vs Opus 4.5 (87.0%) | behind Opus 4.7 (94.2%) | Vellum |
| MMLU-Pro | 88.3% | improved | frontier tier | Vellum |
| AIME 2025 | 85.0% | improved | frontier tier | DataCamp |
| Terminal-Bench 2.0 | 65.4% | +5.6 vs Opus 4.5 (59.8%) | behind Opus 4.7 (69.4%) | Vellum |
| Tau2-bench Retail | 91.9% | +3.0 vs Opus 4.5 (88.9%) | frontier tool use | Vellum |
| OSWorld-Verified | 72.7% | +6.4 vs Opus 4.5 (66.3%) | behind Opus 4.7 (78.0%) | Vellum |
| ARC-AGI-2 | 68.8% | +31.2 vs Opus 4.5 (37.6%) | strong novel-puzzle | Morph |
| HLE (with tools) | 53.1% | +9.7 vs Opus 4.5 (43.4%) | behind Opus 4.7 (54.7%) | Vellum |
| MRCR v2 (long context) | 76.0% | improved | frontier | Morph |
| LMArena Elo | 1490 | improved | behind Opus 4.7 (1503) | OpenLM |
| LMArena Coding Elo | 1535 | improved | behind Opus 4.7 (1554) | OpenLM |
| Artificial Analysis Index | 53 | improved | behind Opus 4.7 (57) | AA |
(MATH-500, LiveCodeBench, Aider Polyglot, IFEval, BBH, SimpleQA carry no clean published Opus-4.6 figure and are null.)
Output speed is ~45.9 tokens/sec with time-to-first-token ~1.76s in high-effort mode (Artificial Analysis). Anthropic labels comparative latency "moderate"; for the compare engine this sits in the slow tier relative to Sonnet/Haiku, though its TTFT is markedly lower than Opus 4.7's adaptive max-effort latency. Fast Mode (beta, 6x price) is available for low-latency needs. It is a deliberate model suited to hard work and batch, not snappy chat.
| Surface | Cost | Notes |
|---|---|---|
| API input | $5 / 1M tok | Identical to Opus 4.7/4.5 |
| API output | $25 / 1M tok | Identical |
| Cached input (read/hit) | $0.50 / 1M tok | 0.1x base |
| Cache write (5m / 1h) | $6.25 / $10 per 1M tok | 1.25x / 2x base |
| Batch (in/out) | $2.50 / $12.50 per 1M tok | 50% off both |
| Fast Mode (beta) | $30 in / $150 out per 1M tok | 6x premium for low latency |
| Web search tool | $10 / 1,000 searches | plus token costs |
| Direct UI | $20/mo Pro · $100/mo Max 5x · $200/mo Max 20x | claude.ai |
| Free tier | none for Opus on API | one-time API trial credits only |
| Rate limits | Tiered (Tier 1–4 + Enterprise) | Priority Tier supported |
Proprietary, no open weights or self-hosting. First-party via the Claude API and Claude Platform on AWS, plus Amazon Bedrock (global and regional endpoints), Google Vertex AI (global, multi-region, regional), and Microsoft Foundry. Regional/multi-region endpoints carry a 10% premium; first-party US-only routing via inference_geo: "us" adds 1.1x. Data residency options include US and global.
Governed by Anthropic's RSP v3.0 and deployed under ASL-3 protections. No training on API inputs by default; opt-out and zero-retention available. Compliance: SOC 2 Type II, ISO 27001:2022, ISO/IEC 42001:2023, HIPAA (BAA available), GDPR. No forced content-moderation classifier; refusal calibration is mature with a slightly warmer tone than Opus 4.7.
SDKs in Python, TypeScript, Java, Go, Ruby, and C#. Works with the Claude Agent SDK, Claude Code, LangChain, LlamaIndex, Vercel AI SDK, and Pydantic AI; selectable in Cursor, GitHub Copilot, Windsurf, and Replit. Popularity is mainstream and remains high in production due to tokenizer/prompt stability.
For new builds, move to 4.7. For tuned production prompts, plan a controlled migration over a quarter — 4.6 stays fully supported meanwhile.
On rate card, identical; in practice 4.6's older tokenizer can bill up to ~35% fewer tokens for the same text.
Yes, at standard pricing with no premium, plus context compaction for long agents.
Yes — no training on inputs, SOC 2 Type II, ISO 27001/42001, HIPAA BAA, GDPR, data-residency options.
First-party Claude API plus Bedrock, Vertex AI, and Microsoft Foundry with regional endpoints.
1M context at the Opus tier, adaptive thinking with four effort levels, context compaction, and agent teams in Claude Code.
Primary references used to verify this review.
Does not train on API inputs by default
Last verified 2026-05-27