Claude Opus 4.6

GA

by Anthropic · Claude 4 family · best for prior frontier, stable production target

FrontierReasoningCodingMultimodalLong-Context
8.6
AI Panel Score
Value 7.8/10

Claude Opus 4.6 was Anthropic's flagship from February 5, 2026 until Opus 4.7 superseded it in April. It remains fully supported and widely deployed: a frontier all-rounder with SWE-bench Verified 80.8%, GPQA Diamond 91.3%, a 1M-token context at standard pricing, and the distinction of a stable tokenizer that Opus 4.7's new tokenizer broke. For a buyer, the single sentence is this: a still-excellent frontier model whose main remaining advantage over 4.7 is prompt/tokenizer stability and the fact that identical text can cost less than on 4.7.

What's new

  • vs Opus 4.5: 1M-token context window at standard pricing (4.5 was 200k).
  • Adaptive thinking introduced at the Opus tier, plus four explicit effort levels.
  • Context compaction for sustained agentic tasks and agent teams in Claude Code.
  • Topped the Artificial Analysis Intelligence Index at release (53), Anthropic's first #1 on that index.
  • ARC-AGI-2 jumped to 68.8% (from 37.6% on Opus 4.5).
  • Held Opus pricing at $5/$25 — the rate set by Opus 4.5 in November 2025.

Benchmarks

BenchmarkScoreSource
Humanity's Last Exam53.1%vellum.ai 2026-02-05T00:00:00.000Z
MMMU77.3%vellum.ai 2026-02-05T00:00:00.000Z
MMLU-Pro88.3%vellum.ai 2026-02-05T00:00:00.000Z
AIME 202585%datacamp.com 2026-02-05T00:00:00.000Z
HumanEval95%morphllm.com 2026-02-05T00:00:00.000Z
TAU-bench91.9%vellum.ai 2026-02-05T00:00:00.000Z
LMArena Elo1490openlm.ai 2026-05-28T00:00:00.000Z
GPQA Diamond91.3%vellum.ai 2026-02-05T00:00:00.000Z
Terminal-Bench65.4%vellum.ai 2026-02-05T00:00:00.000Z
MRCR Long Context76%morphllm.com 2026-02-05T00:00:00.000Z
LMArena Coding Elo1535openlm.ai 2026-05-28T00:00:00.000Z
SWE-bench Verified80.8%vellum.ai 2026-02-05T00:00:00.000Z
Artificial Analysis Index53artificialanalysis.ai 2026-02-05T00:00:00.000Z

AI Panel Review

Six personas, six verdicts — the same panel that reviews every product on TopReviewed.

Decision Maker8.5/10
Opus 4.6 is a solid production target, but the strategic question is now migration timing, not continued use.

Opus 4.6 remains a strong production model, and Anthropic's own docs recommend migrating to 4.7. Pricing is identical, infra fit is unchanged, and the capability lift on 4.7 is real on agentic coding. The argument for staying is stability: Opus 4.7's new tokenizer and more-literal instruction following can shift behavior on tuned prompt suites, and 4.6's tokenizer keeps cost models intact. For most buyers the right call is a controlled migration over a quarter rather than treating 4.6 as a long-term destination.

Strategic Fit 8Vendor Risk 7Roadmap Confidence 8
Pros
  • Frontier, stable tokenizer, multi-cloud, flat price
Cons
  • Superseded by 4.7
  • migrate eventually
Right for: existing 4.6 production with tuned prompts
Avoid if: starting a new build (use 4.7)
Domain Strategist8.5/10
Opus 4.6 was the model that briefly gave Anthropic the #1 intelligence-index slot — its legacy is the price-parity ladder.

Opus 4.6's market role was to take the top of the Artificial Analysis Index for the first time while holding the $5/$25 price Opus 4.5 had set, proving Anthropic could lead on raw intelligence without raising prices. That positioning compounded ecosystem trust. Today its differentiation is mostly historical: 4.7 owns the coding narrative and GPT-5.5 leads the index. Its remaining strategic value is as the stable rung on a price-parity ladder that lets teams adopt Opus capability without repeated cost renegotiation.

Competitive Positioning 8Differentiation 7Market Timing 8
Pros
  • Held the index lead at flat price
  • trusted
Cons
  • Differentiation eroded by 4.7
Right for: stability-first adopters
Avoid if: you need the current capability frontier
Finance Lead8.5/10
Identical rate card to 4.7 — and because 4.6 keeps the old tokenizer, the same work can actually bill less here.

On headline rates Opus 4.6 is a wash with 4.7 at $5/$25, with the same cache and batch discounts and the same 1M-context flat pricing. The wrinkle that favors 4.6 is the tokenizer: Opus 4.7's new tokenizer can use up to 35% more tokens for identical text, so for workloads where 4.6's output quality is sufficient, staying can be cheaper per equivalent task for a quarter or two while planning migration. There is no cost penalty to remaining on 4.6 short-term, and a real (if modest) cost reason to do so.

Cost Efficiency 8Pricing Transparency 9Value per Dollar 8
Pros
  • Flat rates, old tokenizer can be cheaper, full discounts
Cons
  • Capability left on the table vs 4.7
Right for: cost-sensitive teams not yet needing 4.7
Avoid if: you need 4.7's coding lift now
Domain Practitioner8.5/10
Opus 4.6 was the first 'throw the whole repo at it' model and it still is — but new code I write on 4.7.

For builders, Opus 4.6 was a great model and still is, but 4.7 has materially better coding numbers at the same price, so new code goes to 4.7. Where 4.6 stays useful is integrations tuned tightly to its instruction-following quirks, which do not always port cleanly. Adaptive thinking and four effort levels work well, tool use is identical to 4.7, and the 1M context made this generation the first practical whole-repo model. The honest framing: 4.6 is the previous best, 4.7 is the current best.

API Ergonomics 9Tool/Agent Support 9Reliability 9
Pros
  • Stable tooling, whole-repo context, four effort levels
Cons
  • Behind 4.7 on agentic coding
Right for: maintaining tuned 4.6 integrations
Avoid if: building fresh agents (use 4.7)
Power User8.5/10
On casual chat most users can't tell 4.6 from 4.7 — and some prefer 4.6's slightly warmer voice.

For a consumer chat product, Opus 4.6 still delivers a polished experience: high conversation quality, calibrated refusals, working vision, and moderate latency with a lower time-to-first-token than 4.7's max-effort mode. The voice has a slight Anthropic warmth that some users prefer to 4.7's more direct, literal style. The May 2025 reliable knowledge cutoff is starting to feel dated for current events, but web search fills the gap. Most end users will not perceive a difference between 4.6 and 4.7 in casual use.

Output Quality 8.5Speed 7Everyday Usefulness 8.5
Pros
  • Polished, warmer tone, lower TTFT than 4.7
Cons
  • Dated cutoff
  • superseded
Right for: chat surfaces valuing tone
Avoid if: you need the newest knowledge or coding edge
Skeptic8/10
A genuinely strong model now mostly notable for being cheaper-per-task than its own successor.

Opus 4.6's benchmarks are real and were briefly best-in-class, so there is no hype problem with the model itself. The skeptical point is positioning: it has been superseded across nearly every eval by 4.7 at the same list price, and its main live advantage is the accident that its older tokenizer bills fewer tokens than 4.7's. The ARC-AGI-2 jump to 68.8% was impressive but is a single benchmark, and like all Opus models the latency is poor for interactive use. As a data point it is honest; as a purchase decision for new work it loses cleanly to 4.7.

Claim Accuracy 8.5Weakness Severity 7Hype vs Reality 8
Pros
  • Numbers are real and well-sourced
Cons
  • Superseded on nearly everything
  • slow
Right for: skeptics optimizing token cost short-term
Avoid if: you would otherwise just use 4.7

Strengths

  • Frontier all-rounder: top scores across coding, science, math, vision, and agentic benchmarks.
  • 1M context at standard pricing with no premium.
  • ARC-AGI-2 68.8% is a meaningful novel-problem-reasoning signal.
  • Stable tokenizer and prompt behavior — existing prompt suites work without re-tuning.
  • Adaptive thinking and four effort levels give granular cost-quality trade-offs.

Limitations

  • Clearly behind Opus 4.7 on agentic coding (SWE-bench Pro 53.4% vs 64.3%, OSWorld 72.7% vs 78.0%).
  • Mid-pack on long browse-and-synthesize loops; superseded by 4.7 on most evals.
  • May 2025 reliable cutoff misses late-2025 events without web search.
  • Moderate latency; not ideal for sub-second chat.
  • Anthropic's official guidance is to migrate to Opus 4.7 for new builds.

Best use cases

  • Production systems already integrated against Opus 4.6 where prompt and tokenizer stability beat the 4.7 capability lift.
  • Long-context document workflows where the 1M-token window is the differentiator.
  • Hard reasoning benchmarks (GPQA, AIME, ARC-AGI-2) where Opus-tier scores justify the cost.
  • Agent teams and Claude Code workflows tuned to Opus 4.6 tooling.

Deep dive

The full research notes behind this review — verified against primary sources.

Architecture

Anthropic discloses no parameter count, layer count, or attention mechanism — null/unknown. Disclosed: a 1M-token context window at standard pricing, 128k synchronous max output (300k via batch beta), adaptive thinking with four effort levels, and context compaction for long-running agents. It uses the standard pre-4.7 Claude tokenizer, so cost models and prompt suites built for Opus 4.5/4.6 remain stable — a meaningful operational advantage over migrating to Opus 4.7's new tokenizer.

Capabilities

Coding (9.3): SWE-bench Verified 80.8%, HumanEval 95%, LMArena coding Elo 1535 — frontier, just behind Opus 4.7's agentic-coding lead. Reasoning (9.3): GPQA Diamond 91.3%, MMLU-Pro 88.3%, ARC-AGI-2 68.8%, HLE with tools 53.1%, AA Index 53 (top of the index at release). Math (8.8): AIME 2025 85.0%. Agentic/tool use (9.3): Terminal-Bench 2.0 65.4%, OSWorld 72.7%, Tau2-bench retail 91.9% / telecom 99.3%, plus context compaction and agent teams. Long-context (9.2): 1M tokens at standard pricing, MRCR v2 76.0%. Multilingual (9.0): MMMLU 91.1%. Vision (8.5) and document/OCR (8.3): MMMU-Pro 77.3% with tools; solid but below Opus 4.7's high-res pipeline. Instruction-following (9.0): strong, with a slightly less literal style than 4.7. Function-calling (9.3): robust. Safety calibration (9.3): ASL-3. Realtime-data (7.0): May 2025 cutoff plus web search/fetch.

Benchmark analysis

Benchmark Score vs Predecessor vs Successor Source
SWE-bench Verified 80.8% ~flat vs Opus 4.5 (80.9%) behind Opus 4.7 (87.6%) Vellum
SWE-bench Pro 53.4% improved behind Opus 4.7 (64.3%) Vellum
GPQA Diamond 91.3% +4.3 vs Opus 4.5 (87.0%) behind Opus 4.7 (94.2%) Vellum
MMLU-Pro 88.3% improved frontier tier Vellum
AIME 2025 85.0% improved frontier tier DataCamp
Terminal-Bench 2.0 65.4% +5.6 vs Opus 4.5 (59.8%) behind Opus 4.7 (69.4%) Vellum
Tau2-bench Retail 91.9% +3.0 vs Opus 4.5 (88.9%) frontier tool use Vellum
OSWorld-Verified 72.7% +6.4 vs Opus 4.5 (66.3%) behind Opus 4.7 (78.0%) Vellum
ARC-AGI-2 68.8% +31.2 vs Opus 4.5 (37.6%) strong novel-puzzle Morph
HLE (with tools) 53.1% +9.7 vs Opus 4.5 (43.4%) behind Opus 4.7 (54.7%) Vellum
MRCR v2 (long context) 76.0% improved frontier Morph
LMArena Elo 1490 improved behind Opus 4.7 (1503) OpenLM
LMArena Coding Elo 1535 improved behind Opus 4.7 (1554) OpenLM
Artificial Analysis Index 53 improved behind Opus 4.7 (57) AA

(MATH-500, LiveCodeBench, Aider Polyglot, IFEval, BBH, SimpleQA carry no clean published Opus-4.6 figure and are null.)

Speed & latency

Output speed is ~45.9 tokens/sec with time-to-first-token ~1.76s in high-effort mode (Artificial Analysis). Anthropic labels comparative latency "moderate"; for the compare engine this sits in the slow tier relative to Sonnet/Haiku, though its TTFT is markedly lower than Opus 4.7's adaptive max-effort latency. Fast Mode (beta, 6x price) is available for low-latency needs. It is a deliberate model suited to hard work and batch, not snappy chat.

Pricing analysis

Surface Cost Notes
API input $5 / 1M tok Identical to Opus 4.7/4.5
API output $25 / 1M tok Identical
Cached input (read/hit) $0.50 / 1M tok 0.1x base
Cache write (5m / 1h) $6.25 / $10 per 1M tok 1.25x / 2x base
Batch (in/out) $2.50 / $12.50 per 1M tok 50% off both
Fast Mode (beta) $30 in / $150 out per 1M tok 6x premium for low latency
Web search tool $10 / 1,000 searches plus token costs
Direct UI $20/mo Pro · $100/mo Max 5x · $200/mo Max 20x claude.ai
Free tier none for Opus on API one-time API trial credits only
Rate limits Tiered (Tier 1–4 + Enterprise) Priority Tier supported

Deployment & access

Proprietary, no open weights or self-hosting. First-party via the Claude API and Claude Platform on AWS, plus Amazon Bedrock (global and regional endpoints), Google Vertex AI (global, multi-region, regional), and Microsoft Foundry. Regional/multi-region endpoints carry a 10% premium; first-party US-only routing via inference_geo: "us" adds 1.1x. Data residency options include US and global.

Safety & privacy

Governed by Anthropic's RSP v3.0 and deployed under ASL-3 protections. No training on API inputs by default; opt-out and zero-retention available. Compliance: SOC 2 Type II, ISO 27001:2022, ISO/IEC 42001:2023, HIPAA (BAA available), GDPR. No forced content-moderation classifier; refusal calibration is mature with a slightly warmer tone than Opus 4.7.

Ecosystem & tooling

SDKs in Python, TypeScript, Java, Go, Ruby, and C#. Works with the Claude Agent SDK, Claude Code, LangChain, LlamaIndex, Vercel AI SDK, and Pydantic AI; selectable in Cursor, GitHub Copilot, Windsurf, and Replit. Popularity is mainstream and remains high in production due to tokenizer/prompt stability.

Buyer questions

Should I stay on 4.6 or move to 4.7?

For new builds, move to 4.7. For tuned production prompts, plan a controlled migration over a quarter — 4.6 stays fully supported meanwhile.

Is 4.6 cheaper than 4.7?

On rate card, identical; in practice 4.6's older tokenizer can bill up to ~35% fewer tokens for the same text.

Does it have the 1M context?

Yes, at standard pricing with no premium, plus context compaction for long agents.

Is it secure for enterprise?

Yes — no training on inputs, SOC 2 Type II, ISO 27001/42001, HIPAA BAA, GDPR, data-residency options.

Which clouds host it?

First-party Claude API plus Bedrock, Vertex AI, and Microsoft Foundry with regional endpoints.

What did 4.6 introduce?

1M context at the Opus tier, adaptive thinking with four effort levels, context compaction, and agent teams in Claude Code.

Comparable models

Claude Opus 4.7: Direct successor; same price, better on agentic coding and vision, but a new tokenizer that bills more per text.
Claude Sonnet 4.6: Same family; 60% cheaper input, ~1.2 pts behind on SWE-bench Verified, faster.
GPT-5.5 / Gemini 3.1 Pro: Competing flagships; GPT-5.5 leads the AA Index, trade-offs vary by workload.

Sources

Primary references used to verify this review.

Model specs

Input price
$5 / Mtok
Output price
$25 / Mtok
Cached input
$0.50 / Mtok
Batch (in/out)
$2.50 / $12.50
Context window
1M tokens
Max output
128K tokens
Knowledge cutoff
2025-05
Released
2026-02-04
Modalities
text, image → text
Output speed
~45.9 tok/s
License
Proprietary
Clouds
Bedrock, Vertex AI, Azure AI Foundry

Does not train on API inputs by default

Last verified 2026-05-27