by Anthropic · Claude 4 family · best for legacy snapshot, maintenance only
Claude Opus 4.1, released August 5, 2025, is the oldest Claude Opus model still GA — a drop-in upgrade to Opus 4 with better multi-file refactoring and detail tracking, but at the legacy $15/$75 price that Opus 4.5 cut to $5/$25 three months later. For a buyer, the single sentence is this: a competent but obsolete model with no cost-quality case for new work, since every newer Opus (4.5, 4.6, 4.7) is both cheaper and better — keep it only for snapshot-pinned compliance or reproducibility needs.
| Benchmark | Score | Source |
|---|---|---|
| MMLU | 88.8% | datacamp.com 2025-08-05T00:00:00.000Z |
| MMMU | 77.1% | datacamp.com 2025-08-05T00:00:00.000Z |
| LMArena Elo | 1425 | openlm.ai 2026-05-28T00:00:00.000Z |
| GPQA Diamond | 80.9% | llm-stats.com 2025-08-05T00:00:00.000Z |
| LMArena Coding Elo | 1475 | openlm.ai 2026-05-28T00:00:00.000Z |
| SWE-bench Verified | 74.5% | anthropic.com 2025-08-05T00:00:00.000Z |
| Artificial Analysis Index | 33 | artificialanalysis.ai 2025-08-05T00:00:00.000Z |
Six personas, six verdicts — the same panel that reviews every product on TopReviewed.
“There is no strategic case for new deployment on Opus 4.1 — it's a maintenance-only target in mid-2026.”
For a decision maker in mid-2026, Opus 4.1 is maintenance-only. The November 2025 Opus 4.5 release cut the tier price 3x while raising capability; 4.6 added 5x context; 4.7 added an agentic-coding lift at the same price. There is no strategic argument for new deployment on 4.1. Continued use is reasonable only in narrow compliance or reproducibility scenarios where a pinned snapshot is contractually required. The migration plan should target Opus 4.7 with a Sonnet 4.6 fallback for cost-sensitive routes.
“Opus 4.1's only remaining role is as the 'before' in Anthropic's price-cut story — a baseline, not a contender.”
In market terms, Opus 4.1 is now mostly a reference point: the $15/$75 baseline that made Opus 4.5's 3x cut look dramatic. It validated Anthropic's coding direction with named partners (GitHub, Rakuten, Windsurf), which had strategic value at the time, but it holds no current differentiation. Its niche popularity tier reflects reality — it persists only where snapshot-pinning or contract terms keep it alive. Strategically there is nothing to position around.
“The most expensive supported Claude at $15/$75 — three times the price of better models. Finance should drive it out.”
Opus 4.1 is the most expensive currently-supported Claude at $15/$75, three times the rate of 4.5/4.6/4.7, and even its batch rate ($7.50/$37.50) is 3x Opus 4.7's batch. There is no scenario where 4.1 wins on cost-to-quality against any newer Opus generation, and its writing quality is broadly comparable to Sonnet 4.6 at $3/$15. Finance should actively drive migration off 4.1 wherever it remains in production — the value-per-dollar is the worst in the set.
“The 32k output cap alone disqualifies it for modern agent work — start on 4.7, not here.”
Developers building today should not start on Opus 4.1. The 32k max output cap alone is a meaningful constraint relative to the 128k on 4.6/4.7, SWE-bench Verified 74.5% is well below 4.5's 80.9%, and the per-token cost is 3x higher. For existing 4.1 integrations the migration path to 4.7 is the right one; the minor prompt re-tuning is a small cost relative to the capability and cost delta. Tool use is the same shape across the family, so migration is mostly validation, not rebuild.
“On casual chat it still feels fine — but the cost behind it is impossible to justify versus newer models.”
For an end user, the experience on Opus 4.1 is similar to newer Opus models on casual chat. Latency is moderate-to-slow. The user-visible gaps appear on hard reasoning, vision detail, and very long sessions, where 4.6 and 4.7 are meaningfully better, and the 32k output cap can truncate long generations. Refusal calibration is slightly older-generation. The January 2025 reliable cutoff is now well over a year stale. There is no reason to route a consumer experience to 4.1 given the cost.
“A year-old model at triple the price of better successors — the clearest 'do not buy new' in the Claude lineup.”
There is little to debunk because Anthropic does not market Opus 4.1 as current — it is openly a legacy model. The honest skeptical verdict is simply that it is dominated on every axis that matters: 3x the price of 4.5/4.6/4.7, a smaller 32k output cap, weaker SWE-bench, a dated January 2025 cutoff, and only secondary-source benchmark coverage (reflected in the medium research confidence). Its one genuine merit is as a pinned, reproducible snapshot for compliance — a real but narrow use. For everyone else, it is the clearest "do not buy new" in the Claude lineup.
The full research notes behind this review — verified against primary sources.
Anthropic discloses no parameter count, layer count, or attention mechanism — null/unknown. Disclosed: a 200k-token context window, a 32k max output (materially smaller than the 128k on Opus 4.6/4.7), and extended thinking with token budgets (up to 64k). It uses the standard pre-4.7 Claude tokenizer. Opus 4.1 is a point upgrade over Opus 4 rather than a new architecture.
Coding (8.3): SWE-bench Verified 74.5%, LMArena coding Elo 1475 — solid for 2025, well behind every newer Opus. Reasoning (8.3): GPQA Diamond ~80.9%, MMLU 88.8%, AA Index ~33. Math (8.0): AIME class strong but no clean 2025 figure published for 4.1 specifically. Agentic/tool use (8.0): full first-party tool suite; agentic benchmarks predate the big Opus 4.5/4.6 jumps. Long-context (7.5): 200k window. Multilingual (8.5): broad coverage. Vision (8.3) and document/OCR (8.0): MMMU ~77.1%. Instruction-following (8.3): good for its generation. Function-calling (8.5): robust. Safety calibration (9.0): ASL-3, mature though older-generation refusal behavior. Realtime-data (6.5): January 2025 cutoff is the oldest among current Opus models; web search/fetch mitigate.
| Benchmark | Score | vs Predecessor | vs Successor | Source |
|---|---|---|---|---|
| SWE-bench Verified | 74.5% | +2.0 vs Opus 4 (72.5%) | behind Opus 4.5 (80.9%) | Anthropic |
| GPQA Diamond | 80.9% | improved vs Opus 4 (79.6%) | behind Opus 4.5 (87.0%) | llm-stats |
| MMLU | 88.8% | ~flat vs Opus 4 | behind Opus 4.5 (90.8%) | DataCamp |
| MMMU | 77.1% | improved | behind Opus 4.5 (80.7%) | DataCamp |
| LMArena Elo | 1425 | improved | behind Opus 4.5 (1468) | OpenLM |
| LMArena Coding Elo | 1475 | improved | behind Opus 4.5 (1510) | OpenLM |
| Artificial Analysis Index | ~33 | improved | behind Opus 4.5 (43) | AA |
(MMLU-Pro, AIME 2025, MATH-500, HumanEval, LiveCodeBench, Aider Polyglot, IFEval, BBH, Tau-bench, Terminal-Bench, MRCR, SimpleQA, HLE carry no clean published Opus-4.1 figure and are null. GPQA Diamond and AA Index figures are from secondary aggregators, hence medium research confidence.)
Output speed is roughly ~45 tokens/sec with time-to-first-token near ~1.8s — slow tier, consistent with the Opus family. The 32k max output cap also constrains very long single-pass generations relative to the 128k on Opus 4.6/4.7. It is a deliberate model, but with none of the latency or efficiency improvements of later Opus generations.
| Surface | Cost | Notes |
|---|---|---|
| API input | $15 / 1M tok | 3x newer Opus generations |
| API output | $75 / 1M tok | 3x newer Opus generations |
| Cached input (read/hit) | $1.50 / 1M tok | 0.1x base |
| Cache write (5m / 1h) | $18.75 / $30 per 1M tok | 1.25x / 2x base |
| Batch (in/out) | $7.50 / $37.50 per 1M tok | 50% off both; still 3x Opus 4.7 batch |
| Web search tool | $10 / 1,000 searches | plus token costs |
| Direct UI | $20/mo Pro · $100/mo Max 5x · $200/mo Max 20x | claude.ai |
| Free tier | none for Opus on API | one-time API trial credits only |
| Rate limits | Tiered (Tier 1–4 + Enterprise) | Priority Tier supported |
Proprietary, no open weights or self-hosting. First-party via the Claude API and Claude Platform on AWS, plus Amazon Bedrock, Google Vertex AI, and Microsoft Foundry. Opus 4.1 predates the 4.5-generation global/regional endpoint split and the inference_geo parameter, so it always uses standard routing. Data residency options are more limited than on 4.5+ models.
Governed by Anthropic's RSP v3.0 and deployed under ASL-3 protections (ASL-3 was activated with the Opus 4 generation, of which 4.1 is a point release). No training on API inputs by default; opt-out and zero-retention available. Compliance: SOC 2 Type II, ISO 27001:2022, ISO/IEC 42001:2023, HIPAA (BAA available), GDPR. Refusal calibration is mature but reflects an older generation than 4.5+. No forced content-moderation classifier.
SDKs in Python, TypeScript, Java, Go, Ruby, and C#. Works with the Claude Agent SDK, Claude Code, LangChain, LlamaIndex, Vercel AI SDK, and Pydantic AI; validated at release with GitHub, Rakuten, and Windsurf. Popularity is niche and declining, persisting mainly in snapshot-pinned deployments.
Only to maintain a snapshot-pinned integration for compliance or reproducibility; otherwise no.
It predates the November 2025 Opus 4.5 price reset that cut the tier to $5/$25.
Opus 4.7 for capability, or Sonnet 4.6 for a 5x cheaper option with comparable writing quality.
Mostly prompt validation — tool use and API shape are consistent across the family; expect minor re-tuning.
Yes — no training on inputs, SOC 2 Type II, ISO 27001/42001, HIPAA BAA, GDPR; ASL-3.
Better multi-file refactoring and detail tracking, and SWE-bench Verified from ~72.5% to 74.5%.
Primary references used to verify this review.
Does not train on API inputs by default
Last verified 2026-05-27