Command A+

GALatest Flagship

by Cohere · Command A family · best for sovereign self-hosted enterprise agents

ReasoningMultimodalOpen-Weights
7.4
AI Panel Score
Value 7.5/10

Command A+ is Cohere's enterprise wager: a 218B-parameter mixture-of-experts (25B active) released 2026-05-20 as open weights under a plain Apache-2.0 license, deliberately priced with no public per-token rate card because Cohere wants sovereign and private-deployment contracts, not metered API traffic. It runs on a single NVIDIA B200 (or two H100s) at 4-bit, accepts text and image input across 48 languages, and posts a standout agentic-tool-use result (tau-squared-Bench Telecom 85, up from 37 on its predecessor). It is not a frontier flagship — Artificial Analysis places it mid-pack at 37 — but for regulated buyers who must keep inference inside their own perimeter, a governable, self-hostable, cleanly-licensed model beats a benchmark point.

What's new

  • Cohere's first Command A+ tier: a 218B/25B-active MoE succeeding the 111B dense Command A Reasoning, released open-weights under Apache-2.0
  • Big jump on agentic tool use: tau-squared-Bench Telecom 85% (from 37%) and Terminal-Bench Hard 25% (from 3%) versus the predecessor
  • Multilingual coverage expands to 48 languages (from 23), with text-and-image input for document and screenshot workflows
  • Efficiency-first serving: runs on 1x NVIDIA B200 or 2x H100 at W4A4, with BF16 and FP8 checkpoints published
  • No public per-token price — positioned for sovereign, private, and Model Vault managed deployments rather than open API traffic

Benchmarks

BenchmarkScoreSource
MMMU75.1%cohere.com 2026-05-20T00:00:00.000Z
Artificial Analysis Index37cohere.com 2026-05-20T00:00:00.000Z

AI Panel Review

Six personas, six verdicts — the same panel that reviews every product on TopReviewed.

Decision Maker7.5/10
The only frontier-adjacent model you can run entirely inside your own perimeter under Apache-2.0 - for regulated buyers that beats a benchmark point every time.

Command A+ is a procurement instrument as much as a model. Cohere priced it for sovereign and private deployment rather than token traffic, which is exactly why there is no public rate card: the buyer Cohere wants signs a deployment contract and runs the Apache-2.0 weights on hardware it controls. For banks, defense, healthcare, and government that cannot send data to a US hyperscaler API, that is the whole ballgame, and the single-B200 footprint makes it operationally realistic. The capability is mid-pack on open benchmarks (AA Index 37), so this is not the pick when raw intelligence decides — it is the pick when data residency, auditability, and license cleanliness do.

Strategic Fit 7.5Vendor Risk 7.5Roadmap Confidence 7
Pros
  • Apache-2.0 with full self-host rights
  • runs on a single B200
  • genuine data-sovereignty story
Cons
  • Mid-pack raw capability
  • no public pricing to model
Right for: regulated enterprises that must keep inference inside their own perimeter
Avoid if: you are choosing on benchmark leadership alone
Domain Strategist7.5/10
Cohere stopped chasing the frontier and planted a flag where the hyperscalers can't follow: your data center, your license, your data.

The strategy is legible and differentiated. While Chinese labs commoditize open capability and OpenAI/Anthropic own the frontier API, Cohere is monetizing the one segment those players structurally cannot serve well — the regulated enterprise that will not, or legally cannot, run workloads on a public API. Apache-2.0 weights plus a single-GPU footprint plus 48-language coverage plus tool-use and RAG heritage is a coherent enterprise bundle. The weakness is reach: a no-public-price, sales-led motion scales slowly and cedes the developer mindshare that GLM and Qwen are winning. This is a durable niche, not a market-share play.

Competitive Positioning 8Differentiation 8Market Timing 7.5
Pros
  • Clear sovereign-deployment wedge
  • strong multilingual and tool-use fit
  • single-GPU serving
Cons
  • Slow sales-led distribution
  • low developer visibility
Right for: mapping where enterprise-AI value accrues outside the API majors
Avoid if: your thesis is developer-led, API-first adoption
Finance Lead7.5/10
There's no per-token line item because you buy the deployment, not the tokens - for high, steady volume inside your walls that can be the cheapest frontier-adjacent option going.

The economics invert the usual API calculus. With Apache-2.0 weights and a single-B200 (or dual-H100) serving target at 4-bit, a heavy, steady internal workload amortizes fixed hardware cost far below what metered API tokens would cost at the same volume — and the spend is capex you control, not a variable bill that scales with success. The catch is that there is no published token price to benchmark against, so the model only pencils out if you have the volume and the ops maturity to run it; for spiky or low-volume use, a metered frontier API is cheaper and simpler. Value here is real but conditional.

Cost Efficiency 8Pricing Transparency 6Value per Dollar 7.5
Pros
  • No per-token bill
  • capex you control
  • efficient single-GPU footprint
Cons
  • No public price to model
  • only economical at real volume
Right for: high-volume internal workloads with in-house serving ops
Avoid if: your usage is low, spiky, or you lack GPU operations
Domain Practitioner7/10
Tool use and RAG are where Cohere has always been strong, and tau-squared-Bench Telecom 85 shows it - just don't reach for it as your hardest-problem coder.

For the enterprise-agent builder this is comfortable ground: first-class tool use, structured output, a big jump on agentic telecom tasks (85 vs the predecessor's 37), and 48-language coverage for multinational deployments. Image input and MMMU 75.1 cover document and screenshot workflows. The gaps are honest: Terminal-Bench Hard 25 says it is not a frontier coding agent, the 128K context trails same-month 1M rivals, and running the weights well is a real MLOps commitment. As a RAG-and-tools workhorse inside a regulated stack it fits; as a general coding brain it does not lead.

API Ergonomics 7.5Tool/Agent Support 8.5Reliability 7.5
Pros
  • Strong tool-use and RAG behavior
  • broad multilingual
  • image input
Cons
  • Weak hard-coding numbers
  • 128K context
  • self-host ops overhead
Right for: enterprise agent and RAG pipelines with governance needs
Avoid if: you need best-in-class coding or a giant context window
Power User6/10
It's built for enterprise deployments, not for me - there's no slick consumer app, and on a hard question the frontier models are simply sharper.

Command A+ is not aimed at the individual daily driver, and it shows. There is no polished first-party consumer surface, access is through the API or a self-hosted deployment, and on the hardest analytical or coding questions it lands a clear tier below the frontier flagships (AA Index 37). What it does well for a hands-on user is multilingual work and grounded, tool-augmented answering where hallucination control matters. For most people wanting a chat companion, GLM-5.2 or a frontier API delivers more; Command A+ earns its keep inside an enterprise workflow, not on your personal desktop.

Output Quality 6.5Speed 6.5Everyday Usefulness 6
Pros
  • Strong grounded and multilingual answering
  • image input
Cons
  • No consumer app
  • mid-tier on hard problems
Right for: users working inside an enterprise Cohere deployment
Avoid if: you want a top-tier personal assistant
Skeptic7/10
Honest positioning, for once: Cohere never calls it a frontier model, and AA Index 37 confirms it isn't - the pitch is deployment, not benchmarks.

There is refreshingly little to debunk here because Cohere does not overclaim. The headline improvements — tau-squared-Bench Telecom 85 up from 37, Terminal-Bench Hard 25 up from 3 — are large but measured against its own weak predecessor, so they signal progress, not leadership. Independent placement (AA Index 37) puts it below every open flagship in this month's cohort, and the "no public price" framing, while strategically sound, also conveniently avoids a value comparison it would lose on raw capability. The residual risks are ordinary: undisclosed knowledge cutoff, an Apache-2.0 headline that still warrants a read of the model-card conditions, and benchmarks reported first-party. Nothing dishonest; just a mid-tier model with an unusually smart go-to-market.

Claim Accuracy 8Weakness Severity 6.5Hype vs Reality 7.5
Pros
  • Restrained, verifiable claims
  • Apache-2.0 allows independent checking
Cons
  • Gains measured off a weak predecessor
  • mid-pack absolute capability
Right for: buyers who value honest positioning over benchmark theater
Avoid if: you need the measured capability leader

Strengths

  • Apache-2.0 open weights that run on a single B200 (or two H100s) at 4-bit — genuine sovereign-deployment economics for a 218B model
  • Standout agentic tool use (tau-squared-Bench Telecom 85) rooted in Cohere's RAG and enterprise-agent heritage
  • Broad multilingual coverage (48 languages) with text-and-image input for multinational document workflows
  • Restrained, verifiable positioning — Cohere does not overclaim frontier status

Limitations

  • Mid-pack raw capability: AA Index 37 trails every open flagship in the June cohort, and Terminal-Bench Hard 25 shows weak hard-coding
  • No public per-token price, so value is hard to model and only pencils out at real self-hosted volume
  • 128K context and undisclosed knowledge cutoff trail same-month 1M-context rivals
  • Running the weights well is a real MLOps commitment; there is no polished consumer surface

Best use cases

Regulated enterprises — finance, healthcare, defense, government — that must keep inference inside their own perimeter and want a clean Apache-2.0 license they can audit and self-host. Multilingual document and support automation across the model's 48 languages, and grounded RAG-plus-tools agent pipelines where Cohere's heritage shows. High, steady internal workloads where a single-B200 deployment amortizes below metered API costs. Not the pick when raw benchmark leadership, best-in-class coding, or a giant context window is the deciding factor — GLM-5.2 or a frontier API wins those.

Deep dive

The full research notes behind this review — verified against primary sources.

Architecture

A 218B-parameter sparse mixture-of-experts activating 25B parameters per token, with 128 experts (8 routed plus 1 shared per token) and sliding-window attention interleaved roughly 3:1 with global attention layers — a design tuned for efficient long-document serving on a small GPU footprint. The 25B active path is what lets a 218B model run on a single B200 at 4-bit. Context is 131,072 tokens of input with up to 65,536 tokens of generation. Cohere does not publish a training-token count or a knowledge cutoff, and whether the vision path is native or adapter-based is undisclosed. The weights ship in BF16 with FP8 and W4A4 quantizations for lighter deployment.

Capabilities

Agentic tool use is the design center and the standout score (8.0): tau-squared-Bench Telecom 85 (up from the predecessor's 37) reflects Cohere's long enterprise-agent and RAG heritage, and function calling is first-class (8.5). Multilingual coverage (8.5) across 48 languages is a durable Cohere strength for multinational deployments. Vision (7.5) and document handling (7.0) rest on text-and-image input and MMMU 75.1, covering screenshot and document-image workflows. Reasoning (7.0) is available via the model's reasoning output mode and is benchmark-anchored by the mid-pack AA Index (37). Coding (6.5) is the honest weak spot — Terminal-Bench Hard 25 says this is not a frontier coding agent. Math is left unscored (no AIME/MATH disclosure). No built-in web search, so real-time data scores 0; grounding comes from Cohere's RAG tooling, not in-model retrieval.

Benchmark analysis

Benchmark Score Note Source
tau-squared-Bench Telecom 85% agentic tool use; up from 37% on Command A Reasoning Cohere blog
Terminal-Bench Hard 25% agentic coding; up from 3% Cohere blog
MMMU 75.1 multimodal understanding Cohere blog
MMMU Pro 63 harder multimodal split Cohere blog
AA Intelligence Index 37 independent aggregate Cohere blog

MMMU and the Artificial Analysis Index enter this catalog's comparable columns as clean, standard metrics. The tau-squared-Bench Telecom and Terminal-Bench Hard figures are large improvements but are measured against Cohere's own weak predecessor and on harder/domain-specific splits that do not line up with the generic tau-bench / Terminal-Bench columns used elsewhere, so they are reported here in prose rather than in the comparable benchmark map. Research confidence is medium: benchmarks are first-party and pricing is undisclosed.

Speed & latency

Cohere reports up to 63% higher output tokens/second and a 17% latency reduction versus Command A Reasoning, but publishes no absolute throughput or time-to-first-token figures, and no independent tracker has measured the model at verification time — so speed fields are null. The 25B active-parameter path and single-GPU serving target imply competitive interactive latency for the class; validate on your own hardware, since self-hosted performance depends entirely on your deployment.

Pricing analysis

Surface Cost Notes
API (per token) Not published no public rate card
Model Vault Contract managed private inference
Sovereign / private deploy Contract customer-controlled infrastructure
Self-host (weights) Free Apache-2.0, your own hardware

There is deliberately no per-token price. Cohere's motion is deployment contracts and self-hosting, not metered API traffic — which is the whole strategy for the regulated buyer who will not send data to a public endpoint. All pricing fields are null for that reason.

Deployment & access

Open weights on Hugging Face (CohereLabs/command-a-plus-05-2026-bf16) under Apache-2.0, with BF16, FP8, and W4A4 checkpoints. The headline is footprint: one NVIDIA B200 or two H100s at W4A4 serve the full model, unusually light for 218B parameters thanks to the 25B active path. Managed inference is available through Cohere's Model Vault, and private/sovereign deployments run entirely inside customer infrastructure with customer-chosen data residency. There is no consumer app and no public token API in the usual sense — access is enterprise-contract or self-host.

Safety & privacy

Cohere's positioning is enterprise-governance-first: Apache-2.0 weights that customers run inside their own perimeter, no training on customer data by default, and SOC 2 / GDPR-aligned enterprise terms. There is no single named safety framework analogous to a frontier lab's RSP, and no formal safety-level label is published for this release; content moderation and refusal behavior are the deployer's to configure when self-hosting. The knowledge cutoff is undisclosed. For regulated buyers, the sovereign-deployment model is itself the governance story — data never leaves the perimeter.

Ecosystem & tooling

Cohere's platform tooling (Python and TypeScript SDKs, the v2 Chat API, native tool-use and RAG primitives), managed inference via Model Vault, and self-serve Apache-2.0 weights on Hugging Face with standard serving engines. LangChain and LlamaIndex integrations carry over from the Command line. Cohere's own North enterprise platform is the flagship product built on it, and the developer community is enterprise-weighted rather than hobbyist. Popularity is growing within regulated verticals; it is deliberately not a consumer-facing brand.

Buyer questions

What does Command A+ cost?

There is no public per-token price. Cohere sells it as a private or sovereign deployment and via managed Model Vault inference, so pricing is contract-based; you can also self-host the Apache-2.0 weights on your own hardware at no license cost.

Can I really run it on a single GPU?

Yes — Cohere targets one NVIDIA B200 (or two H100s) at 4-bit (W4A4), unusually light for a 218B-parameter model because the MoE activates only 25B parameters per token.

Is it actually open-weights?

Yes, the weights are on Hugging Face under Apache-2.0 (CohereLabs/command-a-plus-05-2026-bf16), with FP8 and W4A4 quantizations. Read the model-card terms before shipping, as with any release.

How capable is it versus the frontier?

Mid-pack on public benchmarks (Artificial Analysis Index 37, MMMU 75.1) with a standout on agentic tool use (tau-squared-Bench Telecom 85). It is not a frontier flagship; it is a governable, self-hostable enterprise model.

Does it handle images and multiple languages?

Yes — it accepts text and image input and covers 48 languages, up from 23 in Command A Reasoning, which fits multinational document and support workflows.

Who is this for?

Regulated enterprises — finance, healthcare, defense, government — that must keep inference inside their own perimeter and want a clean license, plus teams already running Cohere for RAG and tool use.

What is the context window?

128K tokens of input with up to 64K tokens of generation — comfortable for enterprise documents, though below the 1M-token windows some June rivals ship.

Comparable models

GLM-5.2Z.ai

The open-weights capability leader — far higher on the AA Index (51 vs 37) and cheaper to access, but a 753B multi-node model versus Command A+'s single-B200 footprint, and without the sovereign-deployment contract and support Cohere sells.

Mistral Medium 3.5Mistral

The closest European enterprise alternative — comparable governance posture and stronger raw benchmarks, but modified-MIT rather than clean Apache-2.0 and without Cohere's single-GPU sovereign-deployment focus.

Nemotron 3 UltraNVIDIA

The other June open-weights release — much stronger on coding and long context (SWE-bench Verified 70.7, 1M context) but text-only and multi-node to host, where Command A+ adds vision and runs on one B200.

Sources

Primary references used to verify this review.

Model specs

Input price
— / Mtok
Output price
— / Mtok
Cached input
Batch (in/out)
Context window
131K tokens
Max output
66K tokens
Knowledge cutoff
Undisclosed
Released
2026-05-19
Modalities
text, image → text
Output speed
Not profiled
License
Open weights (Apache-2.0)
Clouds
First-party API

Does not train on API inputs by default

Last verified 2026-07-02