by Cohere · Command A family · best for sovereign self-hosted enterprise agents
Command A+ is Cohere's enterprise wager: a 218B-parameter mixture-of-experts (25B active) released 2026-05-20 as open weights under a plain Apache-2.0 license, deliberately priced with no public per-token rate card because Cohere wants sovereign and private-deployment contracts, not metered API traffic. It runs on a single NVIDIA B200 (or two H100s) at 4-bit, accepts text and image input across 48 languages, and posts a standout agentic-tool-use result (tau-squared-Bench Telecom 85, up from 37 on its predecessor). It is not a frontier flagship — Artificial Analysis places it mid-pack at 37 — but for regulated buyers who must keep inference inside their own perimeter, a governable, self-hostable, cleanly-licensed model beats a benchmark point.
| Benchmark | Score | Source |
|---|---|---|
| MMMU | 75.1% | cohere.com 2026-05-20T00:00:00.000Z |
| Artificial Analysis Index | 37 | cohere.com 2026-05-20T00:00:00.000Z |
Six personas, six verdicts — the same panel that reviews every product on TopReviewed.
“The only frontier-adjacent model you can run entirely inside your own perimeter under Apache-2.0 - for regulated buyers that beats a benchmark point every time.”
Command A+ is a procurement instrument as much as a model. Cohere priced it for sovereign and private deployment rather than token traffic, which is exactly why there is no public rate card: the buyer Cohere wants signs a deployment contract and runs the Apache-2.0 weights on hardware it controls. For banks, defense, healthcare, and government that cannot send data to a US hyperscaler API, that is the whole ballgame, and the single-B200 footprint makes it operationally realistic. The capability is mid-pack on open benchmarks (AA Index 37), so this is not the pick when raw intelligence decides — it is the pick when data residency, auditability, and license cleanliness do.
“Cohere stopped chasing the frontier and planted a flag where the hyperscalers can't follow: your data center, your license, your data.”
The strategy is legible and differentiated. While Chinese labs commoditize open capability and OpenAI/Anthropic own the frontier API, Cohere is monetizing the one segment those players structurally cannot serve well — the regulated enterprise that will not, or legally cannot, run workloads on a public API. Apache-2.0 weights plus a single-GPU footprint plus 48-language coverage plus tool-use and RAG heritage is a coherent enterprise bundle. The weakness is reach: a no-public-price, sales-led motion scales slowly and cedes the developer mindshare that GLM and Qwen are winning. This is a durable niche, not a market-share play.
“There's no per-token line item because you buy the deployment, not the tokens - for high, steady volume inside your walls that can be the cheapest frontier-adjacent option going.”
The economics invert the usual API calculus. With Apache-2.0 weights and a single-B200 (or dual-H100) serving target at 4-bit, a heavy, steady internal workload amortizes fixed hardware cost far below what metered API tokens would cost at the same volume — and the spend is capex you control, not a variable bill that scales with success. The catch is that there is no published token price to benchmark against, so the model only pencils out if you have the volume and the ops maturity to run it; for spiky or low-volume use, a metered frontier API is cheaper and simpler. Value here is real but conditional.
“Tool use and RAG are where Cohere has always been strong, and tau-squared-Bench Telecom 85 shows it - just don't reach for it as your hardest-problem coder.”
For the enterprise-agent builder this is comfortable ground: first-class tool use, structured output, a big jump on agentic telecom tasks (85 vs the predecessor's 37), and 48-language coverage for multinational deployments. Image input and MMMU 75.1 cover document and screenshot workflows. The gaps are honest: Terminal-Bench Hard 25 says it is not a frontier coding agent, the 128K context trails same-month 1M rivals, and running the weights well is a real MLOps commitment. As a RAG-and-tools workhorse inside a regulated stack it fits; as a general coding brain it does not lead.
“It's built for enterprise deployments, not for me - there's no slick consumer app, and on a hard question the frontier models are simply sharper.”
Command A+ is not aimed at the individual daily driver, and it shows. There is no polished first-party consumer surface, access is through the API or a self-hosted deployment, and on the hardest analytical or coding questions it lands a clear tier below the frontier flagships (AA Index 37). What it does well for a hands-on user is multilingual work and grounded, tool-augmented answering where hallucination control matters. For most people wanting a chat companion, GLM-5.2 or a frontier API delivers more; Command A+ earns its keep inside an enterprise workflow, not on your personal desktop.
“Honest positioning, for once: Cohere never calls it a frontier model, and AA Index 37 confirms it isn't - the pitch is deployment, not benchmarks.”
There is refreshingly little to debunk here because Cohere does not overclaim. The headline improvements — tau-squared-Bench Telecom 85 up from 37, Terminal-Bench Hard 25 up from 3 — are large but measured against its own weak predecessor, so they signal progress, not leadership. Independent placement (AA Index 37) puts it below every open flagship in this month's cohort, and the "no public price" framing, while strategically sound, also conveniently avoids a value comparison it would lose on raw capability. The residual risks are ordinary: undisclosed knowledge cutoff, an Apache-2.0 headline that still warrants a read of the model-card conditions, and benchmarks reported first-party. Nothing dishonest; just a mid-tier model with an unusually smart go-to-market.
Regulated enterprises — finance, healthcare, defense, government — that must keep inference inside their own perimeter and want a clean Apache-2.0 license they can audit and self-host. Multilingual document and support automation across the model's 48 languages, and grounded RAG-plus-tools agent pipelines where Cohere's heritage shows. High, steady internal workloads where a single-B200 deployment amortizes below metered API costs. Not the pick when raw benchmark leadership, best-in-class coding, or a giant context window is the deciding factor — GLM-5.2 or a frontier API wins those.
The full research notes behind this review — verified against primary sources.
A 218B-parameter sparse mixture-of-experts activating 25B parameters per token, with 128 experts (8 routed plus 1 shared per token) and sliding-window attention interleaved roughly 3:1 with global attention layers — a design tuned for efficient long-document serving on a small GPU footprint. The 25B active path is what lets a 218B model run on a single B200 at 4-bit. Context is 131,072 tokens of input with up to 65,536 tokens of generation. Cohere does not publish a training-token count or a knowledge cutoff, and whether the vision path is native or adapter-based is undisclosed. The weights ship in BF16 with FP8 and W4A4 quantizations for lighter deployment.
Agentic tool use is the design center and the standout score (8.0): tau-squared-Bench Telecom 85 (up from the predecessor's 37) reflects Cohere's long enterprise-agent and RAG heritage, and function calling is first-class (8.5). Multilingual coverage (8.5) across 48 languages is a durable Cohere strength for multinational deployments. Vision (7.5) and document handling (7.0) rest on text-and-image input and MMMU 75.1, covering screenshot and document-image workflows. Reasoning (7.0) is available via the model's reasoning output mode and is benchmark-anchored by the mid-pack AA Index (37). Coding (6.5) is the honest weak spot — Terminal-Bench Hard 25 says this is not a frontier coding agent. Math is left unscored (no AIME/MATH disclosure). No built-in web search, so real-time data scores 0; grounding comes from Cohere's RAG tooling, not in-model retrieval.
| Benchmark | Score | Note | Source |
|---|---|---|---|
| tau-squared-Bench Telecom | 85% | agentic tool use; up from 37% on Command A Reasoning | Cohere blog |
| Terminal-Bench Hard | 25% | agentic coding; up from 3% | Cohere blog |
| MMMU | 75.1 | multimodal understanding | Cohere blog |
| MMMU Pro | 63 | harder multimodal split | Cohere blog |
| AA Intelligence Index | 37 | independent aggregate | Cohere blog |
MMMU and the Artificial Analysis Index enter this catalog's comparable columns as clean, standard metrics. The tau-squared-Bench Telecom and Terminal-Bench Hard figures are large improvements but are measured against Cohere's own weak predecessor and on harder/domain-specific splits that do not line up with the generic tau-bench / Terminal-Bench columns used elsewhere, so they are reported here in prose rather than in the comparable benchmark map. Research confidence is medium: benchmarks are first-party and pricing is undisclosed.
Cohere reports up to 63% higher output tokens/second and a 17% latency reduction versus Command A Reasoning, but publishes no absolute throughput or time-to-first-token figures, and no independent tracker has measured the model at verification time — so speed fields are null. The 25B active-parameter path and single-GPU serving target imply competitive interactive latency for the class; validate on your own hardware, since self-hosted performance depends entirely on your deployment.
| Surface | Cost | Notes |
|---|---|---|
| API (per token) | Not published | no public rate card |
| Model Vault | Contract | managed private inference |
| Sovereign / private deploy | Contract | customer-controlled infrastructure |
| Self-host (weights) | Free | Apache-2.0, your own hardware |
There is deliberately no per-token price. Cohere's motion is deployment contracts and self-hosting, not metered API traffic — which is the whole strategy for the regulated buyer who will not send data to a public endpoint. All pricing fields are null for that reason.
Open weights on Hugging Face (CohereLabs/command-a-plus-05-2026-bf16) under Apache-2.0, with BF16, FP8, and W4A4 checkpoints. The headline is footprint: one NVIDIA B200 or two H100s at W4A4 serve the full model, unusually light for 218B parameters thanks to the 25B active path. Managed inference is available through Cohere's Model Vault, and private/sovereign deployments run entirely inside customer infrastructure with customer-chosen data residency. There is no consumer app and no public token API in the usual sense — access is enterprise-contract or self-host.
Cohere's positioning is enterprise-governance-first: Apache-2.0 weights that customers run inside their own perimeter, no training on customer data by default, and SOC 2 / GDPR-aligned enterprise terms. There is no single named safety framework analogous to a frontier lab's RSP, and no formal safety-level label is published for this release; content moderation and refusal behavior are the deployer's to configure when self-hosting. The knowledge cutoff is undisclosed. For regulated buyers, the sovereign-deployment model is itself the governance story — data never leaves the perimeter.
Cohere's platform tooling (Python and TypeScript SDKs, the v2 Chat API, native tool-use and RAG primitives), managed inference via Model Vault, and self-serve Apache-2.0 weights on Hugging Face with standard serving engines. LangChain and LlamaIndex integrations carry over from the Command line. Cohere's own North enterprise platform is the flagship product built on it, and the developer community is enterprise-weighted rather than hobbyist. Popularity is growing within regulated verticals; it is deliberately not a consumer-facing brand.
There is no public per-token price. Cohere sells it as a private or sovereign deployment and via managed Model Vault inference, so pricing is contract-based; you can also self-host the Apache-2.0 weights on your own hardware at no license cost.
Yes — Cohere targets one NVIDIA B200 (or two H100s) at 4-bit (W4A4), unusually light for a 218B-parameter model because the MoE activates only 25B parameters per token.
Yes, the weights are on Hugging Face under Apache-2.0 (CohereLabs/command-a-plus-05-2026-bf16), with FP8 and W4A4 quantizations. Read the model-card terms before shipping, as with any release.
Mid-pack on public benchmarks (Artificial Analysis Index 37, MMMU 75.1) with a standout on agentic tool use (tau-squared-Bench Telecom 85). It is not a frontier flagship; it is a governable, self-hostable enterprise model.
Yes — it accepts text and image input and covers 48 languages, up from 23 in Command A Reasoning, which fits multinational document and support workflows.
Regulated enterprises — finance, healthcare, defense, government — that must keep inference inside their own perimeter and want a clean license, plus teams already running Cohere for RAG and tool use.
128K tokens of input with up to 64K tokens of generation — comfortable for enterprise documents, though below the 1M-token windows some June rivals ship.
The open-weights capability leader — far higher on the AA Index (51 vs 37) and cheaper to access, but a 753B multi-node model versus Command A+'s single-B200 footprint, and without the sovereign-deployment contract and support Cohere sells.
The closest European enterprise alternative — comparable governance posture and stronger raw benchmarks, but modified-MIT rather than clean Apache-2.0 and without Cohere's single-GPU sovereign-deployment focus.
The other June open-weights release — much stronger on coding and long context (SWE-bench Verified 70.7, 1M context) but text-only and multi-node to host, where Command A+ adds vision and runs on one B200.
Primary references used to verify this review.
Does not train on API inputs by default
Last verified 2026-07-02