by xAI · Grok 4 family · best for runtime reasoning toggle with long context and live X data
Grok 4.20 is xAI's reasoning model released to full GA on 2026-03-10 (beta 2026-02-17), and was the flagship between Grok 4 / Grok 4 Fast and the newer Grok 4.3. Its defining traits: a runtime reasoning toggle (exposed as separate grok-4.20-0309-reasoning and grok-4.20-0309-non-reasoning slugs plus a reasoning-effort parameter), a large context window, the best non-hallucination rate of any model at its release on Artificial Analysis's Omniscience benchmark, and xAI's live-X data access. The single sentence a buyer needs: it is a legacy-but-supported flagship whose niche today is a reasoning on/off switch plus long context with live data — most workloads are now better served by the cheaper, newer-trained Grok 4.3. Provider: xAI. Released: 2026-03-10. Status: GA. Context: 1M tokens (see note). Max output: undisclosed. Modalities: text + image in, text out. Knowledge cutoff: November 2024. Headline price: $1.25 / $2.50 per 1M tokens (repriced from launch $2 / $6).
-reasoning / -non-reasoning slugs and an effort parameter. One integration, two operating modes.| Benchmark | Score | Source |
|---|---|---|
| IFEval | 81% | Artificial Analysis (IFBench)2026-04-30T00:00:00.000Z |
| MATH-500 | 87.3% | xAI launch / secondary coverage2026-03-10T00:00:00.000Z |
| TAU-bench | 93% | Artificial Analysis (tau-2-Bench Telecom; ~5pts below 4.3's 98)2026-04-30T00:00:00.000Z |
| LMArena Elo | 1491 | LMArena / LMSYS (grok-4.20-beta1, Mar-Apr 2026, top-4; +31 May mover)2026-04-30T00:00:00.000Z |
| GPQA Diamond | 78.5% | xAI launch / secondary coverage2026-03-10T00:00:00.000Z |
| Artificial Analysis Index | 49 | artificialanalysis.ai 2026-05-28T00:00:00.000Z |
Six personas, six verdicts — the same panel that reviews every product on TopReviewed.
“A capable, supported legacy flagship — but if my workload fits 1M tokens, I migrate to the cheaper, newer Grok 4.3.”
Strategically, Grok 4.20 is now a continuity choice rather than a new bet. It retains real value where a runtime reasoning toggle, long context, and best-in-class non-hallucination matter together — but xAI itself recommends 4.3 as the default, signaling where investment and roadmap attention go. Vendor risk is the same as the family: thin disclosure, no certs, no published safety framework. Lock-in is low (SDK compatibility). The decision is simple: keep 4.20 only for the specific niches; otherwise plan migration. Roadmap confidence in 4.20 specifically is declining as 4.3 absorbs its use cases.
“Its one durable edge is the lowest hallucination rate at launch — but the market has moved to 4.3, and that's where the moat now lives.”
In positioning terms, Grok 4.20 carries the same structural moat as the family — live X data — but the differentiation has migrated to Grok 4.3, which is cheaper, fresher, and adds video. 4.20's distinct selling point is reliability: the best-in-class non-hallucination rate gives it a credible pitch to error-sensitive verticals. But market timing works against it: launched March, superseded by April's 4.3, repriced to match it. As a standalone competitive play it has limited runway; its role is to retain users until they migrate. Differentiation versus non-xAI rivals rests entirely on the X-data and reliability angles.
“Now repriced to match 4.3 — so there's no cost reason to choose 4.20 over the newer model unless you specifically need its niche.”
At launch, 4.20's $2 / $6 made it a clear loser to 4.3's $1.25 / $2.50. xAI has since repriced 4.20 down to $1.25 / $2.50 / $0.20 cached on the docs card, erasing the cost penalty — but that also erases any financial reason to prefer it, since 4.3 is the same price with newer training and a higher AA Index. The lingering finance risk is the source conflict: AA still bills it mentally at $2 / $6 with $1.10 cached, so anyone modeling from aggregators will misprice. Rule for finance: confirm the live docs.x.ai rate, then ask whether 4.3 wouldn't simply be the better-value choice at identical pricing.
“The reasoning toggle is the real win — ship one integration, flip a boolean between fast and deep, and parse cleaner output thanks to strict adherence.”
For builders, 4.20's standout is the runtime reasoning toggle: one integration covers both a fast conversational mode and a deep reasoning mode, switched by a parameter or slug rather than a model swap. Strict prompt adherence means less defensive parsing code — when you say "JSON only," you get JSON. Tool calling, structured outputs, and the live-X search tool all work cleanly. The 1M context simplifies long-context plumbing. Friction: SDK/docs polish trails OpenAI and Anthropic, reasoning visibility is summary-only, and coding is better done elsewhere. Reliability is decent; rate limits are spend-tiered.
“It was the standard Grok through April — looser, opinionated, live X data — but consumer surfaces have since moved everyone to 4.3.”
For everyday users on grok.com or X Premium, Grok 4.20 was the default Grok experience through April 2026: the looser personality, ready opinions, and live-X integration that define the brand. Its 2M/1M context is invisible to typical chat use. Once 4.3 became the consumer default, users migrated automatically, so most direct 4.20 use today is via teams pinned to a specific API slug for stability. As a daily driver it was fine; it is simply no longer the one most people touch.
“The 2M-context headline is already contradicted by xAI's own current docs (1M), and the pricing it launched at quietly evaporated — read the live card, not the launch post.”
Adversarially, Grok 4.20 is a case study in why xAI's thin disclosure matters. Its marquee launch claim — 2M context — now conflicts with xAI's own docs card showing 1M for the reasoning/non-reasoning slugs, while AA and OpenRouter still show 2M; nobody should cite a Grok context number without checking the live source. The launch pricing ($2 / $6) silently dropped to $1.25 / $2.50, so any cost analysis older than a few weeks is wrong. Architecture is undisclosed, no SWE-bench, no safety framework. The one claim that holds up well is the non-hallucination leadership (AA-Omniscience 78%) — that is independently sourced and genuinely strong.
grok-4.20-0309 slugs who don't need video input.The full research notes behind this review — verified against primary sources.
As with all of the Grok line, xAI discloses essentially nothing about internals — architecture type, parameters, attention, layers, training tokens, and compute are all undisclosed and marked null/unknown. What is documented is behavioral: a switchable reasoning mode (the headline architectural-level feature), a 1M-token context on the current docs card, native image input, and a November 2024 knowledge cutoff. The reasoning toggle is the genuinely distinctive design choice here — most reasoning models from peers are either always-on or a separate SKU, whereas Grok 4.20 exposes both modes behind one model family.
| Benchmark | Score | vs Predecessor | vs Top Competitor | Source |
|---|---|---|---|---|
| Artificial Analysis Intelligence Index | 49 | Up from Grok 4 (~42) | Below GPT-5.5 (~60), Gemini 3.1 Pro (~57); below Grok 4.3 (53) | Artificial Analysis |
| GPQA Diamond | ~78.5% | Up from Grok 4 ceiling | Behind Claude Opus 4.7 / GPT-5.5 | Launch coverage |
| MATH-500 | ~87.3% | Improved | Competitive | Launch coverage |
| AA-Omniscience (non-hallucination) | 78% | Best in class at release | Best in class at release; still leads vs Grok 4.3 | Digital Applied |
| GDPval-AA (tool-use Elo) | 1179 | Baseline | Mid-pack (Grok 4.3 reaches 1500) | AA launch article |
| IFBench | 81% | n/a | Competitive | Artificial Analysis |
| LMArena Elo (beta1 proxy) | ~1491 | Up (+31 May mover, top-4) | Behind Claude Opus 4.6 (~1504) | LMArena tracker |
(MMLU-Pro, AIME 2025, SWE-bench Verified, HumanEval, LiveCodeBench, Aider Polyglot, and MMMU were not consistently published by xAI for Grok 4.20; rows left null. Coverage is thinner than peer models.)
Artificial Analysis measures ~171.4 output tokens/sec with a ~13.24s time-to-first-token — notably faster to first token than Grok 4.3's ~19.7s, because in non-maximal reasoning modes there is less upfront thinking. In non-reasoning mode it behaves as a fast conversational model. Overall latency tier: medium. The reasoning toggle is the practical lever — flip it off for speed, on for quality, without swapping models.
| Surface | Cost | Notes |
|---|---|---|
| API input | $1.25 / 1M tok | docs.x.ai canonical (repriced from $2.00) |
| API output | $2.50 / 1M tok | docs.x.ai canonical (was $6.00) |
| Cached input | $0.20 / 1M tok | docs.x.ai (AA's older card shows $1.10) |
| Batch | n/a | No documented batch endpoint |
| Direct UI (SuperGrok) | $30 / mo | Standalone Grok |
| Direct UI (X Premium / Premium+) | $8 / $40 mo | Bundled in X app |
| SuperGrok Heavy | $300 / mo | Power-user tier |
| Free tier | $0 | grok.com + free X with daily caps |
| Rate limits | tiered by spend/plan | Per docs.x.ai |
Pricing conflict (the v1-flagged discrepancy, now characterized): This is the model where xAI docs and third parties diverge most. xAI's docs.x.ai card is canonical at $1.25 / $2.50 / $0.20 cached. Artificial Analysis's Grok 4.20 page still shows the launch-era $2.00 / $6.00 with $1.10 cached and a 2M context; OpenRouter shows $1.25 / $2.50 but 2M context. The v1 file used the AA/launch $2 / $6 figures — corrected here to the canonical docs values, with the conflict noted rather than hidden.
Proprietary, API-only, no open weights, not self-hostable. OpenAI-SDK and Anthropic-SDK compatible. Resold via OpenRouter. Unlike Grok 4.3, Grok 4.20 is not confirmed as a distinct Azure AI Foundry SKU (Foundry leads with Grok 4 Fast variants and Grok 4.3), so cloud_platforms is empty here. Rate limits are spend/plan-tiered.
Same posture as the rest of the Grok line: no published safety framework, governance via Acceptable Use Policy, content moderation tightened after January 2026. Training-on-inputs: API only via irreversible opt-in data sharing; X consumer surface trains by default with no opt-out (trains_on_inputs: true, data_optout_available: false). The distinctive governance positive for 4.20 is its release-time best-in-class non-hallucination rate (AA-Omniscience 78%), which it still leads — relevant for low-error-tolerance summarization (legal, medical) where factual reliability matters more than the latest knowledge. No verified SOC2/HIPAA/ISO/GDPR certs on the direct API.
Python/TypeScript SDKs plus OpenAI- and Anthropic-SDK compatibility; LangChain and Vercel AI SDK integrations; resold on OpenRouter. Surfaces: Grok on X, grok.com, SuperGrok. Popularity is growing but actively draining toward Grok 4.3 as the recommended default.
xAI docs list $1.25 / $2.50 / $0.20 cached — repriced down from the launch $2 / $6. Artificial Analysis still shows the old $2 / $6 with $1.10 cached; trust docs.x.ai.
xAI's docs card lists 1M for the reasoning/non-reasoning slugs; AA and OpenRouter still show 2M, and the multi-agent sibling is 2M. Verify against your account's live limits.
Either pick the -reasoning vs -non-reasoning slug, or set the reasoning-effort parameter — one integration, two modes.
For almost everything, 4.3: same price, newer training, video, higher scores. Keep 4.20 only for its non-hallucination edge or if you're pinned for stability.
API: only via irreversible opt-in data sharing. X consumer surface: by default, no opt-out.
No SOC2/HIPAA/ISO certs are publicly verified on the direct API; route via a managed cloud if you need them.
xAI's own newer flagship: same price, newer cutoff (Dec 2025), adds video, higher AA Index (53 vs 49), much higher agentic Elo (1500 vs 1179); 4.20 only wins on non-hallucination rate and (per AA/OpenRouter) raw context size.
Larger verified context with stronger multimodal breadth and higher AA Index; loses on live-X access.
Better hard coding and reasoning ceiling and published safety; pricier and no real-time data.
Primary references used to verify this review.
Last verified 2026-05-27