Llama 4 Scout

GALatest Scout

by Meta · Llama 4 family · best for single-GPU long-context open-weights deploy

Open-WeightsMultimodalCost-OptimizedLong-ContextEdge / On-Device
7.5
AI Panel Score
Value 9.5/10

Llama 4 Scout is the small, deployable member of Meta's Llama 4 herd, released April 5, 2025. It is a 109B-total / 17B-active Mixture-of-Experts model (16 experts), natively multimodal, with a headline 10,000,000-token context window — the largest of any openly available model at release. The one-sentence buyer takeaway: it is the only model in 2026 that combines a 10M context, native vision, and single-GPU deployability, making it the obvious open-weights pick when context size and on-prem economics matter more than peak intelligence.

What's new

  • 10M-token context window — the largest of any openly available model at release.
  • Fits on a single H100-class GPU with INT4 quantization, making it the most deployable Llama 4 variant.
  • First Llama small/mid-tier to ship as MoE (16 experts) rather than dense, sharing the same 17B-active speed profile as Maverick.
  • Natively multimodal via the same early-fusion vision tower architecture as Maverick.

Benchmarks

BenchmarkScoreSource
MMLU79.6%Meta / llm-stats aggregator2025-04-05T00:00:00.000Z
MMMU69.4%Meta Llama 4 model card2025-04-05T00:00:00.000Z
MATH-50050.3%Meta (MATH-Hard)2025-04-05T00:00:00.000Z
MMLU-Pro74.3%Meta Llama 4 model card2025-04-05T00:00:00.000Z
HumanEval82%llm-stats aggregator (approx)2025-04-05T00:00:00.000Z
GPQA Diamond57.2%Meta Llama 4 model card2025-04-05T00:00:00.000Z
LiveCodeBench32.8%community aggregator2025-04-10T00:00:00.000Z
Artificial Analysis Index14Artificial Analysis2026-05

AI Panel Review

Six personas, six verdicts — the same panel that reviews every product on TopReviewed.

Decision Maker8.5/10
The easiest 'go open weights' call I can make: one H100, 10M context, every cloud carries it. Just test the long-context cliff before you bet on it.

Scout is the lowest-friction open-weights adoption in 2026. The deployment story is genuinely simple — pull from Hugging Face, quantize to INT4, run on a single GPU — and every major cloud and inference provider carries it, so vendor risk is minimal. For organizations that need on-prem data control without frontier capability, it is close to ideal. The strategic optionality is high and the risk surface small. The one decision-maker-level caveat is the long-context quality cliff: the 10M number is real for retrieval but not for deep reasoning, so do not architect around 10M of usable comprehension without testing your workload.

Strategic Fit 9Vendor Risk 9Roadmap Confidence 6
Pros
  • trivially deployable
  • multi-cloud
  • sovereign
  • 10M context
Cons
  • long-context cliff
  • uncertain Llama roadmap
Right for: on-prem/sovereign teams needing big context cheaply
Avoid if: you need frontier reasoning or guaranteed comprehension at extreme context
Domain Strategist7.5/10
Scout owns 'biggest context that fits on one GPU.' That's a defensible, specific square — even if rivals are closing in on quality.

Positioning, Scout's wedge is the unique intersection of 10M context, native vision, and single-GPU deploy — no other open model in 2026 offers all three. Against closed long-context models (Gemini Flash) it is the open-weights answer; against other open small models (Qwen 3 30B-A3B, Mistral Small 3) it wins on context and multimodality, loses on some reasoning. Market timing rides the same sovereignty/cost tailwinds as Maverick. The durability risk is real — competitors are catching up on context, and the comprehension cliff undercuts the headline — but the deployability story keeps it relevant.

Competitive Positioning 8Differentiation 8Market Timing 7
Pros
  • unique context+vision+single-GPU combo
Cons
  • comprehension cliff dents the headline
  • rivals closing
Right for: teams who genuinely need huge context cheaply
Avoid if: you only need 128K — cheaper dense models suffice
Finance Lead9/10
Cheapest serious open-weights model on the market, and the 10M context can delete an entire RAG pipeline's cost. The math is unambiguous at scale.

Scout is the strongest pure unit-economics story in the Meta lineup. DeepInfra runs it at $0.08/$0.30; self-hosted on a single rented H100 at $2–3/hour, it beats any closed API by 20–100x at volume. The hidden lever is the 10M context: workloads that previously required a vector DB, embedding compute, and a re-ranker can sometimes collapse into a single Scout call, removing whole line items from the bill — though the prefill cost of a truly huge context must be modeled (it is not free). Above ~100M tokens/month, self-hosted Scout dominates on $/Mtok.

Cost Efficiency 10Pricing Transparency 8Value per Dollar 10
Pros
  • cheapest serious open multimodal
  • single-GPU
  • context can replace RAG infra
Cons
  • huge-context prefill cost is real
  • needs utilization to amortize self-host
Right for: high-volume, long-context, cost-sensitive workloads
Avoid if: low volume where managed simplicity wins
Domain Practitioner7.5/10
The smallest Llama 4 with the biggest party trick. Fine-tune on one box, serve on one GPU — just keep evals for the long-context cliff.

Builders love Scout because it is the most accessible Llama 4: fine-tuning fits on a single 8xA100 box, inference on one H100, and the 10M context lets you skip RAG indexing for small-to-mid corpora and just dump everything in. Native support across Transformers, vLLM, llama.cpp, Ollama, SGLang, and MLX makes local iteration fast. The catches are familiar: provider chat-template inconsistencies, the long-context quality cliff (write evals, do not trust the 10M number blindly), and tool-use that trails Maverick on multi-step loops. Genuinely fun and forgiving to build with.

API Ergonomics 8Tool/Agent Support 7Reliability 8
Pros
  • single-GPU serve
  • single-box fine-tune
  • huge context simplifies RAG
Cons
  • long-context cliff
  • provider template drift
  • weaker multi-step tools
Right for: builders who own their stack and want big context cheaply
Avoid if: you need rock-solid agent chains out of the box
Power User6.5/10
It quietly disappears into the background — fast, polite, multilingual. Not the model you pick when the chatbot itself is the show.

For consumer-facing chat Scout is the model that gets out of the way: good latency (sub-second on Groq), sensible refusal rates, clean casual conversation. It is not a flagship-personality model — Claude, GPT-5, or even Maverick feel smarter in extended use — but for embedded SaaS chat, support assistants, and in-app coaches where the model is one feature among many, users get a reliable, polite, multilingual helper. As with all Llama 4, long context-heavy sessions surface the comprehension degradation.

Output Quality 6Speed 8Everyday Usefulness 7
Pros
  • fast
  • polite
  • multilingual
  • reliable backstage
Cons
  • no personality wow
  • long-context comprehension fades
Right for: embedded assistants
Avoid if: the chatbot is the product
Skeptic5.5/10
Ten million tokens is the number on the box. Fiction.LiveBench puts Llama 4 near the bottom on actually understanding long text — buy the GPU story, not the 10M.

Adversarially, Scout's headline is its biggest overclaim. The 10M context is real for needle retrieval and demos, but independent comprehension evals (Fiction.LiveBench, RULER-style) rank Llama 4 near the bottom — so "10M context" oversells usable reasoning by a wide margin. Same family caveats apply: no reasoning mode, mid-pack coding and math, a 2024 cutoff, and a teacher model (Behemoth) that never shipped. The defensible, honest pitch is "the cheapest open multimodal model that fits on one GPU" — a genuinely good thing. The "10M context champion" framing should be treated as marketing until you have tested comprehension on your own data.

Claim Accuracy 5Weakness Severity 6Hype vs Reality 5
Pros
  • single-GPU and cheap is genuinely true
  • retrieval works
Cons
  • 10M comprehension oversold
  • mid-pack reasoning
  • Behemoth unshipped
Right for: skeptics who value it as a cheap single-GPU model
Avoid if: you believe the 10M headline at face value

Strengths

  • 10M-token context window — unmatched in open weights at release.
  • Single-GPU deployable (one H100 at INT4) — runs on a ~$30K box or a rented GPU.
  • Native vision; strong DocVQA (91.6) and ChartQA (85.3) for the size.
  • Cheapest serious open-weights multimodal model on most providers (~$0.08 in).
  • Day-zero availability on Groq, Together, Fireworks, DeepInfra, Bedrock, Vertex.

Limitations

  • Long-context comprehension degrades well before 10M tokens (Fiction.LiveBench / RULER-style evals); the headline is retrieval capacity, not usable reasoning depth.
  • Trails Maverick by ~6 points on MMLU-Pro and ~12 on GPQA Diamond — not a frontier reasoner.
  • No native reasoning mode; loses to DeepSeek R1, o-series, Claude extended thinking on hard math.
  • 8K output ceiling on most managed providers; filling the 10M context spikes prefill cost and latency.

Best use cases

Long-document RAG over legal contracts, technical specs, and multi-PDF research where dumping the corpus into one prompt beats building a retrieval pipeline. Whole-codebase analysis where the full repo fits in context. Single-GPU self-hosted chatbots and assistants where Maverick's 8xH100 node is overkill. Edge and sovereign deployments — a quantized Scout runs on one modern workstation GPU, keeping data fully on-prem.

Deep dive

The full research notes behind this review — verified against primary sources.

Architecture

Scout is a sparse MoE transformer: 109B total parameters across 16 routed experts plus a shared expert, ~17B active per token. It uses the iRoPE attention scheme — chunked RoPE local attention (8K chunks) in three of four layers and a full-context NoPE (no positional embedding) layer every fourth, which is what makes the 10M window mechanically possible. Scout adds QK-normalization (RMS norm of query/key states, no learnable params) and temperature-scaled softmax in NoPE layers to preserve attention over very long sequences. Vision is early-fused. Meta discloses total/active params, expert count, the iRoPE design, and a training-token scale of up to ~40T; exact layer counts, compute, and full data recipe are not published. The released checkpoint is Instruct-tuned and is the only Llama 4 variant designed to fit a single server-grade GPU.

Capabilities

Scout is a competent generalist with one standout trick — context length. Text reasoning (cap_reasoning 5.5) is roughly on par with Llama 3.1 70B: MMLU 79.6, MMLU-Pro 74.3, GPQA Diamond 57.2 — solid for its tier but not frontier. Math (5.0) and coding (5.5) are mid-pack open-weights; HumanEval ~82, LiveCodeBench ~33. Multilingual (7.5) is a genuine strength across 200 languages. Vision and document OCR (6.5 / 7.0) are strong for the size, with ChartQA 85.3, DocVQA 91.6, and MMMU 69.4. The long-context score (4.5) is the honest tension at the heart of this model: needle-in-haystack retrieval over the 10M window looks near-perfect, but independent comprehension benchmarks (Fiction.LiveBench, RULER-style) show meaningful quality degradation long before the ceiling — the 10M number is a retrieval capacity, not a reasoning-over-10M guarantee. Function-calling and instruction-following (6.0 each) are reliable for single steps, weaker on long agent chains. No reasoning mode, no real-time data (0.0).

Benchmark analysis

Benchmark Score vs Predecessor (3.3 70B) vs Top Competitor Source
MMLU 79.6 -6.4 (3.3 70B 86.0) trails Maverick (85.5) llm-stats
MMLU-Pro 74.3 +5.4 (3.3 70B 68.9) trails Maverick (80.5) Meta
GPQA Diamond 57.2 +6.7 (3.3 70B 50.5) competitive at tier Meta
MATH (Hard) 50.3 comparable trails Maverick (61.2) Meta
HumanEval ~82 ~ 3.3 70B (88.4) competitive llm-stats
MMMU (vision) 69.4 new (3.3 70B no vision) strong for size Meta
ChartQA 85.3 new competitive Meta
DocVQA 91.6 new strong Meta
Artificial Analysis Index 14 = (3.3 70B 14) above open non-reasoning median (13) AA

Speed & latency

Median output speed is ~106 tokens/sec across providers, with time-to-first-token around 0.56s on DeepInfra and 0.72s on Google Vertex. Because only ~17B params are active, it generates at small-model speed despite the 109B pool. Groq is the throughput leader at ~449 tokens/sec, giving a snappy sub-second interactive feel. Latency tier is fast; the practical caveat is that filling the 10M context dramatically raises prefill latency and cost, so the long-context superpower is best used selectively.

Pricing analysis

Surface Cost Notes
API input (representative) ~$0.08–$0.11 / 1M tok DeepInfra / Groq floor
API output (representative) ~$0.30–$0.34 / 1M tok
DeepInfra $0.08 in / $0.30 out cheapest mainstream
Groq $0.11 in / $0.34 out fastest (449 tps)
Fireworks $0.15 in / $0.60 out
Together ~$0.18 in / ~$0.59 out
Amazon Bedrock ~$0.22 blended on-demand
Google Vertex AI available 0.72s TTFT
Self-hosted 1x H100 80GB (INT4) ~55–60GB VRAM to hold 109B at INT4
Rate limits provider-specific often 1000+ RPM on managed tiers

Open weights mean no single Meta price; the figures above are the May 2026 inference market. Scout is the cheapest serious open-weights multimodal model on most providers.

Deployment & access

Open weights under the Llama 4 Community License. Download from Hugging Face (meta-llama/Llama-4-Scout-17B-16E-Instruct). The headline deployment property: it fits a single server-grade GPU. At INT4 the full 109B parameter set needs roughly 55–60GB, so one H100 80GB serves it comfortably; aggressive GGUF quants (Q4) run on 24–48GB consumer cards for hobbyist/edge use, though storing all 109B params is what sets the floor. Managed availability spans AWS Bedrock, Google Vertex AI, Azure AI Foundry, OCI, and IBM watsonx. Inference providers include Together, Fireworks, Groq, DeepInfra, OpenRouter, and Novita. Self-host economics are the strongest selling point — a single rented H100 at $2–3/hour serves millions of tokens/day. The Llama 4 Community License permits commercial use but requires a separate Meta license above 700M MAU and forbids training non-Llama models on outputs.

Safety & privacy

Identical posture to Maverick: the weights carry no built-in moderation, and Meta offers Llama Guard 4 (12B multimodal) plus Prompt Guard 2 (22M/86M) as optional pre/post filters for the Llama 4 line. "Trains on inputs" is not applicable when self-hosted; managed-provider terms vary; Meta's own terms do not train on your data. No model-level compliance certifications — those attach to your host or infrastructure. Refusal calibration is moderate and tunable, which is the intended advantage for regulated and sovereign deployments. Governance under Meta's Frontier AI Framework.

Ecosystem & tooling

Native support across Hugging Face Transformers, vLLM, llama.cpp, Ollama, SGLang, MLX, plus LangChain and LlamaIndex. Available on Bedrock, Vertex AI, Azure AI Foundry, OCI, and IBM watsonx, and on Together, Fireworks, Groq, DeepInfra, OpenRouter, and Novita. Used inside Meta's consumer AI surfaces. Popularity is mainstream — the go-to open-weights pick when single-GPU deploy or very large context is the requirement.

Buyer questions

What does Scout cost?

No single Meta price; representative inference is ~$0.08–$0.11 input and ~$0.30–$0.34 output per 1M tokens (DeepInfra/Groq cheapest). Self-hosting on one rented H100 runs $2–3/hour.

Can it really run on one GPU?

Yes — INT4 fits the 109B params in ~55–60GB, so a single H100 80GB serves it; aggressive GGUF quants run on 24–48GB cards for light use.

Is the 10M context usable?

For retrieval, largely; for reasoning across the full window, no — comprehension degrades well before 10M. Chunk and test on your data.

Does it do vision?

Yes, natively (early fusion) with strong DocVQA/ChartQA; it does not generate images.

How does it compare to Maverick?

Scout is smaller, cheaper, single-GPU, and has a bigger context; Maverick is smarter with 128 experts but needs a full node. Both run at ~17B-active speed.

What about safety/compliance?

No built-in moderation; add Llama Guard 4 / Prompt Guard 2. Certifications come from your host/infra, not the model.

Any license limits?

Commercial use allowed; separate Meta license required above 700M MAU; cannot train non-Llama models on its outputs.

Comparable models

Llama 4 Maverick — same family and 17B-active speed; smarter (more experts) but needs an 8xH100 node and offers a smaller 1M context. Scout wins on single-GPU deploy and context length.
Llama 3.3 70B — dense, text-only, 128K context; similar text quality but no vision and a far smaller window. Scout is usually the upgrade for new builds.
Qwen 3 30B-A3B / Mistral Small 3 — open small models competitive on text at a similar price; Scout wins on context length and native vision, may lose on specific reasoning benchmarks.

Sources

Primary references used to verify this review.

Model specs

Input price
$0.10 / Mtok
Output price
$0.34 / Mtok
Cached input
Batch (in/out)
Context window
10M tokens
Max output
8K tokens
Knowledge cutoff
2024-08
Released
2025-04-04
Modalities
text, image → text
Output speed
~106.1 tok/s
License
Open weights (Llama-4-Community)
Clouds
Bedrock, Vertex AI, Azure AI Foundry, GCP, OCI, IBM watsonx

Does not train on API inputs by default

Last verified 2026-05-27