Muse Spark

GALatest Spark

by Meta · Muse family · best for frontier reasoning for consumers, not yet for builders

FrontierReasoningMultimodal
7.3
AI Panel Score
Value 6.0/10

Muse Spark is the model that ended the Llama era: Meta Superintelligence Labs' first frontier release (2026-04-08) is proprietary, closed-weights, and consumer-first — free to billions on meta.ai, the Meta AI app, and (since June 23) Ray-Ban Meta glasses, with the API stuck in private preview and no rate card. A natively multimodal reasoning model with visual chain of thought and a parallel-multi-agent "Contemplating" mode that scores 58% on Humanity's Last Exam, it reportedly reaches Llama 4 Maverick parity with over an order of magnitude less compute. Catalog it for what it signals as much as what it scores: Meta now competes at the frontier on its own terms, and builders are not invited yet.

What's new

  • First model of the Muse family from Meta Superintelligence Labs — a ground-up overhaul that supersedes the Llama line for frontier work
  • Meta's first closed-weights frontier release: no downloadable weights, breaking a four-generation open tradition
  • Contemplating mode: multiple agents reasoning in parallel, Meta's answer to competitors' deep-reasoning modes (HLE 58%, FrontierScience 38% in this mode)
  • Natively multimodal with visual chain of thought — text, image, and speech input — plus tool use and multi-agent orchestration support
  • Free at consumer scale from day one; shipped on-glasses with Ray-Ban Meta on 2026-06-23; API remains private preview with no public pricing

Benchmarks

BenchmarkScoreSource
Humanity's Last Exam58%ai.meta.com 2026-04-08T00:00:00.000Z
Artificial Analysis Index43artificialanalysis.ai 2026-07-03T00:00:00.000Z

AI Panel Review

Six personas, six verdicts — the same panel that reviews every product on TopReviewed.

Decision Maker6.5/10
You can't buy it, license it, or build on it — Muse Spark is a strategy update from Meta, not a procurement option.

There is no purchasing decision to make, which is itself the decision-relevant fact. If your organization standardized on Llama for self-hosted AI, Muse Spark tells you Meta's frontier investment has moved somewhere you cannot follow: Llama 4 is now the effective end of the open line, Behemoth never shipped weights, and "Llama 5 does not exist." Plan open-weights roadmaps around GLM, Qwen, and Nemotron instead, and track the Muse API preview for when it opens. The private-preview structure suggests Meta is testing enterprise appetite; until a rate card and terms exist, any dependency here is imaginary.

Strategic Fit 5.5Vendor Risk 7Roadmap Confidence 7
Pros
  • Clear signal for roadmap planning
  • safety documentation precedes API
Cons
  • Zero procurement path
  • Llama succession uncertainty is now real
Right for: strategy teams recalibrating open-model roadmaps
Avoid if: you need anything actionable this quarter
Domain Strategist8.5/10
The most strategically legible release of 2026: Meta stopped subsidizing the ecosystem and started competing for the frontier itself.

Muse Spark reorders the market map. Meta spent four generations commoditizing its rivals' complement with open weights; MSL's first act reverses the strategy — closed weights, consumer-first distribution through a three-billion-user funnel, and hardware differentiation via Ray-Ban glasses that no competitor can match. The compute-efficiency claim (Maverick parity at a tenth-plus less compute) is aimed at investors as much as researchers. The vacuum it creates in open weights is already being filled by Chinese labs and NVIDIA, which changes the geopolitics of the open ecosystem. Distribution is the moat here, not benchmarks — and it's a very good moat.

Competitive Positioning 8.5Differentiation 9Market Timing 8
Pros
  • Unmatched consumer distribution
  • hardware wedge
  • coherent strategic pivot
Cons
  • Ceded open-ecosystem leadership
  • builder goodwill spent
Right for: anyone modeling where frontier AI competition goes next
Avoid if: you mistake consumer reach for platform availability
Finance Lead6.5/10
Free is not a price, it's an absence of one — there is no TCO to model because there is nothing to buy.

For budget purposes this row is empty: no API pricing, no enterprise tier, no committed-use discounts, nothing to negotiate. The consumer-free positioning is financed by Meta's ad business, not by you, and it tells you nothing about what API pricing will look like if it ships. The finance-relevant angle is indirect: if you built cost models around free Llama weights continuing to improve, that assumption just expired — self-hosted TCO projections should now assume the open Meta line is frozen at Llama 4, with the open-weights alternatives (GLM-5.2, Nemotron) as the new baseline curves. Value score reflects consumer value (high) against enterprise value (currently zero).

Cost Efficiency 6Pricing Transparency 5Value per Dollar 7
Pros
  • Genuinely free at consumer scale
  • no spend to justify
Cons
  • No pricing signal for planning
  • invalidates Llama-based TCO assumptions
Right for: flagging roadmap risk in self-host cost models
Avoid if: you were hoping to budget for it — you can't
Domain Practitioner5.5/10
Tool use, orchestration, visual chain of thought — all the features I'd love to build on, none of them behind an endpoint I can call.

From a builder's chair this is a shop window, not a shop. The capabilities Meta describes — parallel multi-agent Contemplating, visual chain of thought, native tool use — are exactly the primitives agent developers want, and there is no API, SDK, structured-output spec, or rate limit documentation to evaluate any of it. The private preview has no public entry path. What can be assessed: the consumer product is polished, multimodal interactions feel genuinely native rather than stitched, and the reasoning mode produces visibly deliberate answers. Practitioners should note the pattern from wavespeed and others: every "Muse Spark API" tutorial in circulation is speculation. Watch the preview; build on something else.

API Ergonomics 3Tool/Agent Support 6Reliability 7.5
Pros
  • Impressive native multimodality where you can touch it
  • strong reasoning ceiling
Cons
  • No API, docs, or access path
  • everything else is hearsay
Right for: keeping a watchlist entry warm
Avoid if: you have a deadline
Power User8.5/10
The best free AI on the planet right now — frontier reasoning in my pocket and on my face, for exactly zero dollars.

For daily consumer use, Muse Spark is a step change: frontier-tier reasoning, image understanding that genuinely reasons over what it sees, and voice interaction, free, inside apps billions already have. Contemplating mode on a hard question produces answers competitive with paid frontier subscriptions. The Ray-Ban integration is the sleeper feature — hands-free multimodal AI that works — and the June glasses launch made it the first frontier model people wear. Frictions: no power-user controls (no system prompts, no API escape hatch, no conversation export to speak of), Meta-grade content policies, and the knowledge that your usage feeds Meta's flywheel. As a free product it's extraordinary; as a power tool it's a walled garden.

Output Quality 8.5Speed 8Everyday Usefulness 8.5
Pros
  • Frontier quality at zero cost
  • true multimodal + voice
  • glasses integration
Cons
  • No customization or export
  • consumer terms
  • walled garden
Right for: anyone who wants maximum free AI
Avoid if: you need control, privacy terms, or an API
Skeptic7/10
Two cherry-picked numbers in best-case mode, everything else undisclosed — and the everyday index says the ceiling isn't the product.

The claims are carefully constructed. HLE 58% and FrontierScience 38% are quoted exclusively in Contemplating mode — the compute-maximal configuration — while the independently measured everyday index (AA 43) lands mid-pack among 2026 flagships, below GLM-5.2's 51. The order-of-magnitude efficiency claim is unverifiable without disclosed compute budgets. Architecture, parameters, cutoff: all withheld, from the company that used to publish everything. The GPQA/MMMU-Pro numbers circulating in coverage trace to no Meta source. What survives scrutiny: the HLE result is on Meta's letterhead, the safety report is real and substantive, and the consumer product demonstrably works at scale. Muse Spark is real and good; the framing wants you to believe it is the frontier leader, and the verifiable record says "frontier-adjacent, brilliantly distributed."

Claim Accuracy 7Weakness Severity 6.5Hype vs Reality 6.5
Pros
  • Headline numbers are first-party and specific
  • safety documentation is unusually complete
Cons
  • Best-case-mode benchmarking
  • total architecture opacity
  • unverifiable efficiency claim
Right for: calibrating the gap between Meta's story and the measured product
Avoid if: you'd take launch-post numbers as everyday performance

Strengths

  • Frontier-tier reasoning ceiling: HLE 58% in Contemplating mode, with a genuinely novel parallel-multi-agent approach
  • Native multimodality with visual chain of thought and speech input, deployed to billions for free
  • Claimed order-of-magnitude compute-efficiency gain over Llama 4 Maverick — if accurate, the most important engineering result of the release
  • Published safety report and framework evaluation ahead of any API exposure

Limitations

  • No API, no pricing, no weights: builders cannot use it at any price today
  • Capability disclosure is two benchmark numbers in maximum-reasoning mode; everyday-mode performance (AA 43) is a visible step down from the frontier ceiling
  • Architecture, parameters, context ceiling, and knowledge cutoff all undisclosed
  • Ends Meta's open-weights tradition — the strategic anchor that made Llama the default self-host choice — with no stated succession plan for open frontier releases

Best use cases

For consumers: free access to a frontier-class reasoning model with vision and voice, on the phone and on glasses — the strongest no-cost AI product on the market. For everyone else, the use case is strategic awareness: competitive analysts tracking Meta's pivot, enterprises reassessing Llama roadmap risk now that Meta's best work is closed, and builders deciding whether to wait for the API or commit elsewhere. Until general API access ships, Muse Spark cannot be a component in anyone's stack.

Deep dive

The full research notes behind this review — verified against primary sources.

Architecture

Undisclosed — a sentence that has never before applied to a Meta frontier model. No parameter counts, no architecture type, no attention details, no training scale; the safety report (arxiv 2606.12429) is the only technical artifact. What Meta states: natively multimodal reasoning with visual chain of thought, tool use, and multi-agent orchestration, achieving Llama 4 Maverick parity "with over an order of magnitude less compute." Independent trackers measure a 262K-token context window. Knowledge cutoff, output ceiling, and tokenizer are all unpublished. Ignore the 1M-context figure circulating from one aggregator — it does not match Artificial Analysis or any Meta statement.

Capabilities

Reasoning (8.5) is the demonstrated headline: 58% on Humanity's Last Exam in Contemplating mode is a frontier-tier result, and the parallel-multi-agent approach is architecturally distinctive. Vision (8.5) rests on native multimodality and visual chain of thought — the model reasons over images rather than captioning them — with speech input rounding out the everyday surface. Long context (7.0) reflects the measured 262K window, mid-pack for 2026. Agentic capability (6.5) is scored down not for the model but for access: tool use and orchestration exist, but without a public API nobody outside Meta can build agents on it. Coding, math, multilingual, OCR, creative, instruction-following, and function-calling are left unscored — Meta has published nothing, and no independent evaluation surface exists without an API. That null column is itself the honest description of Muse Spark today.

Benchmark analysis

Benchmark Score Note Source
Humanity's Last Exam 58% Contemplating mode Meta announcement
FrontierScience Research 38% Contemplating mode Meta announcement
AA Intelligence Index 43 measured without Contemplating-style depth Artificial Analysis

Meta's disclosure is thin: two headline numbers in the launch post, both in the maximum-reasoning mode. Secondary sources circulate GPQA Diamond and MMMU-Pro figures that do not appear in any primary Meta material and are excluded here. The gap between HLE 58 (Contemplating) and AA 43 (standard operation) is worth understanding — the model's ceiling and its everyday behavior are different products.

Speed & latency

Unmeasured — with no public API, independent trackers cannot benchmark throughput or first-token latency, and Meta publishes none. Anecdotally, standard responses in the consumer apps are chat-snappy while Contemplating mode visibly takes its time (parallel agents are compute-hungry by design). On-glasses inference via Ray-Ban Meta is server-routed, not local.

Pricing analysis

Surface Cost Notes
meta.ai / Meta AI app Free consumer, at billions-scale
Ray-Ban Meta glasses Free with device since 2026-06-23
API Private preview no rate card, select partners only
Weights Not available first closed Meta frontier model

There is no way to pay Meta for this model today, and no way to build on it without an invitation. All pricing fields in this row are null for that reason — not unknown, nonexistent.

Deployment & access

Consumer surfaces only: meta.ai on web, the Meta AI app, WhatsApp/Instagram/Messenger integration points, and Ray-Ban Meta glasses. No public API, no cloud-platform distribution, no weights, no self-hosting. The private API preview has select users; wavespeed.ai and others document the absence of any general access path as of mid-2026. For builders, Muse Spark is a competitor's demo, not an option — the deployable Meta lineage remains Llama 4 (open weights, one tier down).

Safety & privacy

Evaluated under Meta's Advanced AI Scaling Framework across frontier risks, behavioral alignment, and adversarial robustness, with results "within safe margins" per the launch material and a published safety report (arxiv 2606.12429) — more safety documentation than capability documentation, notably. Consumer deployment implies integrated content moderation. Input-training and retention policies follow Meta's consumer terms, which enterprises would not accept — moot until an API exists. No compliance certifications apply to a consumer-only product.

Ecosystem & tooling

Consumer-complete, developer-empty: meta.ai, the Meta AI app, in-app assistants across WhatsApp/Instagram/Messenger, and Ray-Ban Meta glasses (first smart glasses on a superintelligence-lab model, 2026-06-23). No SDKs, no framework integrations, no third-party products — the private API preview has produced nothing public. Popularity is mainstream by raw reach (billions of surfaced users), niche by builder adoption (zero possible). The safety report (arxiv 2606.12429) is the only artifact the technical community can engage with.

Buyer questions

Can I use Muse Spark via API?

Not yet — the API is a private preview for select users with no public rate card or sign-up path. Every "Muse Spark API pricing" article in circulation is speculation. Watch Meta's developer channels for GA.

What does it cost?

Nothing, for consumers: free on meta.ai, the Meta AI app, and Ray-Ban Meta glasses. There is no paid tier and no enterprise offering to price.

Is it open-weights like Llama?

No — this is Meta's first closed frontier release, and the launch language positions Muse as the successor to Llama for frontier work. Llama 4 remains the open line; Llama 4 Behemoth's weights never shipped and no Llama 5 exists.

How good is it really?

Verifiably: HLE 58% and FrontierScience 38% in Contemplating mode (Meta-published), AA Intelligence Index 43 in standard operation (independent). The ceiling is frontier-tier; everyday behavior is a step below the top closed flagships.

What is Contemplating mode?

Meta's deep-reasoning configuration — multiple agents reasoning in parallel before answering. It rolls out gradually across surfaces and is where the headline benchmark numbers come from.

Should enterprises rethink Llama plans because of this?

Yes, at the roadmap level: Meta's frontier effort has moved to a closed line, so long-term self-host strategies should evaluate GLM-5.2, Nemotron 3, and the open Qwen line as the continuing open-weights frontier.

What about privacy?

Consumer Meta terms apply — assume interactions feed product improvement. There is no enterprise data agreement because there is no enterprise product.

Comparable models

The predecessor and the last open Meta frontier model — buildable, hostable, and documented where Muse Spark is none of those; Muse claims parity at a fraction of the compute and adds the reasoning/multimodal ceiling.

Gemini 3.5 FlashGoogle

The closest consumer-scale multimodal comparison — behind on reasoning ceiling but with a full public API, pricing, and enterprise surface, which is exactly what Muse Spark lacks.

GPT-5.5OpenAI

The frontier standard Muse Spark is measured against in coverage — verifiable benchmarks, mature API ecosystem, and paid tiers against Meta's free distribution and closed access.

Sources

Primary references used to verify this review.

Model specs

Input price
— / Mtok
Output price
— / Mtok
Cached input
Batch (in/out)
Context window
262K tokens
Max output
— tokens
Knowledge cutoff
Undisclosed
Released
2026-04-07
Modalities
text, image, audio → text
Output speed
Not profiled
License
Proprietary
Clouds
First-party API

Last verified 2026-07-02