by Meta · Muse family · best for frontier reasoning for consumers, not yet for builders
Muse Spark is the model that ended the Llama era: Meta Superintelligence Labs' first frontier release (2026-04-08) is proprietary, closed-weights, and consumer-first — free to billions on meta.ai, the Meta AI app, and (since June 23) Ray-Ban Meta glasses, with the API stuck in private preview and no rate card. A natively multimodal reasoning model with visual chain of thought and a parallel-multi-agent "Contemplating" mode that scores 58% on Humanity's Last Exam, it reportedly reaches Llama 4 Maverick parity with over an order of magnitude less compute. Catalog it for what it signals as much as what it scores: Meta now competes at the frontier on its own terms, and builders are not invited yet.
| Benchmark | Score | Source |
|---|---|---|
| Humanity's Last Exam | 58% | ai.meta.com 2026-04-08T00:00:00.000Z |
| Artificial Analysis Index | 43 | artificialanalysis.ai 2026-07-03T00:00:00.000Z |
Six personas, six verdicts — the same panel that reviews every product on TopReviewed.
“You can't buy it, license it, or build on it — Muse Spark is a strategy update from Meta, not a procurement option.”
There is no purchasing decision to make, which is itself the decision-relevant fact. If your organization standardized on Llama for self-hosted AI, Muse Spark tells you Meta's frontier investment has moved somewhere you cannot follow: Llama 4 is now the effective end of the open line, Behemoth never shipped weights, and "Llama 5 does not exist." Plan open-weights roadmaps around GLM, Qwen, and Nemotron instead, and track the Muse API preview for when it opens. The private-preview structure suggests Meta is testing enterprise appetite; until a rate card and terms exist, any dependency here is imaginary.
“The most strategically legible release of 2026: Meta stopped subsidizing the ecosystem and started competing for the frontier itself.”
Muse Spark reorders the market map. Meta spent four generations commoditizing its rivals' complement with open weights; MSL's first act reverses the strategy — closed weights, consumer-first distribution through a three-billion-user funnel, and hardware differentiation via Ray-Ban glasses that no competitor can match. The compute-efficiency claim (Maverick parity at a tenth-plus less compute) is aimed at investors as much as researchers. The vacuum it creates in open weights is already being filled by Chinese labs and NVIDIA, which changes the geopolitics of the open ecosystem. Distribution is the moat here, not benchmarks — and it's a very good moat.
“Free is not a price, it's an absence of one — there is no TCO to model because there is nothing to buy.”
For budget purposes this row is empty: no API pricing, no enterprise tier, no committed-use discounts, nothing to negotiate. The consumer-free positioning is financed by Meta's ad business, not by you, and it tells you nothing about what API pricing will look like if it ships. The finance-relevant angle is indirect: if you built cost models around free Llama weights continuing to improve, that assumption just expired — self-hosted TCO projections should now assume the open Meta line is frozen at Llama 4, with the open-weights alternatives (GLM-5.2, Nemotron) as the new baseline curves. Value score reflects consumer value (high) against enterprise value (currently zero).
“Tool use, orchestration, visual chain of thought — all the features I'd love to build on, none of them behind an endpoint I can call.”
From a builder's chair this is a shop window, not a shop. The capabilities Meta describes — parallel multi-agent Contemplating, visual chain of thought, native tool use — are exactly the primitives agent developers want, and there is no API, SDK, structured-output spec, or rate limit documentation to evaluate any of it. The private preview has no public entry path. What can be assessed: the consumer product is polished, multimodal interactions feel genuinely native rather than stitched, and the reasoning mode produces visibly deliberate answers. Practitioners should note the pattern from wavespeed and others: every "Muse Spark API" tutorial in circulation is speculation. Watch the preview; build on something else.
“The best free AI on the planet right now — frontier reasoning in my pocket and on my face, for exactly zero dollars.”
For daily consumer use, Muse Spark is a step change: frontier-tier reasoning, image understanding that genuinely reasons over what it sees, and voice interaction, free, inside apps billions already have. Contemplating mode on a hard question produces answers competitive with paid frontier subscriptions. The Ray-Ban integration is the sleeper feature — hands-free multimodal AI that works — and the June glasses launch made it the first frontier model people wear. Frictions: no power-user controls (no system prompts, no API escape hatch, no conversation export to speak of), Meta-grade content policies, and the knowledge that your usage feeds Meta's flywheel. As a free product it's extraordinary; as a power tool it's a walled garden.
“Two cherry-picked numbers in best-case mode, everything else undisclosed — and the everyday index says the ceiling isn't the product.”
The claims are carefully constructed. HLE 58% and FrontierScience 38% are quoted exclusively in Contemplating mode — the compute-maximal configuration — while the independently measured everyday index (AA 43) lands mid-pack among 2026 flagships, below GLM-5.2's 51. The order-of-magnitude efficiency claim is unverifiable without disclosed compute budgets. Architecture, parameters, cutoff: all withheld, from the company that used to publish everything. The GPQA/MMMU-Pro numbers circulating in coverage trace to no Meta source. What survives scrutiny: the HLE result is on Meta's letterhead, the safety report is real and substantive, and the consumer product demonstrably works at scale. Muse Spark is real and good; the framing wants you to believe it is the frontier leader, and the verifiable record says "frontier-adjacent, brilliantly distributed."
For consumers: free access to a frontier-class reasoning model with vision and voice, on the phone and on glasses — the strongest no-cost AI product on the market. For everyone else, the use case is strategic awareness: competitive analysts tracking Meta's pivot, enterprises reassessing Llama roadmap risk now that Meta's best work is closed, and builders deciding whether to wait for the API or commit elsewhere. Until general API access ships, Muse Spark cannot be a component in anyone's stack.
The full research notes behind this review — verified against primary sources.
Undisclosed — a sentence that has never before applied to a Meta frontier model. No parameter counts, no architecture type, no attention details, no training scale; the safety report (arxiv 2606.12429) is the only technical artifact. What Meta states: natively multimodal reasoning with visual chain of thought, tool use, and multi-agent orchestration, achieving Llama 4 Maverick parity "with over an order of magnitude less compute." Independent trackers measure a 262K-token context window. Knowledge cutoff, output ceiling, and tokenizer are all unpublished. Ignore the 1M-context figure circulating from one aggregator — it does not match Artificial Analysis or any Meta statement.
Reasoning (8.5) is the demonstrated headline: 58% on Humanity's Last Exam in Contemplating mode is a frontier-tier result, and the parallel-multi-agent approach is architecturally distinctive. Vision (8.5) rests on native multimodality and visual chain of thought — the model reasons over images rather than captioning them — with speech input rounding out the everyday surface. Long context (7.0) reflects the measured 262K window, mid-pack for 2026. Agentic capability (6.5) is scored down not for the model but for access: tool use and orchestration exist, but without a public API nobody outside Meta can build agents on it. Coding, math, multilingual, OCR, creative, instruction-following, and function-calling are left unscored — Meta has published nothing, and no independent evaluation surface exists without an API. That null column is itself the honest description of Muse Spark today.
| Benchmark | Score | Note | Source |
|---|---|---|---|
| Humanity's Last Exam | 58% | Contemplating mode | Meta announcement |
| FrontierScience Research | 38% | Contemplating mode | Meta announcement |
| AA Intelligence Index | 43 | measured without Contemplating-style depth | Artificial Analysis |
Meta's disclosure is thin: two headline numbers in the launch post, both in the maximum-reasoning mode. Secondary sources circulate GPQA Diamond and MMMU-Pro figures that do not appear in any primary Meta material and are excluded here. The gap between HLE 58 (Contemplating) and AA 43 (standard operation) is worth understanding — the model's ceiling and its everyday behavior are different products.
Unmeasured — with no public API, independent trackers cannot benchmark throughput or first-token latency, and Meta publishes none. Anecdotally, standard responses in the consumer apps are chat-snappy while Contemplating mode visibly takes its time (parallel agents are compute-hungry by design). On-glasses inference via Ray-Ban Meta is server-routed, not local.
| Surface | Cost | Notes |
|---|---|---|
| meta.ai / Meta AI app | Free | consumer, at billions-scale |
| Ray-Ban Meta glasses | Free with device | since 2026-06-23 |
| API | Private preview | no rate card, select partners only |
| Weights | Not available | first closed Meta frontier model |
There is no way to pay Meta for this model today, and no way to build on it without an invitation. All pricing fields in this row are null for that reason — not unknown, nonexistent.
Consumer surfaces only: meta.ai on web, the Meta AI app, WhatsApp/Instagram/Messenger integration points, and Ray-Ban Meta glasses. No public API, no cloud-platform distribution, no weights, no self-hosting. The private API preview has select users; wavespeed.ai and others document the absence of any general access path as of mid-2026. For builders, Muse Spark is a competitor's demo, not an option — the deployable Meta lineage remains Llama 4 (open weights, one tier down).
Evaluated under Meta's Advanced AI Scaling Framework across frontier risks, behavioral alignment, and adversarial robustness, with results "within safe margins" per the launch material and a published safety report (arxiv 2606.12429) — more safety documentation than capability documentation, notably. Consumer deployment implies integrated content moderation. Input-training and retention policies follow Meta's consumer terms, which enterprises would not accept — moot until an API exists. No compliance certifications apply to a consumer-only product.
Consumer-complete, developer-empty: meta.ai, the Meta AI app, in-app assistants across WhatsApp/Instagram/Messenger, and Ray-Ban Meta glasses (first smart glasses on a superintelligence-lab model, 2026-06-23). No SDKs, no framework integrations, no third-party products — the private API preview has produced nothing public. Popularity is mainstream by raw reach (billions of surfaced users), niche by builder adoption (zero possible). The safety report (arxiv 2606.12429) is the only artifact the technical community can engage with.
Not yet — the API is a private preview for select users with no public rate card or sign-up path. Every "Muse Spark API pricing" article in circulation is speculation. Watch Meta's developer channels for GA.
Nothing, for consumers: free on meta.ai, the Meta AI app, and Ray-Ban Meta glasses. There is no paid tier and no enterprise offering to price.
No — this is Meta's first closed frontier release, and the launch language positions Muse as the successor to Llama for frontier work. Llama 4 remains the open line; Llama 4 Behemoth's weights never shipped and no Llama 5 exists.
Verifiably: HLE 58% and FrontierScience 38% in Contemplating mode (Meta-published), AA Intelligence Index 43 in standard operation (independent). The ceiling is frontier-tier; everyday behavior is a step below the top closed flagships.
Meta's deep-reasoning configuration — multiple agents reasoning in parallel before answering. It rolls out gradually across surfaces and is where the headline benchmark numbers come from.
Yes, at the roadmap level: Meta's frontier effort has moved to a closed line, so long-term self-host strategies should evaluate GLM-5.2, Nemotron 3, and the open Qwen line as the continuing open-weights frontier.
Consumer Meta terms apply — assume interactions feed product improvement. There is no enterprise data agreement because there is no enterprise product.
The predecessor and the last open Meta frontier model — buildable, hostable, and documented where Muse Spark is none of those; Muse claims parity at a fraction of the compute and adds the reasoning/multimodal ceiling.
The closest consumer-scale multimodal comparison — behind on reasoning ceiling but with a full public API, pricing, and enterprise surface, which is exactly what Muse Spark lacks.
The frontier standard Muse Spark is measured against in coverage — verifiable benchmarks, mature API ecosystem, and paid tiers against Meta's free distribution and closed access.
Primary references used to verify this review.
Last verified 2026-07-02