by NVIDIA · Nemotron 3 family · best for self-hostable 1M-context reasoning
Nemotron 3 Ultra is NVIDIA's frontier-adjacent open-weights flagship: a 550B-parameter hybrid MoE (55B active) released 2026-06-04 under the permissive OpenMDW-1.1 license, with a full 1M-token context and — unusually for mid-2026 — benchmarks you can actually reproduce, because the weights are public. It posts strong, verifiable coding and reasoning numbers (SWE-bench Verified 70.7, LiveCodeBench 89.0, MMLU-Pro 86.8) and a class-leading long-context result (RULER 94.7 at 1M tokens). It is text-only and multi-node to serve, but arriving just as Meta closed its open frontier line, it makes NVIDIA the credible Western open-weights anchor — and every self-hosted copy runs best on NVIDIA silicon.
| Benchmark | Score | Source |
|---|---|---|
| MMLU-Pro | 86.8% | huggingface.co 2026-06-04T00:00:00.000Z |
| GPQA Diamond | 87% | huggingface.co 2026-06-04T00:00:00.000Z |
| LiveCodeBench | 89% | huggingface.co 2026-06-04T00:00:00.000Z |
| SWE-bench Verified | 70.7% | huggingface.co 2026-06-04T00:00:00.000Z |
Six personas, six verdicts — the same panel that reviews every product on TopReviewed.
“NVIDIA just filled the open-weights frontier gap Meta left - verifiable benchmarks, a 1M context, and no vendor holding your inference hostage.”
Nemotron 3 Ultra is the strategically safe open-weights bet of mid-2026: it comes from the one vendor whose survival does not depend on selling you tokens, its benchmarks are first-party but independently checkable because the weights are public, and it lands exactly as Meta abandons the open frontier. The capability is real (SWE-bench Verified 70.7, MMLU-Pro 86.8) and RULER 94.7 at 1M tokens is a genuine long-context edge. The costs are operational: 550B parameters is multi-node serving, the OpenMDW-1.1 license is newer and warrants legal review, and it is text-only, so multimodal stacks need a second model. For an org building a self-hosted frontier layer, this belongs on the shortlist beside GLM-5.2.
“The chip company shipping the open frontier model is the most on-brand move of the year: NVIDIA sells the shovels and now hands out the map.”
NVIDIA releasing a frontier-adjacent open-weights model is vertically brilliant — every self-hosted deployment of Nemotron 3 Ultra runs best on NVIDIA silicon, so the model is a demand generator for the hardware. Landing it weeks after Muse Spark closed Meta's open line, and alongside Chinese labs like GLM, Qwen, and Kimi, positions NVIDIA as the credible Western open-weights anchor at a moment when that role was suddenly vacant. The hybrid Mamba-2 / LatentMoE architecture also advertises NVIDIA's research depth and its hardware's efficiency story. The strategic risk is minimal because NVIDIA is not competing for token revenue; the model's job is to keep the open ecosystem thriving on GPUs.
“No license fee and no per-token bill - but a 550B model is a real hardware line item, so the savings show up only at serious scale.”
The direct costs are attractive: OpenMDW-1.1 open weights mean no license fee and no metered API, so at high, sustained volume the amortized cost per token on owned or rented GPUs undercuts frontier APIs substantially. The offsetting reality is the footprint — 550B total parameters (55B active) is multi-node inference, so the fixed cost floor is high and only justified by volume. Where a first-party API price is desired, aggregators like DeepInfra host it; that route restores metered pricing at the cost of the sovereignty benefit. Value is strong for large self-hosted deployments and mediocre for anyone who cannot keep the cluster busy.
“Strong coding numbers you can actually verify, a real 1M context that holds up on RULER, and day-one vLLM support - this one is a pleasure to deploy.”
For builders this is one of the more satisfying open releases: SWE-bench Verified 70.7 and LiveCodeBench 89.0 translate to genuinely useful coding help, the reasoning toggle lets you trade latency for depth, and RULER 94.7 at 1M tokens means the long window is real rather than nominal. NVIDIA ships first-class support across vLLM, SGLang, and TensorRT-LLM, so standing it up is straightforward if you have the GPUs. The friction is purely infrastructural — 550B parameters demand a multi-node or high-end single-node setup — and it is text-only, so screenshot-driven flows need a companion vision model. Solid, honest, buildable.
“Frontier-adjacent answers with a huge memory, if you can reach it - there's no app, so you're going through an aggregator or your own server.”
Accessed through DeepInfra or a self-hosted endpoint, Nemotron 3 Ultra gives a heavy daily user strong technical answers, dependable coding help, and the standout ability to hold an enormous context without losing the thread (RULER 94.7 at 1M). The optional reasoning mode produces visibly more careful answers on hard questions. The everyday frictions are access and scope: there is no polished consumer app, so you live in API tooling, and it is text-only — no images, no voice — which sends multimodal questions elsewhere. For a technical power user who works mostly in text and wants an open, long-memory model, it delivers; for a general assistant it is narrower than the frontier consumer products.
“Rare thing for 2026: an open model whose headline numbers you can actually reproduce - the weights are right there, so the claims survive scrutiny.”
Where this month's Chinese releases lean on vendor-run or contested suites, NVIDIA's numbers sit on a public model card backed by downloadable weights, so SWE-bench Verified 70.7, LiveCodeBench 89.0, and RULER 94.7 are independently checkable rather than take-our-word-for-it — that alone puts Nemotron 3 Ultra in a higher trust bracket than Kimi K2.7-Code or the unverified Qwen3.7-Max card. The legitimate residue: first-party benchmarks still deserve replication, IMOAnswerBench-with-tools and RULER are young metrics with thin comparison sets, the OpenMDW-1.1 license is new enough that "open" needs a careful read, and NVIDIA has an obvious incentive to make a GPU-hungry model look great. But nothing here smells like benchmark theater; publishing the weights is the strongest possible answer to a skeptic.
Self-hosted, long-context coding and reasoning agents at organizations with the GPU capacity to run a 550B MoE — the SWE-bench and RULER numbers translate directly to multi-file coding work and whole-repository or book-length context that stays coherent. Teams building a sovereign, open-weights frontier layer who want verifiable capability from a Western vendor rather than a Chinese lab or a closed API. Reserved-GPU or high, steady-volume workloads where owning the deployment beats metered tokens. Not the pick where vision, a light single-GPU footprint, or a turnkey consumer app is required.
The full research notes behind this review — verified against primary sources.
A 550B-parameter hybrid mixture-of-experts activating 55B per token, described by NVIDIA as a "Mamba2-Transformer Hybrid Latent Mixture of Experts (LatentMoE) with Multi-Token Prediction" — Mamba-2 state-space layers interleaved with select attention layers and MoE feed-forward blocks, a design that keeps compute manageable across a 1M-token window. The hybrid attention/state-space stack does not map to a single conventional attention type, so that field is left null and described here. Context is 1M tokens; the model is text-only in and out. Pre-training data runs to September 2025 with post-training through May 2026. Expert count, layer depth, and a hard output ceiling are not broken out on the card. Weights ship in BF16 with FP8 for lighter deployment.
Long context (9.3) is the standout and it is verifiable: RULER 94.7 at 1M tokens is among the strongest disclosed long-context results in the catalog, meaning the full window degrades gracefully rather than collapsing. Coding (8.8) rests on SWE-bench Verified 70.7 and LiveCodeBench v6 89.0 — strong, checkable numbers — and reasoning (8.8) on MMLU-Pro 86.8 and GPQA 87.0, with an optional thinking mode that trades latency for depth. Math (8.5) is anchored by IMOAnswerBench-with-tools 92.3. Agentic behavior and function calling (both 8.5) reflect tool-use tuning and the multi-token-prediction design. Instruction following is strong (8.5). It is text-only, so vision and OCR score 0, and there is no built-in web search (real-time data 0). Multilingual (7.5) is solid but English-centric — not the headline.
| Benchmark | Score | Note | Source |
|---|---|---|---|
| SWE-bench Verified | 70.7 | agentic software engineering | NVIDIA model card |
| LiveCodeBench v6 | 89.0 | competitive coding | NVIDIA model card |
| MMLU-Pro | 86.8 | broad reasoning | NVIDIA model card |
| GPQA | 87.0 | graduate-level science | NVIDIA model card |
| RULER (1M) | 94.7 | long-context retention at 1M tokens | NVIDIA model card |
| IMOAnswerBench (with tools) | 92.3 | olympiad math with tools | NVIDIA model card |
SWE-bench Verified, LiveCodeBench, MMLU-Pro, and GPQA enter this catalog's comparable columns as standard, primary-source metrics. RULER (a long-context benchmark distinct from the MRCR column used elsewhere) and IMOAnswerBench-with-tools (not the AIME 2025 exam) are reported in prose rather than mapped, to avoid mismatched compare rows. Because the weights are public, every one of these numbers is independently reproducible — the reason this row carries high research confidence.
NVIDIA publishes no median throughput or first-token latency, and no independent tracker had measured it at verification, so speed fields are null. The 55B active-parameter path and the Mamba-2 hybrid design target efficient long-context serving, and NVIDIA's own TensorRT-LLM path is built for high throughput on its GPUs — but real-world numbers depend entirely on your deployment. The optional reasoning mode adds latency when enabled; disable it for interactive, latency-sensitive paths.
| Surface | Cost | Notes |
|---|---|---|
| Weights (self-host) | Free | OpenMDW-1.1 license |
| NVIDIA NIM | Deployment | managed containerized inference |
| DeepInfra | Metered | hosted aggregator API |
| First-party token API | None | NVIDIA does not sell tokens directly |
NVIDIA's business is hardware, not tokens, so there is no first-party per-token rate card — pricing fields are null. Self-hosting is free under the license; a metered API is available through DeepInfra and other hosts if you would rather rent than run.
Open weights on Hugging Face (nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16) under OpenMDW-1.1, in BF16 and FP8. Serving is first-class across vLLM, SGLang, and TensorRT-LLM, with NVIDIA NIM providing a managed containerized path and DeepInfra offering a hosted API. The practical constraint is scale: 550B total parameters (55B active) is multi-node or high-end multi-GPU inference — this is a data-center model, not a workstation one — though the MoE active path keeps per-token serving cost reasonable for the class. Text-only, so multimodal pipelines pair it with a separate vision model.
NVIDIA ships the model with published evaluation results on the card and positions it within its broader trustworthy-AI and NeMo Guardrails tooling rather than a single named safety framework, so the safety-framework field is null. Self-hosting means NVIDIA does not see inputs and cannot train on them; content moderation and refusal behavior are the deployer's to configure. The OpenMDW-1.1 license is newer than Apache/MIT and warrants a legal read before commercial deployment, though it permits commercial use and self-hosting. No compliance certifications attach to an open-weights release — governance travels with wherever you run it.
Open weights on Hugging Face with day-one support across NVIDIA's serving stack — vLLM, SGLang, and TensorRT-LLM — plus availability through NVIDIA NIM and hosted inference on DeepInfra. Python tooling is standard and OpenAI-compatible endpoints are the norm on the hosts. The model completes the Nemotron 3 family (Nano, Super 120B-A12B, and Ultra), giving deployers a size ladder on a common architecture. Community adoption is growing quickly among self-hosters, helped by NVIDIA's obvious interest in a thriving open ecosystem on its hardware.
Yes — the weights are on Hugging Face under the OpenMDW-1.1 license, a permissive open model license. Read the terms before commercial deployment, but self-hosting and commercial use are permitted.
It is a 550B-parameter MoE activating 55B per token, so plan for multi-node or high-end multi-GPU inference. NVIDIA ships first-class support for vLLM, SGLang, and TensorRT-LLM, and DeepInfra hosts it if you prefer a metered API.
Class-leading on disclosure: RULER at 1M tokens scores 94.7, meaning the full window stays usable rather than degrading — one of the strongest verifiable long-context results in the catalog.
SWE-bench Verified 70.7, LiveCodeBench v6 89.0, MMLU-Pro 86.8, and GPQA 87.0, all on the public model card. Coding and reasoning are its strengths.
No — it is text-only for both input and output. Pair it with a vision model if your pipeline needs screenshots or documents-as-images.
Pre-training data runs to September 2025, with post-training through May 2026. Use retrieval for anything more recent.
Nemotron leads on disclosed SWE-bench Verified and RULER long-context and comes from a Western vendor; GLM-5.2 is lighter to host, MIT-licensed, and higher on the broad AA Index. Both are top open-weights choices; the trade is footprint and license versus long-context and provenance.
The other headline open-weights release of June — lighter to host (40B active vs 55B) and under a clean MIT license, and stronger on the broad AA Index (51), while Nemotron leads on disclosed SWE-bench Verified (70.7) and RULER long-context (94.7 at 1M).
The established open-weights workhorse — Apache-2.0 and far lighter to serve, but a generation behind on coding and long-context, where Nemotron's 1M window and verifiable SWE-bench numbers pull ahead.
The open Meta model Nemotron effectively succeeds in the frontier-open role — Maverick still ships weights and multimodality, but Meta moved its frontier work to the closed Muse line, leaving NVIDIA's Nemotron as the actively-advancing Western open option.
Does not train on API inputs by default
Last verified 2026-07-02