Nemotron 3 Ultra

GALatest Ultra

by NVIDIA · Nemotron 3 family · best for self-hostable 1M-context reasoning

FrontierReasoningCodingOpen-WeightsLong-Context
8.5
AI Panel Score
Value 8.5/10

Nemotron 3 Ultra is NVIDIA's frontier-adjacent open-weights flagship: a 550B-parameter hybrid MoE (55B active) released 2026-06-04 under the permissive OpenMDW-1.1 license, with a full 1M-token context and — unusually for mid-2026 — benchmarks you can actually reproduce, because the weights are public. It posts strong, verifiable coding and reasoning numbers (SWE-bench Verified 70.7, LiveCodeBench 89.0, MMLU-Pro 86.8) and a class-leading long-context result (RULER 94.7 at 1M tokens). It is text-only and multi-node to serve, but arriving just as Meta closed its open frontier line, it makes NVIDIA the credible Western open-weights anchor — and every self-hosted copy runs best on NVIDIA silicon.

What's new

  • Completes the Nemotron 3 family (Nano / Super 120B-A12B / Ultra) with the top-tier 550B/55B-active flagship
  • Hybrid Mamba-2 + LatentMoE architecture with multi-token prediction — a state-space/attention/MoE blend tuned for throughput at long context
  • Full 1M-token context with a class-leading RULER score of 94.7 at 1M — the long window is genuinely usable, not nominal
  • Verifiable primary-source benchmarks on the public model card: SWE-bench Verified 70.7, LiveCodeBench v6 89.0, MMLU-Pro 86.8, GPQA 87.0
  • Open weights under OpenMDW-1.1 with day-one vLLM / SGLang / TensorRT-LLM support and NVIDIA NIM distribution — landing as Meta's Muse Spark ended the open Llama frontier

Benchmarks

BenchmarkScoreSource
MMLU-Pro86.8%huggingface.co 2026-06-04T00:00:00.000Z
GPQA Diamond87%huggingface.co 2026-06-04T00:00:00.000Z
LiveCodeBench89%huggingface.co 2026-06-04T00:00:00.000Z
SWE-bench Verified70.7%huggingface.co 2026-06-04T00:00:00.000Z

AI Panel Review

Six personas, six verdicts — the same panel that reviews every product on TopReviewed.

Decision Maker8/10
NVIDIA just filled the open-weights frontier gap Meta left - verifiable benchmarks, a 1M context, and no vendor holding your inference hostage.

Nemotron 3 Ultra is the strategically safe open-weights bet of mid-2026: it comes from the one vendor whose survival does not depend on selling you tokens, its benchmarks are first-party but independently checkable because the weights are public, and it lands exactly as Meta abandons the open frontier. The capability is real (SWE-bench Verified 70.7, MMLU-Pro 86.8) and RULER 94.7 at 1M tokens is a genuine long-context edge. The costs are operational: 550B parameters is multi-node serving, the OpenMDW-1.1 license is newer and warrants legal review, and it is text-only, so multimodal stacks need a second model. For an org building a self-hosted frontier layer, this belongs on the shortlist beside GLM-5.2.

Strategic Fit 8Vendor Risk 8Roadmap Confidence 8
Pros
  • NVIDIA-backed open weights
  • verifiable strong benchmarks
  • class-leading long context
Cons
  • Multi-node to host
  • custom license needs review
  • text-only
Right for: organizations standing up a self-hosted frontier-class model layer
Avoid if: you need multimodality or a light single-GPU footprint
Domain Strategist8.5/10
The chip company shipping the open frontier model is the most on-brand move of the year: NVIDIA sells the shovels and now hands out the map.

NVIDIA releasing a frontier-adjacent open-weights model is vertically brilliant — every self-hosted deployment of Nemotron 3 Ultra runs best on NVIDIA silicon, so the model is a demand generator for the hardware. Landing it weeks after Muse Spark closed Meta's open line, and alongside Chinese labs like GLM, Qwen, and Kimi, positions NVIDIA as the credible Western open-weights anchor at a moment when that role was suddenly vacant. The hybrid Mamba-2 / LatentMoE architecture also advertises NVIDIA's research depth and its hardware's efficiency story. The strategic risk is minimal because NVIDIA is not competing for token revenue; the model's job is to keep the open ecosystem thriving on GPUs.

Competitive Positioning 8.5Differentiation 8.5Market Timing 9
Pros
  • Fills the Western open-frontier vacuum
  • hardware-aligned strategy
  • research showcase
Cons
  • English-centric
  • ecosystem still forming
Right for: reading where open-weights leadership migrates post-Meta
Avoid if: you expect a multimodal or consumer play — this is infrastructure
Finance Lead8/10
No license fee and no per-token bill - but a 550B model is a real hardware line item, so the savings show up only at serious scale.

The direct costs are attractive: OpenMDW-1.1 open weights mean no license fee and no metered API, so at high, sustained volume the amortized cost per token on owned or rented GPUs undercuts frontier APIs substantially. The offsetting reality is the footprint — 550B total parameters (55B active) is multi-node inference, so the fixed cost floor is high and only justified by volume. Where a first-party API price is desired, aggregators like DeepInfra host it; that route restores metered pricing at the cost of the sovereignty benefit. Value is strong for large self-hosted deployments and mediocre for anyone who cannot keep the cluster busy.

Cost Efficiency 8Pricing Transparency 7Value per Dollar 8
Pros
  • No license or per-token cost when self-hosted
  • aggregator route available
  • efficient 55B active path
Cons
  • Multi-node hardware floor
  • only economical at scale
Right for: high-volume self-hosted or reserved-GPU workloads
Avoid if: you cannot keep a multi-GPU cluster utilized
Domain Practitioner8.5/10
Strong coding numbers you can actually verify, a real 1M context that holds up on RULER, and day-one vLLM support - this one is a pleasure to deploy.

For builders this is one of the more satisfying open releases: SWE-bench Verified 70.7 and LiveCodeBench 89.0 translate to genuinely useful coding help, the reasoning toggle lets you trade latency for depth, and RULER 94.7 at 1M tokens means the long window is real rather than nominal. NVIDIA ships first-class support across vLLM, SGLang, and TensorRT-LLM, so standing it up is straightforward if you have the GPUs. The friction is purely infrastructural — 550B parameters demand a multi-node or high-end single-node setup — and it is text-only, so screenshot-driven flows need a companion vision model. Solid, honest, buildable.

API Ergonomics 8.5Tool/Agent Support 8.5Reliability 8.5
Pros
  • Verifiable strong coding
  • genuine 1M context
  • excellent serving-engine support
Cons
  • Heavy to host
  • text-only
  • reasoning adds latency
Right for: teams with GPU capacity building long-context coding and reasoning agents
Avoid if: you need vision in-model or light hardware
Power User7.5/10
Frontier-adjacent answers with a huge memory, if you can reach it - there's no app, so you're going through an aggregator or your own server.

Accessed through DeepInfra or a self-hosted endpoint, Nemotron 3 Ultra gives a heavy daily user strong technical answers, dependable coding help, and the standout ability to hold an enormous context without losing the thread (RULER 94.7 at 1M). The optional reasoning mode produces visibly more careful answers on hard questions. The everyday frictions are access and scope: there is no polished consumer app, so you live in API tooling, and it is text-only — no images, no voice — which sends multimodal questions elsewhere. For a technical power user who works mostly in text and wants an open, long-memory model, it delivers; for a general assistant it is narrower than the frontier consumer products.

Output Quality 8Speed 7.5Everyday Usefulness 7
Pros
  • Strong technical answers
  • huge reliable context
  • open and inspectable
Cons
  • No consumer app
  • text-only
  • access via API or self-host only
Right for: technical daily drivers who work in text and want an open model
Avoid if: you want images, voice, or a turnkey app
Skeptic8/10
Rare thing for 2026: an open model whose headline numbers you can actually reproduce - the weights are right there, so the claims survive scrutiny.

Where this month's Chinese releases lean on vendor-run or contested suites, NVIDIA's numbers sit on a public model card backed by downloadable weights, so SWE-bench Verified 70.7, LiveCodeBench 89.0, and RULER 94.7 are independently checkable rather than take-our-word-for-it — that alone puts Nemotron 3 Ultra in a higher trust bracket than Kimi K2.7-Code or the unverified Qwen3.7-Max card. The legitimate residue: first-party benchmarks still deserve replication, IMOAnswerBench-with-tools and RULER are young metrics with thin comparison sets, the OpenMDW-1.1 license is new enough that "open" needs a careful read, and NVIDIA has an obvious incentive to make a GPU-hungry model look great. But nothing here smells like benchmark theater; publishing the weights is the strongest possible answer to a skeptic.

Claim Accuracy 8.5Weakness Severity 7.5Hype vs Reality 8
Pros
  • Public weights make claims verifiable
  • benchmarks primary-source and specific
Cons
  • First-party numbers await replication
  • new license
  • vendor hardware incentive
Right for: buyers who verify claims and value reproducibility
Avoid if: you distrust any first-party benchmark regardless of openness

Strengths

  • Class-leading, verifiable long context: RULER 94.7 at 1M tokens, with public weights so the number is reproducible
  • Strong, primary-source coding and reasoning: SWE-bench Verified 70.7, LiveCodeBench 89.0, MMLU-Pro 86.8, GPQA 87.0
  • NVIDIA-backed open weights fill the Western open-frontier vacuum left by Meta's closed Muse line
  • Excellent serving-engine support (vLLM, SGLang, TensorRT-LLM) plus NIM and DeepInfra hosting options

Limitations

  • Heavy to host: 550B parameters (55B active) is multi-node inference, a high fixed-cost floor
  • Text-only — no vision, OCR, or audio; multimodal stacks need a second model
  • No first-party API pricing and a newer OpenMDW-1.1 license that needs legal review
  • English-centric multilingual coverage; optional reasoning mode adds latency when enabled

Best use cases

Self-hosted, long-context coding and reasoning agents at organizations with the GPU capacity to run a 550B MoE — the SWE-bench and RULER numbers translate directly to multi-file coding work and whole-repository or book-length context that stays coherent. Teams building a sovereign, open-weights frontier layer who want verifiable capability from a Western vendor rather than a Chinese lab or a closed API. Reserved-GPU or high, steady-volume workloads where owning the deployment beats metered tokens. Not the pick where vision, a light single-GPU footprint, or a turnkey consumer app is required.

Deep dive

The full research notes behind this review — verified against primary sources.

Architecture

A 550B-parameter hybrid mixture-of-experts activating 55B per token, described by NVIDIA as a "Mamba2-Transformer Hybrid Latent Mixture of Experts (LatentMoE) with Multi-Token Prediction" — Mamba-2 state-space layers interleaved with select attention layers and MoE feed-forward blocks, a design that keeps compute manageable across a 1M-token window. The hybrid attention/state-space stack does not map to a single conventional attention type, so that field is left null and described here. Context is 1M tokens; the model is text-only in and out. Pre-training data runs to September 2025 with post-training through May 2026. Expert count, layer depth, and a hard output ceiling are not broken out on the card. Weights ship in BF16 with FP8 for lighter deployment.

Capabilities

Long context (9.3) is the standout and it is verifiable: RULER 94.7 at 1M tokens is among the strongest disclosed long-context results in the catalog, meaning the full window degrades gracefully rather than collapsing. Coding (8.8) rests on SWE-bench Verified 70.7 and LiveCodeBench v6 89.0 — strong, checkable numbers — and reasoning (8.8) on MMLU-Pro 86.8 and GPQA 87.0, with an optional thinking mode that trades latency for depth. Math (8.5) is anchored by IMOAnswerBench-with-tools 92.3. Agentic behavior and function calling (both 8.5) reflect tool-use tuning and the multi-token-prediction design. Instruction following is strong (8.5). It is text-only, so vision and OCR score 0, and there is no built-in web search (real-time data 0). Multilingual (7.5) is solid but English-centric — not the headline.

Benchmark analysis

Benchmark Score Note Source
SWE-bench Verified 70.7 agentic software engineering NVIDIA model card
LiveCodeBench v6 89.0 competitive coding NVIDIA model card
MMLU-Pro 86.8 broad reasoning NVIDIA model card
GPQA 87.0 graduate-level science NVIDIA model card
RULER (1M) 94.7 long-context retention at 1M tokens NVIDIA model card
IMOAnswerBench (with tools) 92.3 olympiad math with tools NVIDIA model card

SWE-bench Verified, LiveCodeBench, MMLU-Pro, and GPQA enter this catalog's comparable columns as standard, primary-source metrics. RULER (a long-context benchmark distinct from the MRCR column used elsewhere) and IMOAnswerBench-with-tools (not the AIME 2025 exam) are reported in prose rather than mapped, to avoid mismatched compare rows. Because the weights are public, every one of these numbers is independently reproducible — the reason this row carries high research confidence.

Speed & latency

NVIDIA publishes no median throughput or first-token latency, and no independent tracker had measured it at verification, so speed fields are null. The 55B active-parameter path and the Mamba-2 hybrid design target efficient long-context serving, and NVIDIA's own TensorRT-LLM path is built for high throughput on its GPUs — but real-world numbers depend entirely on your deployment. The optional reasoning mode adds latency when enabled; disable it for interactive, latency-sensitive paths.

Pricing analysis

Surface Cost Notes
Weights (self-host) Free OpenMDW-1.1 license
NVIDIA NIM Deployment managed containerized inference
DeepInfra Metered hosted aggregator API
First-party token API None NVIDIA does not sell tokens directly

NVIDIA's business is hardware, not tokens, so there is no first-party per-token rate card — pricing fields are null. Self-hosting is free under the license; a metered API is available through DeepInfra and other hosts if you would rather rent than run.

Deployment & access

Open weights on Hugging Face (nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16) under OpenMDW-1.1, in BF16 and FP8. Serving is first-class across vLLM, SGLang, and TensorRT-LLM, with NVIDIA NIM providing a managed containerized path and DeepInfra offering a hosted API. The practical constraint is scale: 550B total parameters (55B active) is multi-node or high-end multi-GPU inference — this is a data-center model, not a workstation one — though the MoE active path keeps per-token serving cost reasonable for the class. Text-only, so multimodal pipelines pair it with a separate vision model.

Safety & privacy

NVIDIA ships the model with published evaluation results on the card and positions it within its broader trustworthy-AI and NeMo Guardrails tooling rather than a single named safety framework, so the safety-framework field is null. Self-hosting means NVIDIA does not see inputs and cannot train on them; content moderation and refusal behavior are the deployer's to configure. The OpenMDW-1.1 license is newer than Apache/MIT and warrants a legal read before commercial deployment, though it permits commercial use and self-hosting. No compliance certifications attach to an open-weights release — governance travels with wherever you run it.

Ecosystem & tooling

Open weights on Hugging Face with day-one support across NVIDIA's serving stack — vLLM, SGLang, and TensorRT-LLM — plus availability through NVIDIA NIM and hosted inference on DeepInfra. Python tooling is standard and OpenAI-compatible endpoints are the norm on the hosts. The model completes the Nemotron 3 family (Nano, Super 120B-A12B, and Ultra), giving deployers a size ladder on a common architecture. Community adoption is growing quickly among self-hosters, helped by NVIDIA's obvious interest in a thriving open ecosystem on its hardware.

Buyer questions

Is Nemotron 3 Ultra really open-weights?

Yes — the weights are on Hugging Face under the OpenMDW-1.1 license, a permissive open model license. Read the terms before commercial deployment, but self-hosting and commercial use are permitted.

How hard is it to run?

It is a 550B-parameter MoE activating 55B per token, so plan for multi-node or high-end multi-GPU inference. NVIDIA ships first-class support for vLLM, SGLang, and TensorRT-LLM, and DeepInfra hosts it if you prefer a metered API.

How good is the long context?

Class-leading on disclosure: RULER at 1M tokens scores 94.7, meaning the full window stays usable rather than degrading — one of the strongest verifiable long-context results in the catalog.

What are its best benchmarks?

SWE-bench Verified 70.7, LiveCodeBench v6 89.0, MMLU-Pro 86.8, and GPQA 87.0, all on the public model card. Coding and reasoning are its strengths.

Does it do images?

No — it is text-only for both input and output. Pair it with a vision model if your pipeline needs screenshots or documents-as-images.

What is the knowledge cutoff?

Pre-training data runs to September 2025, with post-training through May 2026. Use retrieval for anything more recent.

How does it compare to GLM-5.2?

Nemotron leads on disclosed SWE-bench Verified and RULER long-context and comes from a Western vendor; GLM-5.2 is lighter to host, MIT-licensed, and higher on the broad AA Index. Both are top open-weights choices; the trade is footprint and license versus long-context and provenance.

Comparable models

GLM-5.2Z.ai

The other headline open-weights release of June — lighter to host (40B active vs 55B) and under a clean MIT license, and stronger on the broad AA Index (51), while Nemotron leads on disclosed SWE-bench Verified (70.7) and RULER long-context (94.7 at 1M).

Qwen3-235B-A22BAlibaba Cloud

The established open-weights workhorse — Apache-2.0 and far lighter to serve, but a generation behind on coding and long-context, where Nemotron's 1M window and verifiable SWE-bench numbers pull ahead.

The open Meta model Nemotron effectively succeeds in the frontier-open role — Maverick still ships weights and multimodality, but Meta moved its frontier work to the closed Muse line, leaving NVIDIA's Nemotron as the actively-advancing Western open option.

Sources

Primary references used to verify this review.

Model specs

Input price
— / Mtok
Output price
— / Mtok
Cached input
Batch (in/out)
Context window
1M tokens
Max output
— tokens
Knowledge cutoff
2025-09
Released
2026-06-03
Modalities
text → text
Output speed
Not profiled
License
Open weights (custom-OpenMDW-1.1)
Clouds
First-party API

Does not train on API inputs by default

Last verified 2026-07-02