AI-native cloud built for training and inference on NVIDIA GPU clusters
Nebius is an AI-focused cloud platform for teams training and serving machine learning models on NVIDIA GPU infrastructure.
AI Panel Score
6 AI reviews
Reviewed
AI Editor ApprovedApproved and published by our AI Editor-in-Chief after full panel analysis.Nebius is an AI-native cloud platform that provides NVIDIA GPU infrastructure for training and serving machine learning models. It serves AI startups, research teams and enterprises that need large GPU clusters without building their own data centers. Pricing is usage-based: on-demand H100 GPUs start at $3.85 per GPU-hour, preemptible L40S capacity starts at $0.74 per GPU-hour, and multi-month reserved clusters cut on-demand rates by up to 35 percent, with a $25 minimum first payment. Core capabilities include non-virtualized GPU clusters with InfiniBand networking, managed Kubernetes and Slurm orchestration, S3-compatible object storage, and Token Factory, a managed inference service that runs more than 60 open-source models such as DeepSeek and Qwen behind an OpenAI-compatible API, with fine-tuning pipelines and dedicated endpoints backed by a 99.9 percent uptime SLA. It fits teams running sustained training or high-volume inference workloads. Alternatives include CoreWeave, Lambda, Together AI and hyperscalers such as AWS and Google Cloud.
Nebius gives AI teams GPU infrastructure through a cloud console that deploys from zero to clusters in minutes. Users provision GPU virtual machines or multi-node clusters, orchestrate distributed training with managed Kubernetes or Slurm-based clusters, keep datasets and checkpoints in S3-compatible object storage or shared filesystems, and pay per GPU-hour on-demand, at preemptible rates, or through multi-month reservations that cut on-demand pricing by up to 35 percent.
The platform is built on Nebius-designed hardware with non-virtualized NVIDIA GPUs (HGX B300, B200, H200, H100, RTX PRO 6000 and L40S) connected over InfiniBand, and layers MLOps tooling on top: Managed MLflow for experiment tracking, pre-built JupyterLab and vLLM applications, a Container Registry, managed PostgreSQL, and Nebius Echo, an AI agent inside the console that helps manage infrastructure. Nebius Token Factory handles inference, serving 60+ open-source models including DeepSeek-V4-Pro, Qwen3-235B-A22B and GPT OSS 120B through an OpenAI-compatible API, with fine-tuning and LoRA workflows, dedicated endpoints carrying a 99.9% uptime SLA, autoscaling, a zero-retention mode, and SOC 2 Type II, HIPAA and ISO 27001 certifications.
Nebius is aimed at AI startups, research groups and enterprises that need large-scale GPU capacity without operating their own data centers. Billing is usage-based: preemptible L40S capacity starts at $0.74 per GPU-hour, on-demand H100s cost $3.85 per GPU-hour, the minimum first payment is $25, and managed Kubernetes, networking and public IPs are free. In the GPU cloud category it competes with CoreWeave, Lambda and Together AI, as well as hyperscalers such as AWS and Google Cloud.
Developers manage resources through the web console, a CLI, a Terraform provider or a gRPC API, and Token Factory exposes REST endpoints compatible with OpenAI client libraries. Data centers are located in Finland, France and the US, giving teams EU or US regional deployment options for dedicated inference endpoints.
AI agent built into the Nebius console that helps users manage and operate their cloud infrastructure.
Provisions non-virtualized NVIDIA HGX B300, B200, H200 and H100 GPUs as single VMs or multi-node clusters connected over InfiniBand, deployable from the console in minutes.
Manages all cloud resources declaratively via an official Terraform provider, a CLI and a gRPC API, with a Container Registry for images.
Reserved-capacity endpoints with a 99.9% uptime SLA, custom autoscaling, speculative decoding and optional regional deployment in EU or US data centers.
Managed inference service that serves 60+ open-source models, including DeepSeek-V4-Pro, Qwen3-235B-A22B and GPT OSS 120B, with transparent per-token pricing.
Exposes Token Factory inference and fine-tuning through an OpenAI-compatible API, so existing OpenAI client libraries work by swapping the endpoint and key.
Trains models on custom datasets via the dashboard or API, with LoRA adapters supported on all models and full supervised fine-tuning where feasible, then deploys checkpoints directly to endpoints.
Hosted MLflow clusters for experiment tracking and model registry, alongside pre-built JupyterLab and vLLM applications.
Runs containerized workloads with native GPU and InfiniBand support, offered at no charge on top of consumed compute resources.
Deploys Slurm workload manager clusters for scheduling large-scale distributed training jobs across GPU nodes.
Inference mode where requests and outputs are never stored, backed by SOC 2 Type II, HIPAA and ISO 27001 certifications plus SSO and RBAC for team governance.
AWS S3-compatible object storage for datasets and model artifacts at $0.0147 per GiB-month, alongside shared filesystem and WEKA filesystem options for training I/O.
Interruptible GPU capacity at the lowest per-hour rates for fault-tolerant training and batch workloads.
Pay-as-you-go GPU capacity with no commitment for teams that need guaranteed availability.
Multi-month reservations of large-scale GPU clusters for sustained training workloads, priced through sales.
Per-token managed inference for developers and enterprises who want models served without managing GPUs.
Microsoft and Meta already underwrote this vendor bet, so the negotiation is about reserved capacity.
“Nebius is a Nasdaq-listed GPU cloud carrying multi-year capacity contracts from both Microsoft and Meta. For a compute buyer, vendor viability is about as settled as this category gets.”
Microsoft committed $17.4 billion over five years to this vendor's capacity, and Meta followed at up to $27 billion. My first question — will they exist in three years — answered itself. That kind of commitment doesn't happen to shaky vendors.
The offer is clean: H100s at $3.85 per GPU-hour, Token Factory for per-token inference, a $25 minimum first payment. CoreWeave is the obvious comparison; Nebius answers with published pricing and NVIDIA in its investor base since December 2024. Both are real options — this one's been Nasdaq-listed since October 2024.
But the anchor tenants compete with you for the same racks, so reserved capacity is where the negotiation lives. Pilot on-demand this quarter with real workloads. Reserve only what the training roadmap actually needs.
Matches CoreWeave on published pricing and Blackwell access while adding EU data-residency options.
SOC 2 Type II, HIPAA and ISO 27001 cover the compliance file; the Yandex spin-out history still invites a board question.
A $25 minimum first payment and zero-to-cluster-in-minutes provisioning make the pilot nearly free to start.
One vendor covers training clusters and per-token inference through Token Factory.
Nasdaq-listed with five-year Microsoft ($17.4B) and Meta (up to $27B) capacity contracts on the books.
AI teams that need large-scale GPU capacity without hyperscaler contracts.
Buyers who require a vendor with decades of operating history.
A credible multi-year GPU supply line, if you reserve before the anchor tenants eat the capacity.
“For multi-year compute planning, Nebius offers what hyperscalers won't: published GPU rates, reservable GB300-class racks, and EU-or-US data residency. The strategic question is capacity allocation, not capability.”
Reserving GB300 NVL72 racks for a 2027 training roadmap is a supply bet, not a product bet. Nebius passes the supply test: data centers in Finland, France and the US, an H100-to-B300 fleet already on the price list. Multi-month reservations cut on-demand rates by up to 35 percent, which is where multi-year planning actually pencils.
The integration surface stays on standards — Terraform provider, managed Kubernetes and Slurm, S3-compatible storage — so the path back to AWS or CoreWeave remains open. If the roadmap includes serving, the OpenAI-compatible inference API keeps application code portable too. That's the right shape for a three-year commitment.
However, the same Microsoft and Meta contracts that validate this vendor — $17.4 billion and up to $27 billion over five-year terms — will absorb enormous capacity. Reserve early and keep a second supplier qualified. Single-sourcing GPUs in this market is the one unforced error.
A top-tier neocloud alongside CoreWeave, separated by price transparency and an EU-US footprint rather than a moat.
Non-virtualized GPUs, InfiniBand fabric and managed Slurm map directly onto large-scale training programs.
Terraform, Kubernetes, Slurm and S3-compatible storage keep the stack on portable standards.
Anchor-tenant contracts validate the vendor but will compete for the same future capacity you want reserved.
Reservable GB300 NVL72 racks and multi-month terms support genuine multi-year capacity planning.
Infrastructure leaders planning multi-year training capacity across EU and US regions.
Organizations that mandate a single hyperscaler for all workloads.
Published GPU rates and a $25 entry make this the rare procurement-friendly AI cloud.
“Nebius publishes every GPU rate and starts billing at a $25 minimum first payment. The 35% reserved discount is real money but sits behind a sales call.”
Every number is on the pricing page. On-demand H100s run $3.85 per GPU-hour, preemptible drops to $2.15, and object storage is $0.0147 per GiB-month. Managed Kubernetes, networking and public IPs cost nothing on top.
One scenario worth modeling: 64 H100s × $3.85 × 730 hours is roughly $180K a month on-demand. Multi-month reservation takes up to 35% off — call it $117K. Preemptible cuts about 44% for interruption-tolerant jobs. Lambda quotes similar sticker rates; the difference is Nebius publishes its discount structure.
The catch: the reserved discount sits behind a sales conversation, and GB300-class racks are sales-only. Entry runs the other direction: $25 minimum first payment, usage-based billing, no seat licenses anywhere. Procurement clears this one fast.
A $25 minimum first payment and usage-based billing clear procurement without seat-count negotiations.
On-demand and preemptible need no commitment, but the 35% reserved discount requires multi-month terms via sales.
Every rate is public, from $3.85 H100 hours to $0.0147 per GiB-month storage.
Per-GPU-hour and per-token meters make cost-per-training-run directly computable.
Free managed Kubernetes, networking and public IPs strip out line items hyperscalers charge for.
Finance teams that want GPU costs modeled from published rates.
Buyers who need every discount available without a sales cycle.
Non-virtualized GPUs, managed Slurm, and honest preemptible rates cover most of what cluster work needs.
“Nebius ships the pieces cluster engineers actually touch: managed Slurm, GPU-aware Kubernetes, InfiniBand fabric, and a WEKA option for training I/O. Preemptible H100s at $2.15 reward teams with real checkpointing discipline.”
Non-virtualized GPUs on InfiniBand is the detail that matters — no hypervisor between NCCL and the fabric when a 512-GPU job starts timing out. Managed Slurm ships as a first-class deployment next to GPU-aware Kubernetes, the same call CoreWeave made with SUNK, and the right one for training shops.
Storage covers the actual workflow: S3-compatible buckets for datasets, shared filesystems or WEKA for checkpoint I/O, Managed MLflow for run tracking. A Terraform provider and CLI keep cluster config in git instead of console clicks. That's operability, not demo material.
The yellow flag is preemptible economics: $2.15 H100s are only cheap if checkpoint-restart is muscle memory, because interruptible means interrupted. Public docs on multi-node failure modes also read thinner than the hardware story. Budget on-demand at $3.85 until the pipeline survives a kill test.
Managed Slurm and GPU-aware Kubernetes mean day three is job scheduling, not cluster assembly.
Provisioning docs are solid; multi-node failure-mode guidance looks thinner based on what's public.
Preemptible interruptions and sales-gated flagship hardware are the main fights; the fabric stays out of the way.
Non-virtualized GPUs, WEKA filesystems and a gRPC API give advanced teams real knobs to turn.
Terraform provider, CLI, Managed MLflow and S3-compatible storage slot into existing training pipelines.
ML engineers who run distributed training on Slurm or Kubernetes.
Teams that lack checkpointing discipline for preemptible capacity.
The console respects your time, which is not something GPU clouds usually bother with.
“Zero-to-cluster in minutes is the pitch, and the $25 entry means you can test it without a meeting. Nebius Echo, an AI helper inside the console, suggests someone thought about the person clicking around at 11pm.”
A $1 trial credit sounds like a gag until you notice it's for per-token inference, where a dollar buys a real test drive. Same energy on compute — $25 minimum first payment, no sales call, clusters up in minutes per their pitch. Small numbers that tell you who this was built for.
Nebius Echo is the detail that stands out: an AI agent inside the console for managing your own infrastructure. That's the difference between fixing a quota problem at 11pm and filing a ticket. Together AI feels comparable on the inference playground side; the full cloud console is where Nebius carries more.
The catch: everything is web-only, so there's no mobile way to check whether a long training run is still alive. Docs are engineer-first, and Kubernetes newcomers will feel the learning curve for a week or two. For a GPU cloud, that's a short list of complaints.
Pre-built JupyterLab and vLLM apps plus an in-console AI agent show attention to the actual user.
Fine for anyone who knows Kubernetes or Slurm; steeper for teams new to cluster tooling.
Web-only, which is the category norm for GPU clouds; scored neutral.
A $25 entry and $1 Token Factory trial credit make the first hour fully self-serve.
A 99.9% uptime SLA on dedicated endpoints and Nasdaq-grade operations read as dependable.
Hands-on builders who want GPU infrastructure without a sales call.
People who need mobile access to monitor long-running jobs.
Real revenue and real customers, with two names carrying most of the story.
“The Microsoft and Meta contracts are verifiable, the pricing is public, and the hardware is real. Customer concentration and a crowded neocloud field keep the score honest.”
Two customers — Microsoft at $17.4 billion, Meta at up to $27 billion — anchor most of the contracted backlog. Great validation. Also a dependency: if either slows at renewal, the growth story compresses fast.
Credit where due. Rates are published down to $0.74 an hour for preemptible L40S, which most GPU clouds still won't do. The exit path is honest too — OpenAI-compatible API, Terraform provider, S3-compatible storage. Leaving is config work, not a rewrite.
The differentiation question stays open, however. CoreWeave, Lambda and Crusoe sell the same NVIDIA racks, and 'AI-native cloud' is the label all of them use. What Nebius has is execution plus Nasdaq-grade disclosure since October 2024. Based on what's visible, that holds. Watch the renewal cycle.
CoreWeave, Lambda and Crusoe sell the same NVIDIA racks; transparency is the main separator.
OpenAI-compatible API, Terraform and S3-compatible storage make leaving config work, not a rewrite.
Multi-year contracted backlog is real but concentrated in two customers with renewal risk.
'AI-native cloud' is generic, but published rates and named certifications back most claims.
Nasdaq disclosure since October 2024 is solid; the operating history under the Nebius name is short.
Teams that value published pricing and a verifiable exit path.
Buyers who need differentiation beyond execution and transparency.
Common questions answered by our AI research team
On-demand pricing is $3.85 per GPU-hour for HGX H100, $4.50 for H200 and $7.15 for B200, while preemptible L40S capacity starts at $0.74 per GPU-hour. Reserving clusters for multi-month terms saves up to 35%, and the minimum first payment is $25.
Token Factory serves 60+ open-source models, including DeepSeek-V4-Pro, Qwen3-235B-A22B, GPT OSS 120B and 20B, Kimi-K2.6 and GLM-5.1, plus embedding models. You can also fine-tune models with LoRA and deploy your own checkpoints to its endpoints.
Yes. Nebius Token Factory holds SOC 2 Type II, HIPAA and ISO 27001 certifications, and offers a zero-retention mode where requests and outputs are never stored. Data centers in Finland, France and the US support EU or US data residency.
Nebius deploys from zero to clusters in minutes through its web console, with managed Kubernetes or Slurm clusters for orchestration. Infrastructure can also be provisioned via the CLI or Terraform provider, and Managed Kubernetes itself is free.
Yes. Token Factory exposes an OpenAI-compatible API, so existing OpenAI client libraries work by changing the base URL and key. Cloud resources are managed with an official Terraform provider, a CLI and a gRPC API, plus GPU-aware managed Kubernetes.
Company
Nebius Group N.V.Founded
2024Pricing
From $1/moFree Trial
Available




AI cloud infrastructure company based in Amsterdam, operating GPU data centers for training and running AI models. Listed on Nasdaq as NBIS.