AI inference infrastructure built on custom RDU chips
SambaNova is an AI inference platform for running large language models and agentic AI workloads.
AI Panel Score
6 AI reviews
Reviewed
AI Editor ApprovedApproved and published by our AI Editor-in-Chief after full panel analysis.Working with SambaNova typically means choosing between two access paths: calling models through SambaCloud, a serverless inference service with OpenAI-compatible APIs, or deploying SambaStack, a full-stack combination of RDU hardware and software installed on-premises or in a private cloud. In either case, the workflow centers on sending inference requests against pre-loaded open-source models and getting responses back, with SambaNova's dataflow architecture handling the underlying compute rather than general-purpose GPU infrastructure.
The distinguishing element is the Reconfigurable Dataflow Unit (RDU), a chip SambaNova designed specifically for inference rather than adapting general-purpose hardware. The company markets the RDU on tokens-per-watt efficiency and speed for agentic AI workloads, and has published benchmark claims around serving large models such as DeepSeek-R1 671B. SambaManaged extends this into a managed infrastructure service for data centers that want the hardware operated on their behalf, and SambaRack packages the hardware (including the SN50 RDU generation) at rack scale for organizations building their own data center capacity.
SambaNova positions its sovereign AI offering for governments and regions that want AI infrastructure hosted within their own borders, with named data center partners including Argyll in the UK, Infercom and OVHcloud in the EU, and SouthernCrossAI in Australia. Beyond sovereign deployments, its customer base includes enterprises, national research labs (such as Argonne, Oak Ridge, and Lawrence Livermore), and telecom operators like SoftBank. Pricing for SambaCloud is usage-based per published rate cards, while SambaStack and SambaRack deployments are sold on a contact/quote basis; direct competitors in AI inference infrastructure include Groq, Cerebras, and NVIDIA-based cloud inference providers.
SambaCloud exposes an OpenAI-compatible API, allowing existing applications built against OpenAI's API format to be pointed at SambaNova's endpoints with minimal changes. Documentation, quickstarts, and API references are published at docs.sambanova.ai, and the platform is also available through AWS Marketplace.
SambaNova's purpose-built AI chip designed for agentic inference, optimized for speed and tokens-per-watt efficiency.
A dedicated forum where developers can discuss, troubleshoot, and share knowledge about building on SambaNova's platform.
A dataflow processing approach underlying SambaNova's chips and software stack that structures how computation is executed.
Serverless AI inference cloud service offering OpenAI-compatible APIs for fast inference on open-source language models.
Rack-scale hardware system for AI inference workloads, including the SN50 RDU generation.
Full-stack enterprise AI platform combining RDU hardware and software, deployable on-premises or in the cloud.
Provides an API interface compatible with OpenAI's format so developers can integrate SambaCloud inference into existing applications.
Enables nations and regions to host AI inference infrastructure within their own borders via partners like Argyll (UK), Infercom (EU), OVHcloud (EU), and SouthernCrossAI (Australia).
Quickstarts, API reference material, and integration guides for building applications on SambaNova's platform.
Gives developers early access to new SambaNova models and platform capabilities before general release.
Free short-video training courses covering AI fundamentals, dataflow concepts, and how to get started with SambaCloud.
Managed AI infrastructure service that operates and maintains data center deployments on behalf of customers.
For developers testing SambaNova Cloud's API without a payment method; suited for experimentation and small-scale prototyping.
Pay-as-you-go plan billed per million tokens, with prices varying by model — for example $0.22/$0.59 per 1M input/output tokens for gpt-oss-120b, $0.60/$1.20 for Meta-Llama-3.3-70B-Instruct, and $3.00/$4.50 for DeepSeek-V3.1. No single flat monthly price exists since SambaNova Cloud is priced per model per million tokens rather than as a flat plan fee.
For enterprises and data centers needing dedicated capacity, higher throughput, and negotiated contract terms. Specific pricing for this plan is not publicly listed and requires contacting SambaNova's sales team.
Custom silicon, real customers, but a hardware bet you're making alongside them.
“SambaNova has national labs and telecom customers, plus an OpenAI-compatible API that lowers switching cost. The RDU chip is a genuine differentiator, but betting on non-GPU silicon carries its own risk.”
Argonne, Oak Ridge, Lawrence Livermore, SoftBank. That's a real customer list, not logos on a landing page.
Two things stand out. One: the RDU chip is a legitimate architectural bet against Nvidia's GPU dominance, and DeepSeek-V3.1 671B benchmarks back the tokens-per-watt claim. Two: sovereign AI partnerships in the UK, EU, and Australia give this a government-procurement angle Groq and Cerebras don't emphasize as hard.
The tradeoff: custom silicon vendors have a mixed survival record. Betting infrastructure on RDUs instead of GPUs means less community tooling, fewer engineers who've touched it, more vendor lock-in if SambaNova stumbles.
$0.22/$0.59 per 1M tokens at the entry end is competitive. Free tier removes the pilot-cost objection entirely.
Groq and Cerebras chase the same inference-speed niche; SambaNova's sovereign-AI angle is a differentiator but unproven at broad market scale.
Government sovereign AI deployments read as credible; niche chip architecture may raise eyebrows versus safer Nvidia-based choices.
OpenAI-compatible API and free $5 signup credit mean integration in minutes per their own docs.
Custom RDU architecture and agentic inference focus advance capability, not just cost-cutting on existing GPU spend.
National lab and telecom customers plus multi-region sovereign partnerships signal staying power, though no public funding data that I could find.
Teams running agentic or open-source LLM workloads who want inference speed and sovereign deployment options.
Avoid if your stack depends on GPU-specific tooling or you can't tolerate architecture lock-in risk.
Custom silicon buys real inference speed, but you're betting on RDU as a durable second architecture next to GPUs.
“SambaNova's OpenAI-compatible API lowers migration friction from day one. The three-year question is whether the RDU ecosystem matures as fast as CUDA's alternatives, or whether you're stuck with a single vendor's roadmap.”
The OpenAI-compatible API on SambaCloud is the right call architecturally — point your existing client at a new endpoint, no rewrite. That de-risks adoption at the prototype stage. Pricing per model (gpt-oss-120b at $0.22/$0.59 per 1M tokens, DeepSeek-V3.1 at $3/$4.50) is transparent enough to model costs before committing.
The real bet is the RDU. If dataflow architecture holds its tokens-per-watt advantage on frontier models like DeepSeek-V3.1 671B, you get durable cost and latency wins over GPU-based Groq or Cerebras. If it doesn't keep pace, you're locked into SambaStack's on-prem hardware with a much smaller ecosystem than NVIDIA's.
Sovereign deployment partners (Argyll, OVHcloud, SouthernCrossAI) are a genuine differentiator for regulated workloads. But SambaRack and SambaStack pricing is quote-only — no public rate card means real procurement lock-in risk before you've seen total cost of ownership.
Named research lab customers (Argonne, Oak Ridge, Lawrence Livermore) and sovereign AI deals with national partners signal credible enterprise traction against Groq and Cerebras.
OpenAI-compatible API matches how teams already build, letting existing tooling port over with minimal rewrite.
AWS Marketplace availability plus OpenAI API compatibility eases entry into existing cloud and app stacks.
SambaStack on-prem deployment ties you to RDU hardware refresh cycles (SN40 to SN50) outside the GPU ecosystem's gravity.
Purpose-built RDU chip with published benchmarks on DeepSeek-R1 671B shows genuine hardware differentiation, not just repackaged GPU cloud.
Enterprises and government buyers who need data-sovereign inference and can tolerate a custom-silicon vendor relationship.
You need a GPU-portable stack with a large existing tooling ecosystem and no appetite for quote-based procurement.
Per-token pricing is honest. Enterprise tier is not — that's the tax on real usage.
“Free tier and Developer tier rate cards are public, down to $0.22/$0.59 per 1M tokens on gpt-oss-120b. Scale up, and you hit Enterprise, quote-only, same as Groq and Cerebras.”
$0.22/$0.59 per 1M input/output tokens at the entry end. $3/$4.50 for DeepSeek-V3.1. Real numbers, published, no sales call. Developer tier gives $5 free credit, expiring in 3 months. Fine for prototyping, thin for a real eval cycle.
Team running agentic workloads at volume lands in Enterprise: dedicated capacity, negotiated contract, custom pricing. That's where the TCO story breaks down — no rate card, no way to model year-1 spend without a rep. SambaStack and SambaRack are quote-only too, standard for on-prem hardware deals but still a black box.
Compare to Groq and Cerebras — same pattern, quote-gated at scale. NVIDIA-based cloud inference at least has more public benchmarking to anchor cost-per-token claims. SambaNova's OpenAI-compatible API keeps migration cost low. That's a real number: near-zero switching cost if you're already on OpenAI's format.
Available via AWS Marketplace for usage-based billing, cutting procurement friction versus a pure direct-sale model.
No published term length or auto-renewal data for Enterprise/SambaStack; usage-based SambaCloud tiers imply month-to-month flexibility.
Free and Developer tiers publish per-model token rates; Enterprise, SambaStack, SambaRack are quote-only.
Tokens-per-watt and DeepSeek-R1 671B benchmark claims give a measurable axis, though third-party verification is absent.
Token costs are cheap at small scale but Enterprise contract terms and hardware deployment costs are invisible pre-sales-call.
Teams prototyping on open-source models who can measure cost per million tokens directly.
You need a fixed 3-year budget number before talking to a sales rep.
OpenAI-compatible endpoints and real token pricing, but RDU quirks will surface by week two.
“SambaCloud is a drop-in swap for OpenAI-format apps with published per-model token pricing instead of opaque tiers. The real question is what breaks when you hit the RDU's dataflow architecture instead of a GPU stack you already know how to debug.”
Pointing an existing OpenAI client at SambaCloud takes minutes per their own claim, and the rate card is refreshingly explicit: $0.22/$0.59 per 1M tokens at the entry end, $3/$4.50 on DeepSeek-V3.1. That's the kind of pricing transparency Groq and Cerebras customers will recognize and appreciate.
Day-3 reality is murkier. Free tier gives you $5 credit expiring in 3 months — fine for a prototype, tight for a real evaluation cycle. Rate limits scale with tier (20 RPM free, 240 on Llama 3.3 70B once a card's on file), which means load-testing before committing spend is mandatory, not optional.
The RDU architecture is the actual unknown. Dataflow chips behave differently under load than GPU clusters everyone's already debugged for years. Docs cover quickstarts and API reference, but nothing I could find suggests deep troubleshooting content for dataflow-specific failure modes. SambaOrchestrator's auto-scaling and load balancing sound solid on paper — untested against NVIDIA-based providers' maturity.
OpenAI-compatible API lowers initial friction, but RDU-specific behavior under sustained load is unproven from the outside.
Quickstarts and API reference exist at docs.sambanova.ai but no changelog, suggesting thinner ongoing practitioner content.
$5 credit expiring after 3 months and tiered rate limits (e.g. +50% RPM) force early capacity planning.
SambaOrchestrator, SambaManaged, and SambaRack SN50 give a real scaling path from playground to rack-scale deployment.
OpenAI-compatible API format means minimal code changes for teams already building against that spec.
Teams already on OpenAI's API format who want cheaper per-token pricing and a sovereign-deployment option.
You need battle-tested GPU-stack tooling and can't risk time debugging an unfamiliar dataflow architecture.
Fast chips, real docs, but this is a developer tool wearing an enterprise suit.
“SambaNova sells speed and sovereignty, backed by an OpenAI-compatible API that's genuinely easy to drop in. The day-to-day feel is more spec sheet than product, though.”
First ten minutes here is signup, $5 free credit, and a playground — that part's fine, no payment method needed, and if you've used an OpenAI-style API before you're not relearning anything. The OpenAI-compatible endpoint is the real onboarding win. Point your existing app at it, swap a key, done.
But this isn't a polished daily-use product in the Groq or Cerebras sense of 'watch tokens fly and feel good about it.' It's infrastructure. Pricing is per-model-per-million-tokens, which means every time you switch from a small model to a frontier one you're doing math again — $0.22/$0.59 versus $3/$4.50. That's honest pricing, not friendly pricing.
Mobile doesn't exist here, which is fine — nobody's running inference from their phone — but it means there's no casual-check-in layer, just dashboards and docs. The learning curve stretches once you go from SambaCloud to SambaStack or SambaRack, where 'quote basis' means sales calls, not self-serve. Fine for enterprise buyers. Rough if you wanted a quick weekend project.
Docs, playground, and SambaAcademy videos exist, but no changelog listed means unclear ongoing craft investment.
Easy to start with SambaCloud, but SambaStack/SambaRack/SambaOrchestrator stack adds real complexity at scale.
Platforms listed as web-only; category norm for inference APIs, but a real gap for on-the-go monitoring.
Free tier with no payment method and OpenAI-compatible API make first integration fast.
Named national lab and telecom customers (Argonne, SoftBank) suggest production-grade trust, though no public uptime data given.
Developers and enterprises who want fast open-source model inference without building on raw GPU infrastructure.
You want a self-serve path into on-premises deployment without talking to a sales team.
Custom silicon is a real bet. Custom silicon is also how Cerebras and Groq keep score.
“SambaNova's RDU chip is a genuine differentiator, not vaporware — but so was every custom-silicon story before the market consolidated around GPUs anyway. The OpenAI-compatible API is the safety net here.”
"Fastest AI inference platform" is the H1. Every inference vendor says this — Groq says it, Cerebras says it. The DeepSeek-R1 671B benchmark claims are specific enough to be checkable, which is more than most.
The pattern worry: custom-chip AI infra has a graveyard. Graphcore didn't make it. Habana got folded into Intel and quieted down. SambaNova's national research lab customers (Argonne, Oak Ridge) and sovereign deals in the UK and Australia are real anchors, not vapor — that's the strongest signal here.
Exit is the good news. OpenAI-compatible API means SambaCloud migration is genuinely low-friction — swap the endpoint, keep the code. SambaStack, the on-prem full-stack version, is the opposite: hardware lock-in, quote-only pricing, no clean exit once racks are installed.
$5 in free credits expiring after 3 months is thin for real evaluation.
RDU architecture is a real technical wedge vs. GPU-based providers, but Groq and Cerebras claim the same speed lane.
OpenAI-compatible API on SambaCloud is low-friction; SambaStack on-prem hardware is not.
Sovereign deals with named partners (Argyll, OVHcloud, SouthernCrossAI) and lab customers suggest funding runway; no public funding figures given.
"Fastest" superlative is unverified in isolation, but benchmark specifics (DeepSeek-R1 671B) give it something to check against.
Custom-silicon AI infra has a mixed history (Graphcore, Habana); national lab and telecom customers are a genuine counter-signal.
Enterprises or governments needing sovereign, on-prem inference who can tolerate quote-based pricing and hardware lock-in.
You want transparent pricing and zero lock-in risk from day one.
Common questions answered by our AI research team
Yes. SambaNova's APIs are OpenAI compatible, letting you port your application to SambaNova in minutes.
SN40-16 is the fourth-generation system optimized for low-power inference (average 10 kWh) and running many models simultaneously. SN50 is the fifth-generation system optimized for fast agentic inference at lower cost while running the largest models like gpt-oss-120b and DeepSeek.
Yes. SambaStack is a deployable full-stack platform for on-premises or private cloud use, and SambaNova powers sovereign AI data center partners in the UK, EU, and Australia to keep AI within national borders.
SambaCloud supports MiniMax M2.7, DeepSeek models including the 671-billion-parameter DeepSeek-V3.1, Meta's Llama 3.3 70B Instruct, and OpenAI's gpt-oss-120b, plus MiniMax-M3, DeepSeek-V3.2, and Google's gemma-4-31B-it in preview.
SambaOrchestrator simplifies managing AI workloads across data centers, letting teams monitor and manage model deployments and scale automatically to meet user demand, alongside features like auto scaling, load balancing, monitoring, and model management.
Company
SambaNova Systems, Inc.Founded
2017Pricing
Usage-basedFree Plan
Available




SambaNova Systems is a Palo Alto-based company that designs custom AI chips and full-stack systems for enterprise generative AI training and inference.