Energy-first AI cloud with on-demand NVIDIA and AMD GPUs and managed inference
Crusoe is an AI cloud platform for teams training and serving large models on NVIDIA and AMD GPUs powered by clean energy.
AI Panel Score
6 AI reviews
Reviewed
AI Editor ApprovedApproved and published by our AI Editor-in-Chief after full panel analysis.Crusoe is an energy-first AI cloud platform providing GPU compute, orchestration, and managed inference for teams building and serving large AI models. It serves AI labs, startups, and enterprises running training and inference at scale. Pricing is usage-based: on-demand GPUs start at $1.50 per GPU-hour for the NVIDIA L40S, H100 instances cost $3.90 per GPU-hour, Managed Inference starts at $0.05 per 1M input tokens, and reserved capacity is quote-based. Core capabilities include an OpenAI-compatible Managed Inference API, MemoryAlloy KV cache technology delivering up to 9.9x faster time-to-first-token, Managed Kubernetes and Slurm orchestration, and fault-tolerant AutoClusters for training. Data centers in Texas, Nevada, Virginia, Iceland, and Norway run on renewable energy, with 99.98% uptime and ISO 27001 and ISO 42001 certifications. Crusoe best fits teams that need large dedicated GPU clusters without hyperscaler lock-in; alternatives include CoreWeave, Lambda, Nebius, and Together AI.
Teams provision GPU instances on Crusoe Cloud through the console or the Terraform provider, choosing from NVIDIA GB200 NVL72, HGX B200, H200, H100, A100, and L40S or AMD MI355X and MI300X hardware. Training workloads run on clusters orchestrated with Crusoe Managed Kubernetes or Managed Slurm, while production models are served either on owned instances or through Crusoe Managed Inference, a pay-per-token API hosting open models such as DeepSeek V3, Kimi K2.6, Llama 3.3 70B, Qwen3 235B, and GPT-OSS 120B.
Several capabilities distinguish the platform. MemoryAlloy, Crusoe's proprietary KV cache technology, powers an OpenAI-compatible inference engine with time-to-first-token up to 9.9x faster. AutoClusters provide fault-tolerant, self-healing GPU clusters for long training runs, and Crusoe Command Center gives a unified operations view across infrastructure. The company also offers Crusoe Edge Zones for inference close to end users, Crusoe Spark modular AI data centers, S3-compatible object storage, persistent and shared disks, and VPC networking with InfiniBand interconnects, all with $0 network ingress and egress charges.
Crusoe fits AI labs, startups, and enterprises that need large GPU capacity without hyperscaler lock-in; customers include Windsurf, Cursor, HeyGen, Together, and Modal. Pricing is usage-based, with on-demand GPUs from $1.50 per GPU-hour, spot pricing for fault-tolerant workloads, and reserved capacity contracts for guaranteed scale. It competes with GPU cloud providers such as CoreWeave, Lambda, and Nebius, and with inference platforms like Together AI and Fireworks AI.
Technically, the platform is built around open architecture: an OpenAI-compatible inference API, a documented REST API, Terraform provider, GitHub repositories, and a cookbook of code examples. Data centers in Texas, Nevada, Virginia, Iceland, and Norway run with 100% renewable energy coverage, and the platform claims 99.98% uptime backed by ISO 27001 and ISO 42001 certifications.
Pay-per-token inference service hosting open models like DeepSeek V3, Kimi K2.6, Llama 3.3 70B, and GPT-OSS 120B, with support for bringing your own fine-tuned model.
On-demand access to NVIDIA GB200 NVL72, HGX B200, H200, H100, A100, and L40S plus AMD MI355X and MI300X accelerators.
High-performance edge infrastructure that places inference capacity closer to end users for lower latency.
Modular AI data centers that can be deployed rapidly to add capacity where power is available.
Official Terraform provider, REST API references, GitHub repositories, and a code cookbook for infrastructure-as-code automation.
VPC networks, InfiniBand cluster interconnects, and load balancers with $0 ingress and egress charges.
Unified operations platform for monitoring and managing AI infrastructure across the Crusoe environment.
Managed Kubernetes control plane for running GPU clusters, billed at $0.10 per cluster-hour.
Managed Slurm scheduling for batch training jobs on large GPU clusters.
Proprietary KV cache technology behind Crusoe's inference engine that delivers time-to-first-token up to 9.9x faster.
Fault-tolerant, self-healing GPU clusters that keep long training runs going through hardware failures.
Object storage at $0.06 per GiB-month alongside persistent disks, shared disks, and a container registry.
Pay-as-you-go GPU and CPU compute billed hourly with no commitment, starting at $1.50 per GPU-hour.
Pay-per-token model serving for teams that want inference without managing GPUs, from $0.05 per 1M input tokens.
Discounted preemptible capacity for fault-tolerant and batch workloads.
Custom contracts for teams that need guaranteed GPU capacity at scale.
The startups spending most on GPUs already chose Crusoe, and that shortens the board conversation.
“Crusoe pairs on-demand NVIDIA and AMD GPUs with owned clean power and a $10B-plus valuation behind it. Strong pilot candidate for teams that want hyperscaler-grade capacity without hyperscaler contracts.”
Windsurf, Cursor, and HeyGen already run production workloads on Crusoe. When AI-native startups vote with their training budgets, my diligence gets shorter. The $1.375B Series E in October 2025 — valuation north of $10B — settles the runway question for a 36-month bet.
The differentiation isn't the GPU menu — everyone rents H100s now. It's AutoClusters keeping week-long training runs alive through hardware failures, plus owned power generation behind the data centers. Nebius and Lambda compete on hourly price; Crusoe's moat is megawatts, and boards understand energy.
The catch: reserved capacity and the newest silicon route through sales, and they're pouring capex into a handful of giant anchor-tenant campuses — concentration risk if one tenant sneezes. Pilot on-demand first, H100s at $3.90/GPU-hour with zero egress fees. Standardize after a 90-day run holds and the reserved quote beats it.
Owned power differentiates against Nebius, Lambda, and CoreWeave, which compete mostly on price.
ISO 27001 and ISO 42001 certifications plus a 99.98% uptime record keep board questions short.
On-demand instances and an OpenAI-compatible endpoint deliver without a sales cycle, though reserved capacity needs one.
Purpose-built GPU cloud with Kubernetes, Slurm, and inference covers train-to-serve without hyperscaler contracts.
A $1.375B Series E at a $10B-plus valuation, with Windsurf, Cursor, and HeyGen in production.
Infrastructure buyers who need large GPU capacity without hyperscaler lock-in.
Teams who need a broad managed-services catalog beyond GPU compute.
Crusoe turns energy ownership into the strategic edge hyperscalers can't easily copy.
“Vertical integration from power generation to an OpenAI-compatible API makes Crusoe a credible neocloud anchor for training and inference. Its open integration surface keeps the three-year lock-in risk unusually low.”
Megawatts, not GPUs, are the scarce input for the next decade, and Crusoe is the one neocloud in this evaluation that owns its power. Crusoe Spark modular data centers deploy wherever power is available; the Abilene campus is sized at 1.2 gigawatts. That's a supply curve AWS negotiates for — Crusoe builds it.
The integration surface stays open on purpose: Terraform provider, S3-compatible storage, an OpenAI-compatible inference API. Adoption doesn't paint us into a corner; switching costs stay measured in weeks. Five regions — Texas, Nevada, Virginia, Iceland, Norway — cover most latency and residency maps.
The tradeoff is service depth. AWS surrounds its GPUs with hundreds of managed services; Crusoe brings Kubernetes, Slurm, and Managed Inference — enough for a capable platform team, thin for an org that expects IAM depth and a marketplace ecosystem. Slot it as the training-and-inference substrate beside a hyperscaler, not as a replacement.
Energy-first vertical integration separates it from Nebius, Lambda, and other GPU renters.
GB200 NVL72 through MI300X coverage with Slurm and Kubernetes matches how large training orgs operate.
Terraform provider, REST API, and OpenAI-compatible endpoints slot into existing IaC and serving stacks.
Open APIs and S3-compatible storage keep exit costs low, but anchor-tenant capex concentrates risk.
Owning generation, Crusoe Spark modular data centers, and a 1.2GW campus is systems-level positioning, not reselling.
Platform organizations who run large training and inference fleets on open tooling.
Enterprises who expect a full managed-services ecosystem around their compute.
Zero egress and published GPU rates make Crusoe's math unusually easy to model.
“On-demand rates are public and competitive, with $0 network egress removing the classic hidden cost. Spot and reserved pricing require a sales conversation, which is where the real discounts sit.”
Egress is where GPU clouds claw the discount back. Crusoe charges $0 — ingress and egress both. For checkpoint-heavy training, that line item alone reshapes the model.
List rates sit mid-market: H100 $3.90/GPU-hr, H200 $4.29, MI300X $3.45, L40S from $1.50. Managed Kubernetes adds $0.10 per cluster-hour — rounding error. Sample math: 64 H100s × $3.90 × 730 hours = $182K a month on-demand, before reserved discounts. Managed Inference runs $0.05-$1.74 per 1M input tokens; price it against Together AI per model before committing traffic.
Two gaps, however. Spot and reserved rates aren't published — the deepest discounts live behind a sales call — and there's no free tier or trial, so you model from list prices you can't test at $0. Procurement will also want the auto-renewal terms in writing; the pricing page doesn't say.
Usage-based billing is simple, but enterprise terms require sales engagement.
Hourly on-demand carries no commitment; reserved capacity is custom-contract only.
On-demand and per-token rates are published, but spot and reserved pricing aren't.
Per-GPU-hour and per-token meters map cleanly onto workload cost models.
$0 ingress and egress plus $0.10 cluster-hours strip out the usual ancillary line items.
Finance teams who want GPU spend modeled from published per-hour rates.
Budget owners who need committed pricing without entering a sales cycle.
Self-healing clusters and real Slurm support show Crusoe was built by people who run GPUs.
“AutoClusters, Managed Slurm, and a Terraform provider cover the workflows platform engineers actually operate. The gaps are evaluation friction and an inference catalog limited to open models.”
Dead nodes at hour 70 of a training run — that's the tax every GPU platform team budgets for. AutoClusters is Crusoe's answer: self-healing clusters that swap failed hardware without killing the job. Whether pagers fire at 3 a.m. depends on exactly this layer.
Orchestration ships in both flavors real platforms run: Managed Slurm for batch training, Managed Kubernetes at $0.10 per cluster-hour for serving. CoreWeave's SUNK covers similar ground; Crusoe's Terraform provider and cookbook repos suggest a team that scripts everything. Serving through the OpenAI-compatible endpoint means changing a base URL, not rewriting clients.
The friction sits at the edges, however. No free trial means the first smoke test costs real money, and the Managed Inference catalog is open-weights only — DeepSeek V3, Llama 3.3 70B, Kimi K2.6 — so closed-model traffic still routes elsewhere. Your own fine-tuned weights do deploy, which softens it.
AutoClusters and managed orchestration remove the failure-handling toil that dominates cluster operations.
REST references, a cookbook of code examples, and GitHub repos target engineers, not buyers.
No trial and sales-gated reserved capacity add friction at evaluation and scale-up.
InfiniBand interconnects, spot instances, and BYO fine-tuned models reward advanced teams.
Slurm, Kubernetes, Terraform, and an OpenAI-compatible API slot into standard MLOps stacks.
ML platform engineers who operate large training clusters with Slurm or Kubernetes.
Teams who need proprietary frontier models behind one inference endpoint.
A GPU cloud that won't charge you for leaving earns trust fast.
“Zero egress fees and a single-pane Command Center make the daily experience feel respectful. Just know there's no free tier, so kicking the tires costs actual money.”
GPU clouds usually lose people at the invoice, not the terminal. Crusoe charging $0 for egress means you can pull checkpoints and datasets out without the exit toll — rare enough to notice, boring enough to love.
Crusoe Command Center puts the whole fleet on one screen instead of the tab-farm most platforms make you run. The docs lead with a cookbook of runnable examples, not marketing prose. RunPod is cheaper for hobby jobs but scruffier; Lambda feels closest in spirit, minus the owned-power story.
The catch is the front door. No free tier, no trial — your first impression costs money, with L40S time starting at $1.50 an hour. Console's web-only, which is fine; nobody manages an H100 cluster from a phone. Three months in, the win here looks like boring reliability: 99.98% uptime, if it holds.
Command Center consolidation and cookbook-first docs point to a team that uses its own console.
Standard Kubernetes, Slurm, and OpenAI-compatible APIs mean existing skills transfer directly.
Web-only console is the right call for GPU infrastructure; mobile isn't a real use case.
No free plan or trial; the first hour costs money.
A 99.98% uptime record plus self-healing AutoClusters is a calm-pager profile.
Hands-on builders who run steady GPU workloads and hate surprise bills.
Tinkerers who want a free tier for weekend experiments.
Real infrastructure and an honest exit path, wrapped in a few asterisk-worthy claims.
“Crusoe's open APIs and owned power make it more durable than the average GPU reseller. The pivot history and benchmark-style marketing claims still warrant watching.”
A pivot this clean deserves a second look. Crusoe mined bitcoin until March 2025, then sold that entire business to NYDIG and went all-in on AI cloud. It worked — the valuation passed $10B by October 2025 — but a company that exits one market fast can exit another.
The marketing has tells. 'Up to 9.9x faster' time-to-first-token from MemoryAlloy — 'up to' is doing heavy lifting. '100% renewable coverage' is an accounting construct, not a description of the grid. Neither is disqualifying; both earn an asterisk.
The exit story is genuinely good: OpenAI-compatible API, S3-compatible storage, Terraform — leaving costs a weekend, not a quarter. CoreWeave survived to a 2025 IPO; plenty of GPU resellers didn't, and Crusoe owning its power is a real difference. The yellow flag is capex concentrated around a few anchor tenants. Watch that.
Owned power generation is a real moat that CoreWeave and Lambda don't have.
OpenAI-compatible API, S3-compatible storage, Terraform, and $0 egress make leaving cheap.
A $10B valuation and Series E cash are real, but capex-heavy anchor-tenant bets concentrate risk.
Published pricing is straight, but 'up to 9.9x' and '100% renewable coverage' are asterisk phrases.
Named customers and a 99.98% uptime claim are strong, but the pure AI-cloud record only dates to 2025.
Cautious buyers who value cheap exit paths over managed-service breadth.
Risk-averse teams who require a decade of single-market vendor history.
Common questions answered by our AI research team
On-demand GPUs start at $1.50 per GPU-hour for the NVIDIA L40S, with H100 at $3.90 and H200 at $4.29 per GPU-hour. AMD MI300X runs $3.45 per GPU-hour, and network ingress and egress are free.
Crusoe Managed Inference serves open models including DeepSeek V3, Kimi K2.6, Llama 3.3 70B, Qwen3 235B, and GPT-OSS 120B, and supports bringing your own fine-tuned model. Token pricing starts at $0.05 per 1M input tokens.
Yes. Crusoe Managed Kubernetes runs GPU clusters at $0.10 per cluster-hour, and Managed Slurm handles batch training job scheduling. AutoClusters add fault-tolerant, self-healing operation for long training runs.
Crusoe holds ISO 27001 and ISO 42001 certifications and publishes a trust center at trust.crusoe.ai. The platform reports 99.98% uptime with enterprise support, and customers like Windsurf cite that uptime on production clusters.
Yes. Managed Inference exposes an OpenAI-compatible API backed by MemoryAlloy KV cache technology, with time-to-first-token up to 9.9x faster. A Terraform provider, REST API references, and a code cookbook support automation.
Crusoe builds and operates energy-first AI data centers and a GPU cloud for AI workloads. Founded in 2018, the company is headquartered in Denver, Colorado.