Together AI logo

Together AI Review

Visit

Open-source AI platform for building and deploying machine learning models

Together AI is a cloud platform for training, fine-tuning, and deploying open-source AI models.

Together AI·Founded 2022·From $0/moFree PlanFree TrialLLM PlatformsAI APIsAI CloudAI DevOps

AI Panel Score

8.0/10

6 AI reviews

Reviewed

AI Editor Approved

What is Together AI?

Together AI is a cloud platform for training, fine-tuning, and deploying open-source AI models. Developers use its infrastructure to run open-weight models in production through serverless inference, batch inference, dedicated model or container inference, and accelerated compute, with an OpenAI-compatible base URL that lets existing clients switch with a one-line change. Pricing is usage-based with a free plan and free trial; dedicated inference starts at $3.99 per hour, on-demand GPU clusters at $3.49 per hour, and standard fine-tuning at $0.48 per million tokens. Additional capabilities include managed storage, a code sandbox, and the Together Kernel Collection for workload-specific GPU optimization. TopReviewed's six-seat AI review panel scored it 8.0/10, praising batch inference that scales to 30 billion tokens per model while noting that the catalog's relevance depends on Meta, Mistral, and DeepSeek release cadence. It best fits AI platform teams running open-source models in production.

About Together AI

Together AI is a cloud-based platform that specializes in open-source artificial intelligence models and infrastructure. The platform provides developers and organizations with tools to train, fine-tune, and deploy various AI models without requiring extensive machine learning infrastructure expertise.

The platform offers access to a wide range of open-source models including large language models, image generation models, and other AI capabilities. Users can fine-tune these models on their own data or use pre-trained models through API endpoints. Together AI handles the underlying infrastructure, including GPU clusters and scaling requirements.

The service targets developers, AI researchers, and companies looking to integrate AI capabilities into their applications without building infrastructure from scratch. It competes with other AI platform providers by focusing specifically on open-source models rather than proprietary solutions.

Together AI offers both API access for inference and training capabilities for custom model development. The platform aims to make open-source AI models more accessible by providing managed infrastructure and simplified deployment options.

Features

AI

  • Batch Inference

    Runs asynchronous large-scale inference jobs at lower cost than real-time calls.

  • Dedicated Model Inference

    Reserves dedicated model capacity for production inference workloads.

  • Embeddings & Reranking

    Generates vector embeddings and reranks documents by relevance to power RAG pipelines.

  • Fine-tuning

    Supports full fine-tuning and LoRA adapters to train custom models on proprietary data.

  • Function Calling & JSON Mode

    Enables structured outputs, tool use, and guaranteed JSON-formatted model responses.

  • Sandbox

    Offers a secure code interpreter and execution environment for running AI-generated code.

  • Serverless Inference

    Provides pay-per-token API access to 200+ open-source and frontier AI models.

Analytics

  • Evaluations

    Benchmarks and evaluates model performance across tasks and datasets.

Core

  • Accelerated Compute / GPU Clusters

    Provisions on-demand H100/H200/GB200/B300 GPU clusters for training and inference.

  • Dedicated Container Inference

    Deploys custom containers on dedicated infrastructure for specialized workloads.

  • Instant Clusters

    Provisions on-demand GPU clusters with SLURM scheduling for training jobs.

  • Managed Storage

    Provides persistent storage designed for AI workloads and training data.

Integration

  • Framework & Tool Integrations

    Connects with frameworks like LangGraph, CrewAI, AutoGen, DSPy, Vercel AI SDK, and Composio for building agents and apps.

  • OpenAI-Compatible API

    Acts as a drop-in replacement for the OpenAI SDK to ease migration for developers.

Security

  • Enterprise Deployments

    Delivers private deployments, SLAs, and compliance features for enterprise customers.

Preview

Together AI desktop previewTogether AI mobile preview

Pricing Plans

Popular

Serverless Inference

Contact sales

Pay-per-token access to hundreds of open-source chat, vision, image, audio, video, transcription, embedding, rerank and moderation models; most teams start here before moving to dedicated capacity.

  • Price per 1M input/output tokens varies by model (e.g. DeepSeek V4 Flash $0.14 input)
  • Batch API discounted pricing available
  • Image models priced per image (e.g. FLUX.1 [schnell] $0.0027)
  • Video models priced per video (e.g. Kling 2.1 Standard $0.18)
  • Audio/transcription priced per minute or per 1M characters
  • No upfront commitment, pay as you use

Provisioned Throughput

Contact sales

Reserve dedicated capacity in PTUs (provisioned throughput units) for predictable, high-volume workloads with guaranteed tokens-per-minute.

  • Fixed capacity per PTU, TPM depends on model and token type
  • Estimate PTUs and costs via calculator
  • Cost savings vs. commercial model list pricing at scale
  • Input/cached/output TPM per PTU pricing

Dedicated Inference (On-Demand)

$4/hourly

Single-tenant GPU instances for deploying custom models with guaranteed performance, paid hourly on-demand per GPU.

  • Single-tenant GPU instances (no sharing)
  • Support for custom models
  • Autoscaling and traffic spike handling
  • NVIDIA HGX H100 from $3.99-$5.49/GPU/hr
  • NVIDIA HGX B200 from $8.19-$8.99/GPU/hr

Dedicated Inference (Reserved)

Contact sales

Reserved GPU capacity (7-180+ days) for lower hourly rates than on-demand; pricing requires contacting sales for longer terms and newer hardware.

  • Reserved terms from 7-30 days up to 181+ days
  • Lower per-GPU-hour rates than on-demand (e.g. H100 as low as $3.19/hr)
  • Available for H100, H200, B200; B300/GB200/GB300 contact sales
  • Guaranteed dedicated hardware

GPU Clusters

$4/hourly

On-demand pay-as-you-go GPU cluster capacity billed hourly for large-scale compute and training needs.

  • Hourly billing per GPU
  • NVIDIA HGX H100 at $3.99/hr
  • NVIDIA HGX H200 at $5.99/hr
  • NVIDIA HGX B200 at $8.19/hr
  • GB200/GB300/B300 availability varies

Sandbox / Code Interpreter

Contact sales

Usage-based pricing for VM sandboxes and secure code execution for LLM-generated code.

  • Compute priced per vCPU ($0.0446/hr) and GiB RAM ($0.0149/hr)
  • Code Interpreter sessions at $0.03 per 60-minute session
  • Customizable sandbox deployments for dev environments

Managed Storage

$0/monthly

High-bandwidth, parallel filesystem storage colocated with compute, billed per GiB per month.

  • Shared Filesystem storage at $0.16/GiB/month
  • Colocated with compute for low latency

Fine-Tuning

Contact sales

Usage-based pricing to fine-tune open-source models for production, priced per 1M tokens processed based on model size and method.

  • Supervised Fine-Tuning and Direct Preference Optimization (LoRA or Full)
  • Pricing tiers by model size (Up to 16B, 17B-69B, 70-100B)
  • Specialized model pricing for DeepSeek, GLM, Kimi, Llama 4, Qwen3.5, gpt-oss
  • Minimum job charge from $4.00 up to $60.00 depending on model

Enterprise / Contact Sales

Contact sales

Custom pricing for advanced hardware (NVIDIA HGX H200/B300, GB200/GB300 NVL72) and large-scale reserved deployments; requires contacting sales.

  • Custom hardware and capacity planning
  • Dedicated support team
  • Tailored contracts for reserved GPU capacity
  • Access to newest NVIDIA hardware

AI Panel Reviews

The Decision Maker

The Decision Maker

Strategic bet, vendor viability, timing, adoption approval
7.9/10

Together AI is the open-source inference cloud the board can sign off on without long explanations.

Vipul Ved Prakash sold Topsy to Apple in 2013 and is back with $534M raised — NVIDIA, Salesforce Ventures, and Kleiner Perkins on the cap table. The vendor question's settled; the harder call is whether you bet your inference stack on a $3.3B startup or the hyperscaler your CFO already pays.

Open-source AI infrastructure is the layer where the cloud margins shift over the next decade. Together is the cleanest pure-play, and Vipul Ved Prakash is the right founder for it — Topsy, acquired by Apple in 2013.

The runway math holds. $305M Series B at $3.3B in February 2025, with NVIDIA and Salesforce Ventures on the cap table. The product depth follows — Serverless Inference, Dedicated Inference at $3.99/hr per H100, Fine-Tuning down to $0.48 per million tokens. That's a real stack, not a thin API.

The catch: AWS Bedrock and Azure AI Foundry sit inside the cloud commit your CFO already signed. Together's defense is speed — Llama 4 and DeepSeek-R1 ship there before any hyperscaler catalogs them, and that lead matters this year. Pilot it where open-source freshness is the requirement. Don't standardize the org until renewal.

Competitive Positioning7.7

Differentiated against AWS Bedrock and Azure AI Foundry on open-model freshness, but narrower scope than a full hyperscaler.

Reputation Risk8.0

NVIDIA, Salesforce Ventures, and Kleiner Perkins on the cap table makes the vendor easy to defend in a board review.

Speed to Value8.2

Serverless Inference at $0.02 per million input tokens for some models means a pilot can be wired up in days, not weeks.

Strategic Fit7.8

Pure-play open-source inference advances open-model strategy; less of a fit if the company has already standardized on a single hyperscaler stack.

Vendor Viability7.5

$534M raised across four years and Series B at $3.3B in February 2025 funds at least 24 months of runway, but it is still a startup competing with three hyperscalers.

Pros

  • NVIDIA, Salesforce Ventures, and Kleiner Perkins on the cap table de-risks the vendor question for the board.
  • Full inference stack from Serverless Inference to Dedicated H100 to Fine-Tuning under one contract.
  • Newest open-source models like Llama 4 and DeepSeek-R1 ship faster than on hyperscalers.
  • $305M Series B at $3.3B in February 2025 funds at least 24 months of execution.

Cons

  • AWS Bedrock and Azure AI Foundry are bundled inside cloud commits most enterprises already signed.
  • Open-source moat narrows if Meta or DeepSeek slow their open-weights release cadence.

Right for

Teams running multi-model open-source inference who need an alternative to AWS Bedrock.

Avoid if

Buyers committed to a single hyperscaler with cloud spend already signed.

The Domain Strategist

The Domain Strategist

Craft and strategy in the product's domain — adapts identity per category, same lens
8.3/10

Together AI bet on the kernel layer rather than the API, and that is the right architectural call.

Together Kernel Collection is the substrate worth studying — vendor-owned GPU kernels that compound across inference, fine-tuning, and training. The full-stack pricing ladder from $0.02-per-million-token serverless to $3.49/hr H100 clusters lets teams scale without a vendor change.

Together's positioning is the inference-and-training cloud for open-source models, and the architecture follows. Together Kernel Collection — GPU kernels claiming up to 90% faster pre-training — is the layer that compounds across inference, fine-tuning, and dedicated clusters. Fireworks AI counters with FireAttention. Anyscale leans on Ray.

Pricing reflects the full-stack ambition. H100 on-demand at $3.49/hr, Llama-class inference from $0.02 per million tokens, Batch Inference scaling to 30 billion tokens per model. Teams can graduate from serverless to dedicated to reserved capacity without changing vendors — the shape an AI platform group actually wants.

The catch is open-source dependence. The catalog rides the Llama, Qwen, and DeepSeek release cadence; if Meta's open-weights commitment narrows, the moat thins to the kernel work alone. The $305M Series B at a $3.3B valuation in 2025 buys runway, but durability lives in the substrate, not the model selection.

Category Positioning8.5

Clear top-tier in the open-source AI cloud segment alongside Fireworks AI and Replicate, with the deepest research-team lineage.

Domain Fit8.5

Serverless, dedicated, and reserved-cluster tiers map cleanly to how senior AI platform teams actually graduate workloads.

Integration Surface8.0

OpenAI-compatible API, standard endpoints, Hugging Face catalog integration, and Batch API support across serverless and private deployments.

Long-term Implications7.8

Open-source model dependence is a real 3-year constraint; the catalog's relevance rides Meta, Mistral, and DeepSeek release cadence.

Strategic Depth8.5

Together Kernel Collection is real substrate work — vendor-owned GPU kernels claiming up to 90% pre-training speedup, not an API skin.

Pros

  • Together Kernel Collection delivers vendor-owned GPU optimization that compounds across inference, fine-tuning, and training.
  • Full-stack pricing ladder lets teams graduate from $0.02-per-million-token serverless to reserved H100 clusters without changing vendors.
  • Batch Inference scales to 30 billion tokens per model — handles workloads that break most serverless APIs.
  • Founder lineage from Stanford and ETH Zurich shows in the architectural choices, not just the marketing.

Cons

  • Open-source model dependence ties the catalog's relevance to Meta, Mistral, and DeepSeek release cadence.
  • No proprietary frontier model means buyers needing GPT-5 or Claude-tier reasoning still pair Together with another vendor.
  • Specialized fine-tuning for models like Kimi K2 climbs to $15 per 1M tokens — premium territory for teams expecting commodity rates.

Right for

AI platform teams who run open-source models in production.

Avoid if

Buyers who need a single proprietary frontier model with vendor-managed safety guarantees.

The Finance Lead

The Finance Lead

Money, total cost of ownership, contracts, procurement math
8.3/10

H100 reserved drops from $3.49 to $2.55/hr if you commit four months — Together's discount curve is honest.

Together publishes every tier on its pricing page, from $0.02 per million tokens for inference to $2.55/hr for an H100 on a 4-6 month reservation. The catch is Specialized Fine-Tuning — minimums up to $60 per job mean small experiments aren't free.

Pricing is fully published. Every tier, every GPU, every per-token rate. Inference starts at $0.02 per 1M input tokens. H100 on-demand at $3.49/hr — same shape as Lambda Labs, cheaper than AWS Bedrock provisioned throughput. Procurement won't push back.

The reserved-GPU curve is where the math gets honest. H100 drops to $2.99/hr at one week, $2.55/hr at 4-6 months. Six-day minimum. One H100 reserved four months runs about $7,300 — versus $10,000 on-demand. 27% saving compounds across a fleet, but the 6-day floor punishes spiky workloads.

Two line items matter. Specialized Fine-Tuning carries minimum charges — $20 for DeepSeek-R1 LoRA, up to $60 for GLM-5. Small experiments aren't free. However, Managed Storage charges zero egress, which offsets a year of S3 transfer for inference-heavy teams. Read the contract, not the marketing.

Billing & Procurement8.3

Usage-based with web checkout for serverless removes most procurement friction.

Contract Flexibility8.0

On-demand and reserved are both available; only the 6-day reserved minimum limits flexibility.

Pricing Transparency9.0

Every tier and GPU rate is published; serverless and on-demand require no sales call.

ROI Clarity7.9

Per-token and per-hour rates make inference ROI directly measurable; the 60% optimization claim is unverified.

Total Cost of Ownership7.8

Modeling is feasible across compute, storage, and fine-tuning, though minimum charges complicate small jobs.

Pros

  • Every pricing tier is published on the website with no sales call required for serverless or on-demand.
  • Reserved GPU pricing drops 27% from on-demand when you commit to a 4-6 month term.
  • Managed Storage charges zero egress fees, which is unusual in cloud infrastructure.
  • Per-token inference starts at $0.02 per 1M input tokens, competitive with hyperscaler list rates.

Cons

  • Specialized Fine-Tuning carries minimum charges from $20 to $60 per job, so small experiments are not cheap.
  • Reserved GPU clusters require a six-day minimum reservation, which punishes intermittent workloads.
  • The published 60% cost-reduction claim from workload-specific optimization is not independently verifiable.

Right for

Teams who run mixed-model inference and want every rate public before signing.

Avoid if

Teams who need spiky GPU access without committing to a six-day reservation.

The Domain Practitioner

The Domain Practitioner

Daily hands-on reality in the product's domain — adapts identity per category, same lens
8.1/10

Point your OpenAI client at api.together.ai/v1 and DeepSeek-R1 answers — multi-model inference without the SDK juggling.

Together AI runs as an OpenAI-compatible endpoint with hundreds of open-weight models behind one base URL, plus serverless Batch Inference and dedicated H100 clusters when you outgrow shared inference. The catch is the long tail — provider-specific features like reasoning traces or strict tool calling don't always pass through cleanly.

The integration test is a one-liner. Set OPENAI_API_BASE to https://api.together.ai/v1, prefix the model with deepseek-ai/ or meta-llama/, keep the OpenAI client. Compare wiring up Replicate's prediction-polling API — Together is the lowest-effort swap for a Python codebase already on OpenAI.

Batch Inference is the sleeper feature. Asynchronous, 30 billion tokens per model, priced below interactive — for embeddings backfills or eval sweeps it's the right shape. Dedicated Inference at $3.99/hr for an H100 80GB undercuts the ops cost of self-hosting vLLM once you count engineer time. Together Kernel Collection gets cited as the moat.

The friction is the long tail. Reasoning-mode toggles for DeepSeek-R1, prompt caching, structured-output strict mode — features the OpenAI Chat Completions surface doesn't always express, and the docs lag a release behind. However, for the 80% case of multi-model inference across open weights, this is the path of least resistance.

Day-3 Reality8.0

Once integrated it disappears from the stack — daily friction shows up only at provider-specific feature edges.

Documentation Practitioner-Fit7.5

Docs cover the platform broadly but lag new model launches and the Together Kernel Collection details by a release.

Friction Surface7.5

Reasoning-mode and strict tool-call semantics don't always express cleanly through the Chat Completions surface.

Power-User Depth8.5

Batch, Dedicated Inference, fine-tuning to 100B parameters, and GPU clusters all sit on the same control plane.

Workflow Integration8.5

OpenAI-compatible base URL means existing clients, retry logic, and observability tooling work unchanged.

Pros

  • OpenAI-compatible base URL means existing clients work with a one-line change.
  • Batch Inference scales to 30 billion tokens per model for asynchronous workloads.
  • Dedicated H100 80GB at $3.99/hr undercuts self-hosting vLLM once ops time is counted.
  • Fine-tuning supports open-weight models up to 100B parameters with LoRA or full SFT.

Cons

  • Provider-specific features like reasoning traces or strict tool calling don't always pass through the OpenAI surface cleanly.
  • Documentation lags model and feature releases by a release or two.
  • Specialized fine-tuning on DeepSeek-R1 LoRA starts at $10 per 1M tokens with a $20 minimum charge.

Right for

Backend engineers who run open-weight models behind an OpenAI-compatible client.

Avoid if

Teams who need only proprietary frontier models like GPT-5 or Claude Opus.

The Power User

The Power User

Daily human experience, onboarding, polish, learning curve, reliability
8.0/10

Together's pricing page lists every number on one screen, and that small thing tells you a lot.

The playground works without a signup, the pricing page lists every number, and the OpenAI-shaped endpoint means your client code just works. The catch is the docs lag the model catalog by a release.

The playground at api.together.xyz/playground is the small thing the team got right. No credit card to sign in, 200+ open-source models in a dropdown, paste a prompt and watch tokens stream. Hugging Face Inference Endpoints makes you wire up a deployment first.

The pricing page is where the team earns trust. H100 on-demand at $3.49 an hour, Llama-class inference from $0.02 per million input tokens, Code Interpreter at $0.03 per 60-minute session — every number on one page. Modal makes you log in to see GPU hourly.

But the docs lag a release behind the model catalog. A new model lands Tuesday, the structured-output flag for it shows up in the docs the following week. Worth it for a $305M Series B cloud where NVIDIA is on the cap table. Painful if you're chasing a model that dropped this morning.

Daily Polish8.0

Pricing page consolidates every number on one screen and the playground works without a signup.

Learning Curve7.5

First ten minutes are fast, but the docs trail the catalog by a release which slows the day-thirty fight.

Mobile Parity7.5

Mobile is essentially read-only, but for a dev-infrastructure API this is category norm.

Onboarding Experience8.2

No credit card to reach the playground, OpenAI-compatible base_url means existing code runs in minutes.

Reliability Feel7.8

Full-stack ambition with autoscaling and dedicated GPUs at $3.99 an hour, but uptime depends on open-source model release cadence.

Pros

  • Playground at api.together.xyz/playground works without a signup or credit card.
  • OpenAI-compatible base_url lets existing client code call 200+ open-source models without rewrites.
  • Pricing page lists every number on one screen — H100 at $3.49/hr, inference from $0.02/M tokens, Code Interpreter at $0.03 per session.
  • $305M Series B at a $3.3B valuation in 2025 with NVIDIA on the cap table signals real durability.

Cons

  • Docs trail the model catalog by a release — new models land in the dropdown before the reference page.
  • Mobile is essentially read-only; the playground renders on a phone but nobody would write code there.
  • Catalog depends on open-source release cadence from Meta, DeepSeek, and Qwen — a narrowing of open weights would thin the offering.

Right for

Developers who want to swap between open-source models without changing their stack.

Avoid if

Teams who need polished docs the same day a new model launches.

The Skeptic

The Skeptic

Contrarian. Watch-outs, deal-breakers, broken promises, category patterns
7.6/10

OctoAI got swallowed by NVIDIA in September 2024 — same category, same cap table, fewer survivors.

Together AI is the largest pure-play left in open-source inference, and the $305M Series B in February 2025 buys real time. The yellow flag is the category's body count — OctoAI absorbed by NVIDIA, MosaicML by Databricks, and Together has NVIDIA on its own cap table.

Two acquisitions in eighteen months. OctoAI absorbed by NVIDIA around $165M in September 2024, commercial service off by October 31. MosaicML by Databricks at $1.3B in 2023. Same category, fewer survivors. Together is the largest pure-play left.

What I found holds up. $305M Series B at $3.3B in February 2025, led by General Catalyst and Prosperity7. Serverless Inference from $0.02 per million input tokens. H100 reserved at $2.55/hr. Together Kernel Collection is the moat the docs actually try to defend. Real product.

But NVIDIA is on the cap table. So is Salesforce Ventures. Both have absorbed peers in this category — NVIDIA bought OctoAI, Lepton AI followed. The graveyard pattern doesn't predict Together's outcome. It does mean the question shifts: durable cloud, or attractive tuck-in once the kernel work matures.

Competitive Differentiation7.5

Together Kernel Collection is real engineering work but Fireworks AI and Anyscale chase the same kernel-layer moat.

Exit Portability8.2

OpenAI-compatible API at api.together.ai/v1 means migration off looks mechanical, not catastrophic.

Long-term Viability7.5

$305M Series B at $3.3B in February 2025 buys runway; NVIDIA on the cap table is double-edged.

Marketing Honesty8.0

Pricing page lists every GPU tier, per-token rate, and minimum charge — claims are quantified, not aspirational.

Track Record Match6.8

Open-source inference cloud category has visible failures — OctoAI absorbed by NVIDIA, MosaicML by Databricks.

Pros

  • $305M Series B at $3.3B in February 2025 led by General Catalyst — runway is real.
  • Pricing fully published — H100 reserved drops to $2.55/hr, inference from $0.02 per million tokens.
  • OpenAI-compatible API at api.together.ai/v1 means migration off looks mechanical, not catastrophic.
  • Together Kernel Collection is a real engineering moat, not a thin API wrapper.

Cons

  • Open-source inference category has a visible graveyard — OctoAI absorbed by NVIDIA, MosaicML by Databricks.
  • NVIDIA sits on Together's cap table and has already acquired two peers in the same category.
  • Catalog rides Llama and DeepSeek release cadence — open-weights momentum is the substrate.

Right for

Teams who need the largest open-source inference pure-play with real Series B runway.

Avoid if

Buyers who already have an AWS commit covering Bedrock open-weights inference.

Buyer Questions

Common questions answered by our AI research team

Pricing

How is Provisioned Throughput priced?

Provisioned Throughput uses PTUs (provisioned throughput units), where each PTU represents fixed capacity with a price per PTU per minute. It includes a 99% uptime SLA and drop-in API compatibility for production workloads with no infrastructure to manage.

Pricing

Does Together AI offer batch inference pricing?

Yes. Batch Inference lets you cost-effectively process massive workloads asynchronously, scaling to 30 billion tokens per model with any serverless model or private deployment, with dedicated batch API pricing per model shown on the pricing page.

Features

Can I fine-tune open-source models on Together AI?

Yes. Together AI offers fine-tuning of open-source models for production workloads using the latest research techniques, helping improve accuracy, reduce hallucinations, and control model behavior without managing training infrastructure.

Features

Does Together AI provide code sandboxes for AI apps?

Yes. The Sandbox offering provides fast, secure code sandboxes at scale for setting up full-scale development environments for AI apps and agents, supporting workflows like creating, connecting, and running commands in sandboxes via SDK.

Product Information

  • Founded

    2022
  • Pricing

    From $0/mo
  • Free Trial

    Available
  • Free Plan

    Available

Platforms

web

About Together AI

Build what's next on the AI Native Cloud. Full-stack AI platform for inference, fine-tuning, and GPU clusters — powered by cutting-edge research.

Resources

Documentation
API
Blog

Also in LLM Platforms