ngrok AI Gateway logo

ngrok AI Gateway Review

Visit

One gateway for every AI model, public or self-hosted

ngrok AI Gateway is a hosted routing and observability layer for calling multiple AI models through a single API.

AI Panel Score

7.6/10

6 AI reviews

Reviewed

About ngrok AI Gateway

To use the gateway, a developer changes their SDK's baseURL to https://gateway.ngrok.ai and swaps in an ngrok API key, without deploying any additional infrastructure. Requests can specify a primary model, including a self-hosted one, along with a list of fallback models from public providers, and the gateway supports streaming responses. It works with the OpenAI SDK, Anthropic SDK, and Vercel AI SDK.

The gateway connects to self-hosted or local LLMs over private connectivity, without requiring public IPs or open inbound ports. It supports bring-your-own-key (BYOK) setups for OpenAI, Anthropic, or custom providers, so users keep their existing rates and are billed directly by those providers, while managing all keys in one place. Access control lets teams issue separate keys per app or developer, each scoped to specific providers and models, rather than sharing one key with full access. Observability rolls up token usage, latency, and errors across every routed call, broken down by app, developer, or model, which the site notes is missing from individual provider dashboards. The gateway can also be configured entirely through APIs, for use with coding agents, Terraform, a CLI, or custom tooling.

The product is aimed at developers and teams that call multiple AI models, including their own self-hosted ones, and want centralized routing, failover, and cost visibility. Pricing is usage-based at $0.05 per million tokens routed, in addition to the cost of inference; using ngrok-provided keys passes inference through at cost, while bring-your-own-key setups are billed directly by the provider. There are no subscriptions or commitments; users purchase credits upfront that are drawn down as requests are routed.

Features

AI

  • Multi-Provider & Self-Hosted Model Routing

    Routes requests across public providers, custom endpoints, and self-hosted models, including specifying a primary model with fallback models in the same request.

Analytics

  • Usage Observability

    Rolls up tokens, latency, and errors across every routed call, showing what was spent broken down by app, developer, or model.

Automation

  • Automatic Failover

    Reroutes requests to a healthy, predefined alternative model or key when the original one slows or fails.

  • Automatic Retries

    Automatically retries failed requests instantly so applications keep running without needing custom error-handling code.

Core

  • Bring Your Own Key (BYOK)

    Lets users drop in their existing OpenAI, Anthropic, or custom provider keys to route through the gateway while paying providers directly at their existing rates, managed in one place.

  • Unified Gateway URL

    Lets developers point existing OpenAI, Anthropic, or Vercel AI SDK clients at a single baseURL (gateway.ngrok.ai) to route requests without deploying separate infrastructure for each provider.

Integration

  • Local LLM Connectivity

    Connects privately to self-hosted or local LLMs reachable from the AI Gateway without exposing public IPs or inbound ports.

  • Programmable Configuration

    Allows the AI Gateway to be configured entirely through APIs, callable from coding agents, Terraform, CLI, or custom tooling.

Security

  • Scoped Access Keys

    Issues separate access keys per app or developer with configurable permissions for which providers and models each key is allowed to call.

Preview

ngrok AI Gateway desktop previewngrok AI Gateway mobile preview

Pricing Plans

Pay as you go

$0/usage-based per million tokens

Usage-based pricing for routing through ngrok AI Gateway; buy credits upfront and draw down as you route, no subscriptions or commitments

  • $0.05 per million tokens flat routing fee plus cost of inference
  • Use ngrok keys with inference passed through at cost, or bring your own provider keys billed directly by provider
  • One gateway URL for any SDK (OpenAI, Anthropic, Vercel AI SDK)
  • Route to public providers, custom endpoints, and self-hosted local models
  • Access control with scoped API keys per app or developer
  • Built-in observability: tokens, latency, and errors across every call

AI Panel Reviews

The Decision Maker

The Decision Maker

Strategic bet, vendor viability, timing, adoption approval
7.6/10

ngrok bets its networking name on being the boring layer between your app and every model.

One baseURL swap gets you routing, failover, and cost visibility across public and self-hosted models. Pricing is honest but the moat is thin.

ngrok knows private connectivity. That's the real asset here — routing to self-hosted LLMs without opening inbound ports is the one thing the obvious alternative, a raw multi-provider SDK setup, doesn't solve well.

The pricing is clean: $0.05 per million tokens routed, no subscription, BYOK keeps your existing provider rates. That's easy to defend to a board because it's usage-based and reversible if it doesn't work out.

Two gaps: the $0.05 rate lives on the homepage rather than a dedicated pricing page, and there's no changelog to judge momentum beyond the blog. Scoped keys and automatic failover are real, shipped features — not vaporware claims. But you're routing every request through a new hop; that's a dependency, not just a convenience.

Competitive Positioning7.8

Private connectivity to local LLMs plus per-key model scoping goes beyond what a plain multi-provider SDK setup offers.

Reputation Risk6.8

No public IP/inbound port exposure for self-hosted models is a real security plus, but no compliance or SLA detail is published.

Speed to Value8.2

Single baseURL swap with existing OpenAI, Anthropic, or Vercel AI SDK clients means no infrastructure deployment to start routing.

Strategic Fit7.5

Centralized routing and observability across self-hosted and public models advances multi-model architecture, not just cost-cutting.

Vendor Viability7.0

Docs and API exist per evidence; the ngrok blog posts AI Gateway updates, but there's no dedicated changelog to gauge shipping cadence.

Pros

  • Flat $0.05/million token fee with BYOK passthrough at provider rates
  • Private connectivity to self-hosted models without public IPs or open ports
  • Scoped access keys per app or developer limit blast radius of a leaked key
  • Automatic failover and retries reduce custom error-handling code

Cons

  • No changelog, though the ngrok blog shipped AI Gateway updates as recently as April 2026
  • Single flat pricing tier with no visible breakdown for high-volume discounts
  • New network hop for every AI call is a dependency, not free

Right for

Teams running self-hosted models alongside public providers who want one key-managed routing layer.

Avoid if

Skip it if you call a single provider and don't need failover or self-hosted routing.

The Domain Strategist

The Domain Strategist

Craft and strategy in the product's domain — adapts identity per category, same lens
7.6/10

A control-plane play from an infrastructure company that already owns your tunnels.

ngrok is putting a routing and observability layer in front of every model call, priced at $0.05 per million tokens. The architecture is sound; the lock-in risk is that your failover logic now lives in someone else's control plane.

The baseURL swap is the whole pitch: point your OpenAI, Anthropic, or Vercel AI SDK client at gateway.ngrok.ai and you get failover, retries, and scoped keys without deploying anything. That's the right shape for how teams actually run multi-model stacks today — nobody wants to hand-roll retry logic across three provider SDKs. BYOK support means you're not forced into ngrok's margin on inference, only the $0.05/million routing fee.

The real question is where this sits in three years. If your failover rules, fallback chains, and access policies all live in ngrok's config, then migrating off means rebuilding a control plane you didn't have to think about before. Private connectivity to self-hosted models without open inbound ports is a genuine security win, though — that's infra ngrok already knows how to build.

Programmable via Terraform, CLI, and API means it fits GitOps workflows, not just dashboards. No public pricing page or changelog listed is a minor diligence gap for a product touching every inference call you make.

Category Positioning7.5

Usage-based at $0.05 per million tokens with no subscription undercuts the commitment model of the obvious alternative for this job.

Domain Fit8.2

A single baseURL swap across OpenAI, Anthropic, and Vercel AI SDK matches how teams already structure multi-provider calls.

Integration Surface8.0

Terraform, CLI, and API-driven config plus private connectivity to self-hosted LLMs fit existing infra-as-code workflows.

Long-term Implications7.0

Centralizing routing and access control in ngrok's control plane is convenient now but a migration cost later if you outgrow it.

Strategic Depth7.5

Failover, retries, and scoped keys cover the core reliability problem well, but the ceiling is routing — not model evaluation or prompt management.

Pros

  • BYOK keeps existing provider rates while centralizing key management
  • Private connectivity to self-hosted models without open inbound ports
  • Fully programmable via Terraform, CLI, and API for GitOps workflows

Cons

  • Failover and access policy logic becomes dependent on ngrok's control plane
  • Pricing lives on the homepage with no dedicated pricing page or changelog, limiting roadmap diligence
  • Observability is rollup-level, not a substitute for per-provider deep debugging

Right for

Teams running multiple LLM providers and self-hosted models who want unified failover and cost visibility without building it themselves.

Avoid if

Avoid if you need deep model-specific debugging tools or want to avoid adding a routing dependency between your app and inference providers.

The Finance Lead

The Finance Lead

Money, total cost of ownership, contracts, procurement math
7.9/10

$0.05 per million tokens routed. No subscription. Math is simple, that's rare.

Flat routing fee plus pass-through inference cost, prepaid credits, no contract. The real cost driver is your token volume, not ngrok's markup.

$0.05/million tokens for routing. Inference billed at cost with ngrok keys, or direct from provider with BYOK. No subscriptions, no seats, no tiers to compare. Team routing 500M tokens/month pays $25/month in gateway fees. Year 3 at 2x volume: $600/year in fees. Inference cost dwarfs it either way.

Prepaid credits, drawn down as you route. No auto-renewal clause to hunt for — because there's no term. That's the opposite of the hostage contract you usually sign for infrastructure tooling.

ROI is measurable: observability rolls up tokens, latency, errors per app, per developer, per model. That's the number you'd otherwise piece together from separate provider dashboards. Tradeoff: $1 of starting credit is all you get before a $5 minimum top-up, so real load testing starts on your dime. And BYOK billing lands on two invoices — ngrok's routing fee, provider's inference bill — not one clean line item.

Billing & Procurement7.0

Prepaid credits simplify one side, but BYOK setups split billing across ngrok and the provider directly.

Contract Flexibility8.8

No subscriptions, no commitments, prepaid credits with no renewal window to track.

Pricing Transparency9.0

Single flat rate, $0.05/million tokens, published with no sales call required.

ROI Clarity7.5

Token, latency, and error rollups per app/developer/model give a concrete usage number, though savings from failover aren't quantified.

Total Cost of Ownership8.2

Routing fee stays trivial next to inference cost; no seat creep since pricing is usage-based, not per-user.

Pros

  • Flat $0.05/million token fee with no subscription tier games
  • BYOK keeps existing provider rates, billed direct
  • Scoped access keys per app or developer limit blast radius

Cons

  • Only $1 of starting credit before a $5 minimum top-up
  • BYOK means two invoices to reconcile, not one
  • No published overage or rate-limit ceiling

Right for

Teams routing meaningful token volume across multiple providers who want one bill line for routing and clean usage dashboards.

Avoid if

Avoid if you're a single-provider shop with no self-hosted models and no need for cross-provider failover.

The Domain Practitioner

The Domain Practitioner

Daily hands-on reality in the product's domain — adapts identity per category, same lens
7.8/10

Swap a baseURL, get failover and per-key scoping — the ngrok pedigree helps with the private connectivity piece.

One-line SDK swap, private tunnels to self-hosted models, and $0.05/M token routing fee with no subscription. Docs cover the SDK swap and the management API, but there's no dedicated pricing page or changelog, so day-3 confidence depends on how the API actually behaves under load.

Changing baseURL to gateway.ngrok.ai and dropping in an access key is the whole onboarding story. That's the right instinct — no sidecar, no extra infra, works with OpenAI SDK, Anthropic SDK, and Vercel AI SDK out of the box. The private connectivity to self-hosted models without opening inbound ports is the standout: that's ngrok's actual core competency leaking into a new product, not a bolted-together feature.

What's unproven is the daily-fight surface. Fallback model lists per request and automatic retries sound clean in the docs answer, and the management API is documented — POST to api.ngrok.ai/access-keys — but there's no changelog and no dedicated pricing page, just the rate on the homepage. Scoped keys per app/developer is the kind of access-control feature that matters at week 3, not day 1, and it's here.

$0.05 per million tokens plus pass-through inference cost is easy to model, no subscription lock-in. Tradeoff: new accounts get $1 in credit and then a $5 minimum purchase, so kicking the tires at real volume means buying in.

Day-3 Reality7.5

One-line SDK swap is real and the management API is documented, but no changelog means unknowns about retry/failover edge cases under sustained load.

Documentation Practitioner-Fit6.8

Q&A answers read practitioner-written and specific and the management API is documented, but there's no changelog, so depth under load is unverified.

Friction Surface7.3

$1 of starting credit means the first real friction point is a $5 minimum purchase before you can test routing behavior against your own traffic patterns.

Power-User Depth7.9

Scoped per-app/per-developer keys, BYOK, and full API/Terraform configurability suggest real headroom beyond the basic baseURL swap.

Workflow Integration8.2

Works with existing OpenAI/Anthropic/Vercel AI SDK clients and is configurable via API, Terraform, and CLI — fits agent and IaC workflows without new tooling habits.

Pros

  • Private connectivity to self-hosted/local LLMs without public IPs or open ports
  • No subscription — usage-based credits at $0.05 per million tokens routed
  • Works unmodified with OpenAI, Anthropic, and Vercel AI SDK clients
  • Scoped access keys per app/developer for real least-privilege setups

Cons

  • Only $1 of starting credit to test failover before buying credits
  • No changelog, though the management API at api.ngrok.ai is documented
  • BYOK billing split across ngrok and provider dashboards adds a reconciliation step

Right for

Teams already juggling multiple providers or self-hosted models who want one baseURL and centralized cost visibility.

Avoid if

Avoid if $1 of starting credit isn't enough to validate failover and retry behavior before committing production traffic.

The Power User

The Power User

Daily human experience, onboarding, polish, learning curve, reliability
7.6/10

Change one line of code, get a safety net under every model you call.

Swap your baseURL, keep your SDK, and suddenly your self-hosted models and OpenAI calls share one bill and one dashboard. It's the kind of infrastructure tool that's boring in the best way, if the boring parts actually hold up.

The pitch is dead simple: point your OpenAI, Anthropic, or Vercel AI SDK client at gateway.ngrok.ai, swap the key, done. No infrastructure to deploy. That's a real ten-minute setup, not a sales-call setup, and the docs presence backs that up.

What sells me is the boring stuff nobody screenshots: automatic retries, failover to a healthy model when one chokes, scoped keys per app instead of one master key everyone shares. That's the stuff that saves you at 2am, not the stuff that wins the demo. Observability rolling up tokens, latency, and errors by app or developer is genuinely missing from most provider dashboards, so that's a fair claim.

$0.05 per million tokens plus inference cost, no subscription, is easy to reason about. The tradeoff: $1 of starting credit, then a $5 minimum purchase, so you're paying to find out if the failover logic behaves the way the docs say under real load.

Daily Polish7.0

Scoped keys and rolled-up observability by app/model show real daily-use thinking, but no changelog and no dedicated pricing page means limited proof of ongoing care.

Learning Curve7.5

Familiar SDKs plus programmable config via API, CLI, and Terraform suggest it scales from quick test to full automation without a cliff.

Mobile Parity6.5

Platform listed as web only; this is developer infrastructure configured via API, CLI, and Terraform, so mobile isn't really the point.

Onboarding Experience8.5

Change baseURL, swap key, route — the docs describe a setup with zero new infrastructure, about as low-friction as this category gets.

Reliability Feel7.8

Automatic failover and instant retries are named features aimed directly at reliability, though no uptime SLA or status page is published.

Pros

  • Drop-in with existing OpenAI, Anthropic, Vercel AI SDK clients
  • Scoped per-app, per-model access keys instead of one shared key
  • Usage-based at $0.05/million tokens with no subscription lock-in
  • Private connectivity to self-hosted models without open ports

Cons

  • Just $1 of starting credit to test failover before paying
  • Rate published on the homepage only, no dedicated pricing page or changelog, so track record is thin
  • Mobile is a non-factor since this is pure developer infrastructure

Right for

Teams juggling multiple AI providers and self-hosted models who want one bill and one dashboard.

Avoid if

You only call a single provider and don't need failover, key scoping, or cross-model observability.

The Skeptic

The Skeptic

Contrarian. Watch-outs, deal-breakers, broken promises, category patterns
6.9/10

ngrok knows infra. A gateway isn't infra, it's a dependency.

Solid feature list, clean pricing, but this is ngrok extending its brand into a category where the failure mode is different: your model traffic runs through their URL. No dedicated pricing page and no changelog, though the management API at api.ngrok.ai is documented, which backs the 'programmable configuration' claim.

$0.05 per million tokens is easy to understand. BYOK means you're not trapped on their inference markup, which is the right call and rare in this category. Scoped keys and per-model permissions are real, specific, checkable claims — not vapor.

But the switch cost is baked in structurally. Your baseURL becomes gateway.ngrok.ai. Every SDK call routes through them. If ngrok.ai deprioritizes this (it's an extension of a tunneling company's brand, not their core product), reverting means touching every client config again. That's not catastrophic, just not free.

There's no changelog and no dedicated pricing page, though the documented management API at api.ngrok.ai does back the 'configurable entirely through APIs' claim. The self-hosted LLM routing over private connectivity is a genuine differentiator versus point-to-point setups. Whether ngrok sustains this beyond a launch feature is the open question, not the tech itself.

Competitive Differentiation7.0

Self-hosted model routing over private connectivity plus BYOK billing is a real gap versus typical provider-locked gateways.

Exit Portability6.0

Single baseURL swap-in is easy to adopt but equally sticky to unwind across every client.

Long-term Viability7.0

Backed by an established company name, but there's no changelog and pricing lives on the homepage rather than a dedicated page.

Marketing Honesty7.2

Claims map to named features (BYOK, scoped keys, failover) rather than vague superlatives.

Track Record Match6.5

Specific mechanics described (fallback lists, private connectivity) but no usage numbers or uptime data to verify claims.

Pros

  • BYOK keeps existing provider rates, no markup
  • Scoped per-app, per-model access keys
  • Private connectivity for self-hosted LLMs, no open ports

Cons

  • No changelog and no dedicated pricing page, just the rate on the homepage
  • Single baseURL creates real migration friction later
  • Only $1 of starting credit before committing to a $5 minimum purchase

Right for

Teams already juggling multiple model providers who want centralized failover and cost visibility.

Avoid if

You need to verify uptime and support history before routing production traffic through a third party.

Buyer Questions

Common questions answered by our AI research team

Pricing

How much does ngrok AI Gateway cost?

ngrok AI Gateway charges a flat $0.05 per million tokens for routing and observability, plus the cost of your inference. Use ngrok keys and inference is passed through at cost, or bring your own keys and your provider bills you directly. You buy credits up front with no subscriptions or commitments.

Integration

Which SDKs work with ngrok AI Gateway?

It works with the OpenAI SDK, Anthropic SDK, and Vercel AI SDK. Just change your baseURL to https://gateway.ngrok.ai and swap your API key to start routing.

Features

Can I connect my own self-hosted LLMs?

Yes. You can route to any local LLM reachable from your AI Gateway over private connectivity, without wrangling public IPs or inbound ports, just like any other model.

Setup

How do I set up the gateway with my app?

Change your client's baseURL to https://gateway.ngrok.ai and set your API key to your NGROK_AI_ACCESS_KEY. From there you can specify a primary model and fallback models in your request, with no infrastructure deployment needed.

Security

Can I restrict which models a key can access?

Yes. You can issue separate scoped access keys for each app or developer and set exactly which providers and models each key is allowed to call, instead of sharing one key that opens everything.

Also in AI APIs