AI Observability and Security for enterprise agents and ML models
Fiddler AI is an AI observability and security platform for enterprises deploying AI agents and machine learning models.
AI Panel Score
6 AI reviews
Reviewed
Fiddler AI works as a control plane sitting across the lifecycle of AI agents and predictive ML models, from development through production. Teams use it to get hierarchical visibility into agentic systems, from the application level down to session, agent, trace, and span, and to evaluate agents against curated golden and challenger datasets before deployment. In production, the platform continuously monitors for drift, performance degradation, and policy violations, and can enforce rules on agent behavior in real time rather than only reporting on issues after the fact.
A core differentiator the site highlights is Fiddler Centor Models (previously called Fiddler Trust Models), purpose-built task-specific evaluation models that run entirely within the customer's environment with no external API calls or add-on fees, returning results in under 100ms across 80+ out-of-the-box and customizable metrics. This is positioned against the cost of using external LLM-as-a-Judge services for evaluation, which the company refers to as a "Trust Tax." The platform also offers Guardrails for detecting hallucinations, toxicity, PII/PHI leakage, and prompt injection in real time, plus a Control Plane for Coding Agents that adds inline PII/PHI and secret redaction and fleet-wide cost and token visibility by integrating through an existing LLM gateway. Governance features align with frameworks including NIST AI RMF, ISO/IEC 42001, GDPR, HIPAA, NAIC, SR 11-7, and the EU AI Act. The platform is framework-agnostic, supporting LangGraph, LangChain, Strands, OpenTelemetry, and custom implementations, with integrations including Amazon SageMaker AI and NVIDIA NIM.
Fiddler AI is built for enterprises operating AI agents and ML models in regulated or high-stakes environments, including government, healthcare, insurance, and financial services organizations, based on the industry pages and customer case studies (Nielsen, U.S. Navy, Integral Ad Science, Mastercard, Ally, DTCC, and others) referenced on the site. Pricing is structured across Starter, Business, and Premium plans, with specific pricing available via the company's pricing page or by contacting sales; a free tier for Guardrails on LLM applications is also offered.
Deployment options include cloud, VPC, and air-gapped environments, with AWS GovCloud support for government use cases. The platform holds SOC 2 Type II certification and HIPAA compliance.
Purpose-built, task-specific evaluation models that run entirely in-environment with no external API calls, returning results in under 100ms with 80+ out-of-the-box and customizable metrics.
Helps mitigate bias in AI models and supports building a responsible AI culture.
Delivers unified observability at enterprise scale from application to span level across agents and predictive models.
Provides hierarchical visibility from application to session, agent, trace, and span, and evaluates agents with curated golden and challenger datasets in pre-production.
Compares the annual cost of evaluating agent traces with external LLMs versus Fiddler Centor Models to show cost savings.
Monitors traditional ML models with drift detection, performance monitoring, and explainability.
Offers centralized control and accountability for enterprise AI governance, aligned with NIST AI RMF, ISO/IEC 42001, GDPR, HIPAA, NAIC, SR 11-7, and the EU AI Act.
Provides standardized telemetry, reliable evaluation, continuous monitoring, enforceable policy, and auditable governance across first-party, third-party, and coding agents from creation to production.
Provides a Tier-1 integration with Amazon SageMaker AI and an integration with NVIDIA NIM for safeguarding agentic applications.
Supports LangGraph, LangChain, Strands, OpenTelemetry, and custom implementations for connecting agent frameworks to the platform.
Provides inline PII/PHI and secret redaction plus fleet-wide developer, cost, and token visibility for coding agents by integrating through an existing LLM gateway.
Supports Cloud, VPC, and air-gapped environments for secure, in-environment deployments, including AWS GovCloud support.
Detects hallucinations, toxicity, PII/PHI leakage, and prompt injection attacks in real time with under 100ms latency.
Real-time guardrails to detect harmful exposure, listed at no cost on Fiddler's pricing page. Powered by contextual and task specific Fiddler Centor Models.
Ship faster and maximize ROI. Fiddler publishes the rate as "$0.002 per trace", shown here as the equivalent $2 per 1,000 traces. Everything in Free, plus unified observability for agentic and predictive systems.
Governance and safety at enterprise scale. Everything in Developer, plus enterprise-grade guardrails and infrastructure. Fiddler publishes no price for this tier; its call to action is Contact sales.
A real control plane for agent risk, though the enterprise-grade build of it is the tier you negotiate.
“Fiddler AI bundles agent observability, guardrails, and governance mapping into one platform built for regulated industries. The catch is that the enterprise-grade guardrails and scalability sit a tier above the one you can price yourself.”
Three tiers — Free, Developer, Enterprise — and only the top one is quoted rather than priced. The Centor Models pitch is the interesting part: in-environment evaluation, sub-100ms, claiming up to 98% lower TCO than routing evals through external LLM-as-a-judge calls. If that number holds up under contract, it's a real architectural edge, not a marketing line.
Governance mapping to NIST AI RMF, ISO/IEC 42001, GDPR, HIPAA, SR 11-7, and the EU AI Act is the kind of checkbox regulated buyers actually need, and air-gapped/GovCloud deployment backs it up. Framework-agnostic support for LangGraph, LangChain, and OpenTelemetry means you're not locked into one agent stack.
Tradeoff: this is infrastructure for teams already running agents in production, not a starter tool for teams still deciding whether to build one. SOC 2 Type II and HIPAA compliance are there. Pilot it against your actual eval spend before you sign Enterprise.
In-environment Centor Models avoiding external API calls, claiming up to 98% lower eval TCO, is a concrete architecture difference versus routing evals through foundation model APIs.
SOC 2 Type II and HIPAA compliance plus air-gapped deployment cover the regulated-industry bar, and the subscription terms are published on the site rather than held back for a call.
Sub-100ms guardrail latency and a free Guardrails tier let teams test value fast, but full Control Plane rollout likely needs integration work first.
Governance mapping to EU AI Act and NIST AI RMF advances actual compliance posture, not just cost-cutting on existing monitoring.
Three-tier pricing page, named integrations (SageMaker, NVIDIA NIM), and a named rebrand (Trust Models to Centor Models) all point to active product work.
Regulated enterprises running AI agents in production who need audit-ready governance and real-time guardrails.
Skip it if you're still prototyping agents and don't yet have compliance requirements to satisfy.
A control plane built for the audit binder, not just the model card, mapped to seven named frameworks.
“Fiddler AI treats governance as a first-class output, not an afterthought bolted onto a monitoring dashboard. That's the right instinct for anyone who has to produce evidence for a regulator, not just a demo for engineering.”
What matters to me isn't the under-100ms guardrail latency, it's that Fiddler maps its governance layer explicitly to NIST AI RMF, ISO/IEC 42001, GDPR, HIPAA, NAIC, SR 11-7, and the EU AI Act by name. Most observability vendors give you dashboards and leave the framework translation as homework for compliance. Fiddler seems to have decided that mapping is part of the product, which changes who internally can actually use it.
Air-gapped and VPC deployment plus AWS GovCloud support tells me they've sat across the table from a government or financial services security review before, not just a data science team. SOC 2 Type II and HIPAA compliance are table stakes at this point, but pairing them with in-environment evaluation models matters for anyone whose data-sharing agreements forbid sending prompts to a third-party API at all.
The tradeoff: hierarchical trace-level visibility down to span level is powerful for post-incident review, but it also means my audit trail is only as good as our instrumentation discipline going in. Governance tooling doesn't replace the process work of deciding who signs off on a policy violation.
Positioning against the 'Trust Tax' of external LLM-as-a-Judge services stakes out a cost-and-control argument few observability vendors make this directly.
Air-gapped, VPC, and GovCloud deployment options match how regulated-industry compliance teams actually gate AI systems into production.
Tier-1 SageMaker integration and NVIDIA NIM support plus LLM gateway integration for coding agents cover both ML and agentic estates in one plane.
Framework-agnostic support for LangGraph, LangChain, and OpenTelemetry limits lock-in, but a control plane this deep becomes the system of record you can't easily rip out.
80+ metrics and framework mapping to seven named regimes shows governance built as core architecture, not a reporting layer stapled on top.
Compliance leaders in regulated sectors who need governance evidence mapped to named frameworks, not just model performance dashboards.
Your organization hasn't yet standardized agent instrumentation, since the audit value here depends on consistent trace-level data going in.
Developer publishes a real rate, $0.002 per trace, and the contract auto-renews at then-current price.
“Fiddler publishes a unit rate on Developer — $0.002 per trace — so the meter runs on volume rather than seats. The terms auto-renew at then-current price and want 30 days' notice to cancel.”
Read the terms before the pricing page. Subscriptions auto-renew at Fiddler's then-current price. Cancelling takes 30 days' notice, self-serve, from a Change/Cancel Membership page. Fourteen-day trial, then monthly billing. Fees are non-refundable.
The rate is public. Developer bills $0.002 per trace. Their own Large preset runs 100,000 traces a day — $200 a day, roughly $73K a year. Enterprise adds VPC and on-premise deployment, and I couldn't find a price for it.
Against Arize they sell cost shape: flat steps rather than a per-evaluation line. The TCO Calculator for Evaluations does that math for you. But you set the inputs — it defaults to 50,000 tokens per trace and folds in a missed-incident cost you estimate yourself. Recompute it on your own token counts.
Card or PayPal at signup keeps the entry tier clear of procurement; Enterprise adds a named CSM and customized onboarding.
Thirty days' notice and a self-serve Change/Cancel Membership page, against auto-renewal at then-current price.
Developer lists $0.002 per trace and Free is listed at no cost; only Enterprise routes through sales.
The TCO Calculator for Evaluations produces a figure, but its total includes a missed-incident cost the buyer estimates.
Developer spend is computable straight from the trace rate, but in-environment evaluation adds infrastructure you provision.
Finance owners who need a unit price before they budget an AI observability line.
Buyers who need a fixed annual figure that cannot move at renewal.
Maps to seven frameworks on paper — but mapping isn't the same as an audit trail my examiners will accept.
“Fiddler gives me governance language that lines up with NIST AI RMF, ISO/IEC 42001, and SR 11-7, which reads well in a board deck. What I actually need on day 90 is whether that mapping produces evidence artifacts my auditors can pull without a Fiddler engineer walking them through it.”
Alignment to NIST AI RMF, ISO/IEC 42001, GDPR, HIPAA, NAIC, SR 11-7, and the EU AI Act is the right list for a model risk committee. But 'aligned with' is a marketing verb until I see the actual control mapping document — which framework clause maps to which dashboard, which gets exported as a PDF for my examiner. SOC 2 Type II and HIPAA certs on the vendor itself are table stakes I'll verify independently regardless.
What worries me for the compliance function specifically: this looks built by ML engineers for ML engineers, then governance got layered on top. Trace-level telemetry is an engineering concept; my job needs it translated into an incident log with timestamps, remediation owner, and sign-off fields. I couldn't find evidence of a purpose-built compliance reporting workspace distinct from the observability dashboards.
Air-gapped and GovCloud deployment is genuinely useful for the DTCC and Navy-type use cases named on the site — that's a real differentiator for regulated procurement. The tier that carries enterprise-grade governance and on-premise deployment is also the one Fiddler leaves unpriced, so the build I'd actually put in front of an examiner still starts with a sales call.
Guardrails firing under 100ms is a strong operational number, but a compliance officer needs the audit output, not just the block event.
Docs and a pricing page exist, but framework mapping reads high-level; I'd want clause-by-clause control documentation for an actual audit.
Free and Developer both carry a published price, so only the tier holding enterprise-grade governance opens a budget cycle with a sales conversation.
80+ metrics and hierarchical trace-to-span visibility give real depth once instrumented, useful for escalating findings up a governance chain.
Framework-agnostic support for LangGraph, LangChain, and OpenTelemetry helps engineering, but I found no dedicated compliance-officer workflow separate from the ML monitoring views.
A compliance officer at a bank, insurer, or government agency already running production agents who needs technical guardrails to point to during an examination.
Skip this if you need a self-serve audit trail generator rather than an observability platform your engineering team configures on your behalf.
Serious control plane for agents, but this isn't a tool you poke around in on day one.
“Fiddler AI is built for people whose job is watching dashboards, not clicking through a friendly setup wizard. That's fine, but it means the first ten minutes are going to feel like a briefing, not a welcome.”
Nobody's onboarding into hierarchical trace-to-span visibility in an afternoon. This is the kind of tool where week one is spent wiring up LangGraph or OpenTelemetry integrations before you see a single useful chart, and the site's own framework-agnostic pitch confirms that setup is the job, not a formality.
The Guardrails free tier is the one place I'd actually poke around without asking anyone's permission, and it's smart of them to let a PII/toxicity check run standalone under 100ms before you're locked into anything. That's a real day-one moment. Everything past that — Centor Models, the coding agent control plane, air-gapped deployment — reads like stuff you configure once with an implementation team and then live inside for months.
Mobile isn't mentioned anywhere I looked, which for a monitoring and governance dashboard people check between meetings feels like a real gap, not a nitpick. Whether the 80+ metrics library stays legible three months in versus becoming another dashboard nobody trusts depends entirely on how well the alerting is tuned, and that's invisible from the outside.
80+ out-of-the-box metrics and a TCO calculator suggest real product investment, but I found no detail on dashboard-level micro-copy or empty states.
Hierarchical visibility down to session, agent, trace, and span is powerful but demands upfront instrumentation discipline to pay off.
Platforms are listed as web only, and I couldn't find any mention of a mobile app for checking alerts on the go.
The free Guardrails tier gives an easy entry point, but full platform setup clearly requires framework integration work before value shows up.
Sub-100ms guardrail detection and SOC 2 Type II certification point to a platform built for production trust, not a beta feel.
Enterprise teams already running agents in production who need a dedicated person watching the control plane daily.
Skip it if you want a lightweight tool you can glance at from your phone between meetings.
Renamed the trust models mid-flight. Worth asking why.
“Centor Models used to be called Fiddler Trust Models. A rename this late in a product's life is either a rebrand for clarity or a sign the positioning wasn't landing.”
Fiddler's been an ML monitoring vendor for years before agents existed as a category. That's a real track record, not a pivot story invented for this cycle. But the Centor Models rename — previously Trust Models — makes me want the changelog behind that decision. Naming churn on your flagship differentiator is a tell worth watching, not dismissing.
Customer list includes Mastercard, DTCC, U.S. Navy. Regulated-industry names lend weight the '98% cost reduction' marketing math doesn't earn on its own — that figure needs an auditor, not a landing page.
SAP-style control planes tend to calcify. Once you've wired LangGraph and OpenTelemetry traces through their hierarchy, span-to-session, unwinding that instrumentation is real engineering work, not a config change. Fine if Fiddler keeps shipping. Less fine if the agent-observability lane gets crowded and they get acquired for the ML monitoring base instead.
In-environment evaluation models avoiding external LLM calls is a real architectural choice, not just a feature checkbox, against the obvious LLM-as-judge alternative.
OpenTelemetry and framework-agnostic support help, but hierarchical span/trace instrumentation this deep is costly to unwind if you switch.
Docs and blog are visible and the company predates the current agent wave, though I couldn't find a changelog to confirm active cadence.
Named customers (Mastercard, DTCC, Navy) ground the claims, but the 98% TCO figure is company math, not third-party verified.
SOC 2 Type II and HIPAA certifications plus named framework alignment are concrete; the Centor/Trust Models rename raises a question about product stability.
Enterprises already committed to agent infrastructure who need audit-grade governance mapped to named regulatory frameworks.
Skip if you want to validate the platform quickly before wiring in deep trace instrumentation.
Common questions answered by our AI research team
Fiddler Centor Models are purpose-built models that power Fiddler's guardrails and evaluations directly in-environment, avoiding external LLM API calls while remaining versatile and secure.
Fiddler's guardrails detect hallucinations, toxicity, PII/PHI leakage, and prompt injection attacks in under 100ms, with in-environment enforcement latency under 80ms.
Yes. Fiddler provides inline enforcement that detects and redacts PII, PHI, and secrets on the same request, before a prompt reaches the model and before a response reaches the developer.
Yes. Fiddler AI Observability can run securely within Amazon SageMaker Studio as part of its partnership as an AWS Preferred Partner.
Yes. Fiddler Centor Models run in-environment with no external LLM API calls, reducing evaluation TCO by up to 98% compared to using foundation models.
Company
Fiddler Labs Inc.Founded
2018Pricing
Usage-basedFree Trial
AvailableFree Plan
Available




Fiddler AI is a San Francisco-based provider of an AI observability platform that monitors, explains, and governs machine learning models in production.