by OpenAI · GPT-5 family · best for legacy flagship on a migration path
GPT-5 is OpenAI's original unified reasoning flagship, released 2025-08-07 — the first single SKU with a configurable reasoning effort dial (minimal/low/medium/high) and the model that established the 400K context standard for the family. It set state-of-the-art at launch on AIME 2025 (94.6% no tools), SWE-bench Verified (74.9%), Aider Polyglot (88%), and MMMU (84.2%). It remains GA and attractively priced at $1.25/$10, but a 2024-09 knowledge cutoff and the newer GPT-5.4/5.5 generation make it a legacy choice. The one-sentence buyer's take: a strong model now firmly in the migration lane — fine for in-flight builds, wrong for new ones.
| Benchmark | Score | Source |
|---|---|---|
| Humanity's Last Exam | 42% | vellum.ai 2025-08-07T00:00:00.000Z |
| MMMU | 84.2% | openai.com 2025-08-07T00:00:00.000Z |
| AIME 2025 | 94.6% | openai.com 2025-08-07T00:00:00.000Z |
| GPQA Diamond | 89.4% | vellum.ai 2025-08-07T00:00:00.000Z |
| Aider Polyglot | 88% | openai.com 2025-08-07T00:00:00.000Z |
| SWE-bench Verified | 74.9% | vellum.ai 2025-08-07T00:00:00.000Z |
Six personas, six verdicts — the same panel that reviews every product on TopReviewed.
“A capable legacy flagship — the only real question is when to migrate and to which target, not whether.”
GPT-5 is a legacy flagship — still GA, still capable, but not the right choice for new builds. The strategic question for any team still on it is how soon to migrate and to which target: GPT-5.4 is the obvious quality-priority default (better capability, ~2x price), or GPT-5.4 mini for cost-priority workloads (better-than-GPT-5 coding at $0.75/$4.50). The chat-latest alias deprecation forces a 2026-10-23 deadline for teams using it; dated snapshots remain longer but won't see updates. The one durable reason to stay is fine-tuning, which the newer generation lacks. Plan the migration, don't drift.
“Its market is shrinking to fine-tuning and inertia — the newer generation owns the new-build narrative on both quality and cost.”
Strategically, GPT-5's market has narrowed to two pockets: teams that fine-tune (the 5.4/5.5 generation does not offer it) and teams whose migration cost exceeds the upgrade value. On the open market it has no competitive wedge — GPT-5.4 beats it on quality at ~2x price and GPT-5.4 mini beats it on coding at lower cost. Its positioning is purely transitional. Market timing has fully passed; it is a 2025 model competing in a 2026 frontier. Differentiation survives only in the fine-tuning niche.
“Genuinely cheap for the capability, but GPT-5.4 mini at $0.75/$4.50 with better coding is the comparison most migrations actually make.”
$1.25/$10 is genuinely cheap for the capability level, which explains the stickiness. Cached input drops 90%, Batch halves both sides. For high-volume backend pipelines, GPT-5 may still pencil out below GPT-5.4 base. But GPT-5.4 mini at $0.75/$4.50 with comparable-or-better coding is the real comparison — most cost-aware migrations move from GPT-5 to mini, not to GPT-5.4 base. The cost case for staying is shrinking; migration is the right financial call within a quarter unless fine-tuning ties you here.
“Recognizably the modern API shape, minus the modern tools — no computer use, no apply_patch. For code agents, that gap is real.”
Developers on GPT-5 recognize the entire current API shape — Responses, structured outputs, reasoning-effort dial. What's missing is the modern tool surface: no computer use, no apply_patch, no skills. For code-edit agents specifically, the move to the GPT-5.4 family is high-value because apply_patch is transformative for edits. SDK quality is identical to the current generation; promotion paths are clean. The chat-latest deprecation is the most concrete reason to migrate now — dated snapshots are safer but not future-proof. Fine-tuning support is the one thing developers would lose by moving.
“ChatGPT users no longer meet GPT-5 directly — where it surfaces in unmigrated apps it's good, but the stale knowledge shows.”
End users on ChatGPT no longer encounter GPT-5 directly — it has been replaced by 5.4 and 5.5 in default routing. Where it surfaces is in API-backed products that haven't migrated. Quality remains good for general chat; the 2024-09 knowledge cutoff occasionally surfaces in time-sensitive questions. Refusal patterns are slightly more conservative than the current generation. For most users the model is effectively invisible — they're getting GPT-5.4 or GPT-5.5 in ChatGPT and don't realize GPT-5 is still GA in the API.
“The launch-day SOTA headlines are 20 months old — quoting GPT-5's 2025 benchmarks today is the polite way to say 'don't build here.'”
The adversarial read: GPT-5's benchmark sheet is a 2025 time capsule. The AIME 94.6% and SWE-bench 74.9% were genuinely SOTA at launch, but GPT-5.4 nano — four tiers down and cheaper — now beats GPT-5 mini on coding, which tells you where the base sits. The model is still sold partly on launch-day prestige. The honest reasons to use it today are narrow and unglamorous: sunk-cost migration and fine-tuning. The alias deprecation (2026-10-23) is a forcing function the marketing soft-pedals. Nothing here is misleading — it's just an old model wearing its launch medals.
gpt-5-chat-latest alias is deprecated, shutdown 2026-10-23 — only dated snapshots stay stable.The full research notes behind this review — verified against primary sources.
OpenAI discloses no parameter counts, layer counts, or dense/MoE structure for GPT-5 — all null. It is a unified reasoning model with text + image input and text-only output, o200k_base tokenizer, and a 400K context window. The architectural milestone it introduced was conceptual rather than disclosed: collapsing chat and reasoning into one SKU with an effort dial, a pattern every subsequent GPT-5.x model inherited.
GPT-5 was the launch flagship and remains capable, justifying its 8.8 math and 8.5 reasoning scores — AIME 2025 94.6% (no tools) and GPQA Diamond 89.4% (Pro) were SOTA at release. Coding (8.0) is anchored by SWE-bench Verified 74.9% and Aider Polyglot 88% — strong for 2025 but now behind even GPT-5.4 nano on newer coding evals. Agentic (7.8) is capped by the absence of native computer use, apply_patch, and skills (all added in GPT-5.4). Long-context (8.0) reflects the 400K window. Multilingual (8.5), instruction-following (8.7), and function-calling (8.5) are solid. Vision (8.2) and document/OCR (8.0) handle image input. Real-time data (5.5) is the weak axis: a 2024-09 cutoff is now ~20 months stale, so time-sensitive answers lean heavily on the web-search tool.
| Benchmark | Score | vs Successor (GPT-5.4) | vs Top Competitor (at release) | Source |
|---|---|---|---|---|
| AIME 2025 (no tools) | 94.6% | superseded | leader at release | OpenAI |
| GPQA Diamond (Pro) | 89.4% | below GPT-5.4 (92.8%) | leader at release | Vellum |
| SWE-bench Verified | 74.9% | below newer gen | leader at release | Vellum |
| Aider Polyglot | 88% | superseded | leader at release | OpenAI |
| MMMU | 84.2% | below GPT-5.4 MMMU-Pro context | leader at release | OpenAI |
| HLE (Pro, with tools) | 42% | peer with newer no-tools figures | leader at release | Vellum |
AIME/GPQA/HLE figures reflect GPT-5 Pro extended reasoning where noted; SWE-bench and Aider are GPT-5 with thinking. MMLU-Pro, MATH-500, HumanEval, LiveCodeBench, IFEval, BBH, Tau-bench, SimpleQA, and LMArena have no clean GPT-5-base figure recorded here and are null.
GPT-5 serves at roughly 90 tokens/sec at high reasoning on OpenAI's API, with time-to-first-token around 1 second at low/default effort. latency_tier is medium — interactive for chat, slower at high reasoning effort. It is generally faster to first token than the heavier GPT-5.5 because the model is smaller and older, but it lacks the serving-infrastructure improvements that shipped with GPT-5.5.
| Surface | Cost | Notes |
|---|---|---|
| API input | $1.25 / 1M tok | |
| API output | $10.00 / 1M tok | |
| Cached input | $0.125 / 1M tok | 90% discount |
| Batch (in/out) | $0.625 / $5.00 | 50% off, 24h SLA |
| Direct UI | $20/mo (Plus) | superseded in default routing |
| Free tier | none | |
| Fine-tuning | available | GPT-5 supports fine-tuning (newer 5.4/5.5 do not) |
| Deprecation | gpt-5-chat-latest shutdown 2026-10-23 |
dated snapshots remain |
| Rate limits | ~10,000 RPM / ~30M TPM | tier-dependent |
Reasoning-token note: hidden reasoning tokens bill as output tokens at higher effort tiers.
API-only via the Responses API; not open-weights, license Proprietary, not self-hostable. Cloud-managed via Azure OpenAI and Azure AI Foundry; OpenRouter proxies it. Data residency covers US and EU. Notably, GPT-5 supports fine-tuning where the newer GPT-5.4/5.5 generation does not — a real reason some teams remain on it. The tool surface predates computer use, apply_patch, and skills.
Governed by OpenAI's Preparedness Framework. No training on API inputs by default; opt-out and zero-retention available for enterprise. Compliance covers SOC2, GDPR, CCPA, and HIPAA (BAA). Content moderation is built in. Refusal behavior is slightly more conservative than the current GPT-5.4/5.5 generation.
First-party SDKs in Python, TypeScript, Java, Go, and .NET, plus the OpenAI Agents SDK. Framework integrations span LangChain, LlamaIndex, Vercel AI SDK, and Pydantic AI. It was the ChatGPT default through late 2025 and powered Microsoft Copilot during that window. Popularity tier: growing (declining as migration proceeds), still widely embedded in API-backed products.
No — use GPT-5.4 (quality) or GPT-5.4 mini (cost). GPT-5 is a legacy flagship for in-flight workloads.
The base model is not on the deprecation list, but the gpt-5-chat-latest alias shuts down 2026-10-23. Pin a dated snapshot if you stay.
Two reasons: a validated in-flight workload where migration cost exceeds the benefit, or fine-tuning (which GPT-5.4/5.5 do not offer).
The cutoff is 2024-09 — about 20 months. Time-sensitive queries need the web-search tool.
Mostly prompt re-validation; the API shape is identical across the family, so promotion is a model-name change plus testing.
No, not by API default; enterprise opt-out and zero-retention exist.
recommended cost-priority successor; cheaper at $0.75/$4.50 with better coding.
recommended quality-priority successor; better capability at ~2x the price plus computer use and apply_patch.
generational peer from the 2025 cohort; comparable role and pricing.
generational peer; comparable pricing and capability for the era.
Primary references used to verify this review.
Does not train on API inputs by default
Last verified 2026-05-27