OpenAI API Scale Tier SLAOpenAI · AI models · OpenAI · AI models
Promise vs reality
Downtime used against the allowance
0 major and 0 minor incidents. 0 min degraded (not counted). History covers 14 of 365 days.
Major incident minutes per month
Allowed per month: 44 min
- Promised
- 99.9%
- Observed, 365d
- 100%
- Major incident time
- 0 min
- Credit on first breach
- n/a
Commitments and credits
- Scale Tier: GPT-5.5Scale Tier token units for GPT-5.5 (Enterprise customers, min 30-day purchase). Latency SLA: 99% > 50 tokens per second, calculated as p50 request latency on a per 5 minute basis.99.9%
Monthly uptime below, credit of the bill
No credit schedule published for this tier.
- Scale Tier: GPT-5.4 miniScale Tier token units for GPT-5.4 mini (Enterprise customers, min 30-day purchase). Latency SLA: 99% > 100 tokens per second, calculated as p50 request latency on a per 5 minute basis.99.9%
Monthly uptime below, credit of the bill
No credit schedule published for this tier.
- Scale Tier: GPT-5.4 (excludes long-context >272K)Scale Tier token units for GPT-5.4 (excludes long-context >272K) (Enterprise customers, min 30-day purchase). Latency SLA: 99% > 50 tokens per second, calculated as p50 request latency on a per 5 minute basis.99.9%
Monthly uptime below, credit of the bill
No credit schedule published for this tier.
- Scale Tier: GPT-5.2Scale Tier token units for GPT-5.2 (Enterprise customers, min 30-day purchase). Latency SLA: 99% > 50 tokens per second, calculated as p50 request latency on a per 5 minute basis.99.9%
Monthly uptime below, credit of the bill
No credit schedule published for this tier.
- Scale Tier: GPT-5.1Scale Tier token units for GPT-5.1 (Enterprise customers, min 30-day purchase). Latency SLA: 99% > 50 tokens per second, calculated as p50 request latency on a per 5 minute basis.99.9%
Monthly uptime below, credit of the bill
No credit schedule published for this tier.
- Scale Tier: GPT-5Scale Tier token units for GPT-5 (Enterprise customers, min 30-day purchase). Latency SLA: 99% > 50 tokens per second, calculated as p50 request latency on a per 5 minute basis.99.9%
Monthly uptime below, credit of the bill
No credit schedule published for this tier.
- Scale Tier: GPT-5 miniScale Tier token units for GPT-5 mini (Enterprise customers, min 30-day purchase). Latency SLA: 99% > 80 tokens per second, calculated as p50 request latency on a per 5 minute basis.99.9%
Monthly uptime below, credit of the bill
No credit schedule published for this tier.
- Scale Tier: GPT-4.1 (excludes long-context >128K)Scale Tier token units for GPT-4.1 (excludes long-context >128K) (Enterprise customers, min 30-day purchase). Latency SLA: 99% > 80 tokens per second, calculated as p50 request latency on a per 5 minute basis.99.9%
Monthly uptime below, credit of the bill
No credit schedule published for this tier.
- Scale Tier: GPT-4.1 mini (excludes long-context >128K)Scale Tier token units for GPT-4.1 mini (excludes long-context >128K) (Enterprise customers, min 30-day purchase). Latency SLA: 99% > 90 tokens per second, calculated as p50 request latency on a per 5 minute basis.99.9%
Monthly uptime below, credit of the bill
No credit schedule published for this tier.
- Scale Tier: GPT-4.1 nano (excludes long-context >128K)Scale Tier token units for GPT-4.1 nano (excludes long-context >128K) (Enterprise customers, min 30-day purchase). Latency SLA: 99% > 100 tokens per second, calculated as p50 request latency on a per 5 minute basis.99.9%
Monthly uptime below, credit of the bill
No credit schedule published for this tier.
- Scale Tier: GPT-4.1 fine tuningScale Tier token units for GPT-4.1 fine tuning (Enterprise customers, min 30-day purchase). Latency SLA: 99% > 80 tokens per second, calculated as p50 request latency on a per 5 minute basis.99.9%
Monthly uptime below, credit of the bill
No credit schedule published for this tier.
- Scale Tier: GPT-4.1 mini fine tuningScale Tier token units for GPT-4.1 mini fine tuning (Enterprise customers, min 30-day purchase). Latency SLA: 99% > 90 tokens per second, calculated as p50 request latency on a per 5 minute basis.99.9%
Monthly uptime below, credit of the bill
No credit schedule published for this tier.
- Scale Tier: o3Scale Tier token units for o3 (Enterprise customers, min 30-day purchase). Latency SLA: 99% > 80 tokens per second, calculated as p50 request latency on a per 5 minute basis.99.9%
Monthly uptime below, credit of the bill
No credit schedule published for this tier.
- Scale Tier: o4-miniScale Tier token units for o4-mini (Enterprise customers, min 30-day purchase). Latency SLA: 99% > 90 tokens per second, calculated as p50 request latency on a per 5 minute basis.99.9%
Monthly uptime below, credit of the bill
No credit schedule published for this tier.
- Scale Tier: GPT-4oScale Tier token units for GPT-4o (Enterprise customers, min 30-day purchase). Latency SLA: 99% > 80 tokens per second, calculated as p50 request latency on a per 5 minute basis.99.9%
Monthly uptime below, credit of the bill
No credit schedule published for this tier.
- Scale Tier: GPT-4o miniScale Tier token units for GPT-4o mini (Enterprise customers, min 30-day purchase). Latency SLA: 99% > 90 tokens per second, calculated as p50 request latency on a per 5 minute basis.99.9%
Monthly uptime below, credit of the bill
No credit schedule published for this tier.
- Scale Tier: GPT-4o mini fine tuningScale Tier token units for GPT-4o mini fine tuning (Enterprise customers, min 30-day purchase). Latency SLA: 99% > 90 tokens per second, calculated as p50 request latency on a per 5 minute basis.99.9%
Monthly uptime below, credit of the bill
No credit schedule published for this tier.
- Scale Tier: o1Scale Tier token units for o1 (Enterprise customers, min 30-day purchase). Latency SLA: 99% > 80 tokens per second, calculated as p50 request latency on a per 5 minute basis.99.9%
Monthly uptime below, credit of the bill
No credit schedule published for this tier.
- Scale Tier: o3-miniScale Tier token units for o3-mini (Enterprise customers, min 30-day purchase). Latency SLA: 99% > 90 tokens per second, calculated as p50 request latency on a per 5 minute basis.99.9%
Monthly uptime below, credit of the bill
No credit schedule published for this tier.
Terms
- Measured
- Uptime window not stated; credits are for 'the calendar month of that Scale Tier token unit purchase'. Latency SLA calculated as p50 request latency on a per 5 minute basis (older enterprise agreements: per minute).
Not covered
- Credit amounts / schedule not published
- Long-context requests excluded for GPT-4.1 family (>128K) and GPT-5.4 (>272K)
- Only models released before GPT-5.6 (newer models use Reserved Tier)
- Enterprise customers only; via sales order form
Higher reliability: Scale Tier traffic offers a 99.9% uptime SLA and prioritized compute. ... GPT-5.5 50,000 TPM $750.00 per unit/day N/A 99.9% 99% > 50 tokens per second
Published: 99.9% uptime SLA plus per-model latency SLA for every listed model. FAQ: 'What happens if the latency and uptime SLA are both violated? You will be credited with the greater of the two SLA amounts for the calendar month' - but the credit amounts themselves are not published. Page says 'Scale Tier is available on models released before GPT-5.6. For GPT-5.6 and future model releases, see Reserved Tier'. Reserved Tier FAQ (https://openai.com/api-reserved-tier/): 'SLA for the service tier you use for your requests, such as Fast mode or Standard, will apply' - no separate number. openai.com pages were read via a real browser (curl is blocked by Cloudflare).
Evidence: incidents, last 365 days
SLA changes
- Sep23OpenAI API Scale TierGPT-5.5 Scale Tier latency SLA listed as '99% > 100 tokens per second' in the 2026-06-11 snapshot; the live page on 2026-09-23 lists '99% > 50 tokens per second'. Exact change date unknown (between those dates).Source
- Aug10OpenAI API Scale TierLatency SLA methodology changed from 'average request latency on per minute basis across the month' (e.g. GPT-4o 99% > 25 tokens/s in Aug 2024) to 'p50 request latency on a per 5 minute basis'. Change occurred between the 2025-04-04 and 2025-08-10 archived snapshots; uptime stayed 99.9%.Source
Other ai models SLAs
Amazon BedrockAWS99.9% / 99.178%
Vertex AI (Gemini API on Vertex)Google Cloud99.9% / 100%
OpenAI API Scale TierOpenAI99.9% / 100%
Claude API Priority TierAnthropic99.5% / 98.8%
Claude API (Standard tier)Anthropicnone / 98.76%
Claude Enterprise (claude.ai)Anthropicnone / 98.62%
Cohere API (SaaS: Command, Embed, Rerank)Coherenone / 100%
Cohere private deployments / Model Vault / NorthCoherenone / n/aNo public SLA
Hugging Face Inference EndpointsHugging Facenone / n/aNo public SLA
Summaries of published SLAs; the contract you sign governs. Logos via logo.dev; trademarks belong to their owners.