Skip to content

OpenAI API Scale Tier SLAOpenAI · AI models · OpenAI · AI models

AI modelsChecked 15h ago

Promise vs reality

Downtime used against the allowance

0 min of major incidents20 min allowed in 14 days

0 major and 0 minor incidents. 0 min degraded (not counted). History covers 14 of 365 days.

Major incident minutes per month

Allowed per month: 44 min

All incidents, weighted: 99.824% · Downtime
Promised
99.9%
Observed, 365d
100%
Major incident time
0 min
Credit on first breach
n/a

Commitments and credits

  • Scale Tier: GPT-5.5Scale Tier token units for GPT-5.5 (Enterprise customers, min 30-day purchase). Latency SLA: 99% > 50 tokens per second, calculated as p50 request latency on a per 5 minute basis.99.9%

    Monthly uptime below, credit of the bill

    No credit schedule published for this tier.

  • Scale Tier: GPT-5.4 miniScale Tier token units for GPT-5.4 mini (Enterprise customers, min 30-day purchase). Latency SLA: 99% > 100 tokens per second, calculated as p50 request latency on a per 5 minute basis.99.9%

    Monthly uptime below, credit of the bill

    No credit schedule published for this tier.

  • Scale Tier: GPT-5.4 (excludes long-context >272K)Scale Tier token units for GPT-5.4 (excludes long-context >272K) (Enterprise customers, min 30-day purchase). Latency SLA: 99% > 50 tokens per second, calculated as p50 request latency on a per 5 minute basis.99.9%

    Monthly uptime below, credit of the bill

    No credit schedule published for this tier.

  • Scale Tier: GPT-5.2Scale Tier token units for GPT-5.2 (Enterprise customers, min 30-day purchase). Latency SLA: 99% > 50 tokens per second, calculated as p50 request latency on a per 5 minute basis.99.9%

    Monthly uptime below, credit of the bill

    No credit schedule published for this tier.

  • Scale Tier: GPT-5.1Scale Tier token units for GPT-5.1 (Enterprise customers, min 30-day purchase). Latency SLA: 99% > 50 tokens per second, calculated as p50 request latency on a per 5 minute basis.99.9%

    Monthly uptime below, credit of the bill

    No credit schedule published for this tier.

  • Scale Tier: GPT-5Scale Tier token units for GPT-5 (Enterprise customers, min 30-day purchase). Latency SLA: 99% > 50 tokens per second, calculated as p50 request latency on a per 5 minute basis.99.9%

    Monthly uptime below, credit of the bill

    No credit schedule published for this tier.

  • Scale Tier: GPT-5 miniScale Tier token units for GPT-5 mini (Enterprise customers, min 30-day purchase). Latency SLA: 99% > 80 tokens per second, calculated as p50 request latency on a per 5 minute basis.99.9%

    Monthly uptime below, credit of the bill

    No credit schedule published for this tier.

  • Scale Tier: GPT-4.1 (excludes long-context >128K)Scale Tier token units for GPT-4.1 (excludes long-context >128K) (Enterprise customers, min 30-day purchase). Latency SLA: 99% > 80 tokens per second, calculated as p50 request latency on a per 5 minute basis.99.9%

    Monthly uptime below, credit of the bill

    No credit schedule published for this tier.

  • Scale Tier: GPT-4.1 mini (excludes long-context >128K)Scale Tier token units for GPT-4.1 mini (excludes long-context >128K) (Enterprise customers, min 30-day purchase). Latency SLA: 99% > 90 tokens per second, calculated as p50 request latency on a per 5 minute basis.99.9%

    Monthly uptime below, credit of the bill

    No credit schedule published for this tier.

  • Scale Tier: GPT-4.1 nano (excludes long-context >128K)Scale Tier token units for GPT-4.1 nano (excludes long-context >128K) (Enterprise customers, min 30-day purchase). Latency SLA: 99% > 100 tokens per second, calculated as p50 request latency on a per 5 minute basis.99.9%

    Monthly uptime below, credit of the bill

    No credit schedule published for this tier.

  • Scale Tier: GPT-4.1 fine tuningScale Tier token units for GPT-4.1 fine tuning (Enterprise customers, min 30-day purchase). Latency SLA: 99% > 80 tokens per second, calculated as p50 request latency on a per 5 minute basis.99.9%

    Monthly uptime below, credit of the bill

    No credit schedule published for this tier.

  • Scale Tier: GPT-4.1 mini fine tuningScale Tier token units for GPT-4.1 mini fine tuning (Enterprise customers, min 30-day purchase). Latency SLA: 99% > 90 tokens per second, calculated as p50 request latency on a per 5 minute basis.99.9%

    Monthly uptime below, credit of the bill

    No credit schedule published for this tier.

  • Scale Tier: o3Scale Tier token units for o3 (Enterprise customers, min 30-day purchase). Latency SLA: 99% > 80 tokens per second, calculated as p50 request latency on a per 5 minute basis.99.9%

    Monthly uptime below, credit of the bill

    No credit schedule published for this tier.

  • Scale Tier: o4-miniScale Tier token units for o4-mini (Enterprise customers, min 30-day purchase). Latency SLA: 99% > 90 tokens per second, calculated as p50 request latency on a per 5 minute basis.99.9%

    Monthly uptime below, credit of the bill

    No credit schedule published for this tier.

  • Scale Tier: GPT-4oScale Tier token units for GPT-4o (Enterprise customers, min 30-day purchase). Latency SLA: 99% > 80 tokens per second, calculated as p50 request latency on a per 5 minute basis.99.9%

    Monthly uptime below, credit of the bill

    No credit schedule published for this tier.

  • Scale Tier: GPT-4o miniScale Tier token units for GPT-4o mini (Enterprise customers, min 30-day purchase). Latency SLA: 99% > 90 tokens per second, calculated as p50 request latency on a per 5 minute basis.99.9%

    Monthly uptime below, credit of the bill

    No credit schedule published for this tier.

  • Scale Tier: GPT-4o mini fine tuningScale Tier token units for GPT-4o mini fine tuning (Enterprise customers, min 30-day purchase). Latency SLA: 99% > 90 tokens per second, calculated as p50 request latency on a per 5 minute basis.99.9%

    Monthly uptime below, credit of the bill

    No credit schedule published for this tier.

  • Scale Tier: o1Scale Tier token units for o1 (Enterprise customers, min 30-day purchase). Latency SLA: 99% > 80 tokens per second, calculated as p50 request latency on a per 5 minute basis.99.9%

    Monthly uptime below, credit of the bill

    No credit schedule published for this tier.

  • Scale Tier: o3-miniScale Tier token units for o3-mini (Enterprise customers, min 30-day purchase). Latency SLA: 99% > 90 tokens per second, calculated as p50 request latency on a per 5 minute basis.99.9%

    Monthly uptime below, credit of the bill

    No credit schedule published for this tier.

Terms

Measured
Uptime window not stated; credits are for 'the calendar month of that Scale Tier token unit purchase'. Latency SLA calculated as p50 request latency on a per 5 minute basis (older enterprise agreements: per minute).

Not covered

  • Credit amounts / schedule not published
  • Long-context requests excluded for GPT-4.1 family (>128K) and GPT-5.4 (>272K)
  • Only models released before GPT-5.6 (newer models use Reserved Tier)
  • Enterprise customers only; via sales order form
Higher reliability: Scale Tier traffic offers a 99.9% uptime SLA and prioritized compute. ... GPT-5.5 50,000 TPM $750.00 per unit/day N/A 99.9% 99% > 50 tokens per second
OpenAI, SLA. Read the SLA

Published: 99.9% uptime SLA plus per-model latency SLA for every listed model. FAQ: 'What happens if the latency and uptime SLA are both violated? You will be credited with the greater of the two SLA amounts for the calendar month' - but the credit amounts themselves are not published. Page says 'Scale Tier is available on models released before GPT-5.6. For GPT-5.6 and future model releases, see Reserved Tier'. Reserved Tier FAQ (https://openai.com/api-reserved-tier/): 'SLA for the service tier you use for your requests, such as Fast mode or Standard, will apply' - no separate number. openai.com pages were read via a real browser (curl is blocked by Cloudflare).

Evidence: incidents, last 365 days

No incidents on file for this service.

SLA changes

  • Sep23OpenAI API Scale TierGPT-5.5 Scale Tier latency SLA listed as '99% > 100 tokens per second' in the 2026-06-11 snapshot; the live page on 2026-09-23 lists '99% > 50 tokens per second'. Exact change date unknown (between those dates).Source
  • Aug10OpenAI API Scale TierLatency SLA methodology changed from 'average request latency on per minute basis across the month' (e.g. GPT-4o 99% > 25 tokens/s in Aug 2024) to 'p50 request latency on a per 5 minute basis'. Change occurred between the 2025-04-04 and 2025-08-10 archived snapshots; uptime stayed 99.9%.Source

Other ai models SLAs

Promised in the SLAObserved on the vendor's status pagePartial historyShortfall

Summaries of published SLAs; the contract you sign governs. Logos via logo.dev; trademarks belong to their owners.

Weekly: SLA changes from cloud, data and AI vendors, Fridays.