Skip to content

Vertex AI (Gemini API on Vertex) SLAGoogle Cloud · AI models · Google Cloud · AI models

AI modelsSLA of Jun 30, 2026Checked 18h ago

Promise vs reality

Downtime used against the allowance

0 min of major incidents5 h allowed in 208 days

0 major and 0 minor incidents. 0 min degraded (not counted). History covers 208 of 365 days.

Major incident minutes per month

Allowed per month: 44 min

All incidents, weighted: 99.488% · Downtime
Promised
99.9%
Observed, 365d
100%
Major incident time
0 min
Credit on first breach
10%

Commitments and credits

  • Gemini Online Inference API (generateContent / streamGenerateContent)Gemini Online Inference API on Gemini Enterprise Agent Platform (95% for models designated for shorter availability)99.5%

    Monthly uptime below, credit of the bill

    1. Below 99.5%10%
    2. Below 99%25%
    3. Below 95%50%
  • Provisioned Throughput latency (Gemini 2.5 Pro/Flash/Flash-lite, global endpoint)Monthly Latency Target Attainment (p50 TPS vs. target: 60/80/110 TPS)99%

    Monthly uptime below, credit of the bill

    1. Below 99%10%
    2. Below 95%25%
    3. Below 90%50%
  • Vertex AI Platform: Training, Deployment, Batch PredictionVertex AI Platform SLA (https://cloud.google.com/vertex-ai/sla, last modified February 12, 2026)99.9%

    Monthly uptime below, credit of the bill

    1. Below 99.9%10%
    2. Below 99%25%
    3. Below 95%50%
  • Vertex AI Platform: Custom Model Online Prediction (2+ nodes), PipelinesVertex AI Platform SLA99.5%

    Monthly uptime below, credit of the bill

    1. Below 99.5%10%
    2. Below 99%25%
    3. Below 95%50%

Terms

Measured
Calendar month per Project; Downtime = more than 5% server-side 5xx Error Rate over 5+ consecutive minutes
Credit cap
50% of the amount due for the Covered Service for the month
Claim window
Within 30 days from the time Customer becomes eligible
How to claim
Notify Google technical support; provide log files/identifying info showing Downtime Periods and when they occurred

Not covered

  • Features designated pre-general availability
  • Features excluded from the SLA in the Documentation
  • Errors caused by factors outside Google's reasonable control
  • Errors from Customer's or third-party software or hardware
  • Requests using Grounding with Google Search
  • Deadlines set shorter than the server default (deadline_exceeded)
Training, Deployment, and Batch Prediction >= 99.9% ... 99% to < 99.9% 10% 95% to < 99% 25% < 95% 50%
Google Cloud, Last modified June 30, 2026. Read the SLA

Product renamed on the SLA page to 'Gemini Online Inference API on Gemini Enterprise Agent Platform'. The latency tier's uptime field holds the 99% latency attainment, not availability.

Evidence: incidents, last 365 days

No incidents on file for this service.

Other ai models SLAs

Promised in the SLAObserved on the vendor's status pagePartial historyShortfall

Summaries of published SLAs; the contract you sign governs. Logos via logo.dev; trademarks belong to their owners.

Weekly: SLA changes from cloud, data and AI vendors, Fridays.