Vertex AI (Gemini API on Vertex) SLAGoogle Cloud · AI models · Google Cloud · AI models
Promise vs reality
Downtime used against the allowance
0 min of major incidents5 h allowed in 208 days
0 major and 0 minor incidents. 0 min degraded (not counted). History covers 208 of 365 days.
Major incident minutes per month
Allowed per month: 44 min
- Promised
- 99.9%
- Observed, 365d
- 100%
- Major incident time
- 0 min
- Credit on first breach
- 10%
Commitments and credits
- Gemini Online Inference API (generateContent / streamGenerateContent)Gemini Online Inference API on Gemini Enterprise Agent Platform (95% for models designated for shorter availability)99.5%
Monthly uptime below, credit of the bill
- Below 99.5%10%
- Below 99%25%
- Below 95%50%
- Provisioned Throughput latency (Gemini 2.5 Pro/Flash/Flash-lite, global endpoint)Monthly Latency Target Attainment (p50 TPS vs. target: 60/80/110 TPS)99%
Monthly uptime below, credit of the bill
- Below 99%10%
- Below 95%25%
- Below 90%50%
- Vertex AI Platform: Training, Deployment, Batch PredictionVertex AI Platform SLA (https://cloud.google.com/vertex-ai/sla, last modified February 12, 2026)99.9%
Monthly uptime below, credit of the bill
- Below 99.9%10%
- Below 99%25%
- Below 95%50%
- Vertex AI Platform: Custom Model Online Prediction (2+ nodes), PipelinesVertex AI Platform SLA99.5%
Monthly uptime below, credit of the bill
- Below 99.5%10%
- Below 99%25%
- Below 95%50%
Terms
- Measured
- Calendar month per Project; Downtime = more than 5% server-side 5xx Error Rate over 5+ consecutive minutes
- Credit cap
- 50% of the amount due for the Covered Service for the month
- Claim window
- Within 30 days from the time Customer becomes eligible
- How to claim
- Notify Google technical support; provide log files/identifying info showing Downtime Periods and when they occurred
Not covered
- Features designated pre-general availability
- Features excluded from the SLA in the Documentation
- Errors caused by factors outside Google's reasonable control
- Errors from Customer's or third-party software or hardware
- Requests using Grounding with Google Search
- Deadlines set shorter than the server default (deadline_exceeded)
Training, Deployment, and Batch Prediction >= 99.9% ... 99% to < 99.9% 10% 95% to < 99% 25% < 95% 50%
Product renamed on the SLA page to 'Gemini Online Inference API on Gemini Enterprise Agent Platform'. The latency tier's uptime field holds the 99% latency attainment, not availability.
Evidence: incidents, last 365 days
No incidents on file for this service.
Other ai models SLAs
ServiceDowntime vs allowed
Amazon BedrockAWS99.9% / 99.178%
Vertex AI (Gemini API on Vertex)Google Cloud99.9% / 100%
Azure OpenAI in Foundry ModelsMicrosoft Azure99.9% / n/a
Claude API Priority TierAnthropic99.5% / 98.8%
Claude API (Standard tier)Anthropicnone / 98.76%
Claude Enterprise (claude.ai)Anthropicnone / 98.62%
Cohere API (SaaS: Command, Embed, Rerank)Coherenone / 100%
Cohere private deployments / Model Vault / NorthCoherenone / n/aNo public SLA
Hugging Face Inference EndpointsHugging Facenone / n/aNo public SLA
Promised in the SLAObserved on the vendor's status pagePartial historyShortfall
Summaries of published SLAs; the contract you sign governs. Logos via logo.dev; trademarks belong to their owners.