Vertex AI 429 “Resource exhausted” errors
Your Gemini calls on Vertex AI (which Google now calls Gemini Enterprise Agent Platform) fail with 429 Resource exhausted, often while your quota dashboard says you have room to spare. This page explains why that happens, what Google recommends, and how to shift traffic to a second provider while Vertex AI is failing, without changing your application code.
Why you get 429s inside your quota
For standard pay-as-you-go traffic, Vertex AI serves Gemini from a pool of capacity shared with other customers. Google describes a 429 as a sign that "your requests are coming in faster than the service can handle them at that moment". What runs out is capacity in the shared pool, not only your own quota. So even a low-volume workload can see 429s, and they often arrive in bursts.
That is why the reports look alike: 429s in production despite being within quota, and errors from a handful of calls to Gemini 2.5 Flash.
What Google recommends, and when it is enough
Try these first. Google's guide to reducing 429 errors lists five fixes:
- Retry with exponential backoff and jitter. Most SDKs already do this.
- Use the global endpoint, which "routes your traffic across a fleet of regions where there may be more availability".
- Buy Provisioned Throughput, the only option isolated from the shared pool.
- Use context caching and shorter prompts to send fewer tokens.
- Smooth out traffic spikes.
If the global endpoint and retries keep your error rate acceptable, you may not need anything else. The rest of this page is for when they don't:
- The global endpoint still returns 429s at times. Users report bursts there too, and retrying against the same pool doesn't help while it lasts.
- Provisioned Throughput costs too much for your volume.
- You'd rather not depend on one vendor's capacity. A comparable model from another provider has capacity of its own.
Add a second provider with sturnus
sturnus is a small open-source proxy that runs next to your application. It speaks the OpenAI API and sends each request to whichever of your configured providers is fastest and healthy right now. Here, Gemini on the Vertex AI global endpoint is paired with an OpenAI model:
listen = "127.0.0.1:4000"
# Vertex AI global endpoint. Auth uses GKE Workload Identity,
# so no API key is needed.
[provider.vertex]
vertex_ai = { project_id = "my-gcp-project", location = "global" }
[provider.openai]
base_url = "https://api.openai.com/v1"
api_key = "${OPENAI_API_KEY}"
[model]
flash = [
{ provider = "vertex", model = "google/gemini-3.8-flash" },
{ provider = "openai", model = "gpt-6-luna" },
]
Your code then asks for the flash alias:
from openai import OpenAI
client = OpenAI(base_url="http://127.0.0.1:4000/v1", api_key="unused")
response = client.chat.completions.create(
model="flash",
messages=[{"role": "user", "content": "Hello"}],
)
The models under one alias should be interchangeable for your use case. Test your prompts against both before you mix them. If you'd rather stay with Gemini, Gemini through Google AI Studio is another option: it's a separate service with its own quota, and sturnus supports it too.
What happens when Vertex AI starts returning 429s
sturnus keeps two moving averages for each provider: time to first token, and success rate. A 429 lowers Vertex AI's success rate, and sturnus divides latency by success rate. So a provider that fails often scores as slower, and its share of new requests shrinks. It never drops below 1% of traffic, and those requests act as probes. When Vertex AI stops returning 429s, its score recovers and it wins traffic back without you doing anything.
sturnus does not retry. The 429 goes back to your client unchanged, and your SDK's retry logic takes over. The OpenAI Python SDK, for example, retries 429s with backoff by default. Each retry is a new routing decision, made after Vertex AI's share has already dropped, so retries tend to land on OpenAI.
Running it on GKE
Run sturnus as a native sidecar container in the same pod as your application, listening on 127.0.0.1:4000. It gets Vertex AI credentials from the GKE metadata server through Workload Identity and refreshes the tokens itself. Give the pod's Kubernetes service account permission to call Vertex AI (for example the roles/aiplatform.user role), and pass the OpenAI key to the sidecar as an environment variable from a Kubernetes Secret.
Each pod's sidecar routes on what it observes, so there is no Redis, database or shared state to run. Prometheus metrics for each provider, and a /status endpoint showing live scores, tell you where traffic is going and why. The README covers deployment details.
What to know before you use it
- It uses the OpenAI API format. sturnus talks to Vertex AI's OpenAI-compatible Chat Completions endpoint. If your code uses the
google-genaiSDK, you would need to switch those calls to an OpenAI SDK. - Traffic can move between models. While Vertex AI is failing, most requests go to the other provider's model. For multi-turn conversations, the
x-session-affinityheader keeps a conversation on one provider until that provider starts failing. - It routes on latency and errors, not cost. If one provider is cheaper, sturnus won't prefer it for that reason.
- Each sidecar learns from its own traffic. It only sees the errors and latency of the requests it sends. When a sidecar starts, every provider gets a small probe share until it has data.
Get sturnus
sturnus is a single static binary, MIT licensed, published as a container for amd64 and arm64.
docker run -v ./config.toml:/config.toml -p 4000:4000 ghcr.io/sturnus-dev/sturnus:latest
Install options, the full configuration reference and metrics are on the home page and in the README. Questions and bug reports go to GitHub issues.