sturnus

Vertex AI 429 “Resource exhausted” errors

Your Gemini calls on Vertex AI (which Google now calls Gemini Enterprise Agent Platform) fail with 429 Resource exhausted, often while your quota dashboard says you have room to spare. This page explains why that happens, what Google recommends, and how to shift traffic to a second provider while Vertex AI is failing, without changing your application code.

Why you get 429s inside your quota

For standard pay-as-you-go traffic, Vertex AI serves Gemini from a pool of capacity shared with other customers. Google describes a 429 as a sign that "your requests are coming in faster than the service can handle them at that moment". What runs out is capacity in the shared pool, not only your own quota. So even a low-volume workload can see 429s, and they often arrive in bursts.

That is why the reports look alike: 429s in production despite being within quota, and errors from a handful of calls to Gemini 2.5 Flash.

What Google recommends, and when it is enough

Try these first. Google's guide to reducing 429 errors lists five fixes:

If the global endpoint and retries keep your error rate acceptable, you may not need anything else. The rest of this page is for when they don't:

Add a second provider with sturnus

sturnus is a small open-source proxy that runs next to your application. It speaks the OpenAI API and sends each request to whichever of your configured providers is fastest and healthy right now. Here, Gemini on the Vertex AI global endpoint is paired with an OpenAI model:

listen = "127.0.0.1:4000"

# Vertex AI global endpoint. Auth uses GKE Workload Identity,
# so no API key is needed.
[provider.vertex]
vertex_ai = { project_id = "my-gcp-project", location = "global" }

[provider.openai]
base_url = "https://api.openai.com/v1"
api_key = "${OPENAI_API_KEY}"

[model]
flash = [
  { provider = "vertex", model = "google/gemini-3.8-flash" },
  { provider = "openai", model = "gpt-6-luna" },
]

Your code then asks for the flash alias:

from openai import OpenAI

client = OpenAI(base_url="http://127.0.0.1:4000/v1", api_key="unused")
response = client.chat.completions.create(
    model="flash",
    messages=[{"role": "user", "content": "Hello"}],
)

The models under one alias should be interchangeable for your use case. Test your prompts against both before you mix them. If you'd rather stay with Gemini, Gemini through Google AI Studio is another option: it's a separate service with its own quota, and sturnus supports it too.

What happens when Vertex AI starts returning 429s

sturnus keeps two moving averages for each provider: time to first token, and success rate. A 429 lowers Vertex AI's success rate, and sturnus divides latency by success rate. So a provider that fails often scores as slower, and its share of new requests shrinks. It never drops below 1% of traffic, and those requests act as probes. When Vertex AI stops returning 429s, its score recovers and it wins traffic back without you doing anything.

sturnus does not retry. The 429 goes back to your client unchanged, and your SDK's retry logic takes over. The OpenAI Python SDK, for example, retries 429s with backoff by default. Each retry is a new routing decision, made after Vertex AI's share has already dropped, so retries tend to land on OpenAI.

Running it on GKE

Run sturnus as a native sidecar container in the same pod as your application, listening on 127.0.0.1:4000. It gets Vertex AI credentials from the GKE metadata server through Workload Identity and refreshes the tokens itself. Give the pod's Kubernetes service account permission to call Vertex AI (for example the roles/aiplatform.user role), and pass the OpenAI key to the sidecar as an environment variable from a Kubernetes Secret.

Each pod's sidecar routes on what it observes, so there is no Redis, database or shared state to run. Prometheus metrics for each provider, and a /status endpoint showing live scores, tell you where traffic is going and why. The README covers deployment details.

What to know before you use it

Get sturnus

sturnus is a single static binary, MIT licensed, published as a container for amd64 and arm64.

docker run -v ./config.toml:/config.toml -p 4000:4000 ghcr.io/sturnus-dev/sturnus:latest

Install options, the full configuration reference and metrics are on the home page and in the README. Questions and bug reports go to GitHub issues.