When the OpenAI API is slow
A request that normally returns in a second or two is taking ten, twenty or sixty. This page covers the causes you can fix yourself, and what to do when the slowness is on the provider's side: send each request to whichever of OpenAI and Azure OpenAI is faster at that moment, using the same model on both.
First, rule out the causes you control
OpenAI's latency optimization guide is the place to start. The changes that make the biggest difference:
- Stream the response. OpenAI calls streaming "the single most effective approach". Without it, you wait for the whole response to be generated before you see any of it.
- Generate fewer tokens. Output length drives latency more than anything else. Ask for shorter answers, or use structured outputs.
- Use a smaller model where the task allows it, and make fewer sequential requests.
- Check reasoning effort. Reasoning models spend time thinking before the first visible token. A lower effort setting returns sooner.
Also check OpenAI's status page. If there's an incident, you'll see it there.
When it's the provider, not you
Sometimes the same request, with the same prompt and model, is simply slower than it was yesterday. Developers report 15-second waits for calls that usually take one or two, and responses that suddenly take far longer than usual. A slowdown like this doesn't always show up as a status-page incident.
Retrying doesn't help here: the request succeeds, just slowly. What helps is a second place to send it.
The same model on two providers
OpenAI's models are also served by Azure OpenAI, on Microsoft's infrastructure, with its own capacity. GPT-6 Luna, for example, is available on both. Because the model is the same, your prompts behave the same wherever a request goes. You don't need to test and tune against a second model.
sturnus is a small open-source proxy that runs next to your application. It speaks the OpenAI API and sends each request to whichever of your configured providers is fastest and healthy right now:
listen = "127.0.0.1:4000"
[provider.openai]
base_url = "https://api.openai.com/v1"
api_key = "${OPENAI_API_KEY}"
[provider.azure]
api_key = "${AZURE_OPENAI_KEY}"
azure_openai = { resource_name = "my-resource", api_version = "2024-10-21" }
[model]
luna = [
{ provider = "openai", model = "gpt-6-luna" },
{ provider = "azure", model = "gpt-6-luna" }, # your Azure deployment name
]
Your code then asks for the luna alias. Only the base URL changes:
from openai import OpenAI
client = OpenAI(base_url="http://127.0.0.1:4000/v1", api_key="unused")
stream = client.chat.completions.create(
model="luna",
messages=[{"role": "user", "content": "Hello"}],
stream=True,
)
You can add more Azure deployments, for example in a second region, as further candidates under the same alias.
How sturnus decides which is faster
For streaming requests, sturnus measures the time to the first chunk. For non-streaming requests, it measures the full response time. It keeps these separately, as moving averages for each provider, along with a success rate. The faster provider gets most of the traffic, and the slower one keeps a share that shrinks the further behind it falls, but never drops below 1%. Those requests keep its measurements fresh, so when it speeds up again, it wins traffic back without you doing anything.
The /status endpoint shows the current averages and each provider's share, and Prometheus metrics record time to first chunk and latency for each provider. You can see which one is faster, and when that changes.
What to know before you use it
- It doesn't make a model faster. sturnus chooses between the providers you give it. If both are slow at once, requests are still slow.
- For streaming requests, it routes on the time to first chunk. That's when your users see the response start. It doesn't measure how fast tokens arrive after that.
- Azure OpenAI isn't identical to OpenAI. Azure applies its own content filtering, which can reject a prompt that OpenAI accepts. sturnus counts those rejections as errors, so a provider that often rejects your prompts gets less traffic. Check the filter settings on your Azure deployment.
- It routes on latency and errors, not cost. If one provider is cheaper for you, sturnus won't prefer it for that reason.
- Each sidecar learns from its own traffic. When a sidecar starts, every provider gets a small probe share until it has data.
Get sturnus
sturnus is a single static binary, MIT licensed, published as a container for amd64 and arm64.
docker run -v ./config.toml:/config.toml -p 4000:4000 ghcr.io/sturnus-dev/sturnus:latest
Install options, the full configuration reference and metrics are on the home page and in the README. Questions and bug reports go to GitHub issues.