sturnus

When the OpenAI API is slow

A request that normally returns in a second or two is taking ten, twenty or sixty. This page covers the causes you can fix yourself, and what to do when the slowness is on the provider's side: send each request to whichever of OpenAI and Azure OpenAI is faster at that moment, using the same model on both.

First, rule out the causes you control

OpenAI's latency optimization guide is the place to start. The changes that make the biggest difference:

Also check OpenAI's status page. If there's an incident, you'll see it there.

When it's the provider, not you

Sometimes the same request, with the same prompt and model, is simply slower than it was yesterday. Developers report 15-second waits for calls that usually take one or two, and responses that suddenly take far longer than usual. A slowdown like this doesn't always show up as a status-page incident.

Retrying doesn't help here: the request succeeds, just slowly. What helps is a second place to send it.

The same model on two providers

OpenAI's models are also served by Azure OpenAI, on Microsoft's infrastructure, with its own capacity. GPT-6 Luna, for example, is available on both. Because the model is the same, your prompts behave the same wherever a request goes. You don't need to test and tune against a second model.

sturnus is a small open-source proxy that runs next to your application. It speaks the OpenAI API and sends each request to whichever of your configured providers is fastest and healthy right now:

listen = "127.0.0.1:4000"

[provider.openai]
base_url = "https://api.openai.com/v1"
api_key = "${OPENAI_API_KEY}"

[provider.azure]
api_key = "${AZURE_OPENAI_KEY}"
azure_openai = { resource_name = "my-resource", api_version = "2024-10-21" }

[model]
luna = [
  { provider = "openai", model = "gpt-6-luna" },
  { provider = "azure",  model = "gpt-6-luna" },  # your Azure deployment name
]

Your code then asks for the luna alias. Only the base URL changes:

from openai import OpenAI

client = OpenAI(base_url="http://127.0.0.1:4000/v1", api_key="unused")
stream = client.chat.completions.create(
    model="luna",
    messages=[{"role": "user", "content": "Hello"}],
    stream=True,
)

You can add more Azure deployments, for example in a second region, as further candidates under the same alias.

How sturnus decides which is faster

For streaming requests, sturnus measures the time to the first chunk. For non-streaming requests, it measures the full response time. It keeps these separately, as moving averages for each provider, along with a success rate. The faster provider gets most of the traffic, and the slower one keeps a share that shrinks the further behind it falls, but never drops below 1%. Those requests keep its measurements fresh, so when it speeds up again, it wins traffic back without you doing anything.

The /status endpoint shows the current averages and each provider's share, and Prometheus metrics record time to first chunk and latency for each provider. You can see which one is faster, and when that changes.

What to know before you use it

Get sturnus

sturnus is a single static binary, MIT licensed, published as a container for amd64 and arm64.

docker run -v ./config.toml:/config.toml -p 4000:4000 ghcr.io/sturnus-dev/sturnus:latest

Install options, the full configuration reference and metrics are on the home page and in the README. Questions and bug reports go to GitHub issues.