The big latency benchmarks measure from the US, which says little for users in Germany. We measure several times a day straight from Frankfurt: how fast the major AI endpoints start responding, what throughput they deliver, and how reliable they are.
Time-to-first-token (TTFT) is the metric most sensitive to server distance, which is why it's the focus. TTFT and throughput are medians of the daily values over the last 30 days; the error rate covers the same period. Sorted by TTFT; throughput is an approximation (~).
30 daily measurements per endpoint, since 2026-09-01. Every series sits on the same 0 to 2,900 ms scale — a higher line really is slower.
| Endpoint | TTFT (median) | History | Throughput | Errors |
|---|---|---|---|---|
| Mistral Small | 322 ms | ~250 | 1.1% | |
| Claude Haiku 4.5 | 701 ms | ~586 | 1.1% | |
| GPT-5 mini | 793 ms | ~99 | 0% | |
| Claude Haiku 4.5 | 803 ms | ~487 | 0% | |
| GPT-5 mini | 884 ms | ~105 | 0% | |
| Gemini 3 Flash | 1,857 ms | ~774.2 | 1.1% | |
| DeepSeek V4 Flash | 2,039 ms | ~214.4 | 11.2% |
The index compares routes, not models. So the model class stays deliberately fixed, the small and fast tier, and what varies is the path: which region it starts from and how many hops it takes. Measuring a large model against a small one would only show that the large one thinks for longer.
Bedrock in Frankfurt. Request and response never leave the EU. This is the reference every other row is measured against.
ModelsClaude Haiku 4.5 (AWS Bedrock)
OpenAI over api.openai.com. Same model class, same prompt, but the route crosses the Atlantic. The gap to the EU row is what this page actually measures.
ModelsGPT-5 mini (OpenAI)
One managed hop, which many EU teams genuinely use because it reaches several providers through a single credential. Not comparable one to one with the direct rows, which is exactly why it is listed separately.
ModelsGPT-5 mini (OpenAI) · Claude Haiku 4.5 (Anthropic) · Gemini 3 Flash (Google) · Mistral Small (Mistral) · DeepSeek V4 Flash (DeepSeek)
The basket is small because every endpoint costs money on every run, every eight hours, indefinitely. Then there are the credentials: one OpenAI key, one gateway key and the Bedrock role. Anything not reachable through those cannot be measured. eu.api.openai.com was on the list and fell off again, because the key is not enabled for OpenAI EU data residency and the request ends in a 401.
Honest and reproducible — real measurements from a Lambda in Frankfurt, identical for every endpoint.
A Lambda function in AWS eu-central-1 (Frankfurt) sends identical mini-requests and times them. So the index measures latency the way German users experience it — not from the US.
TTFT = time to the first visible token (the server-distance-sensitive metric). Throughput = tokens/second (an approximation, ~). Error rate = share of failed calls.
Every 8 hours, two measurements per endpoint (the faster counts, to dampen outliers). Each day gets the median of its runs; the TTFT in the table is the median of those daily values across the 30-day window.
Past latency from an EU location can't be measured after the fact — the head start from day 1 stays. Nobody publishes a from-Germany series like this.
Comparing "direct" and "via AI Gateway" isn't 1:1 — the gateway path has an extra, deliberately labelled hop, but it's a real path many EU developers use. For reasoning models we set minimal reasoning effort so TTFT reflects infrastructure, not thinking time. Values vary with time-of-day load; only the median over days is reliable. The OpenAI EU endpoint (eu.api.openai.com) is missing because our key isn't enabled for it.
Time-to-first-token — the time from the request to the first word of the answer. It shapes an AI's perceived speed the most and is sensitive to server distance. That's why it's our primary metric.
Because location matters: an endpoint in the US is noticeably slower from Germany than one in the EU. Big benchmarks measure from the US and don't reflect the German reality — we measure where your users are.
Past latency from an EU location can't be measured retroactively. Anyone who starts later has missed the past days forever — which is exactly what makes the series valuable.
Those endpoints run through a gateway (a service that routes requests to many models). That's an extra hop — we label it transparently. Direct endpoints (Bedrock Frankfurt, OpenAI) skip that detour.
Our own measurements from AWS eu-central-1 (Frankfurt), several times a day, an identical mini-request per endpoint. TTFT = time to first token; throughput is an approximation. Values vary with load; the median over days is the reliable figure. No warranty; not a substitute for your own load tests.
We bring AI to your infrastructure performantly and compliantly — from endpoint choice to monitoring.
These tools cover related ground.