AI Latency Index

How fast is AI from Germany?

The big latency benchmarks measure from the US, which says little for users in Germany. We measure several times a day straight from Frankfurt: how fast the major AI endpoints start responding, what throughput they deliver, and how reliable they are.

As of 30 September 2026, 03:20Measured from eu-central-1 (Frankfurt)

Time-to-first-token from Frankfurt

Time-to-first-token (TTFT) is the metric most sensitive to server distance, which is why it's the focus. TTFT and throughput are medians of the daily values over the last 30 days; the error rate covers the same period. Sorted by TTFT; throughput is an approximation (~).

30 daily measurements per endpoint, since 2026-09-01. Every series sits on the same 0 to 2,900 ms scale — a higher line really is slower.

EndpointTTFT (median)HistoryThroughputErrors
Mistral Smallvia AI GatewayMistral · AI Gateway322 msMistral Small: 30 readings, 276 to 384 ms~2501.1%
Claude Haiku 4.5EU region · directAWS Bedrock · eu-central-1 (Frankfurt)701 msClaude Haiku 4.5: 30 readings, 641 to 749 ms~5861.1%
GPT-5 minivia AI GatewayOpenAI · AI Gateway793 msGPT-5 mini: 30 readings, 650 to 1,019 ms~990%
Claude Haiku 4.5via AI GatewayAnthropic · AI Gateway803 msClaude Haiku 4.5: 30 readings, 678 to 1,004 ms~4870%
GPT-5 miniUS-global · directOpenAI · api.openai.com (US-global)884 msGPT-5 mini: 30 readings, 679 to 1,247 ms~1050%
Gemini 3 Flashvia AI GatewayGoogle · AI Gateway1,857 msGemini 3 Flash: 30 readings, 1,668 to 2,097 ms~774.21.1%
DeepSeek V4 Flashvia AI GatewayDeepSeek · AI Gateway2,039 msDeepSeek V4 Flash: 30 readings, 1,154 to 2,878 ms~214.411.2%

What is in the basket, and why it is small

The index compares routes, not models. So the model class stays deliberately fixed, the small and fast tier, and what varies is the path: which region it starts from and how many hops it takes. Measuring a large model against a small one would only show that the large one thinks for longer.

  • 1
    Direct, EU region

    Bedrock in Frankfurt. Request and response never leave the EU. This is the reference every other row is measured against.

    ModelsClaude Haiku 4.5 (AWS Bedrock)

  • 1
    Direct, US global

    OpenAI over api.openai.com. Same model class, same prompt, but the route crosses the Atlantic. The gap to the EU row is what this page actually measures.

    ModelsGPT-5 mini (OpenAI)

  • 5
    Via the AI Gateway

    One managed hop, which many EU teams genuinely use because it reaches several providers through a single credential. Not comparable one to one with the direct rows, which is exactly why it is listed separately.

    ModelsGPT-5 mini (OpenAI) · Claude Haiku 4.5 (Anthropic) · Gemini 3 Flash (Google) · Mistral Small (Mistral) · DeepSeek V4 Flash (DeepSeek)

The basket is small because every endpoint costs money on every run, every eight hours, indefinitely. Then there are the credentials: one OpenAI key, one gateway key and the Bedrock role. Anything not reachable through those cannot be measured. eu.api.openai.com was on the list and fell off again, because the key is not enabled for OpenAI EU data residency and the request ends in a 401.

How we measure

Honest and reproducible — real measurements from a Lambda in Frankfurt, identical for every endpoint.

  1. 1
    Vantage point: Frankfurt

    A Lambda function in AWS eu-central-1 (Frankfurt) sends identical mini-requests and times them. So the index measures latency the way German users experience it — not from the US.

  2. 2
    TTFT, throughput, errors

    TTFT = time to the first visible token (the server-distance-sensitive metric). Throughput = tokens/second (an approximation, ~). Error rate = share of failed calls.

  3. 3
    Several times a day, median

    Every 8 hours, two measurements per endpoint (the faster counts, to dampen outliers). Each day gets the median of its runs; the TTFT in the table is the median of those daily values across the 30-day window.

  4. 4
    Forward only

    Past latency from an EU location can't be measured after the fact — the head start from day 1 stays. Nobody publishes a from-Germany series like this.

Comparing "direct" and "via AI Gateway" isn't 1:1 — the gateway path has an extra, deliberately labelled hop, but it's a real path many EU developers use. For reasoning models we set minimal reasoning effort so TTFT reflects infrastructure, not thinking time. Values vary with time-of-day load; only the median over days is reliable. The OpenAI EU endpoint (eu.api.openai.com) is missing because our key isn't enabled for it.

Frequently asked

What is TTFT?

Time-to-first-token — the time from the request to the first word of the answer. It shapes an AI's perceived speed the most and is sensitive to server distance. That's why it's our primary metric.

Why measure from Frankfurt?

Because location matters: an endpoint in the US is noticeably slower from Germany than one in the EU. Big benchmarks measure from the US and don't reflect the German reality — we measure where your users are.

Why can nobody reconstruct this history?

Past latency from an EU location can't be measured retroactively. Anyone who starts later has missed the past days forever — which is exactly what makes the series valuable.

What does "via AI Gateway" mean?

Those endpoints run through a gateway (a service that routes requests to many models). That's an extra hop — we label it transparently. Direct endpoints (Bedrock Frankfurt, OpenAI) skip that detour.

Our own measurements from AWS eu-central-1 (Frankfurt), several times a day, an identical mini-request per endpoint. TTFT = time to first token; throughput is an approximation. Values vary with load; the median over days is the reliable figure. No warranty; not a substitute for your own load tests.

Fast, GDPR-compliant AI connectivity

We bring AI to your infrastructure performantly and compliantly — from endpoint choice to monitoring.