LLM Model Latency via GitHub Models
models.github.ai). GitHub Models
does not expose a region selector, so these numbers are not Azure
per-region latency — they are cross-model speed evidence from one global vantage
(labelled github-global), and they include network distance from wherever the
probe runs.
Response Latency Leaderboard
github-global) · not Azure per-region latency| Rank | Model | Status | p50 (ms) | p95 (ms) | TTFT p50 (ms) | Tokens/sec | Samples | Trend (p50) |
|---|---|---|---|---|---|---|---|---|
| — | deepseek/deepseek-v3-0324 | unavailable | — | — | — | — | — | ▼ 2% |
| — | meta/llama-3.3-70b-instruct | unavailable | — | — | — | — | — | ▼ 23% |
| — | microsoft/phi-4 | unavailable | — | — | — | — | — | ▲ 9% |
| — | openai/gpt-4.1 | unavailable | — | — | — | — | — | ▼ 25% |
| — | openai/gpt-4.1-mini | unavailable | — | — | — | — | — | ▲ 2% |
| — | openai/gpt-4.1-nano | unavailable | — | — | — | — | — | ▼ 18% |
| — | openai/gpt-4o | unavailable | — | — | — | — | — | ▼ 19% |
| — | openai/gpt-4o-mini | unavailable | — | — | — | — | — | ▲ 15% |
| — | openai/gpt-5 | unavailable | — | — | — | — | — | ▲ 24% |
| — | openai/gpt-5-chat | unavailable | — | — | — | — | — | ▲ 0% |
| — | openai/gpt-5-mini | unavailable | — | — | — | — | — | ▼ 12% |
| — | openai/gpt-5-nano | unavailable | — | — | — | — | — | ▼ 89% |
| — | openai/o3 | unavailable | — | — | — | — | — | ▼ 8% |
Azure Per-Region Latency
Why only these regions? A region appears here only when Azure offers the model as a single-region
Standard SKU. Many large regions (for example
westeurope, northeurope, southeastasia,
koreacentral, centralindia) currently offer these models only as
GlobalStandard/DataZoneStandard, which can run in any datacenter in
their geography and so are not region-attributable. See
Status meanings for details; new regions surface automatically
once a Standard SKU appears.
gpt-4.1 · 8 regions
| Rank | Region | Status | p50 (ms) | p95 (ms) | TTFT p50 (ms) | Tokens/sec | Samples |
|---|---|---|---|---|---|---|---|
| #1 ▲ 1 | northcentralus |
available | 867 | 1,471 | 777 | 115.4 | 5/5 |
| #2 ▼ 1 | westus3 |
available | 1,020 | 1,038 | 960 | 98.0 | 5/5 |
| #3 ▲ 3 | switzerlandnorth |
available | 1,089 | 1,466 | 992 | 91.8 | 5/5 |
| #4 ▲ 1 | swedencentral |
available | 1,153 | 1,231 | 1,046 | 86.8 | 5/5 |
| #5 ▼ 1 | eastus |
available | 1,239 | 1,876 | 1,147 | 80.7 | 5/5 |
| #6 ▲ 1 | eastus2 |
available | 1,417 | 1,636 | 1,299 | 70.6 | 5/5 |
| #7 ▼ 4 | westus |
available | 1,503 | 1,760 | 1,439 | 66.5 | 5/5 |
| #8 — | southcentralus |
available | 1,928 | 2,076 | 1,861 | 51.9 | 5/5 |
gpt-4.1-mini · 14 regions
| Rank | Region | Status | p50 (ms) | p95 (ms) | TTFT p50 (ms) | Tokens/sec | Samples |
|---|---|---|---|---|---|---|---|
| #1 ▲ 2 | canadaeast |
available | 1,000 | 1,336 | 951 | 100.0 | 5/5 |
| #2 ▲ 5 | uksouth |
available | 1,198 | 1,515 | 1,116 | 83.5 | 5/5 |
| #3 ▲ 6 | francecentral |
available | 1,206 | 1,321 | 1,124 | 82.9 | 5/5 |
| #4 — | southcentralus |
available | 1,335 | 1,698 | 1,253 | 74.9 | 5/5 |
| #5 ▲ 6 | switzerlandnorth |
available | 1,428 | 1,466 | 1,324 | 70.0 | 5/5 |
| #6 ▼ 5 | westus3 |
available | 1,610 | 2,100 | 1,565 | 62.1 | 5/5 |
| #7 ▲ 1 | swedencentral |
available | 1,615 | 2,069 | 1,508 | 61.9 | 5/5 |
| #8 ▼ 6 | northcentralus |
available | 1,634 | 4,077 | 1,518 | 61.2 | 5/5 |
| #9 ▼ 3 | westus |
available | 1,666 | 1,977 | 1,584 | 60.0 | 5/5 |
| #10 ▼ 5 | eastus |
available | 1,868 | 2,060 | 1,826 | 53.5 | 5/5 |
| #11 ▼ 1 | eastus2 |
available | 1,946 | 2,238 | 1,884 | 51.4 | 5/5 |
| #12 — | japaneast |
available | 2,137 | 2,659 | 1,984 | 46.8 | 5/5 |
| #13 — | southindia |
available | 3,445 | 4,327 | 3,215 | 29.0 | 5/5 |
| #14 — | australiaeast |
available | 10,356 | 17,113 | 10,166 | 9.7 | 5/5 |
gpt-4o · 16 regions
| Rank | Region | Status | p50 (ms) | p95 (ms) | TTFT p50 (ms) | Tokens/sec | Samples |
|---|---|---|---|---|---|---|---|
| #1 ▲ 4 | canadaeast |
available | 833 | 978 | 792 | 120.1 | 5/5 |
| #2 ▲ 1 | northcentralus |
available | 851 | 1,004 | 755 | 117.5 | 5/5 |
| #3 ▲ 4 | eastus |
available | 919 | 1,104 | 814 | 108.8 | 5/5 |
| #4 ▲ 9 | francecentral |
available | 960 | 1,050 | 878 | 104.2 | 5/5 |
| #5 ▼ 3 | westus3 |
available | 983 | 1,463 | 897 | 101.7 | 5/5 |
| #6 ▲ 3 | uksouth |
available | 991 | 1,276 | 908 | 100.9 | 5/5 |
| #7 ▲ 3 | eastus2 |
available | 1,017 | 1,398 | 972 | 98.3 | 5/5 |
| #8 ▼ 7 | centralus |
available | 1,027 | 1,136 | 975 | 97.3 | 5/5 |
| #9 ▲ 2 | switzerlandnorth |
available | 1,084 | 1,099 | 987 | 92.2 | 5/5 |
| #10 ▲ 4 | norwayeast |
available | 1,105 | 1,347 | 1,000 | 90.5 | 5/5 |
| #11 ▼ 7 | westus |
available | 1,223 | 1,250 | 1,098 | 81.7 | 5/5 |
| #12 ▼ 6 | southcentralus |
available | 1,319 | 1,400 | 1,237 | 75.8 | 5/5 |
| #13 ▼ 5 | japaneast |
available | 1,339 | 1,870 | 1,185 | 74.7 | 5/5 |
| #14 ▲ 1 | swedencentral |
available | 1,459 | 1,647 | 1,350 | 68.5 | 5/5 |
| #15 ▼ 3 | australiaeast |
available | 1,632 | 1,957 | 1,441 | 61.3 | 5/5 |
| #16 — | southindia |
available | 1,925 | 1,982 | 1,711 | 51.9 | 5/5 |
gpt-5.1 · 9 regions
| Rank | Region | Status | p50 (ms) | p95 (ms) | TTFT p50 (ms) | Tokens/sec | Samples |
|---|---|---|---|---|---|---|---|
| #1 ▲ 1 | centralus |
available | 894 | 1,079 | 815 | 121.9 | 5/5 |
| #2 ▲ 2 | southcentralus |
available | 1,138 | 1,269 | 1,104 | 95.8 | 5/5 |
| #3 ▲ 2 | swedencentral |
available | 1,326 | 1,703 | 1,219 | 82.2 | 5/5 |
| #4 ▲ 2 | switzerlandnorth |
available | 1,338 | 1,382 | 1,240 | 86.0 | 5/5 |
| #5 ▼ 2 | westus3 |
available | 2,083 | 2,315 | 2,033 | 52.3 | 5/5 |
| #6 ▲ 2 | eastus2 |
available | 2,211 | 2,356 | 2,170 | 49.8 | 5/5 |
| #7 — | northcentralus |
available | 2,240 | 2,327 | 2,111 | 49.5 | 5/5 |
| #8 ▲ 1 | eastus |
available | 2,357 | 2,917 | 2,312 | 46.3 | 5/5 |
| #9 ▼ 8 | westus |
available | 2,443 | 2,470 | 2,248 | 45.8 | 5/5 |
How to read this
Each model is sent a small deterministic prompt several times. We record
p50 and p95 round-trip time, TTFT
(time to first token), and output tokens per second, then publish the
medians. p50 is the headline number used for ranking.
These are measurements from a single global vantage — the GitHub Models access endpoint — taken from wherever the probe runs. They mix model speed with network distance, so treat them as cross-model speed evidence, not as Azure per-region latency, an SLA, or a throughput guarantee. Reasoning models spend hidden tokens before their first visible token, so their TTFT is expected to be higher.
available means at least one timed call returned a trustworthy response. unknown means every sample failed, timed out, or returned no tokens.
The Trend column sparkline shows each model’s daily p50 over recent days. A ▼ means the model got faster, a ▲ means it got slower; new models show “collecting” until a few days of history accumulate.
The Rank column shows each row’s speed position (fastest = #1) and how it moved versus the previous snapshot: ▲ 2 means it climbed two places (relatively faster), ▼ 1 means it slipped one, — means unchanged, and new marks a row with no previous ranking.