LLM Model Latency via GitHub Models

Latest snapshot: 2026-09-11T03:47:52.404649+00:00
What this measures: response latency of models served by GitHub Models, called through its single global access endpoint (models.github.ai). GitHub Models does not expose a region selector, so these numbers are not Azure per-region latency — they are cross-model speed evidence from one global vantage (labelled github-global), and they include network distance from wherever the probe runs.

Response Latency Leaderboard

Fastest p50 first · GitHub Models global endpoint (github-global) · not Azure per-region latency
Rank Model Status p50 (ms) p95 (ms) TTFT p50 (ms) Tokens/sec Samples Trend (p50)
deepseek/deepseek-v3-0324 unavailable ▼ 2%
meta/llama-3.3-70b-instruct unavailable ▼ 23%
microsoft/phi-4 unavailable ▲ 9%
openai/gpt-4.1 unavailable ▼ 25%
openai/gpt-4.1-mini unavailable ▲ 2%
openai/gpt-4.1-nano unavailable ▼ 18%
openai/gpt-4o unavailable ▼ 19%
openai/gpt-4o-mini unavailable ▲ 15%
openai/gpt-5 unavailable ▲ 24%
openai/gpt-5-chat unavailable ▲ 0%
openai/gpt-5-mini unavailable ▼ 12%
openai/gpt-5-nano unavailable ▼ 89%
openai/o3 unavailable ▼ 8%

Azure Per-Region Latency

Real region-attributable latency · one table per model · fastest region first · vantage = the probe runner
This is real Azure per-region latency. Each row is a single-region Standard deployment of that model in that Azure region, so the timing is attributable to the region. It still includes network distance from the probe runner's vantage, so read it as relative region speed rather than an SLA. The Rank column shows each region's speed position and how it moved since the previous snapshot.
Why only these regions? A region appears here only when Azure offers the model as a single-region Standard SKU. Many large regions (for example westeurope, northeurope, southeastasia, koreacentral, centralindia) currently offer these models only as GlobalStandard/DataZoneStandard, which can run in any datacenter in their geography and so are not region-attributable. See Status meanings for details; new regions surface automatically once a Standard SKU appears.

gpt-4.1 · 8 regions

Rank Region Status p50 (ms) p95 (ms) TTFT p50 (ms) Tokens/sec Samples
#1 ▲ 1 northcentralus available 867 1,471 777 115.4 5/5
#2 ▼ 1 westus3 available 1,020 1,038 960 98.0 5/5
#3 ▲ 3 switzerlandnorth available 1,089 1,466 992 91.8 5/5
#4 ▲ 1 swedencentral available 1,153 1,231 1,046 86.8 5/5
#5 ▼ 1 eastus available 1,239 1,876 1,147 80.7 5/5
#6 ▲ 1 eastus2 available 1,417 1,636 1,299 70.6 5/5
#7 ▼ 4 westus available 1,503 1,760 1,439 66.5 5/5
#8 southcentralus available 1,928 2,076 1,861 51.9 5/5

gpt-4.1-mini · 14 regions

Rank Region Status p50 (ms) p95 (ms) TTFT p50 (ms) Tokens/sec Samples
#1 ▲ 2 canadaeast available 1,000 1,336 951 100.0 5/5
#2 ▲ 5 uksouth available 1,198 1,515 1,116 83.5 5/5
#3 ▲ 6 francecentral available 1,206 1,321 1,124 82.9 5/5
#4 southcentralus available 1,335 1,698 1,253 74.9 5/5
#5 ▲ 6 switzerlandnorth available 1,428 1,466 1,324 70.0 5/5
#6 ▼ 5 westus3 available 1,610 2,100 1,565 62.1 5/5
#7 ▲ 1 swedencentral available 1,615 2,069 1,508 61.9 5/5
#8 ▼ 6 northcentralus available 1,634 4,077 1,518 61.2 5/5
#9 ▼ 3 westus available 1,666 1,977 1,584 60.0 5/5
#10 ▼ 5 eastus available 1,868 2,060 1,826 53.5 5/5
#11 ▼ 1 eastus2 available 1,946 2,238 1,884 51.4 5/5
#12 japaneast available 2,137 2,659 1,984 46.8 5/5
#13 southindia available 3,445 4,327 3,215 29.0 5/5
#14 australiaeast available 10,356 17,113 10,166 9.7 5/5

gpt-4o · 16 regions

Rank Region Status p50 (ms) p95 (ms) TTFT p50 (ms) Tokens/sec Samples
#1 ▲ 4 canadaeast available 833 978 792 120.1 5/5
#2 ▲ 1 northcentralus available 851 1,004 755 117.5 5/5
#3 ▲ 4 eastus available 919 1,104 814 108.8 5/5
#4 ▲ 9 francecentral available 960 1,050 878 104.2 5/5
#5 ▼ 3 westus3 available 983 1,463 897 101.7 5/5
#6 ▲ 3 uksouth available 991 1,276 908 100.9 5/5
#7 ▲ 3 eastus2 available 1,017 1,398 972 98.3 5/5
#8 ▼ 7 centralus available 1,027 1,136 975 97.3 5/5
#9 ▲ 2 switzerlandnorth available 1,084 1,099 987 92.2 5/5
#10 ▲ 4 norwayeast available 1,105 1,347 1,000 90.5 5/5
#11 ▼ 7 westus available 1,223 1,250 1,098 81.7 5/5
#12 ▼ 6 southcentralus available 1,319 1,400 1,237 75.8 5/5
#13 ▼ 5 japaneast available 1,339 1,870 1,185 74.7 5/5
#14 ▲ 1 swedencentral available 1,459 1,647 1,350 68.5 5/5
#15 ▼ 3 australiaeast available 1,632 1,957 1,441 61.3 5/5
#16 southindia available 1,925 1,982 1,711 51.9 5/5

gpt-5.1 · 9 regions

Rank Region Status p50 (ms) p95 (ms) TTFT p50 (ms) Tokens/sec Samples
#1 ▲ 1 centralus available 894 1,079 815 121.9 5/5
#2 ▲ 2 southcentralus available 1,138 1,269 1,104 95.8 5/5
#3 ▲ 2 swedencentral available 1,326 1,703 1,219 82.2 5/5
#4 ▲ 2 switzerlandnorth available 1,338 1,382 1,240 86.0 5/5
#5 ▼ 2 westus3 available 2,083 2,315 2,033 52.3 5/5
#6 ▲ 2 eastus2 available 2,211 2,356 2,170 49.8 5/5
#7 northcentralus available 2,240 2,327 2,111 49.5 5/5
#8 ▲ 1 eastus available 2,357 2,917 2,312 46.3 5/5
#9 ▼ 8 westus available 2,443 2,470 2,248 45.8 5/5

How to read this

What the latency numbers do and do not mean

Each model is sent a small deterministic prompt several times. We record p50 and p95 round-trip time, TTFT (time to first token), and output tokens per second, then publish the medians. p50 is the headline number used for ranking.

These are measurements from a single global vantage — the GitHub Models access endpoint — taken from wherever the probe runs. They mix model speed with network distance, so treat them as cross-model speed evidence, not as Azure per-region latency, an SLA, or a throughput guarantee. Reasoning models spend hidden tokens before their first visible token, so their TTFT is expected to be higher.

available means at least one timed call returned a trustworthy response. unknown means every sample failed, timed out, or returned no tokens.

The Trend column sparkline shows each model’s daily p50 over recent days. A means the model got faster, a means it got slower; new models show “collecting” until a few days of history accumulate.

The Rank column shows each row’s speed position (fastest = #1) and how it moved versus the previous snapshot: ▲ 2 means it climbed two places (relatively faster), ▼ 1 means it slipped one, means unchanged, and new marks a row with no previous ranking.

Help make this page clearer

Two short answers; GitHub will open a draft for your review. No Forms, token, or automatic submission.

This repository is public. Review the text and image for private details. Opening the draft sends text/context to GitHub in its URL. Pasting or attaching an image uploads it to GitHub immediately.

Describe the desired outcome in plain text. This becomes the issue's ### Objective. After reviewing the issue, a maintainer can add azure-backlog to queue it for a scheduled agent; the website does not add that label or start a run.

Page context included in the draft

Nothing is posted by this website. Sign in to GitHub if needed, review, attach your screenshot, and choose Submit. Closing this panel clears the local draft and screenshot.