All models

Live speed benchmarks for every model across all providers. Numbers update continuously — sorted by latest tokens per second within each provider. New to comparing model speed? Read the Ollama speed comparison explainer.

63 total models benchmarked
63 currently available
7 providers

opencode-zen

Model TPS now TPS 24h avg TTFT Reliability Intelligence Index
1 Ling 3.0 Flash Fin (Free) 400.4 1.1s 0% 22.6
2 Nemotron 3 Ultra (Free) 65.9 881ms 0% 22.9
3 Nemotron 3.5 Lightning (Free) 57.0 651ms 0% 12.9
4 Big Pickle 50.3 697ms 0%
5 MiMo V2.5 (Free) 13.7 2.1s 0% 25.2

ollama-free

Model TPS now TPS 24h avg TTFT Reliability Intelligence Index
1 GPT-OSS 120B 355.3 279.8 339ms 100% 11.6
2 Nemotron 3 Nano 30B (non-reasoning) 285.2 191.3 312ms 100% 6.8
3 Gemma4 31B 148.4 222.1 325ms 100% 19
4 GPT-OSS 20B 102.4 108.7 458ms 100% 9
5 MiniMax M3 82.2 1.0s 0% 29.2
6 Nemotron 3 Super 63.1 91.0 541ms 100% 12.8
7 Nemotron 3 Ultra 4.7 44.1 9.6s 100% 22.9

ollama

Model TPS now TPS 24h avg TTFT Reliability Intelligence Index
1 Gemma4 31B 298.4 271.1 348ms 100% 19
2 DeepSeek V4.1 Flash 275.6 175.9 366ms 99% 39.5
3 GPT-OSS 120B 223.7 217.4 375ms 100% 11.6
4 DeepSeek V4 Flash 0731 186.0 184.4 529ms 100% 34.3
5 GLM 5.2 166.0 106.2 651ms 100% 33.7
6 Nemotron 3 Nano 30B (non-reasoning) 129.1 184.6 801ms 100% 6.8
7 DeepSeek V4 Pro 0813 115.3 140.2 652ms 100% 36
8 Kimi K2.7 Code 104.8 116.2 1.1s 100% 25.8
9 GLM 5.1 102.1 90.0 1.0s 100% 26.1
10 Qwen3.5 397B 95.0 91.6 1.0s 99% 18.4
11 MiniMax M3 93.8 83.2 689ms 99% 29.2
12 GLM 5.3 Flash 93.0 117.4 710ms 100% 41.8
13 GLM 5.3 92.8 134.1 730ms 100% 44.8
14 GPT-OSS 20B 87.5 107.1 12.6s 100% 9
15 Nemotron 3 Super 80.4 86.5 576ms 99% 12.8
16 Mistral Large 3 675B (non-reasoning) 77.2 72.5 719ms 100% 9.3
17 Kimi K3 77.2 81.0 853ms 100% 43.6
18 MiniMax M2.7 70.9 64.1 1.1s 100% 22.8
19 Nemotron 3 Ultra 64.7 40.5 568ms 99% 22.9
20 Kimi K2.6 41.8 45.2 1.7s 100% 27

opencode-go

Model TPS now TPS 24h avg TTFT Reliability Intelligence Index
1 GLM 5.3 Flash 217.3 90.4 589ms 100% 41.8
2 Omen Alpha 164.7 98.6 918ms 100%
3 Hy3 146.2 111.2 1.6s 100% 25.3
4 Kimi K2.7 Code 113.5 57.1 980ms 100% 25.8
5 Hy4 Preview 94.0 107.7 3.2s 100%
6 GLM 5.3 90.2 88.0 824ms 100% 44.8
7 GLM 5.2 78.1 84.2 762ms 100% 33.7
8 GLM 5.1 77.9 82.9 833ms 100% 26.1
9 Longcat 2.0 68.6 53.1 2.7s 100% 19.1
10 MiMo V2.5 Pro 46.9 40.6 2.1s 96% 26
11 Kimi K2.6 35.8 38.0 4.5s 100% 27
12 MiMo V2.5 35.5 37.6 1.1s 100% 25.2

opencode-go-anthropic

Model TPS now TPS 24h avg TTFT Reliability Intelligence Index
1 DeepSeek V4 Flash Vision Exp 202.0 183.1 1.1s 100% 34.8
2 DeepSeek Flash 185.7 183.8 1.3s 100% 34.3
3 DeepSeek V4.1 Flash 174.8 185.3 1.1s 100% 39.5
4 DeepSeek V4 Flash 173.9 179.0 1.1s 100% 34.3
5 Qwen3.7 Max 117.2 111.2 1.5s 100% 29.5
6 MiniMax M3 101.8 91.0 2.0s 100% 29.2
7 DeepSeek V4 Pro 94.1 77.8 1.8s 100% 36
8 Kimi K3 82.7 71.5 2.2s 100% 43.6
9 Qwen3.7 Plus 80.5 50.6 1.1s 50% 25.2
10 MiniMax M2.5 64.8 60.9 983ms 98% 22.8
11 MiniMax M2.7 58.3 61.8 1.2s 100% 22.8
12 Qwen3.8 Flash 42.7 60.3 1.2s 100% 39.8
13 Qwen3.8 Max 37.5 40.3 1.5s 100% 45.4
14 Qwen3.6 Plus 15.7 44.5 1.0s 75% 27

opencode-go-responses

Model TPS now TPS 24h avg TTFT Reliability Intelligence Index
1 GPT 5.6 Luna 156.6 149.2 1.8s 44% 37.3
2 Muse Spark 1.2 Contributor 140.4 94.3 1.7s 92% 39.6
3 Grok 4.6 96.7 100.5 108.6s 75% 44.3
4 Muse Spark 1.3 Contributor 83.4 63.6 3.5s 81% 48.1

opencode-zen-responses

Model TPS now TPS 24h avg TTFT Reliability Intelligence Index
1 Muse Spark 1.3 Contributor (Free) 54.4 6.3s 0% 48.1

How these numbers are measured

Every model is benchmarked with the same prompt and the same measurement method, regardless of provider. That makes the numbers directly comparable. Sampling cadence does vary by provider — about every 10 minutes on Ollama Pro, about every 60 minutes on the capped plans — but cadence affects only how fresh a number is, not how it is measured. Use the compare tool to overlay speed timelines for up to 6 models, including the same model on two different providers. See the methodology page for the full measurement spec.