All models
Live speed benchmarks for every model across all providers. Numbers update continuously — sorted by latest tokens per second within each provider. New to comparing model speed? Read the Ollama speed comparison explainer.
opencode-zen
| Model | TPS now | TPS 24h avg | TTFT | Reliability | Intelligence Index |
|---|---|---|---|---|---|
| 1 Ling 3.0 Flash Fin (Free) | 400.4 | — | 1.1s | 0% | 22.6 |
| 2 Nemotron 3 Ultra (Free) | 65.9 | — | 881ms | 0% | 22.9 |
| 3 Nemotron 3.5 Lightning (Free) | 57.0 | — | 651ms | 0% | 12.9 |
| 4 Big Pickle | 50.3 | — | 697ms | 0% | — |
| 5 MiMo V2.5 (Free) | 13.7 | — | 2.1s | 0% | 25.2 |
ollama-free
| Model | TPS now | TPS 24h avg | TTFT | Reliability | Intelligence Index |
|---|---|---|---|---|---|
| 1 GPT-OSS 120B | 355.3 | 279.8 | 339ms | 100% | 11.6 |
| 2 Nemotron 3 Nano 30B (non-reasoning) | 285.2 | 191.3 | 312ms | 100% | 6.8 |
| 3 Gemma4 31B | 148.4 | 222.1 | 325ms | 100% | 19 |
| 4 GPT-OSS 20B | 102.4 | 108.7 | 458ms | 100% | 9 |
| 5 MiniMax M3 | 82.2 | — | 1.0s | 0% | 29.2 |
| 6 Nemotron 3 Super | 63.1 | 91.0 | 541ms | 100% | 12.8 |
| 7 Nemotron 3 Ultra | 4.7 | 44.1 | 9.6s | 100% | 22.9 |
ollama
| Model | TPS now | TPS 24h avg | TTFT | Reliability | Intelligence Index |
|---|---|---|---|---|---|
| 1 Gemma4 31B | 298.4 | 271.1 | 348ms | 100% | 19 |
| 2 DeepSeek V4.1 Flash | 275.6 | 175.9 | 366ms | 99% | 39.5 |
| 3 GPT-OSS 120B | 223.7 | 217.4 | 375ms | 100% | 11.6 |
| 4 DeepSeek V4 Flash 0731 | 186.0 | 184.4 | 529ms | 100% | 34.3 |
| 5 GLM 5.2 | 166.0 | 106.2 | 651ms | 100% | 33.7 |
| 6 Nemotron 3 Nano 30B (non-reasoning) | 129.1 | 184.6 | 801ms | 100% | 6.8 |
| 7 DeepSeek V4 Pro 0813 | 115.3 | 140.2 | 652ms | 100% | 36 |
| 8 Kimi K2.7 Code | 104.8 | 116.2 | 1.1s | 100% | 25.8 |
| 9 GLM 5.1 | 102.1 | 90.0 | 1.0s | 100% | 26.1 |
| 10 Qwen3.5 397B | 95.0 | 91.6 | 1.0s | 99% | 18.4 |
| 11 MiniMax M3 | 93.8 | 83.2 | 689ms | 99% | 29.2 |
| 12 GLM 5.3 Flash | 93.0 | 117.4 | 710ms | 100% | 41.8 |
| 13 GLM 5.3 | 92.8 | 134.1 | 730ms | 100% | 44.8 |
| 14 GPT-OSS 20B | 87.5 | 107.1 | 12.6s | 100% | 9 |
| 15 Nemotron 3 Super | 80.4 | 86.5 | 576ms | 99% | 12.8 |
| 16 Mistral Large 3 675B (non-reasoning) | 77.2 | 72.5 | 719ms | 100% | 9.3 |
| 17 Kimi K3 | 77.2 | 81.0 | 853ms | 100% | 43.6 |
| 18 MiniMax M2.7 | 70.9 | 64.1 | 1.1s | 100% | 22.8 |
| 19 Nemotron 3 Ultra | 64.7 | 40.5 | 568ms | 99% | 22.9 |
| 20 Kimi K2.6 | 41.8 | 45.2 | 1.7s | 100% | 27 |
opencode-go
| Model | TPS now | TPS 24h avg | TTFT | Reliability | Intelligence Index |
|---|---|---|---|---|---|
| 1 GLM 5.3 Flash | 217.3 | 90.4 | 589ms | 100% | 41.8 |
| 2 Omen Alpha | 164.7 | 98.6 | 918ms | 100% | — |
| 3 Hy3 | 146.2 | 111.2 | 1.6s | 100% | 25.3 |
| 4 Kimi K2.7 Code | 113.5 | 57.1 | 980ms | 100% | 25.8 |
| 5 Hy4 Preview | 94.0 | 107.7 | 3.2s | 100% | — |
| 6 GLM 5.3 | 90.2 | 88.0 | 824ms | 100% | 44.8 |
| 7 GLM 5.2 | 78.1 | 84.2 | 762ms | 100% | 33.7 |
| 8 GLM 5.1 | 77.9 | 82.9 | 833ms | 100% | 26.1 |
| 9 Longcat 2.0 | 68.6 | 53.1 | 2.7s | 100% | 19.1 |
| 10 MiMo V2.5 Pro | 46.9 | 40.6 | 2.1s | 96% | 26 |
| 11 Kimi K2.6 | 35.8 | 38.0 | 4.5s | 100% | 27 |
| 12 MiMo V2.5 | 35.5 | 37.6 | 1.1s | 100% | 25.2 |
opencode-go-anthropic
| Model | TPS now | TPS 24h avg | TTFT | Reliability | Intelligence Index |
|---|---|---|---|---|---|
| 1 DeepSeek V4 Flash Vision Exp | 202.0 | 183.1 | 1.1s | 100% | 34.8 |
| 2 DeepSeek Flash | 185.7 | 183.8 | 1.3s | 100% | 34.3 |
| 3 DeepSeek V4.1 Flash | 174.8 | 185.3 | 1.1s | 100% | 39.5 |
| 4 DeepSeek V4 Flash | 173.9 | 179.0 | 1.1s | 100% | 34.3 |
| 5 Qwen3.7 Max | 117.2 | 111.2 | 1.5s | 100% | 29.5 |
| 6 MiniMax M3 | 101.8 | 91.0 | 2.0s | 100% | 29.2 |
| 7 DeepSeek V4 Pro | 94.1 | 77.8 | 1.8s | 100% | 36 |
| 8 Kimi K3 | 82.7 | 71.5 | 2.2s | 100% | 43.6 |
| 9 Qwen3.7 Plus | 80.5 | 50.6 | 1.1s | 50% | 25.2 |
| 10 MiniMax M2.5 | 64.8 | 60.9 | 983ms | 98% | 22.8 |
| 11 MiniMax M2.7 | 58.3 | 61.8 | 1.2s | 100% | 22.8 |
| 12 Qwen3.8 Flash | 42.7 | 60.3 | 1.2s | 100% | 39.8 |
| 13 Qwen3.8 Max | 37.5 | 40.3 | 1.5s | 100% | 45.4 |
| 14 Qwen3.6 Plus | 15.7 | 44.5 | 1.0s | 75% | 27 |
opencode-go-responses
| Model | TPS now | TPS 24h avg | TTFT | Reliability | Intelligence Index |
|---|---|---|---|---|---|
| 1 GPT 5.6 Luna | 156.6 | 149.2 | 1.8s | 44% | 37.3 |
| 2 Muse Spark 1.2 Contributor | 140.4 | 94.3 | 1.7s | 92% | 39.6 |
| 3 Grok 4.6 | 96.7 | 100.5 | 108.6s | 75% | 44.3 |
| 4 Muse Spark 1.3 Contributor | 83.4 | 63.6 | 3.5s | 81% | 48.1 |
opencode-zen-responses
| Model | TPS now | TPS 24h avg | TTFT | Reliability | Intelligence Index |
|---|---|---|---|---|---|
| 1 Muse Spark 1.3 Contributor (Free) | 54.4 | — | 6.3s | 0% | 48.1 |
How these numbers are measured
Every model is benchmarked with the same prompt and the same measurement method, regardless of provider. That makes the numbers directly comparable. Sampling cadence does vary by provider — about every 10 minutes on Ollama Pro, about every 60 minutes on the capped plans — but cadence affects only how fresh a number is, not how it is measured. Use the compare tool to overlay speed timelines for up to 6 models, including the same model on two different providers. See the methodology page for the full measurement spec.