All models
Live speed benchmarks for every model across all providers. Numbers update continuously — sorted by latest tokens per second within each provider. New to comparing model speed? Read the Ollama speed comparison explainer.
opencode-go
| Model | TPS now | TPS 24h avg | TTFT | Reliability | Intelligence Index |
|---|---|---|---|---|---|
| 1 GPT 5.6 Luna | 410.3 | — | 3.9s | 0% | 51.2 |
| 2 GLM 5.1 | 240.7 | 102.6 | 1.7s | 100% | 40.2 |
| 3 GLM 5.2 | 212.1 | 99.7 | 1.0s | 100% | 51.1 |
| 4 Kimi K2.6 | 194.3 | 103.4 | 769ms | 100% | 44.2 |
| 5 Kimi K2.5 | 180.2 | 118.0 | 6.0s | 100% | 35.4 |
| 6 GLM 5 | 177.9 | 80.6 | 3.0s | 100% | 39.5 |
| 7 Kimi K2.7 Code | 142.7 | 83.9 | 888ms | 100% | 41.9 |
| 8 MiMo V2.5 | 86.5 | 73.8 | 7.9s | 100% | 40.3 |
| 9 Hy3 | 68.2 | 48.9 | 1.9s | 100% | 41.2 |
| 10 MiMo V2.5 Pro | 62.5 | 48.8 | 1.7s | 100% | 42.2 |
| 11 DeepSeek V4 Pro | 46.2 | 41.2 | 1.1s | 100% | 44.3 |
| 12 Kimi K3 | 44.1 | 37.3 | 4.2s | 50% | 57.1 |
ollama
| Model | TPS now | TPS 24h avg | TTFT | Reliability | Intelligence Index |
|---|---|---|---|---|---|
| 1 Nemotron 3 Nano 30B (non-reasoning) | 172.7 | 109.2 | 452ms | 100% | 7.4 |
| 2 GPT-OSS 20B | 170.3 | 100.6 | 388ms | 99% | 14.9 |
| 3 DeepSeek V4 Flash 0731 | 165.9 | 115.3 | 13.3s | 96% | 49.9 |
| 4 GLM 5.2 | 136.9 | 113.0 | 451ms | 100% | 51.1 |
| 5 DeepSeek V4 Flash | 129.1 | 192.5 | 1.8s | 99% | 40.3 |
| 6 DeepSeek V4 Pro | 118.6 | 120.9 | 1.1s | 100% | 44.3 |
| 7 Nemotron 3 Super | 96.4 | 78.5 | 540ms | 100% | 25.4 |
| 8 Kimi K2.7 Code | 90.9 | 106.6 | 755ms | 100% | 41.9 |
| 9 Kimi K3 | 89.1 | 85.2 | 875ms | 100% | 57.1 |
| 10 Gemma4 31B | 85.6 | 79.4 | 365ms | 94% | 29.4 |
| 11 MiniMax M3 | 83.7 | 85.6 | 769ms | 100% | 44.4 |
| 12 Qwen3.5 397B | 67.5 | 72.6 | 913ms | 100% | 33.7 |
| 13 GPT-OSS 120B | 66.7 | 87.8 | 473ms | 100% | 23.8 |
| 14 GLM 5.1 | 60.0 | 89.3 | 936ms | 100% | 40.2 |
| 15 Mistral Large 3 675B (non-reasoning) | 57.7 | 54.9 | 630ms | 99% | 15.9 |
| 16 Kimi K2.6 | 56.0 | 97.3 | 970ms | 100% | 44.2 |
| 17 Nemotron 3 Ultra | 40.3 | 39.9 | 8.6s | 100% | 37.8 |
| 18 MiniMax M2.7 | 37.6 | 48.4 | 1.0s | 100% | 38.1 |
ollama-free
| Model | TPS now | TPS 24h avg | TTFT | Reliability | Intelligence Index |
|---|---|---|---|---|---|
| 1 GPT-OSS 20B | 143.8 | 93.8 | 678ms | 100% | 14.9 |
| 2 Gemma4 31B | 104.5 | 87.9 | 282ms | 92% | 29.4 |
| 3 Nemotron 3 Super | 75.7 | 76.8 | 625ms | 100% | 25.4 |
| 4 MiniMax M3 | 74.2 | 82.9 | 1.2s | 100% | 44.4 |
| 5 Nemotron 3 Nano 30B (non-reasoning) | 54.0 | 114.0 | 835ms | 100% | 7.4 |
| 6 Nemotron 3 Ultra | 51.8 | 43.6 | 689ms | 100% | 37.8 |
| 7 GPT-OSS 120B | 47.0 | 91.2 | 631ms | 100% | 23.8 |
opencode-zen
| Model | TPS now | TPS 24h avg | TTFT | Reliability | Intelligence Index |
|---|---|---|---|---|---|
| 1 Big Pickle | 76.2 | 82.1 | 1.5s | 92% | — |
| 2 DeepSeek V4 Flash (Free) | 64.3 | 83.6 | 1.3s | 92% | 40.3 |
| 3 MiMo V2.5 (Free) | 44.5 | 63.8 | 11.0s | 100% | 40.3 |
| 4 Nemotron 3 Ultra (Free) | 39.6 | 38.2 | 999ms | 92% | 37.8 |
| 5 North Mini Code (Free) | 18.6 | 22.2 | 1.2s | 92% | 19.8 |
opencode-go-anthropic
| Model | TPS now | TPS 24h avg | TTFT | Reliability | Intelligence Index |
|---|---|---|---|---|---|
| 1 MiniMax M3 | 68.7 | 74.3 | 1.6s | 100% | 44.4 |
| 2 Qwen3.8 Max (non-reasoning) | 60.9 | 54.6 | 1.4s | 67% | 24 |
| 3 Qwen3.6 Plus | 58.8 | 58.8 | 1.1s | 38% | 39.6 |
| 4 Qwen3.7 Plus | 58.8 | 44.4 | 1.1s | 25% | 39 |
| 5 Qwen3.5 Plus (non-reasoning) | 58.8 | 44.8 | 1.4s | 25% | 30.6 |
| 6 Qwen3.7 Max | 55.6 | 52.7 | 2.8s | 58% | 46 |
| 7 MiniMax M2.7 | 40.3 | 52.3 | 1.3s | 100% | 38.1 |
| 8 MiniMax M2.5 | 36.6 | 49.0 | 4.6s | 100% | 33.7 |
How these numbers are measured
Every model is benchmarked with the same prompt and the same measurement method, regardless of provider. That makes the numbers directly comparable. Sampling cadence does vary by provider — about every 10 minutes on Ollama Pro, about every 60 minutes on the capped plans — but cadence affects only how fresh a number is, not how it is measured. Use the compare tool to overlay speed timelines for up to 6 models, including the same model on two different providers. See the methodology page for the full measurement spec.