LLM tokens per second — live multi-provider benchmark

Real inference speed, measured continuously. Every row is a live model from Ollama, OpenCode Zen, or OpenCode Go — sorted by tokens per second, benchmarked every ~10 minutes on Ollama Pro and every ~60 minutes on the capped provider plans.

Independent & vendor-neutral — no paid placement. TokenDyno is not affiliated with any provider it benchmarks and receives no payment for ranking or favorable coverage. Read the independence policy →

● live — last benchmark 2m ago
Trend
DeepSeek V4 Flash Ollama Pro 192.5 129.1 1.8s 99% 40.3 3m ago
DeepSeek V4 Pro Ollama Pro 120.9 118.6 1.1s 100% 44.3 3m ago
Kimi K2.5 OpenCode Go 118.0 180.2 6.0s 100% 35.4 42m ago
DeepSeek V4 Flash 0731 Ollama Pro 115.3 165.9 13.3s 96% 49.9 2m ago
Nemotron 3 Nano 30B (non-reasoning) Ollama Free 114.0 54.0 835ms 100% 7.4 58m ago
GLM 5.2 Ollama Pro 113.0 136.9 451ms 100% 51.1 2m ago
Nemotron 3 Nano 30B (non-reasoning) Ollama Pro 109.2 172.7 452ms 100% 7.4 7m ago
Kimi K2.7 Code Ollama Pro 106.6 90.9 755ms 100% 41.9 8m ago
Kimi K2.6 OpenCode Go 103.4 194.3 769ms 100% 44.2 41m ago
GLM 5.1 OpenCode Go 102.6 240.7 1.7s 100% 40.2 42m ago
GPT-OSS 20B Ollama Pro 100.6 170.3 388ms 99% 14.9 8m ago
GLM 5.2 OpenCode Go 99.7 212.1 1.0s 100% 51.1 42m ago
Kimi K2.6 Ollama Pro 97.3 56.0 970ms 100% 44.2 8m ago
GPT-OSS 20B Ollama Free 93.8 143.8 678ms 100% 14.9 59m ago
GPT-OSS 120B Ollama Free 91.2 47.0 631ms 100% 23.8 59m ago
GLM 5.1 Ollama Pro 89.3 60.0 936ms 100% 40.2 3m ago
Gemma4 31B Ollama Free 87.9 104.5 282ms 92% 29.4 4m ago
GPT-OSS 120B Ollama Pro 87.8 66.7 473ms 100% 23.8 9m ago
MiniMax M3 Ollama Pro 85.6 83.7 769ms 100% 44.4 7m ago
Kimi K3 Ollama Pro 85.2 89.1 875ms 100% n=6 57.1 3h ago
Kimi K2.7 Code OpenCode Go 83.9 142.7 888ms 100% 41.9 41m ago
DeepSeek V4 Flash (Free) OpenCode Zen 83.6 64.3 1.3s 92% 40.3 25m ago
MiniMax M3 Ollama Free 82.9 74.2 1.2s 100% 44.4 58m ago
Big Pickle OpenCode Zen 82.1 76.2 1.5s 92% 28m ago
GLM 5 OpenCode Go 80.6 177.9 3.0s 100% 39.5 43m ago
Gemma4 31B Ollama Pro 79.4 85.6 365ms 94% 29.4 3m ago
Nemotron 3 Super Ollama Pro 78.5 96.4 540ms 100% 25.4 7m ago
Nemotron 3 Super Ollama Free 76.8 75.7 625ms 100% 25.4 56m ago
MiniMax M3 OpenCode Go 74.3 68.7 1.6s 100% 44.4 39m ago
MiMo V2.5 OpenCode Go 73.8 86.5 7.9s 100% 40.3 57m ago
Qwen3.5 397B Ollama Pro 72.6 67.5 913ms 100% 33.7 3m ago
MiMo V2.5 (Free) OpenCode Zen 63.8 44.5 11.0s 100% 40.3 25m ago
Qwen3.6 Plus OpenCode Go 58.8 58.8 1.1s 38% 39.6 32m ago
Mistral Large 3 675B (non-reasoning) Ollama Pro 54.9 57.7 630ms 99% 15.9 7m ago
Qwen3.8 Max (non-reasoning) OpenCode Go 54.6 60.9 1.4s 67% 24 2m ago
Qwen3.7 Max OpenCode Go 52.7 55.6 2.8s 58% 46 30m ago
MiniMax M2.7 OpenCode Go 52.3 40.3 1.3s 100% 38.1 40m ago
MiniMax M2.5 OpenCode Go 49.0 36.6 4.6s 100% 33.7 41m ago
Hy3 OpenCode Go 48.9 68.2 1.9s 100% 41.2 42m ago
MiMo V2.5 Pro OpenCode Go 48.8 62.5 1.7s 100% 42.2 41m ago
MiniMax M2.7 Ollama Pro 48.4 37.6 1.0s 100% 38.1 8m ago
Qwen3.5 Plus (non-reasoning) OpenCode Go 44.8 58.8 1.4s 25% 30.6 5h ago
Qwen3.7 Plus OpenCode Go 44.4 58.8 1.1s 25% 39 3h ago
Nemotron 3 Ultra Ollama Free 43.6 51.8 689ms 100% 37.8 53m ago
DeepSeek V4 Pro OpenCode Go 41.2 46.2 1.1s 100% 44.3 44m ago
Nemotron 3 Ultra Ollama Pro 39.9 40.3 8.6s 100% 37.8 6m ago
Nemotron 3 Ultra (Free) OpenCode Zen 38.2 39.6 999ms 92% 37.8 23m ago
Kimi K3 OpenCode Go 37.3 44.1 4.2s 50% n=6 57.1 7h ago
North Mini Code (Free) OpenCode Zen 22.2 18.6 1.2s 92% 19.8 22m ago

Intelligence Index scores from Artificial Analysis.

Ollama Free, OpenCode Zen and OpenCode Go are sampled about hourly to avoid burning through their plan balances. Ollama Pro is sampled about every 10 minutes.

1 model unavailable or stale

Frequently asked questions

What is the best LLM tokens per second leaderboard?

There is no single "best" leaderboard because different tools measure different things. TokenDyno is a live tokens-per-second tracker: it re-benchmarks Ollama Pro, Ollama Free, OpenCode Zen, and OpenCode Go models on a fixed schedule — about every 10 minutes on Ollama Pro, about every 60 minutes on the capped provider plans — using the same prompt and measurement method across providers, so speed numbers are directly comparable. For a capability-only score, see LiveBench; for a combined capability-speed-price view, see Artificial Analysis. See the /vs page for a factual side-by-side.

What is TokenDyno's benchmark methodology?

TokenDyno sends a fixed streaming chat-completion request (max_tokens capped at 300) to each provider and measures generation throughput using a hybrid method: inter-token timing when a response streams smoothly (≥8 chunks, mean gap >50ms), or a wall-clock fallback for burst/chunked delivery. Time-to-first-token, error taxonomy, and reliability are also recorded. Full detail is on the /methodology page.

How often is TokenDyno's data refreshed?

The worker benchmarks continuously using a round-robin priority queue. Cadence follows how each provider bills: Ollama Pro is a flat subscription, so it is sampled about every 10 minutes. Ollama Free, OpenCode Zen, and OpenCode Go run on capped plans where every sample spends budget, so they are sampled about every 60 minutes; the two most expensive OpenCode Go models are sampled about every 4 hours. The leaderboard also polls the API roughly every 60 seconds client-side to keep cells current without a full page reload.