The fastest LLM API by tokens per second

On raw throughput, the fastest LLM APIs are specialist inference providers running custom silicon rather than GPUs — Cerebras, Groq and SambaNova. TokenDyno does not benchmark them, so this page will not quote a number for them. What we do publish is a continuous, first-party measurement of the open-weight hosting platforms developers actually switch between: Ollama Cloud, OpenCode Zen and OpenCode Go — 55 models across 3 platforms, currently led by GPT-OSS 120B at 282.7 tokens/s.

If you came here for a Groq-versus-Cerebras number, we are not the right source and would rather say so than pad the page. If you are choosing between open-weight hosts, everything below is measured.

What does TokenDyno actually measure?

Four platforms, sampled around the clock with real streaming completions rather than synthetic pings: Ollama Cloud (Free and Pro plans tracked separately, because they behave differently), OpenCode Zen and OpenCode Go. For every model we publish tokens per second, time to first token, and a success rate — see how we measure.

We do not benchmark Cerebras, Groq, SambaNova, OpenRouter, or the first-party APIs of OpenAI, Anthropic and Google. No numbers for those appear anywhere on this site. On OpenRouter specifically, our sister site covers why an aggregator's speed is a property of whichever upstream served you: Ollama Cloud vs OpenRouter.

What is the fastest model we measure right now?

Rolling 24-hour averages, updated continuously. The full board carries every model we track.

#ModelProviderAvg tokens/s (24h)
1 GPT-OSS 120B Ollama Free 282.7
2 GPT-OSS 120B Ollama Pro 230.3
3 DeepSeek V4 Flash 0731 Ollama Pro 174.2
4 Nemotron 3 Nano 30B Ollama Pro 156.2
5 Nemotron 3 Nano 30B Ollama Free 152.9
6 Gemma4 31B Ollama Pro 140.4
7 Kimi K2.7 Code Ollama Pro 137.1
8 DeepSeek V4 Pro 0813 Ollama Pro 132.3
9 GPT-OSS 20B Ollama Free 127.5
10 Gemma4 31B Ollama Free 119.7

Excerpt of the live leaderboard, which is sortable and carries every model, plus time-to-first-token and reliability.

Why is "fastest" the wrong single question?

Tokens per second measures generation throughput after the first token arrives. Time to first token measures the wait before it does. A chat UI feels fast when time-to-first-token is low; a long batch job cares about sustained throughput. The two do not always move together, and a model can lead one while trailing the other — which is why the leaderboard publishes both rather than ranking on a single figure.

Reliability is the third axis and the most commonly ignored. A model averaging high throughput but failing a quarter of its requests is slower in practice than a steadier one, so every row carries its measured success rate.

Why not add Groq or Cerebras?

Because continuous measurement is the thing that makes this board worth reading, and it is not free. Every provider we add means sustaining paid usage on it around the clock, forever. We would rather cover a narrow set honestly and continuously than a wide set thinly and let the numbers go stale — which is the failure mode of most "fastest LLM API" pages, including ones that rank well.

Frequently asked questions

What is the fastest LLM API by tokens per second?

On raw throughput the leaders are specialist inference providers running custom silicon rather than GPUs — Cerebras, Groq and SambaNova — which routinely report figures well above general-purpose GPU hosting. TokenDyno does not benchmark those providers, so we will not quote numbers for them. Among the providers we do measure continuously (Ollama Cloud, OpenCode Zen and OpenCode Go), the current fastest model is shown live on our leaderboard, updated around the clock.

Which providers does TokenDyno benchmark?

Ollama Cloud on both its Free and Pro plans, OpenCode Zen, and OpenCode Go. We sample real streaming completions continuously and publish tokens-per-second and time-to-first-token for every model on those platforms. We do not benchmark Cerebras, Groq, SambaNova, OpenRouter, or the first-party APIs of OpenAI, Anthropic and Google.

Why does TokenDyno not benchmark Groq or Cerebras?

Scope, not opinion. TokenDyno was built to compare the open-weight hosting platforms developers actually switch between, and it measures those continuously rather than sampling many providers occasionally. Adding a provider means sustaining paid usage on it around the clock. We would rather cover a narrow set honestly and continuously than a wide set thinly.

Is a higher tokens-per-second number always better?

No. Tokens per second measures generation throughput once output has started; time-to-first-token measures how long you wait before it does. A chat interface feels fast when time-to-first-token is low, while a long batch job cares about sustained throughput. A model can lead on one and trail on the other, which is why we publish both.

How current are these speed numbers?

They are rolling 24-hour averages over samples taken continuously, not a one-off test. Cadence varies by provider because sampling a capped plan spends its budget: the Ollama Pro plan is sampled every 10 minutes, capped plans hourly, and per-token-billed models every four hours.