Fastest LLM Inference in 2026: Live Tokens/sec Leaderboard
Last updated: 2026-06-28
The fastest LLM inference in 2026 comes from specialized inference chips, not GPUs. On Artificial Analysis live data, Cerebras serves gpt-oss-120B at 1,753 output tokens/sec, far ahead of GPU clouds like Together AI (563.5) and Fireworks (690.1) (Source: Artificial Analysis, 2026). But “fastest” depends on the model, hardware, and provider, and these numbers shift daily.
There is no single fastest LLM. Speed is a moving target measured per model, per provider, and per workload. A provider that leads on one model can trail on another, and published “world record” peaks rarely match the sustained speed you get from a live API endpoint. This post ranks the current leaders using only real, sourced output-tokens-per-second figures, explains why the numbers move, and separates marketing peaks from what you actually receive.
What is the fastest LLM inference right now?
The fastest LLM inference right now is delivered by wafer-scale and ASIC inference hardware. As of June 2026, Cerebras leads Artificial Analysis live rankings on gpt-oss-120B at 1,753 output tokens/sec, and Groq leads Llama 3.3 70B at 322 tokens/sec (Source: Artificial Analysis, 2026). The honest answer: it depends on the model and changes constantly.
“Fastest” only means something when you fix three variables: the model, the hardware or provider, and the measurement window. Artificial Analysis, the most-cited independent benchmark, reports output speed as a median over a rolling 72-hour window, sampled multiple times per day (Source: Artificial Analysis, 2026). That is why a leaderboard is never final. The same model can show different speeds across two snapshots taken days apart, which is exactly why a live source beats a static blog claim.
How is inference speed actually measured?
Inference speed is most often reported as output tokens per second: how many tokens a model generates after it starts responding. Artificial Analysis measures this with single-request and parallel-request sampling across a 72-hour rolling window, then publishes the median per provider (Source: Artificial Analysis, 2026). Time-to-first-token is tracked separately and matters for latency-sensitive apps.
Two numbers describe most of “speed.” Output tokens/sec captures throughput once generation starts, the figure that dominates how fast a long answer streams. Time-to-first-token captures how long you wait before anything appears. A chatbot feels fast with low time-to-first-token; a code generator or document summarizer feels fast with high output tokens/sec. The leaderboards below rank output speed, the metric most people mean by “fastest.” For a deeper primer on the unit itself, see our guide to LLM tokens per second.
Which provider has the fastest LLM inference in 2026?
On Artificial Analysis live data (June 2026), Cerebras is the fastest provider on gpt-oss-120B at 1,753 tokens/sec, beating SambaNova (692.7), Fireworks (690.1), Together AI (563.5), and Groq (477.6) (Source: Artificial Analysis, 2026). On Llama 3.3 70B, Groq leads at 322 tokens/sec ahead of SambaNova (282.3). Leadership rotates by model.
The table below is the live leaderboard backbone: sustained output speed from production API endpoints, not one-off records. Every row is a real figure retrieved from the Artificial Analysis provider page for that model in June 2026.
Live output speed by provider and model (June 2026)
| Model | Provider | Output tokens/sec | Source |
|---|---|---|---|
| gpt-oss-120B | Cerebras | 1,753.0 | Artificial Analysis, 2026 |
| gpt-oss-120B | SambaNova | 692.7 | Artificial Analysis, 2026 |
| gpt-oss-120B | Fireworks | 690.1 | Artificial Analysis, 2026 |
| gpt-oss-120B | Together AI | 563.5 | Artificial Analysis, 2026 |
| gpt-oss-120B | Groq | 477.6 | Artificial Analysis, 2026 |
| Llama 3.3 70B | Groq | 322.0 | Artificial Analysis, 2026 |
| Llama 3.3 70B | SambaNova | 282.3 | Artificial Analysis, 2026 |
| Llama 3.3 70B | Google Vertex | 130.4 | Artificial Analysis, 2026 |
| Llama 4 Maverick | Amazon Bedrock | 187.1 | Artificial Analysis, 2026 |
| Llama 4 Maverick | Azure (FP8) | 119.7 | Artificial Analysis, 2026 |
| DeepSeek R1 (0528) | Google Vertex | 150.8 | Artificial Analysis, 2026 |
| DeepSeek R1 (0528) | Azure | 81.6 | Artificial Analysis, 2026 |
Two takeaways. First, the spread between providers serving the same model is large: on gpt-oss-120B, Cerebras runs roughly 3.4x faster than DeepInfra’s standard endpoint, which Artificial Analysis lists near 51.7 tokens/sec (Source: Artificial Analysis, 2026). Second, the leader changes by model: Cerebras tops gpt-oss-120B, Groq tops Llama 3.3 70B. Pick the provider after you pick the model. For a structured side-by-side, see our LLM speed comparison.
Why doesn’t the fastest provider win every model?
No provider wins every model because availability and optimization differ per model. On the Llama 4 Maverick live page in June 2026, Cerebras and Groq endpoints are not currently listed; the top live entry is Amazon Bedrock at 187.1 tokens/sec (Source: Artificial Analysis, 2026). A provider can hold a record on a model it no longer serves live.
This is the single most misread fact in inference rankings. A vendor can publish a stunning benchmark for a model, then quietly stop serving that model at scale, or serve it only to enterprise customers. The live leaderboard reflects what an ordinary API key can reach today. When you see a 2,000-plus tokens/sec headline for a model whose live page tops out near 200, the gap is almost always availability, batch configuration, or a one-time benchmarked endpoint, not a number you can call in production.
How fast are the published record claims?
Record claims sit far above live medians because they measure peak, single-stream, latency-optimized configurations. Cerebras reported Llama 4 Maverick at 2,522 tokens/sec, verified by Artificial Analysis, versus NVIDIA Blackwell’s 1,038 on the same model (Source: Cerebras, 2025). NVIDIA’s own DGX B200 hit over 1,000 tokens/sec per user using speculative decoding and FP8 (Source: NVIDIA, 2025).
These are the headline numbers vendors cite, and they are real, but they answer a different question than the live table. Records measure the best achievable speed under ideal conditions; the live leaderboard measures what production endpoints sustain. Keep them in separate columns.
Peak record claims (2025), still cited in 2026
| Model | Hardware / Provider | Peak output tokens/sec | Source |
|---|---|---|---|
| Llama 4 Maverick (400B) | Cerebras WSE-3 | 2,522 | Cerebras, 2025 |
| Llama 4 Maverick (400B) | NVIDIA Blackwell (AA-measured) | 1,038 | Cerebras / Artificial Analysis, 2025 |
| Llama 4 Maverick (400B) | NVIDIA DGX B200 (8 GPU) | >1,000 per user; 72,000 per server | NVIDIA, 2025 |
| gpt-oss-120B | Cerebras | 3,000 | Cerebras, 2025 |
| Llama 3.1 70B | Cerebras | ~2,100 | Cerebras, 2024 |
| Llama 3.3 70B | Groq LPU | 276 | Groq, 2024 |
| DeepSeek R1 671B (FP16) | SambaNova (16x SN40L) | 198 | TechRadar, 2025 |
| Llama 3 8B | SambaNova SN40L | 1,000 | VentureBeat, 2024 |
NVIDIA’s per-user record carries an important caveat the company states directly: the >1,000 tokens/sec-per-user figure on DGX B200 used FP8, EAGLE3-based speculative decoding, and custom CUDA kernels, and represents single-user latency optimization, distinct from the 72,000 tokens/sec aggregate a full server delivers across many users (Source: NVIDIA, 2025). Peak-per-user and aggregate throughput are different products.
Why is specialized inference hardware faster than GPUs?
Specialized inference chips are faster because they remove the memory-bandwidth bottleneck that limits GPU token generation. Cerebras holds model weights in 44 GB of on-chip SRAM with roughly 21 PB/s of on-chip bandwidth, and Groq’s LPU is a deterministic SRAM-based ASIC, both avoiding the HBM round-trips that pace GPUs (Source: Cerebras, 2025; Groq, 2024).
Token generation is memory-bound, not compute-bound. To produce each new token, a model must read its weights from memory; the faster that read, the faster the token. GPUs store weights in high-bandwidth memory (HBM) sitting beside the chip, so every token pays a bandwidth toll. Wafer-scale and ASIC designs keep weights in on-chip SRAM, which is dramatically faster to read, so they generate tokens faster per user.
How do the chip architectures differ?
The three leading inference chips take different routes to the same goal: keep weights close to the compute. Cerebras uses a wafer-scale engine; Groq uses a deterministic LPU; SambaNova uses a reconfigurable dataflow unit (RDU) that ran full DeepSeek-R1 671B at FP16 on a 16-chip rack (Source: SambaNova / TechRadar, 2025).
Cerebras WSE-3 is a single wafer-scale processor with about 900,000 cores and 44 GB of SRAM, eliminating the chip-to-chip and chip-to-HBM hops entirely for many models (Source: Cerebras, 2025). Groq’s LPU is a vertically integrated, deterministic ASIC whose predictable scheduling sustains speed across input sizes, Groq reported holding roughly 275 to 276 tokens/sec on Llama 3.3 70B across input lengths (Source: Groq, 2024). SambaNova’s RDU emphasizes running large models at full FP16 precision rather than dropping to 8-bit, positioning quality alongside speed. NVIDIA’s Blackwell counters with FP8, speculative decoding, and kernel-level optimization to push GPU per-user speed past 1,000 tokens/sec (Source: NVIDIA, 2025).
What is the speed-versus-quality tradeoff?
The main tradeoff is precision: many fast providers serve quantized (FP8 or lower) versions of a model under the same name, trading some numerical fidelity for speed. Artificial Analysis labels these variants explicitly, such as “Azure FP8” and “DeepInfra Turbo FP8” on its Llama 3.3 70B page (Source: Artificial Analysis, 2026). Same model name, different precision, different speed.
This is why two endpoints labeled “Llama 3.3 70B” can differ in both speed and output. A provider running FP8 or FP4 weights generates tokens faster and cheaper but may diverge slightly from the full-precision model on hard tasks. SambaNova markets staying at 16-bit precision as a quality differentiator against 8-bit competitors (Source: VentureBeat, 2024), while Cerebras has published quality evaluations comparing its outputs against Groq, Together, and Fireworks (Source: Cerebras, 2024). When you compare speed, check the precision label, not just the model name.
Does faster inference cost more?
Not necessarily, and specialized hardware can invert the usual relationship. Cerebras claims gpt-oss-120B runs at roughly 16x the median GPU-cloud speed for under 2x the cost, which it frames as about an 8x advantage in tokens-per-second per dollar (Source: Cerebras, 2025). Speed and cost-efficiency can move together when the architecture fits the workload.
The intuition that “faster equals pricier” comes from GPU scaling, where you rent more cards to go faster. Inference ASICs break that link: they deliver high per-user speed without proportionally higher cost, because the speed comes from architecture rather than raw card count. The practical implication is that for latency-sensitive products, the fastest provider is sometimes also among the most cost-effective per token, but you must verify against current published pricing, which changes as often as the speed rankings.
Why do inference speeds vary so much between snapshots?
Speeds vary because of batching, quantization, hardware, context length, and live demand. Artificial Analysis publishes medians over a rolling 72-hour window sampled multiple times daily, so the same model shows different numbers across dates (Source: Artificial Analysis, 2026). A provider under heavy load, or serving a longer context, will report lower output tokens/sec than its best-case benchmark.
Five forces move the numbers. Batching and concurrency: per-user speed falls as a provider packs more requests onto the same hardware, which is why single-user records and aggregate throughput diverge so sharply. Quantization: FP8 and FP4 endpoints run faster than BF16 or FP16. Hardware: SRAM-based chips outrun HBM-based GPUs on per-user generation. Context and input length: longer prompts can slow generation, though Groq cites stable speed across input sizes as a design strength (Source: Groq, 2024). Demand: load spikes throttle real-world speed. Because all five shift continuously, any single published number is a snapshot, not a constant, which is the core reason to consult a live leaderboard rather than a fixed table. TokenDyno tracks these output-tokens-per-second figures live across providers; see the current standings at TokenDyno.
How should you choose the fastest provider for your use case?
Choose by fixing your model first, then ranking providers by output tokens/sec for that exact model and precision. For gpt-oss-120B, Cerebras leads live at 1,753 tokens/sec; for Llama 3.3 70B, Groq leads at 322 (Source: Artificial Analysis, 2026). Match the metric to your workload: output speed for long generations, time-to-first-token for chat.
A practical sequence works better than chasing a single leaderboard. First, pick the model your product needs on quality grounds. Second, list the providers that serve that exact model at the precision you require. Third, rank them by the metric that matches your workload, output tokens/sec for streaming long answers, time-to-first-token for interactive chat. Fourth, re-check before launch and periodically after, because rankings rotate. For ongoing tracking across models, our LLM leaderboard 2026 and LLM speed comparison keep the current standings in one place.
Frequently asked questions
What is the fastest LLM?
There is no single fastest LLM; speed depends on the model, provider, and hardware. As of June 2026, the fastest sustained inference on Artificial Analysis live data is Cerebras serving gpt-oss-120B at 1,753 output tokens/sec (Source: Artificial Analysis, 2026). Rankings change daily, so check a live source before deciding.
Which provider has the fastest inference?
It depends on the model. On gpt-oss-120B, Cerebras leads live at 1,753 tokens/sec; on Llama 3.3 70B, Groq leads at 322 tokens/sec, ahead of SambaNova at 282.3 (Source: Artificial Analysis, 2026). Specialized inference-chip vendors (Cerebras, Groq, SambaNova) generally top GPU clouds on per-user output speed.
Why do inference speeds vary so much?
Speeds vary with batching, quantization, hardware, context length, and live demand. Artificial Analysis reports medians over a rolling 72-hour window, so figures shift between snapshots (Source: Artificial Analysis, 2026). Per-user speed also drops as providers batch more concurrent requests, which is why single-user records exceed sustained API throughput.
Is specialized hardware always faster than GPUs?
For per-user token generation, inference chips usually lead because they keep weights in fast on-chip SRAM. But NVIDIA Blackwell reached over 1,000 tokens/sec per user on Llama 4 Maverick using FP8 and speculative decoding (Source: NVIDIA, 2025), and GPUs still dominate aggregate, multi-user throughput at 72,000 tokens/sec per server.
Do faster providers sacrifice output quality?
Sometimes. Many fast endpoints serve quantized (FP8 or lower) versions of a model under the same name, which can slightly affect quality on hard tasks (Source: Artificial Analysis, 2026). Check the precision label, not just the model name; some vendors, like SambaNova, deliberately stay at 16-bit precision (Source: VentureBeat, 2024).
How often do the fastest-inference rankings change?
Frequently, often within days. Because Artificial Analysis uses a rolling 72-hour measurement window and providers add, remove, or re-optimize endpoints continuously, the leaderboard is never static (Source: Artificial Analysis, 2026). Treat any published number as a snapshot and verify against a live tracker before relying on it.
Sources
- Artificial Analysis, gpt-oss-120B providers: https://artificialanalysis.ai/models/gpt-oss-120b/providers
- Artificial Analysis, Llama 3.3 70B providers: https://artificialanalysis.ai/models/llama-3-3-instruct-70b/providers
- Artificial Analysis, Llama 4 Maverick providers: https://artificialanalysis.ai/models/llama-4-maverick/providers
- Artificial Analysis, DeepSeek R1 providers: https://artificialanalysis.ai/models/deepseek-r1/providers
- Artificial Analysis (Cerebras vs Blackwell, Llama 4 Maverick): https://x.com/ArtificialAnlys/status/1927896050112811125
- Cerebras, Llama 4 Maverick world-record press release: https://www.cerebras.ai/press-release/maverick
- Cerebras, gpt-oss-120B at 3,000 tokens/sec: https://www.cerebras.ai/blog/cerebras-launches-openai-s-gpt-oss-120b-at-a-blistering-3-000-tokens-sec
- Cerebras, inference 3x faster (Llama 3.1 70B): https://www.cerebras.ai/blog/cerebras-inference-3x-faster
- Cerebras, model quality evaluation (Cerebras, Groq, Together, Fireworks): https://www.cerebras.ai/blog/llama3.1-model-quality-evaluation-cerebras-groq-together-and-fireworks
- Groq, Llama 3.3 70B inference speed benchmark: https://groq.com/blog/new-ai-inference-speed-benchmark-for-llama-3-3-70b-powered-by-groq
- Groq, LPU inference engine benchmark: https://groq.com/blog/groq-lpu-inference-engine-crushes-first-public-llm-benchmark
- NVIDIA, Blackwell breaks 1,000 TPS/user with Llama 4 Maverick: https://developer.nvidia.com/blog/blackwell-breaks-the-1000-tps-user-barrier-with-metas-llama-4-maverick/
- NVIDIA, Blackwell InferenceMAX benchmark results: https://blogs.nvidia.com/blog/blackwell-inferencemax-benchmark-results/
- SambaNova, DeepSeek V3-0324 fastest inference: https://sambanova.ai/blog/deepseek-v3-0324-fastest-inference-in-world
- TechRadar, SambaNova DeepSeek R1 record: https://www.techradar.com/pro/nvidia-rival-claims-deepseek-world-record-as-it-delivers-industry-first-performance-with-95-percent-fewer-chips
- VentureBeat, SambaNova breaks Llama 3 speed record: https://venturebeat.com/ai/sambanova-breaks-llama-3-speed-record-with-1000-tokens-per-second
- vLLM, performance update (v0.6.0): https://vllm.ai/blog/perf-update