Kimi K3 tokens per second
Live benchmark data for Kimi K3 on OpenCode Go. Speed, latency, and reliability — benchmarked every ~4 hours.
What is Kimi K3?
Kimi K3 is a language model served through the OpenCode Go API. This page tracks its real inference speed — tokens per second, time to first token, and 24-hour reliability — sampled every ~4 hours on the same fixed prompt and the same measurement method used for every other model, so the numbers are directly comparable across providers. See the full methodology for how each sample is taken.
- Latest TPS
- 40.2 tok/s
- 24h avg TPS
- 40.2 tok/s
- TTFT
- 2.5s
- 24h reliability
- 100%
- Last tested
- 6m ago
How Kimi K3 compares
By latest measured tokens per second, Kimi K3 ranks 12 of 17 available models on OpenCode Go, and 42 of 49 across all providers benchmarked here. This is a snapshot from the last site build — the live stats above and the live leaderboard show its current position.
Tokens per second over time
Time to first token over time
Reliability
Every benchmark attempt counts here — including timeouts and errors. TPS and TTFT charts above show successful runs only, so you can still see speed data while reliability is red.
Speed is measured as output tokens ÷ generation time. Each benchmark fires one streaming chat completion with a fixed ~300-token prompt, every ~4 hours. Full methodology →