Nemotron 3 Nano 30B (non-reasoning) tokens per second

Live benchmark data for Nemotron 3 Nano 30B (non-reasoning) on Ollama Pro. Speed, latency, and reliability — benchmarked every ~10 minutes.

What is Nemotron 3 Nano 30B (non-reasoning)?

Nemotron 3 Nano 30B (non-reasoning) is a reasoning model served through the Ollama Pro API. This page tracks its real inference speed — tokens per second, time to first token, and 24-hour reliability — sampled every ~10 minutes on the same fixed prompt and the same measurement method used for every other model, so the numbers are directly comparable across providers. See the full methodology for how each sample is taken.

Latest TPS
172.7 tok/s
24h avg TPS
109.2 tok/s
TTFT
452ms
24h reliability
100%
Last tested
7m ago

How Nemotron 3 Nano 30B (non-reasoning) compares

By latest measured tokens per second, Nemotron 3 Nano 30B (non-reasoning) ranks 1 of 18 available models on Ollama Pro, and 7 of 50 across all providers benchmarked here. This is a snapshot from the last site build — the live stats above and the live leaderboard show its current position.

Tokens per second over time

Time to first token over time

Reliability

Every benchmark attempt counts here — including timeouts and errors. TPS and TTFT charts above show successful runs only, so you can still see speed data while reliability is red.

Speed is measured as output tokens ÷ generation time. Each benchmark fires one streaming chat completion with a fixed ~300-token prompt, every ~10 minutes. Full methodology →