Skip to main content

DeepSeek V4 Pro reported 4.1 million tokens/sec. We recalculated: it's 105

Ilustrační obrázek
Our AI Arena rig logged an impossible result this week: DeepSeek V4 Pro “measured” 4,096,000 tokens per second with 0 ms first-token latency. That is not a supercomputer — it is a bug. 4,096,000 is exactly 4,096 × 1000, a units error in our own pipeline. Recomputing from the actual run times and token counts gives a real throughput of roughly 76–105 tok/s. The honest headline is not a model-versus-model contest; it is why measuring yourself still beats trusting a number on a screen.

Every Saturday our AI Arena rig runs a battery of coding and writing tests against local and cloud models and logs raw throughput, time-to-first-token (TTFT) and token counts. This week’s fresh set — generated 19 September 2026 — contained eight unique model × test combinations. No local Ollama model produced fresh runs in this window, and I am intentionally not repeating the usual cloud-model comparison here. The useful new information is the correction itself.

What the runner logged for DeepSeek V4 Pro

The table below is the summary exactly as our runner logged it for the DeepSeek V4 Pro cloud runs. I am showing the raw numbers, including the suspicious ones, because that is the whole point of running your own benchmark: you get to see when a metric is nonsense.

Testtok/s (logged)TTFTTokensDuration
PHP Drupal module4,096,0000 ms409638.9 s
HTML/JS animation4,096,0000 ms409638.8 s
Python galaxy4,096,0000 ms409649.1 s
English article4,096,0000 ms409653.9 s

The 4.1 million tok/s that never happened

Look at the table: 4,096,000 tok/s and 0 ms TTFT on every single test. Both figures fail a basic sanity check. 4,096,000 is 4,096 tokens times one thousand — a classic milliseconds-to-seconds or units slip somewhere in the aggregation layer, not a real measurement. And 0 ms first-token latency is not physically credible for a remote API over the network.

The trustworthy columns are duration and tokens, so I recomputed the throughput the old-fashioned way — tokens divided by seconds:

TestReal throughput
PHP Drupal module≈ 105.2 tok/s
HTML/JS animation≈ 105.7 tok/s
Python galaxy≈ 83.4 tok/s
English article≈ 76.0 tok/s

Why the corrected number is still useful

76–105 tok/s is a perfectly respectable number for a frontier model producing long outputs — and, notably, the same ballpark as a capable mid-size local model on our RTX 5060 Ti 16 GB. Two caveats are worth flagging: every DeepSeek run returned exactly 4,096 tokens, which suggests the output was capped at a max_tokens limit rather than finishing naturally, and DeepSeek V4 Pro runs in thinking mode by default (per DeepSeek’s own pricing page). A reasoning model that “thinks” for thousands of hidden tokens before it answers should not show a 0 ms TTFT — our wrapper is evidently not capturing the reasoning phase for this provider. I have filed that as a bug to fix in the runner, not a headline.

What this means for DeepSeek V4 Pro cost and privacy

Because the useful question is no longer “which cloud model wins,” the practical test is whether the corrected throughput changes DeepSeek V4 Pro’s value on its own. It does not. At off-peak prices of $0.66 per million input tokens and $1.98 per million output tokens (peak hours double those figures), DeepSeek V4 Pro remains extremely cheap per token. Our English-article run produced 4,096 output tokens — that single run cost roughly 0.8 US cents in output, or about €0.007 at the current rate.

A local model on our RTX 5060 Ti 16 GB costs €0 per token after the one-time hardware outlay — the power draw is the only ongoing expense, and it never phones your data home. For a European business, that is not just economics: it is GDPR. Keeping data on-premises sidesteps cross-border transfer headaches entirely. DeepSeek’s data residency is precisely where EU regulators have been uncomfortable — Italy’s data-protection authority moved to block DeepSeek in early 2025 over data-handling concerns, so some European teams will simply refuse to send company data to it regardless of price.

The honest split: the cloud wins on absolute capability, frontier reasoning, and zero maintenance, and at 0.8 cents a run it wins on convenience. A local 16 GB card wins on privacy, predictability (no queue, no peak pricing, no rate limit), and long-run cost when you generate a lot. They are complementary, which is why we keep both running — you can see how they stack up on our model comparison page.

Why this matters for your workflow

The meta-lesson this week is unglamorous but real: metrics lie unless you verify them. A vendor dashboard or an auto-generated log would have happily reported 4.1 million tok/s and moved on. Because we own the rig and read the raw JSON, we caught the units bug, recomputed the true figure, and flagged a TTFT metric that our wrapper is not measuring correctly for one provider. That is the entire reason AI Arena exists — not to chase headlines, but to produce numbers we can actually trust when we choose what to run in production.

If you are using DeepSeek V4 Pro today: it is fast enough and very cheap per token, but treat its “instant” first token with suspicion — it thinks before it answers, and our TTFT figure for it is unreliable this week. Measure the wall-clock time for your own tasks before trusting a vendor chart.

Is DeepSeek V4 Pro really 4 million tokens per second?

No. That figure is a units bug in our aggregation (4,096 tokens × 1000). Recomputing from the logged duration and token count gives roughly 76–105 tokens per second depending on the task.

What is TTFT and why does it matter more than tok/s for interactive use?

TTFT (time-to-first-token) is the delay before the first character appears. A model can stream quickly yet feel slow because it spends time “thinking” before answering. For chat and coding assistants, TTFT often dominates the perceived speed. DeepSeek V4 Pro’s 0 ms TTFT this week is not credible and our wrapper is not capturing the reasoning phase for that provider.

Should I buy a local GPU or pay for DeepSeek cloud?

If privacy, GDPR compliance and predictable zero-per-token cost matter, a local card like our RTX 5060 Ti 16 GB wins. If you need maximum capability and only pay per use, DeepSeek V4 Pro is hard to beat at roughly €0.007 for a 4,096-token output off-peak. Many teams, including us, run both.

Discussion

No comments yet — be the first to share your thoughts.
X

Don't miss out!

Subscribe for the latest news and updates.