Skip to main content

DeepSeek V4 Pro's '4 million tok/s' is a benchmark mirage — Gemini won every test

Ilustrační obrázek
DeepSeek V4 Pro logged a headline-grabbing 4,096,000 tokens per second in our AI Arena harness this week — and we're here to tell you that number is a measurement artifact. The endpoint returned the entire answer as one non-streamed response, leaving the harness with no measurable TTFT and triggering a fallback timing value for the tok/s calculation. Measured by total elapsed time, Gemini 3.6 Flash beat DeepSeek in all four tests we ran. Gemini won these four tests on total elapsed time only; these tests included no quality scoring.

What we actually measured this week

Every Saturday our AI Arena rig — an NVIDIA RTX 5060 Ti with 16 GB of VRAM — runs a fixed battery of tests: a PHP Drupal module, an HTML/JS animation, a Python "galaxy" rendering, and an English article. We record tokens-per-second, time-to-first-token, total duration and token counts, for both local Ollama models and cloud APIs.

This week's fresh batch came back cloud-only — two frontier models, four tests each, eight runs in total. No local models produced fresh results in this window, so the honest comparison today is the DeepSeek API's V4 Pro endpoint versus the Google Gemini API's Gemini 3.6 Flash endpoint. Here is the raw summary table straight from our harness. These are this week's internal AI Arena measurements, not a general ranking of model performance:

ModelTesttok/sTTFT (ms)Duration (s)Tokens
DeepSeek V4 ProPHP Drupal module4,096,000N/A38.94096
DeepSeek V4 ProHTML/JS animation4,096,000N/A38.84096
DeepSeek V4 ProPython galaxy4,096,000N/A49.14096
DeepSeek V4 ProEnglish article4,096,000N/A53.94096
Gemini 3.6 FlashPHP Drupal module240.91721311.0905
Gemini 3.6 FlashHTML/JS animation250.861675822.01317
Gemini 3.6 FlashPython galaxy193.692533426.2163
Gemini 3.6 FlashEnglish article138.861774930.41761

For this run, the DeepSeek API was called with streaming disabled and returned one complete response block. The Google Gemini API was called with streaming enabled. The harness measured duration with a wall-clock timer from request dispatch until the complete response was received; TTFT was the interval from dispatch to the first streamed chunk. Because DeepSeek sent no chunks, its TTFT is recorded as N/A, not as zero.

The "4,096,000 tokens per second" is a measurement artifact

Let's be blunt: no model emits four million tokens a second. In these runs, DeepSeek's endpoint returned the complete response as a single non-streamed block. The harness could not measure a real token interval and used its 1-millisecond fallback denominator for the tok/s field. The displayed value therefore came from 4,096 output tokens ÷ 0.001 seconds = 4,096,000 tok/s. That is a fallback calculation, not division by zero and not a valid throughput measurement. The full wall-clock duration remained measurable at 38.8–53.9 seconds, which is why the throughput figure cannot be treated as a real speed record.

There are two more tell-tale signs. First, DeepSeek V4 Pro returned exactly 4,096 tokens in every run. That repeated count suggests a configured output cap may have limited the response, so we never learned where it would have stopped. Second, Gemini 3.6 Flash, which streams normally, produced variable output lengths from 163 to 1,761 tokens. The two endpoints were therefore exposing different response-timing behaviour, and the DeepSeek figure is a buffer-based calculation rather than a stream measurement.

The number that survives scrutiny is total wall-clock duration. If you divide DeepSeek's 4,096 tokens by its real elapsed time, you get an effective throughput of roughly 105 tok/s on the PHP and animation tests, 83 tok/s on Python, and 76 tok/s on the article. That is not "4 million tok/s" — it is a fairly ordinary effective pace, and slower than Gemini's in this week's internal AI Arena measurements.

Wall-clock time is the ranking that matters

Developers don't feel tokens-per-second; they feel the seconds until the answer is done. Ranked by total elapsed time, the picture flips:

TestDeepSeek V4 ProGemini 3.6 FlashWinner
PHP Drupal module38.9 s11.0 sGemini
HTML/JS animation38.8 s22.0 sGemini
Python galaxy49.1 s26.2 sGemini
English article53.9 s30.4 sGemini

Gemini 3.6 Flash finished first in every test in this two-model harness, usually by a wide margin — approximately 3.5× faster on the Drupal module (38.9 ÷ 11.0). That makes it the winner on total elapsed time in this comparison, not necessarily the fastest or best model generally. This elapsed-time win is not an apples-to-apples speed or quality comparison: the models produced substantially different amounts of output, with DeepSeek returning exactly 4,096 tokens in every run and Gemini returning shorter responses from 163 to 1,761 tokens. The tests included no output-quality scoring, and TTFT, output quality and total cost can all matter alongside completion time.

TTFT: why a model can be "fast" and still feel slow

The other number worth staring at is time-to-first-token. Gemini 3.6 Flash waited between 7.2 and 25.3 seconds before emitting its first token, even though it then streamed at a healthy 139–251 tok/s. Prompt processing, queueing or internal reasoning may contribute to that pause; the API used in this test does not expose a separate reasoning interval, so we cannot attribute it to one specific cause.

In practice, a 25-second silent gap is the difference between an assistant that feels instant and one that feels broken. Our Python galaxy test is the sharpest example: Gemini spent 25.3 seconds before its first streamed token, then produced only 163 tokens in under a second. For an interactive coding session that cadence is fine; for a chat UI with no typing indicator, a user would assume it crashed.

DeepSeek's TTFT is not measurable in this run because its response arrived as one pre-assembled block. You can't optimise around a TTFT you can't actually measure. If you're building interactive tooling, test TTFT yourself on a real stream rather than trusting a spec sheet.

Local rig vs cloud: what it costs you in euros

This is where the European angle bites. Our RTX 5060 Ti runs local models with no per-token fee, but hardware, electricity and capital costs still apply — once the card is bought, the marginal cost is a few cents of electricity per hour. Cloud models charge by the million tokens, and the gap is getting interesting:

  • DeepSeek V4 Pro: we are not stating a current price here. V4 Pro pricing changed during 2026, and the exact current figure varies by provider and time, so the earlier $0.66/$1.98 figures and peak-hour claim should not be treated as current without a dated official price sheet.
  • Gemini 3.6 Flash (Google's published API pricing): the pricing page lists the introductory $0.75 / $3.75 per million input/output tokens used in this comparison; see the official Google Gemini API pricing page. This is the price information checked on 26 September 2026. The approximate €0.68 / €3.41 conversion assumes $1 = €0.91; card issuer rates, taxes and fees can change the amount actually charged.

For context, the output-only estimates for the 4,096-token DeepSeek response and Gemini's 1,761-token version are both small at the listed rates. They exclude input-token charges, however, so the exact request cost also depends on the prompt and any other billed input. Cloud inference is now cheap enough that the economics favour renting unless you run at serious volume or need data you cannot send off-premises. Provider and time can also change the applicable price, especially for DeepSeek V4 Pro.

That last point is the GDPR-shaped elephant in the room. A local model keeps data on your machine, which can matter if you process customer or patient data in the EU. A cloud model may require appropriate GDPR safeguards, including a data-processing agreement where applicable, depending on the data and deployment. AI Act duties likewise depend on the system, deployment and use case; a cloud deployment does not automatically trigger one specific transparency duty. The local option's lack of a per-token fee isn't free of context — it can support data residency, and for many European companies that is the deciding factor, not the price.

What to take away this week

Two models, four tests, and one clear lesson: benchmark the wall-clock time, not the headline number. DeepSeek V4 Pro is a capable model, but in this week's internal AI Arena measurements its "4,096,000 tok/s" tells you nothing about real performance — and the endpoint returned a non-streamed response, making that number possible through a fallback calculation. Gemini 3.6 Flash won every test on elapsed time in this comparison despite a slower streaming rate, because it streams at all. That result says nothing definitive about overall model quality, since we did not score quality and tested only two cloud models. The elapsed-time result is also not an apples-to-apples speed or quality comparison: DeepSeek produced exactly 4,096 tokens in every run, while Gemini produced substantially shorter responses, from 163 to 1,761 tokens.

For a European developer choosing a stack, the decision now looks like this: if you need privacy and run continuous workloads, a local 16 GB card can still be attractive on marginal cost, although hardware, electricity and capital costs remain. If you need frontier reasoning and don't mind a per-token bill, both cloud tiers are inexpensive at hobby scale, but provider pricing should be checked at the time of use. We keep the full history and a side-by-side view at /ai-arena/compare.

Why did DeepSeek V4 Pro show N/A for time-to-first-token?

Because its endpoint returned the whole response as a single non-streamed block. The harness had no incremental tokens to time, so TTFT was not measurable. Its displayed 4,096,000 tok/s came from the 4,096-token output divided by the harness's 0.001-second fallback denominator, not from division by zero or a real speed measurement. The full wall-clock duration was still measurable, so the figure is a response-timing artifact, not a real latency or speed measurement.

Is a local model really cheaper than the cloud?

A local model has no per-token fee, but hardware, electricity and capital costs still apply. Cloud models bill per million tokens, and their actual cost also includes input tokens. The main reasons to go local can therefore include data residency and privacy, not just price.

Which model should I pick for interactive coding?

Gemini 3.6 Flash finished faster in every elapsed-time test this week, but its long initial pause (up to 25 seconds) matters in a UI. Prompt processing, queueing or internal reasoning may contribute to that pause, but this API does not expose a separate reasoning interval. If responsiveness is the goal, test TTFT yourself on a real stream — that, alongside output quality and cost, is what your users will feel.

Discussion

No comments yet — be the first to share your thoughts.
X

Don't miss out!

Subscribe for the latest news and updates.