Each Saturday our AI Arena rig runs four fixed tests against a rotating set of models and logs tokens per second, time to first token (TTFT) and VRAM. This week's batch paired two cloud reasoning models, DeepSeek V4 Pro and Gemini 3.6 Flash, on a PHP Drupal module, an HTML/JS animation, a Python "galaxy" snippet and a long English article. The local Ollama runs on our RTX 5060 Ti 16 GB did not land in the database this week, so I will not quote stale local figures as if they were fresh.
The numbers this week
The table shows the latest run per model and test, exactly as the harness recorded it. TTFT is logged in milliseconds and shown here in seconds.
| Model | Test | tok/s | TTFT | Tokens |
|---|---|---|---|---|
| Gemini 3.6 Flash | PHP Drupal module | 260.89 | 6.9 s | 946 |
| Gemini 3.6 Flash | HTML/JS animation | 261.93 | 18.3 s | 163 |
| Gemini 3.6 Flash | Python galaxy | 183.30 | 21.8 s | 162 |
| Gemini 3.6 Flash | English article | 153.90 | 12.8 s | 2 052 |
| DeepSeek V4 Pro | English article | 128.52 | 15.6 s | 4 096 |
| DeepSeek V4 Pro | PHP Drupal module | 4 096 000* | 0 ms* | 4 096 |
| DeepSeek V4 Pro | HTML/JS animation | 4 096 000* | 0 ms* | 4 096 |
| DeepSeek V4 Pro | Python galaxy | 4 096 000* | 0 ms* | 4 096 |
* logging artifact, not a real measurement. See the next section.
DeepSeek's 4,096,000 tok/s is not a real number
Three of the four DeepSeek V4 Pro runs came back with 4,096,000 tokens per second and a 0 ms first token. No model streams that fast. The API returned the whole 4,096-token answer as a single non-streamed chunk, so the harness measured the gap between two tokens as roughly one millisecond and TTFT as zero. Dividing the total duration by the token count gives a real throughput of about 125 tok/s on the PHP task, 122 tok/s on the HTML/JS animation and 109 tok/s on the Python snippet.
The PHP run also logged a duration of exactly 32,767 ms, which is 215 − 1, the ceiling of a signed 16-bit integer. That figure is a floor, not an exact time. Only the English-article test streamed cleanly for DeepSeek: 128.52 tok/s with a 15.6 s first-token delay. That single clean run is the honest basis for comparison.
Most of the wait happens before the first token
Both models are reasoning models. They do not start typing the moment you press enter. They generate a hidden chain of thought first, then the visible reply. TTFT measures the time from the request to the first visible token, so it captures that thinking phase.
Gemini 3.6 Flash spent 6.9 seconds thinking on the PHP task, 18.3 seconds on the animation and 21.8 seconds on the Python snippet. On those last two tasks the thinking was almost the entire run: 18.3 seconds of an 18.9-second total, and 21.8 seconds of a 22.7-second total. The model then produced a short answer, 163 and 162 tokens, and stopped. The small output is a test-parameter limit, not a speed measure, but it makes the point clearly. Throughput of 260 tok/s did not help when 96 percent of the wall-clock was spent before the first token.
This is the practical story for a developer. Ask an assistant to fix a function and you stare at a cursor for 7 to 22 seconds before anything appears. Raw tokens per second then matter only after the wait. In a coding loop with many small prompts, TTFT is the number you actually feel.
A €0 local model against a per-token cloud bill
This week's batch was cloud-only, so I compare on economics and on what the numbers point toward. A local model on our RTX 5060 Ti 16 GB costs €0 per token, runs offline and keeps every prompt inside your own network, which matters under GDPR. A self-hosted model also avoids the processor and data-transfer paperwork that comes with a US-hosted API. The trade-off is size: 16 GB of VRAM fits models in the 8 to 14 billion parameter range, quantized, and those do not match a frontier reasoning model on hard tasks.
The cloud models bill per token. DeepSeek publishes its prices and I verified them this week. Output costs $1.98 per million tokens off-peak (about €1.82) and $3.96 at peak, while input on a cache miss costs $0.66 per million off-peak. Each of our DeepSeek runs returned 4,096 output tokens, which works out to roughly $0.008, or €0.007, per run. Cheap on its own. DeepSeek V4 Pro also advertises a 1M-token context and a 384K max output.
Gemini 3.6 Flash is Google's low-cost Flash tier, but its pricing page was unreachable from our server this week, so I will not quote a number I could not confirm. Flash-tier models are designed for high-volume, low-cost work.
Here is the rough maths. A card like the one in our rig costs on the order of €500 as a one-off purchase. At €0.007 per 4,096-token generation, a cloud bill would need roughly 70,000 such calls to match that card price. That ignores electricity and ignores the fact that the local model is weaker. For a developer issuing thousands of small prompts a day, a local model for boilerplate plus a cloud model for the hard parts is a sensible split. You can put models side by side on /ai-arena/compare.
Why did Gemini return only 163 tokens on two tests?
The harness caps output length for fairness, and on the short code tasks the model stopped early. Output length is a test parameter, not a measure of speed or quality.
Does a 22-second first token mean the model is slow?
No. Reasoning models spend that time on a hidden chain of thought before the visible reply. The visible text then streams quickly. It is a latency cost, not a throughput one.
How do I see the full results and run these tests myself?
The complete history and methodology are published on /ai-arena, with side-by-side comparisons on /ai-arena/compare.