Argon arrives after a long silence. Google DeepMind went more than seven months without shipping a top-tier flagship, the last one being Gemini 3.1 Pro, and skipped the planned Gemini 3.5 frontier model entirely. Instead of iterating on Flash variants, the company moved straight to the Gemini 4 generation. The timing is not accidental: Anthropic released Claude Sonnet 5.5 on September 28, OpenAI followed with GPT-6 Astra on September 29, and the launch material reads like a lab trying to get back into the conversation.
The benchmark numbers Google published
Google's published benchmark table places Argon first in 13 of the 19 categories it lists, and the headline wins are real inside that table. On the Vals Index, Argon scored 68.9%. Claude Opus 5.5 sits at 67.0%, Claude Fable 5.1 at 65.8% and GPT-6 Astra at 63.1%. The margin over Opus is 1.9 points, which is real but not a rout.
Where Argon looks stronger in Google's own results is automation. The company lists 51.3% on AutomationBench against Opus 5.5's 42.5%, and 77.9% on DeepSWE v1.1 against 74.2% for Opus and 74.1% for GPT-6 Astra. Those are the agentic software-engineering tasks that production pipelines actually lean on, and an 8.8-point gap on AutomationBench is the kind of difference you notice when a workflow either completes or needs a human to finish it. These are vendor-run figures, so they are Google's claims rather than neutral measurements.
Independent results from Artificial Analysis
Artificial Analysis runs its own harness instead of relying on vendor numbers, and its picture of Argon is flatter. On the Artificial Analysis Intelligence Index, Argon scores 53 — a tie with GPT-6 Astra and Claude Fable 5.1, but behind Claude Opus 5.5 at 58 and Claude Sonnet 5.5 at 56. The same evaluation measures the cost of completing a task at about $1.99 with Argon at the promotional price, and records an average of 62,000 output tokens per task against 27,000 for GPT-6 Astra. Longer answers eat into whatever the lower per-token price saves.
On the independent agent benchmarks, Artificial Analysis reports about 77.5% on AutomationBench-AA and about 57% on Terminal-Bench 4. Those numbers come from a different harness than Google's, so they should not be mixed with the vendor figures above: the 51.3% AutomationBench and 77.9% DeepSWE v1.1 results are Google's, while the AutomationBench-AA and Terminal-Bench 4 figures are the independent ones.
Where Argon still trails
Terminal-Bench 4 tells a different story, and here the two sources broadly agree on Argon itself. Google's own table lists Argon at 57.4% and Claude Opus 5.5 at 66.4%, while the independent run lands Argon at roughly 57% on the same benchmark. That is a nine-point deficit on shell-level work, the messy category of navigating file systems, chaining commands and recovering from errors. Google's own summary concedes that OpenAI and Anthropic still win in specialized software, terminal and post-training tasks.
So the honest reading of Google's 13 first places out of 19 categories plus one tie is this: Argon has the widest coverage, not the highest peak. A model that wins broadly and loses narrowly on the specific things certain teams care about is a different product from a model that simply dominates — and Artificial Analysis's index, where Argon ties two rivals rather than beating them, is the counterweight to the vendor table.
One million tokens of output, and the caveat
The output ceiling is the change most developers will feel first. Going from 64,000 to 1,000,000 tokens in a single response means an agent can write a large codebase module, draft a long report or carry an extended reasoning chain without being cut off mid-thought and resumed by the client. Anyone who has watched a long generation die at the token limit knows why this matters.
What the launch material does not provide is independent evaluation of coherence at that length. A raised ceiling is a capability statement, not a quality guarantee, and no third-party long-context benchmark for Argon has been published yet. Artificial Analysis's 62,000-token average output length is a hint about typical behaviour, not a measure of quality at the limit.
Pricing, converted to euros
Introductory API pricing is $2 per million input tokens and $10 per million output tokens, with standard pricing later rising to $4 and $20. Cached input gets a 95% discount, which brings cached reads down to $0.10 per million. Converting at roughly $1.08 to the euro, that puts the intro tier at about €1.85 per million input and €9.26 per million output, rising to €3.70 and €18.52. Cached input lands near €0.09 per million.
Here is how that stacks up against the published list prices of the flagship competitors. The euro column is my own conversion at the same rate, so treat it as an approximation rather than a billing figure.
| Model | Input / 1M | Output / 1M | Approx. EUR (in / out) |
|---|---|---|---|
| Gemini 4 Argon (intro) | $2.00 | $10.00 | €1.85 / €9.26 |
| Gemini 4 Argon (standard) | $4.00 | $20.00 | €3.70 / €18.52 |
| GPT-6 Astra | $10.00 | $50.00 | €9.26 / €46.30 |
| Claude Sonnet 5.5 | $10.00 | $50.00 | €9.26 / €46.30 |
| Claude Fable 5.1 | $10.00 | $50.00 | €9.26 / €46.30 |
| Grok 4.7 | $2.00 | $6.00 | €1.85 / €5.56 |
| DeepSeek-V4.1-Flash (peak) | $0.30 | $1.20 | €0.28 / €1.11 |
| GLM-5.3-Flash | $0.15 | $0.50 | €0.14 / €0.46 |
| Mistral OCR 4.1 | $4.00 per 1,000 pages | €3.70 per 1,000 pages | |
Argon's promotional rate is well below the three frontier rivals in that table, which are all published at $10 input and $50 output per million tokens — Claude Sonnet 5.5, Claude Fable 5.1 and GPT-6 Astra alike. Google is buying adoption now and doubling its own rates later. The more useful comparison for buyers is cost per completed task rather than cost per token: Artificial Analysis measures Argon at about $1.99 per task at the promotional price, and because an average Argon task emits 62,000 tokens, the headline discount narrows once you pay for all that output. DeepSeek and Zhipu continue to undercut everyone by an order of magnitude, though self-hosting GLM-5.3-Flash under its MIT licence shifts the cost from tokens to hardware. On a single RTX 5060 Ti 16 GB in our AI Arena rig, that trade-off tends to pay off only at sustained volume, because the hardware bill is fixed rather than per-token. Open-weight models are also the most straightforward way to keep prompts out of a US provider's infrastructure, but they are not the only route — EU-hosted deployments and European data-processing regions exist for closed models as well, and the deciding factor is usually the contract, not the licence.
Availability in Europe
Argon is being rolled out in phases. The first wave covers internal Google teams and vetted partners under the Fairwind program, which focuses on cybersecurity work. Google says it plans to expand access to paying API customers and to Google AI Ultra subscribers, although it has not given a date. For a European developer reading this today, the practical situation is simple: you cannot build on it yet.
If Argon does reach the Gemini API in the EU, two things will matter for procurement. The first is data residency. Google offers EU-based processing for parts of Vertex AI, and whether Argon falls under those guarantees is unconfirmed. The second is the EU AI Act. Article 50 transparency obligations apply to providers and deployers of systems that generate synthetic content or interact directly with people, and they are now mandatory. The duties are not identical for everyone: machine-readable marking of AI-generated output falls primarily on providers such as Google, while deployers carry their own disclosure duties, for instance when publishing deepfakes or AI-generated text on matters of public interest. Enforcement sits with the EU AI Office and national authorities rather than voluntary codes of practice. The European Commission's AI Act pages and the text of Regulation (EU) 2024/1689 set out the scope and the exemptions. Companies preparing production deployments should ask providers for documentation on how that marking is implemented, not assume it.
For now, the model that leads most of Google's published benchmark table is also the one most Europeans cannot touch. That gap between published capability and delivered access is the part worth watching.
Can I access Gemini 4 Argon through the free Google AI Studio tier?
Not today. The first phase covers internal Google teams and Fairwind cybersecurity partners. Google says it plans to expand access to paying API customers and Google AI Ultra subscribers, but it has not given a date.
Does the 1,000,000-token output limit mean better long-form quality?
Not necessarily. The limit describes the maximum length of a single response. No independent benchmark measuring coherence across that full length has been published, so quality at extreme output sizes is untested by third parties.
Is Argon cheaper than GPT-6 Astra?
On list prices, yes. Argon's promotional rate is $2 and $10 per million tokens against GPT-6 Astra's $10 and $50, and its standard rate of $4 and $20 is still lower. Per completed task the gap is narrower: Artificial Analysis measures about $1.99 per Argon task at the promotional price, with an average of 62,000 output tokens produced per task.