Skip to main content

Gemini 3.7 Flash: $0.75/1M tokens, 65.3% DeepSWE, 1M context — what EU devs need to know

Ilustrační obrázek
Google released Gemini 3.7 Flash on August 13, 2026 — and it is a serious workhorse update, not a cosmetic bump. The API price drops to $0.75 per 1M input tokens until the end of the year, while the agentic coding benchmark DeepSWE v1.1 jumps from 49.0% to 65.3%. Here is what changed, what it costs in EUR, and what it means for European teams running real workloads.

Google's Flash models have quietly become the default engine for production AI work — a bit like the mid-range GPU that nobody gets excited about until they look at the price-performance ratio. Gemini 3.7 Flash, released August 13, pushes that formula further. Google positions it specifically for software engineering, web development, and agentic workflows, and this time the benchmark numbers back the positioning with unusually large jumps over the previous generation, as MarkTechPost reports.

A hosted workhorse with room to work

Like every Gemini before it, 3.7 Flash is strictly a hosted model: you access it through Google AI Studio or Vertex AI, and it is also integrated into Google's Gemini Spark tools for agent-style tasks. There are no open weights and no local runtime. The model supports a 1M-token input context and up to 64K output tokens — an unusually generous output ceiling that matters for agentic work: long plans, large diffs, and extended tool-call sequences fit in one response.

Pricing: the deal, and the expiry date

Until December 31, 2026, the API costs $0.75 per 1M input tokens and $3.75 per 1M output tokens. From January 1, 2027, rates double to $1.50 and $7.50. In other words, Google is running a temporary 50% discount; after the new year, the model costs what Gemini 3.6 Flash cost at list price. In EUR, at roughly $1 ≈ €0.87:

  • Intro rates (until Dec 31, 2026): ≈ €0.65 per 1M input, ≈ €3.26 per 1M output
  • Standard rates (from Jan 1, 2027): ≈ €1.31 per 1M input, ≈ €6.53 per 1M output

For context, here is how the current field compares:

  • GPT-5.6 Sol — $5.00 / $30.00 (≈ €4.35 / €26.10)
  • Grok 4.6 — $2.00 / $6.00 (≈ €1.74 / €5.22)
  • DeepSeek V4-Flash — $0.14 / $0.28 (≈ €0.12 / €0.24)

The benchmarks that matter

Benchmark3.7 Flash3.6 FlashChange
DeepSWE v1.1 (autonomous software engineering)65.3%49.0%+16.3 pts
FrontierCode 1.1 Main (coding)43.6%34.4%+9.2 pts
AutomationBench (long-horizon automation)30.4%17.0%+13.4 pts
GDP.pdf (document understanding)34.0%22.0%+12.0 pts
WebDev Arena (Elo rating)15881538+50

The pattern is clear: the biggest improvements are in agentic benchmarks. AutomationBench nearly doubled, and DeepSWE — a measure of how often a model autonomously resolves a real-world software engineering ticket — improved by a relative 33%. Still, a little skepticism is in order: a 65.3% pass rate means roughly one in three tasks fails. This is an excellent workhorse, not a miracle worker.

The GDP.pdf jump also deserves attention in Europe. Document-heavy workflows — contracts, invoices, regulatory filings — are exactly where EU enterprises already deploy models, and a 12-point improvement in structured document understanding directly changes the economics of those pipelines.

The European angle: availability, AI Act, and data residency

Gemini 3.7 Flash is available in the EU through the same channels as previous Gemini models. Google Cloud operates EU regions on Vertex AI, so companies with data-residency requirements can keep traffic inside the EU — the practical GDPR-compliant path. Gemini models support a broad set of languages, including major European ones; the model card in AI Studio is the definitive source for exact coverage.

There is also a regulatory change to plan around. As of August 2026, EU AI Act obligations for general-purpose AI are in mandatory enforcement: national authorities and the EU AI Office are enforcing binding transparency, governance, and risk-mitigation rules under the AI Act and the 2026 AI Omnibus framework. The phase of voluntary draft guidelines and grace periods is over. If you deploy 3.7 Flash in an EU-facing production system, document the intended use and your risk assessment — and for anything near high-risk territory, keep human oversight in the loop. Agentic systems that act autonomously are exactly what regulators will be watching.

Fast iteration is a feature — and a risk

Gemini 3.7 Flash lands just three weeks after 3.6 Flash. That cadence tells you two things. Google is iterating the Flash line aggressively based on developer feedback — and the flagship Gemini 3.5 Pro is still unreleased, still in partner testing. The Flash family is Google's real production line right now. But rapid release cycles mean production teams must re-run their own evaluations continuously; a benchmark that held in July may not hold in September. Our AI Arena project exists precisely to measure real throughput and long-run agent behavior, not vendor slides.

What you can build with it

The combination of low price, 1M context, 64K output, and better tool use makes 3.7 Flash a credible engine for agents that automate software maintenance, browser workflows, and document processing. The cost math is the headline: at introductory rates, a pipeline processing 10M input and 1M output tokens per day costs about $11.25 — roughly €9.80, or about €295 per month, before Vertex AI hosting. That is within reach of mid-sized European companies that previously wrote off agent experiments as too expensive.

If your compliance rules require on-premise or sovereign deployment, 3.7 Flash is not an option. The European road then leads to commercial platforms like Mistral's enterprise offerings, while open-weights alternatives include Meta's Muse Glimmer and Zhipu's GLM-5.2 — each with its own licensing and geopolitical trade-offs.

Is Gemini 3.7 Flash available in the EU?

Yes. It is served through Google AI Studio and Vertex AI, and Google Cloud's EU regions enable data-residency-compliant production use. Pricing is global; EU customers are billed in EUR through Google Cloud.

Will the $0.75 price stay after 2026?

No. $0.75 per 1M input and $3.75 per 1M output are introductory rates valid through December 31, 2026. From January 1, 2027, they double to $1.50 and $7.50 per 1M tokens.

Is there an open-weights version of Gemini 3.7 Flash?

No. Google ships it only as a hosted API with no downloadable checkpoint. For on-premise or sovereign AI, European companies typically turn to Mistral's enterprise platforms or open-weights models like Muse Glimmer and GLM-5.2.

Discussion

No comments yet — be the first to share your thoughts.
X

Don't miss out!

Subscribe for the latest news and updates.