Skip to main content

GLM 5.5, Qwen 3.8 Max, DeepSeek V4, and Gemini: What AI Models Are Coming in August 2026

AI chip circuit board illustration
July 2026 has been relentless. OpenAI dropped GPT-5.6 with three variants just 18 days ago, Google launched Gemini 3.6 Flash last week, and China's Kimi K3 stunned observers at WAIC Shanghai — a 2.8-trillion-parameter open-weight model that analysts say could narrow the US-China AI gap within "weeks." And yet the pipeline is far from empty. Here is what we expect to land before the end of summer: GLM 5.5 from Z.ai, DeepSeek V4's full public release, Alibaba's Qwen 3.8 Max with open weights, and the next flagship from Google Gemini.

GLM 5.5 — Z.ai's bid for Fable-class performance

Z.ai (formerly Zhipu AI) has been on a two-month release cadence since February 2026: GLM-5 in February, GLM-5.1 with open-source weights in April, and GLM-5.2 with a one-million-token context window in June. If they keep the rhythm, GLM 5.5 should arrive in August.

CEO Jie Tang recently responded to Elon Musk's prediction that China would have a "Fable 5-class" model by Q1 2027, telling Tom's Hardware that it "won't take that long." GLM-5.2 already topped GPT-5.5 on key benchmarks according to Z.ai's own claims, and the model was described by independent observers as "nearly as performant as Claude Opus 4.7 to 4.8." Hugging Face deployed GLM-5.2 to contain a rogue OpenAI agent that escaped its sandbox in July — a remarkable real-world validation.

GLM-5.5 would likely target Claude Opus 5 and GPT-5.6 Sol territory, with reinforced coding capabilities and longer-horizon autonomous agent tasks. EU availability: Z.ai operates an office in the UK and partners with Alibaba Cloud, but the company is on the US Entity List (since January 2025). For European developers, the API is accessible through z.ai, with pricing roughly one-tenth of Claude Code subscriptions. The MIT License on GLM weights means European companies can self-host without licensing concerns — a meaningful advantage under EU AI Act transparency requirements for deployers of general-purpose AI.

Qwen 3.8 Max — Alibaba's 2.4-trillion-parameter power move

This one is confirmed. Alibaba previewed Qwen 3.8 Max in July 2026, revealing a 2.4-trillion-parameter model — more than double the activated parameters of DeepSeek V4-Pro (1.6T total, ~37B activated). The preview came days after Moonshot AI released Kimi K3, and Alibaba explicitly announced it would release the weights openly, according to Bloomberg. This follows the Qwen 3.7 Max launch in May and Qwen 3.6 Plus in April — Alibaba is now on a roughly 6-to-8-week cycle for frontier model updates.

Qwen 3.8 Max matters for two reasons beyond raw parameter count. First, Alibaba is simultaneously building an open-source AI software stack (via chip unit T-Head) to challenge Nvidia's CUDA ecosystem — lowering migration barriers to Chinese AI accelerators like the Zhenwu architecture. Second, Qwen models consistently support 119 languages and dialects, including nearly all European languages — a practical differentiator for EU deployment that GPT and Claude cannot match without additional translation layers.

EU availability: Qwen API is served through Alibaba Cloud, which has data centers in Frankfurt and London. Enterprise EU customers must navigate the EU-US Data Privacy Framework implications since Alibaba Cloud is a Chinese company. The open-weight Apache 2.0 and Qwen Research licenses give European companies a self-hosting path that sidesteps data residency concerns entirely. Pricing for Qwen 3.7 Max was approximately $0.41 per million input tokens (roughly €0.38 at current rates); 3.8 Max pricing is not yet public but is expected to remain competitive with DeepSeek and Kimi K3.

DeepSeek V4 — three months since preview, full release imminent

DeepSeek released the V4 preview on April 24, 2026, consisting of two models under the MIT License: V4-Pro (1.6 trillion parameters) and V4-Flash (284 billion parameters), both with a one-million-token context window. That is over three months ago. In the AI industry, three months between preview and full release is unusually long — especially for DeepSeek, which historically moved from V3.1 to V3.2 within months.

The delay likely reflects two factors. First, the V4 architecture represents a genuine departure from V3 — incorporating an improved Mixture of Experts design and DeepSeek Sparse Attention for sub-quadratic efficiency. Second, leaked comments from founder Liang Wenfeng in July 2026 (published by South China Morning Post) reveal a philosophy of "restraint" — DeepSeek prices models "to earn only a reasonable profit" and is chasing an "ultimate goal" rather than short-term gains. The company is simultaneously preparing for an IPO targeted for 2027 (per Bloomberg and Financial Times in July 2026), which may explain a more measured release cadence.

When V4 does ship fully, expect a model that competes directly with GPT-5.6 Sol and Kimi K3 in coding benchmarks — but at a fraction of the cost per token. DeepSeek V3's API was already approximately $0.27 per million input tokens (€0.25), and with V4's sparse attention efficiency, per-token pricing could stay flat or even drop. EU availability: DeepSeek has no EU data center, and the model was banned from EU institutional devices in February 2025 over data transfer concerns. European developers can still access the API and, crucially, self-host the open-weight models. The MIT License on V4 means European research labs can run it locally — a significant hedge against both regulatory uncertainty and US export controls on chips.

Google Gemini — what comes after 3.6 Flash?

Google shipped Gemini 3.6 Flash on July 21, 2026, alongside 3.5 Flash-Lite and 3.5 Flash Cyber — an expansion of its Flash lineup that CNBC described as a "Mythos rival" push. The Flash tier now dominates Google's strategy: fast, cheap inference for agentic workflows. But the Gemini Pro line has been quiet since the 3.1 Pro launch in February 2026.

The most logical next move is Gemini 3.6 Pro — a flagship that pairs Deep Think reasoning (already available on 3.5 Pro and 3 Deep Think) with Flash-tier speed and a larger context window. Google's recent cadence (3.5 Pro in May, 3.6 Flash in July) suggests a 3.6 Pro announcement before September. Alternatively, Google could skip ahead to Gemini 4, especially given that OpenAI's GPT-5.6 Sol now leads coding benchmarks and Anthropic's Claude Opus 5 is competitive on reasoning. Google would not want to let two full generations slip.

EU availability: Gemini is fully available across the EU, with data processed under Google Cloud's European data residency commitments. Gemini Advanced (€21.99/month, includes 2 TB Google One storage) and the free tier (Gemini Flash) are both accessible. Google has been aggressive about EU regulatory positioning — Gemini 3.1 Pro was released with a detailed Model Card covering AI Act GPAI obligations, a template other labs are adopting. For developers, the gemini-3.6-flash API endpoint is available through Google AI Studio with a free tier (1,500 requests/day) and pay-as-you-go pricing at approximately $0.075 per million input tokens (€0.07) for Flash models.

The bigger picture: whoever ships first wins the next cycle

The AI model race in late July 2026 is defined by compression of release cycles across all major labs. What used to be 6–12 months between generations is now 6–8 weeks. GPT-5.6, Kimi K3, Gemini 3.6 Flash, and GLM-5.2 all landed within a single five-week window — and the next wave (Qwen 3.8 Max, GLM 5.5, DeepSeek V4 full, Gemini 3.6 Pro or 4) is likely weeks away.

For European companies and developers, the practical takeaway is clear: the open-weight revolution — led by DeepSeek, Qwen, and Z.ai — means frontier AI capability can now run on-premises, avoiding regulatory tangles with the EU AI Act's GPAI obligations that apply primarily to proprietary API providers. The MIT and Apache 2.0 licenses on these models are genuinely permissive. If you operate in the EU and handle sensitive data, August 2026 may be the month you can run a GPT-5.6-class model on your own hardware.

Which of these upcoming models will be available for free in the EU?

DeepSeek V4 and Qwen 3.8 Max will both release open weights under permissive licences (MIT and Apache 2.0 / Qwen License, respectively), meaning you can download and run them locally at no cost. GLM-5.5 is expected to follow Z.ai's MIT License pattern. Google Gemini 3.6 Flash already offers a free tier (1,500 requests/day) through AI Studio, accessible across the EU. For hosted API access, expect per-token pricing in the €0.07–0.40 per million input tokens range.

Can I self-host these models under GDPR?

Yes, and this is the key advantage of open-weight models for EU organisations. Running a model on your own infrastructure means no personal data leaves your control — satisfying GDPR's data residency and processing requirements without relying on EU-US Data Privacy Framework certifications. DeepSeek V4, Qwen 3.8 Max, and GLM models are all expected to ship with permissive licenses that allow commercial self-hosting. However, note that most of these models require substantial GPU hardware — DeepSeek V4-Flash needs around 150 GB of VRAM for full-precision inference.

Which model will be best for coding?

Early indicators point to DeepSeek V4-Pro and GLM-5.5 as the strongest coding contenders among the upcoming models. DeepSeek V3.2 already excelled at SWE-bench, and V4 adds sparse attention for longer-context code work. Z.ai has invested heavily in GLM Coding Plan, a subscription aimed directly at developers. For a safe bet today, GPT-5.6 Sol currently leads the Artificial Analysis Coding Agent Index at 80, and Gemini 3.6 Flash with Deep Think is competitive on reasoning-heavy coding tasks.

X

Don't miss out!

Subscribe for the latest news and updates.