What Mistral shipped
Mistral Large 4, nicknamed "le Chonk" during development, is the French company's first flagship update since Large 3 and Medium 3.5. It is a sparse mixture of experts: one trillion total parameters, 49 billion active per token. That split decides what the model costs to serve. Memory has to hold all the weights, while compute follows the 49 billion that fire on each token.
The model is natively multimodal. Reporting on the launch describes a training run of roughly two months on 3,800 to 4,000 NVIDIA Grace Blackwell GPUs drawing about 10 megawatts — vendor-side figures we have not been able to verify independently, and Mistral has not published a technical report we could check them against. Steady 10 MW for 60 days works out to roughly 14.4 GWh of electricity anyway, which is the kind of number that explains why the €3 billion Series D Mistral has raised mattered.
The launch is an API preview on Mistral Studio. The weights are separate and come later. According to The Decoder's report on the launch, Mistral has said the public release is set for October 27, 2026. That date has not been confirmed by a second source, and the licence terms have not been announced at all.
Price per million tokens
Mistral's list price is $1.36 per million input tokens and $4.18 per million output tokens, with cached input at $0.14. For the first two weeks of the launch Mistral halves those rates, putting input at $0.68 and output at $2.09.
The table below is a dated launch snapshot, not a live price comparison. It uses GPT-6 Sol's $2/$10 list price, the rate current when the Mistral Large 4 preview opened. Artificial Analysis reports that GPT-6 Sol has since been replaced by GPT-6.1 Sol; we have not verified the successor's current pricing, so the OpenAI row is the older model's rate and may already be out of date.
| Model | Input / 1M | Output / 1M | Cached input / 1M |
|---|---|---|---|
| Mistral Large 4 | $1.36 | $4.18 | $0.14 |
| Mistral Large 4 (launch discount) | $0.68 | $2.09 | – |
| GPT-6 Sol (launch-snapshot pricing) | $2.00 | $10.00 | – |
| GPT-6 Luna | $0.10 | $0.50 | – |
| Claude Sonnet 5.5 | $2.00 | $10.00 | $0.20 |
| Grok 4.7 | $2.00 | $6.00 | $0.50 |
| DeepSeek-V4.1-Flash (peak) | $0.22 | $0.66 | – |
Mistral publishes in USD. European business customers pay in euros, and the euro amount depends on the exchange rate on the invoice date, so percentages are the more stable comparison. Against the $2/$10 list price used here for GPT-6 Sol and Claude Sonnet 5.5, Mistral Large 4 is 32% cheaper on input and 58% cheaper on output. Against Grok 4.7 it is 32% cheaper on input and 30% cheaper on output. DeepSeek-V4.1-Flash at peak rates is still about six times cheaper on both. GPT-6 Luna is roughly 14 times cheaper on input and eight times cheaper on output.
A worked example makes the spread clearer. Take a batch job that reads 500,000 input tokens and writes 100,000 output tokens, priced at the same launch snapshot.
| Model | Cost for 500k in + 100k out |
|---|---|
| Mistral Large 4 | $1.10 |
| Mistral Large 4 (launch discount) | $0.55 |
| GPT-6 Sol (launch-snapshot pricing) | $2.00 |
| Claude Sonnet 5.5 | $2.00 |
| Grok 4.7 | $1.60 |
| DeepSeek-V4.1-Flash (peak) | $0.18 |
| GPT-6 Luna | $0.10 |
Those figures assume no cache hits, so they are an upper bound for workloads that reuse the same prompt prefix. Where a cached input rate is listed — $0.14 per million tokens for Mistral Large 4, $0.20 for Claude Sonnet 5.5, $0.50 for Grok 4.7 — a repeated prefix costs less than the headline input price. The table does not list cached rates for the remaining models, so cache pricing across this group is not comparable here.
Where it lands on benchmarks
On the Artificial Analysis Intelligence Index, Mistral Large 4 scores 38, according to Artificial Analysis. The same index put Mistral Large 3 at 9, and Claude Opus 5.5 at 58 — so on this measure the gap between the best European open-weight model and Anthropic's top model is still 20 points.
The cyber results are the stronger part of the sheet, but most of them come from Mistral itself. The company reports 82% on CyberGym-E2E-AA for vulnerability reproduction and 93% on CyberBench. Artificial Analysis, an independent outfit, places the model at 50 on its Cyber Index. Mistral further reports 61.7% to 62% on DeepSWE v1.1 for agentic software engineering and 67% on FinWorkBench for financial agents. Vendor-reported scores are produced on the vendor's own harness and settings, so they are not directly comparable with independently produced numbers and should be read as claims rather than as a like-for-like ranking.
Composite indices describe a model in general and say very little about your specific task. A score of 38 against 58 means ML4 is not the model to reach for when the hardest reasoning matters most. The cyber numbers are where Mistral says it competes.
Security work and the refusal problem
Mistral's positioning is explicit: reproducing a vulnerability, auditing a codebase for exploitable paths, validating a proof of concept. Frontier models from US labs frequently decline that work, and security teams end up arguing with a system prompt instead of doing the audit. Mistral is selling ML4 as the model that does not put up that fight.
For European teams the timing matters. Manufacturers of products with digital elements face vulnerability-handling duties under the Cyber Resilience Act, and someone has to do the triage. Benchmarks measure capability, though, not guardrails. What a provider permits in its usage policy is a separate question from what the model can do, and the two get conflated in launch posts.
The same capability works for attackers. That is true of every security tool, and it becomes a deployment question the moment weights are downloadable rather than API-gated.
Open weights planned for October 27
The arithmetic on self-hosting is unforgiving. A trillion parameters at 8-bit precision is about 1 TB of weights. At 4-bit it is roughly 500 GB, before any context or KV cache. On our AI Arena rig with 16 GB of VRAM we cannot load this, so the only route for us is the API. Anyone with data-residency obligations needs a multi-GPU server or an EU-hosted inference provider.
The license for the weights has not been announced. Until it is, how far fine-tuning and redistribution can go is unknown — and the October 27 date rests on a single report rather than on a published release schedule.
Two caveats sit over all of this. Mistral Large 4 is still a preview, so prices, rate limits and benchmark sheets can move before general availability. And the weights are not downloadable today: until the licence is published, self-hosting is a plan, not an option.
The AI Act angle
The AI Act applies in phases, and the phase covering general-purpose models is now in force. Obligations for providers of general-purpose AI models have applied since 2 August 2025, and since 2 August 2026 the European AI Office has held enforcement powers over those providers, including the ability to request information and to impose fines. Most other AI Act obligations fall to national market surveillance authorities, whose staffing and priorities still differ noticeably between member states.
Article 50 transparency duties also apply from 2 August 2026, but they do not all land on the same party. The duty to make sure people are told they are interacting with an AI system, and the duty to mark synthetic audio, image, video and text in a machine-readable way, sit primarily with the provider of the AI system. A company that builds a chatbot on top of the Mistral API is usually a deployer, and the Article 50 duties that fall on deployers are narrower — for example disclosing deep fakes, or AI-generated text published on matters of public interest. Which role you actually hold depends on how the system is documented and marketed and on what you do with it, so calling an API is not automatically the moment every transparency duty becomes yours.
There is one practical advantage to a Paris-based provider. Mistral falls under the European AI Office directly, rather than through a third-country route with a separate compliance story. That simplifies the paperwork question without answering it. Guidance, codes of practice and harmonised standards are still being finalised, so none of this should be read as settled legal advice.
Can I run the open weights myself after October 27?
Only on serious hardware, and only if the release happens on schedule and the licence permits it. A trillion parameters at 8-bit precision is about 1 TB of weights, roughly 500 GB at 4-bit, and that is before context. A single 16 GB or 24 GB card will not load it. Self-hosting means a multi-GPU server or a hosted inference endpoint.
Mistral Large 4 or GPT-6 Luna for high-volume work?
Different classes. On the 500k-in, 100k-out example, GPT-6 Luna comes to about $0.10 against $1.10 for ML4. For classification, extraction or short generation, Luna is roughly ten times cheaper. ML4 earns its price where the cyber and agentic results apply.
Is Mistral Large 4 available in the EU today?
Yes, through the public API preview on Mistral Studio, and Mistral AI is based in Paris. For regulated workloads, confirm the processing region and the data-processing agreement before moving production traffic, because a preview is not the same thing as a contracted enterprise deployment.