Mistral has put a trillion-parameter model on the table. On October 6, 2026 the French lab opened a public preview of Mistral Large 4, nicknamed "Le Chonk": a natively multimodal mixture-of-experts model with about 1 trillion total parameters and 49 billion active per token, available through the API now and promised as open weights by the end of October. Mistral's pitch is that it is the strongest open-weight model from the US or Europe on aggregated benchmarks.
The key word is preview. You can call it today, but you cannot download it, and every benchmark below is Mistral's own. This post separates what the official announcement and early coverage actually state from what remains to be proven, and explains what builders should do in the three weeks before the weights land.

Mistral Large 4 at a glance
| Question | Answer |
|---|---|
| What is it? | A hybrid instruct-and-reasoning mixture-of-experts model, natively multimodal (image input, text output) |
| Size | About 1T total parameters, 49B active per token |
| Training | From scratch on roughly 3,800 to 4,000 Nvidia Grace Blackwell GPUs in Mistral's European data centers (sources differ slightly), over about two months |
| Languages | Trained across 160+ languages, including all official EU languages |
| Availability today | Public preview via the Mistral API and Mistral Studio |
| Open weights | Promised by end of October 2026, reported as October 27 |
| Context window | 1M tokens, per the Mistral docs model page (version v26.10) |
| Price | List $1.36 input / $0.14 cached / $4.18 output per million tokens; the docs page also shows $0.68 / $0.07 / $2.09, a promotional half price that commenters say runs about two weeks |
| License | Custom Mistral license, per VentureBeat; full terms arrive with the weights |
| Headline claim | Best open-weight model from the US or Europe on aggregated benchmarks |
Update (October 6, 2026): Artificial Analysis puts Large 4 at 38
Independent scores have arrived. Artificial Analysis rates Large 4 Preview at 38 (38.4) on its Intelligence Index v4.3.2, tied with OpenAI's GPT-6 Luna and, per OfficeChai, the highest-scoring Western open-weights model on the chart. That is a large jump from Mistral Large 3 (9) and Medium 3.5 (14), but Trending Topics notes it is only eighth among open models overall, behind Xiaomi's MiMo-V2.6-Pro (46.3), Z.ai's GLM-5.3 (about 45) and Moonshot's Kimi K3 (about 44). Closed leaders such as Claude Opus 5.5 (57.6) remain well ahead. For the Xiaomi comparison, see our MiMo-V2.6 launch coverage.
Cost is the catch: the same report lists about $1.13 per task, versus $0.27 for DeepSeek V4.1 Flash, and roughly 200 million output tokens across the index against a median of 81 million, which Trending Topics ties to the higher cost. Speed is about 116 tokens per second. Artificial Analysis has also been revising the index frequently, so scores across versions are not directly comparable. The open weights and license terms are still pending.
Update (October 6, 2026): what the official post and docs add
Mistral's launch post and docs model page fill in several gaps.
- Specs. 1.05T total parameters, 49B active, a 1.6B vision encoder, and a 1M-token context. The docs label it "Open" and "Public Preview," with function calling, structured outputs, batching and agents on the API. It is also on OpenRouter.
- Weights. "Weights drop end of this month," with architecture details, more benchmarks and the post-training method promised alongside. Until then Mistral is red-teaming with cybersecurity leaders, vetted partners and state authorities, who get a version with reduced moderation and expanded cyber capabilities.
- Training. 3,800 Nvidia Grace Blackwell GPUs in Mistral's own European data centers, over 160 languages. The post says it uses the same training, customization and RL environment Mistral sells through Mistral Forge. At the current scale of about 3,000 GPUs for RL, one run produces roughly 33 billion tokens a day, about 16 billion of them trainable. Mistral says the RL run is still in flight with no sign of saturation.
- Money. ML4 is the first milestone funded by the 3 billion euro Series D.
Cyber, in Mistral's words. It scores 82% on one Artificial Analysis Cyber Index test (reproduce a real open-source vulnerability, then patch it), which it says is the highest of any model, and 93% on Cybench. It claims Claude Opus 5.5 and GPT-6 Astra score near zero on that test because they refuse the task. Mistral also says its own refusal rate on malicious cyber prompts is higher than any other open model's. Those two claims sit together awkwardly until the weights are public and independently tested.
More numbers (all vendor-reported).
| Area | Reported result |
|---|---|
| Terminal-Bench 4.0 | 28.3% |
| SWE-Atlas QnA | 59.4% |
| Blind human coding eval (Surge AI, 1 to 5) | 3.74, second of five, behind Claude Opus 5 (4.22), ahead of GLM-5.3 (3.60), Kimi K3 (3.59), GLM-5.2 (3.40) |
| AA-Briefcase (long-horizon knowledge work) | 1,393 Elo, ahead of DeepSeek V4 Pro |
| Finance and legal | Says it beats GPT-6 Astra on Vals.ai finance and legal tasks; tops open models on Harvey's Legal Agent benchmark |
| Safety | B3 attack resistance 93.3%; KORA 1.691 of 2 |
| Science | State of the art among open weights on SciCode-Verified |
How Hacker News reacted. The two threads (about 613 and 319 points) split. Fans liked the open weights, the European hosting and the cyber permissiveness. Critics called it "mediocre" and months late, and noted the Artificial Analysis cost per task. Several flagged the charts: bars sorted so Mistral sits next to the weakest competitor, and axes that do not start at zero. One commenter's reminder fits our advice below: compare on your own tasks, not on a vendor's chart. Others pointed out that the nickname is a Pokemon reference, and that the earlier Large 3 was weak, so this is a real step up.
What exactly did Mistral announce?
Per the Mistral announcement and VentureBeat's coverage, Large 4 is a granular mixture-of-experts model: the network contains many small expert sub-networks, and a router picks only a few for each token. That is how a 1T-parameter model runs with the per-token compute of a 49B one. The model takes text and images in and produces text out. It is not an image or audio generator.
Mistral says it trained the model with input from enterprises in finance, engineering, manufacturing, logistics and the public sector, and positions it for those workloads. The promise of open weights means you can run and customize the model on your own or sovereign infrastructure, which is the entire commercial argument for open weights over a closed API.
The release lands weeks after Mistral's Series D of 3 billion euros (about 24 billion dollars post-money, per VentureBeat), described as the largest tech equity raise in Europe. Funding is not the story here, but it explains the scale: a from-scratch run on thousands of Blackwell GPUs is not a side project.
The benchmark numbers, and how much to trust them
The figures below come from Mistral's announcement as relayed by VentureBeat and a summary of the official post. Treat them as vendor-reported until independent evaluators run the weights.
| Benchmark | Reported Large 4 result | Context |
|---|---|---|
| DeepSWE v1.1 (coding agents) | About 62% (61.7%) | Mistral's Coding Agent Index ranks it ahead of DeepSeek V4 Pro and Qwen3.8 Max |
| Coding Agent Index | 49.8% combined | Mistral says ahead of DeepSeek V4 Pro 0813 and Qwen3.8 Max |
| Cybench (cyber challenges) | Solves 93% of challenges | Mistral contrasts with closed models that refuse such tasks |
| Artificial Analysis Cyber Index | Top five globally, per Mistral | Vendor claim |
| AutomationBench (657 business workflows) | 59.9% | Spreadsheet and slide-deck style deliverables |
| FinWorkBench | 67% | Preliminary result reported by VentureBeat |
| Dense 200 visual grounding | 42% | Mistral says it edges GPT-6-Astra at 41% |
| DIOR-RSVG (aerial imagery grounding) | 73% | Satellite and aerial scenes |
| Harvey Legal Agent Benchmark | 15% pass rate | A reminder that hard legal agent tasks remain mostly unsolved |
| B3 AI Security Benchmark | Resists 93.3% of attacks | Safety claim, vendor-reported |
Three things stand out.
The margins are thin. Leads of a few points on agentic coding indexes sit within the range where harness choice, sampling settings and benchmark version can flip the order. We covered how quickly these leaderboards move in the GLM-5.3 FlashX coverage.
The qualifier does a lot of work. "Best open-weight model from the US or Europe" deliberately excludes Chinese models, and the comparison set in Mistral's own tables includes them as the targets to beat. The claim is honest about the real competition: the strongest open weights have come out of China for a while, a theme we explored in the Qwen audit post.
The cyber result is a positioning choice. Scoring 93% on Cybench while saying the model refuses malicious requests at a higher rate than competing open models is a notable pairing. Closed frontier models often decline offensive-security tasks, and defenders complain about that. Mistral is selling a model that engages with security work but, it says, resists abuse. How that balance holds once the weights are public, where anyone can fine-tune away guardrails, is the obvious open question.
Why a 1T open-weight MoE is not the same as "run it yourself"
A mixture-of-experts model has a deceptive cost profile. Per-token compute follows the 49B active parameters, so inference is fast and API pricing can be low. Memory follows the 1T total parameters, because every expert must be resident in GPU memory in case the router picks it.
Rough arithmetic, ours and not Mistral's: at 8-bit precision a trillion parameters is on the order of 1 TB of weights, and at 4-bit about half that. Neither fits on a single workstation. Compare that with Aleph Alpha's Kolibri, a 78B MoE that needs around 78 GB, or the local-inference work in DwarfStar DS4. Large 4 is an open model for organizations with GPU clusters, not for laptops.
That is not a criticism. Open weights at this scale matter to three groups:
- Regulated and sovereign buyers who need the model inside their own perimeter. Our sovereign AI explainer breaks down the layers, and Large 4 is trained and served in European data centers.
- Inference providers and hosts, who can offer it at competitive prices once the weights ship, as happened with other large open models.
- Researchers and fine-tuners, who can inspect behavior, quantize, distill and adapt it.
If you are a solo developer, the practical benefit is indirect: more competition on price among hosted endpoints.
How does the price compare?
At the list price of $1.36 per million input tokens and $4.18 per million output tokens (half that during the preview promotion), Large 4 sits well below typical closed frontier pricing but is not rock-bottom for open-weight models. Cheaper small models exist, such as the free Ling 3.1 Flash. The relevant comparison is cost per solved task, not per token, and that needs your own evals. For the closed-model side of the market, see our Gemini 4 Argon vs Opus 5.5 vs Grok 4.7 vs GPT-6 Astra comparison.
What is still unverified
Being explicit about the gaps is more useful than repeating the headline:
- Weights are not public. Until the October 27 release, nobody outside Mistral can reproduce the benchmarks, inspect the architecture or confirm the license terms.
- License details. Reports say a custom Mistral license, not Apache 2.0. That could include commercial thresholds or use restrictions. Read it before building a product on it.
- Context length. The Mistral docs page lists 1M tokens, but the launch post does not discuss long-context quality, so test it on your own long inputs.
- Numbers differ between sources. The GPU count is 3,800 on Mistral's page and 4,000 in VentureBeat. Small, but a reminder to cite the official post when it matters.
- Preview means change. Pricing, rate limits and model behavior can shift before general availability.
- Self-reported safety scores. The B3 and KORA safety figures are vendor-reported and not independently audited.
What should builders do before October 27?
A short plan that costs little:
- Run your own evals through the preview API. Take 30 to 50 real tasks from your product, not benchmark prompts, and compare Large 4 against whatever you use now. Record pass rate, latency and cost.
- Test multimodal inputs you actually have. If you process PDFs, drawings or aerial images, the visual-grounding claims are the part most worth checking.
- Probe the security behavior. If you build defensive tools, try representative tasks and see where it helps and where it refuses.
- Plan hosting. Decide now whether you would use Mistral's API, a third-party host, or your own cluster, and what data residency you need.
- Read the license on day one. Check redistribution, fine-tuned derivative and commercial terms.
A sample prompt for a sanity check on structured extraction:
Extract every line item from the attached invoice image as JSON with
fields description, quantity, unit_price, total. If a field is unreadable,
return null for it rather than guessing, and list any rows you are unsure about.
Run the same prompt against your current model, then diff the outputs by hand. Ten minutes of this tells you more than any leaderboard.
How does it fit the open-weights race?
Large 4 arrives in a market where Chinese labs have set the pace on open weights, and where Western and European labs have been pushing back with sovereignty arguments: Aleph Alpha's Kolibri days ago, and the policy backdrop in our Europe AI landscape guide. The more interesting shift is in what Mistral chose to emphasize. It is not claiming to beat the best closed models overall. It is claiming parity-or-better with the leading open competitors in enterprise-relevant workloads (cyber, finance, manufacturing, geospatial imagery) while offering the option of full control. That is a market-segmentation argument as much as a capability one, and it is where open weights have a durable advantage: closed vendors cannot offer on-premises deployment of their best model.
Whether the capability claims hold is a three-week question. Whether the license is permissive enough to matter is the one we would watch most closely.
Bottom line
Mistral Large 4 is a real release, not a rumor: the API is live and the numbers are specific. But it is a preview with vendor-reported results and unpublished weights. Test it on your own tasks now, hold judgment on the leaderboard claims, and read the license when it appears on October 27. We will update this post when the weights ship.
Related reading
- Aleph Alpha Kolibri: a 78B open-weight German-English MoE
- Europe's AI landscape in 2026: EU AI Act, sovereign compute, and the Mistral bet
- What is sovereign AI? Layers, models and trade-offs
- GLM-5.3 FlashX on OpenRouter and Nous Portal
- Qwen audit: 3 billion downloads and censorship findings
- Gemini 4 Argon vs Opus 5.5 vs Grok 4.7 vs GPT-6 Astra
- Official: Mistral Large 4 announcement
Specifications, prices and dates are accurate as of October 6, 2026 and may change before general availability.
