Mistral Large 4 Le Chonk Specs vs the Open Field
Le Chonk is live as a preview API at $1.36/$4.18 per million tokens. Here is how Large 4 sits next to DeepSeek, Claude, and GPT-6 Astra.
Open-weight coding and agent tables in Mistral's October 6 launch post already list Large 4 ahead of DeepSeek V4 Pro 0813 and Qwen3.8 Max on the combined Coding Agent Index, ahead of Kimi K3, MiMo-V2.6-Pro, and DeepSeek V4 Pro on AutomationBench, and ahead of DeepSeek V4 Pro on AA-Briefcase, while a hidden-identity Surge AI coding eval still puts Claude Opus 5 first at 4.22 against Large 4 Preview at 3.74, and Mistral says Claude Opus 5.5 and GPT-6 Astra score near zero on a reproduce-then-patch cyber item because they refuse the task. Mistral Large 4, unofficially ML4 and officially le Chonk, joined that board as a public preview API: a 1 trillion-parameter natively multimodal MoE with 49 billion active parameters, trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in the company's own European datacenters, per Mistral's launch post.
1 trillion parameters, 49 billion active
Le Chonk is a hybrid instruct-and-reasoning MoE with multimodal input. Mistral calls it open-weight hybrid, then says the weights drop at the end of the month. Until then you get a preview on Mistral Studio, served from the same European cluster that trained it.
List price on that launch page is $1.36 per million input tokens and $4.18 per million output tokens.
Artificial Analysis's Mistral provider page currently ranks Mistral Large 4 Preview first in that catalog on the Intelligence Index at 38, with a 524k context window. The same table lists the preview license as proprietary, a 116 token-per-second median output speed, 1.46 seconds to first chunk, and an OpenAI-compatible API with function calling and JSON mode. AA also flags the Intelligence Index figure as an estimate, with an independent evaluation forthcoming.
Training data, Mistral says, spanned more than 160 languages, including every official EU language. The same Forge stack the company sells for customer RL is the one used on this run. At the current post-training scale (about 3,000 GPUs), Mistral reports roughly 33 billion tokens per day, of which around 16 billion trainable completion tokens after filtering. The launch post is explicit that the RL run behind the preview is still in flight.
Try the preview API on Mistral Studio if you want the checkpoint this week. Architecture details are promised with the weight drop at the end of the month. Guillaume Lample told WIRED, "Mistral is still in the race of getting the best model." WIRED's paraphrase is that Mistral is reluctant to be pigeonholed into serving only its domestic market. That is a fair commercial line. The preview runs on Mistral's own European infrastructure at $1.36 per million input tokens and $4.18 per million output tokens, not a file you can air-gap.
49.8 percent on the Coding Agent Index
Mistral's own combined Coding Agent Index puts Large 4 at 49.8%, ahead of DeepSeek V4 Pro 0813 and Qwen3.8 Max. The component scores in the launch post are 61.7% on DeepSWE v1.1, 59.4% on SWE-Atlas-QnA, and 28.3% on Terminal-Bench 4.
On AutomationBench (657 business workflows across Gmail, Google Sheets, Slack, and Salesforce) Mistral reports 59.9%, again ahead of Kimi K3, MiMo-V2.6-Pro, and DeepSeek V4 Pro. On AA-Briefcase it reports 1,393 Elo, ahead of DeepSeek V4 Pro. SciCode-Verified is claimed as state of the art among open-weight models.
Those tables are the Chinese open field Large 4 has to beat if the "best outside China" line is going to stick. WIRED repeats Mistral's claim that Le Chonk is the most capable open-weight model developed outside China, and "very, very close" to some proprietary systems. Mistral also says it trained from scratch rather than closing the gap by distillation.
On Mistral's published agent tables, Le Chonk is already past DeepSeek V4 Pro 0813, Qwen3.8 Max, and Kimi K3 on those named scores. That claim is only as strong as a third-party rerun, which Artificial Analysis has not finished.
4.22 versus 3.74 on Surge coding
A blind Surge AI coding eval, identities hidden, 1โ5 scale, five models: Large 4 Preview scored 3.74, second of five, behind Claude Opus 5 at 4.22, with GLM-5.3 at 3.60, Kimi K3 at 3.59, and GLM-5.2 at 3.40.
The sharper split is cyber. On Artificial Analysis's Cyber Index, Mistral says Large 4 ranks among the global top five and leads open-weight models developed outside China by a wide margin. On a reproduce-then-patch test drawn from real open-source flaws, it scores 82%, which Mistral calls the highest of any model. On Cybench (40 CTF-style tasks) it solves 93%. Several leading closed models, including Claude Opus 5.5 and GPT-6 Astra, score near zero on that same reproduce-and-patch item because they refuse the task.
WIRED's surrounding context is the access fight, not the score. After June restrictions on distributing some OpenAI and Anthropic models, Lample's line was that owning the weights matters even for US firms: if you use a closed model, there is no guarantee it will still be there tomorrow. Mistral is also red-teaming a less-moderated, expanded-cyber build with cybersecurity vendors, vetted partners, and state authorities. The Studio preview is not that build.
Opus still wins the blind coding taste test. It also will not do the reproduce-and-patch job Mistral is using as a billboard.
42 percent on Dense 200, Astra at 41
Visual grounding is the closed-model scalp Mistral chose to print. On Dense 200 it reports Large 4 at 42% against GPT-6 Astra at 41%. The launch post also points at ChartQA Pro and GDP.pdf, without filling in extra numbers here. Demos described in the same article include gigapixel satellite scans, engineering drawings, and evidence retrieval from PDFs.
Third-party legal and finance work is where Astra takes a clearer hit in Mistral's writeup. Vals.ai tasks, Mistral says, show Large 4 exceeding GPT-6 Astra on both legal and financial work. On HarveyAI's Legal Agent benchmark it outperforms all open-source models. Finance Agent v2 and Finch (FinWorkBench) are named as the finance tracks. AA-Briefcase is the long-horizon knowledge-work number already cited above.
Safety tables in the same post: 93.3% attack resistance on Lakera's public B3 benchmark, with Mistral seeing no higher competitor score. KORABench at 1.691 out of 2 among open-source models. Indirect prompt-injection benches are described as saturated relative to GLM-5.2, GLM-5.3, Kimi-K2.6, Kimi-K3, and DeepSeek V4 Pro 0813. Average cyber-prompt refusal on JailbreakBench, StrongREJECT, and AgentHarm is higher than all OSS models Mistral listed, which is the intended product stance: capable on Cybench, still refusier than the open pack on jailbreak suites.
GPT-6 Astra, per the launch post, scores near zero on the reproduce-then-patch cyber item because it refuses the task, and trails Large 4 on Dense 200 (41% against 42%). On the verticals Mistral chose to publish, Large 4 is the one posting the higher figure.
End of October for downloadable weights
The preview API is live on Mistral Studio now, with downloadable weights due at the end of the month.
Mistral will release the weights by the end of the month, "along with more details on the architecture, additional benchmarks, and our post-training methodology." Artificial Analysis still marks the Intelligence Index 38 as an estimate. The public checkpoint is moderated. The partner cyber build is not. RL is still running on the expanded European capacity funded by the โฌ3 billion Series D, which Mistral calls the largest equity round ever raised by a European technology company. Headroom is asserted, not measured in a frozen eval suite.
WIRED writes that the model can be used and customized by anyone. The launch post is narrower: Studio preview now, weights at the end of the month, and a European deployment operated end-to-end under European law when you stay on Mistral's cloud.
Mistral Studio is the access path this week. The launch post dates the weight drop to the end of October, with architecture details, extra benchmarks, and post-training methodology to follow. Artificial Analysis still marks the Intelligence Index 38 as an estimate, with an independent evaluation forthcoming.
Keep reading
AI21 Labs Cuts 60% of Staff, Bets on Maestro
AI21 Labs slashes over 60% of staff, drops foundation models, and pivots to its Maestro agent optimization platform after Nebius acquisition talks collapse.
Alibaba Bans Claude Code Over Security Concerns
Alibaba told staff to remove Anthropic's Claude Code by July 10 over security concerns. Here's what triggered the ban and what it signals.
Anthropic Acquires Stainless: What It Means for AI
Anthropic bought Stainless, the SDK generator behind OpenAI and Cloudflare's client libraries. Here's the strategic play for AI agents.