Mistral Large 4: Europe's 1 Trillion-Parameter Model Launches
Mistral AI releases public preview of Mistral Large 4, a 1 trillion-parameter model trained in Europe with strengths in cybersecurity, coding, and multimodal tasks.

What changed
Mistral AI announced a public preview of Mistral Large 4 (officially called “le Chonk”) on October 6, 2026. The model is a 1 trillion-parameter natively multimodal system with 49 billion active parameters. A preview API is available now on Mistral Studio, with model weights to be released by the end of October 2026.
The model was trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in Mistral’s own datacenters in Europe. The public preview is served on that same infrastructure. A significant share of ML4’s training data was multilingual, spanning more than 160 languages, including every official language of the European Union.
Until the weight release, Mistral says it is red-teaming the model with cybersecurity leaders, vetted partners, and state authorities, who will access the same model with reduced moderation and expanded cyber capabilities. The model will be available across multiple regions worldwide, including a European deployment that Mistral operates end-to-end, independently of other digital service providers and under European law.
Why it matters
Mistral Large 4 represents a European alternative to US-based frontier models. For professionals working in regulated or security-sensitive sectors, this distinction carries practical weight.
The model’s strength in cybersecurity is its most significant differentiator. According to Mistral, on the Artificial Analysis Cyber Index, an independent evaluation of how well AI models find and fix security flaws, ML4 ranks among the top five models globally and leads open-weight models developed outside China by a wide margin. On one test asking a model to reproduce a real vulnerability in open-source software and then patch it, ML4 scores 82%, which Mistral describes as the highest of any model. It also solves 93% of the challenges in Cybench, a set of 40 exercises drawn from security competitions.
The contrast is significant. Mistral states that several leading closed models, including Claude Opus 5.5 and GPT-6 Astra, score near zero on the same vulnerability-reproduction test because they refuse to perform the task. For security teams conducting legitimate vulnerability research and incident response, this refusal creates operational friction. Open-weight deployment means organizations can run ML4 on private cloud or on-premise, giving them control over moderation and access policies. Internal testing also showed the model useful for analyzing malware, prioritizing vulnerabilities, and writing detection rules.
Beyond cybersecurity, Mistral reports strong performance in agentic coding and agentic workflows. On DeepSWE v1.1, the company states ML4 scores 61.7%. On AutomationBench, which evaluates business workflows across apps like Gmail, Google Sheets, Slack, and Salesforce, Mistral reports a score of 59.9%, ahead of several competitors. On AA-Briefcase, which evaluates long-horizon knowledge work, it reaches 1,393 Elo. A blind human evaluation with Surge AI on coding quality ranked ML4 Preview second of five models at 3.74 on a 1-5 scale, behind only Claude Opus 5 at 4.22.
For visual grounding tasks, Mistral states the model brings capabilities to industries where perception is critical such as engineering, manufacturing, and earth observation. According to the company, ML4 is one of the most capable models on visual grounding, in some domains surpassing frontier closed models.
The European infrastructure angle matters for organizations in regulated sectors. This addresses data sovereignty concerns for finance, pharma, public sector, and other mission-critical industries, which Mistral says have collaborated on training.
What to test
Before adopting Mistral Large 4, professionals should verify several claims independently.
Cybersecurity benchmarks: The 82% score on vulnerability reproduction and 93% on Cybench deserve scrutiny. Test whether that performance translates to your own threat landscape and incident types. Verify whether the model’s ability to analyze malware and write detection rules meets your security team’s actual workflow. Compare performance against your current tools, not just other LLMs.
Cost and latency of self-deployment: Open weights enable on-premise deployment, but operating a 1 trillion-parameter model requires significant infrastructure. Calculate the actual cost of running this model versus using an API. Measure inference latency under your expected query volume. Assess whether the sovereignty benefit justifies the operational overhead.
Agentic workflow reliability: Test ML4 on your specific business workflows. Mistral’s benchmarks measure isolated tasks. Your workflows may involve error recovery, tool chains, or domain-specific edge cases. Run evaluations similar to what Mistral conducted with Surge AI, using your own annotators and use cases.
Multimodal capabilities for your domain: If you rely on visual analysis, test ML4 on document types, drawings, and imagery specific to your industry. Mistral highlights engineering, manufacturing, and earth observation as strong domains. Performance may vary significantly outside those areas.
Model refinement and updates: Mistral states the model “continues to improve rapidly as we refine it.” Clarify whether future updates will be released as new model versions or whether the base weights will change. Understand your update policy before deploying to production.
Red-teaming results: Ask Mistral for details on what the cybersecurity leaders and state authorities found during red-teaming. Understanding known limitations or failure modes is critical before deploying to sensitive operations.
The conclusion
Mistral Large 4 addresses a real gap: a frontier-capable open-weight model trained and deployed in Europe, with particular strength in cybersecurity tasks where closed models often refuse to engage. For organizations that value AI sovereignty, on-premise deployment, and unrestricted security research, this model merits serious evaluation.
The cybersecurity performance claims are the strongest rationale for adoption. An 82% score on vulnerability reproduction, paired with open weights and European infrastructure, offers capabilities that US-based closed models do not provide. But that capability comes with infrastructure costs and operational responsibility. Running this model is not the same as using an API.
The broader coding and agentic claims are competitive but not clearly superior to existing options. Test those capabilities against your specific workflows before committing resources. The multimodal strengths in engineering, manufacturing, and earth observation should be verified for your own use cases.
Watch for the weight release at the end of October and for independent third-party benchmarks once the model reaches broader use. Also track how the red-teaming process influences final safety guardrails and whether those guardrails are tunable for different deployment contexts.
For security teams, finance, and public sector organizations in Europe or with data sovereignty requirements, this launch signals a frontier alternative to US providers. For others, it opens the option to compare capabilities head-to-head with closed models while maintaining the control that open weights provide.
AI Tool Herald may earn a commission from some links on this site. It never changes what we report or recommend. Affiliate disclosure
Related stories

Claude Dynamic Workflows: 1,000 Parallel Agents Now Available
Anthropic added dynamic workflows to Claude Managed Agents, enabling up to 1,000 AI agents to run in parallel per execution through managed agent infrastructure.

Microsoft releases Decision-1 routing and classification model
Microsoft releases Decision-1, a Qwen3.5-9B decision-scoring model for routing and classification tasks, available in Foundry and OpenRouter.

Nace AI Open-Sources Drex 1.5 Decision Model for Option Scoring
Nace AI open-sources Drex 1.5, a 9B decision model that scores multiple options in one forward pass, ranking first under 10B parameters on Decision Index.