Mistral says it trained ML4 "from scratch" using 3,800 Nvidia Grace Blackwell GPUs in its data centers in Europe and much of its training data was multilingual
First reported by Mistral ·
Open-weight models now achieve state-of-the-art performance in specialized enterprise tasks, directly challenging closed-model dominance.
Mistral AI has announced the public preview of its new large language model, Mistral Large 4 (ML4), also referred to as 'le Chonk'. This 1 trillion-parameter, natively multimodal model features 49 billion active parameters and was trained from scratch using 3,800 NVIDIA Grace Blackwell GPUs in Mistral's European data centers. ML4 reportedly demonstrates performance competitive with leading open-source models and surpasses them in specific enterprise workloads like cybersecurity, finance, and law, even outperforming some frontier closed models in visual grounding. A significant portion of its training data was multilingual, covering over 160 languages, including all official EU languages, emphasizing AI sovereignty. Mistral plans to release the model weights by the end of the month and is currently conducting real-world red-teaming with vetted partners and authorities. The model is designed for self-deployment and offers control over AI capabilities, particularly for cybersecurity operations where provider-level refusals can be a risk.
Mistral's development of ML4 on European infrastructure and with multilingual data underscores a growing trend towards AI sovereignty and distributed AI development, challenging the concentration of power in US-based AI labs. The model's strong performance in cybersecurity, particularly its ability to perform tasks that some closed models refuse, positions it as a critical tool for organizations requiring greater autonomy and less restrictive AI capabilities for sensitive operations. This release signals that open-weight models are rapidly closing the gap in highly specialized, enterprise-critical domains.
The availability of a 1 trillion-parameter, multimodal open-weight model with enterprise-grade performance, including advanced reasoning and visual grounding, lowers the barrier for sophisticated AI deployment. Organizations can now potentially achieve performance parity with leading proprietary solutions without vendor lock-in, especially for tasks requiring nuanced understanding and operational control. The upcoming release of weights by month's end will enable further customization and independent evaluation, potentially accelerating innovation in fields like cybersecurity and scientific research.
AI-written summary. May contain errors.