Mistral releases Mistral Large 4, dubbed "le Chonk", a 1T-parameter open-weight model for general agentic capabilities, trained on 4,000 Grace Blackwell GPUs
First reported by Thedeepview ·
Open-weight models like Mistral Large 4 offer an alternative to proprietary AI systems, giving users more control and independence from single providers.
French AI lab Mistral has launched a public preview of Mistral Large 4 (ML4), nicknamed "le Chonk," a 1-trillion-parameter, open-weight model designed for general agentic capabilities. Mistral claims this model is comparable to leading closed models while being more efficient than open-weight models three times its size. It was trained on 4,000 NVIDIA Grace Blackwell GPUs at a significantly lower cost than competitors. The model excels in cybersecurity applications, coding, domain-specific knowledge work, and multi-modality, with Mistral highlighting its open-weight nature as a safeguard against model deprecation for enterprise users. ML4 is available via Mistral's API and for on-premise or private cloud deployment, with its weights scheduled for release on October 27th. The model is too large for local desktop or laptop use.
Mistral's release of a 1T-parameter open-weight model directly challenges the dominance of closed-source frontier models, potentially democratizing access to highly capable AI. By offering a model that rivals proprietary systems in performance yet remains open, Mistral aims to provide enterprises with greater control over their AI deployments, particularly in sensitive areas like cybersecurity where vendor lock-in and service discontinuation are significant risks. The use of specialized hardware like 4,000 NVIDIA Grace Blackwell GPUs suggests a trend towards optimized, cost-effective training of massive models, signaling a potential shift in the economics of AI development.
This move impacts companies that rely on AI for critical functions, offering them more flexibility and reducing dependency on a few large providers. The open-weight nature means organizations can host and modify the model themselves, ensuring continuity and security without fear of a service being pulled or altered unexpectedly. The availability of such a powerful, yet open, model could spur innovation in agentic AI applications and cybersecurity defenses, as developers can build upon and inspect its architecture.
AI-written summary. May contain errors.