Mistral’s new 1T model aims to leapfrog closed and open rivals
First reported by TechCrunch ·
Mistral's ML4 model will be available as an open-weight option in three weeks, offering a choice beyond current proprietary systems.
French AI lab Mistral AI has launched Mistral Large 4 (ML4), a new large multimodal model with an estimated one trillion parameters. Nicknamed "Le Chonk," the model aims to compete with both closed-source and open-weight AI models from rivals. Currently, ML4 is accessible via a public endpoint, with plans to release its weights in three weeks following safety testing. Mistral VP Science Pierre Stock stated that the company is working with partners and governments to ensure the open-weight model is used for defense rather than malicious attacks. ML4 was trained using approximately 4,000 NVIDIA GPUs, a significantly smaller number compared to competitors. Mistral anticipates ML4 will excel in specific areas like cybersecurity and finance, and potentially chip design, benefiting from its optimized training and multimodal capabilities. The company has received substantial backing from ASML and Samsung.
Mistral's strategy with ML4 signifies a move to bridge the gap between proprietary, closed AI models and fully open-weight alternatives, addressing enterprise and institutional demand for both advanced capabilities and auditability. By training the model on a comparatively smaller GPU cluster, Mistral demonstrates a potential for more efficient large model development, challenging the assumption that massive compute is the sole determinant of performance. This approach could influence future AI infrastructure investments and research directions, particularly in Europe.
The planned release of ML4's weights, after a safety-testing period, positions Mistral as a key player in the open-source AI ecosystem, potentially fostering wider adoption and innovation across industries like cybersecurity, finance, and chip design. The company's focus on specific, high-value use cases, supported by major industry backers, suggests a targeted market strategy aimed at outperforming generalist closed models in critical applications. This development highlights the increasing geopolitical significance of AI development and the emergence of distinct regional approaches to AI governance and deployment.
AI-written summary. May contain errors.