Mistral Releases Large 4 with Mixture-of-Experts Architecture
Mistral AI has debuted Mistral Large 4, a 1-trillion-parameter open-source model using a mixture-of-experts architecture, alongside details on its future roadmap.

Mistral AI SAS has launched Mistral Large 4, representing its most capable large language model to date. Initially, the LLM is available in a public preview via the company’s cloud platform, with plans to release the model’s weights later this month.
Architecture and Capabilities
Mistral Large 4 utilizes a mixture of experts architecture comprising 1 trillion parameters. However, it activates only 49 billion parameters at a time, offering greater hardware efficiency than utilizing the entire model for every prompt. The system supports question-answering across more than 160 languages.
On the AA Cyber Index, a set of benchmarks evaluating LLMs on finding and fixing software vulnerabilities, the model secured a top-five position. It achieved an 82% score in patching open-source projects, placing it ahead of competing open-source alternatives. Computer vision is another notable strength; the model scored 1% higher than GPT-6 Astra on Dense200, a benchmark measuring the detection of objects of interest in images.
While trailing behind frontier models like Astra on mainstream coding benchmarks, Mistral Large 4 outperforms several leading open-source models, including Qwen3.8 Max and DeepSeek V4 Pro. It also scored higher than DeepSeek V4 Pro on AutomationBench and AA-Briefcase, benchmarks covering tasks ranging from simple automation to complex knowledge work that would take a human weeks to finish.
Training Infrastructure and Methodology
Mistral trained the model using 3,800 Grace Blackwell chips, each combining two Nvidia Blackwell graphics cards with a single central processing unit. The training process leveraged trial-and-error exercises known as rollouts, where the developing model attempts tasks without human supervision, followed by analysis from a specialized AI model to refine the LLM.
Engineers built a software stack capable of running tens of thousands of rollouts in parallel, drawing from modules like code sandboxes, a web-access search engine, and evaluation tests. This setup generated 33 billion tokens per day, with slightly under half used for the training workflow to refine Large 4. The rollout and training workflows operated asynchronously, preventing delays in one process from slowing down the other. Mistral did not pause the training run upon creating the current version, anticipating that the workflow will produce larger and more capable versions in the coming months. Long-term plans include using Mistral Large 4 to develop a series of models tailored for specific use cases.
What it means for developers
Developers can experiment with Mistral Large 4 and other leading artificial intelligence models affordably through a pay-per-token API via a single key at https://apixoai.online. As Mistral continues to roll out new iterations and specialized variants based on its 1-trillion-parameter architecture, developers gain access to efficient, multilingual tools for complex knowledge work, software vulnerability patching, and computer vision tasks without managing disparate infrastructure.
Source: Mistral launches open-source Mistral Large 4, details AI roadmap — SiliconANGLE AI. Written by the Apixo team from that report.
One key for Claude, GPT, GLM, DeepSeek and more. Pay per token with crypto.
Get your API key

