A few days ago, I was being deliberately provocative
I published “Mistral hosts GLM-5.3, an admission of falling behind or a sovereignty play?”. The headline was meant to provoke. Looking at Mistral Medium 3.5 scoring 14 on the intelligence index while Chinese open-weight models were well ahead, the gap was hard to ignore.
But the article was more nuanced than its headline. Hosting other models, providing European infrastructure and adding specialist tools were also ways to capture value differently. Mistral could adapt without turning every announcement into another round of one-upmanship.
Now Le Chonk is here. To me, it reinforces that interpretation. Distributing models and continuing to develop your own are compatible strategies. Mistral had not necessarily accepted a permanent place at the back.
Large 4 specifications, without mixing up the numbers
The Mistral technical listing provides the following specifications for v26.10.
| Characteristic | Announced specification |
|---|---|
| Architecture | Granular multimodal MoE |
| Total parameters | 1.05 trillion |
| Active parameters | 49 billion |
| Vision encoder | 1.6 billion parameters |
| Context | 1 million tokens |
| API identifier | mistral-large-4 |
| Integration | Function calls, structured outputs, documents, conversations and batch |
| Availability | Public API preview, weights announced for late October |
MoE means mixture of experts. Instead of using every parameter for every token, a routing mechanism selects part of the network. Tokens are the units the model processes and the usual billing unit. This separates the network’s total capacity from the amount of computation activated at each step.
An expert is not necessarily a miniature accounting or legal agent. It is a subnetwork learned during training. And 49 out of 1,050, roughly 4.7%, does not mean that running the model costs 4.7% of running a dense model. Attention, communication between accelerators, memory and routing also contribute.
A large model still needs substantial hosting resources
Having 49 billion active parameters does not turn Large 4 into a 49-billion-parameter model to load into memory. All the weights must remain accessible, potentially across several machines. As a purely arithmetic illustration, 1.05 trillion values at two bytes each represent roughly 2.1 TB. At four bits per value, that is still around 525 GB, before quantization overhead, buffers and the context cache.
Those calculations are not a deployment recommendation from Mistral. They simply explain why I would not put Le Chonk in the category of models you casually install on a small desktop box. That is a different discussion from local AI alternatives on workstations and an RTX 3090.
Vision and context need to support actual work
A vision encoder converts images into representations the model can process. For my workflows, the potential value is reading a document that mixes text and diagrams, understanding a screenshot or locating an interface element. Understanding an image is different from generating one, let alone producing a video.
Likewise, a large context window creates room for documents, conversation history and tool results. It does not guarantee reliable retrieval of every detail, or make filling the entire window fast and inexpensive. I would rather test whether it finds the relevant information inside a real dossier than stop at the headline capacity.
Some architectural details remain to be documented
I will not invent a layer count, the number of experts selected per token, an attention mechanism or a distributed-cache recipe. The references consulted do not establish those details. The announcement points to the weights release for further technical information. The exact licence and self-hosting terms will also need examination then. Announced open weights are not yet files available for download.
The benchmark jump is clear, but real-world testing comes next

This selection shows Large 4 Preview at 38, against Medium 3.5 at 14. These are two models, not a before-and-after measurement of everything Mistral can do. Nor does 38 mean “almost three times as intelligent”.
The Artificial Analysis listing for DeepSeek V4.1 Flash in max mode reports 39. A one-point gap is promising. The models are close on this index, but that does not establish equivalence on every task, language or tool sequence.
One useful clarification is that the screenshot contains DeepSeek V4 Pro 0813 at 36. The V4.1 Flash comparison comes from its separate listing. This chart is a selection, not the complete ranking of all AI models.
In its announcement, Mistral highlights cyber defence, finance, industry and visual grounding. Those are useful signals when choosing my tests, not results I have independently reproduced.
A strategy that may have been more rational than it looked
Mistral says it trained the model from scratch on 3,800 NVIDIA Grace Blackwell GPUs in its European data centres. The announcement also describes reinforcement post-training using tool environments and checks on outcomes. Mistral’s announcement.
Post-training helps shape the behaviour of an already trained model. For an agent, a correct sentence is not enough. It needs to select a tool, interpret the result, recover from an error and continue without losing the objective. Rewarding successful trajectories can therefore matter as much as expanding the network.
I think taking your time can make sense in an ecosystem where methods and tools improve this quickly. But that does not let me conclude that Mistral borrowed a Chinese MoE recipe or calculate how much training money it saved. That would be a story the sources do not establish.
The European landscape is moving too. Aleph Alpha has just introduced Kolibri. More announcements and more options are welcome. I am not going to create a European podium using incompatible tests. What matters to me is having more serious choices.
What this could mean for businesses in France
For me, the value goes beyond a leaderboard position. A capable alternative from a European operator, with deployment options and complementary tools, can influence an architecture decision. That was already the argument in my earlier article about GLM hosted by Mistral.
Choosing Mistral does not automatically make an application GDPR- or AI Act-compliant. Data processing, contracts, access, retention and obligations tied to the use case still need examination. The European framework distinguishes responsibilities according to systems and risks.
In my workflows, the model must not decide its own permissions either. An agent reading documents does not automatically need permission to delete them or send them outside the organisation. System safeguards and human approval still need to be implemented, even with a strong cybersecurity score.
I will compare it with DeepSeek in my Hermes agents
I am very optimistic and intend to try Mistral in place of DeepSeek V4.1 in my workflows, especially my Hermes agents. I want to start within a test environment and expand if the results hold up. At publication, I have no personal Large 4 measurements yet.
I will reuse the same tasks, documents and tools to compare what actually matters to me.
| What I will test | What I will measure |
|---|---|
| Searching my documents | Correct answers, supporting sources and invented information |
| Tool sequences | Valid arguments, failures, retries and permission handling |
| Coding work | Passing tests and corrections required |
| Visual documents | Extraction and understanding of useful elements |
| Long sessions | Elapsed time, context loss and human interventions |
| Cost per completed task | Tokens, cache, retries and the resulting bill |
A model that costs less per token can cost more if it needs three attempts. A slightly more expensive model may be worthwhile if it finishes correctly and saves half an hour of corrections. That is the comparison I want to publish.
At the time of consultation, the pricing page displays $0.68 per million input tokens, $0.07 for cached input and $2.09 for output. Service mode and conditions must be checked when running the tests. I will record the rate actually applied, the model version and settings instead of promising a theoretical monthly bill today.
I will update this article with my impressions, real-world tests and cost estimates. I want to know whether Large 4 can become a practical, lasting alternative for businesses in France. The announcement is encouraging. Actual use comes next. And frankly, well done to Mistral for this return.
Sources and methodology
Written on October 6, 2026. This article distinguishes vendor announcements, a screenshot from Mistral AI’s announcement thread on X, my interpretation and a future testing protocol. It does not claim completed personal Hermes tests with Large 4.
- Mistral, Large 4 announcement dated October 6, 2026. Availability, training, positioning and weights release schedule.
- Mistral, Large 4 v26.10 model listing. Architecture, parameters, vision, context and API features.
- Mistral pricing. Displayed prices at consultation, subject to change.
- Artificial Analysis, DeepSeek V4.1 Flash. Score of 39 for max mode, consulted October 6.
- Mistral AI announcement thread on X. This thread contains the ranking screenshot reproduced in this article.
- Artificial Analysis models and Intelligence Index. Chart attribution. The chart shared in Mistral’s thread shows a selection rather than all models.
- Aleph Alpha, Kolibri announcement. Announcement dated October 5, 2026.
- European Commission, AI Act framework. Responsibilities associated with AI system use.
All references were consulted on October 6, 2026. The cover and LinkedIn illustration were generated with AI under my direction. The cover’s small modules evoke MoE rather than depicting Large 4’s technical design. The original chart is reproduced separately with attribution. Memory estimates are my theoretical calculations, not inference measurements.



