🔍 Read the full analysis: Mistral Large 4 Still Trails The AI Frontier’s Leading Edge on ThorstenMeyerAI.com
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
Mistral released Mistral Large 4 as an API preview on October 6, 2026. Artificial Analysis gave it an Intelligence Index score of 38, below leading US models and several Chinese alternatives; its weights have not yet been released. A reviewer also reported hallucinations in personal use, but that was not a controlled comparison.
Mistral AI released Mistral Large 4 in public API preview on October 6, but an Artificial Analysis benchmark snapshot published the next day placed it below leading US models and several Chinese competitors. The result adds evidence that the French company’s latest model has not yet matched the market’s top performers, while its weights remain unavailable ahead of a planned release later in October.
Mistral describes Large 4 as a mixture-of-experts model with one trillion total parameters, of which 49 billion are active. The preview accepts text and images. The company said it trained the model on its own infrastructure in Europe and is continuing to improve it. Those details establish the scale and current status of the release, but they do not independently demonstrate how reliably it performs on demanding work.
In the Artificial Analysis Intelligence Index snapshot available on October 7, 2026, Mistral Large 4 Preview scored 38. That matched OpenAI’s GPT-6 Luna at maximum reasoning effort, sat one point below DeepSeek V4.1 Flash at maximum effort, and trailed Z.ai’s GLM-5.3 at 45 and Moonshot AI’s Kimi K3 at 44. US models scored higher: OpenAI’s GPT-6.1 Sol received 52, Google’s Gemini 4 Argon 53 and Anthropic’s Claude Opus 5.5 58.
The index is an aggregate benchmark, not a direct measure of success on every task. The comparison also includes different reasoning settings and does not hold compute budgets constant. Its scores are index points, not percentages or predicted success rates. The source material says Cohere’s Command A+ scored 13, a counterexample to any claim that every listed competitor outperformed Mistral.
Model Watch · Preview Snapshot · 07 Oct 2026
Mistral Large 4 Still Trails The AI Frontier’s Leading Edge
Mistral’s new API preview pairs trillion-parameter scale with a benchmark score of 38—below several leading US and Chinese models in a dated comparison. The result is a signal to test carefully, not a verdict on every task.
01 / Benchmark snapshot
A visible gap at the top
The October 7 Intelligence Index placed Large 4 Preview at 38. Higher scores appeared for leading US models and several Chinese alternatives listed in the source comparison.
02 / What the release tells us
Scale, access, and open questions
Mistral describes a mixture-of-experts model trained on its own European infrastructure. Those release details establish the product’s scope, but do not by themselves demonstrate reliability on demanding work.
Architecture
One trillion total
Mistral says 49 billion parameters are active. Large total size alone does not show how accurately the model handles complex tasks.
Inputs & capacity
Text, images, long context
The preview accepts text and images. Artificial Analysis reports roughly 512,000 tokens of context; capacity measures what can fit, not reasoning quality.
Regional infrastructure
Trained in Europe
Mistral says the model was trained on company infrastructure in Europe and that improvement work is continuing.
03 / Choosing a model for agentic work
Use the score as a reason to test
Multi-step tasks combine planning, tool use, and carrying information across stages. An early error or unsupported assumption can affect the final result.
“The model was trained on the company’s own infrastructure in Europe and is continuing to improve.”
Mistral Large 4’s benchmark position does not establish parity with the higher-scoring models in this snapshot. Thorsten Meyer said he would not choose the preview for demanding agentic or long tasks when higher-scoring alternatives are available. That is an editorial judgment based on benchmark results and personal experience, not a controlled finding that the model fails at those tasks.
Mistral AI announcement, as described in source material04 / Release timeline
Preview now, weights later
The benchmark concerns a preview build. Mistral scheduled a publicly downloadable weights release for later in October 2026.
API preview
Announced October 6, 2026; access was through an API.
Benchmark snapshot
Artificial Analysis published the cited comparison on October 7.
Weights planned
Mistral scheduled public weights for later in October.
Retest on real work
Compare versions on repeatable tasks, error rates, supervision, and cost.
05 / Evidence limits
What this snapshot cannot settle
The Index is an aggregate comparison with different reasoning settings; compute budgets are not held constant. It does not predict success on every coding, research, or tool-use task. The source reports personal hallucination experiences, not controlled hallucination-rate comparisons. It also gives no underlying cost figures for the DeepSeek comparison. Scores and behavior may change as the preview improves and weights are released.
06 / Key questions
Quick answers
What is Mistral Large 4?
Mistral AI’s latest large model, announced in public API preview on October 6, 2026. It is a mixture-of-experts model with one trillion total and 49 billion active parameters, accepting text and images.
How did it score?
It received 38 Intelligence Index points in the October 7 snapshot. It matched GPT-6 Luna in the settings shown and scored below several listed US and Chinese models.
Can I download the weights?
Not at the time covered by the source. Mistral scheduled the public weights release for later in October 2026.
Does the score prove unreliability?
No. The Index is aggregate, and the reported hallucinations reflect the author’s personal experience, not a controlled model comparison.
Benchmark Gap Shapes Model Choice
For developers and companies choosing a model for multi-step agentic work, the preview’s score is a reason to test carefully rather than assume that a large parameter count or context window translates into dependable execution. Agentic tasks can involve planning, tool use and carrying information through several stages; errors or unsupported assumptions early on can affect the final result.
The source article’s author, Thorsten Meyer, said he would not choose the preview for demanding agentic work or long tasks when higher-scoring alternatives are available. That is an editorial judgment based on benchmark results and personal experience, not a finding that the model necessarily fails at those tasks. Artificial Analysis reports a context capacity of roughly 512,000 tokens, but capacity indicates how much information can be included, not whether the model can reason accurately over it.
The comparison matters to Mistral’s position as a European AI provider as well as to individual buyers. The company says it trained the model on European infrastructure, a relevant development for regional AI capacity. Yet the reported benchmark gap means that this release, as measured in the cited snapshot, does not establish parity with the leading US models or the stronger Chinese models listed.
As an affiliate, we earn on qualifying purchases.
Preview Now, Weights Later
Mistral announced the preview API on October 6. At the time of the October 7 benchmark snapshot, the model was available through an API rather than as publicly downloadable weights. Mistral scheduled a weights release for later in October, so the current evaluation concerns a preview and may not describe the final release.
The benchmark comparison is a dated snapshot, and model scores can change. It lists Claude Opus 5.5, Gemini 4 Argon and GPT-6.1 Sol among the higher-scoring US models, and GLM-5.3 and Kimi K3 among Chinese alternatives. DeepSeek V4.1 Flash scored 39, close to Mistral’s 38; the source article also says Artificial Analysis measured DeepSeek at a much lower cost per task, though the supplied material does not include the underlying cost figures.
Meyer also reported encountering hallucinations while using the preview and said this reduced his confidence in assigning it longer tasks. He described that as personal experience, not a controlled comparative study. The account does not establish that Mistral hallucinates more often than competitors across users or workloads.
“The model was trained on the company’s own infrastructure in Europe and is continuing to improve.”
— Mistral AI, as described in its announcement
large language model training infrastructure
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Preview Results Leave Questions Open
The model’s final performance is not yet established: Mistral says it is still improving the preview, and its weights had not been released at the time covered by the source. The Intelligence Index provides one aggregate comparison, with reasoning settings that are not identical; it does not settle how Large 4 performs on a particular organization’s coding, research or tool-use tasks.
The source material does not provide controlled head-to-head tests of hallucination rates, reliability over long tasks, or the cost figures behind the DeepSeek comparison. Meyer’s report of hallucinations is an account of his own use, not a general rate. It is also unclear from the supplied material how the model’s scores or product behavior may change before or after the planned weights release.
As an affiliate, we earn on qualifying purchases.
Watch for Weights and Retests
The next stated milestone is Mistral’s planned release of publicly downloadable weights later in October 2026. The company says it is continuing to improve the model. Once the weights and any updated version are available, developers can assess the release alongside the preview API and test it on their own workloads.
Further benchmark snapshots may show whether the model’s standing changes, but users weighing it for long or autonomous tasks will need evidence beyond an aggregate index: repeatable tests of task completion, error rates, supervision needs and costs under comparable settings. Until those details are available, the reported score is a useful dated signal, not a final verdict on every use of Mistral Large 4.
As an affiliate, we earn on qualifying purchases.
Key Questions
What is Mistral Large 4?
It is Mistral AI’s latest large model, released in public API preview on October 6, 2026. Mistral describes it as a mixture-of-experts system with one trillion total parameters and 49 billion active parameters, accepting text and images.
How did it score against other models?
Artificial Analysis gave Mistral Large 4 Preview an Intelligence Index score of 38 in a snapshot dated October 7. That was below the listed scores for leading US models and for Z.ai’s GLM-5.3 and Moonshot AI’s Kimi K3, while matching GPT-6 Luna in the settings shown.
Are Mistral Large 4’s weights available to download?
Not at the time covered by the source material. Mistral said the weights were scheduled for release later in October 2026; the preview was then available through an API.
Does the benchmark prove the model is unreliable?
No. The Intelligence Index is an aggregate benchmark and does not predict success or failure on every task. The source article’s report of hallucinations reflects the author’s personal experience, not a controlled comparison of models.
Source: ThorstenMeyerAI.com
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
