Mistral Large 4 Outside The US And China: Strengths, Limits, And Agent Fit
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Mistral Large 4 Outside The US And China: Strengths, Limits, And Agent Fit on ThorstenMeyerAI.com

Before you orderOffer from Amazon

Get the latest gadgets delivered free with Prime

  • Fast, free delivery on millions of items
  • Prime Video, Amazon Music and more included
  • Member-only deals all year
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

Mistral Large 4 scored 38.4 on Artificial Analysis’s Intelligence Index v4.3.2, a sharp improvement over its predecessor and a strong result for a European model. The supplied analysis says it trails current US and Chinese flagships, costs more per benchmark task than two Chinese models that score higher, and may be a poor fit for long-running agents; those conclusions rely on the cited benchmark and attributed hands-on observations.

Mistral released Large 4 as a research public preview, and the model scored 38.4 on Artificial Analysis’s Intelligence Index v4.3.2. The result is a substantial rise from Mistral Large 3’s score of 9, but the supplied analysis says leading US and Chinese models score higher, complicating the claim that Large 4 is ready to compete at the global frontier—especially for agent workloads.

Large 4 is a one-trillion-parameter model with 49 billion active parameters. Mistral describes it as natively multimodal, accepting text and images and producing text, with a 512,000-token context window. It is currently available through Mistral’s API as a research public preview. The source says Mistral has promised to release the weights at the end of October; until then, buyers do not have access to them, and the model’s licence has not been published.

The supplied comparison uses Artificial Analysis Intelligence Index v4.3.2, which includes agentic knowledge work, real-world work tasks, SaaS workflows and coding benchmarks. In that table, Large 4’s 38.4 trails US models scoring as high as 57.6 and several Chinese models, including GLM-5.3 at 44.8 and DeepSeek V4.1 Flash at 39.5. It does exceed the scores listed for Mistral Large 3 and Medium 3.5, at 9 and 14 respectively. The source reports Mistral says reinforcement learning is ongoing, so results could change.

The source lists standard API prices of $1.36 per million input tokens and $4.18 per million output tokens, with cached input at $0.14 per million tokens. It also reports a 50% discount for the first two weeks. In the cited benchmark, Large 4 cost $1.13 per task, compared with $0.25 for GLM-5.3-Flash and $0.27 for DeepSeek V4.1 Flash; both of those models scored above Large 4 in the same comparison. These are benchmark-specific costs, not a guarantee of what every customer workflow will cost.

At a glance
reportWhen: Released yesterday, according to the su…
The developmentMistral released Large 4 as a research public preview, with benchmark results showing a major gain over its prior models but a persistent gap to leading US and Chinese systems.
Mistral Large 4: Not a Frontier Model — Reality Check
AI Dispatch · Reality Check · 7 October 2026

Mistral Large 4: best outside the US and China — and still not a model to run your agents on

The headline is true: France has the most intelligent model outside the US and China. The independent data says the rest: every US and Chinese flagship scores higher, the best by 19 points. It costs 4× more per task than Chinese open models that outscore it, and it’s 2.5× as verbose as the median model.

Artificial Analysis Intelligence Index v4.3.2 — same version, like for like
Claude Opus 5.5 US57.6
Claude Sonnet 5.5 US56.0
Claude Fable 5.1 US53.4
GPT-6 Astra US52.7
Gemini 4 Argon US52.6
GPT-6.1 Sol US51.8
GLM-5.3 CN · open44.8
Kimi K3 CN · open43.6
GLM-5.3-Flash CN · open41.8
DeepSeek V4.1 Flash CN · open39.5
Mistral Large 4 (Preview) FR38.4
GPT-6 Luna US · small model~38
DeepSeek V4 Pro 0813 CN36.0
GLM-5.2 CN33.7
vs US frontier
−19.2 pts

~two-thirds of Opus 5.5. Level with OpenAI’s small model, Luna.

vs China open
8th

Eighth among open models once weights ship — behind seven Chinese ones. Beats GLM-5.2 and V4 Pro, loses to their successors.

vs Canada
n/a

Cohere doesn’t compete at this tier — reported ~14% hallucination at ~9% accuracy, because it declines most questions. A field of one.

The cost problem is worse than the intelligence problem — $ per Index task
Mistral Large 4
$1.13
Index 38.4 · $0.57 launch promo
GLM-5.3-Flash
$0.25
Index 41.8 · 4.5× cheaper
DeepSeek V4.1 Flash
$0.27
Index 39.5 · 4.2× cheaper
Gemini 4 Argon
~$1.99
Index 52.6 · +14 points
Per-token pricing looks competitive ($4.18/M output, well under the $10 median) — but it burns 200M output tokens on the Index vs an 81M median. Cheap tokens × 2.5 as many tokens is not a cheap model.
Why not for agentic or long-running work
The gap compounds
19 pts behind

The Index is now agentic-heavy — Briefcase, GDPval, AutomationBench, Terminal-Bench. Errors multiply across steps: tolerable in chat, fatal over a two-hour run.

AA v4.3.2
Verbosity
200M vs 81M

Output tokens to complete the Index. On an agent, verbosity is cost and latency on every step.

AA
Hallucination is back
observed

Confident false assertions in hands-on use. US frontier has largely moved past this — Gemini 4 Argon: 15%. In fairness Chinese open models are worse (Kimi K3 51%, DeepSeek V4 Pro 94%). In an agent, a fabrication is a wrong premise every later step builds on.

AUTHOR’S TESTING · not an AA figure
✓ What it’s genuinely good at
  • Cyber defence: 50 on the AA Cyber Index; 82% CyberGym-E2E (ahead of Luna’s 78%). Likely top-3 open model on cyber.
  • Documents & images: 19% GDP.pdf (+18 vs Large 3); 100 images per request.
  • Speed: 116 tok/s, 1.46s TTFT — well above median.
  • The jump: Large 3 scored 9 on this Index. 9 → 38 is real progress.
  • Jurisdiction: French parent, EU hosting, weights promised end of October.
▸ Who should actually use it
  • Legally bound buyers (defence, classified, DORA, health data): now the best European option by a wide margin. Wait for the weights, check the licence, pilot on cyber and documents.
  • Everyone else, for agentic or long tasks: don’t. A US frontier model is meaningfully more capable; GLM-5.3-Flash is more capable and 4× cheaper.
  • Note: Preview — Mistral says RL is still running, so scores may move. That changes next month’s decision, not today’s.
The take

Mistral says it has “essentially closed the gap.” It has closed the gap to where the Chinese open-weights field was a few months ago, while that field and the US frontier have both moved on. On every independent measure that matters for agents — intelligence, cost per task, verbosity and factual reliability — Large 4 is not a frontier model. “Most intelligent outside the US and China” is true mainly because almost nobody else outside those two countries is competing. Use it if you have to. Don’t use it because of the headline.

Sources: Artificial Analysis — Mistral Large 4 article & model/provider pages (6 Oct 2026), Index v4.3.2, comparison data; Trending Topics independent-ranking analysis; AA-derived reporting for frontier scores and AA-Omniscience rates (Argon 15%, Kimi K3 51%, DeepSeek V4 Pro 94%); Cohere profile as reported by Suprmind. Mistral Large 4’s AA-Omniscience result isn’t published in text — the hallucination point is the author’s own testing. Preview scores may change. Not investment advice.
thorstenmeyerai.com

The Trade-Off for Agent Work

The release matters because the benchmark cited in the source emphasizes multi-step and agentic tasks, where a model must continue making decisions across a workflow. A weaker result on such work may have more impact than a single incorrect answer: mistakes can shape later steps. That does not establish how Large 4 will perform in every deployed system, but it gives buyers a reason to test it against their own tasks before assigning it long-running work.

Cost adds another procurement question. The supplied analysis says Large 4 used 200 million output tokens across the Index tasks, against a median of 81 million for comparable models. If a workload resembles that benchmark, verbosity could add both expense and latency on top of the listed per-token price. The reported $1.13 per task is below the cited Gemini 4 Argon estimate of about $1.99, but above the cited costs for two Chinese models that scored higher. Those figures are tied to the benchmark methodology and should not be treated as universal service prices.

For organizations seeking a model from outside the US and China, Mistral’s result offers a notable European option. But geographic origin alone does not establish an advantage in capability, price, licensing or reliability. The absence of published weights and licence terms also limits what organizations can yet conclude about deployment flexibility.

Amazon

AI language model API access

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A Sharp Rise, a Wider Gap

Artificial Analysis’s table, as reproduced in the supplied source, places Large 4 below multiple US and Chinese systems. It scores above DeepSeek V4 Pro at 36.0 and GLM-5.2 at 33.7, but below newer entries including GLM-5.3, Kimi K3 and DeepSeek V4.1 Flash. The source characterizes Large 4 as a potential eighth-place model among open-weight systems once its weights are released. That ranking is based on the supplied table and could shift as models or benchmark results change.

The improvement over Mistral’s prior scores is still substantial: Large 3 scored 9, Medium 3.5 scored 14, and Large 4 scored 38.4 on the same Index version, according to the source. That jump is evidence of progress by Mistral, not evidence that it has matched the leading systems. The headline framing of a leading model outside the US and China is also narrower than a claim of global leadership: it compares models by their developers’ locations, not against the full international field.

“Reinforcement learning is still running.”

— Mistral, as reported in the supplied source

Amazon

multimodal AI models with image processing

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Preview Limits and Open Questions

Large 4 is a research public preview, and the supplied source does not provide a complete methodology for independently reproducing its cost and token-use comparisons. It is also unclear how much the scores may change as reinforcement learning continues, or when the promised weights and licence terms will be available beyond the stated end-of-October target.

The source author reports seeing confident false statements during hands-on use. That is an attributed observation, not an Artificial Analysis benchmark finding, and the material provided does not specify the test set, number of trials or error rate. It therefore cannot establish how frequently Large 4 hallucinates in general use. Nor do the benchmark numbers alone show how the model will compare on a buyer’s particular tools, prompts, safeguards or workload.

Amazon

large parameter AI models for developers

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Weights, Licensing and Retests

The next stated milestone is Mistral’s promised release of Large 4’s weights at the end of October. Publication of the licence will be necessary for users to understand what forms of use and redistribution are permitted. Buyers can also watch for updated Artificial Analysis results as Mistral continues reinforcement learning and the public preview develops.

For teams considering the API now, a measured next step is to run a controlled evaluation using representative tasks, recording completion rates, factual errors, output tokens, latency and total cost. That would test whether Large 4’s capability gains and European provider origin outweigh its benchmark gaps and reported token use for a specific workflow. No independent deployment results are provided in the source material.

Amazon

AI model cost per token

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is Mistral Large 4?

It is Mistral’s one-trillion-parameter model with 49 billion active parameters, described as accepting text and images and generating text. The supplied source reports a 512,000-token context window and availability as a research public preview through Mistral’s API.

How did Large 4 score against other models?

Large 4 scored 38.4 on Artificial Analysis Intelligence Index v4.3.2, according to the supplied comparison. That is above Mistral Large 3’s score of 9, but below several listed US and Chinese models.

Is Large 4 available as an open-weight model?

Not yet, according to the source. Mistral has promised weights for the end of October, but the licence is unpublished in the supplied material.

Does the benchmark show that Large 4 is too costly for agents?

No universal conclusion follows from the benchmark alone. The source reports $1.13 per Intelligence Index task and higher task costs than two cited Chinese models that scored above it, but actual costs depend on a customer’s workload, output volume and pricing arrangement.

Source: ThorstenMeyerAI.com

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Why AI Is Essential For Proactive Cybersecurity Strategies In Public And Private Sectors

Google launches Fairwind, an AI-powered program for rapid vulnerability patching, highlighting AI’s importance in proactive cybersecurity for public and private sectors.

Valve Considered a Barebones Steam Machine – So Why Isn’t There One?

Valve considered a minimalistic Steam Machine but has not released one. This report explores why and what it means for gaming hardware.

Belkin Surges In Global Coverage

Belkin’s media mentions have surged, with GDELT reporting 26 mentions in a recent window, indicating rising international attention.

“This is going to be a niche device” – Analysts react to the $1,000+ Steam Machine price reveal

Experts say the new Steam Machine priced over $1,000 will appeal to a niche audience, raising questions about its market impact and future prospects.