AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Mistral Large 4 Stands Out Beyond The US And China, With Limits For Agents on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get tech for your team delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

Mistral Large 4 scored 38.4 on the Artificial Analysis Intelligence Index, making it the highest-scoring model from outside the United States and China in the supplied rankings. It remains behind leading US and Chinese models, and the source’s analysis flags high task costs, verbose outputs and observed hallucinations as potential drawbacks for agent use.

Mistral released Large 4 as a research preview, and it scored 38.4 on Artificial Analysis’ Intelligence Index v4.3.2. That makes it the highest-ranked model in the supplied data from outside the United States and China, but the benchmark places it below leading US and Chinese systems, while cost and reliability concerns may limit its appeal for agent-driven work.

The model has 1 trillion parameters, with 49 billion active, and accepts text and images while generating text. Mistral lists a 512,000-token context window. The model is available as a research preview through Mistral’s API; the company says it plans to release the weights at the end of October. Until then, it is proprietary, and the source says its licence has not been published.

Artificial Analysis’ comparison shows Large 4 at 38.4 points, ahead of DeepSeek V4 Pro at 36.0 and GLM-5.2 at 33.7, but behind several Chinese models, including GLM-5.3 at 44.8, Kimi K3 at 43.6 and DeepSeek V4.1 Flash at 39.5. US models in the table score higher still: Claude Opus 5.5 leads at 57.6. The source describes Large 4 as the top model outside the US and China in this set.

Mistral’s earlier Large 3 scored 9 and Medium 3.5 scored 14 on the same version of the index, according to the source. Large 4’s score is a marked increase, though the comparison does not erase its remaining gap with higher-ranked competitors. Mistral says reinforcement learning is ongoing, so benchmark results could change. API pricing is listed at $1.36 per million input tokens and $4.18 per million output tokens, with cached input at $0.14; the source says a 50% discount applies for the first two weeks.

At a glance
reportWhen: Released yesterday, according to the so…
The developmentMistral released Large 4 as a research preview, with independent benchmark data placing it ahead of other models from outside the US and China but below major US and Chinese competitors.
Mistral Large 4: Not a Frontier Model — Reality Check
AI Dispatch · Reality Check · 7 October 2026

Mistral Large 4: best outside the US and China — and still not a model to run your agents on

The headline is true: France has the most intelligent model outside the US and China. The independent data says the rest: every US and Chinese flagship scores higher, the best by 19 points. It costs 4× more per task than Chinese open models that outscore it, and it’s 2.5× as verbose as the median model.

Artificial Analysis Intelligence Index v4.3.2 — same version, like for like
Claude Opus 5.5 US57.6
Claude Sonnet 5.5 US56.0
Claude Fable 5.1 US53.4
GPT-6 Astra US52.7
Gemini 4 Argon US52.6
GPT-6.1 Sol US51.8
GLM-5.3 CN · open44.8
Kimi K3 CN · open43.6
GLM-5.3-Flash CN · open41.8
DeepSeek V4.1 Flash CN · open39.5
Mistral Large 4 (Preview) FR38.4
GPT-6 Luna US · small model~38
DeepSeek V4 Pro 0813 CN36.0
GLM-5.2 CN33.7
vs US frontier
−19.2 pts

~two-thirds of Opus 5.5. Level with OpenAI’s small model, Luna.

vs China open
8th

Eighth among open models once weights ship — behind seven Chinese ones. Beats GLM-5.2 and V4 Pro, loses to their successors.

vs Canada
n/a

Cohere doesn’t compete at this tier — reported ~14% hallucination at ~9% accuracy, because it declines most questions. A field of one.

The cost problem is worse than the intelligence problem — $ per Index task
Mistral Large 4
$1.13
Index 38.4 · $0.57 launch promo
GLM-5.3-Flash
$0.25
Index 41.8 · 4.5× cheaper
DeepSeek V4.1 Flash
$0.27
Index 39.5 · 4.2× cheaper
Gemini 4 Argon
~$1.99
Index 52.6 · +14 points
Per-token pricing looks competitive ($4.18/M output, well under the $10 median) — but it burns 200M output tokens on the Index vs an 81M median. Cheap tokens × 2.5 as many tokens is not a cheap model.
Why not for agentic or long-running work
The gap compounds
19 pts behind

The Index is now agentic-heavy — Briefcase, GDPval, AutomationBench, Terminal-Bench. Errors multiply across steps: tolerable in chat, fatal over a two-hour run.

AA v4.3.2
Verbosity
200M vs 81M

Output tokens to complete the Index. On an agent, verbosity is cost and latency on every step.

AA
Hallucination is back
observed

Confident false assertions in hands-on use. US frontier has largely moved past this — Gemini 4 Argon: 15%. In fairness Chinese open models are worse (Kimi K3 51%, DeepSeek V4 Pro 94%). In an agent, a fabrication is a wrong premise every later step builds on.

AUTHOR’S TESTING · not an AA figure
✓ What it’s genuinely good at
  • Cyber defence: 50 on the AA Cyber Index; 82% CyberGym-E2E (ahead of Luna’s 78%). Likely top-3 open model on cyber.
  • Documents & images: 19% GDP.pdf (+18 vs Large 3); 100 images per request.
  • Speed: 116 tok/s, 1.46s TTFT — well above median.
  • The jump: Large 3 scored 9 on this Index. 9 → 38 is real progress.
  • Jurisdiction: French parent, EU hosting, weights promised end of October.
▸ Who should actually use it
  • Legally bound buyers (defence, classified, DORA, health data): now the best European option by a wide margin. Wait for the weights, check the licence, pilot on cyber and documents.
  • Everyone else, for agentic or long tasks: don’t. A US frontier model is meaningfully more capable; GLM-5.3-Flash is more capable and 4× cheaper.
  • Note: Preview — Mistral says RL is still running, so scores may move. That changes next month’s decision, not today’s.
The take

Mistral says it has “essentially closed the gap.” It has closed the gap to where the Chinese open-weights field was a few months ago, while that field and the US frontier have both moved on. On every independent measure that matters for agents — intelligence, cost per task, verbosity and factual reliability — Large 4 is not a frontier model. “Most intelligent outside the US and China” is true mainly because almost nobody else outside those two countries is competing. Use it if you have to. Don’t use it because of the headline.

Sources: Artificial Analysis — Mistral Large 4 article & model/provider pages (6 Oct 2026), Index v4.3.2, comparison data; Trending Topics independent-ranking analysis; AA-derived reporting for frontier scores and AA-Omniscience rates (Argon 15%, Kimi K3 51%, DeepSeek V4 Pro 94%); Cohere profile as reported by Suprmind. Mistral Large 4’s AA-Omniscience result isn’t published in text — the hallucination point is the author’s own testing. Preview scores may change. Not investment advice.
thorstenmeyerai.com

The Cost of Agent Work

For companies evaluating models to perform multi-step tasks, an overall benchmark score is only part of the decision. The source notes that the Intelligence Index includes agent-oriented tests such as AA-Briefcase, GDPval-AA, AutomationBench and Terminal-Bench 4.0. Its interpretation is that a capability gap can compound over a long sequence of actions: a weak result at one step may shape the steps that follow.

Cost and output length add practical concerns. Artificial Analysis data cited by the source show Large 4 used 200 million output tokens across the index tasks, against a median of 81 million for comparable models. The source calculates a cost of $1.13 per Intelligence Index task for Large 4, compared with $0.25 for GLM-5.3-Flash and $0.27 for DeepSeek V4.1 Flash. Those two models also scored above Large 4 in the cited table. These figures are benchmark-specific; real-world costs will vary with task mix, prompts and usage.

The source author also reports seeing confident hallucinations during hands-on use. That is an attributed observation, not a result established by the Artificial Analysis score. If similar errors occur in agent workflows, an unsupported claim could become a premise for later actions. Buyers will need to test factual reliability, cost and task completion on their own workloads rather than treating the model’s ranking as a procurement verdict.

Amazon

AI language model API

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A European Model’s Benchmark Jump

The supplied source frames Large 4 as a substantial step for Mistral: its score rose from 9 for Large 3 to 38.4 for Large 4 on the same index version. That is evidence of a large improvement on this benchmark, but not proof that the model has reached the performance of the highest-scoring US or Chinese systems.

The headline that France has the most intelligent model outside the US and China is based on the ranking in Artificial Analysis’ index, as quoted in the source. The same data show the competitive field is uneven: multiple US and Chinese labs have models scoring above Large 4. The source’s comparisons with individual rivals and per-task costs are its analysis of those figures, not claims that every deployment will have the same results.

“Reinforcement learning is still running.”

— Mistral, as reported in the supplied source

Amazon

large language model with image input

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Preview Results Still Subject to Change

Large 4 is a research preview, and Mistral says reinforcement learning is ongoing. The source does not provide a later benchmark result or say when the model’s performance will be reassessed. The planned end-of-October weight release is a company timeline reported in the source, not confirmation that the release has occurred.

The licence for those future weights is not published in the supplied material. It is also unclear how Large 4 performs across customers’ production workloads, how often the reported hallucinations occur, or whether model updates will change its verbosity, accuracy or cost per completed task. The benchmark cost figures are tied to the Intelligence Index tasks and should not be read as a universal price for every agent workflow.

Amazon

AI model cost management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Weights and Further Testing Ahead

The next stated milestone is Mistral’s planned release of Large 4’s weights at the end of October. The licence and final terms will matter to organisations deciding whether they can deploy or adapt the model outside Mistral’s API. Mistral’s ongoing reinforcement learning could also lead to updated benchmark results, though no schedule for new results is given.

For now, the available evidence supports a measured conclusion: Large 4 is a strong improvement for Mistral and ranks highest among the models from outside the US and China in the cited table, but it trails several US and Chinese competitors. Organisations considering it for agents should compare its performance, reliability and full task costs against alternatives using their own workloads.

Amazon

AI token usage monitor

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is Mistral Large 4?

Mistral Large 4 is a multimodal model with 1 trillion total parameters, 49 billion active parameters and a 512,000-token context window, according to the supplied source. It accepts text and images and produces text.

How did Large 4 perform on the cited benchmark?

It scored 38.4 on Artificial Analysis Intelligence Index v4.3.2. In the supplied ranking, it is ahead of some earlier Chinese models but behind several newer Chinese models and the listed US models.

Can developers download its weights now?

No. The source says Large 4 is currently a proprietary research preview available through Mistral’s API. Mistral plans to release the weights at the end of October, but the licence is not yet specified in the material.

Why does the source question using it for agents?

The source points to a lower benchmark score than several rivals, higher output-token use and a reported hands-on observation of confident hallucinations. Those concerns warrant task-specific testing; the observation is not an independent benchmark finding.

How much does Large 4 cost?

The listed API prices are $1.36 per million input tokens, $4.18 per million output tokens and $0.14 per million cached input tokens. The source reports a 50% discount for the first two weeks; actual costs depend on usage and task design.

Source: ThorstenMeyerAI.com

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Gig Economy Myth: Flexibility or Just Another Form of Automation?

Falsely promising freedom, the gig economy may be more about automation than flexibility, leaving you to question how much control you truly have.

Communities Across Two Continents Push Back Against Ai’s Growing Data Demands.

Involving communities on two continents reveal growing resistance to AI’s expanding data centers, driven by water scarcity and environmental concerns that demand urgent attention.

Reality Check: Is UBI Truly Unaffordable, or Are We Looking at It Wrong?

Uncover the real costs and funding options behind UBI to see if its perceived un affordability is based on outdated assumptions or something more.

Why Software Does Not Replace Organizations as Easily as Twitter Threads Suggest

Understanding why software can’t fully replace organizations reveals the deep human elements that truly drive success and connection.