🔍 Read the full analysis: Mistral Large 4 Has Yet To Catch Up With The AI Frontier on ThorstenMeyerAI.com
Get tech for your team delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
Mistral introduced Large 4 as a preview API on October 6, 2026. Artificial Analysis rated it 38 on its Intelligence Index, behind several leading US and Chinese models; the model’s weights are scheduled for release later in October. The benchmark is a dated aggregate comparison, not proof of how it will perform on any particular task.
Mistral launched Mistral Large 4 in public API preview on October 6, but an Artificial Analysis Intelligence Index score of 38 puts it below several leading US and Chinese models in a comparison published the next day. The release is a step for Europe’s AI industry, but the current evidence does not establish Large 4 as a leading choice for demanding agentic work; its weights are not yet publicly downloadable.
Large 4 is Mistral’s largest model to date, described as a mixture-of-experts system with one trillion total parameters and 49 billion active parameters. The preview accepts text and images. Mistral says it trained the model on its own infrastructure in Europe and is continuing to improve it. The company has scheduled a release of the model weights for later in October, but that release had not happened as of October 7. This release also figures into debates about Europe’s AI frontier.
Artificial Analysis’s score of 38 on its Intelligence Index matches OpenAI’s GPT-6 Luna at maximum reasoning effort and is just below DeepSeek V4.1 Flash at maximum effort, which scored 39. In the same snapshot, Anthropic’s Claude Opus 5.5 scored 58, Google’s Gemini 4 Argon scored 53, and OpenAI’s GPT-6.1 Sol scored 52. China’s Z.ai GLM-5.3 and Moonshot AI’s Kimi K3 scored 45 and 44, respectively.
The comparison is dated October 7, 2026, and the listed reasoning settings differ, so the results are not evaluations under identical compute budgets. The figures are benchmark index points, not percentages and do not predict the outcome of every task. Artificial Analysis also reports a context capacity of about 512,000 tokens. That measures how much material can fit in a request, not whether the model can reason accurately across it. The model’s broader significance is also discussed in Mistral’s AI sovereignty push.
AI Frontier · Benchmark Snapshot · October 7, 2026
Mistral Large 4 Has Yet to Catch Up With the AI Frontier
Mistral’s new model is available as a public API preview. In a next-day benchmark snapshot, its score trails several leading US and Chinese models. The evidence is a dated aggregate comparison—not a verdict on every task or on the model’s future.
The gap in the benchmark snapshot
Artificial Analysis published these Intelligence Index scores on October 7, 2026. Higher scores indicate a stronger result on this aggregate index; settings vary across entries.
The score is not a percentage and does not predict performance on every coding, research, or agent workflow.
Scores use the same index, but listed reasoning settings differ—so this is not a comparison under identical compute budgets. Developer locations in the source table do not necessarily indicate where API requests are processed.
What the preview offers
Large 4 is Mistral’s largest model to date. Mistral says it was trained on its own European infrastructure and is still being improved.
Mixture of experts
One trillion total parameters, with 49 billion active parameters. Parameter scale alone does not establish task performance.
Text and images
The public preview API accepts text and image inputs, giving developers a way to evaluate it before the planned weights release.
About 512,000 tokens
Context capacity measures how much material can fit in a request. It does not show whether the model reasons accurately across it.
Why agentic work raises the bar
Multi-step coding, research, and tool use depend on decisions that carry forward. An early error can shape everything that follows.
Plan
Break a goal into ordered steps.
Call tools
Gather information or take action.
Interpret
Check what the tools returned.
Carry decisions forward
Use earlier results to guide later work.
“I would not choose the current preview for demanding agentic work or long tasks when stronger alternatives are available.” Thorsten Meyer · Individual assessment, not a controlled evaluation
Preview today, weights later
As of October 7, developers can access Large 4 through a preview API. Its weights are scheduled for release later in October, but they are not yet publicly downloadable.
A weights release could widen access and enable more evaluation. The date, usage terms, and practical requirements have not been specified in the supplied information.
What the score cannot establish
- Task outcomes: The index does not prove success or failure on any particular workflow.
- Like-for-like ranking: Reasoning settings differ, and the snapshot reflects one date.
- Hallucination rates: The author reports personal observations, not controlled comparative results.
- Cost: No figures are provided to confirm a price comparison.
Questions developers should ask
Benchmark results are a reason to test alternatives on the work that matters to you—not a substitute for deployment evidence.
Does the benchmark prove Large 4 cannot handle agentic tasks?
No. It is an aggregate snapshot, not a universal test. The source author’s caution is a personal judgment rather than a controlled finding.
Can developers download the weights now?
Not as of October 7, 2026. The model is available through its preview API; the weights release is planned for later in October.
What should teams evaluate?
Compare models on representative tasks. Check factual support, instruction-following, tool use, and how much human review is needed.
What remains unknown?
The precise release date, weight usage terms, controlled results for long-running tasks, comparative hallucination rates, and verified cost figures.
The Gap in Frontier Benchmarks
The scores matter because developers choosing a model for multi-step coding, research, or tool use need more than a large parameter count or context window. Agentic systems plan, call tools, interpret results, and carry decisions forward. Errors early in that sequence can compound, while a polished final answer may not reveal that the process went wrong.
The source author, Thorsten Meyer, says he would not choose the current preview for demanding agentic work or long tasks when stronger alternatives are available. That is an individual assessment, not a controlled evaluation of all models or workloads. The index offers a reason to test alternatives, but it does not establish that Large 4 will fail at a particular job or that a higher-scoring model will always perform better in a given deployment.
For Mistral, the launch also carries significance beyond model rankings. Training a large model on European infrastructure is a development in the region’s AI capacity. But geographic origin and benchmark standing are separate questions: the former does not by itself show that the preview can match the strongest competitors on demanding work.
As an affiliate, we earn on qualifying purchases.
Preview Today, Weights Later
The immediate product decision is narrower than the full launch story. As of October 7, developers can access Large 4 through a preview API; they cannot yet download its weights. The planned later-October release could change access and allow further evaluation, but it does not change what is available now.
The comparison also includes Canada’s Cohere Command A+, which scored 13 on the same index. That result means the claim that every competing major lab is ahead would be too broad. The supported point is more specific: Large 4 trails several leading US models and stronger Chinese alternatives in this benchmark snapshot. The source notes that the locations in its table refer to developers, not necessarily where individual API requests are processed.
Meyer reports encountering hallucinations during his own use of the preview and says that this reduced his confidence in assigning it longer tasks. He explicitly frames that as personal experience, not a controlled comparison. His account does not establish how often Large 4 hallucinates relative to competitors, and it does not imply that other models never do so.
“The company says it trained the model on its own infrastructure in Europe and continues to improve it.”
— Mistral
As an affiliate, we earn on qualifying purchases.
What the Score Cannot Establish
The available benchmark is a single dated snapshot. Model versions and scores may change, and the reasoning settings in the comparison are not identical. The supplied information does not show how Large 4 performs across a broad set of independent, workload-specific tests for coding, research, or long-running agent workflows.
It is also unclear how the model will change before or after the planned weight release, or what terms and practical requirements will apply to using those weights. Mistral says it is continuing to improve the preview, but the source provides no detailed release schedule beyond later October. The author’s hallucination observations remain anecdotal; no controlled comparative rate is provided. Cost is mentioned as a factor in the source report, but the supplied material ends before presenting cost figures, so no cost comparison can be confirmed here.
As an affiliate, we earn on qualifying purchases.
The Planned Weight Release
The next stated milestone is Mistral’s planned release of Large 4’s weights later in October 2026. Until that happens, the current assessment concerns an API-only preview. Developers considering it for important work can compare it against alternatives on their own tasks, checking not only final answers but also instruction-following, factual support, tool use, and the amount of human review required.
Further benchmark results and evidence from real deployments could clarify whether Mistral’s ongoing improvements narrow the gap. Until those results are available, the October 7 index should be treated as a dated comparison, not a final verdict on the model or its eventual release.
As an affiliate, we earn on qualifying purchases.
Key Questions
What did Mistral announce?
Mistral introduced Large 4 as a public preview API on October 6, 2026. The model accepts text and images; its weights were scheduled for release later in October.
How did Mistral Large 4 score?
Artificial Analysis gave the preview an Intelligence Index score of 38 in its October 7, 2026 snapshot. The score is a benchmark index result, not a percentage or a guarantee of performance on a particular task.
Does the benchmark prove Large 4 cannot handle agentic tasks?
No. The index is an aggregate measure and does not establish success or failure on every workflow. The source author advises against choosing the current preview for demanding agentic work, but presents that as an individual judgment, not a controlled universal finding.
Can developers download the model weights now?
Not as of October 7, 2026. Mistral’s weights release was scheduled for later in October; developers could access the model through the preview API.
What remains unknown about the release?
The supplied information does not provide a precise weights release date, detailed terms for using the weights, or controlled comparative results for hallucinations and long-running tasks. It also gives no cost figures with which to verify a price comparison.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
