📊 Full opportunity report: Is Qwen3.8-Max The Second Best AI? The Latest Data Sparks Debate on ThorstenMeyerAI.com — validation score, market gap, and execution plan.

TL;DR

Alibaba announced the broad availability of Qwen3.8-Max, revealing benchmark data that positions it as the second-best AI model after GPT-5.6. The development sparks debate over its true capabilities and open-weight status.

Alibaba has officially released the full benchmark data for Qwen3.8-Max, confirming it as the second-best AI model after GPT-5.6 based on their own tests. This follows two weeks of speculation after the model was previewed stealthily and then announced publicly, marking a significant milestone in the AI industry.

On August 3, Alibaba published the complete specifications and benchmark results for Qwen3.8-Max, revealing it has 2.4 trillion parameters with roughly 95 billion active ones per query, built on a sparse mixture-of-experts architecture. The model is multimodal, capable of processing text, images, and videos, and outputs text.

The benchmark results, obtained using Alibaba’s own testing framework, rank Qwen3.8-Max as second only to GPT-5.6 Sol at 88.8 on Terminal-Bench 2.1 and top of the table in PaperBench at 93.0. It also performs strongly in multimodal and agentic tasks, with notable improvements over its predecessor, especially in long-horizon agentic tasks, where it significantly outperforms previous models.

Alibaba confirmed that the open weights for Qwen3.8-Max will be available next week, although these are likely to be large, multi-node artifacts unsuitable for individual hosting. A smaller 27-billion-parameter checkpoint, Qwen3.8-27B, designed for local deployment, will also be released, though benchmark data for this version has not yet been published.

At a glance
updateWhen: announced August 3, 2023, with benchmar…
The developmentAlibaba officially released detailed benchmark data for Qwen3.8-Max, confirming its position as a top-performing AI model and revealing new insights into its architecture and performance.
AI DISPATCH · REALITY CHECK Released 3 Aug 2026
Alibaba’s Qwen3.8-Max leaves preview
Second Only to Fable 5?

For fifteen days the claim ran without a benchmark table. Today Alibaba published the table, the active-parameter count, and a weights timeline. The numbers are genuinely strong on the rows Alibaba chose — and twelve to fifteen points behind on the rows it didn’t.

▲ All performance figures: Alibaba’s own harness
2.4T / 95B
Total / active parameters (MoE)
~1M
Context window · 131K max output
Text+Img+Video
Multimodal in · text out
“Next week”
Open weights · licence unpublished
01
Fifteen days from slogan to spec sheet

The claim shipped on a Sunday. The evidence shipped two weeks later. In between, the claim did its work.

17 Jul
Moonshot releases Kimi K3
2.8T parameters; rattles US tech stocks, later suspends new subscriptions under demand.
18 Jul
“kaleb” appears on Code Arena
Anonymous model introduces itself as “Claude” — a distillation artifact — and is identified within a day by a Qwen tokenizer quirk.
19 Jul
WAIC preview: “second only to Fable 5”
No benchmark table, no model card, no licence, no active-parameter count. Paid preview at 10% of standard pricing.
20 Jul
Shares rise as much as 5.4%
The market prices the claim, not the table.
3 Aug
General availability + full benchmark table
95B active confirmed; 2.4T weights and a Qwen3.8-27B checkpoint promised for next week. Licence still unwritten.
02
The table, both halves

“Second only to Fable 5” is true on the rows Alibaba chose and false on the rows it didn’t. Both halves below are from the same release.

Where it leads
Terminal-Bench 2.1 · agentic terminal work
Qwen3.8-Max
86.6
GPT-5.6 Sol
88.8
Fable 5
84.6
OSWorld-Verified · computer use — plus PaperBench 93.0, CAD Bench 91.5
Qwen3.8-Max
86.1
Where it trails — the rows the slogan skips
SWE-bench Pro · deep software engineering
Qwen3.8-Max
67.7
Fable 5
80.0
FrontierSWE · frontier coding agents
Qwen3.8-Max
73.5
Fable 5
88.8
The real jump: one generation of agentic gains vs Qwen3.7-Max
DeepSWE 1.1
21.6 → 56.6
FrontierSWE
40.7 → 73.5
JobBench
31.3 → 53.4
03
Three artifacts, three different facts

“Qwen3.8 is going open-weight” describes three things with very different deployment realities.

Hosted API
Live today

OpenAI- and DashScope-compatible — a base-URL change to A/B against your current backend.

2.4T weights
“Next week” · no licence yet

A multi-node datacenter artifact. At 95B active, no single machine serves it. A flag planted, not a deployment option.

Qwen3.8-27B
Announced · no benchmarks yet

The checkpoint that fits real hardware. Whether the agentic gains survive distillation is the question that decides whether next week matters.

04
Bull and bear

Three Chinese frontier releases in seventeen days, each measured against the same export-controlled model. The contest is real; it is not the same thing as your workload.

Bull
  • The generation jump is real and consistent across a dozen agentic rows, with a stated mechanism: RL-environment scaling.
  • More disclosure than Kimi K3 shipped — full table, active-parameter count, weights timeline.
  • If 2.4T lands under a permissive licence, the ceiling of “open weight” moves permanently.
  • The 27B sibling could become the best local agent model on hardware people already own.
Bear
  • Every number is Alibaba’s harness. Independent testing already tempered Kimi K3’s launch claims substantially.
  • The paying use case still belongs to Fable 5 — twelve to fifteen points on deep software engineering.
  • “Next week” comes from a company that sat on a finished benchmark table for fifteen days.
  • Until the licence text exists, “going open-weight” is a press strategy, not a property of the model.
The claim ran for fifteen days without evidence. Now the evidence exists —
and it says “second only” depends entirely on which row you read.

Implications of Alibaba's Benchmark Data for AI Leadership

The release of detailed benchmark data and open weights positions Alibaba as a serious contender in the AI race, challenging established leaders like OpenAI and Anthropic. The performance gains, especially in agentic and multimodal tasks, could influence future AI development and deployment strategies. However, the selective nature of the benchmarks and the lack of full transparency about licensing and deployment limits remain points of debate among industry experts.

Amazon

AI model benchmark software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background on Alibaba’s AI Model Releases and Industry Position

Over the past two weeks, Alibaba’s AI efforts have been shrouded in secrecy, with models like Kimi K3 and the anonymous ‘kaleb’ appearing briefly before the full specs of Qwen3.8-Max were unveiled. The company’s strategy involved a staged rollout, culminating in the official announcement on August 3, after previewing the model during the July World AI Conference in Shanghai. Prior to this, Alibaba’s models had been less transparent, with limited details about size, architecture, or benchmark performance.

The industry has been closely watching Alibaba’s progress, especially after the brief but impactful release of Kimi K3, which briefly rattled US tech stocks. The recent benchmark disclosures mark a shift toward greater transparency, though some details—like licensing terms—remain unpublished. The model’s architecture, based on Qwen3.5, and its multimodal capabilities reflect ongoing trends toward more versatile and powerful AI systems.

"We are committed to transparency and will release open weights next week, enabling broader access and fostering innovation."

— Alibaba spokesperson

Visual Studio Code Guide for Beginners: Master Programming, Debugging, GitHub Integration, Extensions, AI Tools, Terminal, Deployment, and Professional Development Workflows from Scratch

Visual Studio Code Guide for Beginners: Master Programming, Debugging, GitHub Integration, Extensions, AI Tools, Terminal, Deployment, and Professional Development Workflows from Scratch

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unconfirmed Aspects of Qwen3.8-Max’s Capabilities and Release

It remains unclear how the open weights will be licensed and whether they will be fully open-source or have restrictions. The benchmark results, while impressive, are based on Alibaba’s own testing framework, and independent validation is pending. The performance of the smaller Qwen3.8-27B model, which is more relevant for local deployment, has not yet been benchmarked or released.

Additionally, the long-term viability of the agentic improvements and whether they will survive compression or deployment in real-world applications are still under assessment. The true impact of the model’s multimodal capabilities across diverse tasks is also yet to be confirmed outside Alibaba’s testing environment.

Amazon

AI multimodal processing hardware

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for Benchmark Validation and Open-Weight Availability

Alibaba plans to release the open weights for Qwen3.8-Max next week, which will enable independent testing and deployment. Industry experts and competitors will scrutinize these weights and benchmark results to verify claims. Additionally, benchmark results for the Qwen3.8-27B model are expected soon, providing insight into local deployment potential.

Further analysis from third-party researchers and AI labs will determine whether Qwen3.8-Max’s performance translates into practical advantages and how it compares with other models on a broader set of tasks. The ongoing debate over transparency, licensing, and openness will shape the model’s adoption and industry impact.

Amazon

AI model training and hosting solutions

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What makes Qwen3.8-Max stand out among AI models?

Qwen3.8-Max’s key features include its 2.4 trillion parameters, multimodal capabilities, and strong performance in agentic and multimodal benchmarks, ranking it just behind GPT-5.6 according to Alibaba’s tests.

When will the open weights for Qwen3.8-Max be available?

Alibaba has announced that the open weights will be released next week, enabling independent testing and deployment.

How reliable are Alibaba’s benchmark results?

The benchmarks are based on Alibaba’s own testing framework, and independent validation is pending. The results are promising but should be interpreted with caution until verified externally.

What are the implications for AI competition?

The detailed benchmark data and open-weight plans position Alibaba as a serious contender in AI, potentially challenging existing leaders and influencing future industry standards.

What are the limitations or uncertainties about Qwen3.8-Max?

Uncertainties include licensing details, the performance of the smaller 27B model, and whether the agentic improvements will hold in real-world applications outside Alibaba’s testing environment.

Source: ThorstenMeyerAI.com

You May Also Like

The bank account in the chat. How personal finance became an agentic on-ramp.

OpenAI introduced live bank account integration in ChatGPT for Pro users, marking a major step toward agentic consumer finance tools and transforming fintech intermediation.

7 Best Internal Solid State Drives for Prime Day Deals in 2026

A 2026 Prime Day SSD watchlist ranks SK hynix Gold P31 2TB first while warning buyers to check capacity, form factor and real discounts.

How The EU Court Recognized VPNs As Legal Technical Solutions For Privacy

The EU Court has ruled that VPNs are lawful technical solutions for privacy, marking a significant legal acknowledgment of their role in online privacy.

The Skills Marketplace Nobody Is Building Yet

A comprehensive analysis of the emerging skills standard in AI, its current limitations, and the opportunity for a dedicated marketplace platform.