Second Only To Fable 5: Qwen3.8-Max Finally Shows Its Numbers — And The Claim Gets Complicated

📊 Full opportunity report: Second Only To Fable 5: Qwen3.8-Max Finally Shows Its Numbers — And The Claim Gets Complicated on ThorstenMeyerAI.com — validation score, market gap, and execution plan.

TL;DR

Alibaba announced the full benchmark results for its Qwen3.8-Max model, confirming 2.4 trillion parameters and competitive performance. The open weights will be released next week, marking a significant step in large language model transparency.

Alibaba has officially published the benchmark results for its Qwen3.8-Max model, confirming it as the second-largest publicly disclosed language model after Fable 5. The model features 2.4 trillion parameters and demonstrates strong performance across multiple benchmarks, with open weights expected to ship next week. This marks a key milestone in large language model transparency and accessibility.

On August 3, Alibaba revealed the full specification sheet for Qwen3.8-Max, ending weeks of speculation following its stealth preview and anonymous appearance as ‘kaleb’ on the Code Arena leaderboard. The model is built on the Qwen3.5 architecture, employing a sparse mixture-of-experts design, and is multimodal, supporting text, image, and video inputs.

The core confirmed figure is its 2.4 trillion total parameters, with approximately 95 billion active parameters per query. Benchmark results show it surpasses several competitors, including Claude Opus 4.8 and Claude Fable 5, on terminal and paper benchmarks, and performs well on multimodal and agentic tasks. However, it trails behind GPT-5.6 Sol on certain deep software-engineering benchmarks.

The open weights for the 2.4 trillion-parameter model are scheduled for release next week, though the model’s size means it is primarily a datacenter artifact, not suited for self-hosting. Alibaba also announced a smaller, 27 billion-parameter checkpoint, Qwen3.8-27B, optimized for local deployment and inference, expected to be available soon.

At a glance
updateWhen: announced August 3, 2023; full details…
The developmentAlibaba officially published benchmark data for Qwen3.8-Max, confirming its size and capabilities, after weeks of speculation and stealth previews.
AI DISPATCH · REALITY CHECK Released 3 Aug 2026
Alibaba’s Qwen3.8-Max leaves preview
Second Only to Fable 5?

For fifteen days the claim ran without a benchmark table. Today Alibaba published the table, the active-parameter count, and a weights timeline. The numbers are genuinely strong on the rows Alibaba chose — and twelve to fifteen points behind on the rows it didn’t.

▲ All performance figures: Alibaba’s own harness
2.4T / 95B
Total / active parameters (MoE)
~1M
Context window · 131K max output
Text+Img+Video
Multimodal in · text out
“Next week”
Open weights · licence unpublished
01
Fifteen days from slogan to spec sheet

The claim shipped on a Sunday. The evidence shipped two weeks later. In between, the claim did its work.

17 Jul
Moonshot releases Kimi K3
2.8T parameters; rattles US tech stocks, later suspends new subscriptions under demand.
18 Jul
“kaleb” appears on Code Arena
Anonymous model introduces itself as “Claude” — a distillation artifact — and is identified within a day by a Qwen tokenizer quirk.
19 Jul
WAIC preview: “second only to Fable 5”
No benchmark table, no model card, no licence, no active-parameter count. Paid preview at 10% of standard pricing.
20 Jul
Shares rise as much as 5.4%
The market prices the claim, not the table.
3 Aug
General availability + full benchmark table
95B active confirmed; 2.4T weights and a Qwen3.8-27B checkpoint promised for next week. Licence still unwritten.
02
The table, both halves

“Second only to Fable 5” is true on the rows Alibaba chose and false on the rows it didn’t. Both halves below are from the same release.

Where it leads
Terminal-Bench 2.1 · agentic terminal work
Qwen3.8-Max
86.6
GPT-5.6 Sol
88.8
Fable 5
84.6
OSWorld-Verified · computer use — plus PaperBench 93.0, CAD Bench 91.5
Qwen3.8-Max
86.1
Where it trails — the rows the slogan skips
SWE-bench Pro · deep software engineering
Qwen3.8-Max
67.7
Fable 5
80.0
FrontierSWE · frontier coding agents
Qwen3.8-Max
73.5
Fable 5
88.8
The real jump: one generation of agentic gains vs Qwen3.7-Max
DeepSWE 1.1
21.6 → 56.6
FrontierSWE
40.7 → 73.5
JobBench
31.3 → 53.4
03
Three artifacts, three different facts

“Qwen3.8 is going open-weight” describes three things with very different deployment realities.

Hosted API
Live today

OpenAI- and DashScope-compatible — a base-URL change to A/B against your current backend.

2.4T weights
“Next week” · no licence yet

A multi-node datacenter artifact. At 95B active, no single machine serves it. A flag planted, not a deployment option.

Qwen3.8-27B
Announced · no benchmarks yet

The checkpoint that fits real hardware. Whether the agentic gains survive distillation is the question that decides whether next week matters.

04
Bull and bear

Three Chinese frontier releases in seventeen days, each measured against the same export-controlled model. The contest is real; it is not the same thing as your workload.

Bull
  • The generation jump is real and consistent across a dozen agentic rows, with a stated mechanism: RL-environment scaling.
  • More disclosure than Kimi K3 shipped — full table, active-parameter count, weights timeline.
  • If 2.4T lands under a permissive licence, the ceiling of “open weight” moves permanently.
  • The 27B sibling could become the best local agent model on hardware people already own.
Bear
  • Every number is Alibaba’s harness. Independent testing already tempered Kimi K3’s launch claims substantially.
  • The paying use case still belongs to Fable 5 — twelve to fifteen points on deep software engineering.
  • “Next week” comes from a company that sat on a finished benchmark table for fifteen days.
  • Until the licence text exists, “going open-weight” is a press strategy, not a property of the model.
The claim ran for fifteen days without evidence. Now the evidence exists —
and it says “second only” depends entirely on which row you read.

Implications of Alibaba's Benchmark Disclosure

The release of detailed benchmark data and upcoming open weights positions Alibaba as a major player in large language models, especially in terms of transparency and accessibility. The high performance across multiple benchmarks indicates that Alibaba's approach is competitive with leading models like Fable 5 and GPT-5.6, which could influence the AI landscape by encouraging more open and scalable model deployments. The availability of the smaller 27B checkpoint also signals a focus on practical, local inference use cases, expanding the reach of large language models beyond datacenter environments.

Amazon

large language model AI server

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background on Alibaba's Model Development and Launch Strategy

Over the past two weeks, Alibaba's largest model, initially shrouded in secrecy, was previewed through a stealth release and an anonymous appearance on the Code Arena leaderboard. The company had teased a model "second only to Fable 5" without providing detailed specifications until now. The launch followed a pattern of strategic disclosures, including a preview endpoint and selective benchmarking, designed to generate interest and gauge market response. The full benchmark disclosure today confirms Alibaba's commitment to transparency and competitive positioning in the rapidly evolving large language model space.

"We are committed to open weights and transparency, enabling broader access and innovation in AI development."

— Alibaba spokesperson

Amazon

multimodal AI development kit

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unresolved Questions About Model Licensing and Deployment

It remains unclear what the licensing terms for the open weights will be, as Alibaba has not yet published the license details. The size and infrastructure requirements suggest it will be primarily a datacenter model, limiting self-hosting options. The performance of the smaller 27B checkpoint on real-world tasks and its fidelity compared to the flagship model also require further validation. Additionally, the long-term impact of the agentic improvements and whether they will be maintained in practical applications is still uncertain.

Local LLM Inference Optimization: A Comprehensive Guide to Quantization, Hardware Acceleration, and Efficient Private AI Deployment

Local LLM Inference Optimization: A Comprehensive Guide to Quantization, Hardware Acceleration, and Efficient Private AI Deployment

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for Model Release and Ecosystem Integration

The open weights for Qwen3.8-Max are scheduled to be released next week, enabling researchers and developers to evaluate and deploy the model. Attention will likely turn to how the model performs in real-world applications, especially in local inference scenarios with the 27B checkpoint. Further benchmarking and licensing details are expected soon, potentially shaping how the AI community adopts Alibaba's offerings. Monitoring Alibaba's subsequent updates and ecosystem responses will be crucial for understanding its strategic impact.

Amazon

AI inference server

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

When will Alibaba release the open weights for Qwen3.8-Max?

The open weights are scheduled to be released next week, though exact date details are not yet confirmed.

How does Qwen3.8-Max compare to other large language models?

It outperforms several models on benchmark tests, including Claude Opus 4.8 and Claude Fable 5, but trails behind GPT-5.6 Sol on some deep engineering benchmarks. Its performance is competitive overall, especially in multimodal and agentic tasks.

What are the licensing implications for the open weights?

Licensing details remain unpublished, and it is uncertain whether Alibaba will follow an open-source license like Apache 2.0 or a more restrictive one, which could affect accessibility and use cases.

Will the 27B checkpoint be suitable for local deployment?

Yes, the 27B checkpoint is designed for local inference on high-memory machines and is expected to be available soon, providing a practical option for deployment outside datacenter environments.

What does the model's performance mean for AI development?

The model's strong benchmark results and agentic improvements suggest that Alibaba is making significant strides in scalable, multimodal AI, which could influence future research and commercial deployment strategies.

Source: ThorstenMeyerAI.com

You May Also Like

Today’s Wordle Hints for June 19, 2026

Get the latest confirmed Wordle hints for June 19, 2026, including possible answers and strategies, as provided by The New York Times.

Cadence Design Systems Surges In Global Coverage

Cadence Design Systems experiences a surge in worldwide mentions, indicating expanded global presence and increased market engagement.

Agentic AI And Its Impact On Modern Scientific Computing

OpenAI publishes a webpage titled ‘Scientific computing in the age of agentic AI,’ signaling interest in autonomous AI systems for research, but details remain undisclosed.

Xbox weighs canceling Blade game and shuttering Arkane

Microsoft is reportedly weighing the cancellation of the Blade game and the closure of Arkane Studios, raising questions about its gaming strategy and future projects.