Qwen3.8-Max Review: Alibaba's 2.4T Model and Its Very Big Claims
On August 3, Alibaba published a benchmark table showing its new model beating the best AI systems on earth. The table might even be right. But it's Alibaba's table — and that distinction is the most important sentence in this review.
Qwen3.8-Max is a 2.4-trillion-parameter multimodal MoE with pricing that undercuts everyone. Here's what's real, what's claimed, and what to do about it.
The review in brief
- ✓2.4T total parameters, only 95B active per token (sparse MoE)
- ✓Text, image, and video input; 1M context; 131K output
- ✓Claims: beats GPT-5.6 Sol and Claude Fable 5 on agentic benchmarks
- ✓Pricing: $2 in / $6 out — the aggressive part nobody disputes
What is Qwen3.8-Max?
Qwen3.8-Max is Alibaba's largest model to date, released August 3, 2026: a 2.4-trillion-parameter mixture-of-experts model activating ~95 billion parameters per token, with text/image/video input, a 1M-token context window, and up to 131K output tokens per response. Open weights were scheduled for Hugging Face the week of August 10.

Big model, bigger claims, and a price tag that needs no asterisk at all.
The claims — and the asterisk
The headline number: 86.1 on OSWorld-Verified, ahead of GPT-5.6 Sol Max (83.2), Claude Fable 5 (85.0), and Gemini 3.1 Pro (76.2). Add the highest reported PaperBench score (93.0) and claimed wins across dozens of multimodal benchmarks, and Alibaba is asserting the top of the leaderboard.
The asterisk: at launch, every published number came from Alibaba's own table. Independent leaderboards hadn't scored it yet. That doesn't make the numbers false — Qwen's previous releases have generally benchmarked honestly — but the correct posture is "strong vendor claim, verification pending." When independent runs land, this paragraph gets its answer.
The parts that don't need an asterisk
💡 What to actually do this week
Where it fits
The pattern of Qwen's 2026: aim at agents and multimodal work, win on economics, open the weights, repeat. If the OSWorld numbers survive independent testing, Qwen3.8-Max becomes the default recommendation for cost-sensitive agentic workloads. If they land a few points lower, it's still likely the best price-performance of any frontier-class model. That's a comfortable bet either way — which is presumably why Alibaba made it. The context around China's open-weight strategy is in our open-source LLM ranking, and the direct fight with Moonshot in Kimi K3 vs Qwen3.8-Max. Access instructions: how to use Qwen3.8-Max.
Frequently asked questions
What is Qwen3.8-Max?+
Is Qwen3.8-Max really better than GPT and Claude?+
How much does Qwen3.8-Max cost?+
What is Qwen3.8-Max best at?+
Trust the price, verify the benchmarks, run your own eval. In a launch week, that's the whole review methodology worth having.
Keep Reading
Try Our Free Tools
Want more guides like this?
Join 50K+ readers getting weekly tips on AI, automation & making money online.
Subscribe Free

