Qwen3.8-Max Review: Alibaba's 2.4T Model and Its Very Big Claims
Back to Blog
AI & TechTrendingQwen3.8-MaxAlibaba AI

Qwen3.8-Max Review: Alibaba's 2.4T Model and Its Very Big Claims

Aug 6, 202611 min readClickWise Editorial

On August 3, Alibaba published a benchmark table showing its new model beating the best AI systems on earth. The table might even be right. But it's Alibaba's table — and that distinction is the most important sentence in this review.

Qwen3.8-Max is a 2.4-trillion-parameter multimodal MoE with pricing that undercuts everyone. Here's what's real, what's claimed, and what to do about it.

The review in brief

  • 2.4T total parameters, only 95B active per token (sparse MoE)
  • Text, image, and video input; 1M context; 131K output
  • Claims: beats GPT-5.6 Sol and Claude Fable 5 on agentic benchmarks
  • Pricing: $2 in / $6 out — the aggressive part nobody disputes

What is Qwen3.8-Max?

Qwen3.8-Max is Alibaba's largest model to date, released August 3, 2026: a 2.4-trillion-parameter mixture-of-experts model activating ~95 billion parameters per token, with text/image/video input, a 1M-token context window, and up to 131K output tokens per response. Open weights were scheduled for Hugging Face the week of August 10.

Qwen3.8-Max review — Alibaba's 2.4 trillion parameter multimodal MoE model

Big model, bigger claims, and a price tag that needs no asterisk at all.

The claims — and the asterisk

The headline number: 86.1 on OSWorld-Verified, ahead of GPT-5.6 Sol Max (83.2), Claude Fable 5 (85.0), and Gemini 3.1 Pro (76.2). Add the highest reported PaperBench score (93.0) and claimed wins across dozens of multimodal benchmarks, and Alibaba is asserting the top of the leaderboard.

The asterisk: at launch, every published number came from Alibaba's own table. Independent leaderboards hadn't scored it yet. That doesn't make the numbers false — Qwen's previous releases have generally benchmarked honestly — but the correct posture is "strong vendor claim, verification pending." When independent runs land, this paragraph gets its answer.

The parts that don't need an asterisk

Verifiable on day one
The price$2 input / $6 output per million tokens — roughly a third of comparable Western flagship output pricing
The architecturesparse MoE: 2.4T capacity at 95B-active inference cost, which is how the pricing works
The multimodal demosscreenshot-to-working-app, interactive game generation, 2D floor plan to 3D — reproducible party tricks with real commercial uses
The endurance pitchAlibaba reports 10+ day autonomous coding runs and a 16-day internal engineering project — the agentic angle is the whole strategy
The openness planweights on Hugging Face within a week of launch, continuing Qwen's genuinely open track record

💡 What to actually do this week

Don't rebuild your stack on launch-day claims. Do run your own eval: take 10 real tasks from your workload, run them on Qwen3.8-Max at $2/$6, and compare against your current model. If it matches quality at a third of the cost, the vendor table stops mattering — your table is the one that counts.

Where it fits

The pattern of Qwen's 2026: aim at agents and multimodal work, win on economics, open the weights, repeat. If the OSWorld numbers survive independent testing, Qwen3.8-Max becomes the default recommendation for cost-sensitive agentic workloads. If they land a few points lower, it's still likely the best price-performance of any frontier-class model. That's a comfortable bet either way — which is presumably why Alibaba made it. The context around China's open-weight strategy is in our open-source LLM ranking, and the direct fight with Moonshot in Kimi K3 vs Qwen3.8-Max. Access instructions: how to use Qwen3.8-Max.

Frequently asked questions

What is Qwen3.8-Max?+
Qwen3.8-Max is Alibaba's largest model to date, released August 3, 2026: a 2.4-trillion-parameter mixture-of-experts model activating about 95 billion parameters per token, with text, image, and video input, a 1M-token context window, and up to 131K output tokens.
Is Qwen3.8-Max really better than GPT and Claude?+
Alibaba's own table says yes on several tests — 86.1 on OSWorld-Verified versus 85.0 for Claude Fable 5 and 83.2 for GPT-5.6 Sol Max. But at launch every published number was self-reported; treat it as a strong claim until independent runs land.
How much does Qwen3.8-Max cost?+
$2 per million input tokens and $6 per million output tokens — roughly a third of comparable Western flagship output pricing. Open weights were scheduled for Hugging Face the week of August 10, 2026.
What is Qwen3.8-Max best at?+
Agentic and multimodal work: recreating apps from screenshots, generating interactive games, converting floor plans to 3D, and long autonomous coding runs. Its strongest reported score is 93.0 on PaperBench for research tasks.

Trust the price, verify the benchmarks, run your own eval. In a launch week, that's the whole review methodology worth having.

Want more guides like this?

Join 50K+ readers getting weekly tips on AI, automation & making money online.

Subscribe Free
#Qwen3.8-Max#Alibaba AI#Qwen#Open Source AI#Chinese AI#AI Models 2026

Share this article