Qwen 3.8: Alibaba's 2.4 Trillion-Parameter Bet on Open Weights
Alibaba announced Qwen 3.8, a 2.4T parameter model promised to be 'second only to Fable 5' — but without a single published benchmark. Here's how it stacks up against Kimi K3, GLM 5.2, and DeepSeek V4 in the increasingly competitive open-weight race.
Alibaba just threw its hat into the open-weight ring with a bold promise. On July 19, the company announced Qwen 3.8, a 2.4 trillion-parameter multimodal AI model that they claim is "second only to Fable 5" — Anthropic's top-tier Claude model. But here's the catch: Alibaba hasn't published a single benchmark to back that claim, and the open weights they promised are still "coming soon."
This is not your typical AI launch. Most model announcements arrive with spreadsheets full of numbers, carefully selected to show the new entrant in the best light. Qwen 3.8 shipped with none of that. What it does have is a massive parameter count, a preview API you can access today, and the tantalizing promise that this will become the largest open-weight model ever released.
The Specs: Big Numbers, Sparse Details
Qwen 3.8 is built on a sparse Mixture-of-Experts architecture, which means only a fraction of those 2.4 trillion parameters activate during any given inference pass. Alibaba hasn't disclosed how many parameters are actually active — a critical detail for understanding real-world performance and compute requirements.
What we do know: the model is multimodal, handling text, images, videos, and documents. It's available now through Alibaba's Token Plan subscription at 10% of standard pricing, with Lite plans starting at $6 for 2,500 weekly credits and Pro tiers at $68 for 40,000 credits. The preview also runs on Qoder (Alibaba's coding platform) and QoderWork.
For context, Qwen 3.7-Max — the predecessor that Alibaba has actually benchmarked — scored 92.4% on GPQA Diamond, 80.4% on SWE-bench Verified, and 69.7% on Terminal-Bench 2.0. If Qwen 3.8 significantly outperforms those numbers, it could be competitive with the best models available. But that's a big if right now.
The Competition: A Crowded Field
Qwen 3.8 is entering a market that got very crowded, very fast. Three Chinese labs have shipped serious open-weight contenders in recent months, and each brings something different to the table.
Kimi K3 from Moonshot AI currently holds the title of largest publicly known model at 2.8 trillion parameters — 400 billion more than Qwen 3.8. Released earlier this year, Kimi K3 packs a 1 million token context window and a novel Stable LatentMoE architecture. In practical coding tests, it has shown strong results: one 269-file repository evaluation had Kimi scoring 83/100 versus Qwen 3.8's 80/100, finishing faster with fewer tokens. On the Artificial Analysis Intelligence Index, Kimi K3 scores 57 compared to Qwen 3.7-Max's 46. The open weights for Kimi K3 are already available, giving it a significant head start in developer adoption.
GLM 5.2 from Zhipu AI takes a different approach. At "only" 744 billion parameters, it's smaller than both Qwen 3.8 and Kimi K3, but what it lacks in scale it makes up for in efficiency and results. Released June 16, 2026 with an MIT license, GLM 5.2 has become the darling of the coding community. It scores 62.1 on SWE-bench Pro (beating GPT-5.5's 58.6), 81.0 on Terminal-Bench 2.1, and sits at #3 on Arena's agent leaderboard — the only open model mixing it up with OpenAI and Anthropic's latest. The CEO of Vercel called its coding output "almost shocking." At roughly one-sixth the cost of frontier models, GLM 5.2 has established itself as the open-weight coding benchmark to beat.
DeepSeek V4 from DeepSeek AI focuses on economics. Released April 24, 2026, it's a 1.6 trillion parameter MoE model with a 1 million token context window. The headline numbers are the prices: $1.74 per million input tokens and $3.48 per million output tokens. Compare that to GPT-5.5 at $5/$30 or Claude Opus 4.7 at $5/$25. DeepSeek V4 delivers near-parity with GPT-5.4 on math and Q&A benchmarks at a fraction of the cost. Benchmark scores include SWE-Bench Pro 58.6, GPQA Diamond 90.5, AIME 2026 96.4, and Humanity's Last Exam at 37.7%.
The Open-Weight Question
Here's where things get interesting. Alibaba has historically kept its Max-tier models closed-source. Qwen 3.8 breaking that pattern would be significant — not just because of the scale, but because of what it signals about the competitive pressure Chinese labs are under.
The open-weight promise matters for a few reasons. Self-hosting means developers can fine-tune, modify, and deploy without relying on Alibaba's cloud infrastructure. It eliminates vendor lock-in and API rate limits. For enterprises with data residency requirements, it's often the only viable path to using frontier-level AI.
If Alibaba follows through, Qwen 3.8 would surpass even Kimi K3's open-source offering in strategic significance, given Alibaba's broader developer ecosystem and distribution reach — including a recent deal to power Apple Intelligence in China.
What We Don't Know
The lack of published benchmarks is genuinely unusual. Even models that underperform on certain tasks typically publish numbers and explain the trade-offs. Alibaba's silence suggests either confidence that independent evaluations will validate their claims, or concern that early benchmarks might not tell the story they want.
The "second only to Fable 5" claim is also impossible to verify right now. Fable 5 is currently export-restricted and banned from public discussion in many contexts, making direct comparisons difficult. The few independent tests that have leaked suggest Qwen 3.8 is competitive but not obviously superior to Kimi K3 or GLM 5.2 on practical coding tasks.
The Verdict: Wait and See
Qwen 3.8 is a bet on the future, not a product you can fully evaluate today. The parameter count is impressive, the multimodal capabilities are welcome, and the open-weight promise — if delivered — would be a genuine milestone. But without benchmarks, without weights, and without transparent pricing for standalone API access, it's difficult to assess where this model actually sits in the competitive landscape.
What we can say with confidence: the Chinese open-weight ecosystem has matured dramatically. Six months ago, the discussion was whether Chinese labs could catch up to American frontier models. Today, the question is which Chinese model — Kimi K3, GLM 5.2, DeepSeek V4, or potentially Qwen 3.8 — best fits your specific use case and budget.
For developers choosing today, GLM 5.2 offers proven coding performance at a known price with weights you can download now. Kimi K3 offers the largest scale and established open-weight availability. DeepSeek V4 offers the best economics for high-volume workloads. Qwen 3.8 offers the promise of something potentially larger and more capable — but that promise comes with a wait.
The open-weight AI race isn't just happening. It's already here, and it's moving faster than most people expected.