Alibaba Qwen Releases Qwen3.8 Max: 2.4-Trillion Parameter Flagship with 1M Multimodal Context

Alibaba's Qwen team has launched Qwen3.8 Max, a 2.4T parameter Mixture-of-Experts model offering 1M token context, native video perception, and deep agent tool orchestration.
| Platform | Route Identifier | Context | Max Output | Rate / Tier | Status |
|---|---|---|---|---|---|
| Alibaba DashScope | qwen3.8-max | 1M tokens (1,000,000) | 131K tokens (131,072) | $2.00 / $6.00 per 1M tokens (cached input rate not published by Alibaba) | active |
| OpenRouter | qwen/qwen3.8-max-0902 | 1M tokens (1,000,000) | 131K tokens (131,072) | $2.00 / $6.00 per 1M tokens ($0.25 cached prompt) | active |
Key Takeaways
- check_circleQwen3.8 Max activates 95B parameters out of a 2.4T MoE backbone, outperforming prior generation open models on math, coding, and multilingual reasoning.
- check_circleSupports a 1,000,000-token context window with high-throughput 131,072-token generation limits.
- check_circleAlibaba lists $2.00/M input and $6.00/M output for the International deployment scope. Its Global scope is cheaper at $1.65/M and $4.951/M, and it does not publish a cache-hit rate.
- check_circleAIMI registered the production route on 3 September 2026 at 23:08:24 SAST.
Frontier Multimodal Mixture-of-Experts
Alibaba's Qwen team has launched Qwen3.8 Max as its flagship inference model. Scaling to 2.4 trillion total parameters with 95 billion actively routed per token, Qwen3.8 Max combines massive knowledge density with low per-token compute overhead.
The model excels across complex code generation, database schema extraction, and video visual question answering.
Frequently Asked Questions
What are the specs of Qwen3.8 Max?
Qwen3.8 Max features 2.4T total MoE parameters (95B active), a 1M token context window, and native text, image, and video input capabilities.
Related Frontier Models & Releases
Explore verified specifications, benchmark results, and route pricing across alternative models in this class.
Qwen3.8 Flash Reaches OpenRouter as Qwen Publishes Flash-Next
OpenRouter has added Qwen3.8 Flash, the production model Qwen says is served through QwenCloud with a one-million-token context and built-in tools. Qwen published the related Flash-Next open-weight preview on the same day.
DeepSeek releases V4.1 Flash with native multimodal vision, 552B MoE and lower API rates
DeepSeek has launched DeepSeek-V4.1-Flash, a 552B-parameter mixture-of-experts model featuring a novel Causal Encoder–Decoder architecture with just 8B active input and 16B active output parameters. The release brings native vision, compresses KV cache storage by up to 8x, and slashes off-peak API pricing to $0.15 per million input tokens.
Moonshot AI Releases Kimi K3: 2.8T Parameter Open-Weight Multimodal Frontier Model
Moonshot AI has released Kimi K3, a 2.8-trillion parameter MoE powerhouse featuring 1M context, native video understanding, and leading mathematical reasoning.