Alibaba launched Qwen3.8-Max on Aug. 3, unveiling a 2.4 trillion parameter multimodal model built as a sparse Mixture-of-Experts system with about 95 billion active parameters per token, a 1 million-token context window and API pricing of $2 per million input tokens and $6 per million output tokens. The company says the model runs at more than 4,000 tokens per second per GPU on Nvidia's GB300 NVL72 hardware, or roughly 3,000 words a second on a single chip, and scored 86.1 on OSWorld-Verified, ahead of Anthropic's Claude Fable 5 at 85.0 and OpenAI's GPT-5.6 Sol Max, while broader evaluations place it second overall behind Fable 5. Alibaba said open weights for Qwen3.8-Max and the smaller Qwen3.8-27B were expected in the week after launch, even as earlier materials pointed to a midnight Aug. 15 release for Qwen3.8-27B on ModelScope. The 27B model remains positioned for local deployment on consumer hardware, and Reuters previously reported that Alibaba planned a revenue-sharing mechanism for large commercial users of Qwen3.8-Max as the company shifts Qwen3.8 back toward open-weight distribution.