Alibaba Cloud cuts Qwen3.8-Flash prices one day after launch, undercutting GLM-5.3-Flash list rates

Alibaba Cloud reduced Qwen3.8-Flash API prices one day after launch, cutting China rates per million tokens to 0.8 yuan for input and 2.7 yuan for output, down 20% and 10% from the initial 1 yuan and 3 yuan. The new list prices match GLM-5.3-Flash on input and undercut it by 0.1 yuan on output versus that model’s 0.8 yuan and 2.8 yuan sticker rates, though GLM-5.3-Flash still runs a two-week 50% launch discount at 0.4 yuan input and 1.4 yuan output. Separately, B.AI on August 27 made the open-source multimodal Mixture-of-Experts model available with fully free official API access for programming assistance, agent workflows and document analysis. Qwen3.8-Flash has 125 billion total parameters, 51 billion N-gram Embedding and 6 billion active parameters per token, first fuses GDN and QSA hybrid attention, defaults to a 262,144-token context expandable to 1 million via YaRN, and is framed as an open-source preview of the Qwen4 architecture.

The information on this website is generated using AI and we cannot guarantee its accuracy. Please use it as reference information only.