Zhipu launches open-source GLM-5.3-Flash at one-tenth GLM-5.2’s price

Zhipu has launched and open-sourced GLM-5.3-Flash, the first native multimodal model in its GLM-5 series. The company says its overall performance exceeds GLM-5.2, while programming and Agent evaluations approach Claude Opus 4.8 at just one-tenth of GLM-5.2’s price. The model uses a new base model, introduces a hybrid architecture combining sparse attention and linear attention to the main GLM series for the first time, and was pretrained on 30T tokens of multimodal data. Before the official release, Zhipu tested the anonymous Ox-Alpha model, known in Chinese communities as 牛来, on OpenCode and OpenRouter to gather broad professional user feedback. Ox-Alpha became the most popular model of the week and set new call-volume records on both platforms, with all request traffic supported by domestically produced chips. An early small-sample DeepSWE test briefly showed an 80% result from 10 questions, but performance fell to about 63% after the sample was expanded, prompting testers to correct the initial claim. The revised result remains in the top tier, though it is less striking than 80%. With the weights open-sourced, developers can deploy the model directly through vLLM, SGLang or KTransformers without using the anonymous model’s API.

The information on this website is generated using AI and we cannot guarantee its accuracy. Please use it as reference information only.