Zhipu open-sources GLM-5.3-Flash with 320B parameters and 18B active

Zhipu has formally open-sourced GLM-5.3-Flash after claiming Ox Alpha earlier in the day, placing its model weights on Hugging Face under an MIT license. The model has 320B total parameters but activates only 18B at a time, supports text, images and video natively, and is the first natively multimodal model in the GLM-5 series. Zhipu says its overall performance exceeds GLM-5.2, while its coding and Agent evaluations are close to Claude Opus 4.8 at one-tenth of GLM-5.2’s price. GLM-5.3-Flash uses a new base model, introduces a hybrid of sparse attention and linear attention into the main GLM series for the first time, and was pretrained on 30T tokens of multimodal data. Earlier DeepSWE testing of Ox Alpha reached 80% on 10 questions, but fell to about 63% with a larger sample after the tester corrected the initial claim. The result remains in the first tier, although it is less striking than the original 80% figure. Open weights allow developers to deploy the model directly through frameworks including vLLM, SGLang and KTransformers instead of using an anonymous model API.

本网站上的信息是使用AI生成的,我们无法保证其准确性。 请仅作为参考信息使用。