GLM 5.3 Flash

Released on August 26, 2026

Changes from previous version

Unlike text-only GLM-5.3, Flash is a new multimodal base with far fewer active parameters and about one-tenth the list price, plus native image, video, and file input in the coding loop.

Release Summary

First natively multimodal model in the GLM-5 line: 320B total parameters with 18B active, hybrid sparse plus linear attention, and MIT open weights. Outperforms GLM-5.2 on DeepSWE v1.1 (63.4 vs 46.2) and AutomationBench (48.8 vs 26.2) at Flash API pricing ($0.15/M input, $0.50/M output) with a 1M-token context.

Timeline

August 26, 2026

GLM-5.3-Flash released

Z.ai ships GLM-5.3-Flash to the GLM Coding Plan with 3x quota versus GLM-5.3 and MIT weights on Hugging Face.