GLM-5.3-Flash
320 billion total, 18 billion active parameters, and MIT license: GLM-5.3-Flash is the open mid-tier variant of Z.AI’s 5.3 family for coding and agentic workloads, with native image and video understanding and one million tokens of context. Hybrid attention keeps the inference footprint moderate despite the overall model size. Reasoning is mandatorily active, weights run locally — the sovereignty risk of the cloud is eliminated.