GLM-5.3-Flash
320 billion total, 18 billion active parameters, and MIT license: GLM-5.3-Flash is the open mid-tier variant of Z.AI’s 5.3 family for coding and agentic workloads, with native image and video understanding and one million tokens of context. Hybrid attention keeps the inference footprint moderate despite the overall model size. Reasoning is mandatorily active, weights run locally — the sovereignty risk of the cloud is eliminated.
- Open Weights
- Frontier
- OpenRouter
- Text
- Vision
- Video
- Agentic Orchestrator
- Long Context
- Batch
Sovereign Risk: HIGH Z.AI is headquartered in China, meaning development and potential cloud usage fall under Chinese jurisdiction. The weights are publicly available under the MIT license, enabling local deployment and independent auditing, which significantly reduces sovereignty risk compared to pure cloud operation. For cloud usage, however, Chinese legal and platform risks remain relevant.[web:675][web:677][web:684]
Key metrics
Score · Latency · Cost · Quality
- Total Score Gold
- 80.61
- Routine
- 48.92
- Reasoning
- 31.69
- LLM Judge Avg
- 4.12 / 5
- 100 Coverage
- Avg Task Duration
- 70.05s
- Batch
- Token Rate
- 52.68tok/s
- Output Rate
- P95 Latency
- 212.53s
- Top 5 %
- Total Tokens
- 212600
- Output Volume
- Cost per 1K
- $0.0005
- USD / 1K Requests
- Benchmark Cost
- $0.11
- Total · 212600 tok
Benchmark modules
10 modules · weighted · vs. model median & top performer
Token efficiency & latency
Consumption per module vs. model median