GLM-5.3-Flash

320 billion total, 18 billion active parameters, and MIT license: GLM-5.3-Flash is the open mid-tier variant of Z.AI’s 5.3 family for coding and agentic workloads, with native image and video understanding and one million tokens of context. Hybrid attention keeps the inference footprint moderate despite the overall model size. Reasoning is mandatorily active, weights run locally — the sovereignty risk of the cloud is eliminated.

Zhipu AI Version 5.3-Flash Commercial use permitted MoE 320 B (18 B active) 1000 K Context $0.15 / $0.5 per 1M

  • Open Weights
  • Frontier
  • OpenRouter
  • Text
  • Vision
  • Video
  • Agentic Orchestrator
  • Long Context
  • Batch

Sovereign Risk: HIGH Z.AI is headquartered in China, meaning development and potential cloud usage fall under Chinese jurisdiction. The weights are publicly available under the MIT license, enabling local deployment and independent auditing, which significantly reduces sovereignty risk compared to pure cloud operation. For cloud usage, however, Chinese legal and platform risks remain relevant.[web:675][web:677][web:684]

Key metrics

Score · Latency · Cost · Quality

Total Score Gold
80.61
Routine
48.92
Reasoning
31.69

Rank #1

LLM Judge Avg
4.12
100 Coverage
Avg Task Duration
70.05
Batch
Token Rate
52.68
Output Rate
P95 Latency
212.53
Top 5 %
Total Tokens
212600
Output Volume
Cost per 1K
$0.0005
USD / 1K Requests
Benchmark Cost
$0.11
Total · 212600 tok

Benchmark modules

10 modules · weighted · vs. model median & top performer

GLM-5.3-Flash Best model Ø All models
Code Quality 84.52
CLI Benchmark 91.34
Logical Reasoning 78.43
UX Writing 83.01
Documentation 71.68
Content Transform. 81.32
Cultural Intelligence 81.96
Synthesis Quality 68.33
Tool Execution 90
ToolUse Score 78
Benchmark Cost $0.11

Token efficiency & latency

Consumption per module vs. model median

Token consumption per module

Performance profile