GLM-5.3-Flash (EXL3 High)

GLM-5.3-Flash is a multimodal open-weight model by Z.AI. As a Mixture-of-Experts (MoE) architecture with 320B total and 18B active parameters, it combines efficiency with high performance. It supports a context window of 1 million tokens and processes text, image, and video inputs. The model is particularly optimized for complex agentic and coding tasks and is available under a permissive MIT license.

Zhipu AI Version 5.3-Flash Commercial use permitted MoE 320 B (18 B active) 1000 K Context locally tested

  • Open Weights
  • Server
  • vLLM
  • Text
  • Vision
  • Video
  • Agentic Orchestrator
  • Long Context
  • Interactive

Sovereign Risk: MEDIUM The model was developed by Z.AI, a company headquartered in China (CN). The risk is classified as ‘medium’ rather than ‘high’ because the weights have been released under the very permissive MIT license. This reduces vendor dependency and mitigates some of the risks associated with Chinese jurisdiction. Nevertheless, a residual risk remains regarding training data provenance and potential regulatory influences.

Key metrics

Score · Latency · Cost · Quality

Total Score Silver
77.86
Routine
46.95
Reasoning
30.91

Rank #15

LLM Judge Avg
4
100 Coverage
Avg Task Duration
44.13
Interactive
Token Rate
23.21
Output Rate
P95 Latency
110.13
Top 5 %
Total Tokens
78800
Output Volume
Cost per 1K
$0
USD / 1K Requests
Benchmark Cost
$0
Total · 78800 tok

Benchmark modules

10 modules · weighted · vs. model median & top performer

GLM-5.3-Flash (EXL3 High) Best model Ø All models
Code Quality 80.72
CLI Benchmark 82.66
Logical Reasoning 80
UX Writing 79.45
Documentation 77.35
Content Transform. 79.65
Cultural Intelligence 75.04
Synthesis Quality 63.33
Tool Execution 79.17
ToolUse Score 70.38
Benchmark Cost $0

Token efficiency & latency

Consumption per module vs. model median

Token consumption per module

Performance profile