GLM-5.3-Flash (EXL3)

GLM-5.3-Flash is a multimodal Open Weights model by Z.AI. As a Mixture-of-Experts (MoE) architecture with 320B total and 18B active parameters, it combines efficiency with high performance. It supports a context window of 1 million tokens and processes text, image, and video inputs. The model is specifically optimized for complex agentic and coding tasks and is available under a permissive MIT license.

Zhipu AI Version 5.3-Flash Commercial use permitted MoE 320 B (18 B active) 1000 K Context locally tested

  • Open Weights
  • Server
  • vLLM
  • Text
  • Vision
  • Video
  • Agentic Orchestrator
  • Long Context
  • Unusable

Sovereign Risk: MEDIUM The model was developed by Z.AI, a company headquartered in China (CN). The risk is rated ‘medium’ rather than ‘high’ because the weights have been released under the very permissive MIT license. This reduces vendor dependency and mitigates some of the risks associated with Chinese jurisdiction. Nevertheless, a residual risk remains regarding training data provenance and potential regulatory influences.

Key metrics

Score · Latency · Cost · Quality

Total Score Silver
79
Routine
48.13
Reasoning
30.86

Rank #9

LLM Judge Avg
3.93
100 Coverage
Avg Task Duration
180.75
Unusable
Token Rate
26.48
Output Rate
P95 Latency
537.46
Top 5 %
Total Tokens
260200
Output Volume
Cost per 1K
$0
USD / 1K Requests
Benchmark Cost
$0
Total · 260200 tok

Benchmark modules

10 modules · weighted · vs. model median & top performer

GLM-5.3-Flash (EXL3) Best model Ø All models
Code Quality 80.84
CLI Benchmark 82.67
Logical Reasoning 75.64
UX Writing 77.15
Documentation 79.09
Content Transform. 80.07
Cultural Intelligence 77.84
Synthesis Quality 73.33
Tool Execution 90
ToolUse Score 80.5
Benchmark Cost $0

Token efficiency & latency

Consumption per module vs. model median

Token consumption per module

Performance profile