GLM-5.3-Flash (EXL3, TensorFold)

This EXL3 quantization of Z.AI’s GLM-5.3-Flash runs on TensorFold, a young open inference engine with lossless speculative decoding: accelerated draft tokens correspond exactly to the result of serial decoding. The MoE activates around 18 billion of 320 billion parameters per token, processes text, image, and video, and the weights are available under the MIT license.

Zhipu AI Version 5.3-Flash Commercial use permitted MoE 320 B (18 B active) 1000 K Context

  • Open Weights
  • Server
  • TensorFold
  • Text
  • Vision
  • Video
  • Agentic Orchestrator
  • Long Context
  • Batch

Sovereign Risk: MEDIUM The model was developed by Z.AI, a company headquartered in China (CN). The risk is rated ‘medium’ rather than ‘high’ because the weights were released under the very permissive MIT license. This reduces vendor dependency and mitigates some of the risks associated with Chinese jurisdiction. Nevertheless, a residual risk remains regarding training data provenance and potential regulatory influences. The TensorFold variant uses the same EXL3 weights as the vLLM baseline (glm-5_3-flash-exl3); the risk refers to the weights, not the serving engine.

Key metrics

Score · Latency · Cost · Quality

Total Score Gold
80.94
Routine
49.31
Reasoning
31.63

Rank #4

LLM Judge Avg
4.19
100 Coverage
Avg Task Duration
119.21
Batch
Token Rate
55.51
Output Rate
P95 Latency
354.98
Top 5 %
Total Tokens
345700
Output Volume
Cost per 1K
$0
USD / 1K Requests
Benchmark Cost
$0
Total · 345700 tok

Benchmark modules

10 modules · weighted · vs. model median & top performer

GLM-5.3-Flash (EXL3, TensorFold) Best model Ø All models
Code Quality 85.12
CLI Benchmark 94.34
Logical Reasoning 80.04
UX Writing 79.07
Documentation 82.49
Content Transform. 82.9
Cultural Intelligence 75.96
Synthesis Quality 73.33
Tool Execution 90
ToolUse Score 78.67
Benchmark Cost $0

Token efficiency & latency

Consumption per module vs. model median

Token consumption per module

Performance profile