GLM-5.3-Flash (EXL3, TensorFold)
This EXL3 quantization of Z.AI’s GLM-5.3-Flash runs on TensorFold, a young open inference engine with lossless speculative decoding: accelerated draft tokens correspond exactly to the result of serial decoding. The MoE activates around 18 billion of 320 billion parameters per token, processes text, image, and video, and the weights are available under the MIT license.
- Open Weights
- Server
- TensorFold
- Text
- Vision
- Video
- Agentic Orchestrator
- Long Context
- Batch
Sovereign Risk: MEDIUM The model was developed by Z.AI, a company headquartered in China (CN). The risk is rated ‘medium’ rather than ‘high’ because the weights were released under the very permissive MIT license. This reduces vendor dependency and mitigates some of the risks associated with Chinese jurisdiction. Nevertheless, a residual risk remains regarding training data provenance and potential regulatory influences. The TensorFold variant uses the same EXL3 weights as the vLLM baseline (glm-5_3-flash-exl3); the risk refers to the weights, not the serving engine.
Key metrics
Score · Latency · Cost · Quality
- Total Score Gold
- 80.94
- Routine
- 49.31
- Reasoning
- 31.63
- LLM Judge Avg
- 4.19 / 5
- 100 Coverage
- Avg Task Duration
- 119.21s
- Batch
- Token Rate
- 55.51tok/s
- Output Rate
- P95 Latency
- 354.98s
- Top 5 %
- Total Tokens
- 345700
- Output Volume
- Cost per 1K
- $0
- USD / 1K Requests
- Benchmark Cost
- $0
- Total · 345700 tok
Benchmark modules
10 modules · weighted · vs. model median & top performer
Token efficiency & latency
Consumption per module vs. model median