GLM-5.3-Flash (EXL3)
GLM-5.3-Flash is a multimodal Open Weights model by Z.AI. As a Mixture-of-Experts (MoE) architecture with 320B total and 18B active parameters, it combines efficiency with high performance. It supports a context window of 1 million tokens and processes text, image, and video inputs. The model is specifically optimized for complex agentic and coding tasks and is available under a permissive MIT license.
- Open Weights
- Server
- vLLM
- Text
- Vision
- Video
- Agentic Orchestrator
- Long Context
- Unusable
Sovereign Risk: MEDIUM The model was developed by Z.AI, a company headquartered in China (CN). The risk is rated ‘medium’ rather than ‘high’ because the weights have been released under the very permissive MIT license. This reduces vendor dependency and mitigates some of the risks associated with Chinese jurisdiction. Nevertheless, a residual risk remains regarding training data provenance and potential regulatory influences.
Key metrics
Score · Latency · Cost · Quality
- Total Score Silver
- 79
- Routine
- 48.13
- Reasoning
- 30.86
- LLM Judge Avg
- 3.93 / 5
- 100 Coverage
- Avg Task Duration
- 180.75s
- Unusable
- Token Rate
- 26.48tok/s
- Output Rate
- P95 Latency
- 537.46s
- Top 5 %
- Total Tokens
- 260200
- Output Volume
- Cost per 1K
- $0
- USD / 1K Requests
- Benchmark Cost
- $0
- Total · 260200 tok
Benchmark modules
10 modules · weighted · vs. model median & top performer
Token efficiency & latency
Consumption per module vs. model median