GLM-5.2

GLM-5.2 is Z.AI’s current flagship with 744 billion total and 40 billion active parameters in a MoE architecture, optimized for complex engineering workflows and long-running coding tasks. The context window spans one million tokens; the weights are available as an Open Weights model under the MIT license.

Zhipu AI Version 5.2 Commercial use permitted MoE 744 B (40 B active) 1000 K Context 12/2025 $1.19 / $3.74 per 1M

  • Open Weights
  • Server
  • OpenRouter
  • Text
  • Instruction-Tuned
  • Agentic Orchestrator
  • Interactive

Sovereign Risk: HIGH Z.AI (formerly Zhipu AI) is a Chinese company and subject to China’s National Security Law (NSL), which can enable state access to data. In February 2025, the BSI explicitly warned against the use of Chinese AI cloud services (BSI reference: Warning DeepSeek, 04.02.2025); this risk assessment applies analogously to all Chinese cloud AI providers that process user data on Chinese servers. With purely local inference using the MIT-licensed weights, the Cloud Act-equivalent risk does not apply.

Tool-use profile: 6 assets in detail

Comparison of asset performance (P1/P2/Combined) against the fleet average

Asset performance (radar)

Score Breakdown vs. Fleet Average


Tool-use details

Asset performance, reliability, and runtime profile

CrucibleMark evaluates tool use across 6 independent tests. Click a test name for details.

Reliability

  • Tool Call Valid: No
  • Retry: Not required
  • Hallucination: Not detected

Reliability measures how consistently a model actually executes tool calls: Tool Call Valid schema and format accepted, Retry Required successful only after retry, Hallucination Flag fabricated tools or parameters detected. All three green means production-ready.

Runtime profile

Call 1
1.68
First Request
MCP
0.92
Protocol Latency
Synthesis
55.52
Response Generation
Total
232.5
Sum of All Phases
Token
9926
Input + Output
Cost
$0
Cost per Run

The runtime profile shows the latency and cost metrics for the model run: Call 1 First Request, MCP Protocol Latency, Synthesis Response Generation, Total sum of all phases, supplemented by Token input and output and Cost cost per run.

Tool-use review

Updated · Instruction-Tuned · Agentic Orchestrator

Deployment Verdict

Conditional deploy, because GLM-5.2 shows no reliable end-to-end behavior for MCP pipelines despite strong tool selection: the combined score is weak, and tool calls were not consistently valid.

Tool Execution Profile

GLM-5.2 demonstrates genuine tool intelligence, but not deterministic execution reliability. In the Web Search & Tool Selection test — which checks whether the model chooses web_search over fetch without an explicit hint — the model correctly identifies the right access path and achieves full tool-selection performance. This argues against mere template behavior.

The counterpoint is the URL Construction & Fetch test, which measures whether the model correctly derives a target URL from its own knowledge and then executes fetch. There it fails completely. This is critical for production pipelines, because many agent flows require not only the correct tool but also precise parameter construction. MCP conformance thus appears situational rather than robust. It understands when a search is needed. It fails when it must construct target addresses itself and form the call exactly.

Synthesis Fidelity

How well does it condense tool results? Only limitedly reliable. P2 performance is weak overall, although individual tasks such as HTTP Fetch & Extract and Web Search & Tool Selection show usable condensation. As soon as a task requires multilingual research or error-prone derivation, synthesis quality drops sharply. For architectures in which the model is expected to convert tool output into decision-ready summaries, this is too inconsistent.

Does it stay within the tool result or fall back on training? In the Honeypot EU License Research test — which checks whether current license restrictions are answered from web sources rather than from pre-trained knowledge — GLM-5.2 stays within the tool framework. No hallucination was detected. This is the most important trust signal in this run and prevents a harsher negative verdict.

Error Resilience

In the 404 test, which checks for transparent behavior on a failed tool call, GLM-5.2 does not hallucinate substitute content. This is a production-relevant positive. Response quality remains weak, however: it does not communicate the failure confidently enough to produce a clean fallback path or a clear operator signal. For production this is acceptable, but only with external error handling in the orchestrator.

Operational Profile

Total 232.50s per run. Of that, a second model call at 55.52s and 0.92s MCP latency. Slow for the quality achieved. Costs are local. Economically justifiable only when local inference is strategically more important than throughput.

Conclusion & Recommendation

Suitable for supervised research and orchestration pipelines in which the model is permitted to select tool types, but URL construction, parameter hardening, and error paths are enforced by the system. Not suitable as an autonomous tool agent with free request construction, or for multilingual retrieval synthesis without strong guardrails. Those deploying GLM-5.2 should operate it as a planning frontend with a tightly guided tool layer — not as a freely acting MCP executor.

This analysis was generated automatically based on benchmark data. Model used: GPT-5.4 by OpenAI. The raw data and full methodology are documented in the GitHub project.