GLM 4.6

GLM-4.6 is Zhipu AI’s Frontier language model with a focus on Chinese and English language proficiency. The model operates with a context window of 128,000 tokens and supports tool use for agentic workflows. Due to the Chinese manufacturer jurisdiction, a separate privacy assessment is required, and commercial use is subject to restrictions.

Zhipu AI Version 4.6 Commercial use restricted MoE 355 B (32 B active) 128 K Context 06/2025 $0.43 / $1.75 per 1M

  • Restricted Weights
  • Server
  • OpenRouter
  • Text
  • Instruction-Tuned
  • Batch

Sovereign Risk: HIGH Zhipu AI is a Chinese company and subject to China’s National Security Law (NSL), which may allow state access to data. In February 2025, the BSI explicitly warned against the use of Chinese AI cloud services (BSI reference: Warning DeepSeek, 04.02.2025); this risk assessment applies analogously to all Chinese cloud AI providers that process user data on Chinese servers.

Tool-use profile: 6 assets in detail

Comparison of asset performance (P1/P2/Combined) against the fleet average

Asset performance (radar)

Score Breakdown vs. Fleet Average


Tool-use details

Asset performance, reliability, and runtime profile

CrucibleMark evaluates tool use across 6 independent tests. Click a test name for details.

Reliability

  • Tool Call Valid: No
  • Retry: Not required
  • Hallucination: Not detected

Reliability measures how consistently a model actually executes tool calls: Tool Call Valid schema and format accepted, Retry Required successful only after retry, Hallucination Flag fabricated tools or parameters detected. All three green means production-ready.

Runtime profile

Call 1
6.96
First Request
MCP
0.78
Protocol Latency
Synthesis
31.67
Response Generation
Total
157.61
Sum of All Phases
Token
11512
Input + Output
Cost
$0
Cost per Run

The runtime profile shows the latency and cost metrics for the model run: Call 1 First Request, MCP Protocol Latency, Synthesis Response Generation, Total sum of all phases, supplemented by Token input and output and Cost cost per run.

Tool-use review

Created · Instruction-Tuned

Deployment Verdict

Conditional deploy, because GLM 4.6 shows no reliable end-to-end operation despite solid tool execution in individual areas: the combined score is weak, Tool Calls were not consistently valid, and two core tasks fail completely.

Tool Execution Profile

GLM 4.6 demonstrates genuine tool-selection competence, but no robust execution across the full pipeline. In the Web Search & Tool Selection test, it correctly identifies — without prompting — that a search is needed before a direct fetch. This argues against a purely rigid pattern. EU License Research and HTTP Fetch & Extract also start cleanly in terms of tool usage.

The weakness lies in precise operationalization. In the URL Construction test, which derives the correct target URL from model knowledge and then executes it via fetch, it fails completely. The multilingual research task shows the same picture. This is relevant for MCP pipelines: the model often understands which tool is needed in principle, but does not reliably produce the exact calls on which deterministic downstream processing depends. Since no retry was required, this looks less like a formatting problem and more like a comprehension or precision deficit in the execution step.

Synthesis Fidelity

How well does it consolidate tool results? Only to a limited extent. Consolidation remains inconsistent overall, particularly in EU License Research, Tool Failure Handling (404), and multilingual research. On the positive side, HTTP Fetch & Extract stands out: when usable content is available and the task is clearly structured, GLM 4.6 summarizes cleanly. This level of stability is insufficient for longer or multi-step research chains.

Does it stay within the tool result or fall back on training data? In the Honeypot EU License Research test — which checks whether current license restrictions genuinely come from web sources rather than training — no hallucination was detected. This is an important trust signal. The low synthesis score therefore reflects weak consolidation rather than fabricated content.

Error Resilience

In the 404 test, which checks for transparent handling of a failing Tool Call, GLM 4.6 does not hallucinate substitute content. This is production-ready behavior. Response quality is only moderate, but the critical point is met: the model does not fabricate page content on tool failure, keeping the error surface manageable for downstream systems.

Operational Profile

Total 157.61s. Individual calls 6.96s and 31.67s. MCP latency 0.78s. Slow for the overall performance shown. Costs are local, making it infrastructurally inexpensive, but the runtime is not proportionate to the weak end-to-end quality.

Conclusion & Recommendation

Suitable for supervised pipelines with clearly predefined tools, fixed URL schemas, and downstream validation of results. Not suitable for autonomous MCP orchestration, dynamic URL derivation, multilingual web research, or compliance-adjacent workflows where the synthesis itself must be dependable. If you deploy GLM 4.6, use it as an assistive model within tight guardrails — not as a trusted tool agent.

This analysis was generated automatically based on benchmark data. Model used: GPT-5.4 by OpenAI. The raw data and full methodology are documented in the GitHub project.