GLM 4.6

GLM-4.6 is Zhipu AI’s Frontier language model with a focus on Chinese and English language proficiency. The model operates with a context window of 128,000 tokens and supports tool use for agentic workflows. Due to the Chinese manufacturer jurisdiction, a separate privacy assessment is required, and commercial use is subject to restrictions.

Zhipu AI Version 4.6 Commercial use restricted Dense 128 K Context 06/2025 $0.39 / $1.9 per 1M

  • Restricted Weights
  • Frontier
  • OR
  • Text
  • Instruction-Tuned
  • Batch

Sovereign Risk: HIGH Zhipu AI is a Chinese company and subject to China’s National Security Law (NSL), which may allow state access to data. In February 2025, the BSI explicitly warned against the use of Chinese AI cloud services (BSI reference: Warning DeepSeek, 04.02.2025); this risk assessment applies analogously to all Chinese cloud AI providers that process user data on Chinese servers.

Tool-use profile: 6 assets in detail

Comparison of asset performance (P1/P2/Combined) against the fleet average

Asset performance (radar)

Score Breakdown vs. Fleet Average


Tool-use details

Asset performance, reliability, and runtime profile

CrucibleMark evaluates tool use across 6 independent tests. Click a test name for details.

Reliability

  • Tool Call Valid: No
  • Retry: Required
  • Hallucination: Not detected

Reliability measures how consistently a model actually executes tool calls: Tool Call Valid schema and format accepted, Retry Required successful only after retry, Hallucination Flag fabricated tools or parameters detected. All three green means production-ready.

Runtime profile

Call 1
16.36
First Request
MCP
0.93
Protocol Latency
Synthesis
33.4
Response Generation
Total
304.16
Sum of All Phases
Token
7889
Input + Output
Cost
$0.0057
Cost per Run

The runtime profile shows the latency and cost metrics for the model run: Call 1 First Request, MCP Protocol Latency, Synthesis Response Generation, Total sum of all phases, supplemented by Token input and output and Cost cost per run.

Tool-use review

· Instruction-Tuned

Deployment Verdict

Conditional deploy, because GLM 4.6 produces valid tool calls and does not hallucinate, but the synthesis quality at Combined 75.92 is only viable when downstream validation controls the compression.

Tool Execution Profile

In tool execution, the model appears competent. The tool call was valid, no hallucination was detected, and in Web Search & Tool Selection, which tests the choice between search and direct retrieval without an explicit hint, it made the right decision confidently. This argues against rigid pattern behavior and in favor of genuine context-aware tool selection. Weaker performance shows in URL Construction & Fetch, which derives the correct target URL from internal knowledge and then retrieves it: usable, but not deterministic enough for pipelines that require exact endpoints without a correction step. The fact that a retry was necessary reads more like a protocol or format issue than a comprehension failure. Execution competence is high, but not clean enough for zero-touch orchestration.

Synthesis Fidelity

How well does it compress tool results? Only adequately. P2 of 63.33 is the actual ceiling of this model. In HTTP Fetch & Extract, which pulls structured facts from real page content, it performs solidly. In Multilingual Search & Synthesis, which tests cross-language research and German-language compression, quality drops noticeably. The model finds the sources but does not compress them consistently precisely enough for reliable decision outputs.

Does it stay within the tool result or fall back on training? In the honeypot EU License Research, which tests whether current license restrictions are answered from web sources rather than training knowledge, the model fundamentally stays in the pipeline’s working mode. P2 60 is not strong, but the trust finding is positive: Content Verification State A, no hallucination. For compliance-adjacent retrieval pipelines, that matters more than linguistic elegance.

Error Resilience

In Tool Failure Handling (404), which tests the response to failed retrievals, GLM 4.6 communicates transparently rather than fabricating page content. P2 80 with no hallucination is acceptable for production. The model does not break trust precisely where many tool models become risky.

Operational Profile

Call 1: 16.36s. Call 2: 33.40s. MCP latency: 0.93s. Total: 304.16s. Cost per run: $0.005716. Verdict: slow, but very cost-efficient relative to the tool execution quality demonstrated.

Conclusion & Recommendation

Suitable for MCP pipelines involving web research, retrieval, error handling, and downstream verification of response compression. Not suitable for fully automated decision pipelines where the final response itself must already constitute the reliable truth layer — particularly in multilingual synthesis or URL-precise retrieval logic. Additionally, production use remains justifiable only in tightly controlled environments due to restricted commercial licensing and elevated provenance risk.

This analysis was generated automatically based on benchmark data. Model used: GPT 4.5 by OpenAI. The raw data and full methodology are documented in the GitHub project.