GLM-4.7

GLM-4.7 is Zhipu AI’s flagship model with 355 billion total and 32 billion active parameters in a MoE architecture, optimized for agentic coding, reasoning, and bilingual tasks in Chinese and English. The model supports a switchable thinking system and is available as an Open Weights variant for local deployment or via cloud interfaces.

Zhipu AI Version 4.7 Commercial use permitted MoE 355 B (32 B active) 128 K Context 12/2025 $0.4 / $1.75 per 1M

  • Restricted Weights
  • Server
  • OpenRouter
  • Text
  • Instruction-Tuned
  • Batch

Sovereign Risk: HIGH Zhipu AI / Z.AI is a Chinese company and subject to China’s National Security Law (NSL), which may allow state access to data. In February 2025, Germany’s BSI explicitly warned against the use of Chinese AI cloud services (BSI reference: Warning DeepSeek, 04.02.2025); this risk assessment applies analogously to all Chinese cloud AI providers that process user data on Chinese servers. With purely local inference, the Cloud Act-equivalent risk does not apply.

Tool-use profile: 6 assets in detail

Comparison of asset performance (P1/P2/Combined) against the fleet average

Asset performance (radar)

Score Breakdown vs. Fleet Average


Tool-use details

Asset performance, reliability, and runtime profile

CrucibleMark evaluates tool use across 6 independent tests. Click a test name for details.

Reliability

  • Tool Call Valid: No
  • Retry: Not required
  • Hallucination: Detected

Reliability measures how consistently a model actually executes tool calls: Tool Call Valid schema and format accepted, Retry Required successful only after retry, Hallucination Flag fabricated tools or parameters detected. All three green means production-ready.

Runtime profile

Call 1
8.29
First Request
MCP
0.97
Protocol Latency
Synthesis
30.87
Response Generation
Total
240.8
Sum of All Phases
Token
19726
Input + Output
Cost
$0
Cost per Run

The runtime profile shows the latency and cost metrics for the model run: Call 1 First Request, MCP Protocol Latency, Synthesis Response Generation, Total sum of all phases, supplemented by Token input and output and Cost cost per run.

Tool-use review

Created · Instruction-Tuned

Deployment Verdict

Deploy conditionally, because GLM-4.7 uses tools sensibly in most cases, but invalid tool calls and weak synthesis of results limit confidence in production MCP pipelines.

Tool Execution Profile

The model demonstrates genuine tool selection rather than mere schema-following. In the Web Search & Tool Selection test — which evaluates the choice between search and direct retrieval without explicit guidance — it makes the correct decision consistently. This speaks to usable tool intelligence in open pipelines. In the URL Construction test, which requires deriving the correct target URL from prior knowledge and then fetching it, the model is only partially precise. This is the more significant finding for production, because here a correct intent fails to produce a deterministic call. The overall finding aligns with this: P1 is solid, but tool calls were not consistently valid. This is not a high-level comprehension problem — it is an execution problem at the interface with the MCP protocol.

Synthesis Fidelity

How well does it consolidate tool results? Only to a limited degree. Synthesis frequently remains at a middling level and loses precision during extraction and merging, particularly in HTTP Fetch & Extract and in the EU License Research task. For workflows where the model is expected to convert researched content into concise, reliable decision briefs, this falls short.

Does it stay within tool results or fall back on training data? In the honeypot EU License Research task — designed to verify whether current license restrictions are genuinely sourced from the web rather than from training knowledge — GLM-4.7 stays on the right side. This is the most important trust signal. At the same time, one hallucination was detected across the full run. This is not merely a quality deficiency but a security risk: once a model presents fabricated facts as tool output, the entire tool infrastructure loses its trust anchor.

Error Resilience

In the 404 test, which measures transparent behavior when a tool call fails, GLM-4.7 responds acceptably. It does not fabricate page content and communicates the failure recognizably. This is production-viable. The execution is not elegant, but it is safer than a model that fills gaps with plausible-sounding substitutes.

Sovereignty Profile

Locally operable and therefore of general interest for sovereign deployments. Performance is, however, 0.89 points below the fleet average of 68.17. The operational advantage is less about raw capability than about the ability to bring a large model into proprietary control zones without cloud dependency.

Conclusion & Recommendation

Suitable for local or sovereign MCP pipelines with human oversight, particularly where tool selection matters more than perfect synthesis. Not suitable for compliance, policy, or extraction pipelines in which every response must be strictly derivable from tool results. Anyone deploying GLM-4.7 should apply strict output validation, source binding, and guardrails for tool call formats upstream.

This analysis was generated automatically based on benchmark data. Model used: GPT-5.4 by OpenAI. The raw data and full methodology are documented in the GitHub project.