DeepSeek V4 Pro

DeepSeek V4 Pro is the flagship of the V4 line, designed for reasoning, coding, and agentic workflows. The hybrid attention MoE architecture combines 1.6 trillion total parameters with 49 billion active parameters per token and a context window of one million tokens. The model is available as an Open Weights model under the MIT license, though Chinese jurisdiction in cloud deployments requires a separate privacy assessment.

DeepSeek Version 4 Commercial use permitted MoE 1600 B (49 B active) 1000 K Context 05/2025 $0.87 / $1.74 per 1M

  • Open Weights
  • Frontier
  • OpenRouter
  • Text
  • Agentic Orchestrator
  • Long Context
  • Batch

Sovereign Risk: HIGH DeepSeek is a Chinese company subject to China’s National Security Law (NSL), which may allow state access to data and models. On 04.02.2025, Germany’s BSI explicitly warned against using the DeepSeek cloud service: user data is stored on Chinese servers; use for official or sensitive data is not recommended. This warning applies without restriction to cloud API deployments.

Tool-use profile: 6 assets in detail

Comparison of asset performance (P1/P2/Combined) against the fleet average

Asset performance (radar)

Score Breakdown vs. Fleet Average


Tool-use details

Asset performance, reliability, and runtime profile

CrucibleMark evaluates tool use across 6 independent tests. Click a test name for details.

Reliability

  • Tool Call Valid: No
  • Retry: Not required
  • Hallucination: Not detected

Reliability measures how consistently a model actually executes tool calls: Tool Call Valid schema and format accepted, Retry Required successful only after retry, Hallucination Flag fabricated tools or parameters detected. All three green means production-ready.

Runtime profile

Call 1
4.59
First Request
MCP
1.34
Protocol Latency
Synthesis
27.47
Response Generation
Total
200.37
Sum of All Phases
Token
17179
Input + Output
Cost
$0
Cost per Run

The runtime profile shows the latency and cost metrics for the model run: Call 1 First Request, MCP Protocol Latency, Synthesis Response Generation, Total sum of all phases, supplemented by Token input and output and Cost cost per run.

Tool-use review

Updated · Agentic Orchestrator · Long Context

Deployment Verdict

Conditional deploy, because tool use is strong, but tool calls are not consistently valid and synthesis quality is only moderately reliable for production-grade knowledge pipelines.

Tool Execution Profile

DeepSeek V4 Pro demonstrates genuine tool intelligence. In the Web Search & Tool Selection test — which checks whether the model selects web_search over fetch without being prompted — it chooses the correct tool confidently and achieves full execution reliability. This argues against a rigid retrieval pattern and in favor of usable planning in dynamic MCP pipelines.

It performs weaker on the URL Construction test, which requires precise derivation of a target URL from internal knowledge. There, execution is usable but not deterministic enough for systems that depend on exactly reproducible fetch paths. The central operational flaw remains that the tool call was not valid overall. This is not a collapse of tool capability, but an integration risk at the protocol level. On the positive side, no retry was required — which points less toward a pure formatting issue and more toward inconsistent call precision in individual paths.

Synthesis Fidelity

How well does it consolidate tool results? Only adequately. The P2 performance shows that DeepSeek V4 Pro usually aggregates results correctly, but does not prioritize details stably enough. This is also visible in HTTP Fetch & Extract and URL Construction & Fetch, where tool use is sound but the consolidated output lacks the precision needed for reliable downstream decisions.

Does it stay within tool results or fall back on training data? Here the model is more trustworthy than the P2 score suggests. In the Honeypot EU License Research test — which checks whether current license restrictions are actually retrieved from web sources — it does not hallucinate. For compliance-adjacent research paths, this is a significant positive signal. It therefore tends to remain conservative within the evidence space, even when the summary is not sharp enough.

Error Resilience

In the 404 test — which measures whether a model stays transparent after a failed tool call or fabricates substitute content — DeepSeek V4 Pro does not hallucinate page content. That is the decisive point. The response is nonetheless poorly consolidated and not clean enough communicatively, hence the low synthesis score. For production this is acceptable, because transparency on failure matters more than elegance of phrasing.

Operational Profile

Total 200.37s per run. Slow.
Call latencies 4.59s and 27.47s, MCP 1.34s.
Cost: local. Pricing: $0.435 per 1M input, $0.87 per 1M output. Affordable for Frontier class, but runtime is high relative to output quality.

Summary & Recommendation

Suitable for agentic research and orchestration pipelines where tool selection, long contexts, and cautious error handling matter more than perfect final consolidation. Not the first choice for compliance approvals, precise fact extraction, or strictly deterministic MCP pipelines with hard schema and URL requirements. Deploy only with tight tool call validation, output checks, and a downstream verifier for synthesis quality.

This analysis was generated automatically based on benchmark data. Model used: GPT-5.4 by OpenAI. The raw data and full methodology are documented in the GitHub project.