Gemini 2.5 Pro

Google’s Frontier reasoning model with sparse MoE architecture and configurable Extended Thinking. Gemini 2.5 Pro operates with a one-million-token context window, natively processes text, images, audio, and video, and is exclusively accessible via the Google Cloud API. The focus is on complex reasoning and demanding coding tasks.

Google Version 2.5-pro Commercial use permitted MoE 1000 K Context 01/2025 $1.25 / $10 per 1M

  • Proprietary
  • Frontier
  • OpenRouter
  • Text
  • Vision
  • Audio
  • Video
  • Agentic Orchestrator
  • Interactive

Sovereign Risk: MEDIUM Google DeepMind is a US-based company and subject to the CLOUD Act; the model weights are not publicly accessible.

Tool-use profile: 6 assets in detail

Comparison of asset performance (P1/P2/Combined) against the fleet average

Asset performance (radar)

Score Breakdown vs. Fleet Average


Tool-use details

Asset performance, reliability, and runtime profile

CrucibleMark evaluates tool use across 6 independent tests. Click a test name for details.

Reliability

  • Tool Call Valid: No
  • Retry: Not required
  • Hallucination: Not detected

Reliability measures how consistently a model actually executes tool calls: Tool Call Valid schema and format accepted, Retry Required successful only after retry, Hallucination Flag fabricated tools or parameters detected. All three green means production-ready.

Runtime profile

Call 1
7.63
First Request
MCP
0.92
Protocol Latency
Synthesis
10.05
Response Generation
Total
111.64
Sum of All Phases
Token
7674
Input + Output
Cost
$0
Cost per Run

The runtime profile shows the latency and cost metrics for the model run: Call 1 First Request, MCP Protocol Latency, Synthesis Response Generation, Total sum of all phases, supplemented by Token input and output and Cost cost per run.

Tool-use review

Updated · Agentic Orchestrator

Deployment Verdict

Conditional deploy: tool usage is strong, but synthesis fidelity is still too inconsistent for production tool pipelines. The overall picture is workable, but the invalid tool call and weak condensation limit the confidence level.

Tool Execution Profile

Gemini 2.5 Pro demonstrates genuine tool intelligence, not just rigid pattern matching. On the Web Search & Tool Selection test — which checks whether web_search is chosen over fetch without an explicit hint — it makes the right call reliably. This points to workable orchestration capability in dynamic MCP pipelines. Multilingual Search & Synthesis and EU License Research also perform strongly on execution, supporting the case for solid operational tool usage.

The weakness lies less in tool selection than in the protocol cleanliness of execution. The tool call was not consistently valid, even though no retry was required. This points more to formatting or call-strictness issues than to comprehension problems. On the URL Construction test — which measures whether the model independently derives a target URL and subsequently fetches it — the model operates functionally, but not deterministically enough for fragile automation chains. For robust MCP setups with a validation layer, this is acceptable. For calls passed through directly without guardrails, it is too risky.

Synthesis Fidelity

How well does it condense tool results? Only reliably up to a point. The P2 score of 60 shows that Gemini 2.5 Pro often merges extracted information usably, but does not condense it with consistent precision. This is clearly visible in EU License Research, where condensation is weak despite correct tool usage. By contrast, HTTP Fetch & Extract and Multilingual Search & Synthesis perform considerably more stably.

Does it stay within the tool result or fall back on training? In the Honeypot EU License Research test — which checks whether current license restrictions are answered from web sources rather than from training — the model does not hallucinate. That is the decisive trust signal. The low score here is therefore not a safety breach, but a problem of summarization and source fidelity.

Error Resilience

The model behaves in a production-ready manner when tools fail. On the 404 test — which pits transparent error communication against hallucinated replacement content — Gemini 2.5 Pro does not fabricate page content. This is the correct mode for production systems. Execution itself remains error-prone at P1 40, but the response side stays honest. This can be compensated for with retries and error handling.

Operational Profile

Total 111.64s. Individual calls 7.63s and 10.05s. MCP latency 0.92s. Slow overall. Price: $1.25/1M input, $10.0/1M output. Not inexpensive for Frontier-level, measured against only middling synthesis performance.

Conclusion & Recommendation

Suitable for MCP pipelines with search, fetch, and orchestration components, provided a validation layer checks tool calls and post-processes the final output where necessary. Not suitable for compliance, policy, or other text-critical workflows where the condensation of tool results itself must already be reliably final. Those looking for a model for tool decision-making and broad research can deploy it. Those who need a model that transfers tool results precisely and faithfully into reliable answers should plan for stricter guardrails or chain a more faithful synthesis model downstream.

This analysis was generated automatically based on benchmark data. Model used: GPT-5.4 by OpenAI. The raw data and full methodology are documented in the GitHub project.