Grok 4.5

Grok 4.5 has been xAI’s current flagship since July 2026, with native real-time access to web and X data, a 500,000-token context window, and multimodal text and image input. Reasoning runs server-side at all times and is controlled via configurable reasoning effort levels, with no visible chain-of-thought tags in the response text. Function calling, structured outputs, and prompt caching round out the profile for coding and agentic workflows.

xAI Version 4.5 Commercial use restricted Dense 500 K Context 02/2026 $2 / $6 per 1M

  • Proprietary
  • Frontier
  • xAI
  • Text
  • Vision
  • Interactive

Sovereign Risk: MEDIUM xAI is a US company subject to the CLOUD Act. The operational risk lies not in the distribution of weights (proprietary, not public) but in the use of the cloud API: data is processed server-side, with a default retention period of 30 days unless a zero-data-retention option is activated for enterprise accounts.[web:590]

Tool-use profile: 6 assets in detail

Comparison of asset performance (P1/P2/Combined) against the fleet average

Asset performance (radar)

Score Breakdown vs. Fleet Average


Tool-use details

Asset performance, reliability, and runtime profile

CrucibleMark evaluates tool use across 6 independent tests. Click a test name for details.

Reliability

  • Tool Call Valid: No
  • Retry: Not required
  • Hallucination: Detected

Reliability measures how consistently a model actually executes tool calls: Tool Call Valid schema and format accepted, Retry Required successful only after retry, Hallucination Flag fabricated tools or parameters detected. All three green means production-ready.

Runtime profile

Call 1
2.33
First Request
MCP
1.4
Protocol Latency
Synthesis
12.9
Response Generation
Total
99.77
Sum of All Phases
Token
15470
Input + Output
Cost
$0
Cost per Run

The runtime profile shows the latency and cost metrics for the model run: Call 1 First Request, MCP Protocol Latency, Synthesis Response Generation, Total sum of all phases, supplemented by Token input and output and Cost cost per run.

Tool-use review

Updated

Deployment Verdict

Do not deploy for trust-critical MCP pipelines, because Grok 4.5 hallucinates despite a good overall score and simultaneously fails to produce a consistently valid tool-call path.

Tool Execution Profile

The model demonstrates genuine situational intelligence in tool selection. In the Web Search & Tool Selection test — which checks whether the model chooses web_search over fetch without being prompted — it reliably identifies the appropriate access path. This argues against mere pattern-following. It also uses tools actively and purposefully in Multilingual Search & Synthesis.

Execution precision is weaker. In the URL Construction & Fetch test, which measures the derivation of a target URL from prior knowledge and the subsequent fetch, it performs adequately but not deterministically enough for hard production paths. The global finding “Tool-Call valid: false” is decisive here. The model can plan tool usage but does not consistently deliver a protocol-clean, reliably machine-processable execution path. On the positive side, no retry was required. The problem therefore lies more in first-attempt precision than in mere formatting errors after correction loops.

Synthesis Fidelity

How well does it condense tool results? Inconsistently. In HTTP Fetch & Extract and URL Construction & Fetch, Grok 4.5 condenses cleanly and with sufficient precision. Web Search & Tool Selection is also strong. However, the overall P2 score of 60 shows that this quality does not hold consistently across all task types. Particularly with ambiguous or compliance-adjacent factual situations, the condensation tips from useful summary into uncertain assertion.

Does it stay within the tool result or fall back on training data? No — and that is the critical point. In the Honeypot EU License Research test, which is designed to check whether current license restrictions are actually retrieved from web sources, P2 is 15 and a hallucination was detected. This is not merely a quality deficiency but a security risk. When a model in a tool pipeline outputs invented or training-reconstructed statements as researched findings, the entire infrastructure loses its auditability.

Error Resilience

In the 404 test, which measures transparent handling of failed tool calls, Grok 4.5 stays on the acceptable side. It does not hallucinate page content despite the error. Error communication is therefore production-ready, even if the substantive presentation at P2 60 is only averagely clear. For operational pipelines, transparency matters more here than linguistic elegance.

Operational Profile

Call 1: 2.33s. MCP latency: 1.40s. Call 2: 12.90s. Total: 99.77s.
Price: $2.0 per 1M input tokens, $6.0 per 1M output tokens.
Assessment: on the slow side for the measured utility, pricing is Frontier-typical, but not favorable relative to the trust risk.

Conclusion & Recommendation

Grok 4.5 is suitable for assistive research, exploratory analyst workflows, and human-supervised tool chains where results are visibly reviewed. It is not suitable for compliance, license verification, policy evaluation, autonomous retrieval synthesis, or any pipeline in which tool results are passed on as reliable facts. If you do deploy it, only do so with strict output verification, mandatory source attribution, and downstream validation outside the model.

This analysis was generated automatically based on benchmark data. Model used: GPT 4.5 by OpenAI. The raw data and full methodology are documented in the GitHub project.