Grok 4.3

Grok 4.3 is xAI’s current flagship model with native real-time access to the X platform and a context window of one million tokens. The MoE architecture enables efficient inference on agentic routine tasks, but falls short of specialized models in complex software engineering. Available exclusively via the xAI API at $1.25 per million input tokens and $2.50 per million output tokens.

xAI Version 4.3 Commercial use permitted MoE 1000 K Context 03/2026 $1.25 / $2.5 per 1M

  • Proprietary
  • Frontier
  • API
  • Text
  • Vision
  • Real-Time

Sovereign Risk: MEDIUM xAI is a US company subject to the CLOUD Act. Since the weights are proprietary, there is no risk from the distribution of the weights themselves.

Tool-use profile: 6 assets in detail

Comparison of asset performance (P1/P2/Combined) against the fleet average

Asset performance (radar)

Score Breakdown vs. Fleet Average


Tool-use details

Asset performance, reliability, and runtime profile

CrucibleMark evaluates tool use across 6 independent tests. Click a test name for details.

Reliability

  • Tool Call Valid: No
  • Retry: Not required
  • Hallucination: Detected

Reliability measures how consistently a model actually executes tool calls: Tool Call Valid schema and format accepted, Retry Required successful only after retry, Hallucination Flag fabricated tools or parameters detected. All three green means production-ready.

Runtime profile

Call 1
2.74
First Request
MCP
2.15
Protocol Latency
Synthesis
5.97
Response Generation
Total
65.14
Sum of All Phases
Token
9504
Input + Output
Cost
$0
Cost per Run

The runtime profile shows the latency and cost metrics for the model run: Call 1 First Request, MCP Protocol Latency, Synthesis Response Generation, Total sum of all phases, supplemented by Token input and output and Cost cost per run.

Tool-use review

Deployment Verdict

Created on: 14.06.2026, 16:17:24

Conditional deploy, because Grok 4.3 delivers valid tool calls and does not hallucinate, but synthesis fidelity remains too unreliable for production-grade tool pipelines.

Tool Execution Profile

At the tool execution level, the model performs adequately overall. Tool call valid: true and no retry was required. This indicates clean MCP conformance and argues against format issues at the protocol level. The P1 score of 83.33 reflects stable tool operation, but not precise orchestration at Frontier level.

In terms of tool selection, Grok 4.3 comes across as rule-driven rather than genuinely selective. In the Web Search & Tool Selection test — which requires distinguishing between search and fetch without an explicit hint — it achieves solid execution but no clear strength. In the URL Construction test, which requires deriving the correct target URL from internal knowledge and then executing a fetch, the picture is similar. Both results landing at the same level suggest the model uses tools reliably but does not always identify the most information-efficient path.

Synthesis Fidelity

How well does it condense tool results? Only to a limited degree. The P2 score of 43.33 is this model’s real bottleneck. In HTTP Fetch & Extract and Multilingual Search & Synthesis it still delivers usable condensation. In several other tasks, however, it falls into shallow or incomplete summaries. For pipelines where tool results are merely passed through or lightly normalized, this is tolerable. For compliance, research, or decision-relevant executive summaries, it is too weak.

Does it stay within the tool result or fall back on training data? In the Honeypot EU License Research test — which checks exactly this for current license restrictions — it does not hallucinate. That is the important trust signal. At the same time, P2 is very low there at 20, and the Content Verification State B2 shows: it stays formally on the safe path, but does not process the retrieved content with the required accuracy. This is not a security breach, but it is a verification risk.

Error Resilience

In the 404 test — which checks for transparent handling of a failed tool call rather than fabricated page content — the model stays clean. It does not hallucinate despite the error. However, P2=40 means here as well: error communication is acceptable, but not particularly precise or user-guiding. For production this is manageable, as long as downstream systems handle error states themselves.

Operational Profile

Total 52.75s per run. MCP latency 0.92s. Model calls 2.70s and 5.17s. Overall slow for the synthesis quality achieved. Cost per run: 0.011412 USD. Inexpensive to moderate, but the price-to-performance ratio remains only average due to weak condensation.

Conclusion & Recommendation

Suitable for tool pipelines with clear guardrails, where the model is expected to search, retrieve, and cautiously summarize results. Well suited for simple retrieval, monitoring, and multilingual research workflows with human or rule-based final review. Not suitable as the final synthesis layer for compliance, license assessment, policy interpretation, or other paths where the summary itself carries the decision.

This analysis was generated automatically based on benchmark data. Model used: GPT 4.5 by OpenAI. The raw data and full methodology are documented in the GitHub project.