Tool-use review
Created · Agentic Orchestrator · Long Context
Tool-use profile
Claude Opus 5.5 achieves a Combined Score of 76.9 (Good) in the tool-use benchmark: P1 Execution 90, P2 Synthesis 62.5, fleet average 68.2.
Strongest test: Tool Failure Handling (404) (64.3). Weakest test: Web Search & Tool Selection (91). The spread between these two tests is -26.7 points.
Reliability status: Tool Call Valid No, Retry Not required, Hallucination Detected.
This data-driven auto-review is compiled from the available tool-use benchmark data. Once a detailed LLM-generated analysis (GPT-5.4) is available, it will automatically replace this template. The raw data and full methodology are documented in the GitHub project.