Tool-use review
· Agentic Orchestrator · Long Context
Tool-use profile
GLM-5.3-Flash (EXL3 High) achieves a Combined Score of 70.4 (Good) in the tool-use benchmark: P1 Execution 79.2, P2 Synthesis 63.3, fleet average 68.5.
Strongest test: Web Search & Tool Selection (28.3). Weakest test: HTTP Fetch & Extract (90). The spread between these two tests is -61.7 points.
Reliability status: Tool Call Valid No, Retry Not required, Hallucination Not detected.
This data-driven auto-review is compiled from the available tool-use benchmark data. Once a detailed LLM-generated analysis (GPT-5.4) is available, it will automatically replace this template. The raw data and full methodology are documented in the GitHub project.