Gemini 3.5 Flash

Gemini 3.5 Flash is Google’s fastest Frontier-class model, delivering near Pro-level performance at Flash pricing. With Dynamic Thinking across four configurable levels, a one-million-token context window, and full multimodality for text, images, audio, video, and PDF, the model is well-suited for agentic workflows, coding, and high-throughput workloads.

Google Version 3.5-flash Commercial use permitted MoE 1000 K Context 01/2025 $1.5 / $9 per 1M

  • Proprietary
  • Frontier
  • OpenRouter
  • Text
  • Vision
  • Audio
  • Video
  • Agentic Orchestrator
  • Real-Time

Sovereign Risk: MEDIUM Google DeepMind is a US-based company and subject to the CLOUD Act; the model weights are not publicly accessible.

Tool-use profile: 6 assets in detail

Comparison of asset performance (P1/P2/Combined) against the fleet average

Asset performance (radar)

Score Breakdown vs. Fleet Average


Tool-use details

Asset performance, reliability, and runtime profile

CrucibleMark evaluates tool use across 6 independent tests. Click a test name for details.

Reliability

  • Tool Call Valid: No
  • Retry: Not required
  • Hallucination: Not detected

Reliability measures how consistently a model actually executes tool calls: Tool Call Valid schema and format accepted, Retry Required successful only after retry, Hallucination Flag fabricated tools or parameters detected. All three green means production-ready.

Runtime profile

Call 1
1.31
First Request
MCP
1.12
Protocol Latency
Synthesis
5.25
Response Generation
Total
46.06
Sum of All Phases
Token
6717
Input + Output
Cost
$0
Cost per Run

The runtime profile shows the latency and cost metrics for the model run: Call 1 First Request, MCP Protocol Latency, Synthesis Response Generation, Total sum of all phases, supplemented by Token input and output and Cost cost per run.

Tool-use review

Updated · Agentic Orchestrator

Deployment Verdict

Conditional deploy, because tool use is mostly strong, but tool calls are not consistently valid and synthesis fidelity falls short too often for production-critical pipelines. The overall picture is good, but not stable enough in terms of reliability for unsupervised end-to-end orchestration.

Tool Execution Profile

Gemini 3.5 Flash demonstrates genuine tool selection competence, not just rigid pattern behavior. On the Web Search & Tool Selection test — which checks whether web_search is chosen over fetch without an explicit hint — the model reliably identifies the appropriate strategy. This speaks to usable orchestration logic in dynamic MCP setups.

The execution layer is weaker. On the URL Construction test, which measures correct derivation of a target URL and the subsequent fetch, it performs adequately but not precisely enough for deterministic pipelines. Consistent with this, tool calls were not consistently valid globally. This is not a total failure, but it is an integration risk: the model often plans correctly, yet does not always produce the execution in a protocol-safe manner.

Synthesis Fidelity

How well does it consolidate? Only moderately. P2 performance shows that Gemini 3.5 Flash often aggregates tool results correctly, but loses important precision in condensed output. This is most visible in EU License Research and Multilingual Search & Synthesis, where the research succeeds but the final consolidation remains too coarse. For operational assistants this is acceptable. For compliance, policy, or decision-support flows it is too imprecise.

Does it stay within tool results or fall back on training data? On the EU License Research honeypot — which checks whether current license restrictions are answered from web sources rather than training knowledge — the model formally stayed on the safe side. It does not hallucinate. This is the critical trust anchor. At the same time, the weak synthesis result is a warning signal: no security breach, but insufficient reliable consolidation for regulatory statements.

Error Resilience

On the 404 test, which measures whether a failed tool call is handled transparently or covered up with fabricated page content, Gemini 3.5 Flash responds acceptably. It does not hallucinate substitute content despite the error. Error communication is therefore production-ready, even if it does not consolidate particularly well or proactively reroute.

Operational Profile

Total 46.06s per run. Individual calls 1.31s and 5.25s. MCP latency 1.12s. Fast on individual steps, but long on overall runtime. Price: $1.5/1M input, $9.0/1M output. Not cheap for a Frontier model. Justifiable given the performance only when tool selection matters more than high-precision final synthesis.

Conclusion & Recommendation

Suitable for MCP pipelines with human review, for research orchestration, routing, multi-step web use, and robust error handling without hallucination risk. Not the first choice for unsupervised compliance workflows, license assessments, multilingual executive summaries, or other chains where the final consolidation must serve as a reliable working basis. If you deploy it, do so with strict output validation after the tool layer.

This analysis was generated automatically based on benchmark data. Model used: GPT-5.4 by OpenAI. The raw data and full methodology are documented in the GitHub project.