Gemini 3.5 Flash

Gemini 3.5 Flash is Google’s fastest Frontier-class model, delivering near Pro-level performance at Flash pricing. With Dynamic Thinking across four configurable levels, a one-million-token context window, and full multimodality for text, images, audio, video, and PDF, the model is well-suited for agentic workflows, coding, and high-throughput workloads.

Google Version 3.5-flash Commercial use permitted MoE 1000 K Context 01/2025 $1.5 / $9 per 1M

  • Proprietary
  • Frontier
  • API
  • Text
  • Vision
  • Audio
  • Video
  • Agentic Orchestrator
  • Real-Time

Sovereign Risk: MEDIUM Google DeepMind is a US company and subject to the CLOUD Act; model weights are not publicly available.

Tool-use profile: 6 assets in detail

Comparison of asset performance (P1/P2/Combined) against the fleet average

Asset performance (radar)

Score Breakdown vs. Fleet Average


Tool-use details

Asset performance, reliability, and runtime profile

CrucibleMark evaluates tool use across 6 independent tests. Click a test name for details.

Reliability

  • Tool Call Valid: No
  • Retry: Not required
  • Hallucination: Not detected

Reliability measures how consistently a model actually executes tool calls: Tool Call Valid schema and format accepted, Retry Required successful only after retry, Hallucination Flag fabricated tools or parameters detected. All three green means production-ready.

Runtime profile

Call 1
1.59
First Request
MCP
0.91
Protocol Latency
Synthesis
6.13
Response Generation
Total
51.79
Sum of All Phases
Token
6287
Input + Output
Cost
$0
Cost per Run

The runtime profile shows the latency and cost metrics for the model run: Call 1 First Request, MCP Protocol Latency, Synthesis Response Generation, Total sum of all phases, supplemented by Token input and output and Cost cost per run.

Tool-use review

Created · Agentic Orchestrator

Deployment Verdict

Conditional deploy, because Gemini 3.5 Flash delivers valid tool calls, showed no hallucination during the run, and reliably interfaces with infrastructure with a solid overall profile — however, synthesis quality remains too inconsistent for decision-relevant outputs.

Tool Execution Profile

Tool execution is this model’s clear strength. At P1 90, it selects tools correctly in most cases, produces valid calls, and remains MCP-compliant. Crucially, on the Web Search and Tool Selection test — which checks whether the model chooses between search and direct fetch without an explicit hint — it cleanly switches to web_search. This argues against pure schema-following and in favor of genuine tool selection based on the task at hand.

On the URL Construction test, which checks whether the model can derive a target URL from its own knowledge and then fetch it correctly, it falls back to P1 80. That is workable, but not deterministic enough for pipelines where URL construction needs to be precise. No retry was required. The issue therefore lies not in the protocol format but in the content-level precision of individual execution steps. As a vision-language MoE model, the text-only findings should also not be overstated — they reflect only the linguistic tool-use component, not the model’s primary multimodal strengths.

Synthesis Fidelity

How well does it condense tool results? Only partially convincing. P2 56.67 indicates that Gemini 3.5 Flash often picks up retrieved content correctly but does not reliably convert it into concise, dependable result texts. This is most visible in EU License Research and Multilingual Search & Synthesis, where condensation drops to 40. By contrast, HTTP Fetch & Extract comes in considerably cleaner at 80. Curated extraction suits it better than multi-source synthesis.

Does it stay within tool results or fall back on training data? The trust signal here is stronger than the P2 scores might suggest. In the Honeypot EU License Research test — which checks whether current licensing restrictions are actually answered from web sources rather than training knowledge — no hallucination was detected. Content Verification State A supports this assessment. The model paraphrases weakly, but it does not fabricate.

Error Resilience

On the 404 test, which measures how the model handles a failing tool call, it remains transparent and does not hallucinate substitute content. P2 60 is not a quality score for elegant error reporting, but the operationally decisive point is met: it does not break trust by inventing page data.

Operational Profile

Call 1: 1.56s. MCP latency: 0.88s. Call 2: 6.99s. Total: 56.58s.
Cost/run: 0.022908 USD.
Direct assessment: tool calls fast, total run long, costs moderate. Economically justifiable for the performance shown, but not aggressively cheap.

Conclusion & Recommendation

Suitable for MCP pipelines with clear tool orchestration, web research, fetch extraction, and a robust error path — particularly where non-hallucination matters more than linguistically strong final synthesis. Not the first choice for compliance, policy, or executive reporting workflows where multi-source synthesis must be precise and structurally consistent. Recommended as an execution-layer research and retrieval model behind a downstream validation or editorial stage.

This analysis was generated automatically based on benchmark data. Model used: GPT 4.5 by OpenAI. The raw data and full methodology are documented in the GitHub project.