Upstage Solar Pro4

Solar Pro 4 is Upstage’s agentic flagship model, released on August 6/10, 2026, with an undisclosed parameter count. It is available as a proprietary cloud service (additionally as dedicated/on-premises deployment for enterprise customers) and is designed for multi-step agentic workflows across documents, terminals, and tool calls. The model offers a context window of 524,288 tokens with up to 131,072 output tokens, supports English, Korean, and Japanese for input and output, and allows a configurable ‘Reasoning Effort’ (high for deep analysis, low for real-time chat speed).

Upstage Version pro4 Commercial use permitted Dense 524 K Context 02/2026 $0.03 / $0.12 per 1M

  • Proprietary
  • Frontier
  • OpenRouter
  • Text
  • Long Context
  • Agentic Orchestrator
  • Interactive

Sovereign Risk: MEDIUM Upstage is a South Korean company. The risk is rated as medium, as the exact hosting terms and the potential applicability of foreign laws (e.g., the US CLOUD Act, if hosted via US cloud infrastructure) are not publicly documented. Enterprise customers can contractually arrange dedicated or on-premises deployments, which can significantly reduce the risk for this user group.

Tool-use profile: 6 assets in detail

Comparison of asset performance (P1/P2/Combined) against the fleet average

Asset performance (radar)

Score Breakdown vs. Fleet Average


Tool-use details

Asset performance, reliability, and runtime profile

CrucibleMark evaluates tool use across 6 independent tests. Click a test name for details.

Reliability

  • Tool Call Valid: No
  • Retry: Not required
  • Hallucination: Detected

Reliability measures how consistently a model actually executes tool calls: Tool Call Valid schema and format accepted, Retry Required successful only after retry, Hallucination Flag fabricated tools or parameters detected. All three green means production-ready.

Runtime profile

Call 1
2.91
First Request
MCP
1.44
Protocol Latency
Synthesis
34.4
Response Generation
Total
232.48
Sum of All Phases
Token
12078
Input + Output
Cost
$0
Cost per Run

The runtime profile shows the latency and cost metrics for the model run: Call 1 First Request, MCP Protocol Latency, Synthesis Response Generation, Total sum of all phases, supplemented by Token input and output and Cost cost per run.

Tool-use review

Updated · Long Context · Agentic Orchestrator

Deployment Verdict

Conditional deploy, because tool execution is strong, but one invalid tool call and detected hallucinations limit confidence in production MCP pipelines. The overall impression is good; the safety posture is not.

Tool Execution Profile

Upstage Solar Pro4 demonstrates genuine tool intelligence, not just rigid pattern-matching. On the Web Search & Tool Selection test — which checks whether the model correctly chooses between search and direct retrieval without an explicit hint — it reliably identifies the need for web_search. That is a strong signal for agentic orchestration. On the URL Construction test, which measures correct derivation of a target address followed by a fetch, it remains usable but not deterministic enough for sensitive paths.

P1 of 90 speaks to high operational competence. The catch is protocol compliance: Tool-Call valid is False. The risk therefore lies not in whether the model is willing to use tools in principle, but in whether individual calls are formatted and executed cleanly enough for a strict MCP infrastructure. Since no retry was required, this looks more like a localized validity issue than a systematic comprehension failure.

Synthesis Fidelity

How well does it consolidate tool results? Only limitedly reliable. P2 of 66.67 is the clear separator from the strong execution score. On EU License Research — a live web query about licensing restrictions — it consolidates results only moderately. On HTTP Fetch & Extract, extraction remains solid but not precise enough to serve as a reference answer without downstream validation.

Does it stay within the tool result or fall back on training data? On the honeypot EU License Research, it stays formally within safe territory: no hallucination detected. That matters. At the same time, global hallucination detected is True. This turns the issue from a quality deficiency into a safety risk. When a model outputs fabricated content as a tool result, the entire pipeline loses its auditability.

Error Resilience

This is where the model fails. On the 404 test — which checks for transparent handling of a failed tool call — it hallucinates page content despite the error. P2 of 15 is secondary here. What matters is the finding itself: hallucinated substitute content instead of a clear error signal. That is critical for production, without exception.

Operational Profile

Total 232.48s per run. Slow. Call 1: 2.91s, MCP latency: 1.44s, Call 2: 34.40s.
Cost/run: local. Inexpensive. Attractive relative to performance; sluggish relative to runtime.

Conclusion & Recommendation

Suitable for agentic research and routing pipelines where tool selection matters more than hard reliability of the final answer and a guardrail layer validates every output. Not suitable for compliance, support automation, incident flows, or any pipeline where a tool failure must remain strictly visible as a failure. Anyone deploying Solar Pro4 should make tool output verification, strict 404 abort logic, and response gating mandatory upstream controls.

This analysis was generated automatically based on benchmark data. Model used: GPT-5.4 by OpenAI. The raw data and full methodology are documented in the GitHub project.