Meta Muse Spark 1.2

Muse Spark 1.2 is Meta’s proprietary Frontier model from August 5, 2026, designed for coding and agentic workflows, released by Meta Superintelligence Labs alongside the terminal agent Muse Code, with which it was co-trained. The cloud-only model under US jurisdiction (CLOUD Act) processes text, image, video, and audio with a context of 1,048,576 tokens (max. 131,072 output) and offers configurable reasoning effort up to ‘xhigh’.

Meta Version 1.2 Commercial use permitted Dense 1024 K Context $1.25 / $4.25 per 1M

  • Proprietary
  • Frontier
  • OpenRouter
  • Text
  • Vision
  • Audio
  • Video
  • Long Context
  • Agentic Orchestrator
  • Real-Time

Sovereign Risk: MEDIUM Meta is a US-based company and subject to the CLOUD Act; the model weights are proprietary and not publicly accessible (no self-hosting, no fine-tuning possible). Meta additionally offers a ‘Contributor’ pricing tier in which users agree, in exchange for significantly reduced costs, that their prompts may be used to train future Meta models — under this tier, the actual data risk increases considerably compared to the standard tier.

Tool-use profile: 6 assets in detail

Comparison of asset performance (P1/P2/Combined) against the fleet average

Asset performance (radar)

Score Breakdown vs. Fleet Average


Tool-use details

Asset performance, reliability, and runtime profile

CrucibleMark evaluates tool use across 6 independent tests. Click a test name for details.

Reliability

  • Tool Call Valid: No
  • Retry: Not required
  • Hallucination: Not detected

Reliability measures how consistently a model actually executes tool calls: Tool Call Valid schema and format accepted, Retry Required successful only after retry, Hallucination Flag fabricated tools or parameters detected. All three green means production-ready.

Runtime profile

Call 1
1.75
First Request
MCP
1.4
Protocol Latency
Synthesis
9.81
Response Generation
Total
77.74
Sum of All Phases
Token
18983
Input + Output
Cost
$0
Cost per Run

The runtime profile shows the latency and cost metrics for the model run: Call 1 First Request, MCP Protocol Latency, Synthesis Response Generation, Total sum of all phases, supplemented by Token input and output and Cost cost per run.

Tool-use review

Created · Long Context · Agentic Orchestrator

Deployment Verdict

Conditional deploy, because the overall score is good, but the tool call was not valid — and that means the critical production question cannot be answered with a clean positive.

Tool Execution Profile

The model demonstrates genuine tool-selection competence, but not clean protocol reliability. In the Web Search & Tool Selection test, which checks whether the correct research tool is chosen without a hint, it clearly recognizes that web_search is needed rather than fetch. This argues against a rigid pattern and in favor of situation-aware tool selection. In the URL Construction & Fetch test, which measures independent derivation of a target URL followed by retrieval, it remains usable but not deterministic enough. P1 scores of 100 and 80 therefore indicate: sound decision-making at the planning level, lower precision in concrete execution.

The global finding tool_call_valid=false is critical. Even without a retry being needed, this points to a formal or semantic break in the call — not merely a robustness issue. For MCP pipelines, this means: the orchestration concept is present, but the handoff to infrastructure requires guardrails, schema validation, and tight runtime controls.

Synthesis Fidelity

How well does it consolidate tool results? Solid, but not precise enough for high-quality retrieval pipelines. Individual scores vary noticeably: HTTP Fetch & Extract — structured fact extraction from real page content — lands at 60. Multilingual Search & Synthesis — cross-lingual research with German-language consolidation — likewise at 60. The model can aggregate results, but loses detail and prioritization in the process.

Does it stay within the tool output or fall back on training data? Not reliably enough. In the honeypot EU License Research, which checks whether current license restrictions are sourced from the web rather than from training knowledge, the confidence side drops sharply with P2=40. It does not hallucinate overtly, but it does not bind the tool research tightly enough to the answer. For compliance, policy, or licensing workflows, this is a warning signal.

Error Resilience

The model responds acceptably to tool failures. In the Tool Failure Handling (404) test, which checks whether failed retrievals are communicated transparently rather than replaced with fabricated content, it communicates the error instead of inventing page content. P2=80 and no hallucination finding represent a viable minimum for production. This protects the pipeline against silent misinformation.

Operational Profile

Total 77.74s. Call 1 1.75s, MCP latency 1.40s, Call 2 9.81s. Slow for the quality level shown. Cost per run: listed as local, but the model profile itself is cloud-only. Pricing: $1.25/1M input, $4.25/1M output. Not expensive for Frontier, but the runtime makes it economical only when tool orchestration matters more than throughput.

Conclusion & Recommendation

Suitable for agentic pipelines with human oversight: research workflows, multi-step tool selection, robust error communication. Not suitable for strictly deterministic MCP pipelines where every tool call must be formally correct, and not for compliance-adjacent synthesis tasks where the model must stay strictly within retrieved material. Deploy only with a call validator, structured output verification, and a downstream check of the final answer against the tool artifacts.

This analysis was generated automatically based on benchmark data. Model used: GPT-5.4 by OpenAI. The raw data and full methodology are documented in the GitHub project.