Grok 4.7

Grok 4.7 is xAI’s Frontier flagship as of September 21, 2026, built on a larger base than Grok 4.6 and trained with a focus on multi-hour tasks. 500,000 tokens of context, text and image input, four reasoning levels from low to xhigh at unchanged pricing of 2 / 6 USD per million tokens. According to xAI, nearly doubled performance on long terminal tasks and a new safeguard stack against jailbreaks.

xAI Version 4.7 Commercial use restricted Dense 500 K Context 05/2026 $2 / $6 per 1M

  • Proprietary
  • Frontier
  • xAI
  • Text
  • Vision
  • Interactive

Sovereign Risk: MEDIUM The model is developed and hosted by a US-based company. Due to US jurisdiction, it is potentially subject to the CLOUD Act, which represents a moderate risk of data access by US authorities. Since the weights are proprietary and not distributed, there is no additional risk from disclosure of the weights themselves.

LLM Model Review

Created on

With an overall score of 71.06% and the Speed Profile Badge Interactive DevOps Expert, Grok 4.7 enters as a commercial cloud model from the xAI API like a broad generalist with plenty of self-confidence and even more ambition. The pre-assigned categorization only partially fits: yes, Grok 4.7 is a generalist with Thinking DNA and vision capability, but in the text benchmark it presents itself primarily as an agentically designed Frontier system in a dense transformer architecture, built for planning and longer work chains — not for elegant precision landings in every individual discipline. That’s not consistently sovereign, but it’s certainly not characterless. Sovereign Risk: HIGH — xAI is a US provider, subject to the CLOUD Act, and according to available card data offers no reliable EU safeguard with a public GDPR DPA.

Header Scores: Stability and Reliability

Metric Value Rating Analysis
Timeout Rate 4/49 Sporadic The model shows sporadic dropouts that would require retries in practice. For a proprietary Frontier endpoint, this is not a cosmetic flaw — it’s an API risk.
P95 Response Time 149.17 s Critical Extreme tail latency. The model scatters massively and is unsuitable for time-critical processes.

Grok 4.7 was tested in the endpoint’s factory default behavior; there is no switchable Thinking mode here. This is important for context, because the architecture clearly targets deeper reasoning and agentic decomposition. The catch: a Frontier model of this class must reliably translate its internal complexity into visible quality. That is precisely where Grok 4.7 fails too often.

Architecture and Character: Frontier All-Rounder with Reasoning Ambitions

The metadata says General, Thinking, Vision-Capable. The editorial classification is more precise: agentic use case, Frontier class, dense architecture. That’s more than a label. A dense Frontier model without a known parameter count operates under maximum expectation pressure, because its full capacity is active per request and because it runs exclusively as the manufacturer’s cloud product. There is no bonus here for frugality or weight class. If you sell Frontier, you get measured against Frontier.

You can feel this design. Grok 4.7 visibly thinks in structures, decomposes tasks neatly, holds complex formats together surprisingly well, and rarely seems aimless. At the same time, it lacks — at multiple points — the final sharpness between “I understood the task” and “I executed it precisely, completely, and without collateral damage.” It’s less the picture of a frantic rambler than of a smart system that doesn’t always bring itself cleanly to the point.

The benchmark context also matters: although Grok 4.7 is vision-capable, this run measures exclusively text tasks. The model is therefore only being evaluated on part of its actual product identity. That relativizes the findings somewhat, but it does not excuse weaknesses in the text modules. A multimodal Frontier model must be able to convince without images too.

Performance and Cost: Designed for Interactivity, but Not Reliably Interactive

The Speed Profile Badge Interactive DevOps Expert signals a clear claim: prompt, work-oriented responses for technical dialogues — a model you want to use not only in overnight batch jobs but also during the day in flow. In practice, this benchmark only partially bears that out. Generation feels moderately paced overall rather than genuinely fast, and above all the long latency tail regularly destroys the promise of interactivity.

Then there’s the price. At $2.0 per million input tokens and $6.0 per million output tokens, Grok 4.7 is not absurdly expensive, but also not cheap enough to simply paper over inefficiency. When you buy an agentic Frontier model through the manufacturer’s cloud, you’re paying not only for quality but also for reliability. That is precisely where xAI falls somewhat short of its own claim.

API Cost Profile

Token-economically, Grok 4.7 hasn’t gone off the rails, but in individual modules it works considerably more verbosely than average. Particularly notable are Cultural Intelligence and CLI Benchmark. In Cultural Intelligence, the model produces an average of 1,426 tokens against a fleet median of 257 — a factor of 5.55x compared to the average across all tested models. In the CLI Benchmark, 1,400 tokens face a fleet median of 303, meaning 4.62x as much text.

For a proprietary cloud model, this is not an academic finding. More tokens mean a higher bill, directly. If the extra length delivers qualitative value, that’s acceptable. If it merely expands what could have been half as long and still held up, verbosity becomes a cost factor.

Code Quality: Plenty of Findings, Too Little Distillation

In the Code Quality module, Grok 4.7 lands at 69.55%. That’s not a collapse, but for a Frontier model with an agentic self-image it’s no cause for pride either. Qualitatively, a recurring pattern emerges: the model reliably identifies many vulnerabilities, correctly names standards like SQL Injection, XSS, Session Fixation, CSRF, and IDOR, and delivers the required table structure mostly cleanly. It can analyze. What it more often lacks is the prioritized distillation that distinguishes a security review from a long list of findings.

A good example is the audit of a PHP application’s security analysis. There, Grok 4.7 covers the major topics solidly and even identifies several of the hidden implicit vulnerabilities. At the same time, it misses a core requirement of the task: precisely those five implicit vulnerabilities should have been clearly extracted as their own group. Instead, they partially dissolve into the stream of findings. The model knows a lot, but it doesn’t curate hard enough. A security reviewer needs not just a net, but also a magnet.

There is also a concrete technical flaw: in the same task, the last table row breaks off mid-structure. In the Code Quality section, one output breaks off in the middle of a table — the response is technically truncated, not an error in content. The score deduction results from the incomplete response, not from substantive flaws. Especially in audits that are often copied into downstream processes, this is not a minor detail. Half a table row is enough to halve trust.

The module score is clearly readable as a result: Grok 4.7 is competent enough in security matters to serve as a first-pass scanner. For reliable audit work, it lacks the discipline to cleanly bundle implicit risks and deliver results without structural defects.

CLI and Tool Proximity: The Best Part of the Package

In the CLI Benchmark, Grok 4.7 achieves 90.67%. This is the model’s clear strength and the point where its agentic orientation finally delivers rather than merely claims. Shell-adjacent tasks, technical sequences, command logic, and practical execution suit it visibly better than didactic explanatory pieces or fine editorial formats.

This is consistent with its product identity. xAI positions Grok 4.7 itself with a focus on longer terminal tasks and multi-hour workflows. The benchmark confirms, at least in miniature, that this direction has substance. Anyone looking for a model for DevOps-adjacent interaction, tool preparation, or technical workflow planning gets one here that doesn’t constantly stumble. It’s not surgically concise, but it’s mostly workable. And in the tool domain, that’s often more important than stylistic elegance.

Reasoning and Logic: Correct, but Surprisingly Shallow for a Thinking Model

The biggest disappointment lies in the Logical Reasoning module at 58.96%. For a model classified as Thinking-capable, whose manufacturer explicitly advertises reasoning tiers, that’s not enough. Not because Grok 4.7 is consistently wrong — on the contrary, the logic in the available transcripts is frequently correct. The problem is that the visible response is often too brief, too little exploratory, and didactically too thin.

The transcript for the classic two-guards puzzle illustrates this cleanly. Grok 4.7 finds the right question, correctly explains the mechanism, and chooses the right door. Substantively, that works. But instead of genuinely working through the required exploration of alternative approaches, it delivers a brief utility version. For an instruct model, that would be a legitimate style. For a Frontier system with a Thinking character, it’s a squandered advantage. The model apparently thinks more internally than it converts into usable explanation externally.

Metacognition Compliance (Reasoning): The model refuses to use the explicitly requested <thought> tags in 3/5 metacog tests, with a consistent policy statement. The reasoning content is partially correct in substance — the score deduction results from format non-compliance, not from reasoning errors. For comparison: in the tag-free reasoning_5* tests, the model achieves an average score of approximately 59%, which roughly corresponds to its general reasoning level. CrucibleMark deliberately evaluates native zero-shot instruction compliance as a real-world characteristic — this deduction is methodologically intentional.

There is also another cleanly documented misstep: in a metacognitive task, Grok 4.7 responded in English despite a German-language instruction. This is not merely a stylistic breach but a rule-based failure. When a model drops the language requirement first under combined instructions covering language, format, and reasoning representation, that’s more than a slip. It reveals where the system’s priorities actually lie under load.

Content Transformation: Strong, but Not Linguistically Clean Enough

In the Content Transformation & Adaption module, Grok 4.7 achieves 74.17%. That’s solid overall and includes some of the model’s most convincing individual performances. Particularly strong is the video script transformation: there, Grok 4.7 builds a complete, production-ready sequence with timestamps, hook, pattern interrupt, editor notes, troubleshooting, CTA, and even a double Easter egg. That’s not just formally complete — it’s editorially usable. The model can not only rephrase content but tailor it to a target medium. When it wants to, it works like a conscientious producer.

That’s precisely why the flip side stands out more sharply. In another task in this module, the model ignored the explicit language instruction and responded in English when German was required. The system applied an automatic constraint deduction for this. Content quality is secondary at that point, because the penalty applies regardless of style. In production translation or adaptation pipelines, this is a clear-cut risk.

The language failure is not an isolated outlier. Across multiple tasks in Content Transformation and Reasoning, Grok 4.7 shows a consistent pattern: when faced with simultaneous instructions covering language, format, and reasoning framework, it drops the language requirement as the first condition. That’s not a catastrophe, but it’s exactly the kind of error that hits directly in automated editorial or support flows without human oversight.

UX Writing: Usable, but Not Refined Enough

In the UX Writing & Microcopy module, Grok 4.7 comes in at 73.19%. That’s solid mid-tier work with visible structural understanding, but without the final polish in tone and psychological precision. In one of the available transcripts, it correctly analyzes conversion problems, identifies cognitive overload, jargon, and mobile length, and builds a coherent optimization table. The craft is there.

What’s missing is depth. The required analysis runs too long and list-heavy rather than concise and distilled, psychological principles remain somewhat at the level of “correctly named, not fully internalized,” and the communicative translation for stakeholders remains underdeveloped. Put differently: Grok 4.7 can identify UX problems, but it doesn’t sell the solution with the elegance of an experienced product writer. It’s the analyst, not the fine copywriter.

Documentation Quality: Surprisingly Weak for Long-Context Ambitions

With 62.08% in the Documentation Quality module, Grok 4.7 incurs one of the more problematic weaknesses in its entire profile. This is noteworthy because xAI explicitly loads the model with long-term and long-context competence. A context window of 500,000 tokens is impressive on paper. But a large window does not prove good editorial judgment.

Documentation work demands something many large models underestimate: not just completeness, but prioritization, structural hygiene, and the ability to serve information at the right levels of abstraction. Grok 4.7 often feels heavier here than necessary. It can carry material, but it can’t always arrange it cleanly. For an agentic Frontier model in particular, that’s not enough. A model tasked with steering complex workflows must also convincingly master their documentation.

Cultural Intelligence: Respectable and Mostly Linguistically Clean

In the Cultural Intelligence module, Grok 4.7 achieves 77.84% and shows a pleasantly mature side. Rewriting a toxic job posting into professional, inclusive German succeeds convincingly. The model removes aggressive and gender-coded phrasing, respects the task scope, and stays in the correct register. It doesn’t work brilliantly here, but reliably enough to be useful in everyday contexts.

The cost profile is worth noting here as well, however. Response lengths are significantly above the fleet median. That doesn’t mean quality is poor — it just means: anyone deploying Grok 4.7 for many culture- or tone-related text adaptations via the xAI cloud pays for this verbosity.

Security and Hallucination Profile: Sober Rather Than Fabulist

The security picture is mixed but usable. Grok 4.7 identifies many classic vulnerabilities but doesn’t always prioritize them with the necessary sharpness. It works for triage, initial assessment, and problem-oriented review. For final security-critical sign-offs, the last degree of rigor is missing. That’s not a dismissal — more a clear role assignment: good co-auditor, not a solo final signer.

Hallucinations are not the main adversary in the available transcripts. The model stumbles more on form, language, or completeness than on freely invented nonsense. That’s the better kind of weakness. A model that gets tangled up is easier to manage than one that asserts nonsense with a steady voice.

Data Privacy and Data Sovereignty

For European organizations, Grok 4.7 is a sensitive candidate. The available cards show a calculated Sovereign Risk of HIGH, grounded in US jurisdiction under the CLOUD Act without any discernible EU safeguard. Concretely: US authorities can, under certain conditions, demand access to data — even if the provider claims organizational separation or classifies services differently.

The data location is listed as USA. A reliable data retention period is not specified; the value stands at -1 days, meaning no credible published retention limit. Particularly problematic: a GDPR DPA is not available according to the available data. For organizations that must operate in GDPR compliance, this is not a peripheral detail — it’s a practical procurement obstacle.

The weights provenance risk is rated MEDIUM. The reasoning is plausible: the weights are proprietary and not distributed, so the primary concern is not redistribution but the combination of US development, US hosting, and legal access. Anyone working with sensitive data should factor this in before signing a contract, not add it to the risk file afterward.

Conclusion

Grok 4.7 is an interesting but contradictory Frontier model. It brings agentic structure, strong CLI proximity, solid cultural adaptation, and some very good individual transformation performances. At the same time, it falls short of the final step too often in precisely the areas that would distinguish a great Thinking system: Reasoning remains visibly shallower than expected, documentation is too weak, the API shows sporadic dropouts, and language instructions flip under multiple simultaneous constraints faster than they should.

For DevOps-adjacent assistance, technical workflow planning, and content-complex transformations, Grok 4.7 is genuinely interesting — provided retries, oversight, and cost awareness are built in. For security reviews without human final sign-off, time-critical agent pipelines, or strictly regulated enterprise environments in Europe, the model in its current state is hard to recommend. Across all tests, no notable hallucinations — Grok 4.7 fails more on discipline than on imagination. That’s a better problem than megalomania, but for a Frontier product at this price point, it’s still a problem.

This evaluation was generated automatically based on the benchmark data. Model used: GPT-5.4 by OpenAI. The raw data and full methodology are documented in the GitHub project.