Grok 4.7

Grok 4.7 is xAI’s Frontier flagship as of September 21, 2026, built on a larger base than Grok 4.6 and trained with a focus on multi-hour tasks. 500,000 tokens of context, text and image input, four reasoning levels from low to xhigh at unchanged pricing of 2 / 6 USD per million tokens. According to xAI, nearly doubled performance on long terminal tasks and a new safeguard stack against jailbreaks.

xAI Version 4.7 Commercial use restricted Dense 500 K Context 05/2026 $2 / $6 per 1M

  • Proprietary
  • Frontier
  • xAI
  • Text
  • Vision
  • Interactive

Sovereign Risk: MEDIUM The model is developed and hosted by a US-based company. Due to US jurisdiction, it is potentially subject to the CLOUD Act, which represents a moderate risk of data access by US authorities. Since the weights are proprietary and not distributed, there is no additional risk from disclosure of the weights themselves.

Key metrics

Score · Latency · Cost · Quality

Total Score Silver
71.06
Routine
43.11
Reasoning
27.95

Rank #71

LLM Judge Avg
3.5
100 Coverage
Avg Task Duration
38.77
Interactive
Token Rate
17.39
Output Rate
P95 Latency
149.17
Top 5 %
Total Tokens
116600
Output Volume
Cost per 1K
$0.006
USD / 1K Requests
Benchmark Cost
$0.7
Total · 116600 tok

Benchmark modules

10 modules · weighted · vs. model median & top performer

Grok 4.7 Best model Ø All models
Code Quality 69.55
CLI Benchmark 90.67
Logical Reasoning 58.96
UX Writing 73.19
Documentation 62.08
Content Transform. 74.17
Cultural Intelligence 77.84
Synthesis Quality 53.33
Tool Execution 90
ToolUse Score 71.5
Benchmark Cost $0.7

Token efficiency & latency

Consumption per module vs. model median

Token consumption per module

Performance profile