'Grok-4.7' Arrives With Longer-Running Agents and Modest Gains Over Grok-4.6 at the Same Price

Model version numbers now land the way phone chips used to. A new decimal shows up on a weekday afternoon, the lab posts a chart and a demo, an independent board posts another chart, and people who actually ship software spend the next few days deciding whether anything in their stack should move. 

'Grok 4.7' thread followed that rhythm. 

The SpaceXAI's announcement kept the claim small on paper. Same list price as Grok 4.6, announced less than two months earlier. 

Same advertised serving speed. A larger base model underneath, trained with a longer reinforcement learning pass on tasks that can run for hours. 

The public pitch was not a new context window or a new rate card. It was duration and checking. 

The model is said to stay on hard jobs longer, inspect its own output more carefully, and ship with a thicker refusal stack than 4.6. 

A follow-up put two clips next to each other, 4.7 first and 4.6 second, both asked to build an open world city game. Availability was listed as Cursor, Grok Build, and the Grok API.

Input is still $2 per million tokens. Output is still $6. Cached input is still discounted to $0.50. The context window is still 500,000 tokens. Text and image go in, text comes out. Reasoning effort still runs the same. 

What did change was the kind of work the model is expected to finish without a person in the loop at every step. 

On the lab's own tables, results improved on all fronts. 

Artificial Analysis put 4.7 at 46 on its Intelligence Index, two points above 4.6, and said that score now places the lab among the top four on that particular board. 

Another report suggests that 4.7 landed at 1657 Elo, or 111 above 4.6 at high effort.

This makes it close to Claude Opus 5 and Claude Fable 5.1. 

Paired with Grok Build, the first-party coding harness, the Coding Agent Index rose from 47 to 56 and passed GPT-5.6 Sol, but still behind Fable 5.1, GPT-6 Astra, and Opus 5.

However, those gains are not even across the index. 

Outside agentic knowledge work, 4.7 broadly matches 4.6 high on several remaining tasks, while raw accuracy stayed almost flat. 

Not a surprise since version increment behind the decimal means minor gains and not a uniform lift in every exam.

There is also the usual fog around size. 

Reporting after the launch repeated an earlier remark that 4.7 is about 2.1 trillion parameters against 1.5 trillion for 4.6. The official card did not lock that figure in. 

Pretraining is described as running through June 2026 with later supplemental data. Some coverage mentioned extra engineering text from SpaceX. None of that is something a buyer can verify from the API. What a buyer can verify is the harness. 

Results on Grok Build are not the same object as results on a shared evaluation harness. Native agents inflate or deflate scores depending on tools, retries, and how much of the job is the model versus the wrapper.

Safeguards were listed as a first-class change, not a footnote. The lab said 4.7 is more willing to refuse dual use prompts in cybersecurity and biology while still completing ordinary professional tasks. Vendor benches put a low pass rate on risky hacking prompts and a leading score on a biosafety set. 

Independent users will treat those numbers as a starting point. Refusal rates are easy to game in a launch post and hard to judge until the model is in a real ticket queue.

Summarizing everything, 4.7 is a minor improvement over 4.6, with gains that aren't even across the board. 

For those already on 4.6, the practical test is not whether the lab entered a top-four list. It is whether a multi-hour coding session or a document that has to survive a review is completed better, more consistently, and with less intervention. 

If 4.7 can do that often enough, the upgrade may matter. If not, the benchmark gains alone are unlikely to change much for existing users.

Published