‘Grok-4.6’ Introduced as SpaceXAI Pushes AI Efficiency and Longer-Running Agentic Workflows

The pace of progress in large language models (LLMs) has settled into a steady rhythm of incremental leaps, each one measured against a growing set of specialized benchmarks that try to capture what it means for a system to stay useful across extended stretches of work. 

Labs continue to refine the balance between raw capability, cost of inference, and the ability to sustain multi-step reasoning without drifting or requiring constant human correction. 

But in the tense competition, models not only have to be smarter and better. The demands include the models becoming more efficient in what they do.

As a result of this, smaller upgrades can still shift practical workflows if they reduce the number of tokens or turns needed to reach a usable result.

SpaceXAI released Grok-4.6 as a direct successor to Grok-4.5

The update emphasizes longer agent trajectories and more demanding interactive or visual projects. 

Training involved a longer supplemental run that incorporated curated model-generated data focused on reasoning and advanced technical concepts, higher-quality engineering examples, and refinements to the optimizer and overall recipe. 

Supervised fine-tuning trajectories were regenerated using the previous model and filtered, followed by reinforcement learning across agentic tasks that included knowledge work, general coding, and domain-specific settings such as kernel optimization, web development, and computer-aided design. 

The resulting system is reported to remain engaged across many steps when researching unfamiliar topics, navigating codebases, or converting a high-level product idea into an initial working application, with some observation of increased self-checking before advancing.

Independent composite measures place the high-reasoning configuration of Grok-4.6 at 61 on the Artificial Analysis Intelligence Index, matching the score of OpenAI's powerful GPT-5.6 Sol, but still trailing from Anthropic's Fable 5 Max by a single point. 

On specific agent and knowledge-work evaluations the model shows gains relative to its predecessor: higher Elo on GDPVal-AA, improved CursorBench and FrontierCode scores, and better results on APEX-Agents. 

Performance remains mixed against the broader frontier on pure software-engineering suites such as DeepSWE and Terminal-Bench, where other systems continue to lead. 

Early community rankings on Arena's Code Arena WebDev leaderboard placed the high variant near the top cluster alongside recent competitors, though confidence intervals remain wide as more votes accumulate.

The model runs and was trained on Nvidia's GB300 NVL72 systems connected by NVLink. 

That rack-scale configuration pairs 72 Blackwell Ultra GPUs with Grace CPUs, delivering the dense compute, memory bandwidth, and interconnect performance required for both large-scale training and efficient inference at the reported token pricing. 

API access begins at two dollars per million input tokens and six dollars per million output tokens, unchanged from the prior release, with a faster variant available at twice the cost. 

The model is accessible immediately through Cursor, Grok Build, the SpaceXAI API, and several partner gateways; the first week includes doubled usage quotas in the two primary coding environments.

Elon Musk noted that a further iteration, Grok-4.7, is already in progress and expected within three to four weeks.

He said that Grok-4.7 will come with supplemental training that incorporates a substantial volume of SpaceX company data. 

Early user reports and third-party tests reflect the usual spread of reactions: some developers highlight the combination of speed, multi-step reliability, and cost relative to other frontier offerings, while others flag residual issues around consistency on certain agentic or security-sensitive tasks. 

As with previous releases, the practical value will be determined by how the model behaves across real workloads rather than any single leaderboard snapshot.

Published