How Google's Gemini 4 Argon Shows Astonishing Scores, and the Disappointing Cut for Everyone Who Cannot Use It

Google is about to make the Gemini app a lot thinner for anyone who is not already on a higher paid plan. 

In a support page, Google said that from October 9, personal accounts without a subscription keep only Flash-Lite. Flash and the limited Pro access those accounts still get both go. Google AI Plus, the $4.99 plan that also doubles usage and adds storage, keeps Flash-Lite and Flash but loses Pro. 

Plus subscribers are supposed to get an email saying when that cut reaches their account, so there is no single date for them yet. AI Pro and AI Ultra keep all three models. Deep Think, which had been reserved for the most expensive tiers, moves down to AI Pro. 

The notice covers the Gemini app on personal accounts. It does not, on the published page, change developer API access or workplace accounts.

The timing is awkward. 

The change landed days after Google showed Gemini 4 Argon, the model it is still testing with U.S. government agencies and a small set of partners rather than ordinary users. 

Argon is not what free or Plus accounts are losing. 

What they are losing is the everyday stack: Gemini 3.6 Flash and varying access to Gemini 3.1 Pro on the free tier, and Pro on Plus. Even so, the sequence reads as a company that just published its strongest numbers and then pulled the better consumer models behind a higher bill.

On those numbers, Argon is not a small step. 

In Google's own evaluation table it leads or ties for first in 13 of 18 benchmarks, ahead of GPT-6 Astra and Claude Opus 5.5 by that count. The wins cluster in the work companies actually pay for. 

On AutomationBench, Zapier's test of end-to-end business tasks, Argon scores 51.3%, against 42.5% for Opus 5.5 and 41.4% for Astra. On Vals Finance Agent v2 it scores 65.4%, almost seven points clear of the nearest Claude model. On DeepSWE v1.1, long software jobs drawn from real projects written after each model's training cutoff, it scores 77.9%, ahead of Opus 5.5 at 74.2% and Astra at 74.1%. Harvey's legal-agent test is even more lopsided on Google's sheet, 19.6% against 5.4% for Astra and 3.8% for Opus 5.5. Graph traversal over long context and long-video understanding go the same way. 

Several of those rival figures come from the other labs' own reports or public boards, and Argon's DeepSWE number is self-computed, so the table is not an independent sweep. Even with that caveat, the spread is hard to dismiss. 

On knowledge work and on that particular coding test, Argon is not trading blows. It is clear of the field.

SpaceXAI's strongest public model is Grok 4.7, released September 21 and sold as its best coding and knowledge-work system, at $2 per million input tokens and $6 per million output tokens. Argon's introductory price is the same $2 on input and $10 on output, so the two are close on cost and far apart on who can actually call them. Grok 4.7 is already on the API. Argon is still inside the Fairwind Program. 

But on the overlapping tests, Argon is ahead, not tied. 

SpaceXAI put Grok 4.7 at 71.0% on DeepSWE v1.1 at high effort, behind GPT-5.6 Sol at 72.7% and just ahead of Fable 5.1 at 70.0%. Argon's 77.9% on the same benchmark clears that whole group, though Google computed its own score and Grok's figure is the company's too. On Terminal-Bench 4.0, Grok 4.7's published 38.0% sits well under Argon's 57.4%, and both sit under Opus 5.5 at 66.4%. The Harvey legal-agent number is a coincidence of headlines: both labs print 19.6%, Argon against Astra and Opus, Grok 4.7 against GPT-5.6 Sol and Fable 5.1, so it is the same test name and not a head-to-head. 

Artificial Analysis puts a wider gap on its own index, 53 for Argon, level with GPT-6 Astra, against 46 for Grok 4.7. 

Grok still beats the cheaper OpenAI and Anthropic models on several of SpaceXAI's rows, including electrical engineering and that legal test. It does not reach the sheet Google published for Argon.

It is not clear everywhere. 

Across the agentic coding rows it wins two and finishes last on two others, FrontierSWE v2 and Terminal-Bench 4.0. On FrontierSWE it scores 55.0%, about ten points behind Astra. On Terminal-Bench 4.0 it scores 57.4%, well behind Opus 5.5 at 66.4%. Independent coding checks from Vals have also handed Claude 5.5 narrow edges on vibe coding and code migration. 

The chart that makes Argon look dominant is real, and so are the holes in it.

That gap between the chart and the product is why a lot of people still call Gemini disappointing. 

The model that leads those tables is not the model most people can open. Argon is still a limited preview. The app that ordinary users get has a different reputation, and the complaints are specific. In side-by-side writing tests it keeps falling back on sterile phrasing, symmetrical bullets, and polite corporate filler, the voice of a template rather than a person who has to cancel a service or push back on a bill. 

Asked to catch a planted wrong fact, it can flag the error and then add almost nothing when pressed, while Claude goes and fetches another source. 

People who expected Google's search index to make Gemini the strictest checker keep walking away from that exchange. Usage is part of the sour taste too. Google copied the five-hour allowance window that already frustrates Claude users, and a wrong or useless answer still spends the allowance. For someone on the free tier after October 9, that allowance is attached to Flash-Lite only, the smallest model in the set.

The sharper complaint this week is about the plan itself. Plus is a paid tier. People bought it, some of them on a steep introductory discount, partly because Pro was included. Removing a model from a subscription already sold is not the same thing as charging more for a new model that costs more to run. 

Flash may still be enough for a lot of daily work. The only honest test is to give Flash the same document, the same bug, or the same research question users were giving Pro, and see whether the answer still holds. If it does not, the plan they renew is a different product from the one they signed up for.

Published