Products & Models · Oct 3, 2026
First third-party scores for Gemini 4 Argon: tied with Astra, fewest hallucinations, but more than twice the tokens
Argon's list price is one-fifth of Astra's, but using more than twice as many output tokens per task narrows the gap. To compare the cost of models with the same score, look at cost per task, not list price
Rie Suzuki · Technology Editor

Key points
- According to The Decoder and Trending Topics, Gemini 4 Argon scored 53 on Artificial Analysis's Intelligence Index, the same as GPT-6 Astra. It trailed leader Opus 5.5 (58) and Sonnet 5.5 (56), which contradicts Google's claim that it "outperforms Astra, Fable and Opus"
- Argon reportedly had the fewest hallucinations of the models compared. However, it reportedly used more than twice as many output tokens per task as Astra
- Argon's list price for output is one-fifth of Astra's. If it uses more than twice the output tokens, the difference in output cost shrinks to 2.5 times or less. Two models can have the same score and still cost very different amounts, depending on how many tokens they use
The first third-party evaluation of Gemini 4 Argon has reportedly been published. Google released the model on September 30 to cyber-defense professionals only. According to The Decoder and Trending Topics, Argon scored 53 on Artificial Analysis's Intelligence Index, tying OpenAI's GPT-6 Astra. It reportedly had the fewest hallucinations of the models compared. However, it reportedly used more than twice as many output tokens as Astra to solve each task. All of these figures come from media reports. This newspaper was not able to check Artificial Analysis's comparison page itself.
This article makes one point. Two models can have the same score and still cost very different amounts, depending on how they use tokens. Comparing list prices alone does not tell you what a model will actually cost.
What the reported evaluation shows
Argon's score of 53 puts it at the same level as Astra (about 53) and Fable 5.1 (53). On the same index, Opus 5.5 leads with 58, and Sonnet 5.5, released on September 28, is second with 56 (both previously reported). In its headline, The Decoder says Argon closes the gap with OpenAI and Anthropic but doesn't take a clear lead.
When Google launched Argon, it claimed the model outperforms GPT-6 Astra, Fable and Opus, according to reports by TechCrunch and 9to5Google. The third-party Intelligence Index does not support that claim. Argon tied Astra and Fable, and it scored 5 points below Opus 5.5. However, the secondary reports do not make clear whether the "Opus" Google compared against was Opus 5 or Opus 5.5, or which benchmark the claim referred to.
One-fifth the list price, more than twice the tokens
Argon's list price is $2 per million input tokens and $10 per million output tokens, the same price range as GPT-6 Sol and Sonnet 5.5. GPT-6 Astra costs $10 for input and $50 for output, so on list price alone Argon is one-fifth the cost.
That changes if Argon produces more than twice as many output tokens per task as Astra. Counting output costs alone, the fivefold difference shrinks to 2.5 times or less. If Argon uses three times the output tokens, the difference is about 1.7 times. Argon also has a large output limit of 1 million tokens, which lets it think and write at length. How much more does it write than Astra to reach the same score? That amount directly determines the bill.
Actual costs also depend on the number of input tokens and how well caching works. The reports do not say how much it cost in total to run the full index. Still, the direction is clear. Once the number of tokens is counted, a headline-style comparison like "the same score at one-fifth the price" does not hold up.
This is not the first time the same score has come with very different costs
Artificial Analysis's figures have shown this before. On the late-September index, Grok 4.7 and Xiaomi's MiMo-V2.6-Pro both scored 46. Yet running the index cost $4,967 for Grok 4.7 and $207 for MiMo-V2.6-Pro, a difference of about 24 times. Grok 4.7 used 240 million output tokens to complete the run (GPT-6 Astra at max used 27 million). In that case, the difference in token volume mattered more than the difference in list price.
AI labs have also started marketing their models on cost per task rather than list price. OpenAI itself reported that GPT-6.1 Sol costs $5.47 per task on Terminal-Bench Science, compared with $23.21 for Opus 5.5. Anthropic also promoted Sonnet 5.5 as "up to 30% cheaper per task" without changing its list price. Sol, Sonnet 5.5 and Argon all have exactly the same list price: $2 for input and $10 for output. The only difference among them is how many tokens each uses per task. On that comparison, Argon may come out behind the other two.
How to read "fewest hallucinations"
The finding that Argon had the fewest hallucinations has been reported as one of its strengths. But rates like this need to be read together with how often a model declines to answer. When GPT-6 Sol was released on September 22, its hallucination rate fell from 92% to 60%. Part of that drop, however, came from the share of questions it tried to answer falling from 99% to 83%. Is Argon correctly saying "I don't know" when it doesn't know, or is it simply answering fewer questions? The reports do not yet say.
The link between Argon's lower hallucination rate and its high token use has not been confirmed either. One possibility is that it makes fewer mistakes because it thinks longer and checks its work. If so, it is trading cost for accuracy. Users have to decide how much they will pay for accuracy, and to make that decision they need token counts as well as scores and rates.
The numbers users should look at
For now, Argon is available only through the Fairwind Program. Google has not said when a general API or Ultra will be released. Most users cannot yet measure the costs themselves. That makes this third-party evaluation, published before general availability, all the more important for deciding.
An agent's cost comes from multiplying three things: the model's list price, the number of tokens used per task, and how well the harness reuses context. Astra and Argon both score 53, but their list prices are reportedly 5 to 1 and their token use 1 to 2 or more. Only by multiplying the two can you see what your own work will cost. Put cost per task and output token counts next to the leaderboard scores. That extra step is what choosing a model requires.
Sources: The Decoder, Trending Topics (both media reports; Artificial Analysis's primary data not confirmed)
Editorial cartoon
