Gemini 4 Argon Matches GPT-6 Astra Benchmarks with Lowest Hallucination Rate

Google DeepMind's Gemini 4 Argon matches GPT-6 Astra on key benchmarks while achieving a 15% hallucination rate—the lowest among peers—with a 1M token output limit, aggressive introductory pricing, and a security-first rollout to network defenders.

AI Engineering
AI Engineering
AI Engineering
Gemini 4 Argon Matches GPT-6 Astra Benchmarks with Lowest Hallucination Rate

Overview

Google DeepMind released Gemini 4 Argon, its first non-Flash frontier model in over seven months. Announced by Logan Kilpatrick, the model is initially available to trusted network defenders via the Fairwind Program before expanding to developers, enterprises, and consumers. Paid API and AI Ultra subscribers receive priority access.

Benchmark Results

Argon demonstrates competitive performance across multiple benchmarks:

DeepSWE v1.1 (software engineering): 77.9%, surpassing GPT-6 Astra's 74.1%.

Vals Index : 68.9%, ranking first.

LVBench (long video understanding): 91.7%.

AutomationBench (enterprise workflow automation): 51.3%, leading Claude Opus 5.5 by ~9 percentage points.

Artificial Analysis Intelligence Index : 53, tying GPT-6 Astra (max 53) and leading GPT-6.1 Sol (max 52) by 1 point. This is 23 points higher than Google's previous non-Flash model, Gemini 3.1 Pro Preview.

Independent evaluator Artificial Analysis also measured average output tokens per Intelligence Index task: Argon produces 62K tokens versus Astra's 27K, indicating Argon relies on higher token volume rather than efficiency.

Output Token Limit and Pricing

Argon raises the output token ceiling from 64K to 1M while retaining a 1M context window and multimodal capabilities. Google states the model can generate hundreds of thousands of tokens in a single inference, enabling long-horizon tasks without interruption.

Pricing is split into an introductory period and standard rates:

Introductory: $2 per million input tokens, $10 per million output tokens; cached input receives a 5% discount.

Post-introductory: $4 per million input, $20 per million output.

At introductory pricing, Argon's per-task cost is $1.99 (60% of Astra's $3.26). After the promotion ends, the cost rises to $3.98, making it ~1.2x more expensive than Astra.

Hallucination Rate and Accuracy Trade-offs

Argon's standout feature is its hallucination rate of 15% on the AA-Omniscience evaluation—the lowest among peer models. GPT-6 Astra scores 51%, GPT-6.1 Sol 54%. This suggests Argon prefers to respond "I don't know" rather than fabricate answers.

However, accuracy is 50%, which is 5 points below Gemini 3.1 Pro Preview and 13 points below GPT-6 Astra. Despite lower accuracy, the overall AA-Omniscience score is 42, on par with Astra (43) and Sol (42). The trade-off is favorable for security tools where fabricated answers are dangerous, but may be a drawback in scenarios requiring high factual recall.

Agent Capabilities

Historically a weak point for the Gemini series, agent performance shows marked improvement:

AutomationBench-AA : 77.5%, leading Claude Sonnet 5.5 (max 71.3%) by 6 points.

Terminal Bench 4 : 57%, a 53-point jump over Gemini 3.1 Pro Preview, though still behind Claude Sonnet 5.5 (64%), Claude Opus 5.5 (60%), and GPT-6 Astra (59%).

AA-Briefcase : 1494 Elo overall. Rubric pass rate of 65% is the highest recorded by Artificial Analysis, but analysis quality (1576 Elo) and presentation quality (1308 Elo) are relatively low.

Vals AI tested Argon across 22 benchmarks, with 20 placing in the top five. Argon achieved perfect scores on IOI 2024, 2025, and 2026—a feat matched only by GPT-6 Astra. On Vibe Code Bench, Argon perfectly built 30 applications, exceeding Claude Opus 5 (25) and GPT-6 Astra (24).

Internal Google Use Cases

Google is already deploying Argon internally:

Quantum computing : Optimized spatiotemporal resources of subroutines, improving a published baseline by 40% within minutes.

Datacenter telemetry : Agents autonomously analyzed telemetry and implemented memory optimizations, freeing 300+ TiB with projected total savings of 500 TiB to 1 PiB.

Code migration (C/C++ to Rust) : Migrating codebases ranging from tens of thousands of lines (re2, libgav1) to 800K lines (Fuchsia Zircon kernel). In the libgav1 Rust port, Argon replaced 32K lines of SIMD code with profile-guided experiments, yielding a decoder 2.7x faster than the original Rust version.

Security Focus and Safety Mechanisms

Cybersecurity is the central narrative of this release. Argon is trained to autonomously discover, verify, and patch vulnerabilities. Through the Fairwind Program, it is first provided to trusted defenders. Wiz, using Argon in its Scan for Good initiative, uncovered a high-severity vulnerability in global hospital software that previous frontier models missed. On the CWE-bench v1 vulnerability remediation benchmark, Argon scores 68%, tying for first place.

Google lists four safety mechanisms: refusal of malicious requests, resistance to indirect prompt injection, monitoring for model drift, and sandbox hardening. While not novel, Google claims Argon is the "most robust" generation to date.

Rollout Strategy

Access follows a phased order: network defenders first, then developers, enterprises, and consumers. Paid API and AI Ultra subscribers are prioritized within each phase.

Critical Analysis

The article highlights three caveats:

Token volume over efficiency : Argon completes tasks by generating 2-3x more tokens than competitors; introductory pricing masks this cost structure.

Low hallucination via caution : The 15% hallucination rate comes from a conservative stance—preferring silence over fabrication—which reduces accuracy. This is advantageous for security tooling but may limit utility in other domains.

Security-first rollout : Prioritizing defenders over general users signals that security is the primary product thesis, not an afterthought.

Benchmarks remain proxies; real-world validation will determine whether Argon's paper advantages translate into practical utility.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

model evaluationcybersecurityAI benchmarksGoogle DeepMindagent capabilitieshallucination rateGPT-6 AstraGemini 4 Argon
AI Engineering
Written by

AI Engineering

Focused on cutting‑edge product and technology information and practical experience sharing in the AI field (large models, MLOps/LLMOps, AI application development, AI infrastructure).

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.