Google Tests Gemini 4 'Carbon' Internally as Argon Launch Stalls: 'Feels Like Opus 5.5'

Google is internally testing a new Gemini 4 checkpoint codenamed Carbon, which employees say matches Anthropic's Opus 5.5 for coding, while the already-announced Argon model remains unavailable to most users despite strong benchmarks in long-context software engineering but weaknesses in terminal interaction.

ITPUB
ITPUB
ITPUB
Google Tests Gemini 4 'Carbon' Internally as Argon Launch Stalls: 'Feels Like Opus 5.5'

Google's Gemini 4 'Carbon' Emerges in Internal Testing as Argon Launch Stalls

According to a Business Insider report on October 9, Google employees have begun testing a new Gemini 4 checkpoint codenamed Carbon on the internal Jetski programming platform. One employee described the coding experience as "close to Claude Opus 5.5" , while another simply wrote "Carbon is really good." This internal buzz arrives while Google's officially announced flagship model Argon (released September 30) has yet to reach general users — it is only available through the Fairwind program to trusted security defenders, with paid API customers and Google AI Ultra subscribers listed as "next in line" and ordinary users told to wait longer.

Carbon Exposure: Internal Praise vs. External Availability Gap

Internal documents reveal at least three Gemini 4 derivative codenames under test: Argon (the announced flagship), Barium (whose Barium-B variant was ultimately released as Argon), and the newly surfaced Carbon . The lack of a fixed mapping between internal codenames and public names — demonstrated by Barium-B becoming Argon — means Carbon could ship as an Argon update, a standalone Gemini 4 variant, or something else entirely. The employee comparison to Opus 5.5 (Anthropic's model built for complex, long-horizon code agents) rather than the older Opus 5 suggests Google may be targeting the same agentic coding niche.

Argon's Real Performance: Strengths in Long-Context Engineering, Gaps in Terminal Interaction

Official benchmarks show Argon's strengths and weaknesses clearly:

DeepSWE v1.1 (long-cycle real-world software engineering): Argon 77.9% vs. Opus 5.5 74.2% and GPT-6 Astra 74.1% — Argon leads.

CWE-bench v1 (cybersecurity defense): Argon 68% , tied for first place.

Output token limit raised from 64k to ~100k , enabling much longer reasoning chains.

FrontierSWE v2 (building projects from scratch): Argon 57.4% — notably lower.

Terminal-Bench 4.0 (Linux terminal operation): Argon at the bottom of the pack; Opus 5.5 scores 66.4% .

An internal tester noted early Argon felt "roughly equivalent to Claude Opus 5" on some coding tasks. This pattern reveals Argon excels at deep reasoning over large codebases but lags in interactive, terminal-driven engineering workflows — exactly the gap Carbon's Opus 5.5 comparison hints Google is trying to close.

Tooling Strategy: Google's Real Moat Lies Beyond Model Codenames

Google is simultaneously building a full stack from model to execution environment to user entry points:

Antigravity 2.0 (desktop app, announced at I/O 2024): supports multi-agent parallel tasks, background scheduling, and model selection including Argon with 256K / 512K / 900K context tiers and low/medium/high thinking-difficulty settings ( TestingCatalog, Oct 9 ).

Gemini API Managed Agents : agents run in isolated Linux environments, use tools, execute code, and persist recoverable file state.

Gemini Agent (Google Cloud, announced recently): covers Q&A, task handling, content generation, and coding inside Workspace and other surfaces, with cross-device context retention.

The strategic logic: real software engineering is rarely "write a function" but rather understanding existing systems, modifying multiple files, running tests, and iterating on failures. The developer experience is the joint output of model + execution system. Google aims to differentiate with "system as product" rather than "model as product."

Conclusion: Persuasion Comes from Delivery, Not Codenames

Community reaction is split. Optimists see the rapid internal iteration (Carbon testing before Argon ships) as evidence Google has overcome internal hurdles. Skeptics demand usable products first: "Model not even usable, next version already hyped." The meaningful proof for Carbon will be reproducible, real-world codebase tests: cross-file editing, self-correction after failure, human revision effort, time and cost per task. Those metrics — not a new element codename — will show whether Google is truly closing the gap. Users don't need to know whether the backend runs Argon, Barium, or Carbon; they need to know the task they hand off is reliably completed.

References: Business Insider tweet https://x.com/BusinessInsider/status/2108643821802205478 | TestingCatalog article https://www.testingcatalog.com/gemini-4-argon-hints-emerge-as-google-tests-carbon-checkpoint/
Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

AI Agentsmodel evaluationGoogle AICarbonsoftware engineering benchmarksClaude Opus 5.5ArgonGemini 4
ITPUB
Written by

ITPUB

Official ITPUB account sharing technical insights, community news, and exciting events.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.