Google Tests Gemini 4 'Carbon' Internally as Argon Launch Stalls: 'Feels Like Opus 5.5'
Google is internally testing a new Gemini 4 checkpoint codenamed Carbon, which employees say matches Anthropic's Opus 5.5 for coding, while the already-announced Argon model remains unavailable to most users despite strong benchmarks in long-context software engineering but weaknesses in terminal interaction.
Google's Gemini 4 'Carbon' Emerges in Internal Testing as Argon Launch Stalls
According to a Business Insider report on October 9, Google employees have begun testing a new Gemini 4 checkpoint codenamed Carbon on the internal Jetski programming platform. One employee described the coding experience as "close to Claude Opus 5.5" , while another simply wrote "Carbon is really good." This internal buzz arrives while Google's officially announced flagship model Argon (released September 30) has yet to reach general users — it is only available through the Fairwind program to trusted security defenders, with paid API customers and Google AI Ultra subscribers listed as "next in line" and ordinary users told to wait longer.
Carbon Exposure: Internal Praise vs. External Availability Gap
Internal documents reveal at least three Gemini 4 derivative codenames under test: Argon (the announced flagship), Barium (whose Barium-B variant was ultimately released as Argon), and the newly surfaced Carbon . The lack of a fixed mapping between internal codenames and public names — demonstrated by Barium-B becoming Argon — means Carbon could ship as an Argon update, a standalone Gemini 4 variant, or something else entirely. The employee comparison to Opus 5.5 (Anthropic's model built for complex, long-horizon code agents) rather than the older Opus 5 suggests Google may be targeting the same agentic coding niche.
Argon's Real Performance: Strengths in Long-Context Engineering, Gaps in Terminal Interaction
Official benchmarks show Argon's strengths and weaknesses clearly:
DeepSWE v1.1 (long-cycle real-world software engineering): Argon 77.9% vs. Opus 5.5 74.2% and GPT-6 Astra 74.1% — Argon leads.
CWE-bench v1 (cybersecurity defense): Argon 68% , tied for first place.
Output token limit raised from 64k to ~100k , enabling much longer reasoning chains.
FrontierSWE v2 (building projects from scratch): Argon 57.4% — notably lower.
Terminal-Bench 4.0 (Linux terminal operation): Argon at the bottom of the pack; Opus 5.5 scores 66.4% .
An internal tester noted early Argon felt "roughly equivalent to Claude Opus 5" on some coding tasks. This pattern reveals Argon excels at deep reasoning over large codebases but lags in interactive, terminal-driven engineering workflows — exactly the gap Carbon's Opus 5.5 comparison hints Google is trying to close.
Tooling Strategy: Google's Real Moat Lies Beyond Model Codenames
Google is simultaneously building a full stack from model to execution environment to user entry points:
Antigravity 2.0 (desktop app, announced at I/O 2024): supports multi-agent parallel tasks, background scheduling, and model selection including Argon with 256K / 512K / 900K context tiers and low/medium/high thinking-difficulty settings ( TestingCatalog, Oct 9 ).
Gemini API Managed Agents : agents run in isolated Linux environments, use tools, execute code, and persist recoverable file state.
Gemini Agent (Google Cloud, announced recently): covers Q&A, task handling, content generation, and coding inside Workspace and other surfaces, with cross-device context retention.
The strategic logic: real software engineering is rarely "write a function" but rather understanding existing systems, modifying multiple files, running tests, and iterating on failures. The developer experience is the joint output of model + execution system. Google aims to differentiate with "system as product" rather than "model as product."
Conclusion: Persuasion Comes from Delivery, Not Codenames
Community reaction is split. Optimists see the rapid internal iteration (Carbon testing before Argon ships) as evidence Google has overcome internal hurdles. Skeptics demand usable products first: "Model not even usable, next version already hyped." The meaningful proof for Carbon will be reproducible, real-world codebase tests: cross-file editing, self-correction after failure, human revision effort, time and cost per task. Those metrics — not a new element codename — will show whether Google is truly closing the gap. Users don't need to know whether the backend runs Argon, Barium, or Carbon; they need to know the task they hand off is reliably completed.
References: Business Insider tweet https://x.com/BusinessInsider/status/2108643821802205478 | TestingCatalog article https://www.testingcatalog.com/gemini-4-argon-hints-emerge-as-google-tests-carbon-checkpoint/
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
ITPUB
Official ITPUB account sharing technical insights, community news, and exciting events.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
