Gemini 4 Pro Leaks as 'gemini-3.8-flash' on Arena, Beats Astra and Fable Across Benchmarks
Google's unreleased Gemini 4 Pro model appears anonymously on LMSYS Chatbot Arena as 'gemini-3.8-flash', reportedly surpassing OpenAI's GPT-6 Astra and Anthropic's Claude Fable 5.1 across coding, agent, reasoning, and computer-use benchmarks while offering competitive pricing and advanced features like 10M token context and persistent memory.
Over the past few days, a new model named gemini-3.8-flash quietly appeared on the LMSYS Chatbot Arena. Developers and AI researchers who tested it quickly concluded this is almost certainly Google DeepMind's next-generation flagship Gemini 4 Pro , disguised under a Flash-series name. The previous Gemini 3.8 Flash launched in early September, so a second model bearing the exact same name is highly unusual.
Benchmark Results Show Broad Superiority
A widely circulated benchmark screenshot indicates Gemini 4 Pro achieves comprehensive leads over GPT-6 Astra and Claude Fable 5.1:
DeepSWE v1.1 (AI agent coding tasks): 88% — nearly 2 points above Astra.
GDPval-AA v2 (real-world knowledge work): 2064 Elo — the only model above 2000.
Terminal-bench 2.1 (terminal coding): 95.3% — highest score.
OSWorld-2.0 (computer-use capability): 86.8% — beats both Astra and Fable.
If these numbers hold, Gemini 4 Pro becomes the current “strongest model on earth” across agentic coding, reasoning, and computer-control tasks.
Pricing and Technical Specifications
Pricing is reported at $2.25 per million input tokens and $11.25 per million output tokens , making it the most cost-effective option among the “big three” frontier models. Researcher Qwinah discovered the backend exposes a 10 million token input limit , 256k token output limit , permanent cross-session memory , and built-in web access without requiring an API .
Hands-On Demos Reveal Step-Change in Creative & 3D Capabilities
Developers shared extensive real-world tests:
UI/UX Design : In 14 minutes, Gemini 4 Pro built a “creative showcase” page with a “scroll-as-brush-stroke” interaction where lines deepen smoothly on scroll.
Cyberpunk/Retro-Future Worksite : Combined 3D grid, data dashboards, and an interactive 3D model on the hero screen.
Classic “Pelican on a Bicycle” SVG/3D Test : Output includes color theming, day/night toggle, headlights, anatomical annotations, and cadence control — completion quality matches top-tier models.
Pixel-Art 3D Pagoda : Immersive, explorable environment.
Airbus H145 Helicopter : Full 3D model generated in ~10 minutes.
3D Flight Simulator : Visual fidelity far exceeds the public Gemini 3.8 Flash, confirming the Arena model is a different, more capable system.
Game Generation : Produces playable games with complete interaction logic and polished visuals — e.g., a Minecraft-style voxel builder and a modern 3D kart racer — in minutes.
Developer Pankaj Kumar summarized: “Speed is very fast; SVG and 3D generation noticeably improved; complex requirements understood in a single prompt.”
Recursive Self-Improvement (RSI) as the Strategic Driver
The article connects the leap to Recursive Self-Improvement (RSI) . Google DeepMind Chief Strategist Jasjeet Sekhon recently stated at a Berkeley event that RSI has become a key pillar of AI investment logic. If the RSI loop sustains, model improvement could accelerate beyond exponential curves. Rumors suggest Gemini 4 finished pre-training early precisely because DeepMind closed the RSI loop during training. On 14 October, Google publicly released the Dream-RSI research, which explores how agents can continuously refine their own search strategies from experience — confirming DeepMind’s public commitment to RSI.
Competitive Context: The “Big Three” Race
OpenAI fields Astra ; Anthropic fields Fable 5.1 . Both have pushed competition beyond chat into long-horizon tasks, agents, code, reasoning, and autonomous research. Google’s Gemini 3.5 Pro was cancelled (WSJ reported it “couldn’t even beat Flash”), so Gemini 4 Pro must re-prove Google’s frontier leadership to defend its $4.23 T market cap. The anonymous Arena debut appears to be a live checkpoint test before a formal launch.
Reference
https://x.com/HarshithLucky3/status/2100524699780600136Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
IT Services Circle
Delivering cutting-edge internet insights and practical learning resources. We're a passionate and principled IT media platform.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
