GPT-6.1 Sol Benchmarked: 3x Less Code for SVG but Higher Cost and Latency

Community tests of OpenAI's GPT-6.1 Sol reveal it writes 159 lines for an SVG pelican versus Sonnet 5.5's 517, matches Opus 5.5 in Blender carrier building, but fails at Three.js Eiffel Tower, while costing $5.00 and taking 99 minutes for three racing tracks compared to GPT-6 Sol's $3.90 and 78 minutes.

PaperAgent
PaperAgent
PaperAgent
GPT-6.1 Sol Benchmarked: 3x Less Code for SVG but Higher Cost and Latency

OpenAI announced GPT-6.1 Sol, claiming near-Astra intelligence at one-fifth the price and the best cost-performance ratio at its performance tier. The community immediately stress-tested the model across several coding challenges.

SVG Pelican on a Bicycle: 159 Lines vs 517 Lines

A popular frontend benchmark asked models to create a cycling pelican using pure SVG and vanilla JavaScript. GPT-6.1 Sol delivered a 159-line, 14.9 KB solution featuring a low-saturation seaside illustration style with coordinated pedaling, blinking, and wind effects, plus a “reduce motion” pause toggle. Sonnet 5.5 produced a 517-line, 23.4 KB version with six parallax layers, a helmet, scarf, chain animation, and a speed slider. While aesthetics are subjective, GPT-6.1 Sol completed the requirement with roughly one-third the code, implying a smaller maintenance surface and lower inference cost.

SVG pelican comparison
SVG pelican comparison

Three Racing Tracks: Compliance vs Cost and Speed

The community scaled the test to engineering level: each of three GPT generations (GPT-6 Sol, GPT-6.1 Sol, GPT-5.6 Sol) built three “extreme style” tracks — Dolomites mountain hairpins, cliffside azure coast, rainy neon elevated highway — from a single prompt. Models first planned storyboards and file lists, then generated each file in a separate request, pure code without external assets. All nine tracks rendered successfully.

Aggregated Metrics (Three Tracks Combined)

GPT-6 Sol: Total cost $3.90, total time 1h 18m, output tokens 286,519, code lines 18,605

GPT-6.1 Sol: Total cost $5.00, total time 1h 39m, output tokens 358,667, code lines 22,009

GPT-5.6 Sol: Total cost $5.63, total time 1h 11m, output tokens 360,684, code lines 25,143

GPT-6.1 Sol adhered most closely to the spec — seven hairpin turns on the mountain track, wettest surface on the rainy track — yet it was the most expensive and slowest. GPT-6 Sol was cheapest overall. Token consumption for GPT-6.1 Sol did not decrease versus its predecessor, contradicting the “more efficient” narrative: capability rose but the bill did not shrink.

Racing tracks comparison
Racing tracks comparison

3D Scenes: Carrier Success, Eiffel Tower Failure

In Blender, GPT-6.1 Sol and Opus 5.5 each built an aircraft carrier from keel to bridge, producing layered construction timelapses — a draw, proving GPT-6.1 Sol has genuine 3D workflow competence beyond web animations.

However, a Three.js Eiffel Tower test exposed a weakness: GPT-6.1 Sol performed poorly compared to other models on this common benchmark. The tester noted that while other tasks were acceptable, this standard comparison task failed. The conclusion: GPT-6.1 Sol excels at “small and complete” lightweight scenarios, but struggles when complexity and interaction density increase. Cost-effectiveness flagship does not equal all-round flagship.

Blender carrier build
Blender carrier build
Three.js Eiffel Tower failure
Three.js Eiffel Tower failure

Bug Fixing and Historical Context

Another test threw 105 bugs at GPT-6.1 Sol for repair; results are shown in the accompanying image. The article recalls that the previous GPT-6 Sol was essentially a weakened re-skin of GPT-5.6 Terra — lower price but also reduced capability. The author invites more community benchmarks to confirm whether GPT-6.1 Sol represents genuine progress or another incremental squeeze.

105 bug test results
105 bug test results
Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

code generationSVGAI benchmarkOpenAIThree.jsBlendercost analysisGPT-6.1 Sol
PaperAgent
Written by

PaperAgent

Daily updates, analyzing cutting-edge AI research papers

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.