GPT-5.6 Prices Slashed Up to 80% – What the New Costs Mean for Developers
OpenAI has cut the API prices of its GPT‑5.6 Luna, Terra and Sol models by up to 80%, introduced a faster "Fast mode" for Sol, and outlined how the lower costs reshape model selection, usage patterns, and overall AI budgeting for developers.
OpenAI announced a major price reduction for its GPT‑5.6 family, with discounts of up to 80% for the entry‑level Luna model. Input token price for Luna dropped from $1 to $0.2 per million, and output from $6 to $1.2. Terra’s input and output prices fell 20% to $2 and $12 respectively, while the flagship Sol model kept its list price but added a "Fast mode" that runs 2.5× faster without extra charge.
The new pricing table is:
GPT‑5.6 Luna : $0.2 per million input tokens, $1.2 per million output tokens.
GPT‑5.6 Terra : $2 per million input tokens, $12 per million output tokens.
GPT‑5.6 Sol : $5 per million input tokens, $30 per million output tokens.
Luna is positioned for cost‑sensitive workloads that need tool use, long context handling, and multi‑step workflows, effectively offering agent‑capable capabilities at the price of older, simpler models. Terra targets everyday tasks with a modest 20% price cut, while Sol remains the high‑performance option; its Fast mode delivers up to 2.5× speed for double the standard price, replacing the previous Priority Processing tag.
OpenAI states that the price adjustments also apply to Codex and ChatGPT Work, but subscription fees and total quota limits stay unchanged; the reduced per‑call cost means the same subscription can run more tasks. The company also switched the Auto‑review model in ChatGPT and Codex CLI from GPT‑5.4 to GPT‑5.6 Luna, extending cheaper model usage to code review and other high‑frequency agent tasks.
According to OpenAI, the combined effect of the model upgrade and Luna’s price cut should cut overall costs to roughly one‑tenth of previous levels, allowing cheaper models to handle not only classification or extraction but also code review and backend monitoring that require reasoning and tool calls.
Model selection guidance:
Use Luna for budget‑sensitive, high‑volume tasks.
Assign ordinary daily work to Terra.
Reserve Sol for high‑difficulty tasks; pay double for Fast mode if speed is critical.
Typical Q&A from the announcement: Q: Should I choose Opus 5 or GPT‑5.6 Sol? A: Opus 5 fits product development, discussion, front‑end experience, and documentation; GPT‑5.6 Sol is steadier for back‑end rigor, code review, debugging, and long‑running unattended jobs. Q: Is it worth keeping Opus 5’s Max tier on? A: Generally no; the Extra tier suffices for most tasks. Max is only justified for a few high‑risk inference jobs and can increase cost and output divergence. Q: Why does Opus 5 tend to broaden task scope? A: It is more proactive and fills in missing details, which helps uncover blind spots but may over‑extend small tasks when boundaries are vague. Q: How should teams route models? A: Assign discussion and creation to Opus 5, review and long‑running execution to GPT‑5.6 Sol, then track cost, latency, failure rate, and manual rework. Q: How to use Claude and GPT together in China? A: Use a unified gateway such as Code80 to expose both models, then configure task‑type and permission policies in development tools.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Top Architecture Tech Stack
Sharing Java and Python tech insights, with occasional practical development tool tips.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
