Cut Codex Costs: Assign Strong Model for Planning, Cheap Models for Execution & Review
This tutorial explains Codex's internal model call chain—planning, execution, review, and context compression—and shows how to configure separate models for each role via config.toml and worker.toml, using a powerful model for planning while delegating coding and review to cheaper alternatives, permanently reducing per-task API costs.
Codex is a powerful AI coding agent, but its per-task cost can be high because a single request actually triggers a chain of model calls. The article breaks down this chain into four distinct roles:
Main model : understands the goal, plans the task, coordinates execution, and decides the next step. Errors here propagate downstream, so this role needs strong reasoning.
Sub-agent : performs concrete actions—editing files, writing code snippets, running tests. Tasks are already well-defined, so top-tier reasoning is not required.
Review model : checks code for issues and suggests fixes. The scope is narrow, making it a good fit for a cheaper model.
Compact : compresses context when it grows too long. This step currently cannot be assigned a separate model; it follows the main model.
The core insight: use a high-capability model only for planning, and offload execution and review to lower-cost models . Once configured, Codex applies this division of labor automatically for every task.
Step 1: Configure Main Model and Review Model
Edit the global configuration file:
macOS/Linux: ~/.codex/config.toml Windows: C:\Users\<username>\.codex\config.toml Add the following entries (example uses OpenAI's gpt-5.6 series):
model = "gpt-5.6-sol"
model_reasoning_effort = "high"
review_model = "gpt-5.6-luna" modelis the main planner; model_reasoning_effort = "high" enables deeper reasoning. review_model handles code review with a lighter model.
Step 2: Configure Sub-agent Model
Create a worker configuration file at ~/.codex/agents/worker.toml (create the directory if needed) with:
model = "gpt-5.6-luna"
model_reasoning_effort = "medium"The sub-agent inherits the session's provider, so the model must be supported by the same provider. Using gpt-5.6-luna with medium reasoning effort cuts cost for the high-volume coding and test-running steps.
Using a Third-Party Provider (Volcengine Agent Plan)
If you prefer a non-OpenAI provider that supports the Responses API, you can define a custom provider in config.toml:
model = "<strong-model-in-plan>"
review_model = "<lightweight-model-in-plan>"
model_provider = "volcengine-agent-plan"
model_supports_reasoning_summaries = true
[model_providers.volcengine-agent-plan]
name = "volcengine-agent-plan"
base_url = "https://ark.cn-beijing.volces.com/api/plan/v3"
env_key = "ARK_API_KEY"
wire_api = "responses"Then in worker.toml set the sub-agent to the plan's lightweight model:
model = "<lightweight-model-in-plan>"
model_reasoning_effort = "medium"Note: Agent Plan requires a dedicated API key (not a standard Volcengine Ark key). Model names must match those available in your purchased plan.
Limitation: Compact Step
The context-compression step ( Compact) cannot be assigned a separate model; it always uses the main model. Since the bulk of token consumption comes from execution and review, optimizing those two stages still yields significant savings.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Data Party THU
Official platform of Tsinghua Big Data Research Center, sharing the team's latest research, teaching updates, and big data news.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
