Cut Codex Costs: Assign Strong Model for Planning, Cheap Models for Execution & Review

This tutorial explains Codex's internal model call chain—planning, execution, review, and context compression—and shows how to configure separate models for each role via config.toml and worker.toml, using a powerful model for planning while delegating coding and review to cheaper alternatives, permanently reducing per-task API costs.

Data Party THU
Data Party THU
Data Party THU
Cut Codex Costs: Assign Strong Model for Planning, Cheap Models for Execution & Review

Codex is a powerful AI coding agent, but its per-task cost can be high because a single request actually triggers a chain of model calls. The article breaks down this chain into four distinct roles:

Main model : understands the goal, plans the task, coordinates execution, and decides the next step. Errors here propagate downstream, so this role needs strong reasoning.

Sub-agent : performs concrete actions—editing files, writing code snippets, running tests. Tasks are already well-defined, so top-tier reasoning is not required.

Review model : checks code for issues and suggests fixes. The scope is narrow, making it a good fit for a cheaper model.

Compact : compresses context when it grows too long. This step currently cannot be assigned a separate model; it follows the main model.

The core insight: use a high-capability model only for planning, and offload execution and review to lower-cost models . Once configured, Codex applies this division of labor automatically for every task.

Step 1: Configure Main Model and Review Model

Edit the global configuration file:

macOS/Linux: ~/.codex/config.toml Windows: C:\Users\<username>\.codex\config.toml Add the following entries (example uses OpenAI's gpt-5.6 series):

model = "gpt-5.6-sol"
model_reasoning_effort = "high"
review_model = "gpt-5.6-luna"
model

is the main planner; model_reasoning_effort = "high" enables deeper reasoning. review_model handles code review with a lighter model.

Step 2: Configure Sub-agent Model

Create a worker configuration file at ~/.codex/agents/worker.toml (create the directory if needed) with:

model = "gpt-5.6-luna"
model_reasoning_effort = "medium"

The sub-agent inherits the session's provider, so the model must be supported by the same provider. Using gpt-5.6-luna with medium reasoning effort cuts cost for the high-volume coding and test-running steps.

Using a Third-Party Provider (Volcengine Agent Plan)

If you prefer a non-OpenAI provider that supports the Responses API, you can define a custom provider in config.toml:

model = "<strong-model-in-plan>"
review_model = "<lightweight-model-in-plan>"
model_provider = "volcengine-agent-plan"
model_supports_reasoning_summaries = true

[model_providers.volcengine-agent-plan]
name = "volcengine-agent-plan"
base_url = "https://ark.cn-beijing.volces.com/api/plan/v3"
env_key = "ARK_API_KEY"
wire_api = "responses"

Then in worker.toml set the sub-agent to the plan's lightweight model:

model = "<lightweight-model-in-plan>"
model_reasoning_effort = "medium"

Note: Agent Plan requires a dedicated API key (not a standard Volcengine Ark key). Model names must match those available in your purchased plan.

Limitation: Compact Step

The context-compression step ( Compact) cannot be assigned a separate model; it always uses the main model. Since the bulk of token consumption comes from execution and review, optimizing those two stages still yields significant savings.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

cost optimizationOpenAIAI coding assistantVolcengineCodexmodel configurationmulti-model pipeline
Data Party THU
Written by

Data Party THU

Official platform of Tsinghua Big Data Research Center, sharing the team's latest research, teaching updates, and big data news.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.