Cutting Token Costs in Codex Orchestrator/Subagent Setups via config.toml
Melvin Vivas · X post · 2026-09-07 · Open on X
Topics: AI Agents, Tool Use & MCP, AI Dev Tools & Productivity, LLMOps, Deployment & Monitoring · Level: intermediate
Summary
Melvin Vivas explains how he changed the default model settings in his Codex Astra-Luna Subagents skill after users said it used too many tokens. The root orchestrator (gpt-6-astra) now uses low reasoning effort instead of high. Subagents (gpt-5.6-luna) stay at medium reasoning, and he cut the number of concurrent threads from 6 to 4. He recommends setting up skills and subagents per project, because projects differ in complexity.
Key points
- He suspects the high token use came from the root orchestrator (gpt-6-astra) running at high reasoning effort by default.
- Orchestrator config.toml: model = "gpt-6-astra", model_reasoning_effort = "low" (it was "high").
- Subagent defaults: default_subagent_model = "gpt-5.6-luna", default_subagent_reasoning_effort = "medium". Raise this in the subagent config if subagents perform poorly on their tasks.
- max_concurrent_threads_per_session went from 6 to 4. Change it up or down as needed.
- A common pattern: a strong model as orchestrator at low reasoning effort, with cheaper subagent models doing the actual work.
- Every setting can be configured. Set skills and subagents per project, since different projects have different complexities.
- To update without losing your own changes, pull the latest repo version and copy only the config parts you need instead of overwriting your customized files.
Resources mentioned
- donvito/codex-astra-luna-orchestrator · repo · github.com · free
The creator's Codex skill for an orchestrator-plus-subagents setup (Astra orchestrator, Luna subagents), configured through config.toml.
Also in: Codex Orchestrator Pattern: Astra/Sol Orchestrator with Luna Subagents (Melvin Vivas on X · notes), Open-source Codex skill: Astra/Sol orchestrator with Luna subagents (Melvin Vivas on X · notes), Cost-Efficient Codex Subagents: Astra/Sol Orchestrator + Luna Workers (Melvin Vivas on X · notes), Codex Orchestrator v0.2.1: GPT-6 Sol orchestrator with Luna subagents (Melvin Vivas on X · notes) and 27 more - OpenAI Codex · tool · openai.com · paid
OpenAI's coding agent. In the diagram it writes code, fixes review findings and drives the build loop. The creator also used it to make this video.
Also in: An agent bot that installs and drives Codex on its own (Melvin Vivas on X · notes), Asking a Coder bot to install Codex (Melvin Vivas on X · notes), Sign in with ChatGPT: Setting Usage Limits for Each App (Melvin Vivas on X · notes), Codex Cloud Environments Must Be Saved & Published Before Use (Melvin Vivas on X · notes) and 240 more - GPT-6 Astra · tool · openai.com · paid
The model announced in the quoted launch post, pitched as the developer's most capable model for work, coding, science and cybersecurity, and able to operate a computer.
Also in: Customizing Your Coding Setup with Pi Coding Agent Extensions (Melvin Vivas on X · notes), Pi Agent Council: Ask Multiple LLMs in Parallel and Compare Their Advice (Melvin Vivas on X · notes), Use GPT-6.1 Sol by Default, Save Astra for Emergencies (Melvin Vivas on X · notes), Dots in ChatGPT: always-on AI agents that you hand responsibilities to (Melvin Vivas on X · notes) and 49 more - GPT 5.6 Luna · tool · openai.com · paid
The OpenAI model used inside Codex for the demo. The transcript gives the variant name as 'Soul', which is unclear.
Also in: Set Codex subagent model and reasoning to save usage limits (Melvin Vivas on X · notes), Use GPT-5.6 Luna in Codex for Terminal Tasks (Melvin Vivas on X · notes), Match Reasoning Effort to Task Length in Codex (Astra/Sol) (Melvin Vivas on X · notes), Run Coworker desktop agents cheaply with GPT-5.6 Luna on OpenRouter (Melvin Vivas on X · notes) and 24 more
Try this
- Pull the latest version of the codex-astra-luna-orchestrator repo, or copy just the new config so your customized files aren't overwritten.
- Set the orchestrator's model_reasoning_effort to "low" to reduce token use.
- If subagents perform poorly on their tasks, raise default_subagent_reasoning_effort in the subagent config.
- Change max_concurrent_threads_per_session (now 4) to fit your needs.
- Set up skills and subagents per project, based on how complex each project is.
More in AI Agents, Tool Use & MCP
- Codex Orchestrator: Astra/Sol Orchestrator with Luna Subagents
- JSONL Viewer v1.3.4: Inspecting Nested Codex Subagent Traces
- Cutting Codex Subagent Costs: Luna Orchestrator, Astra Reviewer
- Inspecting Coding-Agent Traces Live with JSONL Viewer (Codex, Claude Code)
- Codex Astra Orchestrator + Luna Subagents Skill Update
- Learn Subagents in Cursor and Codex