Comparing GPT Models in Codex by Intelligence and Cost per Task
Melvin Vivas · X post · 2026-09-09 · Open on X
Topics: LLM Fundamentals, AI Dev Tools & Productivity, LLMOps, Deployment & Monitoring · Level: intermediate
Summary
The creator shares an Artificial Analysis comparison of every GPT model available in Codex. It covers GPT-6 Astra, GPT-5.6 Sol, Terra and Luna, and GPT-5.5 at each reasoning level, showing their intelligence scores and cost per task. He says it is handy for choosing which model to use.
Key points
- Artificial Analysis compares models on its Intelligence Index and on cost per task.
- The comparison covers GPT-6 Astra, GPT-5.6 Sol, Terra and Luna, and GPT-5.5 and GPT-5.5 Pro.
- Each model is listed at its reasoning-effort levels (xhigh, high, medium, low and non-reasoning).
- Use the chart to balance capability against cost when picking a Codex model.
Resources mentioned
- Artificial Analysis – Comparison of AI Models across Intelligence, Performance, and Price · website · artificialanalysis.ai · free
Benchmark site comparing AI models on an intelligence index, speed and price, here filtered to the GPT models in Codex.
Also in: Qwen 3.8 27B Matches GPT 5.6 Luna (max) on Artificial Analysis (Melvin Vivas on X · notes) - OpenAI Codex · tool · openai.com · paid
OpenAI's coding agent. In the diagram it writes code, fixes review findings and drives the build loop. The creator also used it to make this video.
Also in: An agent bot that installs and drives Codex on its own (Melvin Vivas on X · notes), Asking a Coder bot to install Codex (Melvin Vivas on X · notes), Sign in with ChatGPT: Setting Usage Limits for Each App (Melvin Vivas on X · notes), Codex Cloud Environments Must Be Saved & Published Before Use (Melvin Vivas on X · notes) and 240 more
Try this
- Check the Artificial Analysis comparison before choosing a model and reasoning level in Codex.
More in LLM Fundamentals
- Gemma 4 as a Strong Small Model for Local and On-Device Use
- Running Qwen3.8 27B Locally with llama.cpp for Writing
- Three Local GGUF Models That Fit on an RTX 3090 (24GB)
- Open-weight Nemotron models for finance and healthcare
- Qwen3.8 27B Quantization Benchmark: 4-Bit Is Enough for Agentic Coding
- Finding Models to Run Locally on Hugging Face