LoopsBench: Testing Coding Agents on Long-Horizon Tasks
Melvin Vivas · X post · 2026-08-20 · Open on X
Topics: Evaluation (Evals) & Testing, AI Agents, Tool Use & MCP, AI Dev Tools & Productivity · Level: advanced
Summary
LoopsBench evaluates coding agents (a model plus its harness) on 112 long-horizon coding tasks built from real pull requests. Claude Opus 4.7 with Claude Code had the highest resolve rate at 25.00%. GPT-5.5 with Codex reached 21.43%. The low scores show long-horizon coding is still hard for agents.
Key points
- 112 long-horizon coding tasks built from real pull requests.
- The benchmark scores a model and its harness together (for example Opus 4.7 + Claude Code).
- Top resolve rate: Opus 4.7 with Claude Code, 25.00%.
- GPT-5.5 with Codex: 21.43%.
- The paper (10 Aug) frames this as moving from 'harness engineering' to 'loop engineering'.
- The creator hopes the benchmark will be updated with newer models and harnesses.
Resources mentioned
- LoopsBench Leaderboard · website · loopsbench.ai · free
Leaderboard of coding agent resolve rates on LoopsBench. - LoopsBench: From Harness Engineering to Loop Engineering in Coding Agent Evaluation · paper · alphaxiv.org · free
Paper introducing LoopsBench, a benchmark of 112 long-horizon coding tasks from real pull requests. - Claude Code · tool · code.claude.com · paid · recommended by both Bashiri Smith & Melvin Vivas
Build agents and pipelines from the terminal; the guide's main agentic coding tool.
Also in: Create Claude Code Plugins with /plugin-authoring (Melvin Vivas on X · notes), Claude Code mods: customize behavior and UI with plugins (Melvin Vivas on X · notes), AI Engineer Roadmap Overview: From ML Foundations to RAG, Agents & Ops (Bashiri Smith on Facebook · notes), SkillsBento: Free Plugin Marketplace for Codex and Claude Code (Melvin Vivas on X · notes) and 101 more - OpenAI Codex · tool · openai.com · paid
OpenAI's coding agent. In the diagram it writes code, fixes review findings and drives the build loop. The creator also used it to make this video.
Also in: An agent bot that installs and drives Codex on its own (Melvin Vivas on X · notes), Asking a Coder bot to install Codex (Melvin Vivas on X · notes), Sign in with ChatGPT: Setting Usage Limits for Each App (Melvin Vivas on X · notes), Codex Cloud Environments Must Be Saved & Published Before Use (Melvin Vivas on X · notes) and 240 more
Try this
- Read the LoopsBench paper and check the full leaderboard.
- Run several model + harness pairs on a small set of real PR-based tasks and compare their resolve rates.
More in Evaluation (Evals) & Testing
- Idea: A Benchmark for Codex Usage-Limit Consumption
- Why Evaluation Is the Skill That Sets Top AI Engineers Apart
- AI-Written Code Still Needs Human Testing
- Agent Loops Need Real Evals and Feedback
- OCR Testing a Small Vision Model with LLM-Made Ground Truth
- Code Arena Fullstack Benchmark: Kimi K3 Ranks #1