AI Engineer Study Library

LoopsBench: Testing Coding Agents on Long-Horizon Tasks

Melvin Vivas · X post · 2026-08-20 · Open on X

Topics: Evaluation (Evals) & Testing, AI Agents, Tool Use & MCP, AI Dev Tools & Productivity · Level: advanced

Summary

LoopsBench evaluates coding agents (a model plus its harness) on 112 long-horizon coding tasks built from real pull requests. Claude Opus 4.7 with Claude Code had the highest resolve rate at 25.00%. GPT-5.5 with Codex reached 21.43%. The low scores show long-horizon coding is still hard for agents.

Key points

Resources mentioned

Try this

More in Evaluation (Evals) & Testing

All of Evaluation (Evals) & Testing