GPT-6 Astra Tops Vending-Bench, Beating Claude Fable 5.1
Melvin Vivas · X post · 2026-09-09 · Open on X
Topics: Evaluation (Evals) & Testing, Industry Trends & Job Market · Level: beginner
Summary
The creator reacts to a quoted post about the Vending-Bench results. GPT-6 Astra posted the biggest jump in the benchmark's history and ranked #1. It made more money than Claude Fable 5.1 and behaved more ethically. This is the first time an OpenAI model has led Vending-Bench.
Key points
- Vending-Bench tests how well a model runs a business (a vending operation) over time.
- GPT-6 Astra posted the biggest jump in Vending-Bench history.
- It is the first time an OpenAI model ranks #1 on Vending-Bench.
- The top model is no longer the least ethical one: Astra beat Claude Fable 5.1 on both profit and ethics.
Resources mentioned
- Vending-Bench · other · andonlabs.com · free
A benchmark that tests how well LLM agents run a long-running vending business. - GPT-6 Astra · tool · openai.com · paid
The model announced in the quoted launch post, pitched as the developer's most capable model for work, coding, science and cybersecurity, and able to operate a computer.
Also in: Customizing Your Coding Setup with Pi Coding Agent Extensions (Melvin Vivas on X · notes), Pi Agent Council: Ask Multiple LLMs in Parallel and Compare Their Advice (Melvin Vivas on X · notes), Use GPT-6.1 Sol by Default, Save Astra for Emergencies (Melvin Vivas on X · notes), Dots in ChatGPT: always-on AI agents that you hand responsibilities to (Melvin Vivas on X · notes) and 49 more - Claude Fable 5.1 · tool · anthropic.com · paid
An Anthropic Claude model used as the performance reference for Opus 5.5.
Also in: Claude Opus 5.5 Released: Fable 5.1-Level Performance at 40% Lower Cost (Melvin Vivas on X · notes), Claude Fable 5.1 Runs a 38-Hour Unattended ML Task (Melvin Vivas on X · notes), Claude Fable 5.1 Scores 73.4% on CursorBench 3.2 (Melvin Vivas on X · notes), Why Claude Fable 5.1 Costs Less: Cheaper Cache Reads (Melvin Vivas on X · notes)
More in Evaluation (Evals) & Testing
- How to Evaluate a RAG System Before Production (Interview Answer)
- Evaluating Claude Code Plugins with `claude plugin eval`
- local-evals: A Local LLM Eval App Built by Codex Subagents
- Local Evals: Open-Source App for Evaluating Local or OpenAI-Compatible Models
- Building an OCR Evals App with Codex
- Be Skeptical of Model Leaderboards: Muse Spark vs Astra