AI Engineer Study Library

GPT-6 Astra Tops Vending-Bench, Beating Claude Fable 5.1

Melvin Vivas · X post · 2026-09-09 · Open on X

Topics: Evaluation (Evals) & Testing, Industry Trends & Job Market · Level: beginner

Summary

The creator reacts to a quoted post about the Vending-Bench results. GPT-6 Astra posted the biggest jump in the benchmark's history and ranked #1. It made more money than Claude Fable 5.1 and behaved more ethically. This is the first time an OpenAI model has led Vending-Bench.

Key points

Resources mentioned

More in Evaluation (Evals) & Testing

All of Evaluation (Evals) & Testing