Test models on your own workflow, not public benchmarks
Melvin Vivas · X post · 2026-09-29 · Open on X
Topics: Evaluation (Evals) & Testing, LLM Fundamentals · Level: beginner
Summary
Agreeing with a quoted post that is strongly against benchmarks, Melvin argues that the real test of a model is how it does on your own workflow and usage. Models can score well on public benchmarks and still do poorly in real-world use, so build your own evaluations.
Key points
- Public benchmark scores don't reliably predict how a model performs on real tasks.
- The most meaningful benchmark is your own workflow and usage.
- A model can do great on public leaderboards and still be bad at real-world use.
- Before switching models, test candidates on your own representative tasks.
Try this
- Test models on your own workflow and tasks instead of relying only on public benchmark scores.
- Build a small personal eval set from your real daily tasks and use it to compare models.