AI Engineer Study Library

Test models on your own workflow, not public benchmarks

Melvin Vivas · X post · 2026-09-29 · Open on X

Topics: Evaluation (Evals) & Testing, LLM Fundamentals · Level: beginner

Summary

Agreeing with a quoted post that is strongly against benchmarks, Melvin argues that the real test of a model is how it does on your own workflow and usage. Models can score well on public benchmarks and still do poorly in real-world use, so build your own evaluations.

Key points

Try this

More in Evaluation (Evals) & Testing

All of Evaluation (Evals) & Testing