Topic 8 of 16 in the learning path
Evaluation (Evals) & Testing
Measuring quality: test sets, LLM-as-judge, metrics and regression testing.
Reels and posts (34)
- 5 Things AI Engineers Must Evaluate: RAG, Agents, Models, Data, Guardrails · Bashiri Smith, Facebook · 0:33: A quick checklist of the five parts of an AI system that engineers should evaluate, and the two metrics to check for each: RAG, agents, models, datasets and guardrails.
- Test models on your own workflow, not public benchmarks · Melvin Vivas, X: Agreeing with a quoted post that is strongly against benchmarks, Melvin argues that the real test of a model is how it does on your own workflow and usage.
- Jev by TypeSafe AI as an alternative to LLM-as-judge · Melvin Vivas, X: The creator claims that TypeSafe AI's Jev model 'just killed LLM as judge', meaning it could replace the usual approach of grading outputs with an LLM.
- Models Cheating on Terminal-Bench-2.1 · Melvin Vivas, X: The creator quotes a post about AI models cheating on Terminal-Bench-2.1.
- How to Evaluate a RAG System Before Production (Interview Answer) · Bashiri Smith, Facebook · 1:23: This skit shows two candidates answering the same RAG evaluation interview question.
- Evaluating Claude Code Plugins with `claude plugin eval` · Melvin Vivas, X: Claude Code added a `claude plugin eval` command for measuring how much value a plugin or skill adds.
- local-evals: A Local LLM Eval App Built by Codex Subagents · Melvin Vivas, X: Melvin Vivas shares local-evals, an open-source eval app that Codex built from scratch with an Astra orchestrator and Luna subagents.
- GPT-6 Astra Tops Vending-Bench, Beating Claude Fable 5.1 · Melvin Vivas, X: The creator reacts to a quoted post about the Vending-Bench results.
- Local Evals: Open-Source App for Evaluating Local or OpenAI-Compatible Models · Melvin Vivas, X: Melvin Vivas released local-evals, an open-source evaluation app that runs locally.
- Building an OCR Evals App with Codex · Melvin Vivas, X: A short update that he is building an evals app for OCR with Codex.
- Be Skeptical of Model Leaderboards: Muse Spark vs Astra · Melvin Vivas, X: The creator tried Muse Spark 1.3 in OpenCode and found it 'no way close to a GPT model'.
- Personal Smoke Test for New Models: "Make a CRM App" · Melvin Vivas, X: When a new model launches, the creator tries it with one personal smoke-test prompt: "make a crm app".
- Compare Model Outputs with Artificial Analysis MicroEvals · Melvin Vivas, X: Artificial Analysis MicroEvals lets you compare outputs from different models side by side.
- Testing Gemini 3.8 Flash in Cursor with a CRM Smoke Test · Melvin Vivas, X: The creator tested Gemini 3.8 Flash in Cursor using his own CRM app 'smoke test'.
- Evaluation for AI Engineering: Free Full Guide (Resource Share) · Bashiri Smith, Facebook · 0:08: This 8-second reel has no speech.
- Cohere Parse: Pricing vs Parse Bench Score · Melvin Vivas, X: The post points to where Cohere Parse sits on a chart of pricing against Parse Bench score.
- Local Data Studio: Inspect Private Datasets Without Hugging Face Pro · Melvin Vivas, X: Hosting private datasets on Hugging Face needs a Pro subscription.
- Evals Are Real Work: Use Coding Agents to Find and Build Datasets · Melvin Vivas, X: Building evals takes real effort, mostly in finding existing datasets or generating synthetic ones.
- Is OpenAI Evals Still Maintained? · Melvin Vivas, X: Melvin Vivas notes that the OpenAI Evals repo's last commit was 4 months ago and asks whether it is still maintained.
- Running LLM evals with promptfoo on local llama.cpp models · Melvin Vivas, X: The creator found promptfoo, an open-source tool for running evals (tests that measure model quality).
- The 8 Layers of Evaluating Production RAG and Agent Systems · Bashiri Smith, Facebook · 0:07: Building a RAG app or agent demo is only half of an AI engineer's job.
- Pipette: Open-Source Benchmarking for On-Device Models · Melvin Vivas, X: The creator points to Pipette, a newly released open-source suite for evaluating on-device models, built in partnership with Artificial Analysis.
- Idea: A Benchmark for Codex Usage-Limit Consumption · Melvin Vivas, X: Proposes a benchmark that gives Codex the exact same tasks over time and measures how much of the usage limit each run uses.
- Why Evaluation Is the Skill That Sets Top AI Engineers Apart · Bashiri Smith, Facebook · 0:49: The creator argues that many people learn to build RAG systems, agents and AI apps, but the AI engineers who earn the most are paid to evaluate these systems and keep them reliable
- AI-Written Code Still Needs Human Testing · Melvin Vivas, X: Software development isn't just prompting.
- LoopsBench: Testing Coding Agents on Long-Horizon Tasks · Melvin Vivas, X: LoopsBench evaluates coding agents (a model plus its harness) on 112 long-horizon coding tasks built from real pull requests.
- Agent Loops Need Real Evals and Feedback · Melvin Vivas, X: A short lesson: autonomous coding loops are just automated vibe coding unless you fix the evals and feedback loop.
- OCR Testing a Small Vision Model with LLM-Made Ground Truth · Melvin Vivas, X: Melvin Vivas tests the OCR quality of Liquid AI's small LFM2.5-VL-3B vision-language model.
- Code Arena Fullstack Benchmark: Kimi K3 Ranks #1 · Melvin Vivas, X: Code Arena added a fullstack benchmark that ranks AI models on full-stack web development tasks: multi-step reasoning, tool use and building complete apps.
- Test AI-Generated Apps: They Are Not Bug-Free · Melvin Vivas, X: A short reminder that apps built with AI coding tools still have bugs.
- How Anthropic's data team automated 95% of analytics queries with Claude · Melvin Vivas, X: The creator shares Anthropic's announcement that its data team automated 95% of business analytics queries with Claude, calling it the way to automate.
- End-to-end test harnesses are key to automated AI coding · Melvin Vivas, X: The creator argues that a fully automated AI coding pipeline needs a robust end-to-end test harness, because agents need reliable automated checks to confirm their changes.
- Read Benchmark Numbers, Not Highlights: the Muse Spark Lesson · Melvin Vivas, X: In Meta's Muse Spark benchmark table, the model's column is highlighted in blue.
- Gemma 4 on the Arena Leaderboard · Melvin Vivas, X: The creator notes that Gemma 4 has appeared on Arena, the crowdsourced platform that ranks models by side-by-side comparisons.
Read and use (50)
- Artificial Analysis · website · artificialanalysis.ai · free
Independent benchmarking site that compares AI models and providers, including text-to-speech, on quality, speed and price.
Mentioned in: ElevenLabs Eleven v4 & v4 Turbo: Emotive, Multilingual Text-to-Speech Models (Melvin Vivas on X · notes), Be Skeptical of Model Leaderboards: Muse Spark vs Astra (Melvin Vivas on X · notes), Pipette: Open-Source Benchmarking for On-Device Models (Melvin Vivas on X · notes), GLM 5.2 Hits 446 tok/s on Fireworks AI (Melvin Vivas on X · notes) and 2 more - Terminal-Bench · tool · tbench.ai · free
A benchmark that tests how well AI agents complete coding and system tasks in a terminal.
Mentioned in: Claude Sonnet 5.5 release beats Opus 5.5 on Terminal-Bench (Melvin Vivas on X · notes), Devin's SWE-2 Coding Model Is Free for a Limited Time (Until Oct 8/15) (Melvin Vivas on X · notes), Models Cheating on Terminal-Bench-2.1 (Melvin Vivas on X · notes), DeepSeek V4 Flash 0731 Agentic Benchmarks and Use with Hermes Agent (Melvin Vivas on X · notes) and 1 more - donvito/local-evals · repo · github.com · free
Eval app that runs locally and tests local models or any OpenAI-compatible API on JSON extraction and tool calling.
Mentioned in: Melvin Vivas's open-source AI tools, built with Codex (Melvin Vivas on X · notes), local-evals: A Local LLM Eval App Built by Codex Subagents (Melvin Vivas on X · notes), Measure Your Own Codex Token Burn: Orchestrator+Subagents vs Single Agent (Melvin Vivas on X · notes), Local Evals: Open-Source App for Evaluating Local or OpenAI-Compatible Models (Melvin Vivas on X · notes) - Evaluation Field Guide (baswe.ai engineer accelerator – Ops and Evaluation module) · pdf · drive.google.com · free
Bashiri Smith's 41-page practitioner's guide to evaluating AI systems.
Mentioned in: 5 Things AI Engineers Must Evaluate: RAG, Agents, Models, Data, Guardrails (Bashiri Smith on Facebook · notes), Evaluation for AI Engineering: Free Full Guide (Resource Share) (Bashiri Smith on Facebook · notes), The 8 Layers of Evaluating Production RAG and Agent Systems (Bashiri Smith on Facebook · notes), Why Evaluation Is the Skill That Sets Top AI Engineers Apart (Bashiri Smith on Facebook · notes) - RAGAS · tool · github.com · free
Listed under RAG evaluation tools in the guide's evaluation stack.
Mentioned in: SWE-to-AI Engineer Plan for 2027: LLMs, RAG, Agents, Evals, Job Search (Bashiri Smith on Facebook · notes), Taking a RAG App to Production: Evals, Guardrails, Cost and Tracing (Bashiri Smith on Facebook · notes)
In a shared PDF: Evaluation Field Guide (baswe.ai engineer accelerator – Ops and Evaluation module) - shared in this reel on Facebook, this reel on Facebook; The AI Pivot Field Guide - shared in this reel on Facebook, this reel on Facebook - CursorBench · other · cursor.com · free
Cursor's benchmark for comparing model performance on coding tasks.
Mentioned in: Claude Opus 5.5 in Cursor: top of CursorBench at 40% lower cost (Melvin Vivas on X · notes), Claude Fable 5.1 Scores 73.4% on CursorBench 3.2 (Melvin Vivas on X · notes), Why Claude Fable 5.1 Costs Less: Cheaper Cache Reads (Melvin Vivas on X · notes) - ARC-AGI-3 · dataset · arcprize.org · free
A benchmark from the ARC Prize that tests an AI agent's general reasoning and adaptation in interactive tasks.
Mentioned in: Prime Agent: A Self-Improving RLM Coding Harness with Code-Based Tool Calls (Melvin Vivas on X · notes), Prime Agent: A Self-Improving RLM Coding Harness That Calls Tools in Code (Melvin Vivas on X · notes) - DeepSWE · tool · deepswe.datacurve.ai · free
A software-engineering benchmark used to compare how well coding models perform.
Mentioned in: DeepSWE results: GPT-6 Sol slightly below GPT-5.6 Sol, but cheaper (Melvin Vivas on X · notes), DeepSeek V4 Flash 0731 Agentic Benchmarks and Use with Hermes Agent (Melvin Vivas on X · notes) - Design Arena · website · designarena.ai · free
A leaderboard that ranks AI models on design/UI generation using head-to-head comparisons.
Mentioned in: Kimi K3 ranks #1 on Design Arena for frontend building (Melvin Vivas on X · notes), GLM 5.2 Ranks Above Fable 5 on Design Arena (Melvin Vivas on X · notes) - FrontierCode 1.1 · other · cognition.com · free
A coding benchmark used to compare frontier models.
Mentioned in: Claude Opus 5.5 in Devin: #1 on FrontierCode 1.1 (Melvin Vivas on X · notes), GPT-6 Astra Coming to Devin: Benchmark and Cost Claims (Melvin Vivas on X · notes) - LMArena · website · lmarena.ai · free
Leaderboard that ranks AI models by crowd-sourced head-to-head comparisons, including text-to-image.
Mentioned in: Qwen 3.7 Plus Preview Ranks #16 in the Vision Arena (Melvin Vivas on X · notes), Qwen-Image-2.0-Pro Released: Text-to-Image Model Update (Melvin Vivas on X · notes) - OpenAI Evals · repo · github.com · free · recommended by both Bashiri Smith & Melvin Vivas
Listed under eval frameworks for defining and running evaluations.
Mentioned in: Is OpenAI Evals Still Maintained? (Melvin Vivas on X · notes)
In a shared PDF: Evaluation Field Guide (baswe.ai engineer accelerator – Ops and Evaluation module) - shared in this reel on Facebook, this reel on Facebook - promptfoo · repo · github.com · free · recommended by both Bashiri Smith & Melvin Vivas
Listed under eval frameworks for defining and running evaluations, offline and in CI.
Mentioned in: Running LLM evals with promptfoo on local llama.cpp models (Melvin Vivas on X · notes)
In a shared PDF: Evaluation Field Guide (baswe.ai engineer accelerator – Ops and Evaluation module) - shared in this reel on Facebook, this reel on Facebook - SWE-Bench Pro · dataset · labs.scale.com · free
Benchmark that tests whether models can solve real-world software engineering tasks.
Mentioned in: Open-Source GLM-5.1 Beats GPT-5.4 on SWE-Bench Pro (Melvin Vivas on X · notes), GLM-5.1: Open-Source Model for Long-Running Coding Agents (Melvin Vivas on X · notes) - Agents Last Exam · tool · agents-last-exam.org · free
A hard benchmark for AI agents.
Mentioned in: DeepSeek V4 Flash 0731 Agentic Benchmarks and Use with Hermes Agent (Melvin Vivas on X · notes) - Anthropic blog post: automating business analytics queries with Claude · article · claude.com · free
How Anthropic's data team automated 95% of analytics queries with Claude, including evals, ablations and online validation.
Mentioned in: How Anthropic's data team automated 95% of analytics queries with Claude (Melvin Vivas on X · notes) - Arena.ai · website · x.com · free
Crowdsourced platform and leaderboard (formerly LMArena) for comparing LLMs side by side.
Mentioned in: Gemma 4 on the Arena Leaderboard (Melvin Vivas on X · notes) - Automation Bench · tool · github.com · free
A benchmark for automation and workflow tasks done by agents.
Mentioned in: DeepSeek V4 Flash 0731 Agentic Benchmarks and Use with Hermes Agent (Melvin Vivas on X · notes) - BIRD benchmark · dataset · bird-bench.github.io · free
A large text-to-SQL benchmark that tests how well models turn questions into SQL on real databases.
Mentioned in: Gemini-SQL2: text-to-SQL with Gemini 3.1 Pro, SOTA on BIRD (Melvin Vivas on X · notes) - Claude Code `claude plugin eval` · tool · code.claude.com · free
A Claude Code command that runs plugins and skills against test cases, with and without the plugin, and scores the runs.
Mentioned in: Evaluating Claude Code Plugins with `claude plugin eval` (Melvin Vivas on X · notes) - Code Arena (Arena) · website · arena.ai · free
A leaderboard that ranks AI models on coding tasks, now including full-stack web development.
Mentioned in: Code Arena Fullstack Benchmark: Kimi K3 Ranks #1 (Melvin Vivas on X · notes) - CursorBench 4.0 · other · cursor.com · free
Cursor's benchmark for how well models handle coding tasks.
Mentioned in: GLM 5.3 and GLM 5.3 Flash now in Cursor (Melvin Vivas on X · notes) - Datacurve (@datacurve) on X · person · x.com · free
X account of Datacurve, the company that publishes the Deep SWE leaderboard.
Mentioned in: GLM 5.2 Leads Open-Source Models on Datacurve's Deep SWE Leaderboard (Melvin Vivas on X · notes) - Decision Index · other · huggingface.co · free
A benchmark for decision models, on which GLiDE ranks #1.
Mentioned in: GLiDE: Fastino's Decision Model with Adaptive Thinking (Melvin Vivas on X · notes) - Decision Index 0.2.1 · tool · github.com · free
The official benchmark scorer used to measure decision models across five evaluation areas.
Mentioned in: GLiDE by Fastino Labs: A Post-Trainable Reasoning Decision Model (Melvin Vivas on X · notes) - Deep SWE leaderboard · website · deepswe.datacurve.ai · free
A benchmark leaderboard from Datacurve that ranks models on software-engineering tasks.
Mentioned in: GLM 5.2 Leads Open-Source Models on Datacurve's Deep SWE Leaderboard (Melvin Vivas on X · notes) - DeepEval · tool · github.com · free
Listed under eval frameworks for defining and running evaluations, offline and in CI.
In a shared PDF: Evaluation Field Guide (baswe.ai engineer accelerator – Ops and Evaluation module) - shared in this reel on Facebook, this reel on Facebook - DSBench · tool · github.com · free
A benchmark for data-science agent tasks.
Mentioned in: DeepSeek V4 Flash 0731 Agentic Benchmarks and Use with Hermes Agent (Melvin Vivas on X · notes) - Frontend Code Arena (LMArena) · website · arena.ai · free
Arena leaderboard that ranks models on frontend coding tasks.
Mentioned in: Qwen3.8-Max on the Frontend Code Arena cost-performance frontier (Melvin Vivas on X · notes) - FrontierCode · other · cognition.com · free
Coding benchmark used to compare SWE-2 with other models.
Mentioned in: Cognition's SWE-2 Coding Model Now in Devin (Melvin Vivas on X · notes) - FrontierCode 1.1 Extended · other · devin.ai · free
A coding benchmark on which Devin Fusion reportedly ranks first.
Mentioned in: Devin price cuts and top score on FrontierCode 1.1 Extended (Melvin Vivas on X · notes) - GDPval-AA · other · artificialanalysis.ai · free
Artificial Analysis's version of the GDPval benchmark, which measures models on real-world economically valuable tasks and ranks them by Elo.
Mentioned in: GLM-5.2 on Fireworks: Top Open-Weights Model on GDPval-AA (Melvin Vivas on X · notes) - HallusionBench · dataset · github.com · free
Benchmark for hallucination and visual illusions in vision-language models.
Mentioned in: MiniCPM-V 4.6 1.3B: Small Open-Source Vision/OCR Model for Edge Devices (Melvin Vivas on X · notes) - Hume AI voice benchmarks · website · hume.ai · free
Hume AI's benchmarks for comparing voice and TTS models.
Mentioned in: Gemini 3.8 Flash and Flash-Lite TTS: new text-to-speech models (Melvin Vivas on X · notes) - LoopsBench Leaderboard · website · loopsbench.ai · free
Leaderboard of coding agent resolve rates on LoopsBench.
Mentioned in: LoopsBench: Testing Coding Agents on Long-Horizon Tasks (Melvin Vivas on X · notes) - LoopsBench: From Harness Engineering to Loop Engineering in Coding Agent Evaluation · paper · alphaxiv.org · free
Paper introducing LoopsBench, a benchmark of 112 long-horizon coding tasks from real pull requests.
Mentioned in: LoopsBench: Testing Coding Agents on Long-Horizon Tasks (Melvin Vivas on X · notes) - MicroEvals | Artificial Analysis · website · artificialanalysis.ai · check price
Artificial Analysis tool for running small evals and comparing different models' outputs.
Mentioned in: Compare Model Outputs with Artificial Analysis MicroEvals (Melvin Vivas on X · notes) - MUIRBench · dataset · muirbench.github.io · free
Benchmark for understanding multiple images at once.
Mentioned in: MiniCPM-V 4.6 1.3B: Small Open-Source Vision/OCR Model for Edge Devices (Melvin Vivas on X · notes) - NL2Repo · dataset · github.com · free
Benchmark that tests whether models can build a whole code repository from a natural-language description.
Mentioned in: GLM-5.1: Open-Source Model for Long-Running Coding Agents (Melvin Vivas on X · notes) - OCRBench · dataset · github.com · free
Benchmark for testing OCR ability in multimodal models.
Mentioned in: MiniCPM-V 4.6 1.3B: Small Open-Source Vision/OCR Model for Edge Devices (Melvin Vivas on X · notes) - Onely7/local_data_studio · repo · github.com · free
A local data studio for JSONL/JSON/CSV/TSV/Parquet with scalable previews, image inspection, DuckDB SQL, EDA and Embedding Atlas.
Mentioned in: Local Data Studio: Inspect Private Datasets Without Hugging Face Pro (Melvin Vivas on X · notes) - Parse Bench · dataset · github.com · free
Benchmark for scoring how well models parse documents.
Mentioned in: Cohere Parse: Pricing vs Parse Bench Score (Melvin Vivas on X · notes) - Pipette · tool · liquid.ai · free
An open-source suite for evaluating the capability and speed of on-device models.
Mentioned in: Pipette: Open-Source Benchmarking for On-Device Models (Melvin Vivas on X · notes) - RefCOCO · dataset · github.com · free
Benchmark for referring-expression grounding (finding the object an image description refers to).
Mentioned in: MiniCPM-V 4.6 1.3B: Small Open-Source Vision/OCR Model for Edge Devices (Melvin Vivas on X · notes) - SimpleQA · dataset · openai.com · free
A short-form factual question-answering benchmark that measures factual accuracy; Firecrawl uses it to score its search.
Mentioned in: Firecrawl Keyless: Free Web Search & Scraping for AI Agents, No API Key (Melvin Vivas on X · notes) - Terminal-Bench 2.1 · tool · tbench.ai · free
A benchmark that measures how well AI agents complete coding and terminal tasks.
Mentioned in: Qwen3.8 27B Quantization Benchmark: 4-Bit Is Enough for Agentic Coding (Melvin Vivas on X · notes) - Terra · tool · snorkel.ai · free
An AI model the creator says cheated on Terminal-Bench-2.1.
Mentioned in: Models Cheating on Terminal-Bench-2.1 (Melvin Vivas on X · notes) - Toolation · tool · github.com · free
A tool-use benchmark named in the post (name as written).
Mentioned in: DeepSeek V4 Flash 0731 Agentic Benchmarks and Use with Hermes Agent (Melvin Vivas on X · notes) - TruLens · tool · trulens.org · free
Listed under RAG evaluation tools in the guide's evaluation stack.
In a shared PDF: Evaluation Field Guide (baswe.ai engineer accelerator – Ops and Evaluation module) - shared in this reel on Facebook, this reel on Facebook - Vending-Bench · other · andonlabs.com · free
A benchmark that tests how well LLM agents run a long-running vending business.
Mentioned in: GPT-6 Astra Tops Vending-Bench, Beating Claude Fable 5.1 (Melvin Vivas on X · notes)
Build
- Build a small personal eval set from your real daily tasks and use it to compare models. (from Test models on your own workflow, not public benchmarks)
- A RAG evaluation pipeline that scores retrieval and generation separately, includes questions with no answer in the documents, checks document versions and effective dates, and calibrates its LLM judge against human labels, with the results published in a repo. (from How to Evaluate a RAG System Before Production (Interview Answer))
- Build a small with/without benchmark harness for Codex plugins or skills (from Evaluating Claude Code Plugins with `claude plugin eval`)
- Build a local eval app for comparing models through OpenAI-compatible APIs. (from local-evals: A Local LLM Eval App Built by Codex Subagents)
- Build a local eval harness that compares models on JSON extraction and tool-calling accuracy. (from Local Evals: Open-Source App for Evaluating Local or OpenAI-Compatible Models)
- Build an evals app that compares OCR outputs from different models against ground truth. (from Building an OCR Evals App with Codex)
- Build a CRM app as a benchmark task for comparing AI coding models. (from Personal Smoke Test for New Models: "Make a CRM App")
- Build a small CRM app as a reusable smoke test for comparing coding models. (from Testing Gemini 3.8 Flash in Cursor with a CRM Smoke Test)
- Set up a local eval harness: compare several llama.cpp-hosted models on your own test prompts using promptfoo. (from Running LLM evals with promptfoo on local llama.cpp models)
- Take a basic RAG app or agent and add a repeatable evaluation suite that covers all eight layers, so you can show it works and catch regressions. (from The 8 Layers of Evaluating Production RAG and Agent Systems)
- Build a usage and cost regression benchmark: run a fixed task suite through a coding agent on a schedule and track how much of the usage limit or how many tokens each task consumes. (from Idea: A Benchmark for Codex Usage-Limit Consumption)
- Run several model + harness pairs on a small set of real PR-based tasks and compare their resolve rates. (from LoopsBench: Testing Coding Agents on Long-Horizon Tasks)
- Build an OCR evaluation harness: a frontier model creates the ground truth, a small local VLM makes predictions, and an agent or script scores the differences. (from OCR Testing a Small Vision Model with LLM-Made Ground Truth)
- Build a text-to-SQL analytics agent with Claude and measure it with an eval set, ablations and online validation. (from How Anthropic's data team automated 95% of analytics queries with Claude)