Evals Are Real Work: Use Coding Agents to Find and Build Datasets
Melvin Vivas · X post · 2026-08-27 · Open on X
Topics: Evaluation (Evals) & Testing, AI Dev Tools & Productivity · Level: beginner
Summary
Building evals takes real effort, mostly in finding existing datasets or generating synthetic ones. Coding agents like Claude Code or Codex can help with that work.
Key points
- Searching for eval datasets takes a lot of time.
- Generating synthetic eval data is another option, but it also takes time.
- Use Claude Code or Codex to help find or generate datasets.
Resources mentioned
- Claude Code · tool · code.claude.com · paid · recommended by both Bashiri Smith & Melvin Vivas
Build agents and pipelines from the terminal; the guide's main agentic coding tool.
Also in: Create Claude Code Plugins with /plugin-authoring (Melvin Vivas on X · notes), Claude Code mods: customize behavior and UI with plugins (Melvin Vivas on X · notes), AI Engineer Roadmap Overview: From ML Foundations to RAG, Agents & Ops (Bashiri Smith on Facebook · notes), SkillsBento: Free Plugin Marketplace for Codex and Claude Code (Melvin Vivas on X · notes) and 101 more - OpenAI Codex · tool · openai.com · paid
OpenAI's coding agent. In the diagram it writes code, fixes review findings and drives the build loop. The creator also used it to make this video.
Also in: An agent bot that installs and drives Codex on its own (Melvin Vivas on X · notes), Asking a Coder bot to install Codex (Melvin Vivas on X · notes), Sign in with ChatGPT: Setting Usage Limits for Each App (Melvin Vivas on X · notes), Codex Cloud Environments Must Be Saved & Published Before Use (Melvin Vivas on X · notes) and 240 more
Try this
- Use a coding agent to help find or generate synthetic eval datasets.
More in Evaluation (Evals) & Testing
- Cohere Parse: Pricing vs Parse Bench Score
- Evaluation for AI Engineering: Free Full Guide (Resource Share)
- Local Data Studio: Inspect Private Datasets Without Hugging Face Pro
- Is OpenAI Evals Still Maintained?
- Running LLM evals with promptfoo on local llama.cpp models
- The 8 Layers of Evaluating Production RAG and Agent Systems