Evaluating Claude Code Plugins with claude plugin eval
Melvin Vivas · X post · 2026-09-12 · Open on X
Topics: Evaluation (Evals) & Testing, AI Dev Tools & Productivity, AI Agents, Tool Use & MCP · Level: intermediate
Summary
Claude Code added a claude plugin eval command for measuring how much value a plugin or skill adds. You write test cases, run the plugin against them and score the runs. Then you run the same cases without the plugin to compare. The creator wants to try the same approach with Codex plugins.
Key points
claude plugin evalbenchmarks plugins and skills on your own use cases- Write test cases that reflect your real tasks
- Run the plugin or skill against them and score each run
- Run each case again without the plugin to get a baseline
- Compare the two sets of scores to see whether the plugin helps or needs more work
Resources mentioned
- Claude Code `claude plugin eval` · tool · code.claude.com · free
A Claude Code command that runs plugins and skills against test cases, with and without the plugin, and scores the runs. - Claude Code · tool · code.claude.com · paid · recommended by both Bashiri Smith & Melvin Vivas
Build agents and pipelines from the terminal; the guide's main agentic coding tool.
Also in: Create Claude Code Plugins with /plugin-authoring (Melvin Vivas on X · notes), Claude Code mods: customize behavior and UI with plugins (Melvin Vivas on X · notes), AI Engineer Roadmap Overview: From ML Foundations to RAG, Agents & Ops (Bashiri Smith on Facebook · notes), SkillsBento: Free Plugin Marketplace for Codex and Claude Code (Melvin Vivas on X · notes) and 101 more - OpenAI Codex · tool · openai.com · paid
OpenAI's coding agent. In the diagram it writes code, fixes review findings and drives the build loop. The creator also used it to make this video.
Also in: An agent bot that installs and drives Codex on its own (Melvin Vivas on X · notes), Asking a Coder bot to install Codex (Melvin Vivas on X · notes), Sign in with ChatGPT: Setting Usage Limits for Each App (Melvin Vivas on X · notes), Codex Cloud Environments Must Be Saved & Published Before Use (Melvin Vivas on X · notes) and 240 more
Try this
- Write test cases for your own plugins or skills and run `claude plugin eval`
- Compare the scores with and without the plugin
- Try a similar with/without benchmark for Codex plugins
- Build a small with/without benchmark harness for Codex plugins or skills
More in Evaluation (Evals) & Testing
- Jev by TypeSafe AI as an alternative to LLM-as-judge
- Models Cheating on Terminal-Bench-2.1
- How to Evaluate a RAG System Before Production (Interview Answer)
- local-evals: A Local LLM Eval App Built by Codex Subagents
- GPT-6 Astra Tops Vending-Bench, Beating Claude Fable 5.1
- Local Evals: Open-Source App for Evaluating Local or OpenAI-Compatible Models