local-evals: A Local LLM Eval App Built by Codex Subagents
Melvin Vivas · X post · 2026-09-10 · Open on X
Topics: Evaluation (Evals) & Testing, AI Agents, Tool Use & MCP, AI Dev Tools & Productivity · Level: intermediate
Summary
Melvin Vivas shares local-evals, an open-source eval app that Codex built from scratch with an Astra orchestrator and Luna subagents. The app runs locally and can test local models or any OpenAI-compatible API.
Key points
- local-evals runs evaluations on your own machine.
- It works with local models or any OpenAI-compatible API.
- It was built from scratch (greenfield) by Codex, with Astra orchestrating Luna subagents.
Resources mentioned
- donvito/local-evals · repo · github.com · free
Eval app that runs locally and tests local models or any OpenAI-compatible API on JSON extraction and tool calling.
Also in: Melvin Vivas's open-source AI tools, built with Codex (Melvin Vivas on X · notes), Measure Your Own Codex Token Burn: Orchestrator+Subagents vs Single Agent (Melvin Vivas on X · notes), Local Evals: Open-Source App for Evaluating Local or OpenAI-Compatible Models (Melvin Vivas on X · notes)
Try this
- Try local-evals to run evals against local models or OpenAI-compatible endpoints.
- Build a local eval app for comparing models through OpenAI-compatible APIs.
More in Evaluation (Evals) & Testing
- Models Cheating on Terminal-Bench-2.1
- How to Evaluate a RAG System Before Production (Interview Answer)
- Evaluating Claude Code Plugins with `claude plugin eval`
- GPT-6 Astra Tops Vending-Bench, Beating Claude Fable 5.1
- Local Evals: Open-Source App for Evaluating Local or OpenAI-Compatible Models
- Building an OCR Evals App with Codex