Local Evals: Open-Source App for Evaluating Local or OpenAI-Compatible Models
Melvin Vivas · X post · 2026-09-08 · Open on X
Topics: Evaluation (Evals) & Testing, LLM Fundamentals, AI Dev Tools & Productivity · Level: intermediate
Summary
Melvin Vivas released local-evals, an open-source evaluation app that runs locally. It works with local models or any OpenAI-compatible API. It evaluates document/image-to-JSON, text-to-JSON and tool calling. He built it entirely with Codex, using GPT-6 Astra as orchestrator and Luna as subagents, and tested it with LM Studio and OpenRouter.
Key points
- local-evals runs on your own machine and talks to any OpenAI-compatible endpoint, so you can compare local and hosted models with the same harness.
- Supported eval types: document/image → JSON extraction, text → JSON (structured output), and tool calling.
- It was tested with LM Studio for local models and OpenRouter for hosted models.
- He built the whole app with Codex using an Astra/Luna combo (Astra orchestrates, Luna subagents do the work).
- Structured-output and tool-calling accuracy are useful evals when you pick a small or local model for agent tasks.
Resources mentioned
- donvito/local-evals · repo · github.com · free
Eval app that runs locally and tests local models or any OpenAI-compatible API on JSON extraction and tool calling.
Also in: Melvin Vivas's open-source AI tools, built with Codex (Melvin Vivas on X · notes), local-evals: A Local LLM Eval App Built by Codex Subagents (Melvin Vivas on X · notes), Measure Your Own Codex Token Burn: Orchestrator+Subagents vs Single Agent (Melvin Vivas on X · notes) - LM Studio · tool · x.com · free
Desktop app for downloading and running LLMs locally, with a developer mode that serves models through an API.
Also in: Running LLMs Locally Without an Expensive Rig (Melvin Vivas on X · notes), Adding Vercel AI Gateway as a provider in AIBackends with Devin (Melvin Vivas on X · notes), LoRA Fine-Tune Qwen3.5-2B on Your Tweets with Unsloth Studio (Melvin Vivas on X · notes), AIBackends: An API Layer Between Your App and AI Models (Now with Jev) (Melvin Vivas on X · notes) and 31 more - OpenRouter · tool · openrouter.ai · free
A single OpenAI-compatible API that routes requests to hundreds of models from many providers, with one bill and model fallbacks.
Also in: Customizing Your Coding Setup with Pi Coding Agent Extensions (Melvin Vivas on X · notes), The Jev model is now on OpenRouter (Melvin Vivas on X · notes), Kev-4B model, an alternative to Jev, now available on OpenRouter (Melvin Vivas on X · notes), Space Bunny Alpha: stealth 1M-context flash model on OpenRouter (Melvin Vivas on X · notes) and 42 more - OpenAI Codex · tool · openai.com · paid
OpenAI's coding agent. In the diagram it writes code, fixes review findings and drives the build loop. The creator also used it to make this video.
Also in: An agent bot that installs and drives Codex on its own (Melvin Vivas on X · notes), Asking a Coder bot to install Codex (Melvin Vivas on X · notes), Sign in with ChatGPT: Setting Usage Limits for Each App (Melvin Vivas on X · notes), Codex Cloud Environments Must Be Saved & Published Before Use (Melvin Vivas on X · notes) and 240 more
Try this
- Clone local-evals and run it against a local model in LM Studio or a hosted model via OpenRouter.
- Build a local eval harness that compares models on JSON extraction and tool-calling accuracy.
More in Evaluation (Evals) & Testing
- Evaluating Claude Code Plugins with `claude plugin eval`
- local-evals: A Local LLM Eval App Built by Codex Subagents
- GPT-6 Astra Tops Vending-Bench, Beating Claude Fable 5.1
- Building an OCR Evals App with Codex
- Be Skeptical of Model Leaderboards: Muse Spark vs Astra
- Personal Smoke Test for New Models: "Make a CRM App"