Running LLM evals with promptfoo on local llama.cpp models
Melvin Vivas · X post · 2026-08-27 · Open on X
Topics: Evaluation (Evals) & Testing, AI Safety, Security & Guardrails, LLMOps, Deployment & Monitoring · Level: intermediate
Summary
The creator found promptfoo, an open-source tool for running evals (tests that measure model quality). He uses it with local models served from his llama.cpp installation. promptfoo tests prompts, agents and RAG apps with simple config files, compares models side by side, and also does red-team security testing.
Key points
- promptfoo is an open-source tool for evaluating prompts, agents and RAG pipelines.
- You set up tests in a simple config file instead of writing custom test code.
- It compares results across GPT, Claude, Gemini, DeepSeek and other models.
- It also does red teaming and vulnerability scanning for AI apps.
- It can test local models, such as ones served by llama.cpp, so evals cost nothing in API fees.
Resources mentioned
- promptfoo · repo · github.com · free · recommended by both Bashiri Smith & Melvin Vivas
Listed under eval frameworks for defining and running evaluations, offline and in CI. - llama.cpp · repo · github.com · free
An open-source C/C++ engine for running GGUF models locally. Its llama-server command provides an OpenAI-compatible HTTP server.
Also in: Running LLMs Locally Without an Expensive Rig (Melvin Vivas on X · notes), llama.cpp / Llama-macOS v0.5.0 release (Melvin Vivas on X · notes), Run llama.cpp GGUF Checkpoints in Hugging Face Transformers (Melvin Vivas on X · notes), llama.cpp v0.4.1 release announcement (Melvin Vivas on X · notes) and 28 more
Try this
- Install promptfoo and point it at local models served by llama.cpp to run evals.
- Set up a local eval harness: compare several llama.cpp-hosted models on your own test prompts using promptfoo.
More in Evaluation (Evals) & Testing
- Local Data Studio: Inspect Private Datasets Without Hugging Face Pro
- Evals Are Real Work: Use Coding Agents to Find and Build Datasets
- Is OpenAI Evals Still Maintained?
- The 8 Layers of Evaluating Production RAG and Agent Systems
- Pipette: Open-Source Benchmarking for On-Device Models
- Idea: A Benchmark for Codex Usage-Limit Consumption