AI Engineer Study Library

Running LLM evals with promptfoo on local llama.cpp models

Melvin Vivas · X post · 2026-08-27 · Open on X

Topics: Evaluation (Evals) & Testing, AI Safety, Security & Guardrails, LLMOps, Deployment & Monitoring · Level: intermediate

Summary

The creator found promptfoo, an open-source tool for running evals (tests that measure model quality). He uses it with local models served from his llama.cpp installation. promptfoo tests prompts, agents and RAG apps with simple config files, compares models side by side, and also does red-team security testing.

Key points

Resources mentioned

Try this

More in Evaluation (Evals) & Testing

All of Evaluation (Evals) & Testing