5 Things AI Engineers Must Evaluate: RAG, Agents, Models, Data, Guardrails
Bashiri Smith · Facebook reel · 2026-09-30 · 0:33 · 11,414 views · Open on Facebook
Topics: Evaluation (Evals) & Testing, Retrieval-Augmented Generation (RAG), AI Safety, Security & Guardrails · Level: beginner
Summary
A quick checklist of the five parts of an AI system that engineers should evaluate, and the two metrics to check for each: RAG, agents, models, datasets and guardrails. The creator says evaluation is one of the most important skills for building production AI systems. He points viewers to his free AI Evaluation Guide and his BASWE AI Engineer community.
Key points
- RAG: check retrieval quality (did it find the right context?) and generation quality (is the answer good and grounded?) as separate things.
- Agents: measure task success (was the goal completed?) and tool accuracy (were the right tools called with the right arguments?).
- Models: measure output accuracy and efficiency (e.g., latency and cost).
- Datasets: check coverage (does the data represent the cases you care about?) and accuracy (are the labels and content correct?).
- Guardrails: track missed violations (harmful content that got through) and false blocks (safe content that was wrongly refused). This is a tradeoff between the two errors.
- The creator calls evaluation a top skill for production AI work and claims it pays top AI engineers over $300K.
Resources mentioned
- BASWE.Ai Engineer (Skool community) · community · skool.com · paid
The creator's paid community and program, with an AI learning roadmap (including the full ops and evaluation track), daily calls with engineers and recruiters, resume and portfolio help, and a job-search pipeline.
Also in: Basic RAG Pipeline in 60 Seconds: From Documents to Grounded Answers (Bashiri Smith on Facebook · notes), Pointer to Bashiri Smith's Complete AI Engineer Roadmap for 2026 (Bashiri Smith on Facebook · notes), Step-by-Step Roadmap to a $200K+ AI Engineering Role (Bashiri Smith on Facebook · notes), How to Evaluate a RAG Pipeline: Retrieval vs. Generation (Interview Answer) (Bashiri Smith on Facebook · notes) and 76 more
From the PDF shared here: Evaluation Field Guide (baswe.ai engineer accelerator – Ops and Evaluation module)
Open the original · 41 pages
Bashiri Smith's 41-page practitioner's guide to evaluating AI systems. It covers eval fundamentals (eval-driven development, golden datasets, metric types, basic statistics), deterministic and overlap metrics for LLM outputs, LLM-as-a-judge (rubric design, judge biases, checking judges against human labels), RAG evaluation (retrieval metrics and the RAG triad), agent trajectory evaluation, model benchmarking and fine-tune evaluation, and production evals (tracing, online signals, A/B tests, drift, guardrails, CI eval gates). It ends with a map of the eval tool stack, four portfolio projects and interview signals. Read it once end to end, then use it as a reference, and build one of the portfolio projects.
- Langfuse · tool · langfuse.com · free
Listed under platforms that combine tracing and evaluation.
Also in: Step-by-Step Roadmap to a $200K+ AI Engineering Role (Bashiri Smith on Facebook · notes), 7 Habits to Become an AI Engineer: Books, Tooling, Research & Shipping (Bashiri Smith on Facebook · notes), Taking a RAG App to Production: Evals, Guardrails, Cost and Tracing (Bashiri Smith on Facebook · notes) - LangSmith · tool · docs.smith.langchain.com · free
Tracing, logging and evals; also used to track cost and latency.
Also in: Step-by-Step Roadmap to a $200K+ AI Engineering Role (Bashiri Smith on Facebook · notes), 7 Habits to Become an AI Engineer: Books, Tooling, Research & Shipping (Bashiri Smith on Facebook · notes) - RAGAS · tool · github.com · free
Listed under RAG evaluation tools in the guide's evaluation stack.
Also in: SWE-to-AI Engineer Plan for 2027: LLMs, RAG, Agents, Evals, Job Search (Bashiri Smith on Facebook · notes), Taking a RAG App to Production: Evals, Guardrails, Cost and Tracing (Bashiri Smith on Facebook · notes) - OpenAI Evals · repo · github.com · free · recommended by both Bashiri Smith & Melvin Vivas
Listed under eval frameworks for defining and running evaluations.
Also in: Is OpenAI Evals Still Maintained? (Melvin Vivas on X · notes) - promptfoo · repo · github.com · free · recommended by both Bashiri Smith & Melvin Vivas
Listed under eval frameworks for defining and running evaluations, offline and in CI.
Also in: Running LLM evals with promptfoo on local llama.cpp models (Melvin Vivas on X · notes) - Arize Phoenix · tool · github.com · free
Listed under platforms that combine tracing and evaluation. - Braintrust · tool · braintrust.dev · free
Listed under platforms that combine tracing and evaluation. - DeepEval · tool · github.com · free
Listed under eval frameworks for defining and running evaluations, offline and in CI. - Guardrails AI · tool · github.com · free
Listed under guardrails tools for enforcing quality and safety on every request at runtime. - NeMo Guardrails · tool · github.com · free
Listed under guardrails tools for enforcing quality and safety on every request at runtime. - TruLens · tool · trulens.org · free
Listed under RAG evaluation tools in the guide's evaluation stack. - Weights & Biases Weave · tool · github.com · free
Listed under platforms that combine tracing and evaluation.
Try this
- Download the AI Evaluation Guide from the Google Drive link (or comment "eval" on the video to get it).
- For each AI system you build, pick metrics for all five areas: RAG, agents, models, datasets and guardrails.
- Optionally, join the BASWE AI Engineer community on Skool.