Topic 6 of 16 in the learning path
Retrieval-Augmented Generation (RAG)
Chunking, retrieval, reranking, hybrid/graph RAG and grounding answers in your data.
Reels and posts (29)
- Basic RAG Pipeline in 60 Seconds: From Documents to Grounded Answers · Bashiri Smith, Facebook · 0:55: A quick walkthrough of a basic Retrieval-Augmented Generation (RAG) pipeline that lets an LLM answer questions using private data.
- Using Jev as a Reranker to Augment RAG Retrieval · Melvin Vivas, X: The creator shares a quoted post from Rox.
- How to Evaluate a RAG Pipeline: Retrieval vs. Generation (Interview Answer) · Bashiri Smith, Facebook · 2:23: The video walks through a standard RAG pipeline: ingest documents, split them into chunks, create embeddings, store them in a vector database, run similarity search on the user's q
- LangChain vs. LangGraph Explained with One RAG Chatbot · Bashiri Smith, Facebook · 1:13: The video uses one RAG chatbot to show how LangChain and LangGraph differ.
- RAG Pre-Deployment Checklist: Evals, Grounding, Latency, Cost & Monitoring · Bashiri Smith, Facebook · 0:54: A short dialogue covering what to check before shipping a RAG app to production.
- Wannabe vs $200K+ AI Engineer: RAG, Agents, Context and Fine-Tuning Mistakes · Bashiri Smith, Facebook · 0:51: A quick comparison of how a beginner and a senior AI engineer approach the same tasks.
- How RAG Works: Chunking, Embedding, Vector Storage, and Retrieval · Bashiri Smith, Facebook · 1:12: Bashiri Smith calls RAG one of the most important skills for AI engineers right now.
- Using Jev for RAG as a Similarity Metric and Reranker · Melvin Vivas, X: The post quotes a claim that Jev's semantic matching is more reliable than plain dot-product similarity for RAG.
- jina-ocr-v1: Turning PDFs, Scans and Tables into Markdown · Melvin Vivas, X: Jina AI released jina-ocr-v1, a visual document parser post-trained on DeepSeek.
- Common AI Engineer Mistakes: Debugging RAG, Overbuilt Agents and Weak Evals · Bashiri Smith, Facebook · 1:02: The video compares "wannabe" habits with how experienced AI engineers work.
- Fixing RAG Ranking Problems with a Cross-Encoder Re-ranker · Bashiri Smith, Facebook · 1:10: This short is set up as a mock interview question: how do you fix a RAG app when the right chunk ranks below irrelevant ones?
- 3 Resume-Ready RAG Projects: Hybrid Search, Multi-Modal Docs and Agentic RAG · Bashiri Smith, Facebook · 0:05: The video suggests three RAG portfolio projects you can build in a few hours, each with a resume bullet you can adapt.
- How Amazon's Shopping Assistant Uses RAG: Query Planning, Retrieval, Generation · Bashiri Smith, Facebook · 1:06: A real-world case study of how Amazon's shopping assistant answers product questions like "Are these shoes good for hiking in the rain?" using retrieval-augmented generation.
- How RAG Works Under the Hood: Chunking, Embedding, Storage, Retrieval · Bashiri Smith, Facebook · 1:12: The creator calls RAG one of the most important AI engineering skills right now.
- Taking a RAG App to Production: Evals, Guardrails, Cost and Tracing · Bashiri Smith, Facebook · 1:10: A working RAG demo is only step one.
- RAG Chunking: Balancing Chunk Size for Precision and Context · Bashiri Smith, Facebook · 0:29: The video explains how bad chunking can break a RAG pipeline.
- Why Blind Top-K Retrieval Hurts RAG, and What to Use Instead · Bashiri Smith, Facebook · 1:11: The video explains that fixed Top-K retrieval always sends K chunks to the LLM, even when the lower-ranked chunks are weak, outdated or off-topic.
- Wannabe vs $200K AI Engineer: Fixing RAG, Agent Tools and Model Choice · Bashiri Smith, Facebook · 1:00: A short skit comparing a beginner's habits with a senior AI engineer's habits in three areas of production AI.
- Wannabe vs. Senior AI Engineer: Fixing RAG, Agent Tools and Model Choice Mistakes · Bashiri Smith, Facebook · 1:00: A quick skit that compares three beginner mistakes in production AI with how experienced AI engineers handle them.
- 3 Reasons a "Correct" RAG Pipeline Still Fails in Production · Bashiri Smith, Facebook · 1:03: A RAG pipeline can retrieve perfectly and still give wrong answers.
- 3 Reasons a "Correct" RAG Pipeline Still Fails in Production · Bashiri Smith, Facebook · 1:03: A RAG pipeline can retrieve well and still give bad answers.
- Choosing a Knowledge Strategy: RAG vs Graph vs Fine-Tuning vs CAG vs Long Context · Bashiri Smith, Facebook · 0:23: This short video covers five ways to give an LLM knowledge or change how it behaves, and matches each one to the problem it fits best: RAG, GraphRAG/knowledge graphs, fine-tuning,
- Production RAG Interview: Debugging Retrieval, Latency and Cost Like an Engineer · Bashiri Smith, Facebook · 1:47: This skit shows two candidates answering the same RAG system-design interview questions.
- RAG vs CAG: Retrieval vs Cache Augmented Generation Explained · Bashiri Smith, Facebook · 1:31: The video compares two ways to ground an LLM in your own documents.
- 7 AI Engineering Concept Pairs: RAG vs Fine-Tuning, Agents vs Workflows & More · Bashiri Smith, Facebook · 1:47: A quick comparison of seven pairs of AI engineering concepts that people often mix up.
- AI Engineer Interview Q&A: Graph RAG, Agent Memory, Observability, Guardrails · Bashiri Smith, Facebook · 1:32: A mock AI engineering interview with four questions and short model answers.
- Mistral OCR 4 launch, and the missing Qwen VL comparison · Melvin Vivas, X: Shares the launch of Mistral OCR 4, which outputs structured documents with bounding boxes, block classification and inline confidence scores across 170 languages.
- CocoIndex v1: Incremental Data Pipelines for AI and Agents · Melvin Vivas, X: CocoIndex has reached v1, a redesign of its framework for building incremental data pipelines.
- Q&A Over Your Own PDFs with LlamaIndex, OpenAI and Python (2023) · Melvin Vivas, melvinvivas.com: Melvin Vivas introduces generative text AI, the OpenAI API, LangChain and LlamaIndex (formerly GPT Index).
Watch, free (12)
- RAG from Scratch (freeCodeCamp) · video · youtube.com · free
A freeCodeCamp video about two and a half hours long on building basic RAG (the exact title isn't shown).
Mentioned in: 6 Free Videos to Move from Software Engineer to AI Engineer (Bashiri Smith on Facebook · notes), How to Relearn LLMs & RAG in 2026: A 7-Step Roadmap with Free Resources (Bashiri Smith on Facebook · notes) - RAG From Scratch (LangChain) · video · github.com · free
LangChain's free video series and notebooks that build retrieval-augmented generation step by step, from indexing through retrieval and generation.
Mentioned in: AI Engineer Roadmap Before 2027: Fundamentals, RAG, Agents, Ops, Evals (Bashiri Smith on Facebook · notes), AI Engineer Roadmap: Fundamentals, RAG, Agents, Books & Your First AI Service (Bashiri Smith on Facebook · notes) - 5 Levels of Text Splitting for Retrieval (Greg Kamradt) · video · youtube.com · free
A video by Greg Kamradt that walks through chunking and text-splitting strategies for RAG, from simple to advanced.
Mentioned in: How to Relearn LLMs & RAG in 2026: A 7-Step Roadmap with Free Resources (Bashiri Smith on Facebook · notes) - Advanced Retrieval for AI with Chroma (DeepLearning.AI) · course · deeplearning.ai · free
A DeepLearning.AI short course on advanced retrieval techniques that improve RAG results. Free with a DeepLearning.AI account during its platform beta; certificates are paid.
Mentioned in: How to Relearn LLMs & RAG in 2026: A 7-Step Roadmap with Free Resources (Bashiri Smith on Facebook · notes) - Build Production-Ready Retrieval RAG Pipeline in LangChain | Hybrid Search (BM25), Re-ranking & HyDE (Venelin Valkov) · video · youtube.com · free
A 15-minute video by Venelin Valkov on hybrid retrieval and advanced retrieval tactics.
Mentioned in: How to Relearn LLMs & RAG in 2026: A 7-Step Roadmap with Free Resources (Bashiri Smith on Facebook · notes) - Complete RAG tutorial 2026, with free labs · video · youtube.com · free
A full code-along crash course that goes from setup to a working RAG system.
In a shared PDF: How to become an expert in RAG (BASWE AI Engineer Field Guide) - shared in this reel on Facebook, this reel on Facebook - DeepLearning.AI Short Courses · course · deeplearning.ai · free
Applied courses of 1 to 2 hours on RAG, agents, LangChain and LLMOps.
In a shared PDF: The AI Pivot Field Guide - shared in this reel on Facebook, this reel on Facebook - LangChain Chat with Your Data · course · deeplearning.ai · free
A DeepLearning.AI course on building RAG with LangChain. Free with a DeepLearning.AI account during its platform beta; certificates are paid.
Mentioned in: How to Relearn LLMs & RAG in 2026: A 7-Step Roadmap with Free Resources (Bashiri Smith on Facebook · notes) - LangChain: RAG From Scratch (YouTube playlist) · video · youtube.com · free
Short-video series by Lance Martin that goes from indexing basics up to advanced routing and reranking; the guide calls it the gold standard.
In a shared PDF: How to become an expert in RAG (BASWE AI Engineer Field Guide) - shared in this reel on Facebook, this reel on Facebook - Preprocessing Unstructured Data for LLM Applications · course · deeplearning.ai · free
A DeepLearning.AI course on document processing and chunking. Free with a DeepLearning.AI account during its platform beta; certificates are paid.
Mentioned in: How to Relearn LLMs & RAG in 2026: A 7-Step Roadmap with Free Resources (Bashiri Smith on Facebook · notes) - RAG explained in 20 minutes · video · youtube.com · free
A fast, clear intro with a hands-on project; the quickest way to see the full RAG loop in action.
In a shared PDF: How to become an expert in RAG (BASWE AI Engineer Field Guide) - shared in this reel on Facebook, this reel on Facebook - RAG++: From POC to Production (Weights & Biases) · course · wandb.ai · free
RAG course from Weights & Biases named in the Class Central roundup of the best RAG courses.
In a shared PDF: How to become an expert in RAG (BASWE AI Engineer Field Guide) - shared in this reel on Facebook, this reel on Facebook
Read and use (23)
- LangChain · tool · github.com · free · recommended by both Bashiri Smith & Melvin Vivas
Open-source framework with ready-made integrations and common interfaces for connecting LLMs, embedding models, vector stores and tools.
Mentioned in: SWE-to-AI Engineer Plan for 2027: LLMs, RAG, Agents, Evals, Job Search (Bashiri Smith on Facebook · notes), Basic RAG Pipeline in 60 Seconds: From Documents to Grounded Answers (Bashiri Smith on Facebook · notes), Step-by-Step Roadmap to a $200K+ AI Engineering Role (Bashiri Smith on Facebook · notes), LangChain vs. LangGraph Explained with One RAG Chatbot (Bashiri Smith on Facebook · notes) and 4 more - How to become an expert in RAG (BASWE AI Engineer Field Guide) · pdf · drive.google.com · free
A one-page field guide from BASWE that lays out a 5-stage path to mastering retrieval-augmented generation.
Mentioned in: How RAG Works: Chunking, Embedding, Vector Storage, and Retrieval (Bashiri Smith on Facebook · notes), 3 Resume-Ready RAG Projects: Hybrid Search, Multi-Modal Docs and Agentic RAG (Bashiri Smith on Facebook · notes), How RAG Works Under the Hood: Chunking, Embedding, Storage, Retrieval (Bashiri Smith on Facebook · notes) - Unstructured · tool · unstructured.io · free
Document and data processing for building RAG over your domain documents.
Mentioned in: Basic RAG Pipeline in 60 Seconds: From Documents to Grounded Answers (Bashiri Smith on Facebook · notes)
In a shared PDF: The AI Pivot Field Guide - shared in this reel on Facebook, this reel on Facebook - Amazon Rufus (AI shopping assistant) · article · amazon.science · free
Amazon Science's write-up of how Rufus, Amazon's AI shopping assistant, uses RAG over the catalog, reviews and community Q&A - a real-world RAG case study.
Mentioned in: How Amazon's Shopping Assistant Uses RAG: Query Planning, Retrieval, Generation (Bashiri Smith on Facebook · notes) - arxiv-complete (Hugging Face dataset) · dataset · huggingface.co · free
The full arXiv archive (3.1M papers, every version, in LaTeX, PDF, PS and HTML, 16 TB) hosted on Hugging Face.
Mentioned in: The complete arXiv corpus as a Hugging Face dataset (Melvin Vivas on X · notes) - Class Central: 12 best RAG courses · article · classcentral.com · free
Ranked and tested list of RAG courses; the fastest way to find the Boot.dev, Weights & Biases RAG++, Duke and Google Cloud picks.
In a shared PDF: How to become an expert in RAG (BASWE AI Engineer Field Guide) - shared in this reel on Facebook, this reel on Facebook - CocoIndex · tool · cocoindex.io · free
Open-source framework for building incremental data transformation and indexing pipelines for AI and agents.
Mentioned in: CocoIndex v1: Incremental Data Pipelines for AI and Agents (Melvin Vivas on X · notes) - ColBERT · paper · arxiv.org · free
Late-interaction retrieval model that scores documents by matching their token embeddings against the query's token embeddings.
Mentioned in: Sentence Transformers v6: ColBERT-Style Multi-Vector Models Fully Supported (Melvin Vivas on X · notes) - DeepLearning.AI: Retrieval Augmented Generation · course · learn.deeplearning.ai · paid
Course taught by Zain Hasan with five hands-on modules. Module 1 is free to audit; the full course needs a paid upgrade. The same site hosts free short courses on agentic, knowledge-graph and multimodal RAG.
In a shared PDF: How to become an expert in RAG (BASWE AI Engineer Field Guide) - shared in this reel on Facebook, this reel on Facebook - Google Drive · tool · google.com · free
A cloud file storage service, used as a knowledge-base connector.
Mentioned in: Building a Nano Banana 2 Image-Gen App with Memex Managed AI Connectors (Melvin Vivas on X · notes) - Jina Reader · tool · jina.ai · free
Jina AI's API that converts URLs and documents to LLM-friendly text; it exposes jina-ocr-v1 via `x-respond-with`.
Mentioned in: jina-ocr-v1: Turning PDFs, Scans and Tables into Markdown (Melvin Vivas on X · notes) - jina-ocr-v1 · tool · huggingface.co · free
Jina AI's visual document parser that turns PDFs, scans, tables, charts and invoices into markdown.
Mentioned in: jina-ocr-v1: Turning PDFs, Scans and Tables into Markdown (Melvin Vivas on X · notes) - LangChain docs · docs · docs.langchain.com · free
Docs for LangChain, which the guide calls the most common LLM app framework.
In a shared PDF: The AI Pivot Field Guide - shared in this reel on Facebook, this reel on Facebook - Learn Retrieval Augmented Generation (Boot.dev) · course · boot.dev · paid
Boot.dev's RAG course (one of Class Central's picks): lessons are free to read, interactive exercises and certificates need a paid membership.
In a shared PDF: How to become an expert in RAG (BASWE AI Engineer Field Guide) - shared in this reel on Facebook, this reel on Facebook - llama_index.ipynb (Colab notebook gist) · repo · gist.github.com · free
The author's full Colab notebook source for the LlamaIndex + gpt-3.5-turbo example.
Mentioned in: Q&A Over Your Own PDFs with LlamaIndex, OpenAI and Python (2023) (Melvin Vivas on melvinvivas.com · notes) - LlamaIndex (llama_index) · repo · github.com · free
Open-source data framework that connects LLMs to external data for indexing and question answering.
Mentioned in: Q&A Over Your Own PDFs with LlamaIndex, OpenAI and Python (2023) (Melvin Vivas on melvinvivas.com · notes) - LlamaIndex Custom LLMs guide · docs · github.com · free · open in a browser to verify
LlamaIndex how-to on customizing the LLMPredictor/LLM; the link now appears broken.
Mentioned in: Q&A Over Your Own PDFs with LlamaIndex, OpenAI and Python (2023) (Melvin Vivas on melvinvivas.com · notes) - LlamaIndex docs · docs · docs.llamaindex.ai · free
Docs for LlamaIndex, a data framework for RAG over your own documents.
In a shared PDF: The AI Pivot Field Guide - shared in this reel on Facebook, this reel on Facebook - Mistral OCR 4 · tool · mistral.ai · paid
Mistral's OCR model that produces structured document output with bounding boxes, block types and confidence scores.
Mentioned in: Mistral OCR 4 launch, and the missing Qwen VL comparison (Melvin Vivas on X · notes) - Prompt Engineering with LlamaIndex and OpenAI GPT-3 · article · sausheong.com · free · open in a browser to verify
Sau Sheong's article on using LlamaIndex with GPT-3 to query your own documents.
Mentioned in: Q&A Over Your Own PDFs with LlamaIndex, OpenAI and Python (2023) (Melvin Vivas on melvinvivas.com · notes) - RAG-Expert-Guide-BASWE (RAG cheat sheet) · pdf · drive.google.com · free
The creator's full RAG cheat sheet, shared as a Google Drive document.
Mentioned in: Basic RAG Pipeline in 60 Seconds: From Documents to Grounded Answers (Bashiri Smith on Facebook · notes) - Sau Sheong · person · sausheong.com · free · open in a browser to verify
Engineer and writer who blogs about LLMs and software engineering.
Mentioned in: Q&A Over Your Own PDFs with LlamaIndex, OpenAI and Python (2023) (Melvin Vivas on melvinvivas.com · notes) - Turing Post: 10 RAG Courses in 2026 (Free, Trial, and Paid Options) · article · turingpost.com · free
Agentic, multimodal and knowledge-graph RAG learning paths gathered in one place.
In a shared PDF: How to become an expert in RAG (BASWE AI Engineer Field Guide) - shared in this reel on Facebook, this reel on Facebook
Build
- Build a basic RAG system over your own PDFs or company documents: Unstructured for extraction, LangChain text splitters for chunking, Cohere for embeddings, Pinecone or Weaviate for storage, then an LLM to generate answers. (from Basic RAG Pipeline in 60 Seconds: From Documents to Grounded Answers)
- Benchmark LLM-based retrieval against a dedicated reranker in your own RAG pipeline on latency, cost and accuracy. (from Using Jev as a Reranker to Augment RAG Retrieval)
- Build a RAG evaluation harness that scores retrieval precision and recall plus answer faithfulness, relevance and correctness against a test set with verified answers, then compare fixes such as adding a reranker, changing top-K, or using hybrid search. (from How to Evaluate a RAG Pipeline: Retrieval vs. Generation (Interview Answer))
- Build a RAG chatbot with LangChain (embedding model, vector DB and LLM), then add a LangGraph workflow. It should check retrieval quality, answer if results are good, rewrite the question and search again if they are bad, and pause for human input after too many failed attempts. (from LangChain vs. LangGraph Explained with One RAG Chatbot)
- Upgrade a basic RAG app to production quality with hybrid search, re-ranking, query rewriting and retrieval-quality evals. (from Wannabe vs $200K+ AI Engineer: RAG, Agents, Context and Fine-Tuning Mistakes)
- Build a LangGraph agent with narrowly scoped tools and deterministic steps instead of an open-ended LLM loop. (from Wannabe vs $200K+ AI Engineer: RAG, Agents, Context and Fine-Tuning Mistakes)
- Benchmark retrieval quality of dot-product similarity vs. Jev reranking on a small RAG dataset. (from Using Jev for RAG as a Similarity Metric and Reranker)
- Use jina-ocr-v1 to convert invoices or scanned PDFs to markdown as the ingestion step of a RAG pipeline (from jina-ocr-v1: Turning PDFs, Scans and Tables into Markdown)
- Build an evaluation harness for a RAG app that scores retrieval (whether the right chunks were found) and generation (whether the answer is grounded) separately. (from Common AI Engineer Mistakes: Debugging RAG, Overbuilt Agents and Weak Evals)
- Compare a simple workflow with a multi-agent researcher, writer and checker setup on accuracy, cost and latency. (from Common AI Engineer Mistakes: Debugging RAG, Overbuilt Agents and Weak Evals)
- Hybrid-search RAG combining dense vectors and BM25, with cross-encoder reranking and a pass that checks each citation supports its claim. (from 3 Resume-Ready RAG Projects: Hybrid Search, Multi-Modal Docs and Agentic RAG)
- Multi-modal document RAG: OCR plus LLM extraction from PDFs, tables and scanned documents, with a validation layer and a confidence-gated human review queue. (from 3 Resume-Ready RAG Projects: Hybrid Search, Multi-Modal Docs and Agentic RAG)
- Agentic RAG with a retrieval router that decides what to fetch, rewrites weak queries and retries when results are thin before answering. (from 3 Resume-Ready RAG Projects: Hybrid Search, Multi-Modal Docs and Agentic RAG)
- Build a product Q&A assistant that uses query planning, then retrieves from product specs, customer reviews and community Q&A before generating an answer. (from How Amazon's Shopping Assistant Uses RAG: Query Planning, Retrieval, Generation)
- Upgrade an existing RAG app into a production-level portfolio project, adding a Ragas eval suite, permission-aware retrieval, agent/token limits, prompt caching, model routing to smaller models, and Langfuse tracing. (from Taking a RAG App to Production: Evals, Guardrails, Cost and Tracing)
- Build an HR policy Q&A RAG bot (for example, answering PTO questions) and compare answer quality between fixed Top-5 retrieval and a pipeline that uses reranking, a relevance threshold, dynamic K and metadata filtering. (from Why Blind Top-K Retrieval Hurts RAG, and What to Use Instead)
- Upgrade a basic RAG app with retrieval-quality metrics and a reranker. Compare answer quality against simply retrieving more chunks. (from Wannabe vs. Senior AI Engineer: Fixing RAG, Agent Tools and Model Choice Mistakes)
- Build a model router that sends simple requests to a small model and harder ones to a large model, then measure the savings in cost and latency. (from Wannabe vs. Senior AI Engineer: Fixing RAG, Agent Tools and Model Choice Mistakes)
- Build a policy-QA RAG system with document lifecycle management (versioning, expiry, approval state) and automatic exclusion of superseded documents. (from 3 Reasons a "Correct" RAG Pipeline Still Fails in Production)
- Build a RAG app with a pre-retrieval clarification loop that detects low-intent queries and asks the user to clarify them. (from 3 Reasons a "Correct" RAG Pipeline Still Fails in Production)
- Build a multi-jurisdiction RAG assistant (federal, state and company policy) that uses metadata filtering to pick the applicable source. (from 3 Reasons a "Correct" RAG Pipeline Still Fails in Production)
- Build a policy Q&A RAG system that handles document versions and expiration, and filters by jurisdiction metadata (federal, state or company), with a clarification step for vague queries. (from 3 Reasons a "Correct" RAG Pipeline Still Fails in Production)
- Build a production-style RAG pipeline with modular ingestion, chunking, embedding, indexing, retrieval, re-ranking, generation and evaluation components, then show it from a repo in interviews. (from Production RAG Interview: Debugging Retrieval, Latency and Cost Like an Engineer)
- Add hybrid search, query rewriting, metadata filters and re-ranking to a RAG system, and measure how much each one improves retrieval. (from Production RAG Interview: Debugging Retrieval, Latency and Cost Like an Engineer)
- Optimize a RAG app for scale using caching, batched embeddings, parallel retrieval and routing requests to models by cost. (from Production RAG Interview: Debugging Retrieval, Latency and Cost Like an Engineer)
- Ask questions about your own resume: index your LinkedIn resume PDF and ask things like 'What is X's current role?' (from Q&A Over Your Own PDFs with LlamaIndex, OpenAI and Python (2023))
- A customer Q&A bot that answers questions about a company's services and products from its own documents (from Q&A Over Your Own PDFs with LlamaIndex, OpenAI and Python (2023))
- Natural-language querying over a company's unstructured data for customers and prospects (from Q&A Over Your Own PDFs with LlamaIndex, OpenAI and Python (2023))