1-month plan: the fast track
Build one portfolio project in four weeks: an assistant over your own documents that grows from a prompted LLM call to RAG, evals and tools, and ends as a deployed, traced API while you start applying. About 15 hours a week. The topics follow the learning path; other plans: 2-month plan · 3-month plan.
An unofficial study plan built from the public posts of Bashiri Smith and Melvin Vivas, not affiliated with or endorsed by them. Items under Fill the gap are not from the creators: they cover skills their posts don't. Track your progress in the library app; ticks are saved in your browser.
Week 1: LLMs, prompts and APIs
Goal: Explain how LLMs work and call an LLM API from Python that returns validated JSON, with tokens, latency and cost logged. (about 15 h in total)
Core
- [1hr Talk] Intro to Large Language Models (Andrej Karpathy) · Video · youtube.com
Do: Whole talk (60 min): inference, training stages, fine-tuning into an assistant, tool use, and the closing security part (jailbreaks, prompt injection) · about 1.5 h. - ChatGPT Prompt Engineering for Developers · Course · deeplearning.ai
Do: All 9 lessons (1 h 40 min); run each notebook as you go and skip the quiz · about 2 h. - OpenAI, Claude & Gemini API Tutorial in Python (Machine Learning Plus) · Article · machinelearningplus.com
Do: Whole tutorial incl. its 3 exercises: first calls to OpenAI, Claude and Gemini, error handling, a unified wrapper, multi-turn chat, streaming and cost estimation · about 1.5 h.
From the creators' posts
- How to Relearn LLMs & RAG in 2026: A 7-Step Roadmap with Free Resources · Bashiri Smith, Facebook · 1:54
- 3-Month Roadmap to AI Engineering for Experienced Software Engineers · Bashiri Smith, Facebook · 0:07
- 6 Steps to Build a Resume Project That Lands an AI Engineering Job · Bashiri Smith, Facebook · 2:47
- How GPT Works: From Token Embeddings to Multi-Head Attention · Melvin Vivas, X · 15:11
Build
- Start the capstone: pick a focused use case with 20-50 pages of real documents (ideally from a company you'd like to work for) and write 15-20 real questions with expected answers. Then build a Python CLI that calls an LLM API with a system prompt and returns Pydantic-validated JSON (answer, confidence, sources). It retries on invalid output and logs tokens, latency and cost per call; compare three prompt versions on your questions. (about 7.5 h) Idea from 3-Month Roadmap to AI Engineering for Experienced Software Engineers.
Fill the gap (not from the creators)
- Getting Structured LLM Output (DeepLearning.AI x .txt) · Course · deeplearning.ai
Vendor-neutral 1h21m course (updated Oct 2025) that moves from provider structured-output APIs with Pydantic to Instructor re-prompting and Outlines constrained decoding. Free with a DeepLearning.AI account during its platform beta. (about 2 h)
Optional
- The Complete AI Engineer Roadmap for 2026 (Exact Courses + Step-by-Step) · Video · youtube.com · the creator's own
- Prompting best practices (Claude Platform Docs) · Docs · platform.claude.com
Week 2: Embeddings and RAG
Goal: Explain embeddings and vector search, and run RAG over your own documents that cites sources and says when it doesn't know. (about 14 h in total)
Core
- Vector Databases: from Embeddings to Applications (DeepLearning.AI) · Course · deeplearning.ai
Do: All lessons (1 h 5 min) with the notebooks: vector representations, similarity search, approximate nearest neighbours (HNSW), vector databases, sparse/dense/hybrid search · about 1.5 h. - LangChain: RAG From Scratch (YouTube playlist) · Video · youtube.com
Do: Parts 1-9 (about 52 min: indexing, retrieval, generation and query translation: multi-query, RAG-Fusion, decomposition, step-back, HyDE), coding along with the notebooks in the rag-from-scratch repo. - How to become an expert in RAG (BASWE AI Engineer Field Guide) · PDF · drive.google.com · the creator's own
Do: The one-page guide: do stages 1-2 (mental model, base pipeline) this week and keep stage 3 (hybrid search, reranking, query rewriting) as stretch goals. Stage 4 (evals, cost, monitoring) comes in weeks 3-4 · about 15 min.
From the creators' posts
- How RAG Works Under the Hood: Chunking, Embedding, Storage, Retrieval · Bashiri Smith, Facebook · 1:12
- RAG Chunking: Balancing Chunk Size for Precision and Context · Bashiri Smith, Facebook · 0:29
- Why Blind Top-K Retrieval Hurts RAG, and What to Use Instead · Bashiri Smith, Facebook · 1:11
- How RAG Finds the Right Document Fast: Graph-Based Vector Search · Bashiri Smith, Facebook · 1:08
Build
- Add retrieval to the assistant: ingest and chunk your documents (with overlap), embed them into pgvector or Chroma, and retrieve the top matches. Answer with citations and reply 'I don't know' when nothing relevant comes back. Compare two chunk sizes and two top-k values on your question set and record which chunks were retrieved. (about 9 h) Idea from Basic RAG Pipeline in 60 Seconds: From Documents to Grounded Answers.
Fill the gap (not from the creators)
- Contextual Retrieval (Anthropic Engineering) · Article · anthropic.com
Shows how adding LLM-generated context to each chunk before embedding and BM25 indexing, plus reranking, cut failed retrievals by up to 67%, with a runnable cookbook. (about 30 min) - Step-by-Step Guide to Choosing the Best Embedding Model (Weaviate) · Article · weaviate.io
A 4-step process: define the use case, shortlist from MTEB by task, size, dimensions and max tokens, then build a small hand-labelled set (50-100 documents plus test queries) from your own data and compare models on precision and recall. (about 30 min)
Optional
- Complete RAG tutorial 2026, with free labs · Video · youtube.com
- Build Production-Ready Retrieval RAG Pipeline in LangChain | Hybrid Search (BM25), Re-ranking & HyDE (Venelin Valkov) · Video · youtube.com
Week 3: Evals, then tools and agents
Goal: Score your assistant on a golden test set, then turn it into a tool-calling agent with an MCP tool, without losing quality. (about 14.5 h in total)
Core
- Evaluation Field Guide (baswe.ai engineer accelerator – Ops and Evaluation module) · PDF · drive.google.com · the creator's own
Do: Sections 00-04 (why evals, eval-driven development and golden datasets, output metrics, LLM-as-a-judge, RAG evaluation); skim 05 (agents) and 07 (production) · about 1.5 h. - BASWE: Agent Cheatsheet for AI Engineers · PDF · drive.google.com · the creator's own
Do: The agentic design patterns PDF: Parts 1-2 (do the tool-use exercise) and sections 4.1-4.6 · about 2 h. - Model Context Protocol (MCP) · Docs · modelcontextprotocol.io
Do: 'What is MCP' and 'Architecture', then the 'Build an MCP server' quickstart in Python · about 1.5 h.
From the creators' posts
- How to Evaluate a RAG Pipeline: Retrieval vs. Generation (Interview Answer) · Bashiri Smith, Facebook · 2:23
- LangChain vs. LangGraph Explained with One RAG Chatbot · Bashiri Smith, Facebook · 1:13
- 7 AI Engineering Concept Pairs: RAG vs Fine-Tuning, Agents vs Workflows & More · Bashiri Smith, Facebook · 1:47
- Running Local Models for Agents: Tool Use, Context and Quantization · Melvin Vivas, X · 1:45
Build
- Grow your questions into a 25-30-case golden set, including 3-5 the documents can't answer. Score the week-2 RAG baseline on retrieval hit rate and on LLM-as-judge faithfulness and correctness, spot-checking 10 answers by hand. Then make it a tool-calling agent: search_docs plus one action tool that waits for your approval, with a step limit. Expose search_docs through a small MCP server and re-run the evals to compare. (about 9 h) Idea from How to Evaluate a RAG Pipeline: Retrieval vs. Generation (Interview Answer).
Fill the gap (not from the creators)
- The lethal trifecta for AI agents (Simon Willison) · Article · simonwillison.net
A 10-minute mental model of why an agent that has private data, untrusted content and an outbound channel can be hijacked to leak data, and why guardrail filters alone don't fix it. (about 15 min)
Optional
- Hugging Face Agents Course · Course · huggingface.co
- AI Agents vs AI Workflows: How to Choose for Production (AIBackends blog) · Article · aibackends.com · the creator's own
- promptfoo · Repo · github.com
Week 4: Ship it and start applying
Goal: Ship the assistant as a traced, containerized API, present it as a portfolio project, and start a targeted job search. (about 15 h in total)
Core
- FastAPI · Tool · fastapi.tiangolo.com
Do: Tutorial: First Steps, path/query parameters and request body (Pydantic models), then Deployment > 'FastAPI in Containers - Docker' · about 2 h. - Langfuse · Tool · langfuse.com
Do: The get-started guide for tracing: instrument your /ask endpoint so each request records the retrieval, model and tool calls with latency, tokens and cost · about 1 h. - The Interview Engine: The 7-Step System to Land AI Engineering Interviews · PDF · drive.google.com · the creator's own
Do: All 7 steps (12 pages): one-page resume, tailoring prompt, LinkedIn connections export, job tracker, warm paths, finding emails, outreach email and follow-ups. The PDF has the prompts and email template; the tracker spreadsheet is only in the paid community, so make your own (Jobs, Warm paths, Contacts). The free resume template is linked from Bashiri's roadmap video · about 1.5 h.
From the creators' posts
- RAG Pre-Deployment Checklist: Evals, Grounding, Latency, Cost & Monitoring · Bashiri Smith, Facebook · 0:54
- Taking a RAG App to Production: Evals, Guardrails, Cost and Tracing · Bashiri Smith, Facebook · 1:10
- Apply While You Build: Get Job-Search Feedback on Your AI Project Early · Bashiri Smith, Facebook · 1:12
- Find Hidden AI Jobs: Search by Skills (RAG, LLMs), Not "AI Engineer" Titles · Bashiri Smith, Facebook · 1:03
Build
- Ship it: serve the assistant behind a FastAPI /ask endpoint in Docker, deploy it to a free tier, and trace latency, tokens and cost per request with Langfuse. Publish a README (problem, architecture diagram, eval results, trade-offs) and a 2-minute demo, then add the project with its measured results to your resume. Send your first 10 targeted applications or outreach messages using the Interview Engine. (about 9.5 h) Idea from AI Engineer Roadmap for 2026 in 60 Seconds.
Fill the gap (not from the creators)
- Latency optimization guide (OpenAI API docs) · Docs · developers.openai.com
Seven provider-agnostic principles (process tokens faster, generate fewer tokens, use fewer input tokens, make fewer requests, parallelise, stream to cut perceived wait, don't default to an LLM) with a worked example. (about 30 min)
Optional
- 100 AI Engineering Interview Questions Guide · PDF · drive.google.com · the creator's own
- How Cursor cut agent token costs by 7% without losing quality · Melvin Vivas, X
Later
This plan leaves these topics for afterwards; they are all in the learning path: Programming & ML Foundations; Fine-tuning & Model Customization; AI Safety, Security & Guardrails; AI System Design & Architecture; AI Dev Tools & Productivity; Industry Trends & Job Market.