Q&A Over Your Own PDFs with LlamaIndex, OpenAI and Python (2023)
Melvin Vivas · melvinvivas.com article · 2023-04-08 · Open on melvinvivas.com
Topics: Retrieval-Augmented Generation (RAG), Embeddings & Vector Databases, LLM Fundamentals · Level: beginner
Summary
Melvin Vivas introduces generative text AI, the OpenAI API, LangChain and LlamaIndex (formerly GPT Index). He then walks through a basic Python RAG example in Google Colab. It indexes PDFs (his LinkedIn resume) from a /data folder, saves the vector index to disk and queries it in natural language. It also shows how to swap the default text-davinci-003 for the cheaper gpt-3.5-turbo. The code uses the early 2023 llama_index API (GPTSimpleVectorIndex, LLMPredictor), which has since been removed, so treat it as a lesson in concepts rather than code to copy.
Key points
- LlamaIndex (formerly GPT Index) is a simple interface that connects LLMs to your own external data. It used LangChain internally at the time.
- Basic flow: set OPENAI_API_KEY → SimpleDirectoryReader('data').load_data() → GPTSimpleVectorIndex.from_documents() → save_to_disk('index.json') → load_from_disk → index.query('Question?').
- The defaults were text-davinci-003 as the LLM and OpenAI Ada for embeddings. The embedding model could not be swapped for a cheaper one at the time.
- To cut costs, wrap ChatOpenAI(model_name='gpt-3.5-turbo', temperature=0, max_tokens=256) in an LLMPredictor and pass it through ServiceContext.from_defaults. The author says this is about 10x cheaper than davinci.
- PromptHelper settings: max_input_size=4096, num_output=256, max_chunk_overlap=20. Temperature 0 gives more conservative answers; higher values give more creative ones.
- Build the index once and reload it from disk so you don't pay for embedding API calls again. Only rebuild when your documents change.
- llm_predictor.last_token_usage prints the tokens used by a query, which helps you track cost.
- This code uses the deprecated 2023 llama_index API. Current LlamaIndex uses VectorStoreIndex and Settings instead, so check today's docs before running it.
Resources mentioned
- OpenAI API (announcement) · article · openai.com · free
OpenAI's blog post introducing public API access to its models. - OpenAI API · tool · platform.openai.com · paid · recommended by both Bashiri Smith & Melvin Vivas
OpenAI API documentation, including function calling and structured outputs.
Also in: How to Relearn LLMs & RAG in 2026: A 7-Step Roadmap with Free Resources (Bashiri Smith on Facebook · notes), Livestream: Building a React Native ChatGPT App with Cursor and OpenAI (Melvin Vivas on X · notes), GPT-Live-1: OpenAI's Full-Duplex Voice Model for Voice Agents in the API (Melvin Vivas on X · notes), The 3 Levels of AI Engineering: LLM Apps → Production → Agentic Systems (Bashiri Smith on Facebook · notes) and 4 more - GPT-3 powers the next generation of apps · article · openai.com · free
OpenAI blog post showcasing products built on GPT-3. - OpenAI GPT-3 Models (docs) · docs · developers.openai.com · free
OpenAI docs page on the GPT-3 model family; the link now appears broken. - OpenAI Models (docs) · docs · platform.openai.com · free
OpenAI documentation listing the available models and what they can do. - OpenAI Pricing · website · openai.com · free
OpenAI model pricing page (now redirects to ChatGPT pricing). - LangChain · tool · github.com · free · recommended by both Bashiri Smith & Melvin Vivas
Open-source framework with ready-made integrations and common interfaces for connecting LLMs, embedding models, vector stores and tools.
Also in: SWE-to-AI Engineer Plan for 2027: LLMs, RAG, Agents, Evals, Job Search (Bashiri Smith on Facebook · notes), Basic RAG Pipeline in 60 Seconds: From Documents to Grounded Answers (Bashiri Smith on Facebook · notes), Step-by-Step Roadmap to a $200K+ AI Engineering Role (Bashiri Smith on Facebook · notes), LangChain vs. LangGraph Explained with One RAG Chatbot (Bashiri Smith on Facebook · notes) and 3 more - LlamaIndex (llama_index) · repo · github.com · free
Open-source data framework that connects LLMs to external data for indexing and question answering. - LlamaIndex Custom LLMs guide · docs · github.com · free · open in a browser to verify
LlamaIndex how-to on customizing the LLMPredictor/LLM; the link now appears broken. - Prompt Engineering with LlamaIndex and OpenAI GPT-3 · article · sausheong.com · free · open in a browser to verify
Sau Sheong's article on using LlamaIndex with GPT-3 to query your own documents. - Sau Sheong · person · sausheong.com · free · open in a browser to verify
Engineer and writer who blogs about LLMs and software engineering. - llama_index.ipynb (Colab notebook gist) · repo · gist.github.com · free
The author's full Colab notebook source for the LlamaIndex + gpt-3.5-turbo example. - ChatGPT · tool · chatgpt.com · free · recommended by both Bashiri Smith & Melvin Vivas
OpenAI's AI assistant (web, desktop and mobile apps), used for chat, coding, writing and image generation.
Also in: Sign in with ChatGPT: Setting Usage Limits for Each App (Melvin Vivas on X · notes), Dots Demo: A Voice AI Agent Handling Travel Booking and Feedback Triage (Melvin Vivas on X · notes), ChatGPT Can Commit Code Directly to Your GitHub Repo (Melvin Vivas on X · notes), OpenAI Dots: Agents Inside ChatGPT That Debug, Fix, and Open PRs (Melvin Vivas on X · notes) and 23 more - gpt-3.5-turbo · tool · platform.openai.com · paid
OpenAI chat model; a cheaper alternative to text-davinci-003.
Also in: Natural Language to API Calls with LangChain APIChain and OpenAI (Melvin Vivas on melvinvivas.com · notes) - text-davinci-003 · tool · platform.openai.com · paid
OpenAI GPT-3 completion model; the default LLM in llama_index at the time (now deprecated). - OpenAI Ada embeddings · tool · platform.openai.com · paid
OpenAI embedding model used by llama_index by default. - Google Colab · tool · colab.research.google.com · free · recommended by both Bashiri Smith & Melvin Vivas
Free GPU notebooks.
Also in: Run Notebooks on a Free GPU with Google Colab (T4, 15GB VRAM) (Melvin Vivas on X · notes), Run Local Models on a Free GPU with Google Colab (T4) (Melvin Vivas on X · notes), Free Colab Notebooks for Local Model Fine-Tuning and Inference (Melvin Vivas on X · notes), Intent classification for support using GLiNER2.5-Decide notebook (Melvin Vivas on X · notes) and 12 more - Jupyter Notebook · tool · jupyter.org · free
An interactive notebook format (.ipynb) that mixes code and output. Colab can open these files.
Also in: AI DevBox v1.3.0: a GPU-ready Docker image with coding-agent CLIs (Melvin Vivas on X · notes), Run Local Models on a Free GPU with Google Colab (T4) (Melvin Vivas on X · notes) - donvito (GitHub) · repo · github.com · free
Melvin Vivas's GitHub profile listing the AI projects he is working on.
Also in: Melvin Vivas's GitHub Profile of AI Projects (Melvin Vivas on X · notes), Natural Language to API Calls with LangChain APIChain and OpenAI (Melvin Vivas on melvinvivas.com · notes)
Try this
- Get an OpenAI API key from the OpenAI Platform and set it as the OPENAI_API_KEY environment variable.
- Run the example in Google Colab, a Jupyter Notebook, or as a local Python script.
- Upload the PDFs you want to search into a /data folder.
- Build the index once, save it to disk, and reload it so you don't pay for embedding calls again.
- Use gpt-3.5-turbo instead of text-davinci-003 to cut costs, and check OpenAI's pricing page.
- Print last_token_usage to keep track of token cost per query.
- Open the author's Colab notebook gist to see the full code.
- Ask questions about your own resume: index your LinkedIn resume PDF and ask things like 'What is X's current role?'
- A customer Q&A bot that answers questions about a company's services and products from its own documents
- Natural-language querying over a company's unstructured data for customers and prospects