Running LLMs Locally Without an Expensive Rig
Melvin Vivas · X post · 2026-09-28 · Open on X
Topics: LLM Fundamentals, LLMOps, Deployment & Monitoring · Level: beginner
Summary
A beginner path to running LLMs locally on ordinary hardware. Start with Ollama or LM Studio and a small 3B–8B model on CPU, try quantized GGUF models, then move to llama.cpp once comfortable.
Key points
- You don't need a $5,000 AI rig to run models locally
- Install Ollama or LM Studio as an easy starting point
- Start with small models of 3B–8B parameters
- A CPU is fine for learning
- Try GGUF-format (quantized) models
- Move to llama.cpp once comfortable
- Start small, run something today, learn by experimenting
Resources mentioned
- Ollama · tool · x.com · free · recommended by both Bashiri Smith & Melvin Vivas
Open-source tool for downloading and running LLMs on your own machine with minimal setup.
Also in: Adding Vercel AI Gateway as a provider in AIBackends with Devin (Melvin Vivas on X · notes), Running Qwen3.8-27B Locally on an M5 Max MacBook with Inco Splash (Melvin Vivas on X · notes), AIBackends: An API Layer Between Your App and AI Models (Now with Jev) (Melvin Vivas on X · notes), Coworker: An Open-Source Desktop Agent App for Local and Cloud Models (Melvin Vivas on X · notes) and 12 more - LM Studio · tool · x.com · free
Desktop app for downloading and running LLMs locally, with a developer mode that serves models through an API.
Also in: Adding Vercel AI Gateway as a provider in AIBackends with Devin (Melvin Vivas on X · notes), LoRA Fine-Tune Qwen3.5-2B on Your Tweets with Unsloth Studio (Melvin Vivas on X · notes), AIBackends: An API Layer Between Your App and AI Models (Now with Jev) (Melvin Vivas on X · notes), Post-training Qwen3.5-2B on your own X posts with Unsloth Studio (Melvin Vivas on X · notes) and 31 more - llama.cpp · repo · github.com · free
An open-source C/C++ engine for running GGUF models locally. Its llama-server command provides an OpenAI-compatible HTTP server.
Also in: llama.cpp / Llama-macOS v0.5.0 release (Melvin Vivas on X · notes), Run llama.cpp GGUF Checkpoints in Hugging Face Transformers (Melvin Vivas on X · notes), llama.cpp v0.4.1 release announcement (Melvin Vivas on X · notes), Coworker: An Open-Source Desktop Agent App for Local and Cloud Models (Melvin Vivas on X · notes) and 28 more
Try this
- Install Ollama or LM Studio
- Run a small 3B–8B model on your CPU
- Try GGUF models
- Move to llama.cpp once comfortable
- Run something today and experiment
- Set up a local LLM on your own laptop with Ollama or LM Studio and compare several small GGUF models
More in LLM Fundamentals
- Use GPT-6.1 Sol by Default, Save Astra for Emergencies
- Use Sonnet 5.5 instead of Opus 5.5 for faster video-clipping tasks
- Guide: Building with Claude Sonnet 5.5
- The Jev model is now on OpenRouter
- How LLMs Work: A Motion-Graphics Explainer Made in One Shot with Claude Opus 5.5
- Kev-4B model, an alternative to Jev, now available on OpenRouter