Running Local Models for Agents: Tool Use, Context and Quantization
Melvin Vivas · X video post · 2026-06-18 · 1:45 · 73 views · Open on X
Topics: LLM Fundamentals, AI Agents, Tool Use & MCP, LLMOps, Deployment & Monitoring · Level: intermediate
Summary
Melvin Vivas explains why leaderboard scores don't tell you whether a self-hosted model can work inside an agent workflow like OpenClaw. He covers the three things that matter most: reliable tool calling, enough context window (16k tokens at the very least), and how far you can quantize before quality drops. The main lesson is to plan the whole setup together: model size, quantization level, VRAM, and the context you have left once the model is loaded.
Key points
- Leaderboard scores don't matter if a model can't work inside an agent workflow. Test it in your actual agent loop.
- Tool use comes first: the model has to call the right tool at the right time with the right arguments, then understand the result. Weaker models fail here right away.
- Agents use up context fast. Instructions, tool schemas, past messages, file contents, command outputs and errors all pile up before the actual task is added.
- For OpenClaw, treat a 16k-token context window as the absolute minimum, not an ideal. More is better.
- Quantization compresses model weights by storing them with fewer bits. That uses less VRAM and lets the model run on cheaper hardware.
- If you quantize too aggressively, the model gets worse at following instructions, using tools and staying coherent over long tasks. In an agent, that can break the whole workflow: wrong tool calls, missed instructions, lost context, or confidently going down the wrong path.
- Size your setup as a whole: model size + quantization level + available VRAM + context you can still allocate after the model is loaded. It all has to fit in GPU memory.
- A Mac Mini running a local model won't simply replace Claude or OpenAI. Being useful in practice takes more than the size of the weights.
Resources mentioned
- OpenClaw · tool · github.com · free
An open-source, self-hostable personal AI agent that can run on local or hosted LLMs.
Also in: Run Muse Glimmer 30B Locally on an RTX 3090 with llama.cpp (Melvin Vivas on X · notes), Running Muse Glimmer 30B Locally with llama.cpp and the Hermes Agent (Melvin Vivas on X · notes), Run Muse Glimmer 30B Locally with llama.cpp and Connect It to Hermes Agent (Melvin Vivas on X · notes), Ollama Is Now an Official Provider for OpenClaw (Melvin Vivas on X · notes) and 2 more - Claude Opus 5.5 · tool · claude.ai · paid · open in a browser to verify · recommended by both Bashiri Smith & Melvin Vivas
Anthropic's AI assistant, used throughout the guide to tailor resumes, add live roles to the tracker, match connections to target companies and find hiring managers.
Also in: Generating a Repo Promo Video with a Claude Skill on Sonnet 5.5 vs Opus 5.5 (Melvin Vivas on X · notes), Comparing Coding Models on the Same Task in Devin iOS (Melvin Vivas on X · notes), Customizing Your Coding Setup with Pi Coding Agent Extensions (Melvin Vivas on X · notes), AI Engineer Roadmap Overview: From ML Foundations to RAG, Agents & Ops (Bashiri Smith on Facebook · notes) and 53 more - OpenAI · tool · x.com · free
An AI model provider whose models can be used through managed connectors.
Also in: Creator's favorite OpenAI DevDay announcements (Melvin Vivas on X · notes), OpenAI Agents API with Bring Your Own Sandbox as a backend for a 'software factory' (Melvin Vivas on X · notes), Building Software Factories on the OpenAI Agents API (Codex-Powered) (Melvin Vivas on X · notes), Coworker: Free Local-First Desktop App for Running AI Agents (Melvin Vivas on X · notes) and 11 more - Mac Mini · tool · apple.com · paid
Apple desktop computer popular as always-on local hardware for personal AI agents.
Also in: Hosting a Personal AI Agent: Local Mac Mini vs Cloud VM Privacy Trade-off (Melvin Vivas on X · notes)
Try this
- Judge local models on how they perform inside your agent workflow, not on leaderboard scores.
- Test whether the model calls the right tools with correct arguments and understands what the tools return.
- Use a context window of at least 16k tokens for OpenClaw-style agents, and more if you can.
- Pick a quantization level that doesn't hurt instruction-following and tool use. Avoid compressing too aggressively.
- Plan model size, quantization, VRAM and leftover context together so everything fits in GPU memory.
More in LLM Fundamentals
- Free GLM-5.2 via Hugging Face Inference Providers in coding agents
- Gemini 3.5 Live Translate: Real-Time Speech Translation via the Live API
- GLM-5.2 released: open weights, 1M context, two reasoning levels
- GLM-5.2 by Z.ai: open-weights model with a 1M-token context
- Gemini 3.5 Live Translate: Real-Time Speech Translation via the Gemini Live API
- Free gpt-oss-20b and Gemma 4 26B on OpenRouter