Running Pi Coding Agent with a Local Qwen 3.6 Model
Melvin Vivas · X post · 2026-05-01 · Open on X
Topics: AI Agents, Tool Use & MCP, LLMOps, Deployment & Monitoring, AI Dev Tools & Productivity · Level: intermediate
Summary
The creator shares a report from r/LocalLLaMA about running the lightweight Pi coding agent with a local Qwen 3.6 (35B) model. The quoted post says local models work much better when the agent keeps the prefix cache intact and avoids a huge set of tools and a massive system prompt.
Key points
- You can run a coding agent fully locally by pairing the Pi coding agent with a local Qwen 3.6 35B model.
- Agents that keep invalidating the prefix (KV) cache run slowly on local hardware.
- Keep the system prompt short and the prompt prefix stable so the cache can be reused between turns.
- A small tool set works better for local models than a huge one.
- Lean agent harnesses suit local LLMs better than heavy ones built for cloud models.
Resources mentioned
- Pi · tool · x.com · free
A customizable coding agent that can be extended through its Extensions API and connected to several model providers.
Also in: Sign in with ChatGPT: Setting Usage Limits for Each App (Melvin Vivas on X · notes), Pi reaches v1.0 (Melvin Vivas on X · notes), Claude Code mods: customize behavior and UI with plugins (Melvin Vivas on X · notes), Customizing Your Coding Setup with Pi Coding Agent Extensions (Melvin Vivas on X · notes) and 39 more - Qwen 3.6 35B (MTP) · tool · huggingface.co · free
Qwen 3.6 mixture-of-experts model (35B total, about 3B active parameters) with multi-token prediction for local inference.
Also in: Use a Local Model for Confidential Data with Your Agent (Melvin Vivas on X · notes), Running Qwen 3.6 35B locally on an RTX 3090 for agent tool calling (Melvin Vivas on X · notes), Models That Work With the Hermes Agent for Personal Productivity (Melvin Vivas on X · notes), Run Hermes Agent locally with Qwen 3.6 35B MTP in LM Studio (Melvin Vivas on X · notes) and 4 more - r/LocalLLaMA – Been using Pi coding agent with local Qwen 3.6 35B · community · reddit.com · free
Reddit thread in the local-LLM community about using the Pi agent with a local Qwen 3.6 model.
Try this
- Try running the Pi coding agent with a local Qwen 3.6 model.
- Keep system prompts small and stable and limit tools when using local models.
- Set up a fully local coding agent (Pi + Qwen 3.6 35B) and compare speed with and without prefix-cache reuse.
More in AI Agents, Tool Use & MCP
- Multimodal AI Pointer: Combining Voice, Mouse Pointing and Vision with Gemini
- GPT-Realtime-2 & Realtime-Translate: Reasoning Voice Agents and Live Translation
- GPT-Realtime-2 & GPT-Realtime-Translate: Reasoning Voice Agents and Live Translation
- Cursor SDK: Build Agents on Cursor's Runtime, Harness and Models
- skillsbento: Free Open-Source Agent Skills for Claude
- Every Hugging Face Space Now Has AGENTS.md for Coding Agents