Running Qwen 3.5 locally with Claude Code on an RTX 3090
Melvin Vivas · X post · 2026-03-05 · Open on X
Topics: AI Dev Tools & Productivity, LLM Fundamentals · Level: intermediate
Summary
The creator reports that the open Qwen 3.5 model runs on a consumer RTX 3090 GPU with 24 GB of VRAM. He says it works well as the model behind Claude Code and can use agent skills. This suggests local open models are now a workable option for agentic coding.
Key points
- Qwen 3.5 can run locally on an RTX 3090 with 24 GB VRAM
- Qwen 3.5 works well when connected to Claude Code
- The local model can use Claude Code skills
- A consumer GPU is enough for local agentic coding
Resources mentioned
- Qwen 3.5 · tool · qwen.ai · free
Alibaba Qwen open model family with small on-device variants and toggleable reasoning.
Also in: Workshop AI: Building Apps with Cloud and Local Agents (GLM 5, Qwen 3.5) (Melvin Vivas on X · notes), Qwen3.5 0.8B Does Real-Time Local Video Captioning (Melvin Vivas on X · notes), Qwen3.5 0.8B Real-Time Video Captioning on Mac Studio (Melvin Vivas on X · notes), Qwen 3.5 Is the Local Model That Holds Up in Claude Code (Melvin Vivas on X · notes) and 7 more - Claude Code · tool · code.claude.com · paid · recommended by both Bashiri Smith & Melvin Vivas
Build agents and pipelines from the terminal; the guide's main agentic coding tool.
Also in: Create Claude Code Plugins with /plugin-authoring (Melvin Vivas on X · notes), Claude Code mods: customize behavior and UI with plugins (Melvin Vivas on X · notes), AI Engineer Roadmap Overview: From ML Foundations to RAG, Agents & Ops (Bashiri Smith on Facebook · notes), SkillsBento: Free Plugin Marketplace for Codex and Claude Code (Melvin Vivas on X · notes) and 101 more - NVIDIA GeForce RTX 3090 · tool · nvidia.com · paid
Consumer GPU with 24 GB of VRAM, used here to run a long-context LLM locally.
Also in: Portable Computer Now Runs AI Agents and Models Locally on NVIDIA RTX PCs (Melvin Vivas on X · notes), Long-Context Local LLM on an RTX 3090: .env Config and ~64 tok/s Benchmark (Melvin Vivas on X · notes), Running Ornith-1.5-35B-A3B (Q4_K_M) at 128k Context on an RTX 3090 with llama.cpp (Melvin Vivas on X · notes), Generating AI Video Locally with MiniMax H3 in ComfyUI on an RTX 3090 (Melvin Vivas on X · notes) and 2 more
Try this
- Consider buying a GPU with 24 GB VRAM to run local models
- Try connecting a locally hosted Qwen 3.5 to Claude Code
- Set up a local coding agent: Qwen 3.5 served on your own GPU and driven by Claude Code with skills
More in AI Dev Tools & Productivity
- Kling Video 3.0 Motion Control Is Available on fal
- OpenAI Codex App Now Available on Windows
- Kling Motion Control 3.0: Motion Capture with Better Facial ID Consistency
- v0 vs Claude Code for cloning an app's design
- When one coding model gets stuck, switch to another
- 35% of PRs Are Coded by Cursor's Autonomous Agents