Demo: Qwen3.8-27B Running Locally with Pi and llama.cpp
Melvin Vivas · X video post · 2026-08-14 · 0:28 · 502 views · Open on X
Topics: AI Dev Tools & Productivity, LLM Fundamentals, Industry Trends & Job Market · Level: intermediate
Summary
A short reaction post with a silent 28-second video. It shows output made by Alibaba's Qwen3.8-27B open-weight model, which was run locally with llama.cpp and driven by Pi (@pidotdev). There is no narration or tutorial. Its main value is pointing to a local open-model coding setup you can try yourself.
Key points
- The quoted post calls this the 'first test' of Qwen3.8-27B, a 27-billion-parameter model from Alibaba's Qwen team.
- The model was run locally with llama.cpp, an open-source engine for running LLMs (usually quantized GGUF files) on your own hardware.
- Pi (@pidotdev) was the tool used to drive the model and produce the output.
- The creator's reaction ('wtf') suggests a mid-sized local model gave surprisingly strong results. The video has no speech, so no specific numbers or steps are given.
- Takeaway: open-weight models of about 27B parameters running locally are now an option for agent-style work without a cloud API.
Resources mentioned
- Qwen3.8-27B · tool · huggingface.co · free
A 27B-parameter open-weight model from Alibaba's Qwen family that you can run locally or call through hosted APIs.
Also in: Deploying Qwen3.8 27B on Hugging Face Inference Endpoints (Melvin Vivas on X · notes), Using Hugging Face credits: Jobs, Inference Endpoints and Open Models (Melvin Vivas on X · notes), Running Qwen3.8-27B Locally on an M5 Max MacBook with Inco Splash (Melvin Vivas on X · notes), Ternary Bonsai 2 27B: 9x smaller model keeping 98.2% of benchmark scores (Melvin Vivas on X · notes) and 24 more - Qwen (@Alibaba_Qwen) on X · person · x.com · free
Official X account of Alibaba's Qwen team, which posts model releases and demos.
Also in: Qwen-Audio-3.1: Alibaba's Five-Model Audio Stack (Melvin Vivas on X · notes), Running AI Locally: Rebuilding AIBackends as a Python Library with Open Models (Melvin Vivas on X · notes), Audio-Visual Vibe Coding with Qwen 3.5 Omni: Spoken Specs to Web App (Melvin Vivas on X · notes) - Pi · tool · x.com · free
A customizable coding agent that can be extended through its Extensions API and connected to several model providers.
Also in: Sign in with ChatGPT: Setting Usage Limits for Each App (Melvin Vivas on X · notes), Pi reaches v1.0 (Melvin Vivas on X · notes), Claude Code mods: customize behavior and UI with plugins (Melvin Vivas on X · notes), Customizing Your Coding Setup with Pi Coding Agent Extensions (Melvin Vivas on X · notes) and 39 more - llama.cpp · repo · github.com · free
An open-source C/C++ engine for running GGUF models locally. Its llama-server command provides an OpenAI-compatible HTTP server.
Also in: Running LLMs Locally Without an Expensive Rig (Melvin Vivas on X · notes), llama.cpp / Llama-macOS v0.5.0 release (Melvin Vivas on X · notes), Run llama.cpp GGUF Checkpoints in Hugging Face Transformers (Melvin Vivas on X · notes), llama.cpp v0.4.1 release announcement (Melvin Vivas on X · notes) and 28 more
Try this
- Run Qwen3.8-27B locally with llama.cpp, connect it to Pi, and compare its output on a coding task with a cloud model.
More in AI Dev Tools & Productivity
- Cursor iPhone App Adds Plan Mode
- Cursor Workflow: Start an Agent on Mobile, Continue in the Cloud
- Split AI-agent code into stacked, focused PRs
- Plan with a Strong Model, Execute with a Cheaper One
- Grok access: X Premium+ isn't enough, you need SuperGrok
- ChatGPT Desktop App (with Codex) Now in Preview for Linux