Running GLM 5.3 Flash on Baseten with the Pi Coding Agent
Melvin Vivas · X post · 2026-08-28 · Open on X
Topics: LLMOps, Deployment & Monitoring, AI Dev Tools & Productivity · Level: intermediate
Summary
The creator runs the GLM 5.3 Flash model through Baseten's inference platform and uses it inside the Pi coding agent. He says it generates output faster than he can read it.
Key points
- GLM 5.3 Flash is served on Baseten with very high generation speed.
- He pairs it with the Pi coding agent for coding work.
- Fast hosted open models can power agentic coding tools.
Resources mentioned
- GLM 5.3 Flash · tool · huggingface.co · free
A fast model from Z.ai's GLM family, used here for agentic coding.
Also in: Coworker Desktop App Works with Any OpenAI-Compatible Endpoint (Melvin Vivas on X · notes) - Baseten · tool · x.com · paid
Model inference platform offering dedicated GPU deployments for serving open models.
Also in: Serving Qwen3.8 27B FP8 on an H100 with Baseten dedicated inference (Melvin Vivas on X · notes), Coworker Desktop App Works with Any OpenAI-Compatible Endpoint (Melvin Vivas on X · notes), GLM 5.3 Flash Now Available on Baseten (Melvin Vivas on X · notes), GLM-5.2 Vision on Baseten: Turning Images into Code (Melvin Vivas on X · notes) and 6 more - Pi · tool · x.com · free
A customizable coding agent that can be extended through its Extensions API and connected to several model providers.
Also in: Sign in with ChatGPT: Setting Usage Limits for Each App (Melvin Vivas on X · notes), Pi reaches v1.0 (Melvin Vivas on X · notes), Claude Code mods: customize behavior and UI with plugins (Melvin Vivas on X · notes), Customizing Your Coding Setup with Pi Coding Agent Extensions (Melvin Vivas on X · notes) and 39 more
More in LLMOps, Deployment & Monitoring
- Running Qwen3.8 27B Locally on an RTX 3090 with llama.cpp and the Pi Harness
- Low-Cost Agent Run: DeepSeek V4 Flash via OpenRouter in ohmypi
- Serving Qwen3.8 27B FP8 on an H100 with Baseten dedicated inference
- AIBackends v0.7.0: Prompt Routing with Liquid AI's LFM2.5 Encoder
- GLM 5.3 Flash Now Available on Baseten
- Zero-Shot Prompt Routing with Liquid AI's LFM 2.5-Encoder-350M