Running DeepSeek V4 Pro 0813 on Baseten with the Pi Agent
Melvin Vivas · X post · 2026-08-20 · Open on X
Topics: LLMOps, Deployment & Monitoring, AI Dev Tools & Productivity, Industry Trends & Job Market · Level: intermediate
Summary
The creator runs the new DeepSeek V4 Pro 0813 model through Baseten's inference platform and uses it with Pi as the agent harness. It is a short example of pairing an open model host with a coding agent.
Key points
- DeepSeek V4 Pro 0813 is available on Baseten.
- The creator pairs it with Pi (@pidotdev) for agent work.
- Hosted inference providers let you try new open models without running them yourself.
Resources mentioned
- DeepSeek V4 Pro 0813 · tool · baseten.co · paid
DeepSeek large language model, served here through Baseten Model APIs.
Also in: Running Hermes on DeepSeek V4 Pro via Baseten Model APIs (Melvin Vivas on X · notes) - Baseten · tool · x.com · paid
Model inference platform offering dedicated GPU deployments for serving open models.
Also in: Serving Qwen3.8 27B FP8 on an H100 with Baseten dedicated inference (Melvin Vivas on X · notes), Coworker Desktop App Works with Any OpenAI-Compatible Endpoint (Melvin Vivas on X · notes), Running GLM 5.3 Flash on Baseten with the Pi Coding Agent (Melvin Vivas on X · notes), GLM 5.3 Flash Now Available on Baseten (Melvin Vivas on X · notes) and 6 more - Pi · tool · x.com · free
A customizable coding agent that can be extended through its Extensions API and connected to several model providers.
Also in: Sign in with ChatGPT: Setting Usage Limits for Each App (Melvin Vivas on X · notes), Pi reaches v1.0 (Melvin Vivas on X · notes), Claude Code mods: customize behavior and UI with plugins (Melvin Vivas on X · notes), Customizing Your Coding Setup with Pi Coding Agent Extensions (Melvin Vivas on X · notes) and 39 more
Try this
- Try DeepSeek V4 Pro 0813 through Baseten in your agent setup.
More in LLMOps, Deployment & Monitoring
- How to Reduce Latency in a Production AI Agent (Interview Answer)
- Serving Ornith-1.5-35B-A3B at 128k Context on an RTX 3090 with llama.cpp
- Running Ornith-1.5-35B-A3B (Q4_K_M) at 128k Context on an RTX 3090 with llama.cpp
- Red Hat's Quantized Qwen3.8-2.4T-A95B Variant (NVFP4/FP8)
- One-Click Deploy Link for Qwen3.8 27B on Baseten
- Self-Hosting Qwen3.8 27B FP8 on an H100 for $6.50/hr