Low-Cost Agent Run: DeepSeek V4 Flash via OpenRouter in ohmypi
Melvin Vivas · X post · 2026-08-31 · Open on X
Topics: LLMOps, Deployment & Monitoring, AI Dev Tools & Productivity, LLM Fundamentals · Level: intermediate
Summary
A cost data point: a coding agent session in ohmypi using DeepSeek V4 Flash through OpenRouter cost about $0.30. It shows that cheap open models served through a gateway can make agentic coding very affordable.
Key points
- DeepSeek V4 Flash through OpenRouter cost about $0.30 for the session.
- The agent harness used was ohmypi.
- OpenRouter works as a single gateway to many models, including low-cost ones.
Resources mentioned
- OpenRouter · tool · openrouter.ai · free
A single OpenAI-compatible API that routes requests to hundreds of models from many providers, with one bill and model fallbacks.
Also in: Customizing Your Coding Setup with Pi Coding Agent Extensions (Melvin Vivas on X · notes), The Jev model is now on OpenRouter (Melvin Vivas on X · notes), Kev-4B model, an alternative to Jev, now available on OpenRouter (Melvin Vivas on X · notes), Space Bunny Alpha: stealth 1M-context flash model on OpenRouter (Melvin Vivas on X · notes) and 42 more - DeepSeek V4 Flash · tool · huggingface.co · free
DeepSeek model used as the comparison baseline in the tool-calling benchmark.
Also in: Running Codex with DeepSeek V4 Flash through OpenRouter (Melvin Vivas on X · notes), LFM2.5-2.6B Matches DeepSeek-V4-Flash on Tool Calling; LEAP Fine-Tuning (Melvin Vivas on X · notes), DeepSeek V4 Flash at 90% Off on Nous Portal (Melvin Vivas on X · notes), Qwen3.8-Max on the Frontend Code Arena cost-performance frontier (Melvin Vivas on X · notes) and 3 more - ohmypi · tool · github.com · free
A coding agent that lets you assign different models to different roles.
Also in: Assign Different Models per Role in ohmypi (Like Hermes) (Melvin Vivas on X · notes)
Try this
- Try DeepSeek V4 Flash through OpenRouter in your agent harness to cut costs.
More in LLMOps, Deployment & Monitoring
- Serving Qwen3.8-27B (EXL3) with 262K Context on a 24GB RTX 3090
- Ollama vs vLLM: From Local AI Demo to Production Inference Serving
- Running Qwen3.8 27B Locally on an RTX 3090 with llama.cpp and the Pi Harness
- Serving Qwen3.8 27B FP8 on an H100 with Baseten dedicated inference
- Running GLM 5.3 Flash on Baseten with the Pi Coding Agent
- AIBackends v0.7.0: Prompt Routing with Liquid AI's LFM2.5 Encoder