DeepSeek V4.1 Flash Off-Peak Pricing as a Cheap Fallback Model
Melvin Vivas · X post · 2026-09-13 · Open on X
Topics: LLMOps, Deployment & Monitoring, LLM Fundamentals · Level: beginner
Summary
The creator calls DeepSeek V4.1 Flash good and very cheap, and suggests it as a fallback when other usage limits run out. The quoted post gives off-peak per-token prices, a speed of about 299 tokens/s, and the off-peak hours in Singapore time.
Key points
- Off-peak prices: $0.003 per 1M cached input tokens, $0.15 per 1M uncached input tokens, $0.60 per 1M output tokens.
- Average speed of about 299 tokens/s.
- Cached input is about 50x cheaper than uncached, so reuse prompt prefixes.
- Off-peak hours (Singapore time): before 9AM, 12–2PM and after 6PM on weekdays, plus all day Saturday and Sunday.
- Use it as a fallback model when your main tool's limits run out.
Resources mentioned
- DeepSeek V4.1 Flash · tool · api-docs.deepseek.com · paid
A fast DeepSeek language model available through DeepSeek's official API.
Also in: Using Hugging Face credits: Jobs, Inference Endpoints and Open Models (Melvin Vivas on X · notes), DeepSeek 4.1 Flash speed on the official API: about 325 tokens/s (Melvin Vivas on X · notes), One-Shot Agent Management UI with DeepSeek Harness for $0.19 (Melvin Vivas on X · notes), DeepSeek Harness: Open-Source, Browser-Based Agent for DeepSeek V4.1 (Melvin Vivas on X · notes) and 1 more
Try this
- Schedule heavy DeepSeek API workloads during off-peak hours.
- Design prompts to get cache hits on input tokens.
More in LLMOps, Deployment & Monitoring
- Running Local Models on NVIDIA DGX Sparks (Alex Ellis)
- Running Qwen3.8-27B EXL3 on an RTX 3090 with 262K Context at ~64 tok/s
- llama.cpp v0.4.1 release announcement
- How inference engines work: the full life of an LLM request
- SGLang v0.5.19 release: new models and beam search
- llama.cpp's built-in web UI for testing local models