Running Hermes on DeepSeek V4 Pro via Baseten Model APIs
Melvin Vivas · X post · 2026-08-20 · Open on X
Topics: LLM Fundamentals, LLMOps, Deployment & Monitoring · Level: beginner
Summary
The creator used leftover Baseten Model API credits to run his Hermes setup on DeepSeek V4 Pro 0813, and says he likes it. It shows that you can point an agent harness at any hosted model API provider.
Key points
- Baseten offers hosted Model APIs that you pay for with credits.
- Hermes can be set to use a different backing model, here DeepSeek V4 Pro 0813.
- Check your provider accounts for unused credits before paying for something new.
Resources mentioned
- Baseten · tool · x.com · paid
Model inference platform offering dedicated GPU deployments for serving open models.
Also in: Serving Qwen3.8 27B FP8 on an H100 with Baseten dedicated inference (Melvin Vivas on X · notes), Coworker Desktop App Works with Any OpenAI-Compatible Endpoint (Melvin Vivas on X · notes), Running GLM 5.3 Flash on Baseten with the Pi Coding Agent (Melvin Vivas on X · notes), GLM 5.3 Flash Now Available on Baseten (Melvin Vivas on X · notes) and 6 more - DeepSeek V4 Pro 0813 · tool · baseten.co · paid
DeepSeek large language model, served here through Baseten Model APIs.
Also in: Running DeepSeek V4 Pro 0813 on Baseten with the Pi Agent (Melvin Vivas on X · notes) - Hermes · tool · hermes-agent.nousresearch.com · free
The AI agent the creator uses to automate making explainer videos. It is probably Nous Research's Hermes Agent, but the post does not say so.
Also in: OpenAI DevDay recap: Dots, GPT-6.1 Sol, Codex Cloud, Agents API (Melvin Vivas on X · notes), Coworker: free open-source desktop AI coworker app (Melvin Vivas on X · notes), Agent Monitor: see traces, tokens and costs of your coding agents (Melvin Vivas on X · notes), Coworker: An Open-Source Desktop Agent App for Local and Cloud Models (Melvin Vivas on X · notes) and 13 more
More in LLM Fundamentals
- Try the Free Stealth Model Ox Alpha on OpenRouter
- GPT Image 2 Adds Transparent Background Support in the OpenAI API
- Ornith-1.5: Open-Source LLM Family (9B Dense to 397B MoE)
- Find Discounted Models with OpenRouter's New Filter
- Cost-Aware Model Routing in Codex: Default to Luna, Escalate to Sol
- Qwen 3.8 27B Matches GPT 5.6 Luna (max) on Artificial Analysis