Move Non-Coding Work to Local Models (Hermes + llama.cpp)
Melvin Vivas · X post · 2026-08-21 · Open on X
Topics: LLMOps, Deployment & Monitoring, AI Dev Tools & Productivity, AI Agents, Tool Use & MCP · Level: intermediate
Summary
The creator argues for moving some AI work to local models so you depend less on cloud providers and their changing limits. He plans to start with personal and non-coding tasks, using Hermes connected to a local llama.cpp server, and to keep cloud AI for coding.
Key points
- Running some work on local models reduces your dependence on cloud AI services and their usage limits.
- Start with non-coding and personal tasks, which are lower stakes and suit local models.
- Connect an agent (Hermes) to a model served locally with llama.cpp.
- Keep your paid cloud AI budget for coding, where stronger models matter most.
Resources mentioned
- Hermes · tool · hermes-agent.nousresearch.com · free
The AI agent the creator uses to automate making explainer videos. It is probably Nous Research's Hermes Agent, but the post does not say so.
Also in: OpenAI DevDay recap: Dots, GPT-6.1 Sol, Codex Cloud, Agents API (Melvin Vivas on X · notes), Coworker: free open-source desktop AI coworker app (Melvin Vivas on X · notes), Agent Monitor: see traces, tokens and costs of your coding agents (Melvin Vivas on X · notes), Coworker: An Open-Source Desktop Agent App for Local and Cloud Models (Melvin Vivas on X · notes) and 13 more - llama.cpp · repo · github.com · free
An open-source C/C++ engine for running GGUF models locally. Its llama-server command provides an OpenAI-compatible HTTP server.
Also in: Running LLMs Locally Without an Expensive Rig (Melvin Vivas on X · notes), llama.cpp / Llama-macOS v0.5.0 release (Melvin Vivas on X · notes), Run llama.cpp GGUF Checkpoints in Hugging Face Transformers (Melvin Vivas on X · notes), llama.cpp v0.4.1 release announcement (Melvin Vivas on X · notes) and 28 more
Try this
- Set up a local model server with llama.cpp.
- Connect an agent such as Hermes to it and move personal and non-coding tasks there.
- Keep cloud AI subscriptions for coding.
- Build a local personal assistant: an agent framework connected to a llama.cpp server for everyday non-coding tasks.
More in LLMOps, Deployment & Monitoring
- Novita AI Spot GPU Instances: Cheap GPU Compute
- Why Codex Usage Drains Faster: Prompt Cache Hit Rate
- GPT-5.6 Sol Price Cuts on Vercel AI Gateway (70% Off Input)
- How to Reduce Latency in a Production AI Agent (Interview Answer)
- Serving Ornith-1.5-35B-A3B at 128k Context on an RTX 3090 with llama.cpp
- Running Ornith-1.5-35B-A3B (Q4_K_M) at 128k Context on an RTX 3090 with llama.cpp