Models That Work With the Hermes Agent for Personal Productivity
Melvin Vivas · X post · 2026-07-22 · Open on X
Topics: AI Agents, Tool Use & MCP, LLM Fundamentals, AI Dev Tools & Productivity · Level: intermediate
Summary
The creator lists the LLMs he swaps between inside the Hermes agent for personal productivity tasks (not coding). He picks models by availability, cost and job: a daily driver, a fallback, a coding plan, OpenRouter access, a local model and a separate vision model. He says Hermes' harness is good enough that switching models doesn't feel very different.
Key points
- Daily driver: Grok 4.5.
- Fallback when Grok usage runs out: GPT 5.5.
- GLM 5.2 through the Z.ai coding plan.
- DeepSeek V4 through OpenRouter.
- Qwen 3.6 35B (MTP) runs locally when he isn't using his PC.
- Gemini Flash is a helper model for vision tasks.
- A good agent harness makes the choice of model matter less, so you can swap models for cost or availability.
- His use case is personal productivity, not coding.
Resources mentioned
- Hermes · tool · hermes-agent.nousresearch.com · free
The AI agent the creator uses to automate making explainer videos. It is probably Nous Research's Hermes Agent, but the post does not say so.
Also in: OpenAI DevDay recap: Dots, GPT-6.1 Sol, Codex Cloud, Agents API (Melvin Vivas on X · notes), Coworker: free open-source desktop AI coworker app (Melvin Vivas on X · notes), Agent Monitor: see traces, tokens and costs of your coding agents (Melvin Vivas on X · notes), Coworker: An Open-Source Desktop Agent App for Local and Cloud Models (Melvin Vivas on X · notes) and 13 more - Grok 4.5 · tool · x.ai · paid
An LLM said to be trained in partnership with SpaceXAI, pitched as a general model beyond software engineering.
Also in: Grok 4.5 Works Well as the Model Behind Hermes Agent (Melvin Vivas on X · notes), Agent-Made Video in 10 Minutes: Hermes Agent + Grok 4.5 + Hyperframes (Melvin Vivas on X · notes), Personal Assistant Agent on Hermes: Morning Briefings and Inbox Triage (Melvin Vivas on X · notes), Running Parallel Agents With Different Models in Cursor Mobile (Melvin Vivas on X · notes) and 9 more - GPT 5.5 · tool · openai.com · paid
OpenAI models the creator used as the coding model inside Cursor.
Also in: Running GPT-5.5 via Codex as Hermes's Main Model (Melvin Vivas on X · notes), Conductor Walkthrough: Running Parallel Coding Agents in Isolated Git Worktrees (Melvin Vivas on X · notes), Composer 2.5 as the Default Coding Model in Cursor (Melvin Vivas on X · notes), Which AI coding model to use for which task (Melvin Vivas on X · notes) and 2 more - GLM 5.2 · tool · huggingface.co · free
A GLM-family LLM (the transcript says 'GLM-5-2') that the demo ranked as a strong, affordable model for coding and design.
Also in: Qwen3.8-Max on the Frontend Code Arena cost-performance frontier (Melvin Vivas on X · notes), Open Model GLM 5.2 Helped Mitigate OpenAI-Caused Cyberattack (Melvin Vivas on X · notes), AI News Roundup: Claude Fable 5, Scientist AI, ZCode, NVIDIA RL (Melvin Vivas on X · notes), Trying ZCode by Z.ai with GLM 5.2 on a Mac (Melvin Vivas on X · notes) and 27 more - Z.ai Coding Plan · tool · docs.z.ai · paid
Z.ai's subscription plan for accessing GLM models. - DeepSeek V4 · tool · huggingface.co · free
A DeepSeek LLM, accessed through OpenRouter as the final fallback.
Also in: Multi-Teacher On-Policy Distillation (MOPD) in 2026 (Melvin Vivas on X · notes), Cost-Saving Model Fallback Chain: Grok → Codex → OpenRouter DeepSeek (Melvin Vivas on X · notes) - OpenRouter · tool · openrouter.ai · free
A single OpenAI-compatible API that routes requests to hundreds of models from many providers, with one bill and model fallbacks.
Also in: Customizing Your Coding Setup with Pi Coding Agent Extensions (Melvin Vivas on X · notes), The Jev model is now on OpenRouter (Melvin Vivas on X · notes), Kev-4B model, an alternative to Jev, now available on OpenRouter (Melvin Vivas on X · notes), Space Bunny Alpha: stealth 1M-context flash model on OpenRouter (Melvin Vivas on X · notes) and 42 more - Qwen 3.6 35B (MTP) · tool · huggingface.co · free
Qwen 3.6 mixture-of-experts model (35B total, about 3B active parameters) with multi-token prediction for local inference.
Also in: Use a Local Model for Confidential Data with Your Agent (Melvin Vivas on X · notes), Running Qwen 3.6 35B locally on an RTX 3090 for agent tool calling (Melvin Vivas on X · notes), Run Hermes Agent locally with Qwen 3.6 35B MTP in LM Studio (Melvin Vivas on X · notes), Local Personal AI Agent to Explain Your Investment Portfolio (Melvin Vivas on X · notes) and 4 more - Gemini Flash · tool · deepmind.google · free
Google's fast, low-cost Gemini model, suited to auxiliary tasks like web browsing and vision.
Also in: Testing Gemini 3.8 Flash in Cursor with a CRM Smoke Test (Melvin Vivas on X · notes), Gemini 3.8 Flash Available in Cursor CLI (Melvin Vivas on X · notes), Cut Hermes Agent Token Costs with Gemini Flash as an Auxiliary Model (Melvin Vivas on X · notes)
More in AI Agents, Tool Use & MCP
- Hermes Agent Recommendation
- Run a Separate Hermes Agent per Domain with Docker
- Building an Agent Command Centre with Hermes Agents in Docker
- Grok Build Adds Workflows for Repeatable Multi-Agent Pipelines
- Hermes Agent v0.19.0: faster responses and smart approvals
- Updating Hermes Agent to v0.19.0 ("The Quicksilver Release")