Running a Local Ornith Model as the Backend for a Hermes Agent
Melvin Vivas · X post · 2026-08-20 · Open on X
Topics: AI Agents, Tool Use & MCP, LLMOps, Deployment & Monitoring · Level: intermediate
Summary
Shows a split setup: a Hermes agent runs on a Mac and connects to a PC on the network that serves the Ornith-1.5-35B-A3B model with llama.cpp. The creator rates it as a good alternative to Qwen3.8 27B but notes it has no vision, so a separate vision model is still needed.
Key points
- The agent runs on one machine (Mac) and model inference runs on a GPU PC through a llama.cpp server.
- Ornith-1.5-35B-A3B is a mixture-of-experts model (35B total, about 3B active parameters).
- He considers it a good alternative to Qwen3.8 27B.
- It has no vision support, so multimodal tasks need a separate vision model.
Resources mentioned
- Ornith-1.5-35B-A3B · tool · huggingface.co · free
An open-weight mixture-of-experts language model (35B total, about 3B active parameters), run here as a Q4_K_M GGUF quantization.
Also in: Serving Ornith-1.5-35B-A3B at 128k Context on an RTX 3090 with llama.cpp (Melvin Vivas on X · notes), Running Ornith-1.5-35B-A3B (Q4_K_M) at 128k Context on an RTX 3090 with llama.cpp (Melvin Vivas on X · notes) - Hermes Agent · tool · github.com · free
Nous Research's open-source AI agent with CLI, TUI and desktop interfaces. It now supports hands-free activation with a wake word.
Also in: Inspecting Coding-Agent Traces Live with JSONL Viewer (Codex, Claude Code) (Melvin Vivas on X · notes), Hermes Agent: each bot is its own profile (Melvin Vivas on X · notes), An X research bot built with Hermes (Melvin Vivas on X · notes), Run Your Hermes Agent on Free LFM2.5-2.6B via OpenRouter (Melvin Vivas on X · notes) and 55 more - llama.cpp · repo · github.com · free
An open-source C/C++ engine for running GGUF models locally. Its llama-server command provides an OpenAI-compatible HTTP server.
Also in: Running LLMs Locally Without an Expensive Rig (Melvin Vivas on X · notes), llama.cpp / Llama-macOS v0.5.0 release (Melvin Vivas on X · notes), Run llama.cpp GGUF Checkpoints in Hugging Face Transformers (Melvin Vivas on X · notes), llama.cpp v0.4.1 release announcement (Melvin Vivas on X · notes) and 28 more - Qwen3.8-27B · tool · huggingface.co · free
A 27B-parameter open-weight model from Alibaba's Qwen family that you can run locally or call through hosted APIs.
Also in: Deploying Qwen3.8 27B on Hugging Face Inference Endpoints (Melvin Vivas on X · notes), Using Hugging Face credits: Jobs, Inference Endpoints and Open Models (Melvin Vivas on X · notes), Running Qwen3.8-27B Locally on an M5 Max MacBook with Inco Splash (Melvin Vivas on X · notes), Ternary Bonsai 2 27B: 9x smaller model keeping 98.2% of benchmark scores (Melvin Vivas on X · notes) and 24 more
Try this
- Serve a local LLM on a GPU PC with llama.cpp and point an agent running on another machine at it.
More in AI Agents, Tool Use & MCP
- Agent Harness Efficiency: Scaffolding Beats MCP vs. CLI
- Run Your Hermes Agent on Free LFM2.5-2.6B via OpenRouter
- 5 Agentic Design Patterns: ReAct, Planning, Reflection, Routing, Multi-Agent
- Mixture-of-Agents Setting Was Quietly Using Up Codex/ChatGPT Quota
- Mind Viruses: How Bad Ideas Spread Through Multi-Agent LLM Systems
- Running DeepSeek V4 Pro on Baseten with the Pi Agent Harness