Model Routing: Fine-Tune Your Own Router on Your Prompts
Melvin Vivas · X video post · 2026-09-21 · Open on X
Topics: Fine-tuning & Model Customization, LLMOps, Deployment & Monitoring · Level: advanced
Summary
Melvin Vivas argues that the best way to get model routing right is to fine-tune your own routing model on your own prompts. It then sends each request to the model you want, such as Astra or Luna. He links his earlier post where routing with Jev was not worth it for his workflow: the routed 'team' cost almost as much as Astra alone and hit the time limit.
Key points
- Generic routers may not fit your workflow: in his test, Jev-based routing cost almost as much as using Astra alone
- The routed multi-model 'team' hit the time limit before it finished the task
- Suggested fix: fine-tune your own routing model on your own prompt history
- The router decides which target model (e.g. Astra, Luna or any other) gets each prompt
- Measure cost and completion time before adopting routing in production
Resources mentioned
- Jev (TypeSafe AI) · tool · typesafe.ai · free
TypeSafe's new post-training algorithm for calibrated, confidence-scored decisions, pitched as a replacement for RLHF.
Also in: JevDev: Open-Source UI Tool for Experimenting with Jev (Typesafe.ai) (Melvin Vivas on X · notes), jev-dev: A UI Tool to Keep Run History When Experimenting with Jev (Melvin Vivas on X · notes), JevDev: Open-Source UI for Experimenting with Jev by Typesafe.ai (Melvin Vivas on X · notes), Building a Jev Session-History Tool with Codex and Astra (Melvin Vivas on X · notes) and 7 more - GPT-6 Astra · tool · openai.com · paid
The model announced in the quoted launch post, pitched as the developer's most capable model for work, coding, science and cybersecurity, and able to operate a computer.
Also in: Customizing Your Coding Setup with Pi Coding Agent Extensions (Melvin Vivas on X · notes), Pi Agent Council: Ask Multiple LLMs in Parallel and Compare Their Advice (Melvin Vivas on X · notes), Use GPT-6.1 Sol by Default, Save Astra for Emergencies (Melvin Vivas on X · notes), Dots in ChatGPT: always-on AI agents that you hand responsibilities to (Melvin Vivas on X · notes) and 49 more - Luna · tool · openai.com · paid
The other model/agent the classifier routes tasks to; the post doesn't describe it further.
Also in: GPT-6.1 Sol May Beat Luna for Subagents (Melvin Vivas on X · notes), Picking models for orchestrator and subagent roles in Codex (Melvin Vivas on X · notes), GPT-6 Sol Ultra Subagents Use Up Limits Fast (Melvin Vivas on X · notes), GPT-6 Sol vs Opus 5.5 in a Livestream Comparison (Melvin Vivas on X · notes) and 14 more
Try this
- Fine-tune your own routing model using your own prompts
- Start from the linked post on routing with Jev
- Train a custom prompt router that picks between models (e.g. Astra, Luna) from your own prompt history, and compare its cost and time against a single model
More in Fine-tuning & Model Customization
- Train Your Own LLM Model Router by Fine-Tuning ModernBERT as a Classifier
- Fine-tuning Liquid AI LFM2/LFM2.5 MoE models with the new Halo framework
- Hugging Face LLM Course: Chapter 7.3 for Training Your Own Model
- Base vs fine-tuned Gemma 4 E2B as a model router
- QLoRA fine-tune Gemma 4 E2B with Unsloth as a model router
- Free Colab Notebooks: ModernBERT, Gemma QLoRA, GLiNER & Guardrails