Fine-tuning ModernBERT-base to route tasks between two models
Melvin Vivas · X post · 2026-09-18 · Open on X
Topics: Fine-tuning & Model Customization, AI Agents, Tool Use & MCP, AI System Design & Architecture · Level: intermediate
Summary
The creator is experimenting with fine-tuning ModernBERT-base as a task classifier. The classifier decides whether each task should be delegated to Astra or Luna. This shows a lightweight learned router placed in front of more expensive models or agents.
Key points
- Base model: ModernBERT-base, an encoder well suited to classification fine-tuning.
- Training goal: label each incoming task so it can be delegated to Astra or Luna.
- A small fine-tuned classifier is a cheap, fast way to route work between models or agents.
- The creator asks the audience whether they would use such a router.
Resources mentioned
- ModernBERT-base · tool · huggingface.co · free
An open-source modernized BERT encoder model. The quoted post uses it as the speed baseline for LFM2.5-Encoder.
Also in: Train Your Own LLM Model Router by Fine-Tuning ModernBERT as a Classifier (Melvin Vivas on X · notes), Free Colab Notebooks: ModernBERT, Gemma QLoRA, GLiNER & Guardrails (Melvin Vivas on X · notes), Fine-Tune ModernBERT-base as a Task-Routing Classifier (Melvin Vivas on X · notes), Train ModernBERT as a Prompt Router Between Two Models (Colab Notebook) (Melvin Vivas on X · notes) and 2 more - GPT-6 Astra · tool · openai.com · paid
The model announced in the quoted launch post, pitched as the developer's most capable model for work, coding, science and cybersecurity, and able to operate a computer.
Also in: Customizing Your Coding Setup with Pi Coding Agent Extensions (Melvin Vivas on X · notes), Pi Agent Council: Ask Multiple LLMs in Parallel and Compare Their Advice (Melvin Vivas on X · notes), Use GPT-6.1 Sol by Default, Save Astra for Emergencies (Melvin Vivas on X · notes), Dots in ChatGPT: always-on AI agents that you hand responsibilities to (Melvin Vivas on X · notes) and 49 more - Luna · tool · openai.com · paid
The other model/agent the classifier routes tasks to; the post doesn't describe it further.
Also in: GPT-6.1 Sol May Beat Luna for Subagents (Melvin Vivas on X · notes), Picking models for orchestrator and subagent roles in Codex (Melvin Vivas on X · notes), GPT-6 Sol Ultra Subagents Use Up Limits Fast (Melvin Vivas on X · notes), GPT-6 Sol vs Opus 5.5 in a Livestream Comparison (Melvin Vivas on X · notes) and 14 more
Try this
- Build a task router: fine-tune ModernBERT-base on labeled tasks to choose which of two models or agents handles each one.
More in Fine-tuning & Model Customization
- Kev-0.5B: A Tiny Open-Source Decision Model to Train on a MacBook
- Train ModernBERT as a Prompt Router Between Two Models (Colab Notebook)
- Fine-tuning ModernBERT-base as a task router (quote post)
- Train models locally with the Unsloth Docker image
- Post-training Qwen3.5-2B on your own X posts with Unsloth Studio
- Use Codex to Prepare Fine-Tuning Datasets