Multi-Teacher On-Policy Distillation (MOPD) in 2026
Melvin Vivas · X post · 2026-09-03 · Open on X
Topics: Fine-tuning & Model Customization, Industry Trends & Job Market · Level: advanced
Summary
The creator shares a compilation of model distillation papers, quoting a post about Multi-Teacher On-Policy Distillation (MOPD). MOPD is a new post-training approach where one student model learns capabilities from several specialized RL teacher models. According to the quoted post, it's used in frontier models such as MiMo-V2-Flash, Kimi K3 and DeepSeek-V4. The paper list itself isn't linked in the material.
Key points
- MOPD = Multi-Teacher On-Policy Distillation, an emerging post-training paradigm in 2026.
- A single student model absorbs capabilities from multiple specialized RL-trained teacher models.
- On-policy: the student learns from its own generated outputs, with teachers giving guidance on them.
- According to the quoted post, it's used in MiMo-V2-Flash, Kimi K3 and DeepSeek-V4.
Resources mentioned
- MiMo-V2-Flash · tool · github.com · free
Xiaomi's MiMo model, cited as using multi-teacher on-policy distillation. - Kimi K3 · tool · huggingface.co · free
Moonshot AI's large multimodal LLM with a 1M-token context window and Kimi Delta Attention.
Also in: Qwen3.8-Max on the Frontend Code Arena cost-performance frontier (Melvin Vivas on X · notes), 1-bit Kimi K3 GGUF Running Locally vs Claude Opus 5 and GPT 5.6 (Melvin Vivas on X · notes), Code Arena Fullstack Benchmark: Kimi K3 Ranks #1 (Melvin Vivas on X · notes), Kimi K3 ranks #1 on Design Arena for frontend building (Melvin Vivas on X · notes) and 1 more - DeepSeek V4 · tool · huggingface.co · free
A DeepSeek LLM, accessed through OpenRouter as the final fallback.
Also in: Models That Work With the Hermes Agent for Personal Productivity (Melvin Vivas on X · notes), Cost-Saving Model Fallback Chain: Grok → Codex → OpenRouter DeepSeek (Melvin Vivas on X · notes)
Try this
- Read up on on-policy distillation and multi-teacher distillation, for example the technical reports of the models named.
More in Fine-tuning & Model Customization
- Hugging Face Training Agents series: post-training coding agents
- jsonl-viewer: a tool to view and edit JSONL training datasets
- Fine-Tune LFM2.5-350M with GRPO in TRL for Structured Outputs
- Ready-made Docker image for running coding agents on ML/fine-tuning jobs
- NuExtract3: Open 4B VLM for OCR and Structured JSON Extraction
- jsonl-viewer: Viewing and Editing Fine-Tuning Datasets in JSONL