QLoRA fine-tune Gemma 4 E2B with Unsloth as a model router
Melvin Vivas · X post · 2026-09-20 · Open on X
Topics: Fine-tuning & Model Customization, LLMOps, Deployment & Monitoring · Level: intermediate
Summary
The creator QLoRA fine-tuned Gemma 4 E2B Instruct with Unsloth to classify software tasks and send each one to Luna (cheaper) or Astra (stronger). The notebook runs in Google Colab and is shared in his donvito/notebooks repo. He hasn't plugged it into Codex yet and asks others to try.
Key points
- Model routing: a small classifier model decides which larger model should handle each task.
- Base model: Gemma 4 E2B Instruct, small enough to fine-tune cheaply.
- Method: QLoRA (LoRA adapters on a quantized model) with Unsloth.
- Labels: software tasks are routed to luna or astra.
- Runs in Google Colab; the notebook is in github.com/donvito/notebooks.
- Not yet integrated with Codex, so wiring it up is left open.
Resources mentioned
- donvito/notebooks · repo · github.com · free
The creator's notebooks for fine-tuning and running local models, which you can run in Google Colab, including a GLiNER2.5-Decide intent classification example.
Also in: Run Local Models on a Free GPU with Google Colab (T4) (Melvin Vivas on X · notes), Free Colab Notebooks for Local Model Fine-Tuning and Inference (Melvin Vivas on X · notes), Intent classification for support using GLiNER2.5-Decide notebook (Melvin Vivas on X · notes), Getting started with GLiNER2.5-Decide in a Colab notebook (Melvin Vivas on X · notes) and 3 more - Unsloth · tool · unsloth.ai · free
Open-source library for fast, memory-efficient fine-tuning of open LLMs (LoRA/QLoRA) on a single GPU or in Colab.
Also in: Run Laya Decision models locally with Unsloth on 4GB RAM (Melvin Vivas on X · notes), Unsloth passes 500M model downloads on Hugging Face (Melvin Vivas on X · notes), Run Qwen-Image-2.1 locally on 12GB VRAM with Unsloth GGUFs (Melvin Vivas on X · notes), Base vs fine-tuned Gemma 4 E2B as a model router (Melvin Vivas on X · notes) and 29 more - Gemma 4 E2B Instruct · tool · huggingface.co · free
Small open-weight instruction-tuned Gemma model, used here as the base for a QLoRA fine-tune.
Also in: Base vs fine-tuned Gemma 4 E2B as a model router (Melvin Vivas on X · notes), Free Colab Notebooks: ModernBERT, Gemma QLoRA, GLiNER & Guardrails (Melvin Vivas on X · notes) - Google Colab · tool · colab.research.google.com · free · recommended by both Bashiri Smith & Melvin Vivas
Free GPU notebooks.
Also in: Run Notebooks on a Free GPU with Google Colab (T4, 15GB VRAM) (Melvin Vivas on X · notes), Run Local Models on a Free GPU with Google Colab (T4) (Melvin Vivas on X · notes), Free Colab Notebooks for Local Model Fine-Tuning and Inference (Melvin Vivas on X · notes), Intent classification for support using GLiNER2.5-Decide notebook (Melvin Vivas on X · notes) and 12 more - OpenAI Codex · tool · openai.com · paid
OpenAI's coding agent. In the diagram it writes code, fixes review findings and drives the build loop. The creator also used it to make this video.
Also in: An agent bot that installs and drives Codex on its own (Melvin Vivas on X · notes), Asking a Coder bot to install Codex (Melvin Vivas on X · notes), Sign in with ChatGPT: Setting Usage Limits for Each App (Melvin Vivas on X · notes), Codex Cloud Environments Must Be Saved & Published Before Use (Melvin Vivas on X · notes) and 240 more
Try this
- Open the Colab notebook in donvito/notebooks and run the QLoRA fine-tune.
- Try wiring the router into Codex and share the results with the creator.
- Build a cost-saving model router: a fine-tuned small model that sends each coding task to a cheap or a strong model, connected to Codex.
More in Fine-tuning & Model Customization
- Hugging Face LLM Course: Chapter 7.3 for Training Your Own Model
- Model Routing: Fine-Tune Your Own Router on Your Prompts
- Base vs fine-tuned Gemma 4 E2B as a model router
- Free Colab Notebooks: ModernBERT, Gemma QLoRA, GLiNER & Guardrails
- Fine-tuning Gemma4-E2B on your own tweet style with Unsloth
- Not every problem needs an LLM: small fine-tuned models (Jev by TypeSafe AI)