Topic 10 of 16 in the learning path
Fine-tuning & Model Customization
LoRA, distillation, preference tuning and training your own models.
Reels and posts (63)
- Using an ML agent to train an open-source TTS model on your voice · Melvin Vivas, X: The creator's ML Grok Bot explained how to fine-tune an open-source text-to-speech model on his own voice.
- Run Local Models on a Free GPU with Google Colab (T4) · Melvin Vivas, X · 0:54: Melvin Vivas shows how to use Google Colab to run open, local-style models on a free GPU, without buying hardware.
- Free Colab Notebooks for Local Model Fine-Tuning and Inference · Melvin Vivas, X: The creator has added more AIBackends notebooks to his free GitHub repo.
- Train Your Own LLM Model Router by Fine-Tuning ModernBERT as a Classifier · Melvin Vivas, X · 1:00: Melvin Vivas shows how to build a model router.
- Fine-tuning Liquid AI LFM2/LFM2.5 MoE models with the new Halo framework · Melvin Vivas, X: Melvin Vivas shares Halo, a new fine-tuning framework, along with Liquid AI's two official recipes for fine-tuning their Mixture-of-Experts models.
- Hugging Face LLM Course: Chapter 7.3 for Training Your Own Model · Melvin Vivas, X: The creator quotes a post saying that Hugging Face Learn has 'premium Jev material' and links to chapter 7, section 3 of the free Hugging Face LLM Course.
- Model Routing: Fine-Tune Your Own Router on Your Prompts · Melvin Vivas, X: Melvin Vivas argues that the best way to get model routing right is to fine-tune your own routing model on your own prompts.
- Base vs fine-tuned Gemma 4 E2B as a model router · Melvin Vivas, X: Compares the base model with the fine-tuned one for the creator's Gemma 4 E2B model router.
- QLoRA fine-tune Gemma 4 E2B with Unsloth as a model router · Melvin Vivas, X: The creator QLoRA fine-tuned Gemma 4 E2B Instruct with Unsloth to classify software tasks and send each one to Luna (cheaper) or Astra (stronger).
- Free Colab Notebooks: ModernBERT, Gemma QLoRA, GLiNER & Guardrails · Melvin Vivas, X: Melvin Vivas shares a set of free Google Colab notebooks on his website, mostly about fine-tuning.
- Fine-tuning Gemma4-E2B on your own tweet style with Unsloth · Melvin Vivas, X: The creator is building a notebook that fine-tunes the small Gemma4-E2B model to write in his own tweet style, using Unsloth.
- Not every problem needs an LLM: small fine-tuned models (Jev by TypeSafe AI) · Melvin Vivas, X: The creator praises TypeSafe AI's Jev model.
- Hugging Face AutoTrain: a no-code fine-tuning tool · Melvin Vivas, X: The creator asks Hugging Face to revive AutoTrain, its no-code/low-code tool for fine-tuning models.
- LoRA Fine-Tune Qwen3.5-2B on Your Tweets with Unsloth Studio · Melvin Vivas, X: The creator fine-tuned Qwen3.5-2B with LoRA in Unsloth Studio, using his own popular tweets as the dataset.
- Fine-Tune ModernBERT-base as a Task-Routing Classifier · Melvin Vivas, X: The creator is experimenting with fine-tuning ModernBERT-base to classify incoming tasks and send each one to one of two agents (Astra or Luna).
- Kev-0.5B: A Tiny Open-Source Decision Model to Train on a MacBook · Melvin Vivas, X: Shares Kev-0.5B, a tiny open-source decision model similar to Jev, with a TypeSafe-compatible API.
- Train ModernBERT as a Prompt Router Between Two Models (Colab Notebook) · Melvin Vivas, X: The creator shares a Colab notebook that fine-tunes a ModernBERT classifier on prompts.
- Fine-tuning ModernBERT-base as a task router (quote post) · Melvin Vivas, X: The creator adds 'awesome' to his own post about fine-tuning ModernBERT-base.
- Fine-tuning ModernBERT-base to route tasks between two models · Melvin Vivas, X: The creator is experimenting with fine-tuning ModernBERT-base as a task classifier.
- Train models locally with the Unsloth Docker image · Melvin Vivas, X: The creator recommends Unsloth's Docker image for training models locally because it avoids the trouble of setting up Python and CUDA.
- Post-training Qwen3.5-2B on your own X posts with Unsloth Studio · Melvin Vivas, X: The creator post-trained the small Qwen3.5-2B model with Unsloth Studio, using his own popular X posts as the training and validation data.
- Use Codex to Prepare Fine-Tuning Datasets · Melvin Vivas, X: Tip: when you train a model, ask Codex to help prepare the dataset.
- Melvin Vivas's Hugging Face Profile · Melvin Vivas, X: The creator shares his Hugging Face profile, where he publishes his fine-tuned models.
- GPT-6 Astra Medium in Codex as a Low-Cost Fine-Tuning Assistant · Melvin Vivas, X: The creator recommends GPT-6 Astra on the medium setting in Codex as an assistant for fine-tuning and training small models, because it doesn't use much of your usage allowance.
- Push Fine-Tuned Models to Hugging Face from Unsloth Studio · Melvin Vivas, X: After post-training, Unsloth Studio can upload your model to Hugging Face from its UI.
- Fine-Tuned Qwen3.5-2B LoRA Model That Writes X Posts in Your Style · Melvin Vivas, X: The creator fine-tuned Qwen3.5-2B with LoRA in Unsloth Studio.
- First LoRA Run on Qwen3.5-2B with Codex as Training Companion · Melvin Vivas, X: The creator finished his first fine-tuning run of Qwen3.5-2B to copy his tweet style.
- Using Codex to Generate Synthetic Training Pairs for Fine-tuning Qwen2.5 · Melvin Vivas, X: Melvin Vivas shows that a coding agent (OpenAI Codex) can handle dataset prep for machine learning.
- Unsloth Studio: Fine-tuning Without Colab Notebooks · Melvin Vivas, X: The creator says Unsloth Studio makes fine-tuning easier than the old workflow of running Google Colab notebooks.
- Training a Small Model on Your Own Popular Tweets · Melvin Vivas, X: The creator is preparing data to train a small model on his most popular tweets, which suggests a personal style-tuning project.
- Model Distillation Explained: A Teacher Model Trains a Smaller Student · Melvin Vivas, X: The post explains model distillation: a stronger 'teacher' model generates high-quality examples, and those outputs are used to train a smaller model for one specific task.
- Idea: Fine-Tune a Small Model on Your Best Tweets · Melvin Vivas, X: The creator plans to fine-tune a small model to write in his own style, using his most successful posts as the training dataset.
- Hugging Face Training Agents series: post-training coding agents · Melvin Vivas, X: The creator recommends Hugging Face's finished Training Agents series for learning how to post-train models for coding agents.
- jsonl-viewer: a tool to view and edit JSONL training datasets · Melvin Vivas, X: The creator shares his open-source jsonl-viewer, a tool for viewing and editing .jsonl files.
- Fine-Tune LFM2.5-350M with GRPO in TRL for Structured Outputs · Melvin Vivas, X: The creator points to a Hugging Face blog tutorial on fine-tuning a tiny model for better structured outputs.
- Multi-Teacher On-Policy Distillation (MOPD) in 2026 · Melvin Vivas, X: The creator shares a compilation of model distillation papers, quoting a post about Multi-Teacher On-Policy Distillation (MOPD).
- Ready-made Docker image for running coding agents on ML/fine-tuning jobs · Melvin Vivas, X: For ML work like fine-tuning with coding agents, the creator offers a ready-made Docker image with PyTorch/CUDA and the agents already installed.
- NuExtract3: Open 4B VLM for OCR and Structured JSON Extraction · Melvin Vivas, X: NuExtract3 is NuMind's open-source (Apache 2.0) 4B vision-language model with reasoning.
- jsonl-viewer: Viewing and Editing Fine-Tuning Datasets in JSONL · Melvin Vivas, X: The creator built jsonl-viewer, an open-source tool for viewing and editing .jsonl (JSON Lines) files.
- Fine-Tune a Model on Your Own Coding Agent Traces · Melvin Vivas, X: The creator recommends a Hugging Face live tutorial by Ben Burtenshaw on fine-tuning a coding agent with its own agent traces, so it keeps learning over time.
- Model Routing Fine-Tuned on Your Agent Harness Traces · Melvin Vivas, X: An example of how model routing can work: fine-tune Liquid AI's LFM2.5 Encoder on traces from your agent harness, which record each prompt and the model used for it.
- Fine-tune and Deploy Qwen3.8 27B on Together AI · Melvin Vivas, X: Together AI now supports fine-tuning Qwen3.8 27B and serving it on dedicated inference.
- Quantization-Aware Distillation (QAD) for Better 4-bit GGUF Models · Melvin Vivas, X: Liquid AI released new 4-bit checkpoints trained with Quantization-Aware Distillation (QAD).
- Fine-Tune Muse Glimmer 30B with LoRA or Full-Parameter on Fireworks · Melvin Vivas, X: Muse Glimmer 30B is now on Fireworks' Dedicated Training API for both LoRA and full-parameter fine-tuning.
- jsonl-viewer: view and edit JSONL fine-tuning datasets · Melvin Vivas, X: The creator added a file explorer to jsonl-viewer, his open-source tool for viewing and editing .jsonl files.
- Fine-Tune Muse Glimmer 30B for Free with Unsloth (incl. GRPO) · Melvin Vivas, X: Shares Unsloth's announcement that Meta's Muse Glimmer 30B can now be fine-tuned for free with Unsloth notebooks.
- Smoke-Testing LFM Fine-tuning with Codex and Liquid AI's LEAP · Melvin Vivas, X: The creator is trying out fine-tuning Liquid AI's LFM models with the LEAP framework, using Codex as a coding assistant.
- Free JSONL Viewer for Inspecting Fine-tuning Datasets · Melvin Vivas, X: The creator shares a free, open-source tool he built for viewing and editing .jsonl files.
- Liquid AI Cookbook: Fine-Tuning LFMs with CPT, SFT, DPO and GRPO · Melvin Vivas, X: Melvin Vivas recommends the Liquid AI cookbook as a strong, underrated resource for fine-tuning Liquid's LFM models.
- Loop engineering fits fine-tuning better than coding · Melvin Vivas, X: The creator argues that loop engineering (automated iterate-and-check loops) works better for fine-tuning models than for coding.
- Small Local Models + Fine-Tuning Instead of More Compute (Liquid AI, LEAP) · Melvin Vivas, X: Melvin argues that instead of chasing more compute, you should make small local models work for your specific use case.
- Evolution of Image LoRAs: Stable Diffusion → Flux → Krea 2 · Melvin Vivas, X: A short note on which base models image LoRAs have been trained on over time: first Stable Diffusion, then Flux, and now Krea 2.
- Open-Source JSONL Viewer for Fine-Tuning Datasets · Melvin Vivas, X: The creator built an open-source viewer and editor for .jsonl files using GLM 5.2 in Cursor.
- Krea 2 Open Weights: Raw (Undistilled) vs Turbo (Distilled) Image Models · Melvin Vivas, X · 0:24: This is a short post sharing Krea's announcement that it has released the open weights of its Krea 2 image generation model.
- Paper: Context-Aware RL for Agentic and Multimodal LLMs · Melvin Vivas, X: The creator shares a link to an arXiv paper (2606.17053) titled 'Context-Aware RL for Agentic and Multimodal LLMs'.
- 1-bit Bonsai Image 4B: A Diffusion Model Under 1GB · Melvin Vivas, X: A quoted announcement says the 1-bit Bonsai Image 4B diffusion transformer is only 0.93GB, 8.3x smaller than the full-precision version.
- Unsloth joins the PyTorch Ecosystem · Melvin Vivas, X: The creator congratulates Unsloth on joining the PyTorch Ecosystem.
- ML Intern: Post-Train a Model by Chatting (Hugging Face) · Melvin Vivas, X: Highlights Hugging Face's ML Intern, an agent that lets you train or post-train models through a chat conversation.
- Fine-Tuning Models with Unsloth Studio · Melvin Vivas, X: Unsloth Studio makes fine-tuning models easier.
- Fine-tuning Gemma 3 270M with Unsloth for Work Tasks · Melvin Vivas, X: The creator says he is fine-tuning Google's tiny Gemma 3 270M model for a work use case with Unsloth, and is thinking about fine-tuning Gemma 4 next.
- Installing Unsloth Studio on Windows via WSL · Melvin Vivas, X: The creator reports that he installed Unsloth Studio on Windows using WSL (Windows Subsystem for Linux).
- Looking for a Low-Cost Replacement for Hugging Face AutoTrain · Melvin Vivas, X: The creator notes that Hugging Face AutoTrain has been discontinued and asks for a cheaper alternative, since Google AutoML and Amazon SageMaker cost too much for him.
- Unsloth Studio: Open-Source Web UI to Train and Run LLMs Locally · Melvin Vivas, X · 0:45: Melvin Vivas shares the launch of Unsloth Studio, an open-source web UI from the Unsloth team for training (fine-tuning) and running LLMs on your own machine.
Watch, free (3)
- Training Agents series (Hugging Face) · video · youtube.com · free
Hugging Face live tutorial on fine-tuning a coding agent with its own traces for continual learning.
Mentioned in: Hugging Face Training Agents series: post-training coding agents (Melvin Vivas on X · notes), Fine-Tune a Model on Your Own Coding Agent Traces (Melvin Vivas on X · notes) - Hugging Face LLM Course – Chapter 7.3 · course · huggingface.co · free
A section of Hugging Face's free LLM Course with hands-on training and fine-tuning of models.
Mentioned in: Hugging Face LLM Course: Chapter 7.3 for Training Your Own Model (Melvin Vivas on X · notes) - Hugging Face YouTube channel · channel · youtube.com · free
Hugging Face's official YouTube channel, which hosts the Training Agents sessions.
Mentioned in: Hugging Face Training Agents series: post-training coding agents (Melvin Vivas on X · notes)
Read and use (43)
- Unsloth · tool · unsloth.ai · free
Open-source library for fast, memory-efficient fine-tuning of open LLMs (LoRA/QLoRA) on a single GPU or in Colab.
Mentioned in: Run Laya Decision models locally with Unsloth on 4GB RAM (Melvin Vivas on X · notes), Unsloth passes 500M model downloads on Hugging Face (Melvin Vivas on X · notes), Run Qwen-Image-2.1 locally on 12GB VRAM with Unsloth GGUFs (Melvin Vivas on X · notes), Base vs fine-tuned Gemma 4 E2B as a model router (Melvin Vivas on X · notes) and 30 more - donvito/notebooks · repo · github.com · free
The creator's notebooks for fine-tuning and running local models, which you can run in Google Colab, including a GLiNER2.5-Decide intent classification example.
Mentioned in: Run Local Models on a Free GPU with Google Colab (T4) (Melvin Vivas on X · notes), Free Colab Notebooks for Local Model Fine-Tuning and Inference (Melvin Vivas on X · notes), Intent classification for support using GLiNER2.5-Decide notebook (Melvin Vivas on X · notes), Getting started with GLiNER2.5-Decide in a Colab notebook (Melvin Vivas on X · notes) and 4 more - ModernBERT-base · tool · huggingface.co · free
An open-source modernized BERT encoder model. The quoted post uses it as the speed baseline for LFM2.5-Encoder.
Mentioned in: Train Your Own LLM Model Router by Fine-Tuning ModernBERT as a Classifier (Melvin Vivas on X · notes), Free Colab Notebooks: ModernBERT, Gemma QLoRA, GLiNER & Guardrails (Melvin Vivas on X · notes), Fine-Tune ModernBERT-base as a Task-Routing Classifier (Melvin Vivas on X · notes), Train ModernBERT as a Prompt Router Between Two Models (Colab Notebook) (Melvin Vivas on X · notes) and 3 more - Gemma 4 E2B Instruct · tool · huggingface.co · free
Small open-weight instruction-tuned Gemma model, used here as the base for a QLoRA fine-tune.
Mentioned in: Base vs fine-tuned Gemma 4 E2B as a model router (Melvin Vivas on X · notes), QLoRA fine-tune Gemma 4 E2B with Unsloth as a model router (Melvin Vivas on X · notes), Free Colab Notebooks: ModernBERT, Gemma QLoRA, GLiNER & Guardrails (Melvin Vivas on X · notes) - LEAP (Liquid AI) · tool · leap.liquid.ai · check price
Liquid AI's framework for fine-tuning and customizing its small models.
Mentioned in: Smoke-Testing LFM Fine-tuning with Codex and Liquid AI's LEAP (Melvin Vivas on X · notes), Small Local Models + Fine-Tuning Instead of More Compute (Liquid AI, LEAP) (Melvin Vivas on X · notes), LFM2.5-2.6B Matches DeepSeek-V4-Flash on Tool Calling; LEAP Fine-Tuning (Melvin Vivas on X · notes) - Ideogram 4.0 · tool · ideogram.ai · free
Ideogram's text-to-image model, released with open weights that you can download, fine-tune and run yourself.
Mentioned in: 2.3x Faster Ideogram 4 in ComfyUI with INT8 and SageAttention (Melvin Vivas on X · notes), Ideogram 4.0 Released as an Open-Weights Image Model (Melvin Vivas on X · notes) - melvindave/qwen3.5-2b-melvin-posts-v1-GGUF · tool · huggingface.co · free
Hugging Face model repo with the GGUF version of the Qwen3.5-2B model post-trained to write X posts in Melvin's style.
Mentioned in: Post-training Qwen3.5-2B on your own X posts with Unsloth Studio (Melvin Vivas on X · notes), Fine-Tuned Qwen3.5-2B LoRA Model That Writes X Posts in Your Style (Melvin Vivas on X · notes) - TRL (Hugging Face) · tool · github.com · free
Hugging Face's open-source library for post-training LLMs with SFT, DPO, GRPO and more.
Mentioned in: Fine-Tune LFM2.5-350M with GRPO in TRL for Structured Outputs (Melvin Vivas on X · notes), Liquid AI Cookbook: Fine-Tuning LFMs with CPT, SFT, DPO and GRPO (Melvin Vivas on X · notes) - Unsloth Documentation: Muse Glimmer guide · docs · unsloth.ai · free
Unsloth's guide and free notebooks for fine-tuning and GRPO-training Muse Glimmer 30B.
Mentioned in: Fine-Tune Muse Glimmer 30B for Free with Unsloth (incl. GRPO) (Melvin Vivas on X · notes), Muse Glimmer 30B: Unsloth GGUF Release and Run Guide (Melvin Vivas on X · notes) - Amazon SageMaker · tool · aws.amazon.com · paid
AWS's managed platform for training and deploying ML models.
Mentioned in: Looking for a Low-Cost Replacement for Hugging Face AutoTrain (Melvin Vivas on X · notes) - Ben Burtenshaw · person · x.com · free
Hugging Face engineer who presents the agent fine-tuning tutorial.
Mentioned in: Fine-Tune a Model on Your Own Coding Agent Traces (Melvin Vivas on X · notes) - Bonsai Image 4B (1-bit) · tool · prismml.com · free
1-bit quantized 4B diffusion transformer image model that is only 0.93GB.
Mentioned in: 1-bit Bonsai Image 4B: A Diffusion Model Under 1GB (Melvin Vivas on X · notes) - Context-Aware RL for Agentic and Multimodal LLMs (arXiv 2606.17053) · paper · donvitocodes.com · free
An arXiv research paper on context-aware reinforcement learning for training agentic and multimodal LLMs.
Mentioned in: Paper: Context-Aware RL for Agentic and Multimodal LLMs (Melvin Vivas on X · notes) - Fastino · tool · fastino.ai · free
Company that made the finance and healthcare Nemotron models with NVIDIA.
Mentioned in: Open-weight Nemotron models for finance and healthcare (Melvin Vivas on X · notes) - Fine-tuning LFM2.5-350M for structured outputs with GRPO (Hugging Face blog) · article · huggingface.co · free
Tutorial showing how to fine-tune LFM2.5-350M in 100 GRPO steps with TRL for better structured outputs.
Mentioned in: Fine-Tune LFM2.5-350M with GRPO in TRL for Structured Outputs (Melvin Vivas on X · notes) - Fireworks AI Dedicated Training API · tool · docs.fireworks.ai · paid
Fireworks API for LoRA and full-parameter fine-tuning of open models.
Mentioned in: Fine-Tune Muse Glimmer 30B with LoRA or Full-Parameter on Fireworks (Melvin Vivas on X · notes) - FlashAttention-2 · tool · github.com · free
Optimized attention kernel, used here as the baseline for Unsloth's speed and memory claims.
Mentioned in: Fine-Tune Muse Glimmer 30B for Free with Unsloth (incl. GRPO) (Melvin Vivas on X · notes) - Gemma 3 1B · tool · huggingface.co · free
A small open Gemma 3 model that is practical to fine-tune.
Mentioned in: Google Colab CLI and Agent Skills: Cloud GPUs From Your Terminal (Melvin Vivas on X · notes) - Gemma 3 270M · tool · huggingface.co · free
Google's 270M-parameter open model, designed to be fine-tuned for specific tasks.
Mentioned in: Fine-tuning Gemma 3 270M with Unsloth for Work Tasks (Melvin Vivas on X · notes) - Google AutoML · tool · docs.cloud.google.com · paid
Google Cloud's managed automated model training service.
Mentioned in: Looking for a Low-Cost Replacement for Hugging Face AutoTrain (Melvin Vivas on X · notes) - Halo · tool · github.com · check price
A new fine-tuning framework that has recipes for MoE models like Liquid AI's LFM2 family.
Mentioned in: Fine-tuning Liquid AI LFM2/LFM2.5 MoE models with the new Halo framework (Melvin Vivas on X · notes) - Halo LFM2 MoE cookbook (LFM2-24B-A2B) · docs · github.com · free
Halo cookbook explaining how to fine-tune the LFM2-24B-A2B MoE model.
Mentioned in: Fine-tuning Liquid AI LFM2/LFM2.5 MoE models with the new Halo framework (Melvin Vivas on X · notes) - Hugging Face AutoTrain (autotrain-advanced) · tool · github.com · free
Hugging Face's open-source no-code/low-code tool for training and fine-tuning models.
Mentioned in: Hugging Face AutoTrain: a no-code fine-tuning tool (Melvin Vivas on X · notes) - Hugging Face PEFT · docs · huggingface.co · free
LoRA and QLoRA fine-tuning.
In a shared PDF: The AI Pivot Field Guide - shared in this reel on Facebook, this reel on Facebook - jaredpalmer/kev (GitHub) · repo · github.com · free
Family of Jev-like decision models built on Qwen that you can train and run yourself.
Mentioned in: Kev-0.5B: A Tiny Open-Source Decision Model to Train on a MacBook (Melvin Vivas on X · notes) - Krea 2 · tool · krea.ai · check price
Krea's image-generation model, named as the newest base for image LoRAs.
Mentioned in: Evolution of Image LoRAs: Stable Diffusion → Flux → Krea 2 (Melvin Vivas on X · notes) - Krea 2 Raw · tool · huggingface.co · free
An undistilled open-weights image generation checkpoint from mid-training, meant to be fine-tuned.
Mentioned in: Krea 2 Open Weights: Raw (Undistilled) vs Turbo (Distilled) Image Models (Melvin Vivas on X · notes) - Liquid AI 4-bit QAD checkpoints · tool · liquid.ai · free
4-bit GGUF model checkpoints from Liquid AI, trained with Quantization-Aware Distillation to recover accuracy lost to quantization.
Mentioned in: Quantization-Aware Distillation (QAD) for Better 4-bit GGUF Models (Melvin Vivas on X · notes) - Liquid AI Cookbook · repo · github.com · free
Liquid AI's collection of fine-tuning examples for LFM models across modalities and training methods.
Mentioned in: Liquid AI Cookbook: Fine-Tuning LFMs with CPT, SFT, DPO and GRPO (Melvin Vivas on X · notes) - Liquid Foundation Models (LFM) · tool · liquid.ai · free
Liquid AI's family of efficient foundation models that can be fine-tuned.
Mentioned in: Smoke-Testing LFM Fine-tuning with Codex and Liquid AI's LEAP (Melvin Vivas on X · notes) - melvindave (Melvin Vivas) on Hugging Face · website · huggingface.co · free
The creator's Hugging Face profile with his published models.
Mentioned in: Melvin Vivas's Hugging Face Profile (Melvin Vivas on X · notes) - MiMo-V2-Flash · tool · github.com · free
Xiaomi's MiMo model, cited as using multi-teacher on-policy distillation.
Mentioned in: Multi-Teacher On-Policy Distillation (MOPD) in 2026 (Melvin Vivas on X · notes) - ModernBERT_train_classify.ipynb (donvito/notebooks) · repo · github.com · free
Colab notebook that fine-tunes ModernBERT to classify prompts, used here to build a model router.
Mentioned in: Train Your Own LLM Model Router by Fine-Tuning ModernBERT as a Classifier (Melvin Vivas on X · notes) - NuExtract3 · tool · huggingface.co · free
Open-source 4B reasoning VLM for OCR to Markdown and schema-based JSON extraction.
Mentioned in: NuExtract3: Open 4B VLM for OCR and Structured JSON Extraction (Melvin Vivas on X · notes) - NVFP4 · other · developer.nvidia.com · free
NVIDIA's 4-bit floating-point format for low-precision training and inference.
Mentioned in: Jensen Huang on NVIDIA's Long-Term Commitment to Open Nemotron Models (Melvin Vivas on X · notes) - qwen3.5-2b-melvin-posts-v1-GGUF (Hugging Face model) · repo · huggingface.co · free · open in a browser to verify
The creator's LoRA fine-tune of Qwen3.5-2B in GGUF format, which writes X posts in his style.
Mentioned in: LoRA Fine-Tune Qwen3.5-2B on Your Tweets with Unsloth Studio (Melvin Vivas on X · notes) - Self-Flow (research paper) · paper · arxiv.org · free
Open research paper on a self-supervised approach to training multimodal generative models that combines representation learning and generation, with no external encoder.
Mentioned in: Self-Flow: Training Multimodal Generative Models Without an External Encoder (Melvin Vivas on X · notes) - SFT MoE with Halo notebook (LFM2.5-8B-A1B) - Liquid4All cookbook · repo · github.com · free
Notebook showing supervised fine-tuning of the LFM2.5-8B-A1B MoE model with Halo.
Mentioned in: Fine-tuning Liquid AI LFM2/LFM2.5 MoE models with the new Halo framework (Melvin Vivas on X · notes) - Stable Diffusion · tool · stability.ai · free
Open image-generation model family that was the first popular base for image LoRAs.
Mentioned in: Evolution of Image LoRAs: Stable Diffusion → Flux → Krea 2 (Melvin Vivas on X · notes) - Together AI Fine-tuning Overview (docs) · docs · docs.together.ai · free
Together AI's guide to fine-tuning models on your own data and deploying them.
Mentioned in: Fine-tune and Deploy Qwen3.8 27B on Together AI (Melvin Vivas on X · notes) - Unsloth Docker installation guide · docs · unsloth.ai · free
Unsloth's official guide to installing and using its Docker image for local model training.
Mentioned in: Train models locally with the Unsloth Docker image (Melvin Vivas on X · notes) - Unsloth Documentation - Gemma 4 guide · docs · unsloth.ai · free
Unsloth's guide to running and fine-tuning Gemma 4 models.
Mentioned in: Run Gemma 4 12B on 8GB RAM with Unsloth Dynamic GGUFs (Melvin Vivas on X · notes) - Unsloth Joins the PyTorch Ecosystem · article · unsloth.ai · free
Unsloth's blog post announcing that it has joined the PyTorch Ecosystem.
Mentioned in: Unsloth joins the PyTorch Ecosystem (Melvin Vivas on X · notes)
Build
- Fine-tune an open-source TTS model on your own voice, with an agent that starts a Hugging Face GPU and runs the training. (from Using an ML agent to train an open-source TTS model on your voice)
- Run inference with a small open-weight model on Colab's free T4, then fine-tune it with one of the creator's notebooks. (from Run Local Models on a Free GPU with Google Colab (T4))
- Build a custom LLM router that sends minor tasks (e.g., doc cleanup, spelling fixes) to a cheaper model and harder tasks to a stronger model, using a fine-tuned ModernBERT classifier trained on your own traces. (from Train Your Own LLM Model Router by Fine-Tuning ModernBERT as a Classifier)
- Train or fine-tune your own small model by following the Hugging Face LLM Course. (from Hugging Face LLM Course: Chapter 7.3 for Training Your Own Model)
- Train a custom prompt router that picks between models (e.g. Astra, Luna) from your own prompt history, and compare its cost and time against a single model (from Model Routing: Fine-Tune Your Own Router on Your Prompts)
- Fine-tune a small model to act as a router that picks a cheap or a strong model for each task. (from Base vs fine-tuned Gemma 4 E2B as a model router)
- Build a cost-saving model router: a fine-tuned small model that sends each coding task to a cheap or a strong model, connected to Codex. (from QLoRA fine-tune Gemma 4 E2B with Unsloth as a model router)
- Fine-tune a ModernBERT classifier to route tasks between two models or agents. (from Free Colab Notebooks: ModernBERT, Gemma QLoRA, GLiNER & Guardrails)
- QLoRA fine-tune Gemma 4 E2B on your own posts so it writes in your style. (from Free Colab Notebooks: ModernBERT, Gemma QLoRA, GLiNER & Guardrails)
- Add local prompt and response guardrails with a small moderation model. (from Free Colab Notebooks: ModernBERT, Gemma QLoRA, GLiNER & Guardrails)
- Build a knowledge graph from text with zero-shot GLiNER extraction. (from Free Colab Notebooks: ModernBERT, Gemma QLoRA, GLiNER & Guardrails)
- Fine-tune a small open model (e.g. Gemma4-E2B) with Unsloth on your own tweets or posts so it writes in your style. (from Fine-tuning Gemma4-E2B on your own tweet style with Unsloth)
- Fine-tune a small model (e.g. Qwen3.5-2B) with LoRA on your own posts so it writes in your voice. (from LoRA Fine-Tune Qwen3.5-2B on Your Tweets with Unsloth Studio)
- Fine-tune ModernBERT-base as a router that sends tasks to different agents. (from Fine-Tune ModernBERT-base as a Task-Routing Classifier)
- Train and run a tiny decision model locally on a MacBook. (from Kev-0.5B: A Tiny Open-Source Decision Model to Train on a MacBook)
- Build a cost-saving LLM router: fine-tune ModernBERT to send each prompt to a cheap or a strong model, then measure cost savings and quality. (from Train ModernBERT as a Prompt Router Between Two Models (Colab Notebook))
- Fine-tune ModernBERT-base to classify tasks and route each one to one of two models or agents. (from Fine-tuning ModernBERT-base as a task router (quote post))
- Build a task router: fine-tune ModernBERT-base on labeled tasks to choose which of two models or agents handles each one. (from Fine-tuning ModernBERT-base to route tasks between two models)
- Post-train a small model on your own social posts so it writes in your style. (from Post-training Qwen3.5-2B on your own X posts with Unsloth Studio)
- Build a tool that turns a topic or brief into a post written in your voice. (from Post-training Qwen3.5-2B on your own X posts with Unsloth Studio)
- Fine-tune a small model (for example Qwen3.5-2B) with LoRA on your own posts so it writes in your style. (from Fine-Tuned Qwen3.5-2B LoRA Model That Writes X Posts in Your Style)
- Fine-tune a small model on your tweets to copy your writing style. (from First LoRA Run on Qwen3.5-2B with Codex as Training Companion)
- Fine-tune a small Qwen2.5 model on synthetic task-specific training pairs generated by a coding agent. (from Using Codex to Generate Synthetic Training Pairs for Fine-tuning Qwen2.5)
- Fine-tune a small model on your most popular social posts so it writes in your style. (from Training a Small Model on Your Own Popular Tweets)
- Distill a large model into a small task-specific model by fine-tuning it on teacher-generated examples. (from Model Distillation Explained: A Teacher Model Trains a Smaller Student)
- Fine-tune a small model on your own successful social posts to generate content in your writing style, then open-source it. (from Idea: Fine-Tune a Small Model on Your Best Tweets)
- Fine-tune a small (~350M) model with GRPO in TRL so it reliably outputs JSON matching your own schema, then compare valid-output rates before and after. (from Fine-Tune LFM2.5-350M with GRPO in TRL for Structured Outputs)
- Use a coding agent inside the GPU container to run a fine-tuning job (from Ready-made Docker image for running coding agents on ML/fine-tuning jobs)
- Build a receipt/document parser that uses NuExtract3 to output Markdown or schema-based JSON. (from NuExtract3: Open 4B VLM for OCR and Structured JSON Extraction)
- Collect your own coding agent traces and fine-tune a small model on them. (from Fine-Tune a Model on Your Own Coding Agent Traces)
- Fine-tune LFM2.5 Encoder on your own agent traces to build a personalized model router. (from Model Routing Fine-Tuned on Your Agent Harness Traces)
- Fine-tune Qwen3.8 27B on your own dataset and deploy it to a dedicated endpoint. (from Fine-tune and Deploy Qwen3.8 27B on Together AI)
- Compare a regular Q4 GGUF quantization with a QAD 4-bit checkpoint on the same eval set. (from Quantization-Aware Distillation (QAD) for Better 4-bit GGUF Models)
- Fine-tune Muse Glimmer 30B locally on a 24GB GPU for your own task with Unsloth. (from Fine-Tune Muse Glimmer 30B for Free with Unsloth (incl. GRPO))
- Fine-tune a small Liquid AI LFM model on a custom dataset using LEAP, with a coding agent writing the training scripts. (from Smoke-Testing LFM Fine-tuning with Codex and Liquid AI's LEAP)
- Fine-tune a small LFM model with SFT and then DPO using Unsloth or TRL, following the cookbook. (from Liquid AI Cookbook: Fine-Tuning LFMs with CPT, SFT, DPO and GRPO)
- Fine-tune a small Liquid AI model with LEAP for one specific use case and run it locally or on a phone. (from Small Local Models + Fine-Tuning Instead of More Compute (Liquid AI, LEAP))
- Build your own dev tool (like a JSONL dataset viewer/editor) with an AI coding assistant when existing tools fall short. (from Open-Source JSONL Viewer for Fine-Tuning Datasets)
- Post-train a small model through chat using ML Intern. (from ML Intern: Post-Train a Model by Chatting (Hugging Face))
- Fine-tune a small open model on a Hugging Face dataset and record how it differs from the base model. (from Fine-Tuning Models with Unsloth Studio)
- Fine-tune a small Gemma model with Unsloth for a specific work task. (from Fine-tuning Gemma 3 270M with Unsloth for Work Tasks)