Run GGUF models directly in Hugging Face Transformers
Melvin Vivas · X post · 2026-09-22 · Open on X
Topics: LLMOps, Deployment & Monitoring, LLM Fundamentals · Level: intermediate
Summary
Hugging Face Transformers can now run GGUF models directly. The work brings ggml's Metal kernels into the transformers ecosystem for better compatibility and performance, especially on Apple Silicon.
Key points
- GGUF quantized models can now be loaded and run directly with transformers.
- ggml's Metal kernels are integrated, which speeds up inference on Apple Silicon Macs.
- This connects the llama.cpp/ggml ecosystem with the transformers ecosystem.
Resources mentioned
- Hugging Face Transformers · tool · github.com · free
Open-source library for loading and running pretrained models, which now supports GGUF directly.
Also in: Run llama.cpp GGUF Checkpoints in Hugging Face Transformers (Melvin Vivas on X · notes), AIBackends 0.4.0: Liquid AI LFM2.5 Models on llama.cpp and Transformers (Melvin Vivas on X · notes), Gemma 4 Gets Up to 3x Faster with MTP Drafters (Melvin Vivas on X · notes) - ggml · repo · github.com · free
Tensor library behind llama.cpp and the GGUF format, with Metal kernels for Apple GPUs.
Try this
- Try loading a GGUF model with transformers on a Mac to test the Metal kernel speedup.
More in LLMOps, Deployment & Monitoring
- GPT-6 Prompt Caching: Why It Stretches Your Usage Limits
- Run Qwen-Image-2.1 locally on 12GB VRAM with Unsloth GGUFs
- Run llama.cpp GGUF Checkpoints in Hugging Face Transformers
- Devin Fusion: Multi-Model Routing to Cut Agentic Coding Costs
- Fly.io Sprites Get a Price Cut
- Building a Local Model Server with ONNX Support