Run llama.cpp GGUF Checkpoints in Hugging Face Transformers
Melvin Vivas · X video post · 2026-09-22 · Open on X
Topics: LLMOps, Deployment & Monitoring, LLM Fundamentals · Level: intermediate
Summary
Hugging Face Transformers can now load the same GGUF quantized checkpoints used by llama.cpp. On Mac, ggml kernels give fast local inference. The creator plans to refactor his AIBackends project to use this.
Key points
- GGUF checkpoints made for llama.cpp can now run directly in Hugging Face Transformers.
- You use the same quantized models without converting them.
- On Mac, fast local inference comes from ggml kernels.
- Hugging Face hosts the ggml kernels under the ggml-org organization.
- The creator plans to refactor his AIBackends project to take advantage of this.
Resources mentioned
- Hugging Face Blog: Transformers + llama.cpp quants · article · huggingface.co · free
Hugging Face blog post announcing that llama.cpp GGUF quantized checkpoints can run in Transformers. - ggml-org kernels on Hugging Face · repo · huggingface.co · free · open in a browser to verify
ggml kernels on Hugging Face that power fast local inference on Mac in Transformers. - Hugging Face Transformers · tool · github.com · free
Open-source library for loading and running pretrained models, which now supports GGUF directly.
Also in: Run GGUF models directly in Hugging Face Transformers (Melvin Vivas on X · notes), AIBackends 0.4.0: Liquid AI LFM2.5 Models on llama.cpp and Transformers (Melvin Vivas on X · notes), Gemma 4 Gets Up to 3x Faster with MTP Drafters (Melvin Vivas on X · notes) - llama.cpp · repo · github.com · free
An open-source C/C++ engine for running GGUF models locally. Its llama-server command provides an OpenAI-compatible HTTP server.
Also in: Running LLMs Locally Without an Expensive Rig (Melvin Vivas on X · notes), llama.cpp / Llama-macOS v0.5.0 release (Melvin Vivas on X · notes), llama.cpp v0.4.1 release announcement (Melvin Vivas on X · notes), Coworker: An Open-Source Desktop Agent App for Local and Cloud Models (Melvin Vivas on X · notes) and 28 more - AIBackends · repo · aibackends.com · free
Open-source API server runtime for common AI use cases that supports many models and providers (Ollama, LM Studio, OpenRouter, OpenAI, Anthropic).
Also in: Building a production website with Opus 5.5 and TanStack (Melvin Vivas on X · notes), CamelFlow: open-source visual viewer for Apache Camel routes (Melvin Vivas on X · notes), AI Backends: A Production AI Workflow Engineering Site (Link Share) (Melvin Vivas on X · notes), Demo: Claude Opus 5.5 Generating a Motion-Graphics Video for AIBackends (Melvin Vivas on X · notes) and 16 more
Try this
- Read the Hugging Face blog post on running llama.cpp quants in Transformers.
- Try loading a GGUF checkpoint in Transformers for local inference on a Mac.
- Refactor a local-inference backend so it loads GGUF models through Transformers with ggml kernels.
More in LLMOps, Deployment & Monitoring
- How Cursor cut agent token costs by 7% without losing quality
- GPT-6 Prompt Caching: Why It Stretches Your Usage Limits
- Run Qwen-Image-2.1 locally on 12GB VRAM with Unsloth GGUFs
- Run GGUF models directly in Hugging Face Transformers
- Devin Fusion: Multi-Model Routing to Cut Agentic Coding Costs
- Fly.io Sprites Get a Price Cut