🔍 Read the full analysis: Llama.cpp Quants Come To Transformers on ThorstenMeyerAI.com
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
Hugging Face has added support for loading GGUF quantized checkpoints in Transformers with the familiar from_pretrained API. The initial rollout targets Qwen3.5 models on Apple Silicon and is available on the library’s main branch; broader hardware and architecture support, and a stable release date, have not been announced.
Users can select a GGUF checkpoint hosted on the Hugging Face Hub and pass its file through the gguf_file argument to from_pretrained, then generate text without additional model-loading configuration. Hugging Face says the implementation reuses llama.cpp’s ggml kernels to keep inference performance close to llama.cpp. The stated rollout is limited: it requires an Apple Silicon Mac, a supported PyTorch version, the latest Transformers code and a compatible version of the kernels library.
When weights remain packed for execution on Metal, Transformers can load compatible ggml/Metal layer kernels and use ggml-org/ggml-attn for attention. If that attention kernel is unavailable, the loader falls back to standard “sdpa” attention with a warning; users can also select sdpa directly. Without a compatible quantization kernel, the model can still be loaded by dequantizing its weights, but Hugging Face says that path uses more memory.
The same checkpoints can also be served through transformers serve, which exposes an OpenAI-compatible API on localhost. Clients such as Jan or Pi can connect through a custom provider. Hugging Face says its performance comparison uses llama.cpp as the reference and covers three GGUF checkpoints: a small dense model, a larger dense model and a mixture-of-experts model. The supplied material does not include the benchmark results, so it does not establish a specific speed difference.
GGUF Comes to the Transformers Workflow
The change gives developers already using Transformers and PyTorch a way to work with GGUF checkpoints without changing to a separate inference stack. GGUF is widely used for local model distribution and inference in tools including Ollama, LM Studio and Jan. Support in Transformers could make those checkpoints more accessible to people who want to use its existing model-loading and serving interfaces.
Quantization can reduce the memory needed to store and run a model, which can make local inference more practical on laptops. Hugging Face’s example for Qwen3.5-4B lists a BF16 file at 8.42 GB and a Q4_K_M version at 2.74 GB. Those figures describe file sizes, not a guarantee of runtime memory use or output quality. Hugging Face says the quality effect of more aggressive quantization varies by model and task, and recommends evaluating the chosen checkpoint on the intended workload.
For now, the practical impact is concentrated among Apple Silicon users running supported architectures. The announcement does not set a schedule for support on CUDA, Linux or Windows, or establish when additional model families will work.
From GGUF Files to Transformers
GGUF is a model file format associated with the llama.cpp ecosystem. It packages weights and metadata, which can include tokenizer information and a chat template, and supports multiple quantization levels. Formats such as Q4_K_M use lower precision for much of the model’s weights while retaining higher precision for some tensors, trading a smaller file for potential changes in model quality.
For Unsloth’s Qwen3.5-4B example, the listed sizes are 8.42 GB for BF16, 3.53 GB for Q6_K, 3.14 GB for Q5_K_M and 2.74 GB for Q4_K_M. Hugging Face suggests starting with Q4_K_M and trying Q5_K or Q6_K when more memory is available. These are example file sizes; performance and quality depend on the model, hardware and task.
The feature follows growing interest in running quantized models locally. The source material also describes a recent demonstration by Hugging Face co-founder Julien Chaumond of Qwen3.6 27B running in the Pi coding agent through llama.cpp on a MacBook Pro. That demonstration is separate from the Transformers feature and does not establish equivalent performance for this new integration.
“We’re adding support for running GGUF models efficiently in transformers, so you can use checkpoints sized for your laptop’s memory through the familiar transformers APIs.”
— Hugging Face announcement
Hardware and Model Coverage Remain Limited
The announcement identifies Apple Silicon and Qwen3.5 as the initial targets but does not provide a schedule for CUDA, Linux or Windows support, or specify which additional model architectures will be added. It also does not give a date for the next stable Transformers release containing the feature.
Although Hugging Face says it compared performance with llama.cpp across three checkpoints, the source material provided here does not state the numerical results, test hardware or full conditions. It is therefore not possible to quantify how closely the new path matches llama.cpp from these details. The memory and quality effects of dequantization and quantization will also vary with the kernel, model and user workload.
Stable Release and Broader Support
The next stated milestone is inclusion in a stable Transformers release; until then, users need the main-branch version. Hugging Face has not announced a release date. Users considering the feature will also need to check that their Apple Silicon system, PyTorch version and kernels library are compatible.
Further changes could include additional model architectures or hardware backends, but the announcement gives no dates or confirmed roadmap for those expansions. For now, users can follow the Hugging Face Hub’s GGUF documentation and the kernels library for updates, and test quantized checkpoints on their own tasks to assess memory use and output quality.
Key Questions
How do I load a GGUF checkpoint in Transformers?
On the current main branch, pass the checkpoint file using the gguf_file argument to from_pretrained. The initial implementation also requires compatible software and hardware.
Which hardware and models are supported at launch?
The announced initial support targets Apple Silicon Macs and the Qwen3.5 architecture. The announcement does not specify a timeline for other hardware or model families.
Does Transformers run GGUF models as fast as llama.cpp?
Hugging Face says the implementation reuses llama.cpp’s ggml kernels and compares performance against llama.cpp. The supplied material does not include numerical benchmark results, so it does not establish a specific performance match.
What happens if the quantization kernel is unavailable?
Transformers can fall back to dequantizing the model, which Hugging Face says uses more memory. If the ggml attention kernel cannot be fetched, attention falls back to sdpa with a warning.
When will this reach a stable Transformers release?
The feature is currently on the main branch. Hugging Face has not announced a date for a stable release containing it.
Primary source: Hugging Face · via ThorstenMeyerAI.com
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
