AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Llama.cpp Quants Are Now Within Reach In Transformers on ThorstenMeyerAI.com

Buying for a business?Offer from Amazon

Get business pricing on tech for your team

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

TL;DR

Hugging Face has added support for loading GGUF quantized models through Transformers’ from_pretrained API. The feature is on the main branch, initially targets Apple Silicon and Qwen3.5, and uses llama.cpp’s ggml kernels; broader hardware and architecture support has no announced timeline.

Hugging Face has added GGUF model support to the main branch of its Transformers library, as described in the original analysis, allowing users to load quantized checkpoints through the familiar from_pretrained API. The initial rollout targets Apple Silicon Macs and the Qwen3.5 architecture, and reuses llama.cpp’s ggml kernels to run models locally.

Users select a GGUF checkpoint hosted on the Hugging Face Hub and pass its filename through the gguf_file argument to from_pretrained. Hugging Face says this allows text generation without extra configuration. The same checkpoints can also be served through transformers serve, which exposes an OpenAI-compatible endpoint on the user’s machine for clients configured to connect to that address.

On supported Apple Silicon systems, Transformers can keep weights packed on Metal and load compatible ggml/Metal layer kernels. For attention, it uses ggml-org/ggml-attn when available. If that kernel cannot be fetched, the library falls back to PyTorch’s standard sdpa attention with a warning; users can also request sdpa directly. The loader requires a compatible PyTorch release and version of the kernels library. Without a compatible quantization kernel, it dequantizes the model, which uses more memory.

Hugging Face’s published comparison covers three GGUF checkpoints: a small dense model, a larger dense model and a mixture-of-experts model. The company identifies llama.cpp as its reference for local inference performance. The supplied announcement does not give enough detail here to establish a single performance result across hardware or models, so users should consult the benchmark conditions before drawing comparisons.

At a glance
updateWhen: Available on the Transformers main bran…
The developmentHugging Face added a main-branch feature that lets Transformers load GGUF checkpoints through from_pretrained for local inference.
At a glance
announcementWhen: announced April 2026; available via tra…
The developmentHugging Face announced that the transformers library can now run llama.cpp-style GGUF quantized models natively, using ggml kernels for near-llama.cpp performance on Apple Silicon.

GGUF Joins the Transformers Workflow

The feature connects two common approaches to local model use: GGUF checkpoints, associated with llama.cpp-based tools, and the PyTorch-based Transformers library. Developers already using Transformers can try supported quantized Hub checkpoints without switching to a separate inference application, while the serving option lets local clients use an OpenAI-compatible interface.

Quantization can reduce the memory needed to store model weights. Hugging Face lists Unsloth’s Qwen3.5-4B at 8.42 GB in BF16 and 2.74 GB in Q4_K_M. That difference can make a checkpoint more practical on a laptop, though file size alone does not establish its speed or output quality. Those depend on the model, hardware, quantization level and task.

The change may be useful to developers who want to work with quantized checkpoints while keeping their existing Transformers code and tools. Its immediate reach is narrower for people using other platforms: the announced support focuses on Apple Silicon, and no schedule is given for CUDA, Linux or Windows.

From Separate Tools to One API

GGUF is a file format used to package model weights and metadata, including tokenizer information and, in some files, a chat template. It supports different quantization levels that trade weight precision for a smaller memory footprint. Formats such as Q4_K_M use mixed tensor precision, with most weights represented at four bits and some sensitive tensors kept at higher precision.

Hugging Face lists Qwen3.5-4B at 3.53 GB for Q6_K, 3.14 GB for Q5_K_M and 2.74 GB for Q4_K_M, compared with 8.42 GB in BF16. The company suggests starting with Q4_K_M and trying higher-precision variants if memory allows. It also says quality effects vary by model and task, making evaluation on a user’s own workload necessary.

Before this addition, users generally ran GGUF checkpoints through llama.cpp-derived software rather than loading them within the Transformers stack. Hugging Face’s announcement presents the new path as an option for people who want to use those checkpoints through its APIs. The implementation remains in development on the main branch, ahead of a stable release.

““We’re adding support for running GGUF models efficiently in transformers, so you can use checkpoints sized for your laptop’s memory through the familiar transformers APIs.””

— Hugging Face announcement

Hardware and Model Coverage Still Narrow

The announcement sets the initial scope at Apple Silicon and Qwen3.5. It does not give dates for support on CUDA, Linux or Windows, or say when additional architectures will be added. Although the benchmark set includes a mixture-of-experts model, that alone does not establish broad support for other model families.

The feature is currently available on the Transformers main branch, and Hugging Face has not announced when it will enter a stable release. Performance also depends on the specific device, checkpoint and kernels available. The announcement’s benchmark comparison should be read with those conditions in mind, rather than as a general guarantee that Transformers will match llama.cpp in every setup.

Users will also need to check whether their PyTorch and kernels-library versions are compatible. If the quantization kernel is unavailable, the fallback dequantizes weights and uses more memory. The amount of any quality change from quantization remains dependent on the model and task.

Stable Release and Wider Support

The next practical milestone is a stable Transformers release that includes the feature; no release date has been provided. Until then, users who want to try it need the main-branch version and compatible dependencies. Hugging Face’s GGUF documentation and kernels library are the places to check for changes to format and kernel support.

Beyond that release, the open questions are whether Hugging Face will add more model architectures and support for other hardware platforms, including CUDA systems. The company has not published a schedule for either. Users can test the current implementation with their own checkpoint and workload, paying attention to memory use, generation speed and output quality.

Key Questions

What changed in Transformers?

Users can load supported GGUF checkpoints through from_pretrained by providing a gguf_file argument. The feature is on the main branch and can also serve models through transformers serve.

Which devices and models are supported?

The initial support targets Apple Silicon Macs and the Qwen3.5 architecture. Hugging Face has not announced a schedule for other hardware or model families.

Does Transformers match llama.cpp’s speed?

Hugging Face uses llama.cpp as its performance reference and says the feature reuses ggml kernels. Results depend on the model, hardware and available kernels; the announcement does not establish one speed result for all systems.

Will quantization affect model quality?

It can, and the effect depends on the model and task. Hugging Face recommends evaluating quantized checkpoints on the workload where they will be used.

Primary source: Hugging Face · via ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

AI: RSI Joins The AI Roadmap To ‘Super Intelligence’. AI-RTZ #1219

AI-RTZ #1219 reports recursive self-improvement joins a roadmap to superintelligence. Details remain scarce.

Best AI-Powered Automation Software Compared

Compare leading AI automation tools to find the best fit for your business, weighing ease of use, features, scalability, and cost.

2026’S Leading AI-Powered Home Automation Gadgets: The Ultimate List

Discover the 11 best AI-driven home automation devices in 2026, from versatile hubs to smart thermostats, designed for seamless, intuitive living.

Could ChatGPT Ads’ Expansion In Southeast Asia And Taiwan Reshape AI Advertising?

OpenAI says ChatGPT Ads is expanding to Southeast Asia and Taiwan, adding new markets while leaving launch details, pricing and ad formats unclear.