Hugging Face introduces native support for the GGUF format in the transformers library for running models locally on Apple Silicon
Hugging Face has added native support to the transformers library for the GGUF format for local inference on Apple Silicon, including quantization levels such as Q4_K_M that shrink models to a fraction of their original size.
Hugging Face has added native support to the transformers library for the GGUF format, which was developed by the llama.cpp team and is commonly used for quantized models in tools such as Ollama, LM Studio or Jan. According to the company, this is an effort to make local inference accessible directly through the standard transformers API without users having to turn to standalone engines. transformers approaches the performance of llama.cpp by reusing the same ggml kernels through the kernels library. The initial focus is on Apple Silicon and the Qwen3.5 architecture.
The GGUF format stores model weights and metadata (the tokenizer and, where applicable, a chat template) in a single file and supports various quantization levels that reduce memory requirements at the expense of accuracy. Using the Qwen3.5-4B model from Unsloth as an example, the file size varies as follows: the unquantized BF16 variant is 8.42 GB, the Q6_K variant is 3.53 GB, Q5_K_M is 3.14 GB and Q4_K_M is 2.74 GB. The company recommends the Q4_K_M variant as a starting point, noting that the impact of more aggressive quantization on output quality needs to be verified for the specific task.
Using it requires a Mac with Apple Silicon, a supported version of PyTorch and the current (main-branch, not yet released) version of transformers along with the kernels library. The model is loaded by specifying the Hub identifier and passing the GGUF filename to the gguf_file parameter in the from_pretrained function; the standard transformers API is then used. The same checkpoint can also be exposed through transformers serve as an OpenAI-compatible API and connected to clients such as Jan or Pi. If the required quantization kernel is unavailable, the company says the model is automatically dequantized and run with higher memory requirements.
The source article also includes a performance comparison with llama.cpp using llama-bench on several models, but this part of the text was not available in full. Details can be found in the source article.
Why it matters
Developers working with local models can stay with a single tool (transformers) instead of switching between it and standalone engines such as llama.cpp, while retaining access to the same GGUF checkpoints and quantization levels, which they choose according to the available memory on the machine.
Two audiences, two different impacts
What this means
For individuals
Developers running models locally on a Mac with Apple Silicon can now use the standard API of the transformers library instead of standalone tools such as Ollama or LM Studio while using the same GGUF files and quantization levels.
For a business
Companies developing tools for local AI inference gain an alternative to llama.cpp directly within the transformers library, which may simplify the integration and maintenance of tools that work with quantized models.
DevelopmentCheck the original
Event sources
clearly official source · 1 publisher, 0 independent. We count feeds from the same owner only once.