What is the GGUF file format? Load GGUF with transformers Serve GGUF with your preferred interface Benchmarking against llama.cpp transformers and llama.cpp Beyond GGUF: ggml kernels for more models Fast local inference with Python and PyTorch Reusing ggml's Metal kernels Keeping the CPU and GPU working together Current limitations and next steps Acknowledgments We're adding support for running GGUF models efficiently in transformers, so you can use checkpoints sized for your laptop's memory through the familiar transformers APIs. Pick a GGUF from the Hub, load it with from_pretrained, and start generating on your own machine.
Running AI models on your laptop has become much easier, and llama.cpp has been a big part of that. Its inference engine powers local AI tools such as Ollama, LM Studio, and Jan. Alongside projects like MLX, it has helped make local inference a practical option for everyday use.




