Table of Contents Getting started with Nunchaku Lite Background: SVDQuant and Nunchaku Introducing Nunchaku Lite Native loading in Diffusers Hardware support Getting more speed and lower memory Benchmarks End-to-end latency and memory Image quality Quantizing your own model 1. Inspect what will be quantized 2. Run quantization 3. Package a Diffusers pipeline 4. Load, verify, and push to the Hub Quantizing models with structural rewrites Ready-to-use checkpoints Conclusion Acknowledgements Large diffusion transformers can create stunning images (or even videos, audio snippets, and now text), but loading a modern text-to-image model in BF16 precision often requires 20-30 GB of VRAM, which puts these models out of reach of most consumer GPUs. Quantization is a powerful solution to this problem, and Diffusers already integrates several quantization backends such as bitsandbytes, GGUF, torchao, and Quanto, which we covered in Exploring Quantization Backends in Diffusers.