Hugging Face Blog·· Jul 23SelectedAI score68
Nunchaku Lite 正式接入 Diffusers,4-bit 量化扩散模型推理无需额外推理引擎
Bringing Nunchaku 4-bit Diffusion Inference to Diffusers
AI brief
Nunchaku Lite 现已原生集成到 Diffusers,用户可通过 from_pretrained() 直接加载 4-bit 量化扩散模型,无需单独推理引擎或本地 CUDA 编译。
Why it matters
Nunchaku Lite 现已原生接入 Diffusers,无需额外引擎即可通过 from_pretrained() 加载 4-bit 量化模型,在 RTX 5090 上 VRAM 从 24GB 降至约 12GB,推理加速约 30%。
Source: Hugging Face Blog · huggingface.co