deployment
QLoRA
Quantized Low-Rank Adaptation, combining 4-bit quantization with LoRA to enable fine-tuning of large models on a single consumer GPU. QLoRA loads the base model in 4-bit precision while training LoRA adapters in higher precision.
In practice
QLoRA enables fine-tuning a 65B parameter model on a single 48GB GPU, which would normally require multiple high-end GPUs.