Skip to content
Launchpad Library logo
← All glossary terms

How It Works

QLoRA (Quantized LoRA)

LoRA done on a squashed copy of the model. The frozen base model is stored at 4 bits per weight instead of 16, which cuts the memory needed enough to fine-tune a very large model on a single consumer graphics card, while the small trainable adapters stay at full precision. This is the technique behind most “I fine-tuned a big model on my own machine” claims. The honest caveat: squashing the weights costs some accuracy, and the paper’s own evaluation used another AI model as the judge, which is a weaker test than it sounds.

Origin

Introduced by Tim Dettmers, Artidoro Pagnoni, Ari Holtzman and Luke Zettlemoyer at the University of Washington in “QLoRA: Efficient Finetuning of Quantized LLMs,” published May 2023.

Read the source →