How It Works
LoRA (Low-Rank Adaptation)
A cheap way to fine-tune a big model without retraining it. Instead of updating billions of weights, you freeze the original model and train two small extra matrices alongside it, then add their product back in. The trained result is a small file — often a few megabytes — that you attach to the base model, which is why people can share hundreds of style or task “adapters” for one model and swap between them. Most “custom” image and text models people run at home are LoRAs, not new models.
Origin
Introduced by Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang and Weizhu Chen at Microsoft in “LoRA: Low-Rank Adaptation of Large Language Models,” published June 2021.
