How It Works
Supervised Fine-Tuning (SFT)
Taking a model that has only learned to predict text and training it further on example pairs written or approved by people — a request, and the answer it should have given. It is the step that turns a text predictor into something that follows instructions, and it comes before the preference-based steps (RLHF and its relatives). Its limit is worth knowing: the model learns the style and habits of whoever wrote the examples, so who was hired to write them shapes what the finished assistant sounds like and refuses.
Origin · no single documented coiner
No single coiner — the phrase is ordinary machine-learning vocabulary (“supervised learning” plus “fine-tuning”) that hardened into a named stage of the modern pipeline. The version everyone now cites is the first stage of OpenAI’s InstructGPT paper, “Training language models to follow instructions with human feedback” by Long Ouyang, Jeff Wu and colleagues (March 2022), which labels it SFT and puts it in front of reward modeling and reinforcement learning.
