Alignment: SFT, reward models, PPO and DPO
A pretrained language model is extremely capable but completely unguided. Ask it a question and it might produce more questions, write a short story, or generate toxic content — all as equally valid text continuations. Alignment is the process of taking this raw capability and shaping it into a model that is helpful, honest, and harmless. This lesson covers the three-stage pipeline that turns a base model into ChatGPT or Claude.