Alineación: SFT, modelos de recompensa, PPO y DPO

Un modelo de lenguaje previamente entrenado es extremadamente capaz pero completamente carente de guía. Hágale una pregunta y podría generar más preguntas, escribir una historia corta o generar contenido tóxico, todo como continuaciones de texto igualmente válidas. Alineación es el proceso de tomar esta capacidad bruta y darle forma en un modelo que sea útil, honesto e inofensivo. Esta lección cubre el proceso de tres etapas que convierte un modelo base en ChatGPT o Claude.

Full content is available with a subscription.
Get full access to all courses on the platform for one year with a single payment.
Unlike other platforms that charge per course, here you get everything for one price, and after one year of use there will be no automatic charge for the following year.