การจัดแนว (Alignment): SFT, โมเดลรางวัล, PPO และ DPO

โมเดลภาษาที่ได้รับการพรีเทรนมีความสามารถสูง แต่ยังไม่ได้รับการจัดแนว (alignment) เลย ถามคำถามเดียวอาจได้คำตอบเป็นคำถามต่อ เรื่องสั้น หรือเนื้อหาเป็นพิษ — ทั้งหมดนับว่าเป็นการต่อข้อความที่ "ถูกต้อง" เท่าเทียมกัน การจัดแนว (alignment) คือกระบวนการนำศักยภาพดิบนี้มาปรับให้เป็นโมเดลที่มีประโยชน์ ซื่อสัตย์ และปลอดภัย บทเรียนนี้ครอบคลุมไปป์ไลน์สามขั้นตอนที่เปลี่ยนโมเดลฐานให้กลายเป็นคล้าย ChatGPT หรือ Claude

Full content is available with a subscription.
Get full access to all courses on the platform for one year with a single payment.
Unlike other platforms that charge per course, here you get everything for one price, and after one year of use there will be no automatic charge for the following year.