Quantization และการ inference ที่เร็วขึ้น: GPTQ, AWQ, speculative decoding

การรันโมเดล 70B แบบ FP16 ต้องการหน่วยความจำ GPU ราว 140GB สำหรับน้ำหนักเท่านั้น — ยังไม่นับ KV-cache, activation หรือ batch Quantization แมปน้ำหนักความละเอียดสูงไปยังบิตน้อยลง (INT8, INT4 หรือต่ำกว่า) เพื่อให้โมเดลเดียวกันพอใส่ฮาร์ดแวร์ที่ถูกลงและรันเร็วบน tensor core ที่เหมาะกับเลขจำนวนเต็ม บทเรียนนี้ผ่านวิธี post-training quantization (PTQ) ที่ใช้จริงในโปรดักชัน — GPTQ และ AWQ — และ speculative decoding เทคนิคเสริมที่เร่งการสร้างโดยไม่เปลี่ยนน้ำหนัก

Full content is available with a subscription.
Get full access to all courses on the platform for one year with a single payment.
Unlike other platforms that charge per course, here you get everything for one price, and after one year of use there will be no automatic charge for the following year.