การเสิร์ฟ LLM: vLLM, TGI, KV-cache, การแบทช์

การฝึกอบรมจะสอนแบบจำลองในการทำนายโทเค็นถัดไป การให้บริการ คือทุกสิ่งทุกอย่างที่เกิดขึ้นเมื่อผู้ใช้เรียกใช้งานจริง เช่น การตั้งเวลาคำขอหลายพันรายการพร้อมกัน การเพิ่ม KV-cache tensors เมื่อลำดับยาวขึ้น การจัดกลุ่มรูปร่างที่เข้ากันไม่ได้โดยไม่ทำให้ GPU RAM เสียเปล่า และทำให้เทนเซอร์คอร์อิ่มตัวภายใต้ SLO เวลาในการตอบสนองจริง เฟรมเวิร์กเช่น vLLM (PagedAttention) และ Text Generation Inference (TGI) มีอยู่เนื่องจากลูป PyTorch ไร้เดียงสาไม่สามารถรองรับโหลดที่ใช้งานจริงได้

Full content is available with a subscription.
Get full access to all courses on the platform for one year with a single payment.
Unlike other platforms that charge per course, here you get everything for one price, and after one year of use there will be no automatic charge for the following year.