เสียงและคำพูด: Whisper, TTS และ MusicGen

เสียงเป็นเพียงลำดับอื่น — สัญญาณ 1D สุ่มตัวอย่างที่ 16,000 หรือ 44,100 ครั้งต่อวินาที โมเดลเสียงสมัยใหม่จะแปลงรูปคลื่นดิบให้เป็นการแสดงความถี่เวลาที่เรียกว่าสเปกโตรแกรม จากนั้นใช้สถาปัตยกรรมหม้อแปลงแบบเดียวกับที่ขับเคลื่อน LLM และโมเดลรูปภาพ บทเรียนนี้ครอบคลุมกระบวนทัศน์การสร้างเสียงหลักสามกระบวนทัศน์ ได้แก่ การรู้จำเสียงพูดอัตโนมัติ (Whisper) การแปลงข้อความเป็นคำพูด (TTS) และการสร้างเพลง (MusicGen)

Full content is available with a subscription.
Get full access to all courses on the platform for one year with a single payment.
Unlike other platforms that charge per course, here you get everything for one price, and after one year of use there will be no automatic charge for the following year.