Stable Diffusion : l'architecture complète — UNet, VAE et CLIP

Stable Diffusion (Rombach et al., 2022) a rendu accessible la génération texte-image de haute qualité en combinant trois idées puissantes indépendamment : exécuter la diffusion dans un espace latent compressé (VAE), conditionner des incorporations de texte riches (CLIP) et utiliser un U-Net pour le débruitage. Comprendre comment ces trois composants s'articulent est la clé pour comprendre pourquoi le SD fonctionne et comment le contrôler.

Full content is available with a subscription.
Get full access to all courses on the platform for one year with a single payment.
Unlike other platforms that charge per course, here you get everything for one price, and after one year of use there will be no automatic charge for the following year.