GANs, Mode Collapse & Pix2Pix — What I Built This Semester
Hajra | 22F-3443 | FAST CFD
GANs, Mode Collapse & Pix2Pix — What I Built This Semester
Hajra | 22F-3443 | FAST CFD
This semester I implemented two GAN systems from scratch for my Generative AI course.
For Question 1, I built a DCGAN and a WGAN-GP trained on Pokémon sprites and Anime Faces. DCGAN suffered from mode collapse — the generator kept repeating the same outputs instead of generating diverse images. WGAN-GP fixed this by replacing Binary Cross Entropy with Wasserstein loss and adding a gradient penalty, giving the model a smoother, more stable training signal.
For Question 2, I built a Pix2Pix model for sketch-to-photo translation using the CUHK Face Sketch dataset and an Anime Sketch Colorization dataset. The U-Net generator preserves structural detail through skip connections, while the PatchGAN discriminator focuses on local texture quality. Combined with L1 reconstruction loss, the results were surprisingly convincing.
The biggest lesson? Good loss functions matter more than complex architectures. WGAN-GP did not win because it was bigger — it won because it gave a better learning signal.
메타데이터
- post_id
- bfd448733ecc
- slug
- gans-mode-collapse-pix2pix-what-i-built-this-semester-bfd448733ecc
- url
- https://medium.com/@f223443/gans-mode-collapse-pix2pix-what-i-built-this-semester-bfd448733ecc
- canonical_url
- https://medium.com/@f223443/gans-mode-collapse-pix2pix-what-i-built-this-semester-bfd448733ecc
- author_url
- https://medium.com/@f223443
- status
- ok
- fetched_at
- 2026-06-09 15:37:30