← Back to list

GANs, Mode Collapse & Pix2Pix — What I Built This Semester

Hajra | 22F-3443 | FAST CFD

F223443 Hajra Shehzad · 2026-04-27 11:03 · 0 claps · 0.7 min read
#pix2pix #gans #mode-collapse #wgan
Open on Medium ↗

GANs, Mode Collapse & Pix2Pix — What I Built This Semester

Hajra | 22F-3443 | FAST CFD

This semester I implemented two GAN systems from scratch for my Generative AI course.

For Question 1, I built a DCGAN and a WGAN-GP trained on Pokémon sprites and Anime Faces. DCGAN suffered from mode collapse — the generator kept repeating the same outputs instead of generating diverse images. WGAN-GP fixed this by replacing Binary Cross Entropy with Wasserstein loss and adding a gradient penalty, giving the model a smoother, more stable training signal.

For Question 2, I built a Pix2Pix model for sketch-to-photo translation using the CUHK Face Sketch dataset and an Anime Sketch Colorization dataset. The U-Net generator preserves structural detail through skip connections, while the PatchGAN discriminator focuses on local texture quality. Combined with L1 reconstruction loss, the results were surprisingly convincing.

The biggest lesson? Good loss functions matter more than complex architectures. WGAN-GP did not win because it was bigger — it won because it gave a better learning signal.


메타데이터
post_id
bfd448733ecc
slug
gans-mode-collapse-pix2pix-what-i-built-this-semester-bfd448733ecc
url
https://medium.com/@f223443/gans-mode-collapse-pix2pix-what-i-built-this-semester-bfd448733ecc
canonical_url
https://medium.com/@f223443/gans-mode-collapse-pix2pix-what-i-built-this-semester-bfd448733ecc
author_url
https://medium.com/@f223443
status
ok
fetched_at
2026-06-09 15:37:30