spot_img
HomeResearch & DevelopmentConditional-t3VAE: A New Approach for Fair Image Generation in...

Conditional-t3VAE: A New Approach for Fair Image Generation in Imbalanced Datasets

TLDR: Conditional-t3VAE (C-t3VAE) is a novel generative model that addresses class imbalance in datasets by ensuring equitable latent space allocation for all classes. Unlike previous VAEs and t3VAE, which allocate latent volume based on class frequency, C-t3VAE uses per-class Student’s t-distribution priors and an equal-weight latent mixture for balanced sampling. This leads to significantly lower FID scores and improved generative fairness and diversity, especially in highly imbalanced settings (imbalance ratio ρ > 3), outperforming Gaussian-based VAEs and t3VAE.

Generative models, like Variational Autoencoders (VAEs), are powerful tools for creating new data, such as images. However, a common challenge arises when these models are trained on datasets where some classes are much more frequent than others – a situation known as class imbalance. In such scenarios, VAEs tend to favor the dominant classes, leading to an underrepresentation of rare classes in the generated output. This can result in biased or unfair generations, which is particularly problematic in sensitive applications like facial synthesis or medical imaging.

Traditional VAEs often use simple Gaussian priors, which struggle to accurately model complex, heavy-tailed data distributions and don’t inherently address class imbalance. While the t3VAE model improved upon this by incorporating heavy-tailed Student’s t-distributions and a more robust optimization method (γ-power divergence), it still allocated latent space volume proportionally to the class frequency in the training data. This meant that if a class was rare in the training set, it would still be rare and potentially of lower quality in the generated samples.

Addressing this critical issue, researchers have introduced a new model called Conditional-t3VAE (C-t3VAE). This innovative approach explicitly enforces an equitable allocation of latent space across all classes. Instead of a global prior, C-t3VAE defines a unique Student’s t-distribution prior for each class, spanning both latent and output variables. This clever design prevents majority classes from dominating the latent space, ensuring that every class, regardless of its frequency in the original dataset, receives an equal ‘volume’ of representation.

The C-t3VAE model is optimized using a closed-form objective derived from the γ-power divergence, building on the strengths of its predecessor. For generating class-balanced samples, the model employs an equal-weight latent mixture of Student’s t-distributions, with analytically derived component variances. This mechanism allows for uniform sampling of synthetic data across all classes, effectively mitigating the bias often seen in unconditional generative models.

Extensive experiments were conducted on several datasets, including SVHN-LT, CIFAR100-LT, and CelebA, which represent varying degrees of class imbalance and visual complexity. The results consistently showed that C-t3VAE achieves lower Fréchet Inception Distance (FID) scores – a key metric for generative quality – compared to both t3VAE and traditional Gaussian-based VAEs. This improvement was particularly significant under severe class imbalance conditions.

Beyond overall quality, per-class evaluations using metrics like Precision, Recall, and F1 score further highlighted C-t3VAE’s advantages. It demonstrated superior Recall and F1 scores across highly imbalanced settings, indicating better coverage of data modes and improved diversity, especially for underrepresented ‘tail’ classes. For instance, on the CelebA dataset, C-t3VAE showed better generation for attributes like ‘Mustache’ which are heavily imbalanced.

An interesting finding from the research is the identification of an imbalance ratio threshold, approximately ρ ≈ 3. Below this threshold, Gaussian-based models remain competitive. However, when the imbalance ratio exceeds 3, C-t3VAE substantially improves generative fairness and diversity, with its performance gap widening as the imbalance becomes more extreme. Qualitatively, C-t3VAE also produced sharper and more detailed synthetic images compared to its counterparts.

Also Read:

In conclusion, C-t3VAE offers a principled and flexible framework for fair generative modeling, particularly in real-world scenarios characterized by class imbalance. By ensuring equitable latent space allocation and leveraging the robustness of Student’s t-distributions, it significantly enhances both the fairness and quality of generated samples. This work is a crucial step towards more responsible and unbiased AI generation, with future research aiming to extend its capabilities to multi-label problems and integrate it with advanced Latent Diffusion Models. You can read the full research paper here.

Meera Iyer
Meera Iyerhttps://blogs.edgentiq.com
Meera Iyer is an AI news editor who blends journalistic rigor with storytelling elegance. Formerly a content strategist in a leading tech firm, Meera now tracks the pulse of India's Generative AI scene, from policy updates to academic breakthroughs. She's particularly focused on bringing nuanced, balanced perspectives to the fast-evolving world of AI-powered tools and media. You can reach her out at: [email protected]

- Advertisement -

spot_img

Gen AI News and Updates

spot_img

- Advertisement -