spot_img
HomeResearch & DevelopmentAI's Grasp of Gags: Fine-Tuned Decoders Match Encoders in...

AI’s Grasp of Gags: Fine-Tuned Decoders Match Encoders in Humor Classification

TLDR: A study investigated Large Language Models’ (LLMs) ability to classify humor, finding that a fine-tuned decoder model (GPT-4o) performed as well as the best fine-tuned encoder (RoBERTa) in categorizing English jokes into five humor types and a “no-joke” class. This suggests decoders, traditionally text generators, can achieve high performance in complex classification tasks like humor understanding when specifically trained.

For decades, the idea of a machine truly understanding human language, especially its nuances like humor, has been a distant dream. While Large Language Models (LLMs) have made incredible strides in various natural language processing tasks, the question of whether they genuinely “get” a joke has remained open. A recent research paper, “Decoders Laugh as Loud as Encoders”, delves into this fascinating challenge, offering some surprising insights into how different types of LLMs classify humor.

The study, conducted by Eli Borodach, Raj Dandekar, Rajat Dandekar, and Sreedath Panat from Vizuara AI Labs, set out to compare the performance of various LLM architectures—Encoders, Encoder-Decoders, and Decoders—in classifying English jokes into specific categories. The researchers defined five main types of humor: absurdity, dark, irony, wordplay, and social commentary, along with a “no-joke” category for regular sentences.

The Challenge of Humor for AI

Humor is complex, varying across cultures and relying on subtle linguistic cues, unexpected twists, and shared context. For an AI, this presents a significant hurdle. Previous research often focused on feature engineering (manually identifying elements like ambiguity, incongruity, or sentiment) or automated feature extraction (like word embeddings). More recently, deep learning and transfer learning approaches, utilizing pre-trained models like BERT and RoBERTa, have shown promise.

Methodology: How the Models Were Tested

To assess humor understanding, the researchers collected a dataset of English jokes, carefully categorizing them and manually filtering out ambiguous examples, especially those that could fall into multiple humor types like wordplay. They also included an equal number of non-humorous sentences as negative examples. The final dataset comprised 1392 sentences across six categories.

The study then put various LLMs to the test:

  • Encoders: Models like RoBERTa, BERT, and ALBERT were fine-tuned on the humor dataset for 20 epochs. These models are typically strong at understanding and encoding input text.

  • Encoder-Decoders: Models such as BART and Flan-T5 were evaluated using zero-shot and few-shot learning, meaning they were given minimal or no specific training on the humor classification task.

  • Decoders: Large generative models like Llama, Gemma, Qwen2, Mistral, and GPT-4 were also tested with zero-shot and few-shot learning. Crucially, GPT-4o, a powerful decoder, was also fine-tuned on the dataset using the OpenAI API.

The performance was measured using the F1-macro score, a metric chosen because the number of examples in each humor category was uneven, ensuring fair evaluation across all categories.

Surprising Results: Decoders Keep Up

The findings challenged some conventional wisdom. While encoders, particularly RoBERTa-base, showed excellent performance (Mean F1-macro score of 0.8566), the fine-tuned GPT-4o decoder achieved a remarkably similar F1-macro score of 0.8522. A statistical analysis confirmed that there was no significant difference between the performance of the best fine-tuned encoder (RoBERTa) and the fine-tuned decoder (GPT-4o).

This is particularly noteworthy because decoders are primarily designed for text generation, not classification. Their ability to perform on par with models specifically built for understanding and classifying text suggests a deeper, more generalized understanding of language, even in complex areas like humor, when given targeted fine-tuning.

In contrast, models tested with zero-shot and few-shot learning, including other decoders and encoder-decoders, generally lagged significantly behind the fine-tuned models. This indicates that while large pre-trained models have broad knowledge, specific fine-tuning is still key for nuanced tasks like humor classification.

Also Read:

Conclusion: A Step Towards Truly Humorous AI

The study concludes that fine-tuned decoders can indeed understand and classify humor as effectively as fine-tuned encoders. This marks a significant step forward in AI’s ability to grasp one of the most intricate aspects of human communication. While limitations exist, such as the relatively small dataset and potential biases in humor categories, the research opens new avenues for developing AI that can not only generate text but also genuinely comprehend its underlying meaning, including the subtle art of a good joke.

Ananya Rao
Ananya Raohttps://blogs.edgentiq.com
Ananya Rao is a tech journalist with a passion for dissecting the fast-moving world of Generative AI. With a background in computer science and a sharp editorial eye, she connects the dots between policy, innovation, and business. Ananya excels in real-time reporting and specializes in uncovering how startups and enterprises in India are navigating the GenAI boom. She brings urgency and clarity to every breaking news piece she writes. You can reach her out at: [email protected]

- Advertisement -

spot_img

Gen AI News and Updates

spot_img

- Advertisement -