TLDR: A new research paper introduces CreativityPrism, a holistic benchmark to evaluate Large Language Model (LLM) creativity. It decomposes creativity into three dimensions: quality, novelty, and diversity, using nine tasks across divergent thinking, creative writing, and logical reasoning, with twenty metrics. Evaluating 17 LLMs, the study found a significant performance gap between proprietary and open-source models, and revealed that creativity is multi-dimensional, meaning strong performance in one area doesn’t guarantee it in others, emphasizing the need for comprehensive evaluation.
Large Language Models (LLMs) are increasingly seen as capable of generating creative text, but truly understanding and evaluating their creativity has been a significant challenge. The problem stems from the multifaceted nature of creativity itself, leading to fragmented evaluation methods that vary dramatically across different domains and tasks.
A new research paper, titled “CreativityPrism: A Holistic Benchmark for Large Language Model Creativity,” introduces a comprehensive framework designed to address this challenge. Authored by Zhaoyi Joey Hou, Bowei Alvin Zhang, Yining Lu, Bhiman Kumar Baghel, Anneliese Brei, Ximing Lu, Meng Jiang, Faeze Brahman, Snigdha Chaturvedi, Haw-Shiuan Chang, Daniel Khashabi, and Xiang Lorraine Li, this work proposes a unified approach to assess LLM creativity.
The core idea behind CreativityPrism is that creativity isn’t a single, fixed concept. Instead, it can be broken down into three key dimensions: quality, novelty, and diversity. This framework incorporates nine distinct tasks across three major domains: divergent thinking, creative writing, and logical reasoning. To measure these dimensions, CreativityPrism utilizes twenty specific evaluation metrics, each tailored to capture unique aspects of creativity within its respective task.
Understanding the Dimensions of Creativity
Quality evaluates whether an LLM’s output meets the fundamental requirements of a task. For example, in coding tasks, it checks if the code executes successfully, or in writing, if the sentences are coherent and grammatically correct.
Novelty measures the originality of the generated content or solutions. It assesses how different the LLM’s output is compared to existing or commonly seen examples, rewarding unique and unexpected ideas.
Diversity examines the variation among the LLM’s generated content. This dimension captures the model’s ability to produce a wide range of distinct outputs, rather than repetitive or similar responses.
The Tasks and Domains
The framework’s nine tasks are carefully selected to cover a broad spectrum of creative abilities. Divergent thinking tasks, like the Alternative Uses Test (AUT) or the Divergent Association Task (DAT), challenge models to generate numerous varied ideas. Creative writing tasks, such as generating short stories or completing novel paragraphs, test narrative skill and imaginative expression. Logical reasoning tasks, including mathematical problem-solving and code generation with constraints, assess creativity in finding unconventional solutions under strict rules.
Key Findings from the Evaluation
The researchers evaluated 17 state-of-the-art LLMs, encompassing both proprietary (closed-source) and open-source models, using the CreativityPrism benchmark. The results revealed several significant insights:
- A notable performance gap exists between proprietary and open-source models, particularly in logical reasoning and creative writing tasks. Proprietary models generally outperformed their open-source counterparts.
- Model performance has shown improvement over time, with more recently released models demonstrating increased competitiveness.
- Within the same task or domain, model performances tend to be highly correlated. Similarly, quality and diversity metrics show strong correlations, meaning models excelling in one often do well in the other.
- However, novelty metrics exhibited much weaker correlations across different tasks and domains. This suggests that how novelty is defined and measured can vary significantly, highlighting its complex nature.
- Overall, strong performance in one creativity task or dimension does not necessarily guarantee similar excellence in others. This underscores the critical need for a holistic evaluation framework like CreativityPrism.
Also Read:
- AI’s Inner Drive: Assessing Curiosity in Large Language Models
- Enhancing LLM Training: A New Approach to Data Selection Through Orthogonal Diversity
Looking Ahead
CreativityPrism provides a robust foundation for systematically evaluating machine creativity and guiding the future development of more creative LLMs. While the current benchmark is limited to English text and acknowledges potential biases from LLM-as-a-judge evaluations, it offers a crucial step towards a more comprehensive understanding of AI’s creative capabilities. For more details, you can read the full research paper here.


