TLDR: This research explores using large language models (LLMs) to neutralize personality test items, aiming to reduce social desirability bias. The study used GPT-o3 to rewrite the IPIP-BFM-50, and participants completed either the original or AI-neutralized version. Results showed preserved reliability for most traits (improving for Conscientiousness, declining for Agreeableness and Openness) and maintained the five-factor structure. While correlations with social desirability decreased for several items, the effect was inconsistent, with some items showing an increase. The findings suggest AI neutralization is a promising but imperfect method, requiring further refinement and human involvement for optimal results.
Personality assessments are widely used, but they often face a challenge called social desirability bias. This bias occurs when people answer questions in a way they think will be seen favorably by others, rather than giving truly honest responses. This can distort the real measurement of traits like agreeableness or conscientiousness. Traditionally, reducing this bias has been a labor-intensive process, involving manual rewriting of survey items.
A recent study explored a novel approach: using large language models (LLMs) to automate this item neutralization process. The researchers, Sirui Wu and Daijin Yang, investigated whether an LLM could effectively rewrite personality test items to reduce social desirability bias without compromising the test’s reliability or its underlying structure. Their work, titled Automated Item Neutralization for Non-Cognitive Scales: A Large Language Model Approach to Reducing Social-Desirability Bias, offers valuable insights into the potential and limitations of AI in psychological assessment.
The AI’s Role in Neutralizing Items
The study utilized GPT-o3, a large language model, to rewrite items from the International Personality Item Pool Big Five Measure (IPIP-BFM-50). This particular personality scale measures the Big Five traits: Extraversion, Agreeableness, Conscientiousness, Neuroticism, and Openness to Experience. The researchers carefully designed a prompt for GPT-o3, instructing it to act as an expert psychometrician. The prompt incorporated established strategies for debiasing, such as reducing evaluative language and preserving the original behavioral meaning of the items. It also used advanced prompting techniques like role-playing and chain-of-thought reasoning to guide the AI.
To evaluate the AI-neutralized items, 203 participants were divided into two groups. One group completed the original IPIP-BFM-50, while the other completed the AI-neutralized version. Both groups also completed the Marlowe–Crowne Social Desirability Scale, a standard measure for assessing socially desirable responding. This setup allowed the researchers to compare the two versions directly.
Key Findings: A Mixed Bag of Success
The results showed a nuanced picture of AI-assisted neutralization:
-
Reliability: Overall, the reliability of the personality scales was largely preserved. Extraversion and Neuroticism maintained high reliability in both versions. Conscientiousness actually saw an improvement in reliability in the neutralized form. However, Agreeableness and Openness experienced a decrease in reliability, with Agreeableness dropping to lower levels. This suggests that while the AI generally maintained internal consistency, some traits were more sensitive to the neutralization process.
-
Factor Structure: The study confirmed that the intended five-factor structure of personality traits was maintained in both the original and AI-neutralized versions. This means the AI didn’t fundamentally alter what the test was measuring. However, while the basic structure was consistent, the way items related to these factors (metric invariance) and their intercepts (scalar invariance) did change across versions. This implies that while the traits themselves were still being measured, direct comparisons of average scores between the original and neutralized forms might not be straightforward without further statistical adjustments.
-
Social Desirability Linkage: The core goal was to reduce the correlation with social desirability. For several items, this was successfully achieved, meaning participants were less influenced by social desirability when responding to the AI-neutralized questions. For example, items related to Extraversion and Openness showed reduced links to social desirability when wording shifted from status claims to concrete behaviors. However, the effects were inconsistent. Notably, some Agreeableness items actually showed an *increase* in correlation with social desirability. This was attributed to neutralized wording that introduced hedges or explicit social judgment cues, inadvertently inviting more impression management.
Implications and Future Directions
The study highlights the significant potential of AI-assisted item editing as a viable tool for reducing response bias in psychological assessments. LLMs demonstrate an ability to recognize and reproduce human-like social desirability biases, making them suitable for identifying and mitigating such effects. However, the findings also underscore current limitations. A single, one-time output from a prompt may not always achieve ideal results, suggesting the need for more sophisticated approaches.
Future developments should focus on domain-specific fine-tuning of LLMs, using high-quality human-edited data to improve their ability to generate valid revisions. The researchers also suggest incorporating multi-agent systems or human-in-the-loop validation processes. This would involve iterative refinement where AI generates items, humans review and test them, and the results inform further AI adjustments. Such a collaborative approach, mirroring traditional psychometric standards, could lead to more stable and valid measurement tools.
The study acknowledges limitations, including the use of a single language and instrument, and a low-stakes context, which might not fully capture social desirability effects in high-stakes situations. Future research should explore these areas, including testing in diverse populations and contexts, and evaluating predictive validity with external outcomes.
Also Read:
- Navigating the Future of Healthcare: A Deep Dive into Large Language Models in Medicine
- Explaining Recommendations: How AI Bridges Model Logic and User Understanding
Conclusion
In conclusion, AI-based neutralization offers a promising path to creating fairer psychological assessments by reducing social desirability bias. While it successfully preserved the Big Five construct structure and improved reliability for some traits, inconsistencies in bias reduction and measurement invariance indicate that the technology is still evolving. With continued refinement, iterative processes, and human oversight, large language models can become powerful allies in enhancing the quality and fairness of non-cognitive assessments.


