TLDR: A research paper introduces HAZEL, a fine-tuned Generative AI chatbot developed to enhance the accessibility and readability of heritage guidance documents for Historic England. The study compares HAZEL’s performance to ChatGPT, finding modest improvements in readability and consistency after fine-tuning. However, it also highlights significant limitations in areas requiring cultural sensitivity, technical expertise, and the risk of misinformation, concluding that GenAI should augment, not replace, human heritage professionals.
Generative Artificial Intelligence (GenAI) is increasingly finding its way into various professional fields, and cultural heritage is no exception. A recent research paper, titled “GENERATIVE AI IN HERITAGE PRACTICE : I MPROVING THE ACCESSIBILITY OF HERITAGE GUIDANCE”, explores the potential of integrating GenAI into professional heritage practice to make public-facing guidance documents more accessible. The study was conducted by Jessica Witte, Edmund Lee, Lisa Brausem, Verity Shillabeer, and Chiara Bonacchi.
The core of their research involved developing a specialized GenAI chatbot named HAZEL. This chatbot was fine-tuned specifically to assist with revising written guidance related to heritage conservation and interpretation. The goal was to enhance the clarity and readability of documents published by organizations like Historic England, which provides crucial advice on caring for and protecting historic places.
To assess HAZEL’s effectiveness, the researchers compared its performance against ChatGPT (GPT-4) in a series of tasks related to guidance writing. The evaluation involved both quantitative assessments, using established readability formulas, and qualitative assessments by a professional copyeditor familiar with Historic England’s style guidelines.
The quantitative results indicated that both HAZEL and ChatGPT succeeded in slightly improving the average readability scores of heritage texts compared to the original documents. HAZEL, in particular, showed a modest but consistent improvement, suggesting that fine-tuning an underlying large language model (LLM) can make it more effective for specific domains. However, it’s important to note that even with these improvements, many technical heritage texts remained classified as ‘difficult’ to ‘very difficult’ for a general audience, highlighting the inherent complexity of the subject matter.
Qualitative assessments by the copyeditor revealed that HAZEL’s revisions were, on average, slightly more readable and accessible than ChatGPT’s, and also demonstrated more uniform performance. This consistency is a valuable trait for public sector organizations that need predictable tone and clarity across their documents.
Despite these promising results, the study also identified significant limitations. ChatGPT, for instance, frequently generated “ghost bibliographic citations”—references to non-existent papers or journals—and exhibited an American English bias, which is problematic for a UK-focused organization like Historic England. While HAZEL, being fine-tuned, mitigated some of these issues, it still showed limitations in areas requiring deep cultural sensitivity and advanced technical expertise. The researchers also noted that both models scored lowest on “Overall Suitability” and HAZEL on “Diversity and Inclusion,” indicating that neither system consistently met all of Historic England’s expectations for high-quality guidance literature.
The research underscores that while GenAI tools can automate and expedite certain aspects of guidance writing, such as surface-level copyediting and simplifying language, they cannot replace human heritage professionals for tasks demanding nuanced understanding, creative problem-solving, or subjective expertise. Ethical concerns, including intellectual property rights, embedded biases, and data privacy, also remain critical considerations for the responsible integration of GenAI in the heritage sector.
Future work suggests exploring interactive web content, retrieval-augmented generation (RAG) techniques, and co-designing models with heritage professionals and accessibility specialists. The study also emphasizes the need for open-access, freely available models to ensure equitable participation, especially for smaller, resource-constrained institutions. For more details, you can read the full research paper here: GENERATIVE AI IN HERITAGE PRACTICE : I MPROVING THE ACCESSIBILITY OF HERITAGE GUIDANCE.
Also Read:
- Evaluating AI’s Deep Research Abilities: A New Standard
- Crafting Authentic Game NPCs: How “Deflanderization” Balances Personality and Purpose
In conclusion, fine-tuned GenAI tools like HAZEL offer valuable benefits for improving the readability and accessibility of heritage texts, supporting efforts towards cultural democracy. However, their deployment requires careful consideration of their limitations and a commitment to human oversight to ensure accuracy, cultural sensitivity, and ethical practice.


