spot_img
HomeResearch & DevelopmentAnnif System Leads GermEval-2025 LLMs4Subjects Task with Hybrid AI...

Annif System Leads GermEval-2025 LLMs4Subjects Task with Hybrid AI for Subject Indexing

TLDR: The paper details the Annif system’s winning approach at GermEval-2025’s LLMs4Subjects task, focusing on automated subject indexing. It combines traditional extreme multi-label text classification (XMTC) with efficient large language models (LLMs) for data translation, synthetic data generation, and ranking candidate subjects. This hybrid system, using smaller LLMs for efficiency, achieved first place in both quantitative and qualitative evaluations, demonstrating that a balanced approach of traditional methods augmented by LLMs can outperform LLM-heavy systems.

Automated subject indexing is a crucial process for improving the discoverability of bibliographic databases and digital collections. The GermEval-2025 LLMs4Subjects Shared Task challenged teams to develop innovative systems using large language models (LLMs) for this purpose, with a strong emphasis on computational efficiency.

The Annif system, a long-standing automated subject indexing toolkit, participated in Subtask 2 of this challenge. Building upon its previous success at SemEval-2025, the Annif team refined their approach, integrating several novel aspects to enhance performance and efficiency. Their efforts paid off, as the Annif system secured the 1st rank in both the overall quantitative and qualitative evaluations for Subtask 2.

A Hybrid Approach to Subject Indexing

The core of Annif’s strategy lies in augmenting traditional Extreme Multi-label Text Classification (XMTC) algorithms with the power of efficient LLMs. The task required recommending the most relevant subjects from the extensive GND subject vocabulary (over 200,000 subjects) for TIBKAT bibliographic records, based on their titles and abstracts.

The key improvements introduced by Annif include:

  • Translation of metadata records using a variety of small and efficient LLMs, carefully balancing throughput and quality.
  • Generation of synthetic training data by combining different LLMs.
  • A sophisticated fusion approach involving optimized weights and exponents for combining different models.
  • Utilizing an LLM specifically for ranking candidate subjects within an ensemble.

This hybrid methodology demonstrated that combining LLMs with established natural language processing and machine learning techniques can be highly competitive, even against systems that rely more heavily on LLMs, while also offering flexibility in balancing efficiency and quality.

System Components and Efficiency Focus

The Annif system is built on three main backends:

  • Omikuji (Bonsai-style configuration): An implementation of efficient machine learning algorithms for multilabel classification based on partitioned label trees.
  • MLLM (Maui-like Lexical Matching): A lexical algorithm that matches words and expressions in document text to terms in a subject vocabulary.
  • XTransformer: An XMTC and ranking algorithm based on fine-tuned BERT-style Transformer models.

Given the task’s emphasis on energy and compute efficiency, the team primarily focused on smaller, open-weight LLMs, ranging from 3 billion to 12 billion parameters, with support for both English and German. These included models like Aya 8B, Gemma 3 4B, and Qwen 3 4B, which were found to offer an excellent balance of quality and efficiency. Larger models (24B to 30B) were also included for comparison.

Experimental Setup and Key Findings

The experimental process involved several stages:

Data Translation: The GND vocabulary was translated into English using GPT-4o-mini. For TIBKAT records, various LLMs were tested for translation into German-only and English-only variants. Gemma 3 4B (G4) was selected for English and Aya 8B (A8) for German, chosen for their optimal compromise between quality and efficiency, rather than absolute translation quality.

Synthetic Training Data: To address the relatively small number of original training records, additional synthetic data was generated. Top-ranking LLMs from the translation task created new titles and abstracts in a one-shot approach, based on existing records and their subject labels, with an added random subject term. This significantly improved the performance of the Bonsai models and, to a lesser extent, XTransformer.

Ensemble Projects: The base backends were combined into two types of ensembles: simple ensembles (merging suggestions by averaging scores) and a new LLM ranking ensemble. The LLM ranking ensemble further refined predictions by using an external LLM to score the relevance of candidate subjects. Hyperparameter optimization was used to fine-tune the weights and exponents of these ensembles.

Also Read:

Outstanding Results

The Annif system’s LLM ranking ensemble, specifically using the M24 LLM, achieved the top rank in the quantitative evaluation with an nDCG@20 score of 0.5697. Furthermore, in the qualitative evaluation conducted by subject librarians, Annif also ranked first in both assessment cases, demonstrating its ability to produce highly relevant subject predictions.

The research highlights that efficient local LLMs like Aya 8B, Gemma 3 4B, and Qwen 3 4B provide a strong balance of quality and efficiency. The use of LLM-generated synthetic training data proved beneficial for traditional XMTC models, particularly Bonsai. The novel LLM ranking ensemble consistently improved nDCG scores, offering different compromises between quality and computational efficiency depending on the LLM used.

For more detailed information, the full research paper can be accessed here: Annif at the GermEval-2025 LLMs4Subjects Task: Traditional XMTC Augmented by Efficient LLMs.

Meera Iyer
Meera Iyerhttps://blogs.edgentiq.com
Meera Iyer is an AI news editor who blends journalistic rigor with storytelling elegance. Formerly a content strategist in a leading tech firm, Meera now tracks the pulse of India's Generative AI scene, from policy updates to academic breakthroughs. She's particularly focused on bringing nuanced, balanced perspectives to the fast-evolving world of AI-powered tools and media. You can reach her out at: [email protected]

- Advertisement -

spot_img

Gen AI News and Updates

spot_img

- Advertisement -