spot_img
HomeResearch & DevelopmentCLIMB: A New Framework for Data-Driven Occupation Taxonomies

CLIMB: A New Framework for Data-Driven Occupation Taxonomies

TLDR: CLIMB is a novel framework that automates the creation of high-quality, data-driven occupation taxonomies directly from raw job postings. It uses global semantic clustering to distill core occupations and a reflection-based multi-agent system to iteratively build a coherent hierarchy. The framework outperforms existing methods in coherence, scalability, and efficiency, and uniquely captures specific regional labor market characteristics, making it highly adaptive.

Understanding and organizing the vast and ever-changing landscape of job roles is crucial for everything from helping job seekers find their ideal positions to enabling governments to analyze labor trends. Traditionally, creating these ‘occupation taxonomies’ – hierarchical structures that categorize jobs – has been a slow, manual, and often outdated process. Existing automated methods either struggle to adapt to dynamic regional markets or fail to build coherent structures from messy, real-world job data.

A new framework called CLIMB (CLusterIng-based Multi-agent taxonomy Builder) aims to change this. Developed by Nan Li, Bo Kang, and Tijl De Bie from Ghent University, CLIMB fully automates the creation of high-quality, data-driven occupation taxonomies directly from raw job postings. Unlike previous approaches, CLIMB builds these hierarchies from the ground up, ensuring they are tailored to specific markets and can be easily updated.

The core challenge CLIMB addresses lies in two main areas: first, how to extract a consistent set of core job concepts from a massive and often noisy collection of job advertisements; and second, how to then arrange these concepts into a deep and logically sound hierarchy. CLIMB tackles these problems through a clever multi-stage pipeline.

How CLIMB Works: A Multi-Stage Journey

The CLIMB framework operates in several key stages, each designed to refine the raw data into a structured taxonomy:

First, CLIMB begins with **Job Posting Distillation and Embedding**. Job postings often contain a lot of irrelevant text alongside core occupational details. To make processing efficient, CLIMB uses a smart approach: a small sample of text is labeled by an advanced AI model (acting as an expert) to identify relevant job-related content. This labeled data then trains a lightweight classifier, which can quickly filter out irrelevant text from all job postings. The cleaned descriptions are then converted into numerical representations (embeddings) that computers can understand.

Next is **Semantic Clustering**. Simply grouping jobs by how similar their text embeddings are isn’t enough to capture the nuanced human judgment of what constitutes the ‘same occupation’. CLIMB overcomes this by training a specialized similarity model. It uses an AI model as an ‘HR expert’ to label pairs of job descriptions as either ‘same occupation’ or ‘different occupation’. This expert-labeled data then trains a machine learning classifier (XGBoost) to learn a more sophisticated understanding of job similarity. Finally, an algorithm called Affinity Propagation uses these learned similarities to group job postings into fine-grained occupation clusters, which form the initial ‘leaf nodes’ of the taxonomy.

These raw clusters are then transformed in the **Leaf Node Generation** stage. Since raw cluster titles can be inconsistent, CLIMB uses an AI model to generate concise, canonical titles and descriptions for each occupation cluster. This process abstracts a clear concept from noisy raw text. Further normalization and deduplication steps ensure that each distinct occupation is represented by a single, well-defined node in the taxonomy.

The final and most complex stage is **Hierarchical Taxonomy Construction**. With a solid set of leaf nodes, CLIMB builds the hierarchy upwards, level by level. This is done using a reflection-based multi-agent framework. A ‘Generator’ AI proposes parent concepts to group the current level’s nodes, performing specific-to-general reasoning. An ‘Evaluator’ AI then scrutinizes the Generator’s output for logical consistency, checking for missing nodes, duplicate assignments, or illogical mappings. If flaws are found, the Evaluator provides feedback, and the Generator refines its output. This iterative ‘generate-evaluate’ cycle continues until a coherent structure is validated, ensuring the taxonomy is logically sound at every level.

Also Read:

Demonstrated Impact and Regional Adaptability

The researchers tested CLIMB on three diverse, real-world datasets from Palestine, Botswana, and the USA. The results showed that CLIMB produces taxonomies that are significantly more consistent, unambiguous, and scalable than existing automated methods. It also strikes a superior balance between comprehensiveness (covering most jobs) and efficiency (using relevant labels).

One of CLIMB’s most compelling advantages is its ability to capture unique regional characteristics. For instance, the taxonomy generated for Palestine highlighted the prominence of humanitarian work, with specialized roles like ‘MEAL Specialist’ and ‘WASH Specialist’. In Botswana, the taxonomy reflected key economic drivers, detailing roles in the diamond industry (e.g., ‘Diamond Processing, Grading, and Valuation’) and the booming tourism sector (‘Safari / Tour Guide’). For the USA, CLIMB accurately modeled the distinct structure of the American education system, differentiating between K-12 and higher education sectors. This level of regional specificity is often missing in generic, international standards.

While CLIMB represents a significant leap forward, the authors acknowledge limitations, such as the reliance on AI models as proxy experts and annotators, and the scalability of certain clustering algorithms for extremely large datasets. Nevertheless, CLIMB offers a powerful, automated solution for building adaptive, high-quality occupation taxonomies that can truly reflect the dynamic nature of global and regional labor markets. For more details, you can read the full research paper. Building Data-Driven Occupation Taxonomies: A Bottom-Up Multi-Stage Approach via Semantic Clustering and Multi-Agent Collaboration.

Meera Iyer
Meera Iyerhttps://blogs.edgentiq.com
Meera Iyer is an AI news editor who blends journalistic rigor with storytelling elegance. Formerly a content strategist in a leading tech firm, Meera now tracks the pulse of India's Generative AI scene, from policy updates to academic breakthroughs. She's particularly focused on bringing nuanced, balanced perspectives to the fast-evolving world of AI-powered tools and media. You can reach her out at: [email protected]

- Advertisement -

spot_img

Gen AI News and Updates

spot_img

- Advertisement -