spot_img
HomeResearch & DevelopmentEnhancing LLM Safety Across Languages with SEALGUARD

Enhancing LLM Safety Across Languages with SEALGUARD

TLDR: SEALGUARD is a new multilingual guardrail that significantly improves the detection of unsafe and jailbreak prompts in Large Language Models (LLMs) across Southeast Asian languages. Unlike existing English-centric guardrails that struggle with multilingual inputs, SEALGUARD, built using LoRA adaptation on a multilingual LLM and evaluated on the new SEALSBENCH dataset, achieves superior performance in defense success rate, precision, and F1-score, demonstrating robust safety alignment for diverse linguistic contexts.

Large Language Models (LLMs) are becoming increasingly common in real-world applications, especially in multilingual settings like language learning tutors. While these systems are powerful, they face a significant challenge: ensuring safety across diverse languages. Current safety measures, often called “guardrails,” are primarily designed and trained on English content. This leaves LLM systems vulnerable to harmful or “jailbreak” prompts written in other languages, particularly those with fewer digital resources, such as Southeast Asian languages.

Imagine a scenario where an unsafe prompt in English is successfully blocked by a guardrail, preventing it from reaching the LLM. However, if the same prompt is translated into a language like Lao, it might bypass the guardrail entirely, potentially leading the LLM to generate harmful responses. This highlights a critical gap in multilingual safety alignment for LLM-powered software systems.

To address this pressing issue, researchers have introduced a new multilingual guardrail called SEALGUARD. This innovative system aims to significantly improve safety alignment across various languages, effectively filtering out unsafe and jailbreak prompts in LLM-powered applications. The core idea behind SEALGUARD is to adapt a general-purpose multilingual language model, specifically SeaLLM, into a specialized safety guardrail. This adaptation is achieved using a technique called Low-Rank Adaptation (LoRA), which efficiently fine-tunes the model without altering its entire parameter set, thus preserving its original multilingual capabilities while adding guardrail-specific knowledge.

To rigorously test SEALGUARD’s effectiveness, the researchers developed SEALSBENCH, a comprehensive, large-scale multilingual safety alignment dataset. This benchmark contains over 260,000 prompts in ten languages, including English, Chinese, Indonesian, Vietnamese, Thai, Khmer, Lao, Malay, Burmese, and Tagalog. The dataset includes a mix of safe, unsafe, and jailbreak prompts, allowing for a thorough evaluation of how well guardrails can block harmful inputs while minimizing false alarms on safe ones.

The findings from the evaluation are striking. Existing state-of-the-art guardrails, such as LlamaGuard, showed a substantial drop in performance when confronted with multilingual unsafe and jailbreak prompts. For instance, LlamaGuard’s Defense Success Rate (DSR) decreased by 9% for multilingual unsafe prompts and a significant 18% for multilingual jailbreak prompts, compared to its performance on English-only inputs. OpenAI Moderation also experienced considerable declines in DSR, dropping by 31% and 38% respectively in similar scenarios.

In stark contrast, SEALGUARD demonstrated superior performance. It consistently outperformed existing guardrails in detecting multilingual unsafe and jailbreak prompts. SEALGUARD achieved an impressive F1-score of 98%, which is 34% to 58% higher than the baselines. Its Defense Success Rate (DSR) reached 97%, representing an improvement of 48% to 66% over LlamaGuard and OpenAI Moderation. Furthermore, SEALGUARD maintained a high precision of 99%, indicating its effectiveness in accurately identifying unsafe content while keeping false alarms to a minimum.

An in-depth analysis, known as an ablation study, revealed that the LoRA adaptation strategy is the most crucial component contributing to SEALGUARD’s success, accounting for a 72% boost in its F1-Score. Interestingly, the study also found that the model’s size had only a minor impact on performance, suggesting that even smaller models can achieve promising results with this adaptation approach.

Also Read:

The development of SEALGUARD marks a significant step forward in ensuring the safety of LLM systems in a globally connected, multilingual world. Alongside the innovative guardrail itself, the introduction of the SEALSBENCH benchmark provides a valuable resource for future research in multilingual safety alignment. The researchers have also committed to open science by releasing the pre-trained model, the benchmark dataset, and all experimental scripts, fostering further advancements in the field. You can find more details about this research in the full paper available here.

Karthik Mehta
Karthik Mehtahttps://blogs.edgentiq.com
Karthik Mehta is a data journalist known for his data-rich, insightful coverage of AI news and developments. Armed with a degree in Data Science from IIT Bombay and years of newsroom experience, Karthik merges storytelling with metrics to surface deeper narratives in AI-related events. His writing cuts through hype, revealing the real-world impact of Generative AI on industries, policy, and society. You can reach him out at: [email protected]

- Advertisement -

spot_img

Gen AI News and Updates

spot_img

- Advertisement -