spot_img
HomeResearch & DevelopmentThe Durability of Edited Facts in Language Models After...

The Durability of Edited Facts in Language Models After Fine-Tuning

TLDR: This research paper investigates how fine-tuning affects knowledge previously edited into large language models (LLMs). It finds that edited knowledge is significantly more susceptible to forgetting than intrinsic knowledge acquired during pre-training. While freezing layers associated with edited content can improve retention, current editing methods struggle to make edits as robust as pre-trained knowledge, often leading to the generation of related but incorrect information when edits are forgotten. The study highlights the need for future model editing evaluations to include post-fine-tuning retention.

Large language models (LLMs) are incredibly powerful, storing vast amounts of information. However, this knowledge isn’t static; it often needs updates to correct errors, add new facts, or adjust how the model behaves. Traditional methods like retraining the entire model are computationally expensive and can be inefficient. This is where ‘model editing’ comes in, offering a more precise and cost-effective way to modify specific pieces of knowledge within an LLM.

While model editing has gained traction, a crucial question has remained largely unanswered: what happens to this edited knowledge when the LLM is subsequently fine-tuned for a specific task? Fine-tuning is a common practice to adapt LLMs to various downstream applications, but its interaction with previously edited information has been poorly understood. This research systematically explores how different fine-tuning objectives impact knowledge that has been modified using various model editing techniques.

The study reveals a significant finding: knowledge that has been ‘edited’ into an LLM is considerably more prone to being forgotten during subsequent fine-tuning compared to the knowledge the model acquired during its initial pre-training phase (referred to as ‘intrinsic knowledge’). This highlights a key limitation of current model editing approaches, suggesting that evaluating how well edits hold up under downstream fine-tuning is essential for their real-world deployment.

The researchers investigated three main categories of knowledge editing methods: ‘Brutal-force Fine-tuning’ (a baseline), ‘Locate-then-Edit’ (like ROME and MEMIT), and ‘Meta-learning’ (like MALMEN). They then subjected these edited models to four different fine-tuning tasks: fine-tuning with unstructured text, structured factual text, classification tasks, and supervised fine-tuning (SFT).

A striking observation was that while edited knowledge showed varying retention rates depending on the editing method and fine-tuning task, the intrinsic knowledge remained remarkably stable. Even after the initial model editing, which caused a slight drop in intrinsic knowledge retention, subsequent fine-tuning had little further impact on it. This suggests that the model’s inherent, pre-trained knowledge is far more robust than newly introduced edited knowledge.

Interestingly, the study found that the ‘Locate-then-Edit’ methods (ROME/MEMIT) generally had lower edited knowledge retention rates for unstructured, classification, and SFT tasks. However, for structured tasks, ROME surprisingly showed the highest retention. This might be because ROME conditions the model to process information in a specific format, and when downstream tasks use similar data formats, the edited knowledge is better preserved.

To combat this forgetting, the researchers explored strategies to improve knowledge retention. Their hypothesis was that avoiding fine-tuning the specific layers where the edited knowledge was inserted could help. They tested two layer-specific fine-tuning strategies: freezing early layers and fine-tuning only layers beyond a certain threshold, and freezing all layers except a specific ‘window’ of layers. The results supported their hypothesis: freezing layers associated with the edited content significantly improved the retention of that knowledge. In some cases, if only very late layers were fine-tuned, the edited knowledge retention reached levels comparable to intrinsic knowledge.

Further analysis into how the model’s output changes during fine-tuning provided more insights. When edited knowledge was forgotten, the model didn’t just revert to the original incorrect answer or produce random tokens. Instead, it often generated ‘related-term tokens’—words that belong to the same category as the correct answer, even if they made the sentence factually incorrect. For example, if ‘Apple’ was edited as the developer of ‘Windows Mobile 6.5’ (instead of ‘Microsoft’), after fine-tuning, the model might start predicting ‘Intel’ or ‘Nokia’ instead of ‘Apple’ or ‘Microsoft’. This indicates that the original knowledge was effectively erased, but the model still tries to provide a plausible, albeit incorrect, related answer.

The core reason for the lower retention of edited knowledge, as hypothesized by the researchers, lies in its acquisition. Intrinsic knowledge is learned from diverse expressions in vast, unstructured pre-training data, making it highly robust. Edited knowledge, however, is often enforced on specific layers or through single expression patterns, making it less ‘natural’ and more localized within the model. This creates a trade-off: editing is efficient, but the resulting knowledge is less stable long-term.

While freezing layers offers a solution for preserving single edits, it comes with limitations. It can restrict the model’s ability to learn new information during fine-tuning, decrease training efficiency, and isn’t scalable for scenarios with many edits across multiple layers. Therefore, future research needs to focus on developing more robust editing methods that can withstand fine-tuning without compromising the model’s overall learning capabilities.

Also Read:

This research underscores that for future knowledge editing methods, evaluating the retention of edits should not only be done immediately after editing but also after subsequent fine-tuning. For more details, you can read the full paper here.

Ananya Rao
Ananya Raohttps://blogs.edgentiq.com
Ananya Rao is a tech journalist with a passion for dissecting the fast-moving world of Generative AI. With a background in computer science and a sharp editorial eye, she connects the dots between policy, innovation, and business. Ananya excels in real-time reporting and specializes in uncovering how startups and enterprises in India are navigating the GenAI boom. She brings urgency and clarity to every breaking news piece she writes. You can reach her out at: [email protected]

- Advertisement -

spot_img

Gen AI News and Updates

spot_img

- Advertisement -