spot_img
HomeResearch & DevelopmentUnveiling the Hidden Costs of Fine-Tuning: How Too Much...

Unveiling the Hidden Costs of Fine-Tuning: How Too Much Data Can Harm LLM Knowledge

TLDR: A new study reveals that supervised fine-tuning (SFT) can unexpectedly degrade large language model (LLM) performance, with models fine-tuned on more data (1,920 samples) performing up to 14% worse than those with less (240 samples). The research, using LLaMA-2 and LLaMA-3 families, found that up to 90% of parameter updates during SFT do not enhance knowledge and can be detrimental. Restoring these unnecessary updates significantly improves performance, challenging current fine-tuning practices and suggesting a need for more efficient strategies based on data scale and knowledge mastery.

Large language models (LLMs) are at the forefront of artificial intelligence, capable of understanding and generating human-like text. These models gain a vast amount of world knowledge during their initial pre-training phase. However, this knowledge is further refined and shaped by post-training techniques, with supervised fine-tuning (SFT) being a common method. Despite its widespread use, the precise impact of SFT on a model’s existing knowledge has remained a relatively underexplored area.

Recent research has shed light on this critical aspect, revealing some surprising and counter-intuitive findings. A study evaluating five LLMs from the LLaMA-2 and LLaMA-3 families on a closed-book question answering (CBQA) task uncovered that more fine-tuning data doesn’t always lead to better performance. In fact, models fine-tuned on a larger dataset of 1,920 samples performed up to 14% worse than those trained on a smaller set of just 240 samples. This suggests that there’s an optimal amount of data for fine-tuning, beyond which performance can actually decline.

Furthermore, the study found that the ‘mastery level’ of the knowledge within the fine-tuning data significantly impacts performance. Depending on the category of data used for fine-tuning, performance fluctuations of over 12% were observed. This indicates that not all data is equally beneficial, and the model’s prior familiarity with the information plays a crucial role.

To understand these effects, the researchers conducted a two-level analysis: token-level and parameter-level. At the token level, they looked at how fine-tuning changes the model’s predicted token distributions compared to the pre-trained model. They used KL divergence to quantify these differences. The analysis showed that as the fine-tuning data size increases, the divergence initially decreases, indicating better stability. However, past a certain point, the divergence sharply rises, especially with poorly mastered data, which correlates with a drop in performance. This suggests that excessive shifts in the model’s internal knowledge can be detrimental, leading to what’s known as ‘catastrophic forgetting’ where previously learned knowledge is lost.

The parameter-level analysis provided even more striking insights. By selectively restoring parameters that changed the most during SFT back to their original pre-trained values, the researchers discovered that a significant portion of these updates were unnecessary. Up to 90% of parameter updates during SFT were found not to contribute to knowledge enhancement, and in many cases, they actually degraded performance. Restoring these ‘redundant’ parameters could improve performance by over 10% in some scenarios. This implies that SFT often introduces changes that don’t help the model learn new information or generalize better, and can even harm its existing knowledge.

The study also highlighted that models fine-tuned with larger datasets or data they had low mastery over were more negatively affected by these unnecessary parameter changes. For instance, models trained with 1,920 samples continued to show performance improvements even after restoring 40% of their parameters, suggesting a higher proportion of unnecessary updates compared to models trained with 240 samples. Similarly, fine-tuning with low-mastery data allowed for more parameter restoration and greater performance gains.

Also Read:

These findings challenge conventional wisdom in fine-tuning LLMs, suggesting that simply adding more data or performing SFT without careful consideration of data quality and quantity can be counterproductive. The research offers practical guidance for developing more effective fine-tuning strategies that focus on minimizing unnecessary updates and preserving valuable pre-trained knowledge. While the study focused on LLaMA models, preliminary validations suggest these conclusions may generalize to other model families. For more details, you can refer to the full research paper: Analyzing the Effects of Supervised Fine-Tuning on Model Knowledge from Token and Parameter Levels.

Nikhil Patel
Nikhil Patelhttps://blogs.edgentiq.com
Nikhil Patel is a tech analyst and AI news reporter who brings a practitioner's perspective to every article. With prior experience working at an AI startup, he decodes the business mechanics behind product innovations, funding trends, and partnerships in the GenAI space. Nikhil's insights are sharp, forward-looking, and trusted by insiders and newcomers alike. You can reach him out at: [email protected]

- Advertisement -

spot_img

Gen AI News and Updates

spot_img

- Advertisement -