spot_img
HomeResearch & DevelopmentTeaching Language Models to Speak More Efficiently: The Art...

Teaching Language Models to Speak More Efficiently: The Art of Convention Formation

TLDR: Researchers developed a post-training method to teach large language models (LLMs) to communicate more efficiently by forming ‘ad-hoc conventions,’ similar to how humans shorten phrases over repeated interactions. Using fine-tuning on human conversation data and special ‘planning tokens,’ they significantly improved LLMs’ ability to use concise and consistent language in new text-based communication tasks, bridging a key gap between AI and human conversational efficiency.

In the evolving landscape of artificial intelligence, large language models (LLMs) have demonstrated remarkable capabilities in understanding and generating human-like text. However, a recent study highlights a crucial area where these advanced models still fall short compared to human communication: the ability to form ad-hoc conventions for efficient interaction. Humans naturally adapt their language in multi-turn conversations, often shortening phrases and forming shared understandings to communicate more effectively. This behavior, known as convention formation, is a cornerstone of natural human dialogue, improving both accuracy and efficiency.

A research paper titled “Post-training for Efficient Communication via Convention Formation” by Yilun Hua, Evan Wang, and Yoav Artzi from Cornell University addresses this very challenge. Their work introduces a novel post-training process designed to imbue LLMs with this essential human communication trait. The core idea is to fine-tune models using carefully selected examples of convention formation observed in human interactions.

The Problem: LLMs Don’t Adapt Like Us

Previous research has indicated that current LLMs do not spontaneously exhibit the tendency to adapt their language for increased efficiency over repeated interactions. They often maintain verbose descriptions even when a shorter, mutually understood phrase would suffice. This lack of adaptive communication can make interactions with LLMs feel unnatural and less efficient, especially in multi-turn scenarios.

The Solution: A Targeted Post-Training Approach

The Cornell researchers propose a targeted post-training method that involves several innovative steps. First, they construct unique “preference data” by identifying instances of convention formation in human conversations, specifically from TV series scripts. They create “minimal pairs” of examples: a preferred continuation (showing convention formation) and a dispreferred one (lacking it). For instance, if a character initially describes “the large, red book with a dragon on the cover,” a preferred re-mention might be “the dragon book,” while a dispreferred one would repeat the full description.

To further enhance the model’s reasoning, they introduce a special “planning token” – [remention] – which explicitly marks when a concept is being re-mentioned. This helps the LLM differentiate between initial mentions and subsequent, potentially more concise, references. The training process itself is a two-stage approach: an initial supervised fine-tuning (SFT) phase helps the model learn to use the planning token correctly, followed by a preference optimization stage that encourages the model to generate the more concise, convention-forming responses.

Evaluating the Improvement: New Benchmarks

To rigorously test their method, the researchers developed two new evaluation benchmarks, distinct from their training data, to assess the generalizability of the learned ability:

  • Text-only Reference Game: Inspired by cognitive science studies, this game involves an LLM speaker describing a target item from a set of text-based referents to a listener (another LLM, GPT4o-mini). Over multiple rounds, the same items reappear, allowing the researchers to observe if the speaker’s descriptions become shorter and more consistent, mirroring human behavior.
  • Document-grounded Utterance Completion: This task is closer to real-world LLM applications. The model completes an agent’s response in a dialogue, based on a provided document and conversation history. The goal is to see if the model can concisely re-mention concepts previously discussed, similar to how a human assistant might.

Also Read:

Promising Results and Future Directions

The study’s findings are compelling. Off-the-shelf LLMs consistently failed to form conventions, often increasing message lengths or introducing new words for repeated items. However, the post-trained models, specifically Gemma and Llama, showed significant improvements. They shortened their messages by up to 26% in the reference games and maintained much greater consistency. In the document-grounded task, the post-trained models also substantially outperformed their original versions.

Crucially, the researchers found that all components of their post-training method – including the planning tokens and the specific training stages – were necessary for these improvements. Simple prompting or few-shot examples alone were insufficient to elicit this complex adaptive behavior. Furthermore, the post-training had a minimal impact on the models’ general capabilities, indicating that this specialized skill can be added without compromising broader performance.

While the post-trained LLMs showed remarkable progress, they still haven’t fully closed the gap with human communication efficiency. Humans continue to demonstrate even greater levels of conciseness and consistency. This suggests exciting avenues for future research, perhaps by integrating this convention formation training more deeply into the overall LLM development process or by exploring how models can explicitly reason about the costs and benefits of communication, much like humans do.

This research marks a significant step towards making LLMs more natural and efficient communicators, paving the way for more fluid and human-like interactions. For more details, you can read the full research paper here.

Meera Iyer
Meera Iyerhttps://blogs.edgentiq.com
Meera Iyer is an AI news editor who blends journalistic rigor with storytelling elegance. Formerly a content strategist in a leading tech firm, Meera now tracks the pulse of India's Generative AI scene, from policy updates to academic breakthroughs. She's particularly focused on bringing nuanced, balanced perspectives to the fast-evolving world of AI-powered tools and media. You can reach her out at: [email protected]

- Advertisement -

spot_img

Gen AI News and Updates

spot_img

- Advertisement -