TLDR: IBM Research, in collaboration with the University of Washington, has released ‘Toucan,’ the largest and most comprehensive dataset of real-life, end-to-end tool-calling scenarios. This dataset, featuring 1.5 million task sequences across 2,000 web services, is designed to significantly enhance the ability of AI agents to find, deploy, and execute web-based applications, with early results showing small models fine-tuned on Toucan outperforming much larger frontier models.
IBM Research, in a significant stride for artificial intelligence, has announced the release of ‘Toucan,’ a monumental dataset poised to transform the landscape of AI agent tool-calling. Developed in partnership with the University of Washington (UW), Toucan represents the largest and most comprehensive collection of publicly available, real-life, end-to-end tool-calling scenarios to date, as reported on October 17, 2025.
Tool-calling is a fundamental capability for AI agents, enabling large language models (LLMs) to interact with and utilize web applications beyond their core conversational functions. Historically, the challenge has been the scarcity of high-quality, diverse examples for training LLMs in this complex skill. Toucan addresses this by providing 1.5 million real-life tool-calling task sequences, referred to as ‘trajectories,’ which collectively invoke 2,000 distinct web services.
The scenarios within Toucan are remarkably diverse, encompassing a wide array of complex tasks. These range from intricate business operations like analyzing sales reports and drafting summaries to everyday scheduling activities such as arranging meetings and sending calendar invitations.
Upon its release on Hugging Face, Toucan immediately captured the attention of the AI community, quickly becoming a top trending dataset. An enthusiastic observer noted on LinkedIn, ‘Toucan changes everything. This isn’t another simulated dataset. It captures actual API executions in real environments. Complete interaction chains from start to finish.’
A new pre-print study by the IBM and UW team highlights Toucan’s profound impact. It demonstrates that smaller, open-source models, when fine-tuned using the Toucan dataset, can surpass the performance of frontier models that are many times larger on two prominent benchmarks for agentic tool-use: Berkeley Function Calling Leaderboard version 3 (BFCLv3) and MCP-Universe.
Rameswar Panda, the IBM researcher who spearheaded the Toucan project, emphasized the dataset’s importance: ‘Tool-calling is central to AI agents. How can you train better agents? Through diverse, high-quality examples sourced from the real world.’
Toucan’s design specifically targets teaching AI agents how to effectively call APIs by connecting to MCP servers, which function as topic-based software libraries. The researchers leveraged a fleet of LLMs and MCP server metadata from GitHub and Smithery.ai to curate this extensive and challenging dataset. A notable feature is that approximately a fifth of Toucan’s scenarios necessitate models to call multiple tools concurrently, a design choice aimed at fostering more economical AI agent operation. Zhangchen Xu, a University of Washington graduate student and IBM intern who contributed to the dataset, explained, ‘You can imagine how parallel calling improves efficiency, which can lower the cost of running agentic systems.’
Also Read:
- IBM Unveils Three New AI Agents on Oracle Fusion Applications AI Agent Marketplace, with Further Expansion Planned
- Agentic Context Engineering (ACE) Revolutionizes AI Agent Performance with Evolving Playbooks
This release underscores IBM Research’s ongoing commitment to advancing AI capabilities, particularly in the realm of agentic AI, as highlighted in their broader research initiatives focused on quantum computing and AI.


