spot_img
HomeNews & Current EventsOpenAI's O3 Model Surpasses Newer GPT-5 in Complex Multi-Application...

OpenAI’s O3 Model Surpasses Newer GPT-5 in Complex Multi-Application Office Workflows

TLDR: A recent benchmark, OdysseyBench, reveals that OpenAI’s established o3 model consistently outperforms the newer GPT-5 model in complex, multi-application office tasks, challenging the notion of continuous linear improvement in AI performance across all domains.

In a surprising turn of events for the artificial intelligence community, OpenAI’s o3 model has demonstrated superior performance over its successor, the recently launched GPT-5, when tackling intricate, multi-application office tasks. This revelation comes from a new benchmark called OdysseyBench, designed to simulate realistic, multi-day office workflows.

Developed by researchers at Microsoft and the University of Edinburgh, OdysseyBench aims to move beyond isolated ‘atomic tasks’ to evaluate how AI models handle scenarios that unfold over several days, requiring long-term context and coordination across various applications like Word, Excel, email, and calendar. The benchmark’s focus on real-world office environments provides a more comprehensive assessment of AI agents’ practical utility.

On OdysseyBench-Neo, which features the most demanding, hand-crafted tasks, the o3 model achieved a notable 61.26% success rate. In contrast, GPT-5 recorded a 55.96% success rate, and GPT-5-chat trailed slightly behind at 57.62%. The performance gap became even more pronounced in tasks necessitating the simultaneous use of three applications, where o3 scored 59.06%, while GPT-5 managed only 53.80%. Similar trends were observed on OdysseyBench+, with o3 scoring 56.2%, surpassing GPT-5 at 54.0% and GPT-5-chat at 40.3%. These results highlight o3’s robust reasoning capabilities, particularly in scenarios demanding intricate planning and context management across multiple software environments.

While GPT-5 has shown significant advancements in other areas, such as coding challenges (e.g., 74.9% on SWE-Bench for bug fixes, 88% on Aider Polyglot coding test) and multimodal visual reasoning, its performance on these specific multi-app office tasks suggests that progress isn’t uniformly distributed across all AI capabilities. The findings also imply that while both o3 and GPT-5 represent improvements over older models, the leap from o3 to GPT-5 isn’t as significant in this particular domain, especially considering o3 was officially released only in April.

Also Read:

OpenAI has been actively refining its GPT-5 offerings, including introducing a more natural voice mode, personalization options, and memory integration with services like Gmail and Google Calendar. For developers, GPT-5 offers custom tool calls and verbosity controls. The company has also been adjusting GPT-5’s tone to be warmer based on user feedback and has introduced various model selection options within ChatGPT, including ‘Auto,’ ‘Fast,’ ‘Thinking mini,’ ‘Thinking,’ and ‘Pro,’ alongside legacy models like o3 and o4-mini. Despite these broader enhancements, the OdysseyBench results underscore the specialized strengths of the o3 model in complex office automation.

Karthik Mehta
Karthik Mehtahttps://blogs.edgentiq.com
Karthik Mehta is a data journalist known for his data-rich, insightful coverage of AI news and developments. Armed with a degree in Data Science from IIT Bombay and years of newsroom experience, Karthik merges storytelling with metrics to surface deeper narratives in AI-related events. His writing cuts through hype, revealing the real-world impact of Generative AI on industries, policy, and society. You can reach him out at: [email protected]

- Advertisement -

spot_img

Gen AI News and Updates

spot_img

- Advertisement -