spot_img
HomeResearch & DevelopmentStarBench: A New Benchmark for AI Agents in Turn-Based...

StarBench: A New Benchmark for AI Agents in Turn-Based RPGs

TLDR: StarBench is a new benchmark using Honkai: Star Rail to evaluate vision-language models (VLMs) in two key areas: multimodal decision-making from raw pixels to actions, and agentic information seeking. It features direct control (raw screenshots to low-level actions) and tool-assisted control (higher-level actions with optional textual UI hints), plus an “ask-or-act” diagnostic. Results show VLMs fail in direct control but perform competently with tool assistance, especially with OCR, and benefit from judicious information seeking, highlighting a gap in pixel-to-primitive grounding.

A new research paper introduces StarBench, a groundbreaking benchmark designed to evaluate how well vision-language models (VLMs) can play complex video games like humans. The benchmark, derived from the popular turn-based RPG Honkai: Star Rail, focuses on two key human-like abilities: making decisions based on what’s seen on screen (multimodal decision-making) and knowing when to seek information when stuck (agentic information seeking).

Traditional game agents often rely on simplified interfaces or direct access to game data, which doesn’t fully capture the challenges human players face. Humans must interpret raw visuals, understand user interface elements, and combine on-screen information with their knowledge to perform precise actions. StarBench aims to bridge this gap by testing VLMs in a realistic game client.

Two Ways to Play: Direct vs. Tool-Assisted Control

StarBench evaluates agents under two distinct control regimes, using the same game client, tasks, and metrics to ensure a fair comparison. The first is Direct Control (DC), where the VLM receives only raw screenshots and must output low-level actions like pixel coordinates for clicks and keypresses. This regime tests the model’s ability to parse the UI and localize actions without any semantic hints.

The second is Tool-Assisted (TA) Control. Here, the VLM still sees the screenshot but can express actions using higher-level commands, such as selecting a character, a move type (Basic Attack, Skill, Ultimate), and a target. This mode also provides optional textualized observations from detectors and OCR (Optical Character Recognition) to help the model understand UI elements like HP percentages, skill points, and status effects. This allows the VLM to focus more on strategic decision-making rather than the intricate details of UI manipulation.

The “Ask-or-Act” Challenge: When to Seek Help?

Beyond just playing, StarBench also includes an “ask-or-act” diagnostic. This feature measures whether and when an agent chooses to request brief guidance before proceeding with a battle. Similar to how humans might look up a strategy guide, agents can ask a targeted question to obtain a short textual hint. This diagnostic helps researchers understand if VLMs can judiciously seek information and how that guidance impacts their subsequent performance.

Key Findings: A Gap in Human-like Play

The research reveals significant insights into the current capabilities of VLMs. In the Direct Control regime, contemporary VLMs like GPT-4o-mini, Claude 3.5 Sonnet, and Gemini 1.5 Flash largely failed, achieving 0% success rates across many tasks. This highlights a fundamental challenge in their ability to accurately interpret raw pixels and translate them into precise, low-level keyboard-mouse actions.

However, when given minimal UI grounding through the Tool-Assisted interface, these same VLMs showed substantial improvements. For example, GPT-4o-mini achieved 100% success rates on several Echo of War tasks in TA mode. This suggests that while VLMs struggle with raw pixel-to-primitive control, they can become competent players when the burden of UI manipulation is abstracted away.

The study also found that OCR, which provides textualized UI information, significantly boosts performance even in the tool-assisted mode. Without OCR, success rates dropped, and battles took more steps, indicating that textual cues are crucial for reducing decision friction and ensuring timely, legal actions.

Furthermore, the “ask-or-act” diagnostic demonstrated that seeking brief guidance can be beneficial. VLMs that asked questions, especially GPT-4o-mini, showed measurable uplifts in performance. This suggests that calibrated information seeking is a valuable competency for agentic play.

Also Read:

Conclusion

StarBench establishes a reproducible benchmark for evaluating multimodal agents in a real game client. The findings underscore that while current VLMs struggle with direct pixel-to-primitive control, lightweight tools and strategic information seeking are vital for achieving human-like performance in complex game environments. The paper, titled “StarBench: A Turn-Based RPG Benchmark for Agentic Multimodal Decision-Making and Information Seeking,” can be found here.

Karthik Mehta
Karthik Mehtahttps://blogs.edgentiq.com
Karthik Mehta is a data journalist known for his data-rich, insightful coverage of AI news and developments. Armed with a degree in Data Science from IIT Bombay and years of newsroom experience, Karthik merges storytelling with metrics to surface deeper narratives in AI-related events. His writing cuts through hype, revealing the real-world impact of Generative AI on industries, policy, and society. You can reach him out at: [email protected]

- Advertisement -

spot_img

Gen AI News and Updates

spot_img

- Advertisement -