TLDR: BLM1 is a new AI model that unifies digital-space reasoning with physical-space robotic control. It achieves superior performance by enabling cross-space transfer, cross-task learning, and cross-embodiment generalization through a two-stage training process. This allows it to understand and reason in digital environments while also providing robust, low-level control for various robot types across different manipulation tasks, outperforming existing multimodal and embodied AI models.
In the rapidly evolving landscape of artificial intelligence, a significant challenge has been the creation of models that can seamlessly operate across both digital and physical environments, handle a variety of tasks, and adapt to different robotic forms. Current AI models, such as Multimodal Large Language Models (MLLMs), Embodied Large Language Models (ELLMs), and Vision-Language-Action (VLA) models, often struggle with these aspects, showing limitations in generalizing their knowledge and control capabilities.
Addressing these critical gaps, researchers have introduced a groundbreaking new model called BLM1, or the Boundless Large Model. This innovative multimodal spatial foundation model is designed to unify digital-space reasoning with physical-space robotic control, offering a comprehensive solution for embodied intelligence. BLM1 is the first of its kind to achieve robust cross-space transfer, cross-task learning, and cross-embodiment generalization within a single, cohesive framework.
BLM1’s Core Capabilities
BLM1 is built upon three fundamental capabilities:
- Cross-space transfer: This allows the model to apply knowledge gained in digital domains to real-world physical tasks, enabling advanced embodied perception, spatial reasoning, and direct robotic control.
- Cross-task learning: BLM1 promotes a deep understanding of semantic relationships across different tasks. For instance, its ability to answer questions about object relations and causal structures significantly enhances its planning and execution for complex, long-duration control tasks.
- Cross-embodiment generalization: The model can align the underlying behavior patterns of various robot designs when performing similar tasks, leading to a shared, adaptable policy that works across diverse robotic platforms.
How BLM1 Works: A Two-Stage Training Approach
The effectiveness of BLM1 stems from its unique two-stage training paradigm:
- Stage I: Injecting Embodied Knowledge: In this initial phase, the model undergoes supervised fine-tuning using extensive digital corpora. These datasets are rich in multimodal question-answer pairs that cover embodied perception and reasoning. This stage is crucial for endowing the MLLM backbone with a strong understanding of the physical world while carefully preserving its native language comprehension and instruction-following abilities.
- Stage II: Cross-Embodiment Learning: With the MLLM backbone now equipped with embodied knowledge, Stage II focuses on training a policy module for physical control. This stage introduces an ‘intent-bridging interface’ that extracts high-level semantic information from the MLLM to guide low-level robotic actions. Crucially, the MLLM backbone itself is frozen during this stage, preventing degradation of its general reasoning skills. The policy module, a Diffusion Transformer, is trained on a specially curated cross-embodiment demonstration suite featuring four different robot types (Franka Emika Panda, xArm-6, xArm-7, and WidowX AI) performing six progressively challenging tasks.
Also Read:
- HyPerNav: A New Approach for Robots to Find Objects Using Hybrid Perception
- RobotArena∞: A New Framework for Scalable Robot Evaluation
Impressive Performance Across the Board
Evaluations have shown that BLM1 sets new benchmarks in both digital and physical domains. In digital tasks, it achieves approximately 6% gains over existing MLLMs, ELLMs, VLAs, and General Multimodal Large Models (GMLMs). For physical tasks, BLM1 demonstrates around 3% improvement, showcasing its superior performance in real-world robotic control.
Notably, BLM1 consistently outperforms leading models like GPT-4o and Cosmos-7B in various digital benchmarks, demonstrating enhanced embodied multimodal reasoning, especially in first-person reasoning, planning, and fine-grained action generation. In physical space, BLM1 achieves an average success rate of 75.83%, surpassing all from-scratch policies and even outperforming many pre-trained VLAs. It maintains consistent high performance across diverse robot embodiments and tasks, from simple object picking to complex stacking and upright placement.
The ability of BLM1 to provide accurate reasoning in multiple-choice questions and detailed, non-hallucinatory instructions in free-form question answering further highlights its robust visual understanding and action planning capabilities. For more in-depth technical details, you can refer to the research paper.
By successfully integrating cross-space transfer, cross-task adaptation, and cross-embodiment control within a single model, BLM1 represents a significant leap forward, paving a scalable path toward truly general-purpose embodied intelligence.


