spot_img
HomeResearch & DevelopmentAdvancing Autonomous Catheter Navigation with Multimodal AI

Advancing Autonomous Catheter Navigation with Multimodal AI

TLDR: A new AI model, DINO-CV A, is introduced for autonomous cardiac catheter navigation. It learns from expert demonstrations by combining visual information from a camera and joystick movements, allowing it to understand both the anatomy and control actions. Tested on a robotic setup with a synthetic vascular phantom, the model accurately predicts catheter movements towards a specified goal, demonstrating a significant step towards reducing manual operation and improving safety in cardiac procedures.

Cardiac catheterization is a vital minimally invasive procedure, but it still heavily relies on manual control by physicians. This manual dependency leads to operator fatigue, increased radiation exposure for both patients and staff, and variations in procedural outcomes. While robotic systems exist, they often act as ‘follow-leader’ tools, meaning they still require constant input from the physician and lack true intelligent autonomy.

Addressing these challenges, researchers Pedram Fekri, Majid Roshanfar, Samuel Barbeau, Seyedfarzad Famouri, Thomas Looi, Dale Podolsky, Mehrdad Zadeh, and Javad Dargahi have introduced DINO-CV A. This innovative multimodal, goal-conditioned behavior cloning framework aims to bring autonomous navigation to cardiac catheterization. The core idea behind DINO-CV A is to fuse visual observations (what the camera sees) with joystick kinematics (the robot’s movement data) into a combined understanding. This allows the system to develop policies that are aware of both the visual environment and the physical actions needed to navigate.

The model learns by observing expert demonstrations, predicting actions autoregressively, and using ‘goal conditioning’ to guide the catheter towards specific destinations. This means the AI is shown where the catheter needs to go and learns the best way to get there based on how human experts perform the task.

How DINO-CV A Works

DINO-CV A functions as a vision-action model. It takes a sequence of video frames from the catheterization task along with the corresponding joystick states (translation, rotation, and knob movements) of the robotic platform. These video frames are processed using a pre-trained vision transformer called DINOv2, which helps the model understand the visual context. Simultaneously, the joystick states are converted into a format that can be combined with the visual information.

A crucial component is the ‘causal transformer,’ which processes these fused visual and kinematic inputs over time, understanding both spatial (where things are) and temporal (how things change over time) dependencies. This allows the model to learn a navigation policy that is deeply grounded in the anatomical environment. Unlike previous methods that might rely on external object detection or pre-planned trajectories, DINO-CV A learns the navigation directly from the combined visual and action data.

The ‘goal-conditioning’ mechanism is another key feature. The model is provided with an image representing the target location within the vascular system. This target image helps the model plan its trajectory by projecting the learned behaviors towards the desired surgical goal, making the navigation adaptive and generalizable.

Experimental Setup and Data Collection

To develop and test DINO-CV A, the researchers designed and implemented a custom robotic platform. Since commercial robotic catheterization systems are not easily accessible for research, this setup mimicked a clinical ‘follow-leader’ robot. It included a standard ablation catheter integrated with a robotic actuation unit and a teleoperation interface. An Xbox controller allowed an expert operator to control the catheter’s translation, rotation, and steering. A top-mounted camera provided real-time visual feedback, similar to fluoroscopic monitoring in clinical settings.

A transparent, synthetic vessel phantom, scaled 1:1 to a human anatomical model, was used to simulate the cardiac environment. Expert operators navigated the catheter to nine predefined target points within this phantom, with each task repeated multiple times. During these procedures, synchronized video frames and robot states were recorded, creating a rich dataset for training and evaluating DINO-CV A.

Performance and Insights

The DINO-CV A model was evaluated on its ability to predict joystick states (translation, rotation, and knob movements) towards both partially and entirely unseen destinations. The results showed high accuracy, with low error values and strong R-squared scores across all tasks. Notably, DINO-CV A performed comparably to a kinematics-only baseline model (LSTM-based), which solely relies on the temporal sequence of joystick inputs.

The significant advantage of DINO-CV A, however, lies not just in its accuracy but in its ‘vision-aware’ and ‘kinematics-aware’ policy. While a kinematics-only model might correctly predict actions based on past movements, it lacks an understanding of the actual anatomical environment. DINO-CV A, by fusing visual and kinematic information, learns to associate expert movements with the vascular context, ensuring the catheter follows the correct anatomical trajectory. This is a crucial step towards truly autonomous navigation, where the system understands its position within the vasculature and its destination.

Ablation studies further confirmed the importance of both goal conditioning and multimodal fusion. Removing either the goal information or one of the modalities (vision or states) led to a significant drop in performance, highlighting that these components are essential and complementary for robust catheter steering, especially in novel environments.

Also Read:

Future Directions

While DINO-CV A represents a significant advancement, the researchers acknowledge several limitations. Future work will focus on exploring different positional encoding strategies to handle varying sequence lengths, incorporating richer spatial goal embeddings, and evaluating the model in more complex vascular phantoms with diverse branching structures. Additionally, extending the study to synthetic X-ray data, which is used in clinical catheterization, will help bridge the gap between laboratory evaluation and real-world clinical applicability.

In conclusion, DINO-CV A demonstrates the feasibility of multimodal, goal-conditioned architectures for autonomous catheter navigation. By learning from expert demonstrations in a trajectory and vision-aware manner, this framework offers a promising path towards reducing operator dependency and enhancing the reliability of catheter-based therapies. You can read the full research paper here.

Meera Iyer
Meera Iyerhttps://blogs.edgentiq.com
Meera Iyer is an AI news editor who blends journalistic rigor with storytelling elegance. Formerly a content strategist in a leading tech firm, Meera now tracks the pulse of India's Generative AI scene, from policy updates to academic breakthroughs. She's particularly focused on bringing nuanced, balanced perspectives to the fast-evolving world of AI-powered tools and media. You can reach her out at: [email protected]

- Advertisement -

spot_img

Gen AI News and Updates

spot_img

- Advertisement -