spot_img
HomeResearch & DevelopmentAdvancing Robot Dexterity in Fabric Handling with Semantic Keypoints

Advancing Robot Dexterity in Fabric Handling with Semantic Keypoints

TLDR: CLASP is a novel robotic system designed for general-purpose clothes manipulation, such as folding, flattening, and hanging. Its core innovation lies in using “semantic keypoints” (e.g., “left sleeve”) as a spatial-semantic representation of clothes, which are easily extracted from images and guide robot actions. By integrating Vision Language Models for high-level task planning and a pre-built skill library for low-level execution, CLASP demonstrates superior performance and generalization across diverse clothing types and tasks in both simulations and real-world experiments, significantly advancing the capabilities of home service robots in handling deformable objects.

Imagine a future where your home service robot effortlessly handles your laundry, from folding a T-shirt to hanging a skirt. While this vision has been anticipated for a long time, the complex, high-dimensional geometry of clothes has made it a significant challenge for robots. Most existing methods are limited to specific tasks or types of clothing, struggling to generalize across the vast variety of garments we use daily.

A new research paper introduces a groundbreaking approach called CLASP, which stands for CLothes mAnipulation with Semantic keyPoints. This method aims to achieve general-purpose clothes manipulation, capable of handling different clothes types—like T-shirts, shorts, skirts, and long dresses—and various tasks, including folding, flattening, and hanging. The core innovation behind CLASP is the concept of “semantic keypoints.”

What are Semantic Keypoints?

Semantic keypoints are a sparse yet powerful spatial-semantic representation of clothes. Think of them as specific, meaningful points on a garment, such as “left sleeve,” “right shoulder,” or “hem.” These keypoints are crucial because they are salient for both perception (how the robot sees and understands the clothes) and action (how the robot manipulates them). Unlike previous methods that relied on dense, purely geometric keypoints, semantic keypoints carry explicit meaning and can be described using natural language, making them more intuitive and consistent across different instances of the same clothing type.

How CLASP Works

CLASP uses semantic keypoints to bridge the gap between high-level task planning and low-level action execution. At the high level, it leverages advanced artificial intelligence models known as Vision Language Models (VLMs) to predict task plans based on these semantic keypoints. For example, if the task is to fold a T-shirt, the VLM might identify a plan involving grasping the shoulders and moving them to the center.

At the low level, CLASP executes these plans with the help of a simple, pre-built manipulation skill library. This library contains fundamental actions like “grasp,” “moveto,” “release,” “rotate,” and “pull.” These skills are not hard-coded for specific garments but are parameterized by the semantic keypoints, allowing for flexible and generalizable manipulation.

The process begins with CLASP extracting semantic keypoints from RGB-D images (color and depth information). This extraction is a two-stage process: first, autonomously discovering keypoints on a prototype image of a clothing type, and then precisely matching these keypoints to novel clothes, even if they are crumpled or partially hidden. This stage utilizes powerful vision foundation models like DINOv2, SAM, and OWLv2.

Once the keypoints are extracted, the VLM-powered task planner takes the image, keypoints, and a natural language instruction (e.g., “fold the T-shirt”) to generate a sequence of subtasks. Each subtask specifies a basic skill and the semantic keypoints to interact with. The system also includes a closed-loop mechanism, meaning it verifies the plan before execution and can dynamically replan if unexpected changes occur or if the task isn’t completed as expected.

Also Read:

Impressive Results and Generalization

Extensive experiments, both in simulation and with a real Franka dual-arm robot system, demonstrate CLASP’s effectiveness. In simulations, CLASP outperformed state-of-the-art baseline methods on multiple tasks across diverse clothes types. For instance, it achieved high success rates in folding, flattening, hanging, and placing tasks, even with unseen clothing items, showcasing its strong generalization capabilities.

Real-world experiments on 15 different types of clothes, varying in size, shape, and material, further confirmed CLASP’s robust performance. It achieved an 86% success rate in clothes folding, 66% in flattening, 94% in hanging, and 92% in placing. These results are comparable to, or even surpass, existing task-specific algorithms, but CLASP achieves this across a much broader range of garments and tasks.

While CLASP represents a significant leap forward, the researchers acknowledge some limitations. Extracting semantic keypoints under significant occlusion (e.g., heavily crumpled clothes) remains challenging. Manipulating very large clothes can also be difficult due to robot workspace constraints. Additionally, real-time keypoint extraction for online trajectory optimization is an area for future improvement.

Nevertheless, CLASP’s integration of general semantic keypoint representation with foundation models pretrained on vast internet-scale data makes it a highly effective and general-purpose method for clothes manipulation, bringing us closer to truly intelligent home service robots. You can find more details on their work here: CLASP Research Paper.

Meera Iyer
Meera Iyerhttps://blogs.edgentiq.com
Meera Iyer is an AI news editor who blends journalistic rigor with storytelling elegance. Formerly a content strategist in a leading tech firm, Meera now tracks the pulse of India's Generative AI scene, from policy updates to academic breakthroughs. She's particularly focused on bringing nuanced, balanced perspectives to the fast-evolving world of AI-powered tools and media. You can reach her out at: [email protected]

- Advertisement -

spot_img

Gen AI News and Updates

spot_img

- Advertisement -