CoRL 2025 ①|LimX Dynamics Advances Video Data Training for Autonomous Robotic Manipulation…
LimX Dynamics, as a collaborating research partner, has had two papers accepted at the Conference on Robot Learning (CoRL 2025), a leading…
CoRL 2025 ①|LimX Dynamics Advances Video Data Training for Autonomous Robotic Manipulation Partnering with SUSTech and HKU
[embed]GVF-TAPE Overview

GVF-TAPE CoRL, 2025 Conference on Robot Learning
LimX Dynamics, as a collaborating research partner, has had two papers accepted at the Conference on Robot Learning (CoRL 2025), a leading international academic conference for robotics and machine learning. Here we present the first paper: “Generative Visual Foresight Meets Task-Agnostic Pose Estimation in Robotic Table-top Manipulation (2025)”, which enables robots to learn autonomous manipulation by predicting task execution from videos.
Today, much of the robot learning still relies on costly human demonstrations or specialized hardware, which limits scalability and slows deployment.
To address this, a joint team from Southern University of Science and Technology (SUSTech), LimX Dynamics, and the University of Hong Kong (HKU) have introduced our latest framework: Generative Visual Foresight + Task-Agnostic Pose Estimation (GVF-TAPE). This framework sets out a new paradigm: robots no longer depend on costly human demonstrations, but instead learn directly from unlabeled videos and their own random motion exploration.
The process works in three steps.
- First, robots predict future RGB-D visual outcomes by leveraging an existing generative video model, not just color frames but also depth information that provides spatial awareness. This allows them to “imagine” what completing the task should look like.
![[Figure 1. Generated RGB-D operation video. The first row shows RGB frames, and the second row shows the corresponding depth maps.]](https://miro.medium.com/v2/resize:fit:1400/1*DEy9rxr5sxbjeNGSwjhjJw.png)
[Figure 1. Generated RGB-D operation video. The first row shows RGB frames, and the second row shows the corresponding depth maps.]
- Next, they extract the end-effector poses from these predicted sequences.
- Finally, they execute actions in the real world, directly translating foresight into motions, without requiring human demonstrations.
![[Video 1. Real-World Task Examples]](https://miro.medium.com/v2/resize:fit:1138/1*encW4VkQdW5Fjp0WFfyeXQ.gif)
[Video 1. Real-World Task Examples]
This approach closes the loop from video understanding to real-world execution, offering a far more flexible and efficient way to train robots. Robots can build skills from abundant online video data and their own random motion exploration, then generalize and adapt those skills to various tasks and real-world environments.
Core Breakthroughs:
Low-cost RGB-D Data
Robots generate depth information from 2D videos. No expensive depth sensors are required.
Traditionally, robots need specialized depth sensors to understand objects in three dimensions. GVF-TAPE removes this dependency by generating RGB-D information directly from ordinary 2D video. This effectively gives robots low-cost stereo-based pose estimation, making them more precise in spatially complex environments without extra hardware, and advancing scalable deployment in real-world scenarios.
Scalable learning from motion exploration
Robots teach themselves through motion exploration, creating reusable data across tasks and settings.
GVF-TAPE replaces costly human demonstrations with a task-agnostic pose estimation model trained on real-world random exploration, where robots record diverse 3D end-effector poses to create a reusable dataset, decoupling from specific tasks.
![[Video 2. Robot Autonomous Random Exploration]](https://miro.medium.com/v2/resize:fit:1138/1*3FM-4TK0GNguF3uWrBO0yg.gif)
[Video 2. Robot Autonomous Random Exploration]
Real-time foresight with Rectified Flow
Rectified Flow produces high-quality predictions in only three steps, allowing robots to act quickly and reliably in real time.
Prior video prediction methods often rely on diffusion models, where higher-quality results comes only at the cost of many sampling steps and increases inference time. GVF-TAPE instead adopts Rectified Flow, achieving comparable video quality in just three steps and drastically reducing latency. This level of efficiency is critical for enabling real-time closed-loop control.
Why this matters:
Together, these breakthroughs make robot training faster, at lower cost, and more flexible, paving the way for real-world deployment across industrial applications.
Earlier this year, LimX Dynamics introduced LimX VGM, an approach connecting human operation videos to robot execution by refining existing large video generation model. Building on that foundation, the GVF-TAPE framework pushes the idea further with faster, more scalable, and more adaptable approaches. Together, these milestones highlight LimX Dynamics’ continued progress in turning videos into a universal training resource for embodied AI, helping accelerate the development of autonomous robotic manipulation.
About LimX Dynamics
LimX Dynamics is an embodied intelligence robotics company driving the innovation of full-size general-purpose humanoid robots, and other innovative products.
We are committed to disruptive technology in Embodied AI, with the mission to unlock the generalization of Artificial General Intelligence (AGI) in real world. We focus on three core technologies: hardware design and manufacturing, RL-based motion control, and embodied AI training paradigms.
Our products and technologies are built to serve Innovators, Developers, and System Integrators (IDS), accelerating embodied AI research, development, and real-world deployment across fields such as research, manufacturing, business, and household services. Founded in 2022, LimX Dynamics is headquartered in Shenzhen, China.
For more information, visit our website: www.limxdynamics.com, YouTube channel and LinkedIn Page.
메타데이터
- post_id
- b3a6bb5e507a
- slug
- corl-2025-①-limx-dynamics-advances-video-data-training-for-autonomous-robotic-manipulation-b3a6bb5e507a
- url
- https://medium.com/@limxdynamics/corl-2025-%E2%91%A0-limx-dynamics-advances-video-data-training-for-autonomous-robotic-manipulation-b3a6bb5e507a
- canonical_url
- https://medium.com/@limxdynamics/corl-2025-%E2%91%A0-limx-dynamics-advances-video-data-training-for-autonomous-robotic-manipulation-b3a6bb5e507a
- author_url
- https://medium.com/@limxdynamics
- status
- ok
- fetched_at
- 2026-08-01 01:21:10