Closed-Loop Visual Tracking and Robotic Grasping

An integrated Piper eye-in-hand RGB-D grasping system and a controlled study of when to commit a moving target pose.

Piper manipulator · ROS 2 · RGB-D · MoveIt 2

The system sees a target through a camera on its wrist, updates its 3D pose, and plans a grasp. The research question is when that changing pose is stable enough to commit to a motion plan.

Move-stop target, seed 42: the stability-gated policy commits after the observation settles, then grasps and lifts the cube. This is a 4×-accelerated, Phase 20 configuration replay, not part of the 150 canonical trials. Replay provenance.
  • 150canonical simulation trials120 controlled + 30 simulated RGB-D
  • 10/10gated RGB-D task success5/5 static + 5/5 move-stop
  • 5/10 vs 0/10snapshot vs continuous tracking10 trials per policy · 5 static + 5 move-stop
Period
Aug 2026
My role
Lead Software & Systems Developer
Stack
ROS 2 · Gazebo · MoveIt 2 · RGB-D

When should the robot commit to a target pose?

A wrist-mounted RGB-D camera can update the target continuously, but the arm must eventually plan against one pose. Committing too early can leave the gripper at an old location; insisting on a fresh pose during every step can also make a valid plan expire. I studied that choice inside an end-to-end Piper pipeline, from perception to grasp verification.

The shared system segments the target in RGB, fuses the mask with registered depth, transforms the 3D estimate through the hand–eye calibration into base_link, and passes a target to MoveIt 2. The arm plans PREGRASP, GRASP, and LIFT; task success requires lift and hold evidence, not merely a valid trajectory.

Flow diagram from wrist RGB-D observation through semantic mask, 3D localization, target-update policy, MoveIt 2 planning, grasp execution, and task-level audit
Implemented system flow. Snapshot, continuous tracking, and stability-gated commitment share the downstream planner and evaluator.

One observation, from pixels to a planner target

These three images are from the same successful gated replay at ROS time 25.047 s. The frame identifiers and timestamps match exactly; the overlay is the actual segmentation output, rather than a redrawn illustration.

Wrist RGB camera view of the Piper workspace and target cube at 25.047 seconds
01 / Wrist RGB view
Colorized semantic mask for the target cube from the same camera timestamp
02 / Semantic target mask
Recorded segmentation overlay on the matching wrist RGB frame
03 / Recorded overlay

At the nearest telemetry sample, ROS time 25.000 s, depth fusion marked the observation valid and fresh. The estimated target in base_link was approximately (0.425, 0.000, 0.040) m. This position comes from the nearby state record; it is not claimed to have the exact image timestamp. The frozen segmentation model card reports 97.37% foreground mIoU on a 37-image held-out split from the recorded collection setup.

Three ways to choose the pose

01 / SNAPSHOT

Commit once

Take an observation and keep that target pose while planning and executing. It works when the target remains at the observed location, but can become stale when the cube moves.

02 / TRACKING

Keep updating

Replan from fresh observations. In the controlled ground-truth track it handled move-stop in 20/20 trials. In the simulated RGB-D track, all ten task trials stopped at final task-level planning after the fresh, low-drift PREGRASP conditions were not met; that result is specific to this configuration.

03 / GATED

Wait, then commit

Collect a short run of fresh target estimates, check their spatial spread and duration, then freeze a stable pose for execution. The gate is one policy in the same pipeline, not a separate perception or grasping system.

Matched move-stop seed 42 with snapshot commitment: the replay ended in task failure at the GRASP stage, with no verified lift. This 3×-accelerated configuration replay is illustrative and is not part of the 150 canonical trials.

Two evaluation tracks, separate denominators

The frozen Phase 20 audit includes 120 controlled ground-truth trials (20 per policy–scenario cell) and 30 simulated semantic RGB-D trials (5 per cell). Both tracks compare static and move-stop targets under the same task-level success definition. The chart reports counts directly; the smaller RGB-D cells are not pooled with ground-truth trials.

Two-track bar chart: controlled ground-truth snapshot 20 of 20 static and 0 of 20 move-stop, tracking 20 of 20 both, gated 20 of 20 both; simulated semantic RGB-D snapshot 5 of 5 static and 0 of 5 move-stop, tracking 0 of 5 both, gated 5 of 5 both
Canonical task success from the final audited CSVs. Controlled ground-truth: 120 trials. Simulated semantic RGB-D: 30 trials. Each bar retains its policy–scenario denominator. Controlled counts · RGB-D counts · RGB-D failure stages.

In simulated RGB-D, gated succeeded in all five static and all five move-stop trials. Snapshot's five move-stop failures occurred at grasp. RGB-D tracking's ten task failures occurred at final task-level planning, despite initial plans being available in all ten cases. These observations identify where this frozen system failed; they do not show that continuous tracking cannot handle moving targets in general.

My contribution and the shared foundation

The Piper/RGB-D grasping foundation was collaborative work already in place. Yue Zhang proposed the continuous-tracking and stability-gated research direction and leads the manuscript; Wei Liu provided the Piper and RGB-D platform and collaborated on the physical system. I led ROS 2/Gazebo/MoveIt 2 simulation development and system integration, connecting wrist RGB-D perception, 3D target localization, collision-aware planning, and grasp-and-lift execution. I implemented the target policies and simulated RGB-D evaluation path, and led benchmark and audit tooling, quantitative analysis, and grasp-physics checks. I contributed to physical-robot deployment; the hardware work was collaborative. The research repository documents the implementation and provenance.

Physical demonstration and evidence boundary

This physical Piper video is a qualitative hardware demonstration. It is not included in the 150 canonical simulation trials and does not provide a repeated-trial hardware success estimate.

Replay and figure provenance

View replay logs, checksums, and audit sources

Both Gazebo videos reproduce the final Phase 20 simulated semantic RGB-D configuration for move-stop seed 42. The gated replay logged TRIAL_FINISHED with task success; the matched snapshot replay logged task failure at grasp. RGB, mask, and overlay carry the identical camera timestamp of 25.047 s in the gated replay. The videos are accelerated for viewing. Media provenance and checksums , gated replay events, snapshot replay events, and the research repository make the source and evaluation boundary inspectable.