Accepted to Robotics: Science and Systems 2026

SID: Sliding into Distribution for Robust Few-Demonstration Manipulation

Recover online. Execute with confidence. Realign when needed.

Yicheng Ma1,2,*Wei Yu1,2,*Zhian Su1,2Xidan Zhang1,2Huixu Dong1,2,†
1 Grasp Lab, Zhejiang University2 Torch Kernel Co., Ltd.

* Equal contribution  ·  † Corresponding author

SID motion field sliding out-of-distribution robot states toward the in-distribution region before an execution policy performs manipulation
Recover, execute, and realign. An object-centric motion field moves out-of-distribution states toward demonstrated support; closed-loop SID continues to monitor drift and recover online when needed.
01 · Overview

Distribution shift is a control problem

Generalizing robotic manipulation across object poses, viewpoints, and dynamic disturbances is difficult—especially with only a few demonstrations. End-to-end visuomotor policies are expressive but data-hungry, while planning and optimization do not directly capture the interaction strategies demonstrated by humans.

We propose Sliding into Distribution (SID), a structured framework that learns an object-centric motion field from canonicalized demonstrations. The field iteratively slides the system toward the demonstrated manifold and into the reliable operating region of a lightweight egocentric execution policy. In the closed-loop variant, an auxiliary confidence signal monitors distribution drift online and routes the robot back to field-based realignment when needed.

Across six real-world tasks, SID achieves approximately 90% success under OOD initializations using only two demonstrations, while maintaining robustness to distractors and external disturbances.

01

Online recovery

Continuously brings unseen or drifted states back toward demonstrated support instead of asking the policy to extrapolate.

02

Consistent augmentation

Reprojects point clouds and actions together, preserving their geometric and kinematic relationship.

03

Closed-loop routing

Uses ID confidence to switch between execution, field-based realignment, and recovery as conditions change.

02 · Method

A closed-loop path back in distribution

SID separates global alignment from local interaction, then reconnects them through confidence-aware online recovery.

SID architecture with an object-centric motion field, egocentric point-cloud data augmentation, and an egocentric execution policy
SID overview. The motion field predicts object-centric sliding steps; reprojection produces kinematically consistent ID and OOD samples; the execution policy predicts actions and ID confidence from egocentric point clouds.
A

Object-centric motion field

A smooth descent field over canonicalized SE(3) approach states supplies large corrections far from the demonstrations and naturally vanishes near the demonstrated manifold.

OOD → ID alignment

B

Egocentric augmentation

Hand–eye-calibrated point-cloud reprojection perturbs viewpoints while updating relative actions consistently, creating valid ID training examples and explicit OOD negatives.

Geometric consistency

C

Confidence-aware execution

A lightweight flow-matching policy handles precise manipulation, while an auxiliary ID-confidence head detects drift and enables closed-loop realignment.

Closed-loop recovery

Object-centric representation space anchored by a common SE(3) object pose
Representation space. Segmented images, point clouds, or object poses can share a common SE(3) anchor for pose-aligned distance computation.
Object-centric representation

A shared anchor across modalities

SID defines an object-centric representation space anchored by a target pose in SE(3). Segmented images, point clouds, and explicit object poses can therefore be compared through the same pose-aligned geometry.

In our implementation, the target-object pose is estimated in the wrist-camera frame. Canonicalizing demonstrations around this anchor suppresses scene- and camera-specific variation, giving the motion field a consistent space in which to learn corrective steps.

From alignment to execution

Closed-loop recovery beyond the first handoff

SID-open performs a single field-to-policy handoff. SID-closed instead keeps monitoring the policy’s operating region and can route control back to realignment or recovery throughout execution.

Open-loop and closed-loop SID inference pipelines
Two inference modes. SID-open hands off to the execution policy once the field norm is low. SID-closed additionally monitors ID and pose confidence, routing drifted states back to field alignment or recovery.
03 · Experiments

Robust with two demonstrations

All success rates are measured over 50 real-world trials per task. SID uses 2 raw demonstrations; trainable baselines use 100.

89.7%

average OOD success for SID-closed across six tasks

+25.7 points over MT3 on average
Static evaluation under OOD initializations
TaskBest baselineSID-openSID-closed
Open Drawer769092
Pour Water788690
Hang Tape468288
Hang Cup748690
PnP-Box588492
Multi-PnP-Box528286

Success rate (%). “Best baseline” is the strongest non-SID result for each task.

01

Workspace-wide OOD generalization

SID recovers from varied object poses and wrist-camera viewpoints before executing each task.

OOD

Open Drawer

OOD

Pour Water

OOD

Hang Tape

OOD

Hang Cup

OOD

Pick & Place

OOD · Long horizon

Multi-object PnP

02

Distractors and external disturbances

Target-centric perception filters clutter, while closed-loop confidence detects execution drift and triggers realignment.

Table II

External disturbances

50 trials / task

TaskBest baselineSID-OSID-C
Hang Tape804484
Hang Cup683688
PnP-Box745286
Pour Water784082

85.0% average success for SID-closed, compared with 75.0% for the strongest baseline.

Table III

Distractor objects

50 trials / task

TaskBest baselineSID-OSID-C
Hang Cup728688
PnP-Box748490

89.0% average success for SID-closed, with less than a 10% drop from the corresponding static OOD setting.

Success rate (%). “Best baseline” reports the strongest non-SID method for each task. SID-O and SID-C denote SID-open and SID-closed.

03

Long-horizon reuse and recomposition

Lightweight SID skills can be loaded together and invoked stage by stage—without retraining an end-to-end policy.

04 · Discussion

Why online recovery matters

SID treats distribution shift as a control signal: detect it, realign online, and resume execution inside the policy’s reliable region.

Takeaway

The experiments suggest that few-demonstration robustness comes not only from a stronger execution policy, but from actively recovering its operating distribution. Closed-loop confidence-aware routing is especially important under disturbances, where the robot may need to realign more than once.

Scope and outlook. The current system still assumes reliable object-centric perception and a visible, reachable approach region. Future work can make recovery more object-agnostic and obstacle-aware.

05 · Citation

Cite SID

If you find this work useful, please consider citing our paper.

Accepted toRSS 2026
BibTeX
@misc{ma2026sidslidingdistributionrobust,
  title         = {SID: Sliding into Distribution for Robust
                   Few-Demonstration Manipulation},
  author        = {Yicheng Ma and Wei Yu and Zhian Su and
                   Xidan Zhang and Huixu Dong},
  year          = {2026},
  eprint        = {2605.13428},
  archivePrefix = {arXiv},
  primaryClass  = {cs.RO},
  url           = {https://arxiv.org/abs/2605.13428}
}