Proprioceptive Sketches as Long-Horizon Intent
for Generative Action Policies

*Corresponding author

TL;DR: Alongside each short action chunk, the policy jointly generates a compact, timing-free sketch of the robot's remaining joint-space path. Only the action chunk is executed.

Show denoising

Sketch noise
Action noise
Sketch (lighter → darker along phase) Control points Action chunk Current tool position

Predictions during playback are the ones the robot executed in this episode. The denoising at each replan is re-run offline with the paper's lead η = 1/8 from the episode's recorded camera frames, joint states and initial noise, using the same checkpoint (the live run used a longer lead; final predictions typically agree within a few millimetres). Everything is projected through forward kinematics into an external camera that the policy does not observe.


Abstract: Generative robot policies predict short action chunks but lack explicit long-horizon intent. Recent methods expose longer-horizon structure through language plans, subgoal images, or video forecasts, which are costly to generate and still need to be translated into robot motion. Predicting future robot motions avoids this translation, but a dense, time-indexed trajectory requires numerous parameters to cover the full remaining task, and over a short horizon it largely repeats the action chunk and adds little guidance for action generation. We propose Proprioceptive Action Models (PAM), which jointly generate a compact, timing-free sketch of the robot's remaining joint-space path and a dense executable action chunk within a single transformer denoiser. The sketch parameterizes the path by arc length rather than time, capturing geometric intent invariant to execution timing. Block-causal attention and a staggered denoising schedule maintain directed sketch-to-action dependence, ensuring the action tokens condition on a progressively cleaner sketch throughout sampling. In simulation, PAM improves over its action-only counterparts on Push-T and LIBERO-Long; on four real-world bimanual tasks, it raises success from 47.5% to 75.0%.


The proprioceptive sketch

The sketch indexes the remaining configuration path by arc-length phase instead of time and fits it with a few B-spline control points. It covers the whole remaining task with a fixed size and does not change when the same path is executed faster, slower, or with pauses.

Try it: timing changes the trajectory, not the sketch
Time-indexed remaining path
Phase-indexed sketch
Execution timing
Demonstrated path Sketch Control points Action chunk Current state

Joint sketch and action generation

One transformer denoiser generates both token streams. Actions attend to the sketch but not the reverse, the two streams are trained with independent noise levels, and at inference the sketch is denoised a lead η ahead of the actions.

Block-causal attention, independent training noise levels, and staggered denoising
(a) Block-causal attention. (b) Independent training noise levels. (c) Staggered denoising.
Try it: staggered denoising with lead η
Sketch tokens
Action tokens

Each stream gets K = 16 updates for K + ⌈ηK⌉ network calls. Coverage is image-based Push-T.


Real-world dual-arm manipulation

Dual-UR3 platform, three RGB cameras, 14-D joint and gripper commands at 8 Hz; 20 trials per task and method.

Push-T, LIBERO-Long, and the four real-world tasks at their initial and goal states

PAM-DP rollouts

Successful episodes in real time with the recorded predictions drawn on top. Toggle layers, step through replans (←/→), or replay the denoising where it was recorded.

Sketch Control points Action chunk Current tool position

Handoff pass the loaf over the barrier

Close lid place the pot, then its lid

Fold cloth two coordinated folds

Measure stretch a tape along the loaf

Success rates

Completion time and failure types
Completion time (s), successful episodes
MethodHandoffClose lidFoldMeasure
Diffusion Policy36.2 ± 4.644.6 ± 9.241.9 ± 1.342.8 ± 3.6
B-spline Policy39.9 ± 7.242.2 ± 11.942.0 ± 1.240.6 ± 5.3
PAM-DP (ours)35.8 ± 8.045.3 ± 11.142.4 ± 2.445.2 ± 6.4
Failed episodes out of 80
MethodTotalMissSlipCoord.
Diffusion Policy4222128
B-spline Policy4528152
PAM-DP (ours)201091
Predictions of Diffusion Policy, B-spline Policy, and PAM-DP projected into the external camera view on the four tasks
Action predictions cover the next moment; PAM-DP's sketch extends over the remaining motion.

Simulation benchmarks

PAM-DP is trained from scratch on Push-T; PAM-VLA adds the sketch to the pretrained FLOWER VLA on LIBERO-Long.

Push-T coverage (max / mean of last 10 checkpoints)
MethodStateImage
LSTM-GMM0.67 / 0.610.69 / 0.54
IBC0.90 / 0.840.75 / 0.64
Diffusion Policy0.95 / 0.790.78 / 0.66
B-spline Policy0.92 / 0.740.76 / 0.63
PAM-DP (ours)0.96 / 0.880.82 / 0.72
LIBERO-Long success (%)
MethodSuccess
π085.2
π0.5-KI85.8
OpenVLA-OFT94.5
FLOWER94.9 ± 1.2
PAM-VLA (ours)95.2 ± 0.4

Does the action depend on the sketch?

At a fixed handoff observation and action noise, replacing the sketch with one that takes a different route around the barrier moves the predicted left tool by 10.5 mm RMS at action step 8, compared with 8.5 mm for resampling the action noise and 1.5 mm for a same-route sketch.

Offline sketch intervention: alternative routes, predicted left-gripper motion, and action sensitivity

Video


Citation

@misc{wang2026pam,
  title  = {Proprioceptive Sketches as Long-Horizon Intent for Generative Action Policies},
  author = {Wang, Fangyuan and Huang, Songhao and Sun, Haoxiang and Lyu, Shipeng and He, Chengyang and Duan, Anqing and Zhou, Peng and Navarro-Alarcon, David},
  year   = {2026}
}