Proprioceptive Sketches as Long-Horizon Intent
for Generative Action Policies
- 1Hong Kong Polytechnic University
- 2National University of Singapore
- 3Mohamed bin Zayed University of Artificial Intelligence
- 4Great Bay University
*Corresponding author
TL;DR: Alongside each short action chunk, the policy jointly generates a compact, timing-free sketch of the robot's remaining joint-space path. Only the action chunk is executed.
Abstract: Generative robot policies predict short action chunks but lack explicit long-horizon intent. Recent methods expose longer-horizon structure through language plans, subgoal images, or video forecasts, which are costly to generate and still need to be translated into robot motion. Predicting future robot motions avoids this translation, but a dense, time-indexed trajectory requires numerous parameters to cover the full remaining task, and over a short horizon it largely repeats the action chunk and adds little guidance for action generation. We propose Proprioceptive Action Models (PAM), which jointly generate a compact, timing-free sketch of the robot's remaining joint-space path and a dense executable action chunk within a single transformer denoiser. The sketch parameterizes the path by arc length rather than time, capturing geometric intent invariant to execution timing. Block-causal attention and a staggered denoising schedule maintain directed sketch-to-action dependence, ensuring the action tokens condition on a progressively cleaner sketch throughout sampling. In simulation, PAM improves over its action-only counterparts on Push-T and LIBERO-Long; on four real-world bimanual tasks, it raises success from 47.5% to 75.0%.
The proprioceptive sketch
The sketch indexes the remaining configuration path by arc-length phase instead of time and fits it with a few B-spline control points. It covers the whole remaining task with a fixed size and does not change when the same path is executed faster, slower, or with pauses.
Joint sketch and action generation
One transformer denoiser generates both token streams. Actions attend to the sketch but not the reverse, the two streams are trained with independent noise levels, and at inference the sketch is denoised a lead η ahead of the actions.
Each stream gets K = 16 updates for K + ⌈ηK⌉ network calls. Coverage is image-based Push-T.
Real-world dual-arm manipulation
Dual-UR3 platform, three RGB cameras, 14-D joint and gripper commands at 8 Hz; 20 trials per task and method.
PAM-DP rollouts
Successful episodes in real time with the recorded predictions drawn on top. Toggle layers, step through replans (←/→), or replay the denoising where it was recorded.
Handoff pass the loaf over the barrier
Close lid place the pot, then its lid
Fold cloth two coordinated folds
Measure stretch a tape along the loaf
Success rates
- Higher success on every task: 75.0% vs. 47.5% and 43.8% overall (Fisher exact, p < 0.001).
- The gain comes from the remaining path, not from splines: B-spline Policy never beats Diffusion Policy.
- Fewer failures: missed grasps 22 → 10 and coordination errors 8 → 1 compared with Diffusion Policy.
Completion time and failure types
| Method | Handoff | Close lid | Fold | Measure |
|---|---|---|---|---|
| Diffusion Policy | 36.2 ± 4.6 | 44.6 ± 9.2 | 41.9 ± 1.3 | 42.8 ± 3.6 |
| B-spline Policy | 39.9 ± 7.2 | 42.2 ± 11.9 | 42.0 ± 1.2 | 40.6 ± 5.3 |
| PAM-DP (ours) | 35.8 ± 8.0 | 45.3 ± 11.1 | 42.4 ± 2.4 | 45.2 ± 6.4 |
| Method | Total | Miss | Slip | Coord. |
|---|---|---|---|---|
| Diffusion Policy | 42 | 22 | 12 | 8 |
| B-spline Policy | 45 | 28 | 15 | 2 |
| PAM-DP (ours) | 20 | 10 | 9 | 1 |
Simulation benchmarks
PAM-DP is trained from scratch on Push-T; PAM-VLA adds the sketch to the pretrained FLOWER VLA on LIBERO-Long.
| Method | State | Image |
|---|---|---|
| LSTM-GMM | 0.67 / 0.61 | 0.69 / 0.54 |
| IBC | 0.90 / 0.84 | 0.75 / 0.64 |
| Diffusion Policy | 0.95 / 0.79 | 0.78 / 0.66 |
| B-spline Policy | 0.92 / 0.74 | 0.76 / 0.63 |
| PAM-DP (ours) | 0.96 / 0.88 | 0.82 / 0.72 |
| Method | Success |
|---|---|
| π0 | 85.2 |
| π0.5-KI | 85.8 |
| OpenVLA-OFT | 94.5 |
| FLOWER | 94.9 ± 1.2 |
| PAM-VLA (ours) | 95.2 ± 0.4 |
Does the action depend on the sketch?
At a fixed handoff observation and action noise, replacing the sketch with one that takes a different route around the barrier moves the predicted left tool by 10.5 mm RMS at action step 8, compared with 8.5 mm for resampling the action noise and 1.5 mm for a same-route sketch.
Video
Citation
@misc{wang2026pam,
title = {Proprioceptive Sketches as Long-Horizon Intent for Generative Action Policies},
author = {Wang, Fangyuan and Huang, Songhao and Sun, Haoxiang and Lyu, Shipeng and He, Chengyang and Duan, Anqing and Zhou, Peng and Navarro-Alarcon, David},
year = {2026}
}