Shape the demonstration
A deterministic tactile reflex regulates the gripper at 25 Hz while the human guides the arm. We log the follower’s realized actions, including the controller’s corrections.
Human intent + tactile closed loopController-Shaped Grasping Behavior for Contact Force-Sensitive Manipulation
The Hong Kong University of Science and Technology (Guangzhou)
* Corresponding author · arXiv:2609.25887
FROM THE AUTHORS’ REBUTTAL CLARIFICATION
We initially set out to train tactile-conditioned policies. Adding two tactile-image streams did not reliably help in our setup: ACT degraded substantially, while π0.5 retained task ability without a clear gain.
Yet policies trained on reflex-assisted demonstrations already performed well without tactile input at inference. This changed the question: instead of asking only how to fuse touch into a policy, can high-rate tactile control during collection shape demonstrations so that force-sensitive behavior becomes easier to learn?
Central claim. Tactile feedback can act as a collection-time teacher: it makes a narrow, force-valid contact regime repeatable in the demonstrations, and a tactile-free student can learn that nominal behavior.
CHANGE COLLECTION. KEEP THE POLICY INTERFACE.
A deterministic tactile reflex regulates the gripper at 25 Hz while the human guides the arm. We log the follower’s realized actions, including the controller’s corrections.
Human intent + tactile closed loopACT or π0.5 learns from RGB and robot state. Within each backbone, the main comparison keeps architecture, inputs, and training procedure fixed; the demonstration source changes.
Tactile-free policy learningAt deployment, the policy can control both arm and gripper. A second mode re-engages the tactile arbiter, retaining fast gripper authority for disturbance rejection.
Same policy · optional tactile arbiter
The controller-shaped demonstration distribution — not a staged training schedule.
TACTILE ACCESS ≠ CLOSED-LOOP ACTION SHAPING
Manual collection was not “vision only.” The operator could inspect the live tactile GUI, observe slip events, and watch the robot on site. The difference was where the feedback loop closed.
Human-mediated correction must pass through observation, interpretation, decision, and coarse button actuation. The reflex maps tactile state directly to repeatable, bidirectional gripper corrections at 25 Hz.
Corrections are filtered through a serial human perception–decision–actuation loop. Favorable trials can occur, but fine corrections are difficult to reproduce consistently across the full horizon.
The tactile signal directly affects the executed gripper action. That makes the correction repeatable at contact, lift, transport, and release — and those realized actions are what enter the demonstrations.
We do not invent a standalone millisecond reaction-time number for the operator. The experimentally supported comparison is structural: a human-mediated, coarse command loop versus direct 25 Hz tactile feedback.
MAIN CONTROLLED COMPARISON
30 training demonstrations per source.
20 nominal trials per condition.
Student inference uses no tactile input.
π₀.₅ · cross-backbone validation
Collection-time shaping matters. The same policy backbone learns substantially better nominal grasp behavior from reflex-shaped demonstrations.
30 training demonstrations per source · 20 nominal trials per condition · student inference uses no tactile input.SELECTION ≠ SHAPING
A fresh manual pool was ranked using logged contact quality. ACT trained on its top 30 demonstrations reached 30% stable grasps (6/20), versus 95% (19/20) with reflex-shaped data.
Contact traces were used for offline selection only; they did not actuate the gripper or enter policy training.WHY THE BEHAVIOR TRANSFERS
A sustained teacher-near run of at least 10 frames occurs in 15/20 reflex-data ACT episodes, versus 0/20 with the video-screened manual baseline.
The learned effect appears in grasp trajectory shape, not only final success.FEEDBACK CAN TEACH THROUGH ACTIONS
The tactile signal does not need to enter the student network to influence what the student learns. During collection, feedback changes the executed action distribution; the student then learns from those realized trajectories.
Screening can rank and retain favorable manual trials. It cannot retroactively insert the reflex’s missing corrections into contact, lift, transport, or release.
The controller alters actions while the demonstration is being generated, creating trajectories that are more repeatable and easier for the student to imitate.
BEHAVIOR TRANSFERS · REACTIVITY DOES NOT
Learning from tactile-shaped demonstrations does not transfer the full tactile feedback loop. Under randomized lateral disturbances, the same reflex-data π0.5 policy benefits from a separate 25 Hz gripper arbiter.
The policy keeps arm control. The reflex keeps fast gripper authority.
The 20/20 result is an observed outcome under this disturbance protocol, not a universal safety guarantee. This experiment deliberately separates learned nominal behavior from online high-rate reactivity.
EXPLORATORY GENERALIZATION
π0.5 stable grasps for reflex-data versus manual-data: 8/10 vs. 3/10.
The trend is favorable but not statistically significant (two-sided Fisher p = 0.0698); this is preliminary scope evidence, not a broad generalization claim.SECONDARY TRAINING REFINEMENT
A gripper dynamics loss matches second-order differences in the demonstrated trajectory, improving teacher-near entry and reducing release time without adding tactile policy inputs.
Direct-contact fragile-container manipulation with a narrow viable force window. The experiments cover disposable cups, one unseen paper-cup variant, and one disturbance family. They establish a collection-time role for tactile sensing; they do not claim that tactile-free policies are generally better than tactile-conditioned policies.
THE TAKEAWAY