BETTER CURRICULUM CoRL 2026
10TH CONFERENCE ON ROBOT LEARNING · 2026

What is the Better Curriculum?

Controller-Shaped Grasping Behavior for Contact Force-Sensitive Manipulation

Ziyan Feng · Zizhao Yuan · Yulong Fu · Yuxin He · Zhiyuan Zhang · Zhengjie Zhang · Jinni Zhou · Renjing Xu · Qiang Nie*

The Hong Kong University of Science and Technology (Guangzhou)

* Corresponding author · arXiv:2609.25887

Leader-follower teleoperation with reflex-assisted and manual fragile-cup grasping
Fig. 1. Leader–follower collection for fragile-cup grasping. The tactile reflex closes the fast contact loop while the human guides the task. Full resolution ↗

FROM THE AUTHORS’ REBUTTAL CLARIFICATION

What if touch could teach through the data?

We initially set out to train tactile-conditioned policies. Adding two tactile-image streams did not reliably help in our setup: ACT degraded substantially, while π0.5 retained task ability without a clear gain.

Yet policies trained on reflex-assisted demonstrations already performed well without tactile input at inference. This changed the question: instead of asking only how to fuse touch into a policy, can high-rate tactile control during collection shape demonstrations so that force-sensitive behavior becomes easier to learn?

Central claim. Tactile feedback can act as a collection-time teacher: it makes a narrow, force-valid contact regime repeatable in the demonstrations, and a tactile-free student can learn that nominal behavior.

3.5 gplastic cup
0.3 mmwall thickness
sub-Newtondamage-sensitive contact
Too little grip slips. Too much grip deforms. The viable region is narrow and hard to judge consistently through a human-operated interface.

CHANGE COLLECTION. KEEP THE POLICY INTERFACE.

Make stable contact collectable. Then make it learnable.

01

Shape the demonstration

A deterministic tactile reflex regulates the gripper at 25 Hz while the human guides the arm. We log the follower’s realized actions, including the controller’s corrections.

Human intent + tactile closed loop
02

Learn nominal behavior

ACT or π0.5 learns from RGB and robot state. Within each backbone, the main comparison keeps architecture, inputs, and training procedure fixed; the demonstration source changes.

Tactile-free policy learning
03

Separate fast correction

At deployment, the policy can control both arm and gripper. A second mode re-engages the tactile arbiter, retaining fast gripper authority for disturbance rejection.

Same policy · optional tactile arbiter
System overview of reflex-shaped data collection, tactile-free policy learning and decoupled deployment
Fig. 2. Teleoperated data collection → tactile-free policy learning → decoupled policy/arm and reflex/gripper deployment. Full resolution ↗
WHAT “CURRICULUM” MEANS HERE

The controller-shaped demonstration distribution — not a staged training schedule.

TACTILE ACCESS ≠ CLOSED-LOOP ACTION SHAPING

The operator can see the signal. The reflex can close the loop.

Manual collection was not “vision only.” The operator could inspect the live tactile GUI, observe slip events, and watch the robot on site. The difference was where the feedback loop closed.

Human-mediated correction must pass through observation, interpretation, decision, and coarse button actuation. The reflex maps tactile state directly to repeatable, bidirectional gripper corrections at 25 Hz.

Same tactile access. Different control loop. Click to focus each pathway.
A

Human observation → reaction → action

Indirect · low-bandwidth
GUIlive tactile parameters
+ slip events
On-siterobot motion
+ visible deformation
01Observeread GUI + scene
02Interpretinfer force / slip
03Decideclose or open?
04Buttoncoarse gripper command
05Robotcontact changes

Corrections are filtered through a serial human perception–decision–actuation loop. Favorable trials can occur, but fine corrections are difficult to reproduce consistently across the full horizon.

B

Tactile state → direct gripper correction

Direct · 25 Hz
TACTILE STATEcontact / slip / load
REFLEXbidirectional correction
GRIPPERrealized action
updated contact state

The tactile signal directly affects the executed gripper action. That makes the correction repeatable at contact, lift, transport, and release — and those realized actions are what enter the demonstrations.

The key distinction is not who can see touch. Both modes had tactile access; only the reflex supplied direct high-rate feedback to the gripper.

We do not invent a standalone millisecond reaction-time number for the operator. The experimentally supported comparison is structural: a human-mediated, coarse command loop versus direct 25 Hz tactile feedback.

Same student.
A better teacher in the data.

MAIN CONTROLLED COMPARISON

30 training demonstrations per source.
20 nominal trials per condition.
Student inference uses no tactile input.

Stable grasps on the nominal plastic cup

π₀.₅ · cross-backbone validation

Reflex-shaped dataPolicy only
95%19 / 20
Manual dataPolicy only
5%1 / 20
Contact-quality screened manual dataPolicy only
30%6 / 20

Collection-time shaping matters. The same policy backbone learns substantially better nominal grasp behavior from reflex-shaped demonstrations.

30 training demonstrations per source · 20 nominal trials per condition · student inference uses no tactile input.

SELECTION ≠ SHAPING

Better selection alone did not close the gap.

A fresh manual pool was ranked using logged contact quality. ACT trained on its top 30 demonstrations reached 30% stable grasps (6/20), versus 95% (19/20) with reflex-shaped data.

Contact traces were used for offline selection only; they did not actuate the gripper or enter policy training.

WHY THE BEHAVIOR TRANSFERS

The student recovers the teacher’s closing profile.

A sustained teacher-near run of at least 10 frames occurs in 15/20 reflex-data ACT episodes, versus 0/20 with the video-screened manual baseline.

The learned effect appears in grasp trajectory shape, not only final success.

FEEDBACK CAN TEACH THROUGH ACTIONS

A controller can teach a policy without becoming a policy input.

The tactile signal does not need to enter the student network to influence what the student learns. During collection, feedback changes the executed action distribution; the student then learns from those realized trajectories.

FEEDBACKTactile statecontact · slip · load
TEACHERClosed-loop controllerhigh-rate corrections
WHAT CHANGESExecuted actionsforce-valid trajectories
DATADemonstrationsshaped distribution
STUDENTTactile-free policyRGB + robot state
OFFLINE SELECTION

Can choose what already happened.

Screening can rank and retain favorable manual trials. It cannot retroactively insert the reflex’s missing corrections into contact, lift, transport, or release.

CLOSED-LOOP SHAPING

Can change what happens.

The controller alters actions while the demonstration is being generated, creating trajectories that are more repeatable and easier for the student to imitate.

THE INTELLECTUAL TAKEAWAY Feedback can be distilled into nominal behavior through the demonstrations it creates.

BEHAVIOR TRANSFERS · REACTIVITY DOES NOT

Nominal behavior can be distilled. Fast correction still needs feedback.

Learning from tactile-shaped demonstrations does not transfer the full tactile feedback loop. Under randomized lateral disturbances, the same reflex-data π0.5 policy benefits from a separate 25 Hz gripper arbiter.

SAME REFLEX-DATA π₀.₅ POLICY · DISTURBANCE TEST
Policy only55%11 / 20 retained
Policy + tactile arbiter100%20 / 20 retained

The policy keeps arm control. The reflex keeps fast gripper authority.

The 20/20 result is an observed outcome under this disturbance protocol, not a universal safety guarantee. This experiment deliberately separates learned nominal behavior from online high-rate reactivity.

EXPLORATORY GENERALIZATION

An unseen paper cup

80%vs.30%

π0.5 stable grasps for reflex-data versus manual-data: 8/10 vs. 3/10.

The trend is favorable but not statistically significant (two-sided Fisher p = 0.0698); this is preliminary scope evidence, not a broad generalization claim.

SECONDARY TRAINING REFINEMENT

Sharper entry, faster release

A gripper dynamics loss matches second-order differences in the demonstrated trajectory, improving teacher-near entry and reducing release time without adding tactile policy inputs.

14.6%ACT release-time reduction
22.8%π0.5 release-time reduction

Where the claim is strongest

Direct-contact fragile-container manipulation with a narrow viable force window. The experiments cover disposable cups, one unseen paper-cup variant, and one disturbance family. They establish a collection-time role for tactile sensing; they do not claim that tactile-free policies are generally better than tactile-conditioned policies.

THE TAKEAWAY

Feedback can shape what a policy learns before policy training even begins.

Close the fast loop during collection. Distill nominal behavior through executed actions. Keep online feedback for reactivity that cannot be distilled.