Skip to content

Original Feishu Document · Source Revision 9

💡

Mechanisms Lesson: Many robotic tasks are difficult to specify with a complete reward function, yet it is easy to judge which behavior is better or to take over when failure occurs. This lesson provides a unified explanation of trajectory preferences, reward models, corrective imitation, DAgger, human intervention, and safety gating.

Learning Objectives ​

After completing this lesson, you should be able to derive a Bradley-Terry reward model from pairwise preferences; understand preference optimization and reward model error; explain how DAgger addresses state-distribution shift; distinguish demonstrations, corrections, interventions, and preferences; and design human-feedback collection and safety evaluation.

1. Why Not Handcraft Rewards? ​

Tidiness, comfort, naturalness, and safety are difficult to compress into a single formula. Handcrafted rewards are prone to reward hacking: the model satisfies numerical metrics while violating the true objective.

Human feedback has the advantage of evaluating complex behavior holistically, but it is expensive, noisy, and influenced by context and annotator preferences.

2. Pairwise Preferences ​

Given two trajectory segments , the annotator selects A as better. The reward model assigns each segment a score:

Interpretation: The total score of a trajectory segment of length L is the sum of the local scores assigned by the reward model to the state-action pair at each step.

Derivation: This is the additive trajectory-reward assumption. It facilitates credit assignment, but if overall style or safety cannot be decomposed into individual steps, the model also requires a segment-level network or non-additive aggregation.

Bradley-Terry probability:

Interpretation: The probability that A is preferred to B is obtained by exponentially normalizing the scores of the two trajectories; equivalently, it is the sigmoid of their score difference.

Derivation: Bradley-Terry assumes that the preference odds between two candidates equal the ratio of their exponentiated scores. Dividing both the numerator and denominator by the exponential of A's score yields the sigmoid of the score difference.

Interpretation: Across all annotations stating that a winning segment is preferred to a losing segment, the preference loss minimizes the negative log-probability assigned by the model to that preference.

Derivation: Treating each pairwise annotation as a binary classification observation and maximizing the likelihood under the Bradley-Terry probability is equivalent to this binary cross-entropy loss. It identifies only relative scores: adding the same constant to every R does not change the probability.

3. Statistical Issues in Preference Data ​

IssueImpactMitigation
Annotator inconsistencyThe reward model averages conflicting preferencesMultiple annotations, hierarchical models, confidence estimates
Imbalanced segment difficultyThe model learns only obvious failuresActively select boundary cases
Order and presentation biasPosition affects the selectionRandomize the order
Unobservable outcomesAnnotators cannot assess contact and forceMultiple views, force curves, task outcomes
Distribution shiftNew policies produce new types of trajectoriesContinuous feedback and OOD detection

4. Policy Exploitation of the Reward Model ​

Policy optimization actively searches for weaknesses in the reward model. Even if training-set accuracy is high, out-of-distribution trajectories may receive spuriously high scores.

Use:

  • Reward ensembles and uncertainty estimates.
  • KL constraints to keep the policy close to the reference policy.
  • Actual task outcomes and hard safety constraints.
  • Periodic human review.
  • Adversarial and counterexample data.

5. Direct Preference Optimization and Policy Weighting ​

Instead of explicitly training a reward model, one can directly increase the probability of preferred trajectories. An abstract objective is:

Interpretation: Direct preference optimization increases the winning trajectory's log-probability advantage relative to the reference policy and decreases the losing trajectory's relative advantage, then maximizes the preference likelihood through a sigmoid.

Derivation: For a KL-regularized optimal policy, the log-ratio between the policy and the reference policy is proportional to the implicit reward. Substituting this relationship into the Bradley-Terry preference probability eliminates the explicit reward model. For continuous robot trajectories, the log-probability must be summed over conditional action probabilities, with trajectory length and noise scale handled appropriately.

Robot trajectories are continuous action sequences, so probability computation, length normalization, and action noise are more complex than for language tokens. Preferences can also be converted into trajectory weights for weighted behavior cloning.

6. Corrective Imitation and DAgger ​

Behavior cloning is trained only on the expert's state distribution. DAgger allows the current policy to execute while the expert provides actions for the states it visits:

  1. Train an initial policy.
  2. Deploy the policy so that it visits its own state distribution.
  3. Have the expert label the correct actions for these states.
  4. Aggregate the new data and retrain.

In theory, this can improve sequence error from potentially quadratic compounding to closer to linear growth, but real-robot deployment requires controlling safety risks and expert costs.

7. Human Intervention ​

Intervention data includes:

  • Signs of impending failure before the intervention.
  • The moment at which the intervention is triggered.
  • Expert recovery actions.
  • The point after recovery when control is returned to the policy.

Interventions do not occur randomly. Humans intervene only when they perceive danger or impending failure, creating selection bias. “No intervention” must be distinguished from “safe.”

8. Roles of Corrections and Preferences ​

FeedbackMost directly providesBest suited for
Full demonstrationActions from the initial state to successBase policy
Local correctionRecovery actions from a specific erroneous stateClosed-loop recovery
Pairwise preferenceRelative quality of two behaviorsHolistic evaluation of style, efficiency, and safety
Success labelOutcome of an entire trajectoryValue learning and filtering
Natural-language feedbackCauses of failure and high-level recommendationsSubgoals and data annotation

9. Active Feedback Collection ​

Humans should not be asked to annotate all trajectories at random. Priority can be given to:

  • Samples for which the reward model is uncertain.
  • Samples on which two candidate policies strongly disagree.
  • Samples near the safety boundary.
  • High-value but OOD samples.
  • New tasks and new objects.

Active selection improves label efficiency but changes the sampling distribution, which must be recorded during training and evaluation.

10. Connections to VLAs and Hierarchical Planning ​

Human feedback can be applied at different levels:

  • Low-level actions: correct grasping and contact.
  • Action chunks: compare trajectory smoothness and success.
  • Skills: select a more appropriate recovery skill.
  • High-level plans: compare task orderings.
  • Visual subgoals: determine whether a generated future is plausible.

The granularity of the feedback must match the granularity of the model output; otherwise, credit assignment becomes ambiguous.

11. Minimal Experiments and Credible Evaluation ​

  1. Report the number of annotators, inter-annotator agreement, and the number of samples in each category.
  2. Revalidate the reward model on trajectories from new policies.
  3. Compare random annotation with active annotation.
  4. Hold the feedback budget fixed and compare full demonstrations, corrections, and preferences.
  5. Report safety incidents, intervention rates, and recovery success rates.
  6. Check whether the policy exploits reward-model loopholes.

12. Failure Modes and Diagnostics ​

Failure modeDiagnostic evidenceMitigation
Reward-model exploitationModel scores increase while actual success, safety, or human-review ratings declineGround-truth outcome constraints, adversarial examples, periodic relabeling
Annotator disagreementLow agreement on choices between the same pair of trajectoriesMultiple annotators, hierarchical preferences, retain “uncertain/tie” labels
Intervention selection biasNon-intervention samples are incorrectly interpreted as safe samplesRecord observability and reaction latency; construct non-intervention controls
Feedback-granularity mismatchWhole-trajectory preferences cannot localize a specific contact errorSegment comparisons, phase labels, action-level corrections
Policy-distribution shiftReward-model accuracy drops sharply on trajectories from a new policyContinuous querying, OOD detection, online calibration
Delayed human feedbackIntervention actions are temporally misaligned with triggering eventsPrecise time synchronization, latency sweeps, pre-trigger window annotation

13. Paper Landscape and Boundaries of Evidence ​

WorkFacts established by the paperAuthors' interpretationCourse assessment
DAggerIteratively collects expert actions at states visited by the current policy and aggregates them into the training setCan mitigate the state-distribution shift in imitation learning that causes compounding errorsReal-robot applications must incorporate safety gating, expert costs, and a policy-mixture protocol
Deep RL from Human PreferencesTrains a reward model from pairwise segment preferences and then optimizes the policy against itHuman comparisons can supervise complex behaviors for which rewards are difficult to handcraftReward-model accuracy is not sufficient final evidence; true objectives and exploitability must be validated on new policies
HG-DAggerHumans decide when to intervene and provide corrections, reducing the need for continuous expert annotationHuman gating can balance safety and data efficiencyIntervention triggering is a selective observation process; non-intervention cannot be used directly as a negative label
PEBBLECombines unsupervised pretraining, active preference querying, and reward learning to improve feedback efficiencyActively selecting informative segments can reduce the number of human labels requiredActive sampling changes the training distribution; evaluation must return to an independent task distribution and report the annotation budget
DPODirectly optimizes a policy's relative log-probabilities under a reference policy using preference pairs, without an explicit reward modelPreference optimization can be reformulated as a stable supervised objectiveContinuous robot trajectories require definitions of probability density, length normalization, and action noise; token-based formulas cannot be transferred directly
RECAPUses deployment successes, failures, and correction data to improve a general-purpose VLAFeeding real-robot experience back into training can scale the base policyThe feedback budget and amount of additional data should be held fixed when comparing standard retraining, value weighting, and preference-based methods

14. Exercises ​

  1. Derive the Bradley-Terry preference probability.
  2. Design a robotics example in which a reward model can be easily exploited.
  3. Define a safe real-robot data-collection procedure for DAgger.
  4. Explain the selection bias in intervention data.
  5. Design an active preference-query strategy.
  6. Compare low-level action preferences with high-level plan preferences.

Article text is licensed under the Apache License 2.0