Original Feishu Document · Source Revision 9
💡
Mechanisms Lesson: Many robotic tasks are difficult to specify with a complete reward function, yet it is easy to judge which behavior is better or to take over when failure occurs. This lesson provides a unified explanation of trajectory preferences, reward models, corrective imitation, DAgger, human intervention, and safety gating.
Learning Objectives
After completing this lesson, you should be able to derive a Bradley-Terry reward model from pairwise preferences; understand preference optimization and reward model error; explain how DAgger addresses state-distribution shift; distinguish demonstrations, corrections, interventions, and preferences; and design human-feedback collection and safety evaluation.
1. Why Not Handcraft Rewards?
Tidiness, comfort, naturalness, and safety are difficult to compress into a single formula. Handcrafted rewards are prone to reward hacking: the model satisfies numerical metrics while violating the true objective.
Human feedback has the advantage of evaluating complex behavior holistically, but it is expensive, noisy, and influenced by context and annotator preferences.
2. Pairwise Preferences
Given two trajectory segments , the annotator selects A as better. The reward model assigns each segment a score:
Interpretation: The total score of a trajectory segment of length L is the sum of the local scores assigned by the reward model to the state-action pair at each step.
Derivation: This is the additive trajectory-reward assumption. It facilitates credit assignment, but if overall style or safety cannot be decomposed into individual steps, the model also requires a segment-level network or non-additive aggregation.
Bradley-Terry probability:
Interpretation: The probability that A is preferred to B is obtained by exponentially normalizing the scores of the two trajectories; equivalently, it is the sigmoid of their score difference.
Derivation: Bradley-Terry assumes that the preference odds between two candidates equal the ratio of their exponentiated scores. Dividing both the numerator and denominator by the exponential of A's score yields the sigmoid of the score difference.
Interpretation: Across all annotations stating that a winning segment is preferred to a losing segment, the preference loss minimizes the negative log-probability assigned by the model to that preference.
Derivation: Treating each pairwise annotation as a binary classification observation and maximizing the likelihood under the Bradley-Terry probability is equivalent to this binary cross-entropy loss. It identifies only relative scores: adding the same constant to every R does not change the probability.
3. Statistical Issues in Preference Data
| Issue | Impact | Mitigation |
|---|---|---|
| Annotator inconsistency | The reward model averages conflicting preferences | Multiple annotations, hierarchical models, confidence estimates |
| Imbalanced segment difficulty | The model learns only obvious failures | Actively select boundary cases |
| Order and presentation bias | Position affects the selection | Randomize the order |
| Unobservable outcomes | Annotators cannot assess contact and force | Multiple views, force curves, task outcomes |
| Distribution shift | New policies produce new types of trajectories | Continuous feedback and OOD detection |
4. Policy Exploitation of the Reward Model
Policy optimization actively searches for weaknesses in the reward model. Even if training-set accuracy is high, out-of-distribution trajectories may receive spuriously high scores.
Use:
- Reward ensembles and uncertainty estimates.
- KL constraints to keep the policy close to the reference policy.
- Actual task outcomes and hard safety constraints.
- Periodic human review.
- Adversarial and counterexample data.
5. Direct Preference Optimization and Policy Weighting
Instead of explicitly training a reward model, one can directly increase the probability of preferred trajectories. An abstract objective is:
Interpretation: Direct preference optimization increases the winning trajectory's log-probability advantage relative to the reference policy and decreases the losing trajectory's relative advantage, then maximizes the preference likelihood through a sigmoid.
Derivation: For a KL-regularized optimal policy, the log-ratio between the policy and the reference policy is proportional to the implicit reward. Substituting this relationship into the Bradley-Terry preference probability eliminates the explicit reward model. For continuous robot trajectories, the log-probability must be summed over conditional action probabilities, with trajectory length and noise scale handled appropriately.
Robot trajectories are continuous action sequences, so probability computation, length normalization, and action noise are more complex than for language tokens. Preferences can also be converted into trajectory weights for weighted behavior cloning.
6. Corrective Imitation and DAgger
Behavior cloning is trained only on the expert's state distribution. DAgger allows the current policy to execute while the expert provides actions for the states it visits:
- Train an initial policy.
- Deploy the policy so that it visits its own state distribution.
- Have the expert label the correct actions for these states.
- Aggregate the new data and retrain.
In theory, this can improve sequence error from potentially quadratic compounding to closer to linear growth, but real-robot deployment requires controlling safety risks and expert costs.
7. Human Intervention
Intervention data includes:
- Signs of impending failure before the intervention.
- The moment at which the intervention is triggered.
- Expert recovery actions.
- The point after recovery when control is returned to the policy.
Interventions do not occur randomly. Humans intervene only when they perceive danger or impending failure, creating selection bias. “No intervention” must be distinguished from “safe.”
8. Roles of Corrections and Preferences
| Feedback | Most directly provides | Best suited for |
|---|---|---|
| Full demonstration | Actions from the initial state to success | Base policy |
| Local correction | Recovery actions from a specific erroneous state | Closed-loop recovery |
| Pairwise preference | Relative quality of two behaviors | Holistic evaluation of style, efficiency, and safety |
| Success label | Outcome of an entire trajectory | Value learning and filtering |
| Natural-language feedback | Causes of failure and high-level recommendations | Subgoals and data annotation |
9. Active Feedback Collection
Humans should not be asked to annotate all trajectories at random. Priority can be given to:
- Samples for which the reward model is uncertain.
- Samples on which two candidate policies strongly disagree.
- Samples near the safety boundary.
- High-value but OOD samples.
- New tasks and new objects.
Active selection improves label efficiency but changes the sampling distribution, which must be recorded during training and evaluation.
10. Connections to VLAs and Hierarchical Planning
Human feedback can be applied at different levels:
- Low-level actions: correct grasping and contact.
- Action chunks: compare trajectory smoothness and success.
- Skills: select a more appropriate recovery skill.
- High-level plans: compare task orderings.
- Visual subgoals: determine whether a generated future is plausible.
The granularity of the feedback must match the granularity of the model output; otherwise, credit assignment becomes ambiguous.
11. Minimal Experiments and Credible Evaluation
- Report the number of annotators, inter-annotator agreement, and the number of samples in each category.
- Revalidate the reward model on trajectories from new policies.
- Compare random annotation with active annotation.
- Hold the feedback budget fixed and compare full demonstrations, corrections, and preferences.
- Report safety incidents, intervention rates, and recovery success rates.
- Check whether the policy exploits reward-model loopholes.
12. Failure Modes and Diagnostics
| Failure mode | Diagnostic evidence | Mitigation |
|---|---|---|
| Reward-model exploitation | Model scores increase while actual success, safety, or human-review ratings decline | Ground-truth outcome constraints, adversarial examples, periodic relabeling |
| Annotator disagreement | Low agreement on choices between the same pair of trajectories | Multiple annotators, hierarchical preferences, retain “uncertain/tie” labels |
| Intervention selection bias | Non-intervention samples are incorrectly interpreted as safe samples | Record observability and reaction latency; construct non-intervention controls |
| Feedback-granularity mismatch | Whole-trajectory preferences cannot localize a specific contact error | Segment comparisons, phase labels, action-level corrections |
| Policy-distribution shift | Reward-model accuracy drops sharply on trajectories from a new policy | Continuous querying, OOD detection, online calibration |
| Delayed human feedback | Intervention actions are temporally misaligned with triggering events | Precise time synchronization, latency sweeps, pre-trigger window annotation |
13. Paper Landscape and Boundaries of Evidence
| Work | Facts established by the paper | Authors' interpretation | Course assessment |
|---|---|---|---|
| DAgger | Iteratively collects expert actions at states visited by the current policy and aggregates them into the training set | Can mitigate the state-distribution shift in imitation learning that causes compounding errors | Real-robot applications must incorporate safety gating, expert costs, and a policy-mixture protocol |
| Deep RL from Human Preferences | Trains a reward model from pairwise segment preferences and then optimizes the policy against it | Human comparisons can supervise complex behaviors for which rewards are difficult to handcraft | Reward-model accuracy is not sufficient final evidence; true objectives and exploitability must be validated on new policies |
| HG-DAgger | Humans decide when to intervene and provide corrections, reducing the need for continuous expert annotation | Human gating can balance safety and data efficiency | Intervention triggering is a selective observation process; non-intervention cannot be used directly as a negative label |
| PEBBLE | Combines unsupervised pretraining, active preference querying, and reward learning to improve feedback efficiency | Actively selecting informative segments can reduce the number of human labels required | Active sampling changes the training distribution; evaluation must return to an independent task distribution and report the annotation budget |
| DPO | Directly optimizes a policy's relative log-probabilities under a reference policy using preference pairs, without an explicit reward model | Preference optimization can be reformulated as a stable supervised objective | Continuous robot trajectories require definitions of probability density, length normalization, and action noise; token-based formulas cannot be transferred directly |
| RECAP | Uses deployment successes, failures, and correction data to improve a general-purpose VLA | Feeding real-robot experience back into training can scale the base policy | The feedback budget and amount of additional data should be held fixed when comparing standard retraining, value weighting, and preference-based methods |
14. Exercises
- Derive the Bradley-Terry preference probability.
- Design a robotics example in which a reward model can be easily exploited.
- Define a safe real-robot data-collection procedure for DAgger.
- Explain the selection bias in intervention data.
- Design an active preference-query strategy.
- Compare low-level action preferences with high-level plan preferences.

