Skip to content

Original Feishu Document · Source Revision 9

🔬

Paper Lab: Use temporal abstraction, subgoal representations, memory updates, and closed-loop recovery as a unified framework for comparing traditional hierarchical RL with emerging hierarchical VLA systems.

Unified Comparison Framework ​

DimensionQuestion
High-level variableOption, language, image, skill token, or latent
Low-level policyRL, imitation, VLA, or controller
TerminationFixed duration, learned termination, or completion detection
Future predictionModel-free, value-based, or world model
MemoryContext, external memory, parameters, or fast weights
EvidenceLong-horizon success, compositional generalization, recovery, and real-robot performance

1. Options and Option-Critic ​

Options provide rigorous definitions of initiation, intra-option policies, and termination. Option-Critic learns these components end to end and represents the approach of “learning temporal abstractions through RL.”

Its advantage is a clear mathematical definition; the challenge is that automatically discovered skills may be uninterpretable, irrelevant to the task, or switch too frequently.

2. Feudal / Goal-Conditioned Hierarchy ​

The high level outputs a subgoal in the state space or latent space, and the low-level policy is responsible for reaching it. Key questions are whether the subgoal is achievable, whether it can be reused across tasks, and how to select the high-level timescale.

3. Language Planning and Embodied Chain-of-Thought ​

A VLM generates language sub-tasks or reasoning chains, which a low-level robot policy then executes. Compared with traditional Options, the high-level variables are semantic and interpretable, but they lack explicit initiation and termination conditions as well as guarantees of physical reachability.

Decisive experiments must examine closed-loop feedback rather than merely whether the textual plan appears plausible.

4. Hierarchical Reasoning in π0.5 ​

π0.5 connects high-level semantics, heterogeneous co-training, and low-level actions. It demonstrates the combination of language sub-tasks with a VLA, but this does not by itself prove that the model possesses explicit long-horizon planning or a causal world model.

5. Prompts and Visual Subgoals in π0.7 ​

π0.7 uses diverse prompts and generated visual subgoals to improve policy controllability. It connects:

  • VLA: low-level action generation.
  • World model: future visual generation.
  • Hierarchical planning: high-level prompts and subgoals.
  • Data strategy: mixed-quality data with metadata.

It is necessary to verify whether the visual goals are reachable, whether the low-level policy actually uses them, and whether they are regenerated after failures.

6. WAM-TTT ​

WAM-TTT uses test-time training to write human videos into fast-weight memory. Its key innovation is that the inner loop uses only human-side signals, while the outer loop uses supervision from robot action tasks to update the learning rule.

Key assumptions include human–robot task correspondence, temporal/event alignment, and memory safety.

7. Method Comparison ​

MethodHigh-level variableTerminationMemoryPrimary risk
OptionsDiscrete skillsExplicit βUsually noneSkill discovery
Goal-conditionedState/latent goalGoal-reach detectionOptionalUnreachable subgoals
Language planningLanguageCompletion detectionContextPhysical grounding
π0.7Diverse prompts/visual goalsClosed-loop policyContextGoal hallucination
WAM-TTTHuman demonstrations written into memoryTask policyFast weightsMemory contamination

8. Unified Experimental Protocol ​

  1. Use the same long-horizon task and low-level policy.
  2. Compare a flat baseline, a fixed task graph, learned Options, language subgoals, and visual subgoals.
  3. Introduce intermediate perturbations and failure recovery.
  4. Hold out skill combinations and task orderings.
  5. Provide memory with correct, irrelevant, and incorrect demonstrations.
  6. Report stage-level success, switching errors, recovery rate, and full-task success.

9. Lab Assignment ​

Lab Deep Dive|Unified Formulation, Visualization, and Reproduction ​

Place Options, embodied chain-of-thought, π0.7, and WAM-TTT within the same hierarchical decision-making framework. The high-level variable can be an option, a language step, a visual subgoal, or a memory-retrieval result; the low-level policy executes:

Interpretation: While the high-level condition z_k remains active, low-level actions are sampled from a policy distribution conditioned on the current observation and that high-level condition.

Derivation: Whether z_k is an Option, a language step, a visual subgoal, or retrieved memory, the low-level interface can be unified as a conditional policy. When comparing different approaches, pi_low must be held fixed so that performance differences can be attributed primarily to the high-level representation, selection, and termination mechanisms.

The high level is updated on a slower timescale:

Interpretation: At the k-th high-level decision boundary, the high-level policy selects a new high-level condition z_k based on the history, overall task goal, and current memory.

Derivation: The high level is updated only at boundaries t_k rather than at every motor-control cycle. Conditioning on task history h, goal g, and memory m provides a unified representation of task-graph selection, language-based replanning, visual-goal generation, and test-time memory retrieval.

In words: “The high level selects the next executable condition based on history, goals, and memory, and the low level converts it into continuous actions.” When comparing papers, ask whether the subgoal is reachable, when it terminates, whether failure can be detected, and whether memory actually changes behavior at inference time.

Minimal Experiment: Unified Reproduction Protocol ​

Choose a task with at least five stages that permits objects to be moved mid-execution or grasp failures to be induced. Hold the low-level policy, training data, visual encoder, control frequency, model-call limit, and number of real-robot interaction steps fixed. Replace only the high-level mechanism with one of five alternatives: a fixed task graph, learned Options, language subgoals, visual subgoals, or test-time memory.

Minimum reporting requirements: Full-task success rate, mean number of completed stages, invalid-subgoal rate, termination false-positive and false-negative rates, switching latency, local recovery rate, success rate on unseen skill combinations, number of planning calls, adaptation compute, and the proportion of identical failures attributed to the high level, low level, interface, or completion detector.

  1. Choose a task that requires more than five stages and permits intermediate perturbations.
  2. Compare a flat policy, a fixed task graph, learned Options, language plans, and visual subgoals.
  3. Hold the low-level policy constant and replace only the high-level representation and termination mechanism.
  4. Report stage completion rate, error recovery, repeated loops, planning calls, and final success rate.

Lab Exercises ​

  1. Specify an option’s initiation set, intra-option policy, and termination function.
  2. Explain why a language plan that “looks correct” does not imply that its subgoals are executable.
  3. Design an intervention proving that WAM-TTT uses the current demonstration rather than the task name.
  4. Compare the boundary between π0.7 visual subgoals and explicit world-model planning.
  5. Draw module boundaries for the five methods.
  6. Add initiation and completion conditions for π0.7.
  7. Express WAM-TTT as inner-loop/outer-loop pseudocode.
  8. Design an ablation that tests whether textual CoT improves real closed-loop performance.
  9. Compare visual subgoals with state-space subgoals.
  10. State what each paper actually proves and does not prove.

Primary Failure Modes ​

FailureSymptomUnified diagnosis
Unfair comparisonOne approach uses a stronger low-level policy, longer context, or more model callsHold the low-level policy, data, interaction steps, latency, and inference budget fixed
Unreachable high-level variablesLanguage or visual subgoals are plausible but cannot be achieved by the low levelTest preconditions, reachability, and conditional success rates
Missing termination mechanismThe system fails to switch after completing a subgoal or advances before completionReport termination precision, recall, and switching-time error under a unified protocol
Low level masks the high levelA strong low-level policy completes the task directly, obscuring whether the high-level variable is usedApply high-level-condition permutation, masking, and counterfactual interventions
Textual or visual explanation bypassThe intermediate representation changes, but the actions do notIntervene on z_k under the same observation and measure changes in actions and outcomes
Memory contaminationIncorrect demonstrations cause persistent degradation during test-time adaptationUse incorrect and irrelevant demonstrations, gating, version snapshots, and rollback
Reporting only average success rateIt is impossible to determine whether failures arise from planning, execution, the interface, or verificationDecompose the evidence chain by stage and module, and report confidence intervals

Paper Facts, Authors’ Interpretations, and Course Assessment ​

ApproachPaper factAuthors’ interpretationCourse assessment
Options / Option-CriticFormalizes actions spanning multiple steps and can learn intra-option policies and terminationTemporal abstraction can shorten decision chains and discover skillsProvides the clearest mathematical interface, but automatically discovered skills may still collapse, remain uninterpretable, or be difficult to compose across tasks
Embodied language planningSayCan, Inner Monologue, and related methods use language priors, executability, and environmental feedback for skill-level planningLinguistic knowledge and closed-loop feedback can support long-horizon tasksText-generation ability, skill-library coverage, and gains in real physical closed-loop performance must be distinguished
π0.7A general-purpose VLA uses diverse conditioning forms, including visual subgoals, to improve controllabilityRicher high-level conditions can improve open-world task executionIt remains primarily a VLA; this lab audits only subgoal reachability, utilization, completion detection, and recovery
WAM-TTTHuman-side context is written into fast memory at test time, while robot-task supervision trains the update mechanismHuman demonstrations at test time can help robots adaptRequires strict holdouts, stage alignment, incorrect-demonstration controls, and evidence of rollback capability

Course Cross-Reading ​

D0|Hierarchical Planning, Skills, and Memory

D1|Options, Skills, and Hierarchical Reinforcement Learning

D2|Subgoal Planning and Embodied Reasoning

D3|Memory and Test-Time Adaptation

π0.7: Visual Subgoals and Automated Prompting

Article text is licensed under the Apache License 2.0