Original Feishu Document · Source Revision 15
💡
Mechanisms Lesson: This lesson examines how high-level plans generate subgoals that are achievable, verifiable, and recoverable, and distinguishes among textual chains of thought, task graphs, visual subgoals, and world-model-based counterfactual planning.
Learning Objectives
After completing this lesson, you should be able to define the four requirements for subgoals; construct task graphs and preconditions; explain why language planning is not equivalent to embodied reasoning; compare symbolic, visual, state, and latent subgoals; and design closed-loop replanning and recovery experiments.
1. What Is an Effective Subgoal?
A subgoal must be:
- Observable: The system can determine whether it has been achieved.
- Achievable: The low-level policy is capable of completing it.
- Useful: Completing it increases the value of the overall task.
- Recoverable: If it fails, the system can diagnose the failure and select an alternative path.
“Tidy up the kitchen” is not a subgoal that can be executed by a low-level policy; “the cup is stably placed on the upper shelf” is closer to a verifiable state.
2. Task Graphs and Preconditions
A task graph:
Interpretation: A task graph consists of a set of subgoal nodes and a set of permitted transition edges.
Derivation: Abstract each verifiable task stage as a node, and abstract ordering constraints, preconditions, and recovery paths as directed edges. This produces an inspectable and modifiable high-level plan structure.
Nodes are subgoals, while edges represent permitted transitions. Each node may have preconditions and completion conditions .
If the current state does not satisfy the preconditions, the high-level planner cannot simply output that subgoal. The rules may come from human knowledge or be estimated through world models and value learning.

💡
Interactive Validation|Hierarchical Planning, Skills, and Memory Lab
Change the hierarchy depth, closed-loop replanning frequency, and test-time memory to observe error propagation and adaptability in long-horizon tasks.
Hierarchical Planning, Skills, and Memory Lab
:::3. Subgoal Representations
| Representation | Advantages | Main Problems |
|---|---|---|
| Language | Semantically clear and compositional | Imprecise geometry and contact |
| Visual goal | Encodes spatial layout | May generate unreachable images |
| State goal | Suitable for control and verification | True state is difficult to obtain |
| Skill token | Convenient for discrete planning | Fixed capability boundaries |
| Latent goal | Flexible representation | Difficult to interpret and inspect |
4. Limits of Textual Chains of Thought
Textual CoT produces a language sequence:
Interpretation: Conditioned on observation o and task instruction l, the language model generates a textual reasoning sequence y of length M.
Derivation: An autoregressive model factorizes this joint distribution into a product of per-token conditional probabilities, allowing it to generate coherent plans. However, the training objective constrains only text probabilities. Unless the intermediate variables in y are supervised by real perception, actions, and outcomes, textual correctness does not imply correct physical closed-loop behavior.
It may improve task decomposition, but it does not automatically perform:
- Visual object grounding.
- 3D reachability checks.
- Contact and mechanics checks.
- Post-execution state updates.
- Failure recovery.
Embodied reasoning must constrain intermediate reasoning variables through physical observations and execution outcomes.
5. World Models Provide Counterfactuals
For a candidate subgoal or skill , the world model predicts the future:
Interpretation: Given the current high-level state z_k and candidate subgoal g_k, the skill-level world model predicts the distribution of the next high-level state after subgoal execution terminates.
Derivation: Marginalizing over the low-level actions used to achieve g_k, environmental stochasticity, and execution duration yields the transition between adjacent high-level decision points. This is suitable for comparing which subgoal should be executed first, but prediction quality must be validated on the distribution actually selected by the planner.
The high-level planner can ask, “Should the drawer be opened first, or should the cup be picked up first?” and compare candidate futures. The model must learn at the skill timescale and does not necessarily need to predict every low-level action.
6. Value and Subgoal Ranking
Interpretation: The value of selecting subgoal g in high-level state z equals the reward accumulated while completing that subgoal plus the conditional expectation of the terminal-state value, discounted according to the actual duration.
Derivation: This is the Semi-MDP Bellman decomposition: a single high-level decision spans tau_g low-level time steps, so the subsequent value must be multiplied by gamma raised to the power of tau_g, with the expectation taken over low-level execution outcomes and stochastic durations.
Value estimates the long-term return after a subgoal is completed. High-level planning can combine a world model with value:
Interpretation: From the current set of feasible subgoals, select the subgoal that maximizes the sum of the skill reward and subsequent value predicted by the world model.
Derivation: First use preconditions to identify the feasible set G(z_k), then use the world model to generate a terminal-state distribution for each candidate, and finally rank them by immediate progress and long-term value. Without feasibility filtering, the value model may assign a high score to a subgoal that appears highly rewarding but cannot be executed.
7. Completion Detection
A completion detector:
Interpretation: Given the current subgoal, visual observations up to time t, proprioceptive or force-sensing state, and executed actions, the completion detector estimates the probability that the subgoal has been completed.
Derivation: Defining completion as a binary variable d_t is clearer than expressing it as a probability over the subgoal itself. Observation history is used to determine state changes; q represents proprioceptive or force-sensing information; and action history helps distinguish “not yet attempted” from “attempted but failed.” The threshold should be calibrated according to the costs of switching too early versus too late.
Switching too early causes the next skill to begin before its preconditions are satisfied; switching too late wastes time and may disrupt the achieved result. Completion detection should be jointly trained with skill-termination and task-stage data.
8. Closed-Loop Replanning
Instead of generating a complete plan only once, the high-level planner:
- Generates the current subgoal.
- Executes the low-level policy for a period.
- Checks for completion, failure, and environmental changes.
- Updates memory and the task graph.
- Selects the next subgoal or a recovery path.
This is similar to MPC, except that the planning variables are subgoals or skills rather than continuous actions.
9. Recovery Strategies
| Failure | Recovery Level | Example |
|---|---|---|
| Action deviation | Low-level policy | Realign with the grasp point |
| Skill failure | Retry or replace the skill | Use a different grasp |
| Violated precondition | High-level replanning | First set the object upright |
| Infeasible task | Request human assistance or terminate | Target object is missing |
10. Visual Subgoals
After generating a future goal image , the low-level policy is:
Interpretation: The low-level action is sampled from a policy distribution conditioned jointly on the current image and the goal image.
Derivation: The goal image o_g describes the desired visible outcome, and the policy learns to reduce task-relevant differences between the current and goal observations. If the goal image has inconsistent object identity, is unreachable, or lacks contact information, reducing visual dissimilarity does not guarantee completion of the physical task.
A visual goal must preserve object identity, be geometrically reachable and physically plausible, and permit reliable completion detection. π0.7 can serve as a cross-cutting case study connecting this approach with VLAs and world models.
11. Minimal Experiment: Do Closed-Loop Subgoals Actually Improve Long-Horizon Tasks?
Choose a five-stage tabletop or kitchen task that can be reset precisely, and hold out one object layout and one subgoal ordering during training. Keep the low-level skills, visual backbone, amount of training data, number of model calls, and control budget fixed. Compare five systems: a one-shot textual plan, a rule-based task graph, closed-loop textual replanning, a world model with value-based ranking, and visual subgoals.
Inject object displacement, grasp slippage, target occupancy, and completion-detection noise during the second or third stage. Report whole-task success rate, mean number of stages completed, illegal-subgoal rate, completion-detection precision and recall, number of plan modifications, local recovery rate, recovery time, success rate on unseen compositions, and results under the same inference budget.
- Select a task containing five stages.
- Compare a one-shot textual plan with closed-loop subgoals.
- Add a world-model-based reachability check.
- Add value-based ranking.
- Artificially perturb an intermediate stage.
- Report task success, stage success, recovery rate, and the number of plan modifications.
12. Exercises
- Convert a natural-language task into a task graph.
- Define preconditions and completion conditions for each subgoal.
- Construct an example in which a textual plan is reasonable but physically infeasible.
- Compare action-level MPC with skill-level replanning.
- Design a reachability check for visual subgoals.
- Explain which closed-loop variables an “embodied chain of thought” must include.
13. Major Failure Modes
| Failure | Manifestation | Diagnosis and Correction |
|---|---|---|
| Unobservable subgoal | The planner outputs states such as “tidied up” or “processing complete” that cannot be verified from sensor data | Rewrite the goal in terms of objects, relations, contacts, or task predicates, and explicitly specify the detector inputs |
| Missing precondition | The system is asked to retrieve an object from a closed drawer or begins transporting an object before securing the grasp | Record the illegal-subgoal rate and filter subgoals using a task graph or reachability model |
| Hallucinated textual plan | The language steps are reasonable, but the object does not exist, its location is wrong, or the action cannot be executed | Ground every entity to a visual instance and compare against a non-textual planner with identical state feedback |
| Exploitation of the world model | Search identifies a subgoal that the model predicts will yield high returns but that fails in the real world | Report planning-distribution error, model uncertainty, and verification through real-world rollouts |
| Value shortcut | The high-level planner pursues easy local rewards and undermines the final task | Inspect reward shaping, terminal success, and subgoal-ranking ablations |
| Completion-detection false positive | The system switches to the next stage before its preconditions are satisfied | Report precision, recall, threshold curves, and the cost of premature switching |
| Completion-detection false negative | The system repeatedly manipulates an already completed target, causing drops or collisions | Inspect temporal labels, sensor latency, and hysteresis mechanisms |
| Unreachable visual goal | The goal image looks plausible but violates geometric, identity, or contact constraints | Add object-consistency checks, inverse-dynamics reachability checks, and real closed-loop verification |
| Recovery loop | After a failure, the system repeatedly transitions among the same set of subgoals | Record visitation history, failure causes, recovery budget, and conditions for escalation to a higher level |
14. Paper Facts, Authors’ Interpretations, and Course Assessments
| Work | Paper Fact | Authors’ Interpretation | Course Assessment |
|---|---|---|---|
| SayCan | Combines a language model’s prior over skill sequences with robot affordance or value scores for ranking | Combining language knowledge with executability can support long-horizon tasks | It depends on predefined skills and corresponding values; skill-library coverage, completion detection, and recovery should be evaluated separately |
| Inner Monologue | Writes environmental feedback, such as scene descriptions and success detection, back into the language context and replans repeatedly | Continuous environmental feedback can improve long-horizon plan execution | The core evidence concerns closed-loop feedback, not text length; it must be compared against structured-state planning with the same information and budget |
| Code as Policies | Uses a language model to generate programmatic policies that invoke perception and control APIs | Code structure can compose existing robot capabilities and express spatial logic | Program interpretability does not imply physical correctness; API preconditions, exception handling, and real execution feedback still determine reliability |
| VoxPoser | Uses language models and vision-language models to construct 3D value maps that guide motion planning for manipulation tasks | Language semantics can be grounded into continuous control through spatial value fields | It demonstrates a bridge from symbols to geometry; contact dynamics, occlusion, and value-map calibration still require independent validation |
15. Cross-Reading
Physical Causality and Embodied Chains of Thought
D0|Hierarchical Planning, Skills, and Memory
D1|Options, Skills, and Hierarchical Reinforcement Learning
B2|Model-Based Planning: MPC, CEM, and Trajectory Optimization

