Original Feishu Document · Source Revision 26
🔁
Core assessment: Physical causality changes the structure of the world represented by the model, while embodied chains of thought change how the model uses those structures during execution. Genuine embodied reasoning does not consist of generating linguistic explanations; it completes the physical closed loop of “observe—predict—act—verify—correct.”
This article is a thematic extension of From Robot Control to Flow Matching: The Mathematical Foundations of the π Series.
1. From Conditional Imitation to Causal Intervention
Standard behavior cloning learns:
How to read it: The policy pi theta defines a probability distribution over action chunk A_t, conditioned on the current observation o_t and language instruction l.
Derivation: Behavior cloning treats action chunks in expert data as supervised labels and maximizes the conditional likelihood of expert actions. The policy therefore learns “what the expert typically does upon seeing this observation and instruction” within the training distribution.
It answers: Under similar observations and task conditions, what action does the expert typically execute?
After introducing physical causality, the model must also learn how actions change the world:
How to read it: Given the current state s_t and the deliberate application of action a_t, the causal model p_phi defines a distribution over the next state s_{t+1}.
Derivation: The conditional distribution may incorporate selection bias concerning “the states in which the expert would choose this action.” The do operator treats the action as an external intervention and requires the model to answer what would happen if a different action were taken in the same state. Randomized robot trials or controllable simulations can directly provide this type of interventional data.
Here, indicates that the action is an intervention deliberately applied to the environment by the robot, rather than a correlated variable that happens to be observed in the data.
The core change is that the policy no longer merely fits actions; it must account for state transitions, task utility, risk, and uncertainty when selecting actions.
2. What Is Physical Causality in Robotics?
Physical causality in robotics is not abstract philosophy. It consists of three experimentally testable questions:
- Intervention: After the robot executes an action, does the environment change as predicted?
- Mechanism: Is the state change produced through contact, force, friction, constraints, or object relations?
- Counterfactual: Would the outcome differ if the action, contact point, or object properties were different?
Structured Dynamics
A real system can be written as:
How to read it: The state at the next time step equals the dynamics function F applied to the current state, action, and unmodeled factors xi_t.
Derivation: The next state of a physical system depends not only on what the robot does, but also on the current state and disturbances such as friction, latency, and external forces. Collecting these incompletely observed factors into xi_t yields the most basic controlled state-transition equation.
Here, includes incompletely observed friction, mass, compliance, control latency, and external disturbances.
An object-centric formulation can be written as:
How to read it: The next state of object j is jointly determined by its current state, the robot action, contact relations, physical properties, and disturbances.
Derivation: After decomposing the global state into object-level states, whether the same action changes an object depends on contact and on properties such as the object’s mass, friction, and material. This object-centric decomposition expresses “which action changed which object through what contact” more readily than simply predicting the entire image.
Here:
- : the position, pose, and state of object ;
- : the contact relation between the robot and the object;
- : physical properties such as mass, friction, and material.
The Difference Between Correlation and Causation
Observing that “the cup also moves upward when the hand moves upward” is only a correlation. A causal model must additionally distinguish among the following cases:
- Without contact, moving the hand upward will not move the cup;
- If the grip force is insufficient, the cup may slip;
- Different contact points will change the cup’s rotation;
- A change in the cup’s mass will change the required force and velocity.
📌
Minimum criterion for physical causality: The model must be able to predict the outcomes of interventions and modify its predictions accordingly when physical conditions change, rather than merely restating co-occurrence patterns in the training data.
3. What Is an Embodied Chain of Thought?
An embodied chain of thought does not translate internal reasoning into a passage of text. Instead, the robot repeatedly executes the following loop in the physical environment:
State Memory
The robot maintains a belief state:
How to read it: The current belief state b_t is updated from the previous belief, the current observation, and the previous action.
Derivation: A single-frame observation cannot directly reveal occluded objects, contact stability, or task progress. The system therefore uses the previous belief to retain history, the executed action to predict changes, and the new observation to correct that prediction. Update may be a Bayesian filter, a recurrent network, or a Transformer.
may include object states, contact states, task stages, action history, and uncertainty.
Subgoal Reasoning
How to read it: Based on the current belief state b_t and language task l, the planner selects the current subgoal g_t.
Derivation: Long-horizon tasks cannot rely solely on the original instruction. The planner must combine the completed stages, current physical state, and uncertainty to reduce the global task to a local goal that is currently executable and verifiable.
For example, “place the cup on the tray” can be decomposed into:
Locate cup
→ Select grasp region
→ Establish stable contact
→ Lift and avoid obstacles
→ Move to tray
→ Place
→ Verify cup_on_trayClosed-Loop Actions
How to read it: Action a_t is sampled from a policy distribution conditioned on the current belief state and subgoal.
Derivation: High-level planning specifies the local outcome to be achieved, while the low-level policy generates an action based on the current state. Writing this as a point distribution of pi_theta avoids incorrectly treating the already sampled a_t as an input variable of the distribution.
After every execution step, the robot must inspect the actual outcome rather than assume that the plan has succeeded.
4. How the Two Jointly Change the Policy Model
Physical causality defines “how the world responds to actions,” while the embodied chain of thought defines “how the robot repeatedly uses these response patterns.” When combined, they transform the policy from a direct mapping into model-assisted decision-making:
How to read it: Among candidate action sequences of horizon H, select the one with the highest expected task utility predicted by the world model minus its risk cost.
Derivation: For each candidate action sequence, the world model produces a distribution over future states. Utility U measures progress toward the subgoal, while C measures collision, energy, force, and other risks. Taking the expectation over uncertain predicted futures and then maximizing net utility yields model-based action selection.
Here:
- : an action-conditioned world model or object-state transition model;
- : utility from task progress and eventual success;
- : costs associated with collisions, energy consumption, velocity, torque, and risk.
The model architecture typically adds four components:
- An object-centric state or belief encoder;
- An action-conditioned state-transition model;
- A subgoal or task-graph module;
- An outcome-verification and failure-recovery module.
It does not necessarily require a complete pixel-level world model. For robot control, predicting “whether contact occurs, whether the object moves, and whether task predicates change” is often more direct than generating photorealistic video.
5. Implications for Data, Tokenizers, and Training Objectives
Changes in the Data
| Standard demonstration data | Causal and embodied data |
|---|---|
| Successful trajectories | Successes, failures, disturbances, and recoveries |
| Observations and actions | Observations, actions, contacts, state changes, and outcomes |
| Fixed environments | Deliberate variation of mass, friction, position, and object instances |
| Action labels | Action preconditions, expected consequences, and completion conditions |
Changes in the Tokenizer
Encoding only human or robot motion:
[MOVE_LEFT] [ROTATE_WRIST] [CLOSE_GRIPPER]is upgraded to object-centric causal events:
[APPROACH cup]
[CONTACT cup stable]
[GRASP cup]
[CAUSE cup_follows_hand]
[VERIFY cup_lifted]
[FAIL slip]
[RECOVER regrasp]Tokens represent not only “what was done,” but also:
- The preconditions under which the action occurred;
- The object on which the action operated;
- The state change the action was expected to cause;
- Whether the outcome matched expectations;
- The recovery state entered after failure.
Changes in the Training Objective
How to read it: The total loss is a weighted sum of five terms: behavior cloning, dynamics prediction, physical events, subgoal completion, and failure recovery.
Derivation: Behavior cloning supervises only “what to do.” The remaining auxiliary objectives respectively supervise “what will happen, what event occurred, whether the goal was completed, and how to continue after failure.” Each lambda is a nonnegative weight balanced using a validation set or gradient scales; this equation should not be interpreted as requiring every system to use the same fixed set of weights.
- : imitate expert actions;
- : predict state changes after an action;
- : predict contact, grasp, and release events;
- : determine whether subgoals and task predicates have been completed;
- : recover from failure states.
6. Why a Textual Chain of Thought Is Not Equivalent to Embodied Reasoning
A language model may generate:
The cup may slip, so the velocity should be reduced and the grip force increased.
This statement sounds reasonable, but it constitutes embodied reasoning only when the following information actually enters the control loop:
- The model can observe or estimate contact and slip;
- “Increase the grip force” can be mapped to controls executable by the robot;
- After execution, the robot checks again whether the cup is stable;
- The robot replans if the outcome does not match expectations.
| Textual chain of thought | Embodied chain of thought |
|---|---|
| Outputs linguistic steps | Maintains executable states and subgoals |
| May be a post hoc explanation | Predicts before acting and verifies afterward |
| Does not require real feedback | Must receive visual, tactile, or force feedback |
| Errors do not necessarily affect the next step | Prediction errors must trigger state updates and recovery |
| Linguistic correctness is sufficient | Physical task success is the ultimate metric |
❗
Decision rule: If removing the generated text leaves the robot’s actions completely unchanged, the CoT is likely only an explanatory layer rather than part of the control mechanism.
7. How to Verify That the Model Has Truly Learned Physical Causality
Physical causality cannot be demonstrated using action-prediction error alone. Variables in the training distribution must be actively intervened upon, and the model must be checked for mechanism-consistent changes.
| Intervention experiment | Expected change | Spurious correlation ruled out |
|---|---|---|
| Change object mass | Adjust velocity, acceleration, or grasping strategy | Selecting actions solely by appearance |
| Change friction | Change grip force, contact duration, or grasp type | Treating a fixed trajectory as a mechanism |
| Move the contact point | Predict changes in object translation and rotation | Merely memorizing hand trajectories |
| Change appearance while preserving physical properties | Keep behavior largely unchanged | Leakage from backgrounds and colors |
| Preserve appearance while changing internal weight | Update the belief based on interaction feedback | Relying only on the first visual frame |
| Apply a disturbance during execution | Detect deviation and recover | Open-loop playback of action sequences |
Metrics That Should Be Reported
- Action-consequence prediction error;
- Accuracy of object states and contact events;
- Complete task success rate;
- Recovery success rate and recovery time after disturbances;
- OOD generalization across physical properties;
- Policy sensitivity before and after interventions on causal variables;
- Policy stability under irrelevant visual changes.
The minimal counterfactual test is to preserve similar image semantics while changing only one physical factor and check whether the model changes its action. Then preserve the physical factors while changing only the appearance and check whether the model preserves its action.
8. Relationship to WAM-TTT, World Models, and Behavior Tokenizers
| Method | Primary problem addressed | What remains missing |
|---|---|---|
| WAM-TTT | Writes human videos into fast-weight memory at test time | Explicit object causality, contact, and interpretable recovery states |
| Pixel world model | Predicts future visuals or video | Visual realism does not guarantee physically correct actions and contacts |
| Object world model | Predicts changes in object states, relations, and events | Requires reliable object and contact representations |
| Behavior Tokenizer | Compresses behavior into composable events and skills | Requires action consequences and physical verification to acquire causal semantics |
| Embodied chain of thought | Organizes observation, subgoals, prediction, action, and recovery | Must be integrated into a closed loop with real control and feedback |
They can be combined as:
Human and robot multimodal data
→ Object / Contact / Event Tokenizer
→ Object-centric causal world model
→ Embodied task graph and subgoals
→ Action Expert
→ Visual / force feedback
→ Verification and recoveryWAM-TTT is more closely concerned with “how to rapidly absorb memories of human behavior.” Physical causality and embodied chains of thought further address “why an action works, whether the outcome matches expectations, and how to correct failures.”
9. A Minimal Implementable System
A complete physical causal system is expensive. A pragmatic minimal version does not need to generate high-fidelity video or reconstruct every mechanical parameter first.
Minimal System
- Construct an object-centric belief state from visual and proprioceptive signals;
- Recognize approach, contact, grasp, release, state-change, and failure;
- Predict whether a candidate action will change the state of the target object;
- Use task predicates to select the current subgoal;
- Execute a short action chunk and inspect the outcome;
- Upon predicting failure, perceive again, regrasp, or replan.
The corresponding minimal closed loop is:
How to read it: The system derives a subgoal from the current belief, predicts the respective next state for each candidate action, selects action a_t, observes the actual outcome, and then updates the next belief.
Derivation: Next-state prediction must be conditioned on candidate actions. The system cannot produce a single unique hat s_{t+1} before an action has been specified. A closed loop in which causal predictions enter control decisions is formed only by pairing each candidate action with its predicted outcome, comparing the pairs, and executing one of them.
✅
Section summary: Physical causality moves robots from “what the expert typically does” to “what my action will cause.” Embodied chains of thought move robots from “outputting an action once” to “continually observing, predicting, executing, verifying, and recovering.” They truly become part of robot intelligence only when predictions and real physical feedback jointly change subsequent actions.
10. Major Failure Modes
| Failure | Surface symptom | Evidence that actually needs to be examined |
|---|---|---|
| Mistaking correlation for causation | Action prediction is accurate within the training distribution but fails immediately after changes in mass or friction | Single-variable physical interventions and counterfactual prediction error |
| Actions do not enter the prediction | Different candidate actions yield nearly identical futures | Action permutation, action ablation, and sensitivity curves |
| Visual shortcuts | Actions are selected based on color, background, or operator identity | Paired tests that preserve physics while changing appearance and change physics while keeping appearance similar |
| Incorrect state estimation | Planning logic is reasonable but based on the wrong object, occlusion state, or contact assessment | Belief calibration, object-binding accuracy, and contact-event detection |
| Exploitation of the model by the planner | The model’s average error is small, but search discovers unrealistic high-utility futures | Error on the planning distribution, verification with real rollouts, and uncertainty constraints |
| Linguistic explanations are decoupled from control | The chain of thought is comprehensive, but removing or shuffling the text does not change the actions | Text ablations and state-feedback controls under equal control budgets |
| Completion-detection drift | The system advances task stages prematurely or repeatedly acts on already completed stages | Task-predicate precision, recall, and transition-timing error |
| Recovery policy can only restart | A local failure resets the entire task, causing long-horizon success rates to decline rapidly | Local recovery rate, rollback level, and recovery cost |
11. Minimal Experiment: Does the Model Actually Use Physical Causality?
Choose a repeatable pushing, grasping, or pouring task. In the training data, make object appearance highly correlated with physical properties—for example, red objects are usually light and blue objects are usually heavy. At test time, construct four sets of paired conditions.
- Same appearance, different physics: Keep the image and task semantics approximately unchanged while varying only mass, friction, or the contact point.
- Same physics, different appearance: Keep the dynamics parameters unchanged while varying only color, texture, and background.
- Action intervention: Execute different actions from the same initial state and compare predicted consequences with actual consequences.
- Closed-loop disturbance: Move the object or induce slipping midway through execution, then check whether the system updates its belief, reselects a subgoal, and recovers.
Compare four types of systems: pure behavior cloning, a policy with an auxiliary dynamics loss, a policy that uses a world model without search, and a closed-loop system that uses a world model to compare candidate actions. Fix the amount of training data, visual backbone, policy capacity, control frequency, and inference budget.
Minimum reporting requirements: Action success rate, intervention-consequence prediction error, action sensitivity under changes in physical properties, action invariance under appearance changes, model calibration error, disturbance-recovery rate, and model error on both the planning distribution and the standard test distribution.
Decision criterion: Evidence that physical causality has entered policy decision-making exists only when the model changes its predictions and actions according to the mechanism as physical factors change, preserves its behavior under appearance-only changes, and these predictions actually improve closed-loop outcomes.
12. Paper Facts, Author Interpretations, and Course Assessments
| Work | Paper fact | Author interpretation | Course assessment |
|---|---|---|---|
| CausalWorld | Provides a robotic manipulation environment with controllable variables such as mass, friction, and size for evaluating causal structure and transfer | Systematic interventions help study structural generalization rather than merely fitting a fixed distribution | It provides a valid causal stress test, but controllable simulation variables do not mean that causal identification has been solved on real robots |
| SayCan | Combines a language model’s skill priors with robotic affordance or value scores to select executable skills | Joint ranking based on linguistic knowledge and physical feasibility can support long-horizon tasks | It demonstrates composition and grounding over a given skill library, not the learning of continuous dynamics or automatic discovery of causal mechanisms |
| Inner Monologue | Writes environmental feedback, such as success detection and scene descriptions, back into the language-planning context and replans repeatedly | Closed-loop environmental feedback can improve long-horizon plan execution | The key contribution is that feedback enters replanning; whether natural language is necessary should be compared against a non-text planner receiving equivalent state feedback |
| World-model approaches | Learn action-conditioned transitions that can be used for imagination-based training or online planning | Predicting the future can improve data efficiency and decision quality | High pixel fidelity is not sufficient for physical causality; mechanism correctness on the planning distribution and gains in real closed-loop performance matter more |
13. Exercises and Cross-Reading for This Lesson
- Explain why high action-prediction accuracy still does not prove that a model has learned the causal effects of actions.
- Draw a causal graph for a grasping task that includes states, actions, contacts, physical properties, and outcomes, and identify possible confounders.
- Design one paired test with “the same appearance but different physics” and another with “the same physics but different appearance.”
- Give an example in which a textual chain of thought appears correct but would fail under real control, and design an ablation experiment that would reveal the failure.
- Compare the types of physical questions best addressed by pixel world models, object-centric world models, and task-predicate models.
B0|World Models and Model-Based Planning
D0|Hierarchical Planning, Skills, and Memory
D2 | Subgoal Planning and Embodied Reasoning

