Skip to content

Original Feishu Document · Source Revision 26

🔁

Core assessment: Physical causality changes the structure of the world represented by the model, while embodied chains of thought change how the model uses those structures during execution. Genuine embodied reasoning does not consist of generating linguistic explanations; it completes the physical closed loop of “observe—predict—act—verify—correct.”

This article is a thematic extension of From Robot Control to Flow Matching: The Mathematical Foundations of the π Series.

1. From Conditional Imitation to Causal Intervention ​

Standard behavior cloning learns:

How to read it: The policy pi theta defines a probability distribution over action chunk A_t, conditioned on the current observation o_t and language instruction l.

Derivation: Behavior cloning treats action chunks in expert data as supervised labels and maximizes the conditional likelihood of expert actions. The policy therefore learns “what the expert typically does upon seeing this observation and instruction” within the training distribution.

It answers: Under similar observations and task conditions, what action does the expert typically execute?

After introducing physical causality, the model must also learn how actions change the world:

How to read it: Given the current state s_t and the deliberate application of action a_t, the causal model p_phi defines a distribution over the next state s_{t+1}.

Derivation: The conditional distribution may incorporate selection bias concerning “the states in which the expert would choose this action.” The do operator treats the action as an external intervention and requires the model to answer what would happen if a different action were taken in the same state. Randomized robot trials or controllable simulations can directly provide this type of interventional data.

Here, indicates that the action is an intervention deliberately applied to the environment by the robot, rather than a correlated variable that happens to be observed in the data.

The core change is that the policy no longer merely fits actions; it must account for state transitions, task utility, risk, and uncertainty when selecting actions.

2. What Is Physical Causality in Robotics? ​

Physical causality in robotics is not abstract philosophy. It consists of three experimentally testable questions:

  1. Intervention: After the robot executes an action, does the environment change as predicted?
  2. Mechanism: Is the state change produced through contact, force, friction, constraints, or object relations?
  3. Counterfactual: Would the outcome differ if the action, contact point, or object properties were different?

Structured Dynamics ​

A real system can be written as:

How to read it: The state at the next time step equals the dynamics function F applied to the current state, action, and unmodeled factors xi_t.

Derivation: The next state of a physical system depends not only on what the robot does, but also on the current state and disturbances such as friction, latency, and external forces. Collecting these incompletely observed factors into xi_t yields the most basic controlled state-transition equation.

Here, includes incompletely observed friction, mass, compliance, control latency, and external disturbances.

An object-centric formulation can be written as:

How to read it: The next state of object j is jointly determined by its current state, the robot action, contact relations, physical properties, and disturbances.

Derivation: After decomposing the global state into object-level states, whether the same action changes an object depends on contact and on properties such as the object’s mass, friction, and material. This object-centric decomposition expresses “which action changed which object through what contact” more readily than simply predicting the entire image.

Here:

  • : the position, pose, and state of object ;
  • : the contact relation between the robot and the object;
  • : physical properties such as mass, friction, and material.

The Difference Between Correlation and Causation ​

Observing that “the cup also moves upward when the hand moves upward” is only a correlation. A causal model must additionally distinguish among the following cases:

  • Without contact, moving the hand upward will not move the cup;
  • If the grip force is insufficient, the cup may slip;
  • Different contact points will change the cup’s rotation;
  • A change in the cup’s mass will change the required force and velocity.

📌

Minimum criterion for physical causality: The model must be able to predict the outcomes of interventions and modify its predictions accordingly when physical conditions change, rather than merely restating co-occurrence patterns in the training data.

3. What Is an Embodied Chain of Thought? ​

An embodied chain of thought does not translate internal reasoning into a passage of text. Instead, the robot repeatedly executes the following loop in the physical environment:

State Memory ​

The robot maintains a belief state:

How to read it: The current belief state b_t is updated from the previous belief, the current observation, and the previous action.

Derivation: A single-frame observation cannot directly reveal occluded objects, contact stability, or task progress. The system therefore uses the previous belief to retain history, the executed action to predict changes, and the new observation to correct that prediction. Update may be a Bayesian filter, a recurrent network, or a Transformer.

may include object states, contact states, task stages, action history, and uncertainty.

Subgoal Reasoning ​

How to read it: Based on the current belief state b_t and language task l, the planner selects the current subgoal g_t.

Derivation: Long-horizon tasks cannot rely solely on the original instruction. The planner must combine the completed stages, current physical state, and uncertainty to reduce the global task to a local goal that is currently executable and verifiable.

For example, “place the cup on the tray” can be decomposed into:

text
Locate cup
→ Select grasp region
→ Establish stable contact
→ Lift and avoid obstacles
→ Move to tray
→ Place
→ Verify cup_on_tray

Closed-Loop Actions ​

How to read it: Action a_t is sampled from a policy distribution conditioned on the current belief state and subgoal.

Derivation: High-level planning specifies the local outcome to be achieved, while the low-level policy generates an action based on the current state. Writing this as a point distribution of pi_theta avoids incorrectly treating the already sampled a_t as an input variable of the distribution.

After every execution step, the robot must inspect the actual outcome rather than assume that the plan has succeeded.

4. How the Two Jointly Change the Policy Model ​

Physical causality defines “how the world responds to actions,” while the embodied chain of thought defines “how the robot repeatedly uses these response patterns.” When combined, they transform the policy from a direct mapping into model-assisted decision-making:

How to read it: Among candidate action sequences of horizon H, select the one with the highest expected task utility predicted by the world model minus its risk cost.

Derivation: For each candidate action sequence, the world model produces a distribution over future states. Utility U measures progress toward the subgoal, while C measures collision, energy, force, and other risks. Taking the expectation over uncertain predicted futures and then maximizing net utility yields model-based action selection.

Here:

  • : an action-conditioned world model or object-state transition model;
  • : utility from task progress and eventual success;
  • : costs associated with collisions, energy consumption, velocity, torque, and risk.

The model architecture typically adds four components:

  1. An object-centric state or belief encoder;
  2. An action-conditioned state-transition model;
  3. A subgoal or task-graph module;
  4. An outcome-verification and failure-recovery module.

It does not necessarily require a complete pixel-level world model. For robot control, predicting “whether contact occurs, whether the object moves, and whether task predicates change” is often more direct than generating photorealistic video.

5. Implications for Data, Tokenizers, and Training Objectives ​

Changes in the Data ​

Standard demonstration dataCausal and embodied data
Successful trajectoriesSuccesses, failures, disturbances, and recoveries
Observations and actionsObservations, actions, contacts, state changes, and outcomes
Fixed environmentsDeliberate variation of mass, friction, position, and object instances
Action labelsAction preconditions, expected consequences, and completion conditions

Changes in the Tokenizer ​

Encoding only human or robot motion:

text
[MOVE_LEFT] [ROTATE_WRIST] [CLOSE_GRIPPER]

is upgraded to object-centric causal events:

text
[APPROACH cup]
[CONTACT cup stable]
[GRASP cup]
[CAUSE cup_follows_hand]
[VERIFY cup_lifted]
[FAIL slip]
[RECOVER regrasp]

Tokens represent not only “what was done,” but also:

  • The preconditions under which the action occurred;
  • The object on which the action operated;
  • The state change the action was expected to cause;
  • Whether the outcome matched expectations;
  • The recovery state entered after failure.

Changes in the Training Objective ​

How to read it: The total loss is a weighted sum of five terms: behavior cloning, dynamics prediction, physical events, subgoal completion, and failure recovery.

Derivation: Behavior cloning supervises only “what to do.” The remaining auxiliary objectives respectively supervise “what will happen, what event occurred, whether the goal was completed, and how to continue after failure.” Each lambda is a nonnegative weight balanced using a validation set or gradient scales; this equation should not be interpreted as requiring every system to use the same fixed set of weights.

  • : imitate expert actions;
  • : predict state changes after an action;
  • : predict contact, grasp, and release events;
  • : determine whether subgoals and task predicates have been completed;
  • : recover from failure states.

6. Why a Textual Chain of Thought Is Not Equivalent to Embodied Reasoning ​

A language model may generate:

The cup may slip, so the velocity should be reduced and the grip force increased.

This statement sounds reasonable, but it constitutes embodied reasoning only when the following information actually enters the control loop:

  • The model can observe or estimate contact and slip;
  • “Increase the grip force” can be mapped to controls executable by the robot;
  • After execution, the robot checks again whether the cup is stable;
  • The robot replans if the outcome does not match expectations.
Textual chain of thoughtEmbodied chain of thought
Outputs linguistic stepsMaintains executable states and subgoals
May be a post hoc explanationPredicts before acting and verifies afterward
Does not require real feedbackMust receive visual, tactile, or force feedback
Errors do not necessarily affect the next stepPrediction errors must trigger state updates and recovery
Linguistic correctness is sufficientPhysical task success is the ultimate metric

❗

Decision rule: If removing the generated text leaves the robot’s actions completely unchanged, the CoT is likely only an explanatory layer rather than part of the control mechanism.

7. How to Verify That the Model Has Truly Learned Physical Causality ​

Physical causality cannot be demonstrated using action-prediction error alone. Variables in the training distribution must be actively intervened upon, and the model must be checked for mechanism-consistent changes.

Intervention experimentExpected changeSpurious correlation ruled out
Change object massAdjust velocity, acceleration, or grasping strategySelecting actions solely by appearance
Change frictionChange grip force, contact duration, or grasp typeTreating a fixed trajectory as a mechanism
Move the contact pointPredict changes in object translation and rotationMerely memorizing hand trajectories
Change appearance while preserving physical propertiesKeep behavior largely unchangedLeakage from backgrounds and colors
Preserve appearance while changing internal weightUpdate the belief based on interaction feedbackRelying only on the first visual frame
Apply a disturbance during executionDetect deviation and recoverOpen-loop playback of action sequences

Metrics That Should Be Reported ​

  • Action-consequence prediction error;
  • Accuracy of object states and contact events;
  • Complete task success rate;
  • Recovery success rate and recovery time after disturbances;
  • OOD generalization across physical properties;
  • Policy sensitivity before and after interventions on causal variables;
  • Policy stability under irrelevant visual changes.

The minimal counterfactual test is to preserve similar image semantics while changing only one physical factor and check whether the model changes its action. Then preserve the physical factors while changing only the appearance and check whether the model preserves its action.

8. Relationship to WAM-TTT, World Models, and Behavior Tokenizers ​

MethodPrimary problem addressedWhat remains missing
WAM-TTTWrites human videos into fast-weight memory at test timeExplicit object causality, contact, and interpretable recovery states
Pixel world modelPredicts future visuals or videoVisual realism does not guarantee physically correct actions and contacts
Object world modelPredicts changes in object states, relations, and eventsRequires reliable object and contact representations
Behavior TokenizerCompresses behavior into composable events and skillsRequires action consequences and physical verification to acquire causal semantics
Embodied chain of thoughtOrganizes observation, subgoals, prediction, action, and recoveryMust be integrated into a closed loop with real control and feedback

They can be combined as:

text
Human and robot multimodal data
→ Object / Contact / Event Tokenizer
→ Object-centric causal world model
→ Embodied task graph and subgoals
→ Action Expert
→ Visual / force feedback
→ Verification and recovery

WAM-TTT is more closely concerned with “how to rapidly absorb memories of human behavior.” Physical causality and embodied chains of thought further address “why an action works, whether the outcome matches expectations, and how to correct failures.”

9. A Minimal Implementable System ​

A complete physical causal system is expensive. A pragmatic minimal version does not need to generate high-fidelity video or reconstruct every mechanical parameter first.

Minimal System ​

  1. Construct an object-centric belief state from visual and proprioceptive signals;
  2. Recognize approach, contact, grasp, release, state-change, and failure;
  3. Predict whether a candidate action will change the state of the target object;
  4. Use task predicates to select the current subgoal;
  5. Execute a short action chunk and inspect the outcome;
  6. Upon predicting failure, perceive again, regrasp, or replan.

The corresponding minimal closed loop is:

How to read it: The system derives a subgoal from the current belief, predicts the respective next state for each candidate action, selects action a_t, observes the actual outcome, and then updates the next belief.

Derivation: Next-state prediction must be conditioned on candidate actions. The system cannot produce a single unique hat s_{t+1} before an action has been specified. A closed loop in which causal predictions enter control decisions is formed only by pairing each candidate action with its predicted outcome, comparing the pairs, and executing one of them.

✅

Section summary: Physical causality moves robots from “what the expert typically does” to “what my action will cause.” Embodied chains of thought move robots from “outputting an action once” to “continually observing, predicting, executing, verifying, and recovering.” They truly become part of robot intelligence only when predictions and real physical feedback jointly change subsequent actions.

10. Major Failure Modes ​

FailureSurface symptomEvidence that actually needs to be examined
Mistaking correlation for causationAction prediction is accurate within the training distribution but fails immediately after changes in mass or frictionSingle-variable physical interventions and counterfactual prediction error
Actions do not enter the predictionDifferent candidate actions yield nearly identical futuresAction permutation, action ablation, and sensitivity curves
Visual shortcutsActions are selected based on color, background, or operator identityPaired tests that preserve physics while changing appearance and change physics while keeping appearance similar
Incorrect state estimationPlanning logic is reasonable but based on the wrong object, occlusion state, or contact assessmentBelief calibration, object-binding accuracy, and contact-event detection
Exploitation of the model by the plannerThe model’s average error is small, but search discovers unrealistic high-utility futuresError on the planning distribution, verification with real rollouts, and uncertainty constraints
Linguistic explanations are decoupled from controlThe chain of thought is comprehensive, but removing or shuffling the text does not change the actionsText ablations and state-feedback controls under equal control budgets
Completion-detection driftThe system advances task stages prematurely or repeatedly acts on already completed stagesTask-predicate precision, recall, and transition-timing error
Recovery policy can only restartA local failure resets the entire task, causing long-horizon success rates to decline rapidlyLocal recovery rate, rollback level, and recovery cost

11. Minimal Experiment: Does the Model Actually Use Physical Causality? ​

Choose a repeatable pushing, grasping, or pouring task. In the training data, make object appearance highly correlated with physical properties—for example, red objects are usually light and blue objects are usually heavy. At test time, construct four sets of paired conditions.

  1. Same appearance, different physics: Keep the image and task semantics approximately unchanged while varying only mass, friction, or the contact point.
  2. Same physics, different appearance: Keep the dynamics parameters unchanged while varying only color, texture, and background.
  3. Action intervention: Execute different actions from the same initial state and compare predicted consequences with actual consequences.
  4. Closed-loop disturbance: Move the object or induce slipping midway through execution, then check whether the system updates its belief, reselects a subgoal, and recovers.

Compare four types of systems: pure behavior cloning, a policy with an auxiliary dynamics loss, a policy that uses a world model without search, and a closed-loop system that uses a world model to compare candidate actions. Fix the amount of training data, visual backbone, policy capacity, control frequency, and inference budget.

Minimum reporting requirements: Action success rate, intervention-consequence prediction error, action sensitivity under changes in physical properties, action invariance under appearance changes, model calibration error, disturbance-recovery rate, and model error on both the planning distribution and the standard test distribution.

Decision criterion: Evidence that physical causality has entered policy decision-making exists only when the model changes its predictions and actions according to the mechanism as physical factors change, preserves its behavior under appearance-only changes, and these predictions actually improve closed-loop outcomes.

12. Paper Facts, Author Interpretations, and Course Assessments ​

WorkPaper factAuthor interpretationCourse assessment
CausalWorldProvides a robotic manipulation environment with controllable variables such as mass, friction, and size for evaluating causal structure and transferSystematic interventions help study structural generalization rather than merely fitting a fixed distributionIt provides a valid causal stress test, but controllable simulation variables do not mean that causal identification has been solved on real robots
SayCanCombines a language model’s skill priors with robotic affordance or value scores to select executable skillsJoint ranking based on linguistic knowledge and physical feasibility can support long-horizon tasksIt demonstrates composition and grounding over a given skill library, not the learning of continuous dynamics or automatic discovery of causal mechanisms
Inner MonologueWrites environmental feedback, such as success detection and scene descriptions, back into the language-planning context and replans repeatedlyClosed-loop environmental feedback can improve long-horizon plan executionThe key contribution is that feedback enters replanning; whether natural language is necessary should be compared against a non-text planner receiving equivalent state feedback
World-model approachesLearn action-conditioned transitions that can be used for imagination-based training or online planningPredicting the future can improve data efficiency and decision qualityHigh pixel fidelity is not sufficient for physical causality; mechanism correctness on the planning distribution and gains in real closed-loop performance matter more

13. Exercises and Cross-Reading for This Lesson ​

  1. Explain why high action-prediction accuracy still does not prove that a model has learned the causal effects of actions.
  2. Draw a causal graph for a grasping task that includes states, actions, contacts, physical properties, and outcomes, and identify possible confounders.
  3. Design one paired test with “the same appearance but different physics” and another with “the same physics but different appearance.”
  4. Give an example in which a textual chain of thought appears correct but would fail under real control, and design an ablation experiment that would reveal the failure.
  5. Compare the types of physical questions best addressed by pixel world models, object-centric world models, and task-predicate models.

B0|World Models and Model-Based Planning

D0|Hierarchical Planning, Skills, and Memory

D2 | Subgoal Planning and Embodied Reasoning

E2 | Latent Action and Inverse Dynamics

F0 | Dynamics, Control, and Physical Interaction

Article text is licensed under the Apache License 2.0