Feishu Source Catalog · Source Revision 91
💡
Intended audience: Algorithm, data, and robotics systems practitioners with knowledge of probability and statistics, linear algebra, and foundational deep learning who want to quickly build a comprehensive understanding of Physical AI. This catalog is organized not by company or paper publication date, but by the mathematical objects that models actually learn.
How to Use This Catalog
Physical AI is not a single-track technology path. Direct policies, world models, value learning, hierarchical planning, data representations, control, and systems engineering each address different problems while operating together within the same physical closed loop.
All readers should first complete the “Common Foundations” and then choose one primary track for deeper study. An article’s primary placement only indicates the best entry point; it does not mean that the article is relevant only to that track. Cross-track relationships are explicitly identified after each module.
Unified Course Quality Standards
This course is intended for readers with university-level knowledge of probability and statistics, linear algebra, and foundational deep learning. Every mechanism course and paper lab must meet the following requirements to be considered complete.
| Requirement | Must Answer |
|---|---|
| Learning object | Is the model estimating a distribution, state, value, subgoal, representation, or control law? |
| How to read equations | Every standalone equation must be immediately followed by a blockquote that explains, in plain language, the conditions, random variables, and operators term by term. |
| Derivation | Explain which probabilistic assumption, definition, physical equation, or optimization objective the equation follows from; presenting only the result is not allowed. |
| Visualization | Include at least one computational graph, probabilistic graph, closed-loop diagram, or curve that genuinely reduces the effort required to understand the material. |
| Minimal experiment | Readers must be able to reproduce the key phenomenon in a toy environment rather than simply trust the paper’s conclusions. |
| Position in the literature | Distinguish facts, the authors’ interpretations, and the course’s assessment, and compare against strong baselines from the same track. |
| Closed-loop risks | Discuss distribution shift, contact, latency, missing observations, safety, and failure recovery. |
All equations follow a unified “four-step explanation”: first present the equation, then provide a blockquoted plain-language reading, derive it, and finally explain its meaning within the robot’s closed loop.
Example equation-reading block: Rather than merely reading the symbols aloud, explain “what is conditioned on, which random quantity is being predicted, and which cases are covered by the sum or expectation.”
Inline symbols are explained consistently in each section’s notation table. Every standalone equation that defines a concept, loss, dynamics, update rule, or inference procedure must include its own blockquoted reading and derivation.
00|Common Foundations: The Shared Language of All Tracks
Core questions: What are states, observations, actions, trajectories, and policies? Why does the training distribution change during deployment? How do the loss functions of probabilistic models translate into real robot behavior?
00.1|Common Foundations Core Course
00|Physical AI Common Foundations: What Is an Agent Actually Learning?
Paper and derivation lab: An In-Depth Derivation of Behavior Cloning, Distribution Shift, and Flow Matching
Learning objects: Conditional action distributions, maximum likelihood, mean squared error, multimodal actions, action chunks, and deployment distribution shift.
00.2|Boundaries Between Physical AI Modules
| Module | Primary Mathematical Object | Question Answered |
|---|---|---|
| Policy | Conditional action distribution | What should be done now? |
| World model | Action-conditioned state-transition distribution | What will happen after doing this? |
| Value function | State value and action value | How good will the future be? |
| Planner | Candidate action sequences or subgoals | Which path should be selected? |
| Controller | Feedback control law | How can the action be executed stably? |
Track A|VLA and Direct Policy Learning
Track question: Can robot actions be generated directly from vision, language, and proprioceptive state? This track learns conditional action distributions and is currently the primary direction for general-purpose robot policies and robot foundation models.
A0|Track Core Course: VLA and Direct Policy Learning
A0|VLA and Direct Policy Learning: From Conditional Imitation to Robot Foundation Models
Cross-track relationships: Uses data and representations from Track E; can be further improved using value signals from Track C; and is ultimately executed by the control systems in Track F.
A1|Specialized Lab: Action Representations and FAST
Action Representations: MSE, Autoregressive Tokens, Diffusion, Flow Matching, and FAST
Core question: Does the policy output a conditional mean, a discrete action sequence, a score, or a probability-flow vector field?
A2|Mainstream VLA Lineages and Architectural Evolution
RT, Open X, Octo, OpenVLA, GR00T, and Gemini Robotics
Core questions: Beyond the π series, how do mainstream VLAs combine multitask imitation, vision-language pretraining, cross-embodiment data, and generative action modules? Do capability gains actually come from architecture, data, or post-training?
A3|Physical Intelligence π-Series Paper Labs
| Lab | Article | Topic Within the VLA Track | Related Tracks |
|---|---|---|---|
| A3.1|π0 | How a VLM Connects to Continuous Control | Multimodal conditioning and the Action Expert | E, F |
| A3.2|FAST | Why Actions Need a Tokenizer | Discrete action pretraining and the continuous-control interface | E |
| A3.3|π0.5 | How Open-World Generalization Is Trained | Heterogeneous co-training, hierarchical reasoning, and post-training | D, E |
| A3.4|π0.6* / RECAP | Improving from Successes, Failures, and Corrections | How a general-purpose VLA continues to improve using deployment experience | C, G |
| A3.5|π0.7 | From a General-Purpose Policy to a Controllable Model | Prompts, visual subgoals, automatic hierarchy, and controllability | B, D, E |
| A3.6|OpenPI | OpenPI in Practice and Data Production | VLA training, data formats, and deployment engineering | G |
Recommended sequence: A0 → A1 → A2 → π0 → π0.5 → π0.6* → π0.7. FAST can be studied independently after π0.
Track B|World Models and Model-Based Planning
Track question: Can action-conditioned future changes be learned so that the model can compare “what would happen if we did this”?
B0|Track Core Course: World Models and Model-Based Planning
World Models and Planning: State-Space Models, ELBO, MPC, CEM, and Learning from Imagination
B1-B7|World Model Track Course Tree
| Module | Core Question | Course |
|---|---|---|
| B1|State Spaces and Latent Dynamics | How can a predictive state be constructed from partially observable data? | State Spaces, Filtering, ELBO, and Multistep Rollouts |
| B2|Model-Based Planning | How can actions be searched within a learned model? | MPC, CEM, Trajectory Optimization, and Risk-Aware Planning |
| B3|Learning from Imagination | How can latent rollouts be used to train an actor-critic? | Dreamer, Model Gradients, and Imagination Bias |
| B4|Video World Models | How can visual futures become subgoals and control conditions? | Video Prediction, Visual Planning, and Physical Consistency |
| B5|Decision World Model Lineage | How do internal models produce decisions through imagination, search, and MPC? | From World Models and Dreamer to TD-MPC2 |
| B6|Modern World Model Lineages | How do JEPA, spatial intelligence, and generative simulators converge? | Yann LeCun, Fei-Fei Li, and the Modern World Model Roadmap |
| B7|4D World States and World Action Models | How can actions, future worlds, and time-varying 3D structures be jointly modeled and integrated into the physical closed loop? | From Conceptual Convergence and Geometric Prediction to Executable Futures |
Recommended sequence: Begin with B0 to establish the five lineages and a unified comparison framework. Use B1-B4 to develop the necessary foundations in state, planning, learning from imagination, and video prediction. Study decision world models in depth in B5. Use B6 to understand Yann LeCun’s JEPA approach, Fei-Fei Li’s spatial intelligence approach, and generative simulators. Finally, proceed to B7 for 4D-WAM, whole-body control interfaces, and closed-loop evidence.
Cross-track relationships: Provides subgoals and counterfactual rollouts to Track D and imagined data to Track C. The visual subgoals of π0.7 and WAM-TTT can serve as cross-track case studies.
Track C|Value, Reward, and Experience Learning
Track question: How does a robot determine which states and actions produce better long-term outcomes, and how can it improve its policy using successes, failures, preferences, and corrections gathered during deployment?
C0|Track Core Course: Value, Reward, and Experience Learning
C0|Value, Reward, and Experience Learning: How Physical AI Improves Behavior from Outcomes
C1-C4|Value and Experience Learning Course Tree
| Module | Core Question | Course |
|---|---|---|
| C1|Bellman and Actor-Critic | How do long-term outcomes update a policy? | TD, Policy Gradients, Advantage, and Continuous Control |
| C2|Offline RL | How can a fixed dataset outperform behavior cloning? | Out-of-Distribution Overestimation, CQL, IQL, and AWR |
| C3|Preferences and Human Corrections | How can human feedback become reward signals and recovery supervision? | Reward Models, DAgger, Interventions, and Active Feedback |
| C4|Paper Lab | How do online, offline, corrective, and VLA experience learning compare? | SAC, CQL, IQL, DAgger, and RECAP |
Cross-track relationships: π0.6* / RECAP is primarily placed in the VLA paper labs while also serving as a central case study for this track. World models can provide imagined rollouts for this track.
Track D|Hierarchical Planning, Skills, Reasoning, and Memory
Track questions: How can long-horizon tasks be decomposed into subgoals? How can high-level language or visual plans invoke low-level continuous policies? How can a model remember demonstrations and adapt at test time?
D0|Track Core Course: Hierarchical Planning, Skills, and Memory
D0|Hierarchical Planning, Skills, and Memory: How Physical AI Completes Long-Horizon Tasks
Specialized lab: Physical Causality and Embodied Chain-of-Thought
D1-D4|Hierarchical Intelligence Course Tree
| Module | Core Question | Course |
|---|---|---|
| D1|Options and Hierarchical RL | How can initiable, terminable, and reusable skills be learned? | Options, Semi-MDPs, Skill Discovery, and Composition |
| D2|Subgoals and Embodied Reasoning | How can language plans become reachable and verifiable closed-loop subgoals? | Task Graphs, World Models, Visual Subgoals, and Recovery |
| D3|Memory and Test-Time Adaptation | How can new experience from the current environment be written into the policy? | Context, External Memory, Fast Weights, and WAM-TTT |
| D4|Paper Lab | How do traditional hierarchical RL and emerging hierarchical VLA systems compare? | Options, Embodied Chain-of-Thought, π0.7, and WAM-TTT |
Track E|Perception, State, Data, and Cross-Embodiment Learning
Track questions: How does a robot construct a spatial state from incomplete observations? What can human videos without action labels provide? How can different robots, cameras, coordinate systems, and action spaces be incorporated into joint training?
E0|Track Core Course: Data, Representations, and Cross-Embodiment Learning
E0|Data, Representations, and Cross-Embodiment Learning: What Do Robots Learn Without Action Labels?
Specialized lab: Human Videos Have No Velocity Labels
E1-E6|Perception, Data, and Cross-Embodiment Course Tree
| Module | Core Question | Course |
|---|---|---|
| E1|Video Motion and Object-Centric Representations | How can motion and interaction be extracted without action labels? | Optical Flow, Keypoints, Hand-Object Representations, and Events |
| E1.5|Spatial State Estimation and 3D Geometry | How does a robot infer itself, objects, and uncertainty from observation histories? | Bayes Filters, SE(3), Depth, Object States, and Affordances |
| E3|Latent Actions and Inverse Dynamics | How can behavioral variables be discovered from state changes? | Inverse Models, Latent Actions, Non-Identifiability, and Adaptation |
| E4|Behavior Tokenizers | How can behavior be converted into compositional semantic units? | VQ, Event Boundaries, Hierarchical Tokens, and FAST |
| E5|Cross-Embodiment Alignment and Co-Training | How can different robots share data without conflating their actions? | Adapters, Object-Centric Actions, Data Mixtures, and Transfer |
| E6|Paper Lab | How do the primary mechanisms for incorporating human data into robot training compare? | R3M, MimicPlay, ATM, HumanPlus, and WAM-TTT |
Cross-track relationships: Provides action and semantic supervision for Track A, temporal representations for Track B, and demonstrations and memory inputs for Track D. π0.5 is a case study in heterogeneous co-training, while WAM-TTT is a case study in adaptation from human demonstrations.
Track F|Dynamics, Control, and Physical Interaction
Track question: How can the actions produced by a learned model be applied stably and safely to real systems with mass, friction, contact, compliance, and latency?
F0|Track Core Course: Dynamics, Control, and Physical Interaction
Dynamics, Control, and Physical Interaction: From Robot Equations to Impedance Control
F1-F5|Dynamics, Contact, and Whole-Body Control Course Tree
| Module | Core Question | Course |
|---|---|---|
| F1|Feedback Control and Stability | How can policy outputs be tracked stably under sampling, saturation, and latency? | State Spaces, PD, Lyapunov Stability, Trajectory Tracking, and Policy Interfaces |
| F2|Contact, Impedance, and Tactile Sensing | How can robots make appropriate contact under friction, impacts, and compliant environments? | Contact Constraints, Friction Cones, Force Control, Impedance, and Tactile Closed Loops |
| F3|System Identification and Sim-to-Real | How can real dynamics be estimated, latency handled, and the simulation gap bridged? | Least Squares, Identifiability, Randomization, Online Adaptation, and Latency |
| F4|Humanoids, Mobility, and Whole-Body Control | How can dynamic balance, foothold placement, contact switching, and whole-body coordination be handled? | Floating-Base Dynamics, Centroidal Control, Locomotion RL, and High-Level/Low-Level Interfaces |
| F5|Paper and Engineering Lab | How can classical control, MPC, Residual RL, and robust learning be compared fairly? | From Classical Control to Residual RL and Sim-to-Real |
Cross-track relationships: Tracks A, B, and D must ultimately undergo real-world execution tests through this track. Failures and sensor data generated by this track are fed back into Tracks G and E.
Track G|Systems, Trustworthy Evaluation, and Data Engines
Track question: How can algorithms be transformed into reproducible, deployable, and auditable real-world robotic systems that continuously improve using failure data?
G0|Track Core Course: Systems, Trustworthy Evaluation, and Data Engines
G0|Systems, Trustworthy Evaluation, and Data Engines: How to Prove That Physical AI Actually Works
Engineering lab: OpenPI in Practice and ArcheBase Data Production
G1-G4|Systems and Trustworthy Engineering Course Tree
| Module | Core Question | Course |
|---|---|---|
| G1|Statistical Evaluation and Uncertainty | How can we prove that an improvement is real and quantify the uncertainty in the evidence? | Wilson Intervals, Stratified Estimation, Calibration, and Risk Coverage |
| G2|Data Engines and Reproducibility | How can every fact about data, training, and models remain traceable? | Schemas, Time Synchronization, Versioning, Lineage, and Quality Gates |
| G3|Deployment Safety and Regression Testing | After deployment, how can a model be monitored, degraded gracefully, overridden, rolled back, and continuously kept within its boundaries? | Safety Envelopes, Runtime Monitoring, Canaries, and Incident Regression Testing |
| G4|OpenPI Systems Lab | How can VLA data and checkpoints be turned into rollback-capable robot deployments? | From Data Contracts, Training, and Inference Services to Failure Feedback |
Additional engineering case study: OpenPI in Practice and ArcheBase Data Production
Fast Learning Paths
| Goal | Path | Completion Criterion |
|---|---|---|
| Build a comprehensive understanding in five days | 00 → A0 → B0/C0 → D0/E0 → F0/G0 | Can compare all tracks by their learning objects, supervision signals, inference methods, and closed-loop risks. |
| VLA algorithms | 00 → A0 → π-series labs → C0 → F0/G0 | Can explain VLA training, action generation, experience-based improvement, and real-world execution. |
| World models | 00 → B0 → C0 → D0 → F0/G0 | Can implement latent dynamics and MPC and validate gains in real-world control. |
| Human data | 00 → E0 → E specialized labs → A0 → D0 → G0 | Can design cross-embodiment representations, latent actions, and held-out task experiments. |
| Real-world robotics deployment | 00 → F0 → G0, while also selecting A0 or B0 | Can localize failures across the policy, control, data, and evaluation layers. |

