Skip to content

Feishu Source Catalog · Source Revision 91

💡

Intended audience: Algorithm, data, and robotics systems practitioners with knowledge of probability and statistics, linear algebra, and foundational deep learning who want to quickly build a comprehensive understanding of Physical AI. This catalog is organized not by company or paper publication date, but by the mathematical objects that models actually learn.

How to Use This Catalog ​

Physical AI is not a single-track technology path. Direct policies, world models, value learning, hierarchical planning, data representations, control, and systems engineering each address different problems while operating together within the same physical closed loop.

All readers should first complete the “Common Foundations” and then choose one primary track for deeper study. An article’s primary placement only indicates the best entry point; it does not mean that the article is relevant only to that track. Cross-track relationships are explicitly identified after each module.

Unified Course Quality Standards ​

This course is intended for readers with university-level knowledge of probability and statistics, linear algebra, and foundational deep learning. Every mechanism course and paper lab must meet the following requirements to be considered complete.

RequirementMust Answer
Learning objectIs the model estimating a distribution, state, value, subgoal, representation, or control law?
How to read equationsEvery standalone equation must be immediately followed by a blockquote that explains, in plain language, the conditions, random variables, and operators term by term.
DerivationExplain which probabilistic assumption, definition, physical equation, or optimization objective the equation follows from; presenting only the result is not allowed.
VisualizationInclude at least one computational graph, probabilistic graph, closed-loop diagram, or curve that genuinely reduces the effort required to understand the material.
Minimal experimentReaders must be able to reproduce the key phenomenon in a toy environment rather than simply trust the paper’s conclusions.
Position in the literatureDistinguish facts, the authors’ interpretations, and the course’s assessment, and compare against strong baselines from the same track.
Closed-loop risksDiscuss distribution shift, contact, latency, missing observations, safety, and failure recovery.

All equations follow a unified “four-step explanation”: first present the equation, then provide a blockquoted plain-language reading, derive it, and finally explain its meaning within the robot’s closed loop.

Example equation-reading block: Rather than merely reading the symbols aloud, explain “what is conditioned on, which random quantity is being predicted, and which cases are covered by the sum or expectation.”

Inline symbols are explained consistently in each section’s notation table. Every standalone equation that defines a concept, loss, dynamics, update rule, or inference procedure must include its own blockquoted reading and derivation.

00|Common Foundations: The Shared Language of All Tracks ​

Core questions: What are states, observations, actions, trajectories, and policies? Why does the training distribution change during deployment? How do the loss functions of probabilistic models translate into real robot behavior?

00.1|Common Foundations Core Course ​

00|Physical AI Common Foundations: What Is an Agent Actually Learning?

Paper and derivation lab: An In-Depth Derivation of Behavior Cloning, Distribution Shift, and Flow Matching

Learning objects: Conditional action distributions, maximum likelihood, mean squared error, multimodal actions, action chunks, and deployment distribution shift.

00.2|Boundaries Between Physical AI Modules ​

ModulePrimary Mathematical ObjectQuestion Answered
PolicyConditional action distributionWhat should be done now?
World modelAction-conditioned state-transition distributionWhat will happen after doing this?
Value functionState value and action valueHow good will the future be?
PlannerCandidate action sequences or subgoalsWhich path should be selected?
ControllerFeedback control lawHow can the action be executed stably?

Track A|VLA and Direct Policy Learning ​

Track question: Can robot actions be generated directly from vision, language, and proprioceptive state? This track learns conditional action distributions and is currently the primary direction for general-purpose robot policies and robot foundation models.

A0|Track Core Course: VLA and Direct Policy Learning ​

A0|VLA and Direct Policy Learning: From Conditional Imitation to Robot Foundation Models

Cross-track relationships: Uses data and representations from Track E; can be further improved using value signals from Track C; and is ultimately executed by the control systems in Track F.

A1|Specialized Lab: Action Representations and FAST ​

Action Representations: MSE, Autoregressive Tokens, Diffusion, Flow Matching, and FAST

Core question: Does the policy output a conditional mean, a discrete action sequence, a score, or a probability-flow vector field?

A2|Mainstream VLA Lineages and Architectural Evolution ​

RT, Open X, Octo, OpenVLA, GR00T, and Gemini Robotics

Core questions: Beyond the π series, how do mainstream VLAs combine multitask imitation, vision-language pretraining, cross-embodiment data, and generative action modules? Do capability gains actually come from architecture, data, or post-training?

A3|Physical Intelligence π-Series Paper Labs ​

LabArticleTopic Within the VLA TrackRelated Tracks
A3.1|π0How a VLM Connects to Continuous ControlMultimodal conditioning and the Action ExpertE, F
A3.2|FASTWhy Actions Need a TokenizerDiscrete action pretraining and the continuous-control interfaceE
A3.3|π0.5How Open-World Generalization Is TrainedHeterogeneous co-training, hierarchical reasoning, and post-trainingD, E
A3.4|π0.6* / RECAPImproving from Successes, Failures, and CorrectionsHow a general-purpose VLA continues to improve using deployment experienceC, G
A3.5|π0.7From a General-Purpose Policy to a Controllable ModelPrompts, visual subgoals, automatic hierarchy, and controllabilityB, D, E
A3.6|OpenPIOpenPI in Practice and Data ProductionVLA training, data formats, and deployment engineeringG

Recommended sequence: A0 → A1 → A2 → π0 → π0.5 → π0.6* → π0.7. FAST can be studied independently after π0.

Track B|World Models and Model-Based Planning ​

Track question: Can action-conditioned future changes be learned so that the model can compare “what would happen if we did this”?

B0|Track Core Course: World Models and Model-Based Planning ​

World Models and Planning: State-Space Models, ELBO, MPC, CEM, and Learning from Imagination

B1-B7|World Model Track Course Tree ​

ModuleCore QuestionCourse
B1|State Spaces and Latent DynamicsHow can a predictive state be constructed from partially observable data?State Spaces, Filtering, ELBO, and Multistep Rollouts
B2|Model-Based PlanningHow can actions be searched within a learned model?MPC, CEM, Trajectory Optimization, and Risk-Aware Planning
B3|Learning from ImaginationHow can latent rollouts be used to train an actor-critic?Dreamer, Model Gradients, and Imagination Bias
B4|Video World ModelsHow can visual futures become subgoals and control conditions?Video Prediction, Visual Planning, and Physical Consistency
B5|Decision World Model LineageHow do internal models produce decisions through imagination, search, and MPC?From World Models and Dreamer to TD-MPC2
B6|Modern World Model LineagesHow do JEPA, spatial intelligence, and generative simulators converge?Yann LeCun, Fei-Fei Li, and the Modern World Model Roadmap
B7|4D World States and World Action ModelsHow can actions, future worlds, and time-varying 3D structures be jointly modeled and integrated into the physical closed loop?From Conceptual Convergence and Geometric Prediction to Executable Futures

Recommended sequence: Begin with B0 to establish the five lineages and a unified comparison framework. Use B1-B4 to develop the necessary foundations in state, planning, learning from imagination, and video prediction. Study decision world models in depth in B5. Use B6 to understand Yann LeCun’s JEPA approach, Fei-Fei Li’s spatial intelligence approach, and generative simulators. Finally, proceed to B7 for 4D-WAM, whole-body control interfaces, and closed-loop evidence.

Cross-track relationships: Provides subgoals and counterfactual rollouts to Track D and imagined data to Track C. The visual subgoals of π0.7 and WAM-TTT can serve as cross-track case studies.

Track C|Value, Reward, and Experience Learning ​

Track question: How does a robot determine which states and actions produce better long-term outcomes, and how can it improve its policy using successes, failures, preferences, and corrections gathered during deployment?

C0|Track Core Course: Value, Reward, and Experience Learning ​

C0|Value, Reward, and Experience Learning: How Physical AI Improves Behavior from Outcomes

C1-C4|Value and Experience Learning Course Tree ​

ModuleCore QuestionCourse
C1|Bellman and Actor-CriticHow do long-term outcomes update a policy?TD, Policy Gradients, Advantage, and Continuous Control
C2|Offline RLHow can a fixed dataset outperform behavior cloning?Out-of-Distribution Overestimation, CQL, IQL, and AWR
C3|Preferences and Human CorrectionsHow can human feedback become reward signals and recovery supervision?Reward Models, DAgger, Interventions, and Active Feedback
C4|Paper LabHow do online, offline, corrective, and VLA experience learning compare?SAC, CQL, IQL, DAgger, and RECAP

Cross-track relationships: π0.6* / RECAP is primarily placed in the VLA paper labs while also serving as a central case study for this track. World models can provide imagined rollouts for this track.

Track D|Hierarchical Planning, Skills, Reasoning, and Memory ​

Track questions: How can long-horizon tasks be decomposed into subgoals? How can high-level language or visual plans invoke low-level continuous policies? How can a model remember demonstrations and adapt at test time?

D0|Track Core Course: Hierarchical Planning, Skills, and Memory ​

D0|Hierarchical Planning, Skills, and Memory: How Physical AI Completes Long-Horizon Tasks

Specialized lab: Physical Causality and Embodied Chain-of-Thought

D1-D4|Hierarchical Intelligence Course Tree ​

ModuleCore QuestionCourse
D1|Options and Hierarchical RLHow can initiable, terminable, and reusable skills be learned?Options, Semi-MDPs, Skill Discovery, and Composition
D2|Subgoals and Embodied ReasoningHow can language plans become reachable and verifiable closed-loop subgoals?Task Graphs, World Models, Visual Subgoals, and Recovery
D3|Memory and Test-Time AdaptationHow can new experience from the current environment be written into the policy?Context, External Memory, Fast Weights, and WAM-TTT
D4|Paper LabHow do traditional hierarchical RL and emerging hierarchical VLA systems compare?Options, Embodied Chain-of-Thought, π0.7, and WAM-TTT

Track E|Perception, State, Data, and Cross-Embodiment Learning ​

Track questions: How does a robot construct a spatial state from incomplete observations? What can human videos without action labels provide? How can different robots, cameras, coordinate systems, and action spaces be incorporated into joint training?

E0|Track Core Course: Data, Representations, and Cross-Embodiment Learning ​

E0|Data, Representations, and Cross-Embodiment Learning: What Do Robots Learn Without Action Labels?

Specialized lab: Human Videos Have No Velocity Labels

E1-E6|Perception, Data, and Cross-Embodiment Course Tree ​

ModuleCore QuestionCourse
E1|Video Motion and Object-Centric RepresentationsHow can motion and interaction be extracted without action labels?Optical Flow, Keypoints, Hand-Object Representations, and Events
E1.5|Spatial State Estimation and 3D GeometryHow does a robot infer itself, objects, and uncertainty from observation histories?Bayes Filters, SE(3), Depth, Object States, and Affordances
E3|Latent Actions and Inverse DynamicsHow can behavioral variables be discovered from state changes?Inverse Models, Latent Actions, Non-Identifiability, and Adaptation
E4|Behavior TokenizersHow can behavior be converted into compositional semantic units?VQ, Event Boundaries, Hierarchical Tokens, and FAST
E5|Cross-Embodiment Alignment and Co-TrainingHow can different robots share data without conflating their actions?Adapters, Object-Centric Actions, Data Mixtures, and Transfer
E6|Paper LabHow do the primary mechanisms for incorporating human data into robot training compare?R3M, MimicPlay, ATM, HumanPlus, and WAM-TTT

Cross-track relationships: Provides action and semantic supervision for Track A, temporal representations for Track B, and demonstrations and memory inputs for Track D. π0.5 is a case study in heterogeneous co-training, while WAM-TTT is a case study in adaptation from human demonstrations.

Track F|Dynamics, Control, and Physical Interaction ​

Track question: How can the actions produced by a learned model be applied stably and safely to real systems with mass, friction, contact, compliance, and latency?

F0|Track Core Course: Dynamics, Control, and Physical Interaction ​

Dynamics, Control, and Physical Interaction: From Robot Equations to Impedance Control

F1-F5|Dynamics, Contact, and Whole-Body Control Course Tree ​

ModuleCore QuestionCourse
F1|Feedback Control and StabilityHow can policy outputs be tracked stably under sampling, saturation, and latency?State Spaces, PD, Lyapunov Stability, Trajectory Tracking, and Policy Interfaces
F2|Contact, Impedance, and Tactile SensingHow can robots make appropriate contact under friction, impacts, and compliant environments?Contact Constraints, Friction Cones, Force Control, Impedance, and Tactile Closed Loops
F3|System Identification and Sim-to-RealHow can real dynamics be estimated, latency handled, and the simulation gap bridged?Least Squares, Identifiability, Randomization, Online Adaptation, and Latency
F4|Humanoids, Mobility, and Whole-Body ControlHow can dynamic balance, foothold placement, contact switching, and whole-body coordination be handled?Floating-Base Dynamics, Centroidal Control, Locomotion RL, and High-Level/Low-Level Interfaces
F5|Paper and Engineering LabHow can classical control, MPC, Residual RL, and robust learning be compared fairly?From Classical Control to Residual RL and Sim-to-Real

Cross-track relationships: Tracks A, B, and D must ultimately undergo real-world execution tests through this track. Failures and sensor data generated by this track are fed back into Tracks G and E.

Track G|Systems, Trustworthy Evaluation, and Data Engines ​

Track question: How can algorithms be transformed into reproducible, deployable, and auditable real-world robotic systems that continuously improve using failure data?

G0|Track Core Course: Systems, Trustworthy Evaluation, and Data Engines ​

G0|Systems, Trustworthy Evaluation, and Data Engines: How to Prove That Physical AI Actually Works

Engineering lab: OpenPI in Practice and ArcheBase Data Production

G1-G4|Systems and Trustworthy Engineering Course Tree ​

ModuleCore QuestionCourse
G1|Statistical Evaluation and UncertaintyHow can we prove that an improvement is real and quantify the uncertainty in the evidence?Wilson Intervals, Stratified Estimation, Calibration, and Risk Coverage
G2|Data Engines and ReproducibilityHow can every fact about data, training, and models remain traceable?Schemas, Time Synchronization, Versioning, Lineage, and Quality Gates
G3|Deployment Safety and Regression TestingAfter deployment, how can a model be monitored, degraded gracefully, overridden, rolled back, and continuously kept within its boundaries?Safety Envelopes, Runtime Monitoring, Canaries, and Incident Regression Testing
G4|OpenPI Systems LabHow can VLA data and checkpoints be turned into rollback-capable robot deployments?From Data Contracts, Training, and Inference Services to Failure Feedback

Additional engineering case study: OpenPI in Practice and ArcheBase Data Production

Fast Learning Paths ​

GoalPathCompletion Criterion
Build a comprehensive understanding in five days00 → A0 → B0/C0 → D0/E0 → F0/G0Can compare all tracks by their learning objects, supervision signals, inference methods, and closed-loop risks.
VLA algorithms00 → A0 → π-series labs → C0 → F0/G0Can explain VLA training, action generation, experience-based improvement, and real-world execution.
World models00 → B0 → C0 → D0 → F0/G0Can implement latent dynamics and MPC and validate gains in real-world control.
Human data00 → E0 → E specialized labs → A0 → D0 → G0Can design cross-embodiment representations, latent actions, and held-out task experiments.
Real-world robotics deployment00 → F0 → G0, while also selecting A0 or B0Can localize failures across the policy, control, data, and evaluation layers.

Article text is licensed under the Apache License 2.0