Skip to content

Original Feishu document · Source revision 18

💡

Mechanisms lesson: An offline checkpoint is not a robotics product. A deployment system requires safety envelopes, runtime monitoring, anomaly detection, fallback modes, human takeover, shadow evaluation, regression gates, and traceable rollback. This lesson examines how to translate model uncertainty into system actions.

Learning Objectives ​

After completing this lesson, you should be able to map hazards to constraints and monitors; distinguish model confidence, out-of-distribution detection, and physical safety metrics; design runtime assurance, shadow deployment, canary releases, and rollback; and make release decisions using risk budgets rather than average success rates.

1. Safety Is Not a Loss Term ​

Robot risk can be approximated as:

Interpretation: Expected risk equals the probability of each type of hazard multiplied by its consequence cost, summed over all hazard categories.

Derivation: Express total loss as a random variable partitioned by hazardous events. Applying the law of total expectation yields a weighted sum of each event’s probability multiplied by its conditional loss. This formulation requires sufficiently comprehensive hazard categories and comparable costs; correlated accident chains, unacceptable harm, and hard regulatory constraints cannot be diluted through averaging expected values.

Low-probability, high-consequence events cannot be diluted by large numbers of routine successes. Engineering systems must simultaneously reduce occurrence probability, limit consequences, increase detection rates, and shorten response times. Unacceptable harm should be treated directly as a hard constraint rather than something that task utility may offset.

2. From Hazard Analysis to Runtime Constraints ​

HazardObservable signalsRuntime action
Collision/pinch injuryDistance, force, velocity, human-occupied regionsLimit speed, stop, retreat
Dropped objectTactile slip, gripper width, visionRegrasp, place in a safe zone
Policy loses track of the objectiveConfidence, progress, cyclic trajectoriesRe-perceive, request human assistance
System timeoutInference latency, message ageHold position, switch to a local controller
Sensor anomalyDropped frames, drift, timestampsDegrade operation or terminate

3. Safety Envelope ​

Let the admissible state set be . A simple runtime filter finds the safe action closest to the policy action:

Interpretation: Within the action bounds, find the command closest to the policy’s recommendation while requiring that the model-predicted next state remain within the safe set.

Derivation: Treat the policy action as the desired command, express safety requirements—such as joint limits, velocity, collision distance, and contact force—as constraints on the next state, and then project onto those constraints with the minimum possible modification. If the model or state estimate is incorrect, or constraints are missing, the projection itself does not guarantee safety; independent monitoring and hardware-enforced limits are therefore required.

Constraints may come from joint limits, velocity limits, collision distance, contact force, power, or barrier functions. The safety filter should record how much it modifies each policy action. Large corrections sustained over time indicate that the division of responsibilities between the policy and the safety layer has become misaligned.

4. Control Barrier Function Intuition ​

Use to denote the safe set. For the continuous system , we require:

Interpretation: The sum of the safety function’s rate of change under the uncontrolled dynamics, the component affected by the control input, and a boundary-recovery term must be nonnegative.

Derivation: Taking the time derivative of h(x) along the system x_dot=f(x)+g(x)u gives the gradient of h multiplied by f and g u, respectively—that is, the two Lie derivatives. After adding a monotone function alpha, the constraint requires the system, when approaching the boundary, to remain on or return to the safe side at a sufficient rate, thereby making the safe set forward invariant.

CBF guarantees depend on state observability, bounded model error, and the optimization completing on time. Real systems also require physical emergency stops, independent hardware-enforced limits, and conservative margins. Mathematical feasibility must not be equated with the safety of the product as a whole.

Course whiteboard

💡

Interactive Validation | Action Age, Canary Releases, Rollback, and Fault-Injection Lab

Vary action age, the safety response budget, canary traffic, and fault injection to observe when the system should degrade operation or roll back.

Action Age, Canary Releases, Rollback, and Fault-Injection Lab

:::

5. Model Confidence Is Not Safety ​

Action-distribution entropy, ensemble variance, and VLM token probabilities can serve as uncertainty signals, but danger also depends on speed, contact, proximity to humans, and recoverability. The same low confidence may be acceptable during slow movement through open space but unacceptable when a cutting tool is close to a person.

Triggering rules should therefore be conditioned on the risk context:

Interpretation: The system must intervene when the estimated probability of harm in the current context multiplied by the consequence cost exceeds the risk threshold.

Derivation: The same uncertainty has different consequences in a low-speed open environment and near a person, so probability predictions are combined with contextual costs to produce conditional risk. The threshold should be determined jointly by the safety budget, calibration error, and response capability. High-consequence scenarios should also have hard exclusion zones that do not depend on probability estimates.

6. Four Layers of Online Monitoring ​

  1. System health: Latency, dropped frames, temperature, GPU status, network status, and message age.
  2. Physical health: Force, velocity, collision distance, remaining joint margin, and vibration.
  3. Task health: Progress, repeated actions, timeouts, and whether the target still exists.
  4. Distribution health: Drift in input embeddings, objects, scenes, actions, and failure slices.

Monitoring metrics must be tied to executable responses; otherwise, they are merely dashboard decorations.

6.1 Safety Responses Must Outrun the Time to Harm ​

It is not enough for monitoring to detect an issue eventually. For rapid collisions or dropped objects, the total time required for detection, decision-making, and braking must be shorter than the remaining time between hazard onset and irreversible harm:

Interpretation: The total duration of hazard detection, response selection, and braking must be shorter than the time required for the hazard to develop into harm.

Derivation: Starting from hazard onset, monitoring first incurs detection latency, the safety logic then incurs decision latency, and the actuator finally incurs braking or retreat time. A software response is physically meaningful only if the sum of these three intervals still precedes the moment of harm. Otherwise, the system must reduce speed in advance, increase the safety distance, or use faster local protection.

This inequality should be validated under worst-case rather than average latency and must account for sensor periods, queue jitter, network timeouts, braking distance, and load variations.

6.2 Drift Detection Requires Sequential Statistics ​

For latency, action-boundary, or embedding anomaly scores ell_t, a one-sided CUSUM can accumulate persistent shifts:

Interpretation: The current cumulative anomaly score equals the previous cumulative value plus the new score minus the tolerated drift, truncated at zero. An alert is triggered when the score exceeds threshold H.

Derivation: If the mean normal score is below the reference value nu, the cumulative value will frequently return to zero. If the distribution shifts upward persistently, each step contributes positive drift, and the cumulative value gradually crosses the threshold. H controls the trade-off between false alarms and detection latency and must be calibrated using both normal-operation and fault-injection data.

7. Shadow, Canary, and Staged Releases ​

StageModel authorityPrimary evidence
Offline replayNo control authorityConsistency, latency, action bounds
ShadowPredicts but does not executeDisagreement with the production policy, resource usage
Controlled experimentSafe test site with immediate human takeover availableReal closed-loop behavior and fault response
CanarySmall subset of robots/tasksOnline risk and rollback capability
Expanded rolloutGradually increased coverageContinued compliance with stratified metrics

8. Regression Testing Is Not Merely Rerunning an Old Benchmark ​

The regression suite should include core capabilities, severe historical failures, boundary conditions, degraded sensors, injected latency, hardware versions, and cross-task combinations. Every fix should add a test that reproduces the corresponding failure, creating a cumulative “incident immunity memory.”

9. Release Gates ​

Release gates should include all of the following:

  • A lower bound on primary-task performance, not merely a point estimate.
  • No degradation in any critical safety metric.
  • All regression tests for severe historical failures must pass.
  • Inference latency, resource usage, and timeouts must remain within budget.
  • Data, models, code, firmware, and configurations must be traceable.
  • Drills for stopping, degraded operation, human takeover, and rollback must pass.

10. Incidents and Near Misses ​

Near misses that cause no actual loss should still be recorded. At minimum, each event record should include the timeline, model and hardware versions, observations and actions, whether the monitor triggered, human interventions, direct causes, systemic causes, remediation items, and regression tests. Root-cause analysis must not stop at non-actionable labels such as “model hallucination.”

11. Minimum Deployment Drill ​

Drill system. Build a replay-plus-hardware-in-the-loop environment in which the production policy, new policy, safety filter, independent monitor, and low-level controller all run at their actual frequencies. Every fault injection must record the actual onset time, detection time, response command, and physical stop time.

Fault injectionExpected responseRequired measurements
Camera freeze, timestamp regression, force-sensor driftDegrade perception, hold, or stop safelyDetection rate, false-positive rate, detection latency, and message age
Inference timeout, GPU overheating, network packet lossSwitch to a local controller and prohibit stale actionsWorst-case end-to-end latency, switchover gap, and control continuity
Human entry, reduced collision distance, sudden increase in contact forceLimit speed, retreat, or trigger an emergency stopResponse-time budget, stopping distance, peak force, and minimum safety margin
Target disappearance, cyclic actions, task timeoutRe-perceive, recover, or request human assistanceTask-monitor recall, duration of ineffective loops, and takeover success rate
Slow drift in input embeddings and task distributionsAlert, reduce coverage, or halt rollout expansionCUSUM detection latency, interval between false alarms, and performance on drift slices
Degraded canary metrics and incompatible configurationsAtomically roll back the model, configuration, tokenizer, and control parametersRollback time, state recovery, version consistency, and post-rollback regression results

Release sequence. Offline replay, shadow deployment, controlled experiments, canary releases, and expanded rollout must use progressively increasing authority. Pass, stop, and rollback criteria must be defined in advance for every stage. Shadow disagreement only indicates that the candidate behaves differently; it cannot replace real closed-loop safety testing.

Minimum outputs. A fault-injection coverage matrix, detection-latency distribution, safety response-time budget, risk-slice dashboard, shadow disagreement report, canary decision log, and one-click rollback drill record.

12. Failure Modes ​

Failure modeSystemic causeDecisive checkRemediation
Safety logic runs in the same process as the policyA model hang, memory exhaustion, or GPU failure disables protection at the same timeKill the policy process and observe whether the protection system can still brakeUse an independent process, independent watchdog, and hardware emergency-stop chain
Thresholds were tuned only on the development setReal-world noise, latency, and high-risk subgroups were excluded from calibrationRun fault injection and stratified reliability/risk-coverage testsUse context-conditioned thresholds, conservative margins, and periodic recalibration
The emergency-stop action itself causes an object to be dropped“Stop” does not account for the held object, gravity, or current contact stateDrill shutdowns across different poses, loads, and task stagesDesign state-dependent fallback modes such as holding, placing in a safe zone, and controlled retreat
Average metrics pass while high-risk slices regressCanary traffic is not stratified by hazard contextSlice results by human proximity, tool, speed, and task stageUse risk-weighted sampling and gates that prohibit regression in any critical slice
Rollback switches only the checkpointThe configuration, tokenizer, firmware, and control parameters remain on the new versionsCompare the complete deployment manifest after rollbackUse atomic release packages and transactional rollback
Monitoring raises alerts, but no one is responsibleThere is no owner, response deadline, or automated actionDrill the complete timeline from an overnight alert to recoveryBind every monitor to a runbook, owner, SLO, and automated fallback action

13. Lesson Exercises ​

  1. List five hazard categories and corresponding monitors for a collaborative robotic-arm pick-and-place task.
  2. Construct an example in which the model has high confidence but the physical risk is high.
  3. Specify the release gates for moving from shadow deployment to a canary release.
  4. Design independent safety responses for camera freezes and inference timeouts.
  5. Convert an object-drop incident into an automatically replayable regression test.

Relationship to Other Tracks ​

G3 does not replace the algorithm tracks; it constrains their deployment authority. Track F provides physical limits and low-level protection; G1 provides statistical thresholds; G2 ensures version traceability; and human takeovers and failure data from Track C can become learning signals for the next iteration.

14. Evidence Matrix for Safety and Runtime Assurance Papers ​

WorkFacts from the paperAuthors’ interpretationCourse assessment
Sha et al.|Simplex ArchitectureThe Simplex architecture allows a high-performance but not fully verified controller to coexist with a verified safe controller, with a decision module switching between them when necessary.The authors use an independent safety baseline to limit the risks of online upgrades.The core of runtime assurance is fault isolation and a switchable safe fallback—not adding another loss term to the policy.
Ames et al.|Control Barrier FunctionsCBFs transform the forward-invariance condition for a safe set into constraints that can be added to online control optimization.The authors use barrier inequalities to combine performance control with safety constraints.CBFs cover only hazards that the model and state estimate can express; they must be combined with hardware protection, latency budgets, and fault monitoring.
García & Fernández|Safe RL SurveyThis survey organizes safe reinforcement learning methods into two broad categories: modifying optimization criteria and restricting the exploration process.The authors distinguish safety during training from performance guarantees during deployment.Robot deployment safety cannot be justified solely by training rewards; evidence for exploration constraints, runtime protection, and rollback must be evaluated separately.
Lakshminarayanan et al.|Deep EnsemblesDeep Ensembles use multiple independently trained models to improve predictive uncertainty and calibration.The authors use disagreement among models as a practical approximation of predictive uncertainty.Ensemble variance is a signal, not a safety conclusion; it must be calibrated in conjunction with consequences, physical state, and distribution drift.
Geifman & El-Yaniv|Selective ClassificationSelective classification studies the risk–coverage trade-off when a model rejects a subset of samples.The authors reduce the error risk among accepted samples by ranking predictions by confidence.Robotics systems must map rejection to re-perception, degraded operation, or human takeover while accounting for coverage and response costs.
Mitchell et al.|Model CardsModel Cards proposes documenting intended uses, evaluation conditions, limitations, and performance across groups.The authors use structured documentation to improve deployment transparency and clarify boundaries of responsibility.Physical AI model cards should also be linked to control frequency, hardware, firmware, safety layers, rollback packages, and the history of real-world incidents.

15. Cross-Reading ​

Read together with F1–F3. Runtime safety depends on stability, contact limits, latency, and system identification; it cannot be determined independently from high-level confidence.

Read together with G1. Thresholds, canary releases, and staged rollouts require calibration, risk coverage, and stratified confidence intervals.

Read together with G2. A complete rollback must restore the model, configuration, tokenizer, firmware, and control parameters, while preserving traceability for every event.

Article text is licensed under the Apache License 2.0