Engineering

The five hardest problems in Physical AI safety

3 min readMati Melchior
The five hardest problems in Physical AI safety

Five problems. Each is real, each is unsolved at production scale, and each represents a potential decade-defining company if solved cleanly in the next three years.

Problem 1: Real-time verification of learned policies against physical safety constraints. A robot's neural network produces an action. Before that action moves a motor, can you verify — in real time, with formal guarantees — that the action satisfies a set of physical safety constraints? Not "approximately safe." Not "probably safe." Provably safe, at the speed of inference. Current research (SAFE-SMART, RAIL, conformal prediction methods) shows progress but no production-ready solution exists. The fundamental tension: formal verification is computationally expensive, and real-time control demands decisions on a millisecond budget.

Problem 2: Fleet-level anomaly detection without per-robot false-positive explosion. Monitoring a single robot for anomalous behavior is well understood. Monitoring a thousand robots simultaneously, with a shared model, without drowning operators in false positives, is not. The detection threshold that works for one robot produces N times the alerts for N robots. Naive scaling of anomaly detection is operationally useless. What's needed is population-level statistical monitoring — detecting when the fleet's aggregate behavior deviates from expectations, not when any individual robot deviates.

Problem 3: Continuous certification that survives over-the-air model updates. You certify a safety system. Six months later, the AI vendor pushes a model update that changes the inference behavior. Is the system still certified? From 20 January 2027, when Regulation (EU) 2023/1230 replaces the Machinery Directive, a software modification that alters a safety function can make the modifier a new manufacturer — new risk assessment, new declaration of conformity, new CE marking. The tooling is moving toward this: NVIDIA's Halos AI Systems Inspection Lab, announced 22 June 2026, is ANAB-accredited and helps partners prepare Halos integrations for third-party certification by TÜV Rheinland, TÜV SÜD, UL Solutions, exida, SGS and CertX. But an inspection that prepares you for a certification is not a re-validation that runs after every update. No framework today provides continuous assurance — the ability to re-validate safety properties automatically after each model push, without a full manual re-assessment. ISO/IEC TS 22440, the first international standard aimed squarely at functional safety and AI, was still at committee-draft stage through 2026.

Problem 4: Composable safety guarantees across heterogeneous fleets. One vendor's AMR is rated Performance Level d under ISO 13849-1. Another vendor's cobot has SIL 2. A third vendor's vision system has no safety rating at all. When these three systems operate together in the same workspace, what is the composite safety guarantee? There is no standard methodology for composing safety guarantees across heterogeneous, multi-vendor robotic systems. ISO 10218-1:2025 and ISO 10218-2:2025 raised the bar for robots and robot cells, but neither tells you how to compose ratings across vendors. VDA 5050 standardises the interface between AGVs/AMRs and the master control, and leaves functional safety explicitly out of scope — emergency stop and personnel detection stay on the vehicle, under standards such as ISO 3691-4.

Problem 5: Root-cause attribution after multi-robot coordinated failures. When a fleet-level incident occurs — congestion, deadlock, or safety boundary violation involving multiple robots — identifying root cause is extraordinarily difficult. Each robot has its own log. The fleet manager has its own log. The timestamps may not be synchronized. The causal chain may involve emergent behavior that no single robot's log captures. Building a forensic reconstruction of a multi-robot incident to a standard sufficient for regulatory investigation or insurance claims is an unsolved engineering challenge.

These five problems share a pattern: each exists at the intersection of AI/ML, hardware-enforced safety, and scale. Single-robot solutions don't transfer. Standard approaches break. The company that solves any one of them cleanly doesn't just build a product — it defines a category.

Share

Physical AI Safety Dispatch

Monthly analysis. No spam. One exclusive insight per issue.

One issue per month. Unsubscribe in one click from any email. Privacy policy.

We use cookies

This site uses essential cookies to function and, with your consent, analytics cookies (Google Analytics) to understand how the site is used. Learn more.