AI’s Leap Into the Physical World Demands New Safety Measures, Survey Finds

As artificial intelligence moves from digital screens into the physical world, the stakes are rising. A new survey published in the journal Machine Intelligence Research outlines the security and ethical challenges posed by vision-language models (VLMs) and vision-language-action models (VLAs) that guide embodied systems like autonomous vehicles, drones, and service robots. These models, which connect visual and textual information to physical actions, are increasingly used in critical applications where a mistake can have real-world consequences.

The review, conducted by researchers from the Institute of Automation, Chinese Academy of Sciences, University College London, Minzu University of China, and the China Academy of Electronics and Information Technology, examines the entire pipeline of embodied intelligence—from perception and planning to instruction following and human-robot interaction. The authors argue that the same multimodal capabilities that enable these systems to understand and act also create vulnerabilities that can lead to catastrophic failures.

One of the key concerns is hallucination, where models describe objects that do not exist, potentially causing an autonomous vehicle to brake suddenly for a phantom obstacle or a robot to grasp at nothing. The review also highlights the risks of adversarial attacks, where tiny perturbations in sensor inputs can cause misclassification, and backdoor attacks, where hidden triggers can hijack the system’s behavior. Privacy is another major issue, as persistent sensing can expose sensitive information about individuals.

The authors emphasize that existing safeguards are often inadequate—they are fragmented, benchmark-specific, or too computationally expensive for real-time use. They propose a layered defense strategy that spans the entire system, from sensor input to physical execution. This includes hallucination filtering, cross-modal forgery detection, defenses against adversarial perturbations and backdoors, privacy-preserving computation techniques like differential privacy and homomorphic encryption, and safeguards for navigation and control.

In addition to technical defenses, the survey stresses the importance of human oversight and interpretability. Robots must be able to explain their actions, recognize uncertainty, and fall back to safe modes when they are not confident. The authors call for transparent risk metrics and continuous monitoring to ensure safety in dynamic, real-world environments.

The findings come at a critical time as embodied AI is being deployed in sectors such as transportation, healthcare, and manufacturing. The review provides a practical checklist for developers and regulators to evaluate embodied systems before large-scale deployment. However, the authors caution that strong laboratory results may not translate to noisy, culturally diverse, and resource-constrained settings. They advocate for cross-disciplinary cooperation and testing that measures not just task success, but safe behavior under stress.

This research highlights that the challenge of embodied AI is not just about making models more accurate, but ensuring they remain safe and reliable when sensors, language inputs, and operating conditions are imperfect. As embodied intelligence becomes more prevalent, these considerations will be crucial for building trust and ensuring that AI can be safely integrated into our physical world.

Blockchain verification QR code
Blockchain Registered
This article is registered on the blockchain by Newsramp. Verify this record.