Research Track 03 · Decision Safety
LLM/VLA 机器人决策安全 Decision Safety for LLM/VLA-Based Robots
当机器人开始使用大语言模型、视觉语言模型或视觉语言动作模型时,危险不一定来自电机、传感器或机械结构故障。 它可能来自语义理解、任务分解、世界模型、动作生成和长时序执行之间的链式偏差。
When robots begin to use large language models, vision-language models, or vision-language-action models, hazards do not necessarily originate from motors, sensors, or mechanical faults. They may arise from chained deviations across semantic understanding, task decomposition, world modeling, action generation, and long-horizon execution.
1. 大模型把机器人 SOTIF 推向新阶段 1. Foundation Models Move Robot SOTIF Into a New Phase
在传统自动驾驶 SOTIF 中,我们关心感知误检、漏检、误分类、场景触发条件和功能不足。 这些问题在机器人里仍然存在,但 LLM/VLA 机器人又增加了一层:语言和任务语义。
In traditional automated-driving SOTIF, we worry about perception false positives, false negatives, misclassification, scenario triggering conditions, and functional insufficiencies. These issues still exist in robotics, but LLM/VLA robots add another layer: language and task semantics.
同一句“帮我把这个拿过去”,在不同空间、对象、用户身份、隐私边界和安全约束下,可能对应完全不同的可接受行为。
The same instruction, “help me take this over there,” can imply very different acceptable behaviors depending on the space, object, user identity, privacy boundary, and safety constraint.
2. 决策安全的五个典型风险源 2. Five Typical Sources of Decision Safety Risk
| 风险源 | 在机器人中的表现 | 需要的证据 |
|---|---|---|
| 语义幻觉 | 错误理解用户意图、对象属性或空间规则 | 指令歧义测试、对象/地点语义校验、拒绝策略 |
| 分布外行为 | 面对未见过的物体、姿态、人群或场景组合时做出不稳定动作 | 边界场景集、异常检测、置信度与回退触发记录 |
| 目标偏移 | 完成了字面任务,但违反安全、隐私或社会规范 | 任务约束层、价值冲突案例、人工确认机制 |
| 长时序累积风险 | 单步动作都合理,但多步执行后进入危险状态 | 计划滚动评估、状态可达性分析、运行时监控日志 |
| 多模态攻击面 | 文字、图像、贴纸、语音或环境提示影响机器人行为 | 提示注入测试、环境诱导测试、跨模态一致性检查 |
| Risk Source | How It Appears in Robots | Evidence Needed |
|---|---|---|
| Semantic Hallucination | Misinterpreting user intent, object properties, or spatial rules | Ambiguous-instruction tests, semantic checks for objects and places, refusal strategies |
| Out-of-Distribution Behavior | Unstable action under unseen objects, postures, crowds, or scenario combinations | Boundary-scenario sets, anomaly detection, confidence records, and fallback triggers |
| Goal Misgeneralization | Satisfying the literal task while violating safety, privacy, or social constraints | Task-constraint layers, value-conflict cases, and human-confirmation mechanisms |
| Long-Horizon Risk Accumulation | Each action looks reasonable, but the multi-step plan leads to a hazardous state | Rolling plan evaluation, reachability analysis, and runtime-monitoring logs |
| Multimodal Attack Surface | Text, images, stickers, speech, or environmental prompts influence robot behavior | Prompt-injection tests, environment-induced behavior tests, and cross-modal consistency checks |
3. 一个实用的安全架构 3. A Practical Safety Architecture
对 LLM/VLA 机器人来说,我不建议把大模型直接接到动作执行层。更合理的是建立多层约束:
For LLM/VLA robots, I would not connect a foundation model directly to the action-execution layer. A more defensible approach is to build multiple layers of constraints:
大模型可以参与任务理解和候选计划生成,但不应该拥有无约束的最终执行权。 真正的安全边界需要由任务规则、物理约束、环境状态、用户确认和运行时监控共同决定。
A foundation model may support task understanding and candidate plan generation, but it should not hold unconstrained final authority over execution. The actual safety boundary should be determined jointly by task rules, physical constraints, environment state, user confirmation, and runtime monitoring.
4. 与移动物理 AI 测评的连接 4. Connection to Mobile Physical AI Evaluation
决策安全不能只靠问答 Benchmark。机器人真正执行任务时,语言、视觉、几何、动力学、人群行为和环境规则会叠加在一起。
Decision safety cannot be evaluated only through question-answering benchmarks. When a robot executes real tasks, language, vision, geometry, dynamics, human behavior, and environmental rules are coupled.
因此,机器人 SOTIF 的测评应该包括:自然语言任务变体、对象属性变化、空间约束变化、人员行为扰动、异常提示注入、 以及跨步骤风险积累。
Robot SOTIF evaluation should therefore include natural-language task variants, object-property changes, spatial-constraint changes, human-behavior disturbances, adversarial or abnormal prompt injection, and risk accumulation across steps.
5. 后续可推进的研究与落地方向 5. Research and Deployment Directions
- 建立 LLM/VLA 机器人任务语义风险分类,而不是只看任务成功率。
- Build a semantic-risk taxonomy for LLM/VLA robot tasks, rather than measuring only task success rate.
- 设计可解释的拒绝与确认机制,让机器人在不确定时能合理停下来。
- Design explainable refusal and confirmation mechanisms, so a robot can stop appropriately under uncertainty.
- 把提示注入、环境诱导和多模态误导纳入物理 AI 安全测试。
- Include prompt injection, environmental inducement, and multimodal deception in physical-AI safety tests.
- 把运行时日志设计成安全证据,而不只是调试信息。
- Design runtime logs as safety evidence, not merely as debugging traces.
- 把模型能力边界、物理能力边界和用户授权边界放在同一个安全案例里。
- Place model capability boundaries, physical capability boundaries, and user-authorization boundaries within one safety case.