人类如何评估用于自动列车运行中人员检测的人工智能系统：并非所有的遗漏都相同

Research

arXiv

How humans evaluate AI systems for person detection in automatic train operation: Not all misses are alike

摘要 Abstract

如果要在高安全性领域应用人工智能（AI），则需要可靠地评估其性能。本研究旨在了解人类如何评估自动列车运行中的人员检测AI系统。在三项实验中，参与者观看了靠近铁路轨道移动的人的图像序列。一个模拟的AI对所有检测到的人进行了标记，有时正确，有时不正确。参与者需要为AI的表现提供数值评分，并口头解释他们的评分。这些实验操纵了可能影响人类评分的多个因素：AI错误的类型及其合理性、受影响图像的数量、图像中人数的数量、与轨道相关的人的位置以及诱发人类评价的方法。尽管所有这些因素都会影响人类的评分，但有些效果出乎意料或偏离了规范标准。例如，影响最大的因素是人与轨道之间的相对位置，尽管参与者明确被告知AI无法处理此类信息。综合来看，结果表明人类有时会评估超出AI任务表现的内容。在进行AI系统的安全性审核时，应考虑到AI能力与人类期望之间的这种不匹配。

If artificial intelligence (AI) is to be applied in safety-critical domains, its performance needs to be evaluated reliably. The present study aimed to understand how humans evaluate AI systems for person detection in automatic train operation. In three experiments, participants saw image sequences of people moving in the vicinity of railway tracks. A simulated AI had highlighted all detected people, sometimes correctly and sometimes not. Participants had to provide a numerical rating of the AI's performance and then verbally explain their rating. The experiments varied several factors that might influence human ratings: the types and plausibility of AI mistakes, the number of affected images, the number of people present in an image, the position of people relevant to the tracks, and the methods used to elicit human evaluations. While all these factors influenced human ratings, some effects were unexpected or deviated from normative standards. For instance, the factor with the strongest impact was people's position relative to the tracks, although participants had explicitly been instructed that the AI could not process such information. Taken together, the results suggest that humans may sometimes evaluate more than the AI's performance on the assigned task. Such mismatches between AI capabilities and human expectations should be taken into consideration when conducting safety audits of AI systems.