כתבה
arXiv cs.LG ·
When Compliance Data Masquerades as Evaluation: Measurement Validity for Deployed AI Systems
תקציר מקורי באנגליתarXiv:2609.13642v1 Announce Type: new Abstract: We argue that a recurring failure in the evaluation of deployed AI systems occurs when data collected for operational monitoring or regulatory compliance are interpreted as if they were designed for comparative evaluation. Automated driving provides a concrete example of this problem. U.S. disengagement and crash-reporting regimes produce valuable operational evidence, but differences in reporting scope, exposure, deployment domain, event capture, and comparator construction limit the safety claims that can be supported from these measurements alone. We frame this issue as a measurement-validity problem in AI evaluation rather than as a transportation-specific data limitation. We argue that comparative claims about deployed AI systems require
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית