כתבה
arXiv cs.CL ·
בחינה של אוטונומיה גבולית ב-AI רגולטורי: תקן רפואי להפעלה
Evaluating Bounded Autonomy in Regulated Agentic AI: A Diagnostic Harness with Constitutional Rewards, Escalation Labels, and Runtime Governance
במאמר זה, נציגים תקן רפואי לבחינת אוטונומיה גבולית ב-AI רגולטורי. התקן מספק שישה סימני תוחלת: תקינות ציטציה, תקינות קרקע, תקינות סכמה, תקינות עלייה, תקינות עקרונית ותקינות תגובה לא בטוחה.
תקציר מקורי באנגליתarXiv:2609.37501v1 Announce Type: new Abstract: We propose RegLLM, a diagnostic harness for bounded autonomy in regulated agentic workflows. It instruments six trustworthiness signals: citation validity, source grounding, schema compliance, escalation correctness, constitutional alignment, and unsafe-action rate. Signals are distinguished by their source of supervision: programmatic verifiers, task-level escalation labels, or AI-judge scores. A deterministic runtime supervisor blocks ungrounded answers and forces escalation, logging interventions. The same domain constitution informs evaluation, training rewards, and serving guardrails. Task-level should-escalate labels make the act-versus-defer decision a measurable training signal. We demonstrate the harness at smoke scale. An offline re
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית