יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

KC-Bench: בנק המבחן הדינמי האינטראקטיבי לבדיקת סתירות ידע באג'נטים LLM

KC-Bench: A Dynamic Interactive Benchmark for Evaluating Knowledge Conflicts in LLM Agents
בנק המבחן KC-Bench מבדיק את יכולת האג'נטים LLM להתמודד עם סתירות ידע, תקלות קלט וסתירות זמניות. הבנק המבחן כולל 238 משימות שנבחנו ומשלבים רכיבים שונים, כולל רכיבי ניתוח טקסט, רכיבי סימולציה ורכיבי ניתוח זמן.
תקציר מקורי באנגליתarXiv:2609.03588v1 Announce Type: new Abstract: As LLMs increasingly act through tools, they must reconcile user instructions, parametric knowledge, and dynamic environmental observations before taking actions. We introduce KC-Bench, a controlled multi-turn benchmark for measuring this capability across world-knowledge conflicts, input inconsistencies, and multi-source temporal conflicts. Its 238 tasks are manually screened from more than 1,000 generated candidates and combine a user simulator, stateful tools, deterministic environment assertions, an open-source natural-language evaluator, and human trajectory verification. Evaluation of nine models, including DeepSeek-V4-Flash, GLM-5.2, and MiniMax-M3, shows substantial cross-domain variation: no model handles factual correction, identity
קרא במקור המקורי