כתבה
arXiv cs.AI ·
מדוע סוכן מחקר עמוק שלך נכשל?
Why Your Deep Research Agent Fails? On Hallucination Evaluation in Full Research Trajectory
חוקרים הציגו שיטה חדשה לאבחון נקודות כשל בסוכני מחקר עמוק. השיטה מבוססת על מודל PING Taxonomy, המסווג את ההזיות לארבע קטגוריות. החוקרים פיתחו גם את DeepHalluBench, מאגר נתונים לבדיקת הזיות.
תקציר מקורי באנגליתarXiv:2601.22984v3 Announce Type: replace Abstract: Diagnosing failure patterns in Deep Research Agents (DRAs) remains a critical challenge. Existing benchmarks predominantly rely on end-to-end evaluation, obscuring intermediate hallucinations that accumulate throughout the research trajectory. To bridge this gap, we propose a shift from outcome-based to process-aware evaluation by auditing hallucinations in the full plan-search-summarize trajectory. We introduce the PING Taxonomy, which categorizes DRA hallucinations into four complementary types: Propagation, Intent, Noise-induced, and Grounding. We further instantiate this taxonomy into a fine-grained evaluation framework that decomposes trajectories into atomic actions, claims, and sub-queries for rigorous verification, and we validate
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית