כתבה
arXiv cs.AI ·
כשגורמי חיפוש מדעיים נכשלים: ניתוח נקודות תפנית של חשיפה וניסיון חקירה
Where Scientific Search Agents Fail: Decision-Checkpoint Auditing of Exposure and Inspection Attempts
גורמי חיפוש מדעיים נכשלים להגיע לעבודות מטרה או לחזור עם תשובות נכונות. ניתוח נקודות תפנית חשף נקודות חולשה בגורמי חיפוש רגילים ובגורמי חיפוש קודם-קריאה.
תקציר מקורי באנגליתarXiv:2609.38670v1 Announce Type: new Abstract: Final-answer accuracy does not reveal whether a scientific-search agent failed to encounter a target paper, attempt to inspect it, or return an accepted answer after inspection. We introduce decision checkpoints that record observations and tool actions without benchmark labels during inference, then join target identities and evaluator labels to assign outcome categories from recorded events. Across five conditions on 540 answerable AutoResearchBench Deep questions in a fixed, target-enriched environment, keyword search achieves 24.6\% accuracy, compared with 17.8\% for raw search. The keyword condition has fewer incorrect answers with neither target exposure nor inspection, but more incorrect answers after the target is exposed and left uni
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית