יום חמישי, 20 באוגוסט 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

Reconstruction: A Blind Benchmark for Recovering Research Ideas from Pre-Publication Bibliographies

תקציר מקורי באנגליתarXiv:2608.16645v1 Announce Type: cross Abstract: Can a language model recover the true research idea of a published paper when given only that paper's pre-publication bibliography? We introduce Reconstruction, a blind idea-recovery benchmark that withholds the seed paper and all contemporaneous or future literature, and asks models to propose hypotheses that an independent large language model judge matches against the held-out ground-truth idea. A strict anti-leakage protocol-temporal citation cutoff, anonymous reference IDs, and frozen per-paper bibliographies, which prevents prompt-time leakage of the seed idea. Across six scientific domains and 643 evaluated papers, seven frontier models achieve only modest Match rates (approx. 3-15%). We then evaluate a reference-only multi-agent (to
קרא במקור המקורי