כתבה
arXiv cs.AI ·
הלוטרקיה של הביקורת: הערכת מדד תצפיתי של רעש בביקורת
The Review Lottery: Benchmarking an Observational Estimator of Peer-Review Noise (ICLR 2017-2025)
במאמר זה נבחן מדד תצפיתי של רעש בביקורת, המבוסס על נתוני ביקורת ציבוריים. המדד נבחן ב-9 שנים של ICLR (2017-2025) ומציג תוצאות מעניינות.
תקציר מקורי באנגליתarXiv:2610.06591v2 Announce Type: replace Abstract: How much of a conference accept/reject decision would change if the same paper were reviewed by a different set of reviewers? Running a second independent program committee is the gold standard for answering this, but it is prohibitively expensive: done only twice (NeurIPS 2014 and 2021). We build an observational estimator of this quantity from public review data alone, calibrate it twice, and apply it to nine years of ICLR (2017-2025; 36,113 papers, 134,912 reviews). The estimator decomposes scores with a Bayesian ordered-probit model into paper quality and reviewer noise, maps scores to decisions with a logistic model, and simulates two independent committees (posterior draws B=1,000; committee sizes k=2,3,4). Estimated disagreement ra
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית