יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

הנדסה נקייה, מדידה לא יציבה

Clean Engineering, Unstable Measurement: A Preregistered Reliability Failure of Black-Box LLM Observers on Shared Endpoints
חוקרים בדקו את אמינותם של מודלים של LLM, ומצאו כי הם אינם יציבים. המחקר הראה כי המודלים מגיבים באופן שונה לאותו בקשה, וכי התוצאות אינן עקביות.
תקציר מקורי באנגליתarXiv:2609.04198v1 Announce Type: cross Abstract: Language-model judges now gate training data, score generations, and drive leaderboards. The judge is then a measurement instrument, resting on one rarely stated assumption: the same request, sent to the same model name, reads the same tomorrow. We audited that assumption in two preregistered campaigns with every threshold fixed in advance; neither got past validating its instrument. Across 52,988 audited request attempts, same-window repeat rankings agreed at Spearman 0.400 against a required 0.90, and byte-identical next-day replays agreed at 0.78 against a required 0.99, each time with the execution record at ceiling. Three mechanisms explain the gap: a label-to-meaning mapping that biased readouts as strongly as the signal; candidate ga
קרא במקור המקורי