יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

NxN E-valuation: הסקה ואימות של תאוריות על ידי CRT

NxN E-valuation: Hypothesis Certification via a Conformal CRT Null
מתודולוגיה חדשה לאימות תאוריות באמצעות CRT ומאגרי נתונים גדולים. השיטה מתאימה במיוחד למערכות חקירה בעזרת LLMs.
תקציר מקורי באנגליתarXiv:2608.06621v3 Announce Type: replace Abstract: We propose NxN E-valuation, a handy, e-value-based hypothesis-certification algorithm that lets a hypothesis be verified without building any case-specific certification procedure---such as constructing a dedicated null hypothesis---as long as a large enough dataset is available. The method is especially suited to LLM-based exploration systems, where LLMs are remarkably good at proposing hypotheses but suffer badly from hallucination; this hallucination prevents us from harvesting LLM outputs directly, and existing remedies each fall short. The most common solutions include letting the LLM verify or correct itself circular verification and held-out testing (where false hypotheses can still pass via spurious correlations), among other reme
קרא במקור המקורי