יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

באיזה שלב עוזרת דגם שנייה? סיקור דגמים בווריפיקציה של LLM

When Does a Second Model Help? Cross-Model Review in LLM Verification
באיזה שלב עוזרת דגם שנייה בווריפיקציה של LLM? ניתוח של 30 תרגילים ו-3 דגמים.
תקציר מקורי באנגליתarXiv:2610.01471v1 Announce Type: cross Abstract: Large language models now generate code, documentation, and analyses, and are increasingly used to review such output. We ask when a second review by a different model helps. Building on the author's earlier preprints, which varied context, repetition, and role structure within one model, we test model independence in a controlled experiment: 30 artifacts with 150 planted errors, 10 review conditions, and 900 review sessions with three reviewer models from two developers. In this experiment, (1) a top-tier cross-model reviewer is not significantly different in F1 from same-model review in a fresh session (CCR), which does not establish equivalence; (2) the two find partly different errors (Jaccard 41.2%); and (3) at two review calls, one CC
קרא במקור המקורי