כתבה
arXiv cs.CL ·
When Does a Second Model Help? Cross-Model Review in LLM Verification
תקציר מקורי באנגליתarXiv:2610.01471v2 Announce Type: replace Abstract: Large language models now generate code, documentation, and analyses, and are increasingly used to review such output. We ask when a second review by a different model helps. Building on the author's earlier preprints, which varied context, repetition, and role structure within one model, we test model independence in a controlled experiment: 30 artifacts with 150 planted errors, 10 review conditions, and 900 review sessions with three reviewer models from two developers. In this experiment, (1) a top-tier cross-model reviewer is not significantly different in F1 from same-model review in a fresh session (CCR), which does not establish equivalence; (2) the two find partly different errors (Jaccard 41.2%); and (3) at two review calls, one
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית