כתבה
arXiv cs.AI ·
הגעה להחלטה ללא ראיות מסומנות: ניסויי דחייה או פוסט-אימון רק על פי החלטה
Verdicts Without Annotated Evidence: Rejection Sampling or Label-Only Post-Training for Evidence Recovery?
במאמר זה נבחן כיצד יכול דגם שפה קטן לחזור על ראיות ללא סימון ידי אדם. התוצאות היו טובות, ושני שיטות שונות הצליחו לשפר את הראיות.
תקציר מקורי באנגליתarXiv:2610.06962v1 Announce Type: cross Abstract: In many review workflows the verdict is the only thing retained. The passages behind it are not marked, because that annotation costs far more than recording the decision. We measure how much of that evidence a small language model can recover when it is post-trained on the verdicts alone, with no human evidence labels at any stage. On ContractNLI the human evidence spans are held out until evaluation. Matching the recorded verdict and agreeing with those spans are not the same thing: across six systems the two scores are only weakly related and rank the systems differently, so accuracy is a poor guide when the citations have to be reviewable. Label-only training on the bare verdict reaches accuracy 0.896 and span F1 0.564. Rejection sampli
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית