יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

רגולריזציה של אנטרופיה

Entropy Regularization: A Free Correction to Cross-Entropy for Verified Demonstrations
חוקרים הציגו שיטה חדשה לשיפור דיוק מאשרים במודלים גדולים. השיטה, הנקראת רגולריזציה של אנטרופיה, משתמשת באנטרופיה של שאנון כדי למנוע התפלגות של מסה הסתברותית לפלטים שגויים. הניסויים הראו שיפור בדיוק המאשרים במטלות שונות.
תקציר מקורי באנגליתarXiv:2609.30572v1 Announce Type: new Abstract: Large language models are often post-trained on expert demonstrations using cross-entropy (CE), even when the downstream objective is not to imitate the demonstrated solution but to produce any output accepted by a verifier. This mismatch is seen in verifiable domains with multiple correct solutions, such as mathematical reasoning and code generation, where training data may contain only one expert solution per problem. We show that minimizing cross-entropy can be misaligned with minimizing verifier risk; two policies can assign identical likelihood to the observed demonstrations while placing different probability mass on incorrect outputs. This is formalized through a learning-theoretic counterexample in which CE minimization selects a subo
קרא במקור המקורי