יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

דגימה מגוונת אסטרטגית לאימון עצמי

Strategically Diverse Sampling for Self-Training
חוקרים מציגים שיטה חדשה לדגימה מגוונת אסטרטגית לאימון עצמי. השיטה, GROOT, מייצרת נתונים מגוונים יותר מאשר שיטות דגימה מסורתיות. הניסויים הראו שהשיטה החדשה משפרת את ביצועי המודלים, אפילו כאשר הנתונים המדגמים אינם נכונים.
תקציר מקורי באנגליתarXiv:2609.31571v1 Announce Type: new Abstract: Many LLM training and inference methods, including RL and test-time scaling, depend on repeated sampling, but benefit only when the responses meaningfully differ. Self-training faces the same challenge: training data is typically constructed by sampling IID responses and filtering primarily for correctness, thereby overrepresenting strategies a model already favours. We investigate strategic diversity, or substantive variation among approaches to a problem, as an alternative principle for constructing self-training data. We generate strategically diverse data with two sampling methods: GROOT, a new method which constructs a hierarchical tree of approaches and samples distinct paths, and Verbalized Sampling (VS), adapted to produce an unstruct
קרא במקור המקורי