יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

כל דבר במידה: אופטימיות של עבודה בתחומים ופערים חסימה-לא-תלויי-תחום באמצע-לימוד

Everything in Moderation: Per-Domain Coverage Optima and Alignment-Resistant Domain Gaps in Multi-Domain Mid-Training
במחקר זה נבדקה השפעת תפוצת הנתונים באמצע-לימוד על יכולת הדגמים. נמצא כי תפוצה ממוצעת (10%-40%) היא האופטימלית לכל תחום. נקבע כי פערים חסימה-לא-תלויי-תחום נותרו גם לאחר עבודת התאמה. נמצא כי תפוצה של 0% גורמת ליפוי ביצועים. המחקר עוסק במודל Qwen.
תקציר מקורי באנגליתarXiv:2609.09081v1 Announce Type: new Abstract: Mid-training, the stage between pre-training and alignment, is where a model's per-domain data composition is typically set by data availability rather than principled design. We ask what that decision buys, and whether a later alignment pass can undo it. In a controlled logical-reasoning setting (Qwen3-8B-Base, with a 4B replication; five semantically rule-disjoint KOR-Bench domains) we train 30 allocations spanning the five-domain simplex, 24 sweep configurations plus six withheld from the fit, at five seeds each. Three findings emerge. First, every domain has an interior coverage optimum: the moderate band ($10\%$-$40\%$) is best for all five domains, and a calibrated permutation test for quadratic interiority gives $P\approx0.010$; the fi
קרא במקור המקורי