יום רביעי, 7 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

כלפי התאמה: חוקי גדילה ומסגרת ראשונית

Toward Alignment Scaling Laws: A Framework and First Preregistered Measurements
חברת המחקר הציגה תוכנית חדשה למדידת חוקי גדילה של התאמה. התוכנית כוללת ניסויים ראשונים ומסגרת ראשונית.
תקציר מקורי באנגליתarXiv:2610.08540v1 Announce Type: new Abstract: Whether alignment gets easier or harder as models grow is often argued from isolated findings, as if alignment were one property. We treat it as a family of measurable scaling relations: for each risk category r, the alignment burden needed to hold a fixed safety target is modeled as B_r(N)=a_rN^alpha_r, with N a capability proxy; against a budget proportional to N, scaling helps if alpha_r<1, keeps pace if alpha_r~1, and accumulates alignment debt if alpha_r>1. We give three operationalizations of burden and distinguish observed, audited and true alignment. A toy model, in which corrections consume capability headroom, makes the consequences explicit. We prove that the largest exponent among corrected risks, not an average, sets the long-run
קרא במקור המקורי