יום שישי, 9 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

אינדקס: סקור זה לא פוליסי: מדידת ערך של תיקון תקין

A Score Is Not a Policy: Measuring the Value of Adaptive Revision
במאמר זה, חוקרים חוקרים את ערך של תיקון תקין במערכות סיבוב-על. הם מציעים שיטה למדידת ערך של תיקון תקין, ומדגימים את השיטה במערכת סיבוב-על. המאמר כולל גם דיון בערך של תיקון תקין במערכות סיבוב-על, ומציע שיטה למדידת ערך של תיקון תקין.
תקציר מקורי באנגליתarXiv:2609.00874v2 Announce Type: replace Abstract: As agentic systems become compound systems, increasingly important decisions move above task execution itself: when should a higher-level controller preserve the strategy guiding another process, and when should it revise it? We study this meta-level control problem in a hierarchical latent reasoner whose manager can retain or replace a commitment governing lower-level computation. Across three precommitted training seeds, learned revision timing produces qualitatively different policies, ranging from an almost deterministic early clock to substantially more state conditioned schedule distributions, yet none outperforms the best forced timing policy evaluated on the same frozen checkpoint. This separates state dependence from decision val
קרא במקור המקורי