כתבה
arXiv cs.CL ·
Hermes: למידת תהליכי היגיון הקשורים להקשר
Hermes: Learning Contextual Reasoning Unlocks Test-Time Scaling
Hermes הוא משפחת הרנסות פשוטות וגמישות שמאפשרות למודלים לקבל החלטות על הקצאת הקשר ושימוש במידע. המחקר מציג את Hermes-Learn, שיטה ללמידת יכולות אלו. תוצאות הניסויים מראות כי מודלים מסוגלים לנצל את הגמישות הזו כדי לשפר ביצועים עם חישוב נוסף בזמן המבחן.
תקציר מקורי באנגליתarXiv:2609.38332v1 Announce Type: cross Abstract: Test-time scaling improves model performance by allocating additional compute during inference. Using this compute effectively across multiple context windows requires deciding how to allocate fresh contexts and what information to carry between them. We call a model's ability to make these decisions contextual reasoning. Existing approaches largely prescribe these decisions through their harness; we instead shift them to the model. We introduce 1) Hermes, a family of simple, configurable harnesses that progressively varies model control over context allocation and reuse, and 2) Hermes-Learn, a two-stage framework for learning these capabilities. We find that capable models can exploit this flexibility to scale with additional inference-tim
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית