יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

למה לסמפל מה שאפשר לרשום? אופטימיזציה נאותה לכלי גנומי

Why Sample What You Can Enumerate? Exact Policy Optimization for Genomic Tool Selection
אופטימיזציה נאותה לכלי גנומי: פתרון חדש לבחירת כלים גנומיים. המחברים הציגו פתרון חדש לבחירת כלים גנומיים, המשתמש באופטימיזציה נאותה. הפתרון, הנקרא FGPO, מסוגל לבחור את הכלים הטובים ביותר לכל שאלה גנומית, ומשפר את התוצאות בהשוואה לפתרונות אחרים.
תקציר מקורי באנגליתarXiv:2609.10221v2 Announce Type: new Abstract: Reinforcement learning over a frozen reasoner has become a common recipe for teaching a policy which external tools to invoke. We show that this recipe becomes structurally mismatched in specialist scientific settings where the complete tool-subset space is enumerable. There, a small set of recurring computational capabilities covers the domain, so the space of tool subsets is combinatorial yet small enough to enumerate, and GRPO still estimates an action expectation from a handful of sampled rollouts. Worse, the approximation degrades as training succeeds: as the policy concentrates on preferred subsets it resamples them, sampled rewards collide, and the group-normalized advantage vanishes. On genomic reasoning the fraction of questions yiel
קרא במקור המקורי