כתבה
arXiv cs.LG ·
Q-Steer: הנחיה לאופטימיזציה מולקולרית
Q-Steer: Action-Value Guidance for Molecular Policy Optimization
Q-Steer הוא כלי להנחיה בזמן ריצה, שמשפר את תהליך האופטימיזציה המולקולרית. הוא משתמש ב-PAVS-Q, מודל שאומן מראש, כדי להעריך את התגמול המוענק עבור כל פעולה. Q-Steer משפר את התוצאות ב-18-20 משימות, עם רווחים ממוצעים של +0.033 עד +0.049.
תקציר מקורי באנגליתarXiv:2607.26391v1 Announce Type: new Abstract: Oracle-limited molecular optimization gives reward only after a complete molecule is generated, while each rollout requires many local next-token decisions. This delayed-feedback interface makes molecular policy optimization myopic: an optimizer can learn that a molecule was good without knowing which intermediate actions made it good. We introduce Q-Steer, a rollout-time action-value steering primitive for molecular language models. Q-Steer uses an offline-trained and frozen prefix-action value scorer, PAVS-Q, that estimates the downstream reward of taking a candidate next token under a partial SMILES prefix, then adds a normalized value bonus to sampling logits. The optimizer update rule and online oracle budget are unchanged; the claim is
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית