כתבה
arXiv cs.LG ·
MARCO: Multi-Round Agentic Reinforcement for Conditional Molecular Optimization
תקציר מקורי באנגליתarXiv:2609.36683v1 Announce Type: new Abstract: Molecular optimization is inherently iterative: a candidate is proposed, evaluated against several objectives, and revised while preserving a relationship to the source molecule. Most instruction-following models instead emit one edited molecule, forcing validity, property improvement, and similarity control into a single response. We introduce MARCO, an evaluator-grounded reinforcement-learning framework that trains molecular editors on bounded proposal--feedback--revision trajectories. MARCO aggregates shaped turn rewards into an undiscounted trajectory return for group-relative policy optimization. We evaluate two consequences of this training: Same-1 tests the trained policy under a one-response budget, while Same-5 tests whether the same
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית