יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

הקלון התנהגותי משיג יותר מהפענוח האנטרופי: כשל המתווך-מדריך של שיטות Actor-Critic בטיפול קדחתני

Behavioral Cloning Outperforms Entropy-Regularized RL: Critic-Driven Failure of Actor-Critic Methods on Adaptive Tumor Treatment
הקלון התנהגותי משיג יותר מהפענוח האנטרופי בטיפול קדחתני. ניתוח זה חושף כשל של שיטות Actor-Critic בטיפול קדחתני. המחקר מציע פתרון חדשני בצורת הקלון התנהגותי.
תקציר מקורי באנגליתarXiv:2609.06667v1 Announce Type: new Abstract: Adaptive dosing requires policies that reduce tumor burden without excessive toxicity. Learned dosing policies are typically judged against historical or heuristic comparators, which cannot show whether a policy has found the best behavior available. We instead study a three-population tumor-control ODE in which optimal-control analysis fixes the form of a good schedule -- bang-bang dosing punctuated by a singular arc -- and construct a numerical controller of that form as a proxy for near-optimal behavior. Judged against this reference under a sustained-cure criterion -- 200 consecutive days below 5% carrying capacity -- Soft Actor-Critic (SAC) trained from scratch never reaches cure. Behavioral cloning (BC) of the reference reproduces it (1
קרא במקור המקורי