כתבה
arXiv cs.LG ·
Counterfactual Shapley Credit Assignment
פותח כלי חדש לחלוקת אשמה בלמידת תגמול. המערכת מסתמכת על תורת הסיבה והתוצאה כדי לייחס אשמה וזכות בצורה מדויקת יותר.
תקציר מקורי באנגליתarXiv:2607.16999v1 Announce Type: new Abstract: The Credit Assignment Problem (CAP) is fundamental to developing efficient and explainable Reinforcement Learning (RL) agents. Existing frameworks, whether relying on temporal contiguity or hindsight-conditioned reward reweighting, frequently fail to attribute properly between an agent's policy (skill) and environmental stochasticity (luck). A principled approach to CAP must isolate the true causal drivers of observed outcomes from spurious correlations and environmental randomness. We introduce Counterfactual Shapley Credit Assignment, a novel framework grounded in causal theory that attributes credit and blame via the Counterfactual Shapley Value ($\phi$-value). By redistributing environmental rewards, $\phi$-values enhance temporal credit
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית