כתבה
arXiv cs.CL ·
Which Rollout Taught It That? BehaviorTrace and the Limits of Training-Data Attribution in Online RL
תקציר מקורי באנגליתarXiv:2610.10422v1 Announce Type: cross Abstract: When reinforcement learning teaches a language model a new behavior, can we find the training rollouts that taught it? And when an attribution method says it can, how do we know the answer is real? We study both questions on online RL fine-tuning with GRPO, using a planted behavior with a known cause. We release BehaviorTrace, an open evaluation harness that combines full-gradient sketching, the planted-behavior setup, and controls for gradient magnitude, fluency, headroom, and variation across seeds and generation draws. Across three seeds on Qwen2.5-1.5B, much of the apparent attribution signal comes from confounds. A control that ranks training steps by gradient size alone, with no behavior target, reaches 4.2 to 4.5 times chance and mat
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית