כתבה
arXiv cs.AI ·
What Pretraining and Midtraining Make Learnable from Rewards?
תקציר מקורי באנגליתarXiv:2609.38446v1 Announce Type: cross Abstract: A reward can identify a correct answer while leaving the computation needed for new inputs undetermined. We study how pretraining and midtraining supply the information and computation that make reward adaptation effective. In sequential state computation and contextual memory, we characterize mechanisms that agree on every training reward yet demand different held-out answers. Task-independent source observations resolve this ambiguity. We construct finite sampled Adam paths from specified random initializations through source prediction and reward adaptation in the same parameters, proving how prediction acquires execution or retrieval and rewards learn their task-specific use. Experiments with pretrained Qwen2.5 checkpoints test this div
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית