יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

מה יכולים ללמוד פרימטרינג ומידטרינג מפרסומים?

What Pretraining and Midtraining Make Learnable from Rewards?
במאמר זה נחקר כיצד פרימטרינג ומידטרינג מספקים מידע ומחשבה להתאמה לפרסומים. התוצאות כוללות תאריך: Sequential models trained with correct source and first-operation supervision reach 82.61% success, versus 44.15% for a private-random source control.
תקציר מקורי באנגליתarXiv:2609.38446v1 Announce Type: new Abstract: A reward can identify a correct answer while leaving the computation needed for new inputs undetermined. We study how pretraining and midtraining supply the information and computation that make reward adaptation effective. In sequential state computation and contextual memory, we characterize mechanisms that agree on every training reward yet demand different held-out answers. Task-independent source observations resolve this ambiguity. We construct finite sampled Adam paths from specified random initializations through source prediction and reward adaptation in the same parameters, proving how prediction acquires execution or retrieval and rewards learn their task-specific use. Experiments with pretrained Qwen2.5 checkpoints test this divis
קרא במקור המקורי