יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

אופטימיזציה רגולריזציה של דרגה לפי רנק-מאסקד לפוליצי ללמידה מחדש בזמן המבחן לפי קוד

Entropy-Regularized Rank-Masked Policy Optimization for Test-Time Reinforcement Learning in Code Generation
אופטימיזציה רגולריזציה של דרגה לפי רנק-מאסקד לפוליצי ללמידה מחדש בזמן המבחן לפי קוד. המחקר מציע פתרון לבעיית הלמידה מחדש בזמן המבחן לפי קוד, על ידי פיתוח פונקציית פרסום חדשה. הפונקציית פרסום החדשה מבוססת על חיפוש נרחב בקוד, ומסוגלת לגלות קוד חדש ולהציגו בצורה יעילה. המחקר נערך על ידי צוות מחקר באוניברסיטת [Gemini] והתפרסם בכתב העת [arXiv].
תקציר מקורי באנגליתarXiv:2609.09135v1 Announce Type: new Abstract: Existing methods for test-time reinforcement learning (TTRL) derive rewards from answer-level self-voting on unlabeled test-time tasks with canonical answers, but this breaks down for code generation because programs cannot be compared by surface form and therefore do not directly provide a usable training signal. To make TTRL applicable to code generation, we propose probe-driven TTRL, which constructs output-free probe inputs from the problem statement, executes candidate programs on these probes, and defines a Probe Consensus Reward (PCR) from the resulting behavioral agreement. PCR provides a behavioral training signal for open-vocabulary programs, but it is not a fully reliable verifier and remains susceptible to reward hacking through s
קרא במקור המקורי