יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

אימון מה שמופעל

Train What You Deploy:Token-Faithful Post-Training of a Production Coding
פותח כלי לאימון מחדש של סוכנים לקידוד, המשמר אמינות ומקטין שגיאות. המחקר מציג גישה חדשה לאימון, המבטיחה תוצאות טובות יותר.
תקציר מקורי באנגליתarXiv:2609.04678v1 Announce Type: new Abstract: Existing post-training pipelines for coding and terminal agents suffer severe token and control fidelity errors: simplified training environments mismatch production deployments, and offline token reconstruction from agent logs distorts original prompts and conflates policy calls with background model operations. We present a fidelity-aware training coupling framework that retains trainer-side sampling over original prompts, eliminates spurious model calls via a negotiated training protocol, and restricts loss computation to verifiable token spans with closed-failure guarantees. We further propose Certified Divergence Proximal Policy Optimization (C-DPPO), which establishes tight two-sided TV certification bounds, adaptive-K rules, budget-awa
קרא במקור המקורי