יום חמישי, 8 באוקטובר 2026 LIVE
AI־INFO

כתבה MarkTechPost ·

NVIDIA: PivotOPD מלמד סוגי סוכני AI להתאושש מטעויות קריטיות

NVIDIA PivotOPD Teaches Multi-Turn AI Agents to Recover From Pivotal Mistakes
חוקרי NVIDIA הציגו את PivotOPD, שיטת תרגול על-פוליצי עבור סוכני AI בעלי תגובה רב-סיבתית, לצורך התאוששות מטעויות קריטיות. PivotOPD הציגה את הביצועים הטובים ביותר ב-13 בסיסים, והצליחה להתאושש מ-72.7% מהטעויות הקריטיות, ולעומת זאת 20.3% לשיטת OPD הרגילה.
תקציר מקורי באנגליתNVIDIA researchers, with Princeton University and the University of Maryland, have introduced PivotOPD , an on-policy distillation method for multi-turn LLM agents. PivotOPD on-policy distillation trains an agent to avoid its most damaging early mistake, and to recover when it happens anyway. Against 13 baselines, it posts the best average on ALFWorld, WebShop and Search-based QA for Qwen3-1.7B and Qwen3-8B students. The takeaway: recovery is learnable, and standard OPD rarely teaches it. TL;DR Size: A training method, not a model. Tested on Qwen3-1.7B and Qwen3-8B students, plus a Nemotron-3.5-SFT student on SWE-Bench Verified. Runs on: Trained on NVIDIA H100 nodes. Adds 0 inference cost, so the trained agent runs wherever its base model runs. Performance: First on all 8 per-benchmark ave
קרא במקור המקורי