יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

TRACE: בנק מיומנויות עצמי-מתפתח

TRACE: A Self-Evolving Skill Bank for Consistent, Limit-Aware LLM Agents
TRACE הוא בנק מיומנויות עצמי-מתפתח המשפר את עקביותם של סוכנים LLM. הוא משתמש ב-GPT-5.5 ומשפר את העקביות ב-34.6%. הוא זכה במקום הראשון ב- CAR-bench.
תקציר מקורי באנגליתarXiv:2608.22793v2 Announce Type: replace Abstract: Reliable deployment of LLM agents in user-facing products depends not on raw task-solving ability but on consistency and limit-awareness: behaving the same way across repeated trials, and recognizing when a request cannot, or cannot yet, be safely fulfilled. CAR-bench exposes this reliability gap in the domain of in-car assistants: an LLM-simulated user issues incomplete or ambiguous requests, requiring the agent to resolve uncertainty through multi-turn dialogue and tool use while strictly adhering to domain policies. Even frontier models show a substantial gap between what they can solve at least once (Pass@3) and what they solve consistently across trials (Pass^k). We bridge this gap with TRACE (TRAjectory-Contrastive Evolution), which
קרא במקור המקורי