יום רביעי, 7 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

Transect: שימור תצפית עבור הערכת LLM Agent ארוכת טווח

Transect: Retaining Observability for Long-Horizon LLM Agent Evaluations
Transect הוא חבילה פתוחה המסייעת למעריכים להבין איך התנהלה ריצה ארוכה של סוכן. היא מאפשרת לזהות התנהגויות ראויות לחקירה ולבדוק פרשנויות נגד התעתיק. Transect משתמשת ב-Inspect Scout ומספקת דו''חות ניווטים המקבילים אירועים רשומים, שימוש בטוקנים, פעילות תת-סוכנים ותוויות התנהגות נוצרות על ידי מודלים.
תקציר מקורי באנגליתarXiv:2610.08364v1 Announce Type: new Abstract: Frontier AI evaluations increasingly use open-ended, agentic, long-horizon tasks whose transcripts can span hundreds of pages of outputs and actions from complex multi-agent networks. The observability envelop-the range of what evaluators can reliably infer about an agent's behaviours-is therefore narrowing. Language model assistants can help classify and interpret agent behaviour but also afford human evaluators significant analytical degrees of freedom, threatening the reproducibility and auditability of language-model-based transcript analysis. Transect is an open source package built on Inspect Scout to help evaluators understand how a long agent run unfolded, identify behaviour worth investigating, and check interpretations against the t
קרא במקור המקורי