כתבה
arXiv cs.AI ·
Transect: שימור ניתובים לבירוריות של LLM Agent Evaluations
Transect: Retaining Observability for Long-Horizon LLM Agent Evaluations
Transect שומר ניתובים לבירוריות של LLM Agent Evaluations, כולל שימור ניתובים לבירוריות של LLM Agent Evaluations.
תקציר מקורי באנגליתarXiv:2610.08364v2 Announce Type: replace Abstract: Frontier AI evaluations increasingly use open-ended, agentic, long-horizon tasks whose transcripts can span hundreds of pages of outputs and actions from complex multi-agent networks. The observability envelop-the range of what evaluators can reliably infer about an agent's behaviours-is therefore narrowing. Language model assistants can help classify and interpret agent behaviour but also afford human evaluators significant analytical degrees of freedom, threatening the reproducibility and auditability of language-model-based transcript analysis. Transect is an open source package built on Inspect Scout to help evaluators understand how a long agent run unfolded, identify behaviour worth investigating, and check interpretations against t
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית