כתבה
arXiv cs.LG ·
תכנון מערכת הערכה רב-שלבית ל-LLM לבדיקת AI גנרטיבי בתגלית תרופות
Designing a Robust LLM-Based Evaluation System for Agentic AI in Drug Discovery Through Human Alignment
במאמר זה, נציגים מערכת הערכה רב-שלבית ל-LLM לבדיקת AI גנרטיבי בתגלית תרופות. המערכת כוללת חמש תרומות, כולל הגדרת ארבעה סדרי ערך לבדיקת תוצאות LLM, ואימות של השופט דרך מחקר התאמה אנושי.
תקציר מקורי באנגליתarXiv:2608.21057v2 Announce Type: replace Abstract: Agentic large language model (LLM) systems are reshaping scientific workflows in chemistry and drug discovery, but evaluating their open-ended, tool-augmented outputs remains a fundamental bottleneck. The LLM-as-a-Judge paradigm has emerged as a scalable alternative, but existing drug discovery benchmarks deploy LLM judges without validating their alignment with human experts. In this work, we present an LLM-as-a-Judge evaluation framework for ChatInvent, an agentic drug discovery assistant deployed at AstraZeneca, with five contributions. First, we define four output-quality evaluation dimensions---Completeness, Relevancy, Structural Clarity, and Scope Adherence---alongside deterministic Tool Call Correctness checks. Second, we validate
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית