כתבה
arXiv cs.AI ·
בחירת ספרה או שופט חדש? הערכת תוצאות חריג של מודלי תגובה בבנקאות והשקעות
Better Deck or Different Judge? Evaluating Agentic Harness Gains in Corporate and Investment Banking
במאמר זה, נבחן את התוצאות של חריג של מודלי תגובה בבנקאות והשקעות. החריג, המשתמש במודלי תגובה של 27 מיליארד תווים, כולל חישובים פיננסיים, רטרוספקציה של נתונים ובדיקה של תגובות. המאמר נותן תוצאות חיוביות לשימוש בחריג זה, אך גם חושף חולשות בהערכת תוצאות.
תקציר מקורי באנגליתarXiv:2609.39958v1 Announce Type: new Abstract: Corporate and investment banking teams use presentations to support credit decisions and advise clients on financing and transactions. Producing these decks requires reconciling financial data, tracing sources and turning analysis into a recommendation. We retrospectively study the development of an agentic harness combining a 27B language model, financial calculations, narrative templates and validation checks. LLM judges guide engineering changes and assess the resulting decks, raising the question of whether higher scores reflect better documents or changes in grading. In shared-session text-only grading with template markers removed, five judges score the complete system 20.4 to 33.6 points out of 95 above the same model generating direct
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית