כתבה
arXiv cs.LG ·
מה אחר צריך לתקן? החקירה של חישוב זמן-הצצה זול להפצת תיקונים באמצעות שיחה
What Else Needs Fixing? Exploring Cost-Effective Test-Time Compute for Revision Propagation in Artifacts Generated Through Conversation
במאמר זה, חוקרים חקרו את יכולת ה-LLMs להפצת תיקונים באמצעות שיחה. הם גם חקרו חישוב זמן-הצצה זול להפצת תיקונים. התוצאות הראו שהבסיסים הגיעו לדיוק של 68.3--93%, והשיטה היעילה ביותר הייתה לבחור מבין שלושה דגימות נקודתיות שנבחרו באמצעות LLM-based או medoid selection, שהשיגה דיוק של 2.2--9.7%.
תקציר מקורי באנגליתarXiv:2609.03254v1 Announce Type: cross Abstract: Large Language Models (LLMs) often help users generate artifacts through iterative cycles of generation and revision in conversation. A challenge here is that, when users specify only a local change during revision, LLMs must instead identify the relevant dependencies and propagate the revision to all affected parts of the artifact. This paper studies this ability of LLMs on conversationally generated artifacts, where the artifact context and its dependencies may be embedded in the conversation history. Toward practical use, we also explore cost-effective test-time compute for this new setting. Specifically, we introduce a new benchmark for this setting, and evaluate nine revision methods, including sequential reflection and parallel sampli
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית