כתבה
arXiv cs.LG ·
האם אתה משתן או מזכיר? בדיקה של LLMs בתחום חיפוש קוד אלגוריתמי
Are you Synthesizing or Recalling? Evaluating LLMs on Algorithmic Code Retrieval
בדיקה של LLMs בתחום חיפוש קוד אלגוריתמי. המאמר עוסק בבדיקה של 15 מודלי LLM שונים בתחום חיפוש קוד אלגוריתמי, ומציג תוצאות של 599 בעיות שונות. התוצאות מציגות כי המודלים השונים חופשיים ביכולתם לחפש קוד אלגוריתמי, וכי השיפורים ביכולת החיפוש יכולים להיות משמעותיים.
תקציר מקורי באנגליתarXiv:2610.02438v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated strong performance in code generation, where success depends on both recalling relevant algorithmic knowledge and reasoning about how to apply it. However, existing LLM pipelines are opaque, with no explicit separation between these two components. We argue that for well-known algorithms whose canonical implementations are widely accessible in pretraining corpora, code generation is better measured as \textit{parametric code retrieval}: reproducing a named algorithm from internalised knowledge rather than synthesizing a novel one. We introduce AlgoREval, a benchmark of 599 problems spanning classical 77 algorithms across 14 domains, 7 programming languages, and 4 graph-input representations to ev
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית