כתבה
arXiv cs.LG ·
LatentMD: בדיקת כישלונות גבולות Markdown בטקסט שנוצר על ידי LLM
LatentMD: Benchmarking Markdown Boundary Failures in LLM-Generated Text
LatentMD הוא בנץ'מארק ופרוטוקול הערכה לאבחון כישלונות גבולות Markdown בטקסט שנוצר על ידי מודלי שפה גדולים. הוא מאפשר לגלות פלטים שתוכנם נכון אך גבולותיהם שבורים. הבנץ'מארק כולל 4,179 פרומפטים ו-CLI לציון פלטים של מודלים שונים.
תקציר מקורי באנגליתarXiv:2609.06993v2 Announce Type: replace-cross Abstract: Large language models (LLMs) increasingly generate Markdown that is consumed by renderers, agents, code extractors, and structured downstream pipelines. Yet existing evaluations often conflate content quality with format adherence, leaving Markdown boundary failures under-measured. We introduce LatentMD, a benchmark and evaluation protocol for diagnosing CommonMark-level fence-boundary failures in LLM-generated Markdown. LatentMD separates content correctness from boundary correctness, enabling detection of outputs that are content-correct but boundary-broken. The benchmark contains 4,179 prompts and a CLI for scoring arbitrary model outputs. Across 9 LLMs and roughly 37,600 generations, we find that Markdown boundary failures are w
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית