כתבה
arXiv cs.CL ·
HumorRank: A Tournament-Based Leaderboard for Evaluating Humor Generation in Large Language Models
תקציר מקורי באנגליתarXiv:2604.19786v3 Announce Type: replace Abstract: Evaluating humor in large language models (LLMs) is an open challenge because existing approaches yield isolated, incomparable metrics rather than unified model rankings, making it difficult to track progress across systems. We introduce HumorRank, a tournament-based evaluation framework and leaderboard for textual humor generation. On two public benchmarks (SemEval-2026 MWAHAHA and Humor Transfer Bench), we conduct extensive automated pairwise evaluation across nine models spanning proprietary, open-weight, and specialized systems. Pairwise judgments are produced by LLM judges grounded in the General Theory of Verbal Humor (GTVH): each judge integrates structured comedic analysis into adjudication, jointly yielding a preference decision,
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית