כתבה
arXiv cs.AI ·
מאגר נתונים וביצועים לוויכוחים תחרותיים בסינית
Chinese Competitive Debating Dataset and Benchmark
מאגר נתונים וביצועים לוויכוחים תחרותיים בסינית. המאגר מכיל 148 משחקים, 2,698 שלבים ו-20,542 יחידות חילופי, עם הערכות מומחים. הוא בוחן את הבנתם של מודלי שפה גדולים, כגון Llama, בוויכוחים אינטראקטיביים.
תקציר מקורי באנגליתarXiv:2609.21637v2 Announce Type: replace-cross Abstract: Debate adjudication requires tracking how arguments develop through interaction, yet existing datasets rarely combine fine-grained debate transcripts with professional judgments collected during real competitions under a shared rubric. We introduce a dataset and benchmark for evaluating large language models' understanding of competitive Chinese-language debate at the match, stage, and speaker levels. We organized 182 matches and recruited 120 professional judges, with each match independently adjudicated by three judges using a predefined rubric. After excluding matches with incomplete records, the dataset contains 148 matches, 2,698 stages, and 20,542 exchange units, with manually verified transcripts and segmentation. It preserve
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית