יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

מאגר נתונים וביצועים לוויכוחים תחרותיים בסינית

Chinese Competitive Debating Dataset and Benchmark
מאגר נתונים וביצועים לוויכוחים תחרותיים בסינית. המאגר מכיל 148 משחקים, 2,698 שלבים ו-20,542 יחידות חילופי, עם הערכות מומחים. הוא בוחן את הבנתם של מודלי שפה גדולים, כגון Llama, בוויכוחים אינטראקטיביים.
תקציר מקורי באנגליתarXiv:2609.21637v2 Announce Type: replace-cross Abstract: Debate adjudication requires tracking how arguments develop through interaction, yet existing datasets rarely combine fine-grained debate transcripts with professional judgments collected during real competitions under a shared rubric. We introduce a dataset and benchmark for evaluating large language models' understanding of competitive Chinese-language debate at the match, stage, and speaker levels. We organized 182 matches and recruited 120 professional judges, with each match independently adjudicated by three judges using a predefined rubric. After excluding matches with incomplete records, the dataset contains 148 matches, 2,698 stages, and 20,542 exchange units, with manually verified transcripts and segmentation. It preserve
קרא במקור המקורי