יום חמישי, 8 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

אליס: מבחן גדול-ממדים גרמני לביקורת רב-ממדית של תשובות קצרות

Alice: A Large-Scale German Benchmark for Rubric-Based Multi-Dimensional Automatic Short Answer Scoring
נחשפה ספריית מבחן גדול-ממדים גרמנית לביקורת רב-ממדית של תשובות קצרות, שמכלילה שלושה תחומי עבודה: יכולת למידה, ידע ומיומנויות. המבחן נועד לבדוק את יכולתן של מודלי LLM לבצע ביקורת רב-ממדית של תשובות קצרות.
תקציר מקורי באנגליתarXiv:2610.09661v1 Announce Type: new Abstract: Automatic Short Answer Scoring (ASAS) is central to NLP for Education. However, openly available benchmarks remain scarce, and existing datasets largely address how well students answer a question directly rather than how well they master underlying concepts (knowledge elements) such as thermal energy or epistemic activities (skills) such as reasoning or claim. To address this gap, we introduce Alice, a large-scale, rubric-based German ASAS dataset that is pedagogically aligned and comprises three subtasks: (i) learning performance (Alice-LP), (ii) knowledge elements (Alice-KE), and (iii) skills (Alice-SK). We further formulate rubric-based ASAS as a rubric-retrieval task and benchmark the dataset with a range of language models, from encoder
קרא במקור המקורי