יום שישי, 9 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

CompOrca: תיקון סקאלה של תוויות תקינות למידת-אינסטרוקט

CompOrca: Corpus-Scale Compliance Labelling of Instruction-Tuning Data
CompOrca: תיקון סקאלה של תוויות תקינות למידת-אינסטרוקט. קורפוס של 4.2 מיליון דגימות שנסווגו לתקינות או לא-תקינות.
תקציר מקורי באנגליתarXiv:2609.37807v2 Announce Type: replace-cross Abstract: Studying how fine-tuning shapes refusal and noncompliance behaviour requires knowing which training examples refuse or otherwise fail to fulfil the request. Existing annotations cover evaluation sets, which are far smaller than training corpora. We present CompOrca, compliance labels for all 4,233,923 examples of the OpenOrca corpus. Every example was classified as compliant or noncompliant by five passes of an open-weight LLM judge (LongCat-2.0, 1.6T parameters). The corpus is released as unanimous compliance (94.75%), unanimous noncompliance (1.28%), and nonunanimous rows (3.97%), with the raw vote counts. A single pass flags 2.7-3.2% of the corpus as noncompliant, while only 1.28% is flagged by all five, so the most ambiguous row
קרא במקור המקורי