יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

CompOrca: תיוג מחויבות בקנה מידה גדול

CompOrca: Corpus-Scale Compliance Labelling of Instruction-Tuning Data
CompOrca הוא כלי תיוג מחויבות בקנה מידה גדול, שמסווג דוגמאות אימון כמחויבות או לא מחויבות. הכלי משתמש במודל LongCat-2.0 עם 1.6T פרמטרים. התוצאות זמינות באתר Hugging Face.
תקציר מקורי באנגליתarXiv:2609.37807v1 Announce Type: new Abstract: Studying how fine-tuning shapes refusal and noncompliance behaviour requires identifying training examples that refuse, evade or otherwise fail to fulfil the requested task. But existing annotation covers evaluation sets of a few thousand prompts at most. We present CompOrca, a compliance labelling over the entirety of the 4,233,923-example OpenOrca corpus. Every example was classified as compliant or noncompliant by five independent passes of an open-weight LLM judge (LongCat-2.0, 1.6T parameters), and the corpus is released as unanimous compliance (94.75%), unanimous noncompliance (1.28%), and nonunanimous rows (3.97%) along with the raw vote counts. A single pass flags 2.7-3.2% of the corpus as noncompliant, while only 1.28% is flagged by
קרא במקור המקורי