כתבה
arXiv cs.AI ·
HARDEN: חיפוש אבולוציוני מוגבל לבדיקות קצרות יותר
HARDEN: Constrained Evolutionary Search for Harder, Answer-Preserving Evaluation Cases
HARDEN היא שיטת חיפוש אבולוציוני מוגבל שמטרתה ליצור בדיקות קצרות יותר. השיטה נועדה להתאים את הקלט של בדיקות קיימות לגרסאות קשות יותר, תוך שמירה על תוצאות צפויות. HARDEN נבדקה על פי 3 גרסאות של מודל Qwen3.5 (35B-A3B, 122B-A10B, ו-397B-A17B) והציגה תוצאות טובות יותר מבסיסים יחיד-עבר.
תקציר מקורי באנגליתarXiv:2609.30571v1 Announce Type: new Abstract: Language models are often evaluated on curated benchmarks that underrepresent the complexity of enterprise deployments. We introduce HARDEN, a constrained evolutionary search method to adapt the input of existing evaluation cases into more challenging variants while keeping their expected outputs fixed. HARDEN searches along generated domain-specific complexity axes while enforcing feasibility constraints such as preserving task semantics, realism, and execution validity. Across FinQA, PubMedQA, and ContractNLI and three Qwen3.5 model scales (35B-A3B, 122B-A10B, and 397B-A17B), HARDEN reduces task-model accuracy by 22.7% on average and by up to 49.9% relative to single-pass baselines using the same feasibility checks. These results show that
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית