יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

בדיקת עמידות ללחץ של כריית נתונים אפקטיבית: כאשר הישעות בחישוב משנים את המסקנות של הבסיס

Stress-Testing Efficient Responsible-AI Evaluation: When Compute Savings Change Benchmark Conclusions
בדיקת עמידות ללחץ של כריית נתונים אפקטיבית לבסיסיות רספקטיבית של AI. המאמר עוסק בבדיקת עמידות ללחץ של כריית נתונים אפקטיבית לבסיסיות רספקטיבית של AI. המאמר עוסק בבדיקת עמידות ללחץ של כריית נתונים אפקטיבית לבסיסיות רספקטיבית של AI.
תקציר מקורי באנגליתarXiv:2608.31108v2 Announce Type: replace Abstract: Efficient evaluation changes the protocol used to support claims about model behavior, yet it is rarely tested whether those claims remain stable after the evaluation itself is made cheaper. We stress-test conclusion robustness in responsible-AI benchmarking by evaluating three dense and mixture-of-experts models on BBQ and BBQ-V under seven conditions spanning batching, quantization, benchmark reduction, and their combinations. Rather than treating preserved aggregate accuracy as sufficient, we compare accuracy, bias severity and prevalence, reasoning quality, subgroup behavior, subset-membership stability, runtime, and measured GPU energy against a full-benchmark BF16 baseline. Larger batching keeps accuracy within 0.35 percentage point
קרא במקור המקורי