יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

התמצאות-מודע: זיהוי והגנה על תפקודי תפיסה חשובים להפחתת צריכת אנרגיה של LLM

Reasoning-Aware Compression: Identifying and Protecting Vulnerable Reasoning Circuits for Energy-Efficient LLM Deployment
חברת arXiv:2609.05512v1 הציגה פרקטיקה של התמצאות-מודע להפחתת צריכת אנרגיה של LLM. הם זיהו תפקודי תפיסה חשובים ושימרו אותם ב-FP16. הם גם גילו שהפחתת צריכת אנרגיה עשויה להגדיל את הזמן של ה-LRM.
תקציר מקורי באנגליתarXiv:2609.05512v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) impose substantial energy costs during deployment, yet current compression methods apply uniform quantization across all components, risking damage to critical reasoning circuits. We present a reasoning-aware compression framework that benchmarks quantization conditions across five reasoning benchmarks, GSM8K, FOLIO, MATH-500, ProofWriter, and MuSiQue, with hardware-level GPU energy measurement; profiles per-module INT4 vulnerability across all 196-224 (layer, projection) pairs via a perturbation sweep on a held-out calibration split, then selectively restores the most sensitive circuits to FP16. Three findings emerge. First, INT4 quantization can increase energy by extending reasoning chains; a 25% power reducti
קרא במקור המקורי