כתבה
arXiv cs.AI ·
קליברציה של החלטות שמשנות את העתיד: קוונטיזציה אונ-פוליסי לאחר האימון למודלי שפה גדולים רב-תכליתיים
Calibrate the Decisions That Change the Future: On-Policy Post-Training Quantization for Multimodal Large Language Models
אופציונליות חדשה לקוונטיזציה אונ-פוליסי למודלי שפה גדולים רב-תכליתיים. זה יכול לשפר את הביצועים של המודלים ולמנוע שגיאות.
תקציר מקורי באנגליתarXiv:2609.36828v1 Announce Type: new Abstract: Post-training quantization (PTQ) lowers deployment cost for multimodal large language models, but calibration typically reconstructs fixed sequences with local objectives. This overlooks autoregressive feedback: a quantization-induced token change redirects the prefix and changes future states. Yet on-policy coverage alone is insufficient because many decision mismatches barely affect future generation. We propose OnPTQ, an on-policy framework that calibrates on trajectories visited by the current quantized policy. On shared prefixes, OnPTQ identifies quantization-eroded boundaries, evaluates competing tokens through short counterfactual rollouts, and combines current discrepancy with branch consequence into a Decision--Consequence risk. The
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית