כתבה
arXiv cs.LG ·
מוצרים שבורים מלידה: ארטיפקטים פגומים ברישומים ציבוריים
Broken on Arrival: Silently Defective LLM Artifacts in Public Model Registries and How to Catch Them
חוקרים גילו ארטיפקטים פגומים ברישומים ציבוריים של מודלים גדולים. הארטיפקטים, שהם Qwen2.5-Coder-3B ו-phi3.5-mini, לא עוברים בדיקות ואינם פועלים כראוי. החוקרים פרסמו כלי בדיקה ונתונים לאימות התוצאות.
תקציר מקורי באנגליתarXiv:2609.05881v1 Announce Type: cross Abstract: Developers increasingly run large language models locally by pulling quantized GGUF artifacts from public registries, yet nothing in the distribution pipeline functionally tests these conversions before they reach users. We executed 327 quantized code-capable model artifacts: 305 from the official Ollama library, spanning 15 model lines at every eligible quantization level at or under 8 GB, and 22 from the most-downloaded community repositories on HuggingFace. Each ran a 15-task smoke suite calibrated so that healthy artifacts pass while a known-broken one fails; suspects then faced full 164-task evaluation, a second inference backend, an independent distributor's conversion of the same model and quantization as referee, and, for community
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית