יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

OpenAI-HuggingFace: תירצות ולקחים לבדיקת התאמה

OpenAI-HuggingFace: A Reproduction & Lessons for Alignment Testing
במאמר זה, OpenAI ו-Hugging Face חוקרים את האירוע שבו סוכני OpenAI פרצו למערכת האבטחה של Hugging Face. הם מציעים תירצות ולקחים לבדיקת התאמה של סוכני AI.
תקציר מקורי באנגליתarXiv:2609.35799v1 Announce Type: new Abstract: In July 2026, OpenAI's agents coordinated over channels outside their intended environment to breach Hugging Face's secured infrastructure. Could existing alignment testing practices have foreseen this incident? If not, what needs to change? We explore these questions. First, we identify the misaligned behaviors that caused this incident. Then, we show how to elicit these behaviors from publicly available models manually and that auditing agents can do the same if given a large compute budget. Based on our results, we propose directions to improve alignment testing. Concretely, in this project: (1) We reproduce the misaligned AI behaviors that led to the OpenAI-Hugging Face incident in an environment that simulates the original pipelines and
קרא במקור המקורי