יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

חלוקה, ייעוץ, ניצחון: כביסת יכולות דרך מודלים מתואמים

Divide, Consult, Conquer: Capability Laundering Through Aligned LLMs
חוקרים מצאו פרצת אבטחה במודלים מתואמים. מודל חלש יכול לחלק משימה מסוכנת לתת-משימות תמימות, לייעץ עם מודל חזק יותר ולשלב את התשובות. הם בדקו את GPT-5.5, Claude Opus 4.8 ו-Grok-4.3.
תקציר מקורי באנגליתarXiv:2609.15383v1 Announce Type: cross Abstract: Language model safety is typically evaluated one interaction at a time. We show that a weaker, unaligned model can split a harmful task into benign-looking subproblems, consult a stronger aligned model independently on each, and combine the answers locally. We call this attack capability laundering. Unlike a jailbreak, no single response is a harmful task. We measure consultation-aided uplift using tasks that a raw frontier model solves, the aligned frontier refuses, and the unassisted orchestrator fails. We evaluate GPT-5.5, Claude Opus 4.8, and Grok-4.3 as consultants to four local orchestrators on CyBench, BountyBench, and harmful CBRN requests. On CyBench, Gemma-4-31B recovers 8/14 candidates with GPT-5.5 and 7/9 with Opus, compared wit
קרא במקור המקורי