יום שישי, 9 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

רמזים עוינים במודלים של קבלת החלטות

Adversarial Cues in Decision Models Used as Judges: The Role of Request Presentation
נמצא כי מודלים של קבלת החלטות רגישים לרמזים עוינים. המחקר בדק את השפעת הרמזים על מודל GPT. התוצאות הראו כי המודלים יכולים להיות מושפעים מרמזים קטנים.
תקציר מקורי באנגליתarXiv:2610.11436v1 Announce Type: new Abstract: An answer judge instructed to grade the final commitment should reject an explicitly wrong final value even when an earlier value matches the reference. We show that adding one colon to a candidate can violate this requirement depending on the presentation of the structured judging request. Numeric references certify the error, and paired interventions distinguish the candidate edit from the integration's presentation choices. On 200 previously unused DROP and GSM8K source clusters, the edit increased Jev's false acceptance from 1.0% to 26.0% with three output labels and from 3.0% to 26.5% with the published four-label grading instruction under sorted JSON keys. Both candidate variants were rejected under insertion presentation. These interac
קרא במקור המקורי