כתבה
arXiv cs.CL ·
Sensitive-Topic Leakage Through LLM Routing Metadata: Measurement and Mitigation
תקציר מקורי באנגליתarXiv:2610.09981v1 Announce Type: cross Abstract: LLM routers pick a cheap or expensive model per request by its content, and many gateways and some cloud platforms can log that choice with content logging off. We measure this privacy channel beyond token counts, accounting for noisy labels and repeated prompts. We run pre-registered studies on 1.7 million real requests (WildChat-1M, LMSYS-Chat-1M) with two cost/quality routers and a domain router, survey eleven systems' logging, and test post-processing defenses. At matched length, the shift's direction depends on category and router. For RouteLLM at the 50% operating point, harassment and self-harm requests reach the strong model 19 points less often than comparable ones on prompts unseen in exploration, medical requests (exploratory: LL
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית