כתבה
arXiv cs.AI ·
K-Bench: בסיס מדידה לשכחה של LLM בהפעלות אגנטיות
K-Bench: A Benchmark for LLM Unlearning in Agentic Deployments
בסיס מדידה לשכחה של LLM בהפעלות אגנטיות. המחקר מציג כלי לבדיקת יכולת השכחה של מודלי LLM בהפעלות אגנטיות.
תקציר מקורי באנגליתarXiv:2609.12808v2 Announce Type: replace Abstract: Unlearning benchmarks such as TOFU and MUSE certify forgetting by reading the model's final answer, where a model that refuses to answer already counts as having forgotten. We show that this model-level certificate does not transfer once the model is deployed as an agent. We introduce K-Bench, a benchmark that scores LLM unlearning under agentic deployment. K-Bench inspects all six channels a ReAct agent exposes, including its chain-of-thought (CoT), tool calls and tool observations, and elicited summary. A query counts as leaked if the secret appears in any of them. Each experiment places the secret in exactly one of the agent's three sources (the weights, the prompt, or the retrieval store). The K-Score is computed separately for each s
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית