כתבה
arXiv cs.CL ·
The Age of Curiosity Meets the Age of AI: Benchmarking Child Safety in Large Language Models
תקציר מקורי באנגליתarXiv:2605.25510v3 Announce Type: replace Abstract: Children increasingly have access to Large Language Models (LLMs), which may expose them to responses that are developmentally inappropriate or require age-sensitive safety, guidance, and boundaries. Existing LLM safety evaluations largely focus on general harmful-content avoidance and do not explicitly target child-facing safety. We introduce KIDBench, a benchmark for evaluating child-facing LLM safety for ages 7-11 using a LLM-as-a-Judge rubric grounded in developmental-psychology. KIDBench contains realistic child queries across ten categories, with single-turn prompts and multi-turn child-actor simulations. We compare no-cues prompts with no child context, implicit-cues prompts that suggest a child speaker, and explicit age instructio
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית