כתבה
arXiv cs.AI ·
Are AI Coders Snitches? An Empirical Study of Pretraining Data Detection on Code Large Language Models
תקציר מקורי באנגליתarXiv:2507.17389v2 Announce Type: replace-cross Abstract: Recent advances in code large language models (CodeLLMs) have made them indispensable tools in modern software engineering. However, these models occasionally produce outputs that contain proprietary or sensitive code snippets, raising concerns about potential non-compliant use of training data, and posing risks to privacy and intellectual property. To ensure responsible and compliant deployment of CodeLLMs, training data detection (TDD) has become a critical task. While recent TDD methods have shown promise in natural language settings, their effectiveness on code data remains largely underexplored. This gap is particularly important given code's structured syntax and distinct similarity criteria compared to natural language. To ad
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית