כתבה
arXiv cs.AI ·
איתור נאמנות יעיל וקשיח לקטעי קוד שנוצרו על ידי LLM
Efficient and Scalable Provenance Tracking for LLM-Generated Code Snippets
איתור נאמנות יעיל וקשיח לקטעי קוד שנוצרו על ידי LLM. המחקר מציג פיתוח חדש של טכנולוגיית איתור נאמנות, המשלב חיפוש וקטעי קוד וזיהוי נאמנות. המערכת נבחנה על ידי 10 מיליון קטעי קוד והציגה תוצאות טובות. המחקר יכול להיות רלוונטי לפיתוח כלים לאיתור נאמנות ולשיפור הביצועים של כלים אלו.
תקציר מקורי באנגליתarXiv:2605.28510v3 Announce Type: replace-cross Abstract: Large language models (LLMs) for code completion and generation are increasingly used in software development, yet they may reproduce training examples verbatim and without authorship attribution, raising legal and ethical concerns around plagiarism and license compliance. Classical fingerprint-based plagiarism detectors, such as Winnowing, remain highly effective, yet the inspection requires comparing fragments of code to the entire training set, and their linear-time search makes them impractical for the billion-scale corpora used to train modern code LLMs. To bridge this gap, we introduce SourceTracker, a 300M-parameter encoder tailored for code retrieval, together with a hybrid two-stage provenance-tracking pipeline HybridSource
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית