כתבה
arXiv cs.AI ·
Ranked by the Matcher: A Reproducibility Audit of Knowledge Graph Extraction from Threat Reports
תקציר מקורי באנגליתarXiv:2609.01671v2 Announce Type: replace-cross Abstract: Security teams and researchers choose knowledge-graph extraction tooling for threat reports on the strength of published triple-F1 scores, yet those scores depend on how predicted triples are matched to gold annotations. We could reimplement the stated matching rule for only five of twelve inspected systems. Re-scoring ten system outputs on shared documents under eight protocols reverses eleven of forty-five pairwise orderings; one fixed prediction set spans 0.16--0.70 F1. On an external, human-adjudicated set, no mechanical matcher---lexical, embedding, or entailment---agrees with the reviewers more than seven times in ten; an LLM judge agrees far more often. To separate component effects from matcher rewards, we build CTIForge, wh
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית