כתבה
arXiv cs.LG ·
SIRF: דגם-פנימי לבקרת סיכון תוכן תעשייתי
SIRF: A Spec-Internalized Risk Foundation Model for Industrial Content Risk Control
דגם-פנימי לבקרת סיכון תוכן תעשייתי, המשלב כלים כמו EntiGraph ו-MAGA, ומציע פתרון לבקרת סיכון תוכן תעשייתי בעל דיוק גבוה וזמן תגובה נמוך.
תקציר מקורי באנגליתarXiv:2609.11752v1 Announce Type: cross Abstract: For industrial content risk control, the real deployment constraint is not average accuracy but how much risk can be auto-handled under high precision and second-level latency. We present SIRF (Spec-Internalized Risk Foundation Model), which internalizes a platform's complex policies, synthesized without additional human annotation via EntiGraph, MAGA rewriting and account-level chain-of-thought (CoT), into the weights via continued pretraining (CPT), so rules are applied at high precision under an ultra-low-latency, verdict-only deployment. A controlled same-source comparison (Qwen3-8B-SFT vs. SIRF-8B-SFT, identical policy injection and verdict-only output form, differing only in policy-grounded CPT) attributes the gain to internalization:
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית