יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

SIRF: מודל יסודי לבקרת סיכונים

SIRF: A Spec-Internalized Risk Foundation Model for Industrial Content Risk Control
SIRF הוא מודל יסודי לבקרת סיכונים בתוכן תעשייתי. המודל משלב מדיניות מורכבת במשקלים, ללא צורך בסימון נוסף. SIRF משיג תוצאות טובות יותר מאשר מודלים אחרים, כגון Qwen3-8B-SFT.
תקציר מקורי באנגליתarXiv:2609.11752v1 Announce Type: new Abstract: For industrial content risk control, the real deployment constraint is not average accuracy but how much risk can be auto-handled under high precision and second-level latency. We present SIRF (Spec-Internalized Risk Foundation Model), which internalizes a platform's complex policies, synthesized without additional human annotation via EntiGraph, MAGA rewriting and account-level chain-of-thought (CoT), into the weights via continued pretraining (CPT), so rules are applied at high precision under an ultra-low-latency, verdict-only deployment. A controlled same-source comparison (Qwen3-8B-SFT vs. SIRF-8B-SFT, identical policy injection and verdict-only output form, differing only in policy-grounded CPT) attributes the gain to internalization: S
קרא במקור המקורי