כתבה
arXiv cs.LG ·
FlowBalance: שיפור עצמי מונחה על ידי בודק
FlowBalance: Verifier-Grounded Self-Improvement from On-Policy Reasoning Experience
FlowBalance הוא שיטה לשיפור עצמי של מודלים להיגיון. היא משתמשת בבודק לצורך הדרכה ושיפור הביצועים. FlowBalance שופרת את הביצועים של Qwen3-4B ו-Qwen3-8B.
תקציר מקורי באנגליתarXiv:2609.03241v1 Announce Type: new Abstract: A reasoning model can improve from its own on-policy experience, but this inner loop is fragile: terminal verifiers provide reliable yet sparse supervision, while dense same-model guidance can reinforce false confidence or overconcentrate learning on a narrow solution mode. We introduce FlowBalance, a verifier-grounded self-improvement method that learns a normalized distribution over complete responses. For each on-policy trajectory, a frozen training-time view of the same policy uses privileged context to produce token-level log-probability gains, which are aggregated into a trajectory-level self-guidance score. FlowBalance calibrates this score with the verifier-derived group advantage: guidance is retained on positive-advantage trajectori
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית