יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

TRACE: פיתוח סוכנים לחקירה סיבתית עם תגמולים מלאכותיים

TRACE: Training Reasoning Agents for Causal Exploration with Synthesized Rewards
במאמר זה, פותחים סוכנים לחקירה סיבתית בעזרת תגמולים מלאכותיים. הסוכנים חוקרים את הסיבות לבעיות בתחום הדיאגנוסטיקה. המאמר מציג תוצאות של חקירה של סוכנים בתחום זה, כולל שימוש במודלי Qwen ו-Claude.
תקציר מקורי באנגליתarXiv:2609.10315v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) has advanced language-model reasoning in domains such as mathematics and code, where objective answers are inexpensive to check. Diagnostic reasoning over complex data lacks this advantage: establishing the true cause of an anomaly often requires costly expert investigation and may remain ambiguous after the fact. We ask whether this asymmetry of verification can instead be engineered. We sample an intervention, inject it into a controlled simulator, and generate the observations it would produce. The hidden intervention provides an oracle label and objective reward, while the agent must still investigate noisy, confounded, and distributed evidence. We instantiate this approach in TRACE, a
קרא במקור המקורי