יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

Search-G1: סוכני חיפוש מוצברים על ידי שכר פנימי-בסיסי

Search-G1: Grounded Search Agents via Representation-Based Intrinsic Rewards
סוכני חיפוש שמוצברים על ידי שכר פנימי-בסיסי, כדי להגדיר תפקודיות חיפוש-חיפוש.
תקציר מקורי באנגליתarXiv:2608.07531v3 Announce Type: replace Abstract: Search-augmented language agents should retrieve external information only when necessary and ground their answers in retrieved evidence. Existing external rewards provide either sparse outcome supervision or richer feedback from process annotations and LLM judges. Outcome rewards scale readily but cannot distinguish grounded retrieval from redundant search, whereas richer signals require costly annotation or inference during training. Internal rewards based on policy-side signals such as entropy, likelihood, or information gain are graded and inexpensive to evaluate, yet mainly reflect model confidence rather than evidence grounding. We propose Search-G1, a representation-based intrinsic reward framework that measures the operational gro
קרא במקור המקורי