יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

התאמת סוגיות LLM לשירותי טכני עם הרחבת לוגיקה סמויה, צמצום רעש ודגמי פרסות משולבים

Adapting Technical-Service LLM Agents with Latent Logic Augmentation, Robust Noise Reduction, and Hybrid Reward Modeling
סוגיות LLM לשירותי טכני עם הרחבת לוגיקה סמויה, צמצום רעש ודגמי פרסות משולבים. המחקר מציע פרקטיקה להתאמת סוגיות LLM לשירותי טכני, כולל הרחבת לוגיקה סמויה, צמצום רעש ודגמי פרסות משולבים. הפרקטיקה נבחנה על סוגיות טכני-שירותי בעולם האובייקטיבי.
תקציר מקורי באנגליתarXiv:2603.18074v2 Announce Type: replace Abstract: Technical-service LLM agents are entering production workflows, where value depends on whether engineers adopt generated replies. Service tickets hide decision logic, contain noisy single-reference responses, and make reward evaluation costly, making standard post-training brittle. Existing post-training and LLM-as-a-Judge approaches improve grounding or feedback, but do not jointly model latent decision logic, response diversity, and reward cost. We address this gap by coupling latent logic augmentation, robust noise reduction, and hybrid reward modeling. The framework augments supervised fine-tuning data with Planning-Aware Trajectory Modeling and Reasoning Augmentation, builds dual-filtered Multiple Ground Truths, and trains the policy
קרא במקור המקורי