יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

מסגרת היגוי מבוססת QLoRA

A Verifier-Guided Explainable Reasoning Framework with Gold-Anchored QLoRA, Task-Aware Mixture-of-Experts, and Group-Relative RLVR
חוקרים הציגו מסגרת היגוי מבוססת QLoRA, המשלבת רווח-מומחים ו-RLVR. המסגרת משפרת את היכולת להסביר תוצאות בשאלות חינוכיות. המודל Qwen2.5-3B-Instruct שימש לבדיקת המסגרת.
תקציר מקורי באנגליתarXiv:2609.05221v1 Announce Type: cross Abstract: Large language models (LLMs) show strong reasoning ability, but their explanations can remain inconsistent, weakly grounded, or difficult to verify. We propose a verifier-guided explainable reasoning framework for transparent educational question answering that combines gold-anchored QLoRA, task-aware symbolic routing, and group-relative RLVR. Qwen2.5-3B-Instruct is first adapted with field-weighted QLoRA supervision anchored to authoritative answers. A lightweight router then assigns logic problems to a FOL/Z3 verifier and physics problems to a formula- and unit aware symbolic solver. Verifier feedback is further used to support candidate evaluation, self-revision, and reward construction during RLVR. Candidate responses are evaluated alon
קרא במקור המקורי