כתבה
arXiv cs.AI ·
SkillLift: Learning Dense Rubrics from Sparse Oracles for Efficient Skill Evolution
תקציר מקורי באנגליתarXiv:2609.15396v1 Announce Type: new Abstract: LLM-based agents increasingly rely on persistent skills, i.e., reusable procedural prompts, to adapt without weight updates. Existing skill self-evolution methods directly revise skill text based on execution feedback, but each oracle evaluation requires a full agent rollout, creating a supervision bottleneck that confines search to failure-patching updates. Our key insight is that ranking is a smoother supervision target than absolute outcome regression: identifying which skill is better requires fewer oracle evaluations than predicting exact scores. Building on this insight, we propose SkillLift, which decouples skill search from oracle cost by learning an oracle-aligned rubric as a structured evaluation space. We formalize this as a bileve
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית