כתבה
arXiv cs.LG ·
Learning Source Acquisition Policies by Offline Planning
תקציר מקורי באנגליתarXiv:2609.14299v1 Announce Type: new Abstract: Predicting under an acquisition budget requires choosing feature groups whose value can depend on later queries. O-MPAC transfers finite-horizon risk-cost targets from complete training records into a shared source-action scorer. At inference time, the scorer uses partial observations and source metadata, re-scores after each query, and applies a hard cost mask. We analyze how tied teacher targets and the remaining planning horizon affect the learned decisions. Uniform supervision over tied minima preserves the target distribution under source relabeling. In a five-seed routing experiment, it achieves 0.965 accuracy under both original and context-last orders. On six real tasks, validation selects H1 without action cross-entropy in all thirty
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית