כתבה
arXiv cs.LG ·
Replay-buffer engineering for noise-aware quantum circuit optimization
תקציר מקורי באנגליתarXiv:2604.21863v2 Announce Type: replace-cross Abstract: Deep reinforcement learning for quantum circuit optimization faces three bottlenecks: replay buffers that overlook temporal difference (TD) target reliability, curriculum-based architecture search requiring a full quantum-classical evaluation after every edit, and the discard of noiseless trajectories when retraining under hardware noise. We address these limitations by treating replay as a central algorithmic lever. We introduce ReaPER+, an annealed replay rule that transitions from TD-error prioritization to reliability-aware sampling as value estimates mature. ReaPER+ achieves up to 4x higher sample efficiency than fixed PER, ReaPER, and uniform replay, while matching prior on-policy solution quality with up to 32x fewer interact
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית