כתבה
arXiv cs.LG ·
GRPO-QM: חקירה נוספת לתצפית קוונטית
GRPO-QM: Target Preserving Exploration for Quantum Tomography
למדנות-מבוססת-פרסום עשויה לשנות את ההפצה האחרונה של תצפית קוונטית. GRPO-QM נמנע מכך על ידי למידה של סטרטגיה לחקירה רק עבור ההפצה הקבועה של תצפית קוונטית.
תקציר מקורי באנגליתarXiv:2609.14711v1 Announce Type: new Abstract: Reward-based learning can alter the very posterior distribution that scientific inference aims to estimate. GRPO-QM sidesteps this by learning only an exploration strategy for a stated quantum-tomography posterior: a group-relative policy chooses among reversible physical moves, and an exact Metropolis correction ensures the posterior remains stationary once the policy is fixed. We then examine what learning contributes beyond physical proposal mechanisms and prior knowledge. Reconstruction comparisons suggest that most of the gains over the tested flows come from those two components rather than from learning itself, and a closed-form counterexample explains why: a reward tied to accepted motion can increase even when a physical observable r
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית