יום רביעי, 7 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

מבנה, לא אמונה: סיימפלינג תומפסון המורכבת מ-LLM-Derived Covariance ב-Semi-Bandits

Structure, Not Belief: Correlated Thompson Sampling from LLM-Derived Covariance in Combinatorial Semi-Bandits
במאמר זה, החוקרים פיתחו טכניקה חדשה לסיימפלינג תומפסון המורכבת, המשתמשת ב-LLM-Derived Covariance ב-Semi-Bandits. הם ניסו את הטכניקה במספר תרחישים, כולל תרחישים סימולטוריים ותרחישים אמיתיים.
תקציר מקורי באנגליתarXiv:2610.07470v1 Announce Type: cross Abstract: Combinatorial Thompson sampling (CTS) draws independent posterior samples for every arm, so its exploration dynamics ignore any relation among arms. We study a minimal change to those dynamics: an LLM is queried once for a partition of the arms, the partition becomes a positive-definite correlation matrix $\Sigma$ through an RBF kernel on cluster ranks, and the per-round posterior sample is drawn with covariance $\Sigma$ while the Beta posteriors are updated from real rewards only, so the LLM shapes how the sampler moves, not what it believes. We give a self-contained Bayesian regret bound for the idealized Gaussian sampler whose information gain splits into a $K\log T$ term from the $K$-cluster structure and a ridge term that grows to $d\l
קרא במקור המקורי