כתבה
arXiv cs.LG ·
Structure over Pixels: Learning Variable-Length Visual Programs
תקציר מקורי באנגליתarXiv:2605.27696v3 Announce Type: replace-cross Abstract: Discrete visual tokenizers map images to ordered sequences of tokens, providing a natural representation for structural scene descriptions. Most use a fixed sequence length, while adaptive methods often require post-hoc search or choose among a small set of rates that control the length. We propose STROP, a discrete tokenizer that learns both a visual program and its image-dependent active length. A length head is trained with a four-phase curriculum using local rate-distortion probes against frozen DINOv3 features, then predicts the active prefix in a single forward pass. At a matched rate of about $250$ nominal bits per crop, the adaptive model improves segmentation over a separately trained fixed-length baseline on four benchmark
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית