כתבה
arXiv cs.LG ·
Compressed Video Aggregator: Content-driven Module for Efficient Micro-Video Recommendation
תקציר מקורי באנגליתarXiv:2605.08810v2 Announce Type: replace Abstract: We propose \textbf{Compressed Video Aggregator} (CVA), a lightweight micro-video recommendation module that decouples video information from preference learning. CVA first summarizes frozen VFM frame embeddings into a semantic-consensus anchor through masked mean pooling, projects this anchor into a compact latent space, and refines the projected representation with residual self-attention and feedforward blocks before producing a single video embedding for the recommender. Due to the redundancy in the frame count of the original benchmark and its overly coarse sampling, we used titles to re-select key frames based on CLIP. Experiments on MicroLens and Short-Video show consistent gains with orders-of-magnitude reductions in training time
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית