יום חמישי, 8 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

MUNITE: פלטפורמת חדשה להפקת תמונות וטקסט בצורה זולתית

MUNITE: Unified Multimodal Latent Inference for Any-to-Any Multimodal Generation
MUNITE היא פלטפורמת חדשה להפקת תמונות וטקסט בצורה זולתית. היא מאפשרת הפקת תמונות וטקסט בצורה זולתית, כלומר, היא יכולה להפיק תמונות וטקסט מסוגים שונים של נתונים. MUNITE משתמשת בטכנולוגיית Latent Variable Models כדי להפיק תמונות וטקסט בצורה זולתית. היא גם מאפשרת הפקת תמונות וטקסט בצורה זולתית מסוגים שונים של נתונים.
תקציר מקורי באנגליתarXiv:2610.09866v1 Announce Type: new Abstract: We introduce MUNITE, a latent-variable framework for flexible any-to-any multimodal generation that treats encoding and latent generation as the same inference problem under different amounts of observed evidence. Given any subset of modalities, MUNITE models the conditional distribution over the latent representation associated with the complete observation. Full observation recovers deterministic encoding, no observation recovers the latent marginal, and intermediate subsets define conditional latent inference, all within a single conditional flow model. A shared latent sample captures variation that must remain consistent across generated targets, while modality-specific generative decoders model the remaining uncertainty independently. To
קרא במקור המקורי