כתבה
arXiv cs.LG ·
SCAPES: קודד סמנטי לאוטורג'נטיבי לטקסטורות אקולוגיות
SCAPES: Semantically Conditioned Autoregressive Prior for Environmental Sounds
מודל גנרטיבי לטקסטורות אקולוגיות שמאפשר שליטה סמנטית גבוהה. המודל SCAPES מסוגל לייצר טקסטורות אקולוגיות באיכות גבוהה, כולל טקסטורות של רעשי סביבה, כמו גשם, רוח, וכדומה. המודל SCAPES פועל על המישור הלטנטי המתמשך של קודק נוירלי לאודיו, ומסוגל לייצר טקסטורות אקולוגיות שונות, כולל טקסטורות של רעשי סביבה, כמו גשם, רוח, וכדומה.
תקציר מקורי באנגליתarXiv:2609.04634v1 Announce Type: cross Abstract: As generative audio models grow in complexity, the computational and ecological costs of synthesizing everyday sounds have become increasingly prohibitive, often requiring industrial-scale resources and massive datasets. In this paper, we present SCAPES: a Semantically Conditioned Autoregressive Prior for Environmental Sounds. SCAPES is a lightweight, resource-efficient generative model designed to synthesize high-fidelity environmental textures through high-level semantic control. By operating on the continuous latent manifold of a neural audio codec, our approach bypasses the rigid structural constraints inherent to discrete tokenization. We propose a segmentation strategy that decomposes audio into overlapping segments, enabling a Contin
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית