Stable Audio 3
TL;DR AI
2 min readKey summary
Stable Audio 3 introduces a family of latent diffusion models for generating and editing audio at variable lengths.
It uses a semantic-acoustic autoencoder and adversarial post-training to improve speed, fidelity, and prompt adherence.
The paper also highlights inpainting and other editing capabilities for more flexible audio workflows.
Stability AI released the small and medium model weights and pipelines, making the system more accessible on consumer hardware.
