Open-source family of speech generation models for long-form speech, conversational dialogue, real-time streaming, and sound effects, with multiple weight sizes and companion inference code for researchers and developers.