Open-source 0.9-billion-parameter model for long-form transcription in more than fifty languages, featuring speaker diarization, timestamps, and acoustic event awareness for audio analysis pipelines.