All-in-one pure C++ inference engine powered by ggml for audio models, covering TTS, STT, VAD, voice conversion, and music generation with optimized performance and no Python requirement.