A modular, low-latency pipeline for building local voice agents from open-source speech recognition, language, and synthesis models. Supports developers creating conversational and robotic dialogue applications fully on-device.