An open-source model for understanding speech, environmental sounds, and music through captioning, question answering, and reasoning