Generates expressive speech from short clips to full-length audiobooks, cloning any voice from a 10-second reference while controlling emotion, pacing, and breathing for audiobook creators and voice designers.