A family of universal multimodal embedding models unifying text, image, video and interleaved inputs for state-of-the-art retrieval and understanding.