Efficient inference-serving framework for omni-modal models, unifying text, image, audio, video, and action inputs under one serving architecture for developers deploying multimodal applications.