A unified toolkit for multimodal model inference, evaluation, and post-training, supporting over a dozen models and a hundred benchmarks. Serves researchers extending evaluations and training tasks to the cloud.