Fully open multimodal training framework unifying image, long-video, and spatial understanding, releasing models, datasets, and complete training pipelines for researchers building vision-language systems.