Lightweight multimodal model family covering image and video understanding alongside text-to-image generation, with released weights supporting research, fine-tuning, and deployment by developers.