Open-weight generalist vision-language model for medical text, 2D/3D imaging, and surgical video, emphasizing transparent evaluation and strong performance across multiple medical understanding benchmarks.