Multimodal model for embodied and edge vision delivering fine-grained open-world visual understanding, supporting object localization with abstention on missing targets for robotics developers.