Microsoft's collection of large-scale self-supervised pretraining research spanning language, vision, speech, and multimodal foundation models, with code and weights for research reproduction.