Microsoft's open-source frontier voice AI series providing streaming text-to-speech synthesis and recognition with a CPU-capable inference engine and demos for developers.
Generative speech model designed for daily dialogue, supporting multiple speakers and fine-grained prosody control for conversational synthesis research and development.
Releases the Kronos family of foundation models for financial market candlestick sequences, including multi-scale pretrained weights and hierarchical discrete tokenizers for quantitative researchers and forecast developers.
Google Research pretrained time-series foundation model for multivariate forecasting, releasing weights and evaluation benchmarks for researchers and developers building prediction systems.
Baidu-released open OCR model weights and inference examples for one-shot, long-horizon document parsing, designed for developers handling lengthy and structurally complex documents.
Microsoft's collection of large-scale self-supervised pretraining research spanning language, vision, speech, and multimodal foundation models, with code and weights for research reproduction.
Collection of open-weight Tongyi Qianwen Qwen large language models, including base and chat variants with technical reports, intended for deployment, fine-tuning, and research by developers.
Open-source financial large language models with instruction-tuned weights, evaluation benchmarks, sentiment analysis, and robo-advisory forecasting capabilities for finance researchers and practitioners.
Open 30B-parameter deep-research model from Tongyi focused on long-horizon information seeking with ReAct reasoning, achieving top results on intelligent search benchmarks.
Provides open Chinese LLaMA and Alpaca model weights alongside pretraining and instruction-tuning scripts plus local quantized deployment solutions supporting Chinese LLM research and experimentation.
A family of autoregressive unified multimodal models that decouple visual encoding to support both multimodal understanding and text-to-image generation within one architecture.
Feed-forward 3D foundation model reconstructing scenes from streaming data, released with weights, paper, and long-sequence inference implementation for real-time spatial understanding applications.
Ultra-lightweight open-source text-to-speech model library under 25MB that runs on CPU, offering multiple voices and speed control for developers and applications.
Reconstructs the Mythos architecture as an open theoretical model from published research, implementing recurrent-depth Transformers and sparse mixture-of-experts structures for study.
Official implementation of RWKV large language models with RNN architecture, providing training and inference code plus model weights that combine Transformer-level performance with linear-time inference efficiency.
Open text-and-image-to-video model series released by Zhipu. Supplies CogVideoX weights plus inference and fine-tuning code for researchers deploying video generation systems.
Offers a powerful open-source local music-generation model supporting multilingual lyrics, style control and fast composition across Mac, AMD, Intel and CUDA devices for independent musicians and producers.
Lightweight 27M-parameter general reasoning model using hierarchical inference. Achieves near-perfect results on Sudoku, maze, and ARC benchmarks with only a thousand training examples.
Open-source bilingual Chinese-English conversational LLM from Fudan University, providing model weights, dialogue data, and fine-tuning code with plugin-call support for researchers and developers.
14MB foundation model for tiny devices including phones, wearables, smart homes, and robots, optimized for on-device tool calling and structured extraction tasks.
Provides a large-scale 3D generation model using native and compact structured latents, enabling researchers and artists to create high-fidelity 3D assets with material modeling from images.
Lightweight LLM series optimized for edge deployment, offering 1B/2B weights, training recipes, and fine-tuning guides for developers building on-device applications.
Official Python package for the LTX-2 audio-video generative model, providing inference code and LoRA training scripts for researchers and developers working with synchronized media generation.
Open-weight frontier multimodal foundation model from Moonshot AI, natively supporting long-context and agentic tasks for developers building advanced AI applications.
Meituan LongCat foundation video generation model unifying multiple tasks to produce minute-long videos efficiently, positioned against open-source and commercial high-end generators.
Releases a 20-billion-parameter open-source image foundation model excelling at complex text rendering and precise image editing, supporting both text-to-image generation and instruction-based editing tasks.
Releases Google DeepMind's global medium-range weather and cyclone forecasting models with runnable code and pretrained weights for research and operational meteorological use.
Offers Z.ai's open-source mixture-of-experts language model series designed for agentic workflows, providing long-context understanding and coding capabilities for developers building autonomous applications.
Second-phase Chinese extension of LLaMA-2, releasing Chinese pretraining and instruction-tuned weights plus 64K long-context variants for deployment, fine-tuning and research.
Open-source language models specialized for binary decompilation, offering multiple model sizes plus training scripts and benchmarks to convert x86_64 binaries back into readable C code for reverse engineers.
A family of unified native multimodal model weights supporting text-to-image generation, image editing, and high-resolution up to 4K output for developers and researchers.
An open-source on-device text-to-speech model family supporting instant voice cloning and real-time multilingual synthesis for developers building low-latency voice applications locally with efficient deployment.
Lightweight block-diffusion draft model for speculative decoding acceleration, providing weights across multiple model families along with inference and evaluation support for faster generation.
Generalist foundation model weights for structured data, unifying classification, regression, and missing-value imputation tasks for practitioners working with tabular datasets in production.
Open-source family of speech generation models for long-form speech, conversational dialogue, real-time streaming, and sound effects, with multiple weight sizes and companion inference code for researchers and developers.
Pretrained tabular foundation model from Google Research for regression and classification on mixed-type tabular data, supporting zero-shot prediction for data scientists and practitioners.
Open-source 0.9-billion-parameter model for long-form transcription in more than fifty languages, featuring speaker diarization, timestamps, and acoustic event awareness for audio analysis pipelines.
Open long-horizon audio-visual generation and omnimodal world models for persistent stories and interactive worlds, with inference code and checkpoints for research.
Tencent's lightweight end-to-end vision-language OCR model focused on faster, more accurate recognition of documents, tables, and scene text for developers.
Real-time open-ended video editing model using autoregressive diffusion for causal frame-by-frame edits. Supporting live streams at 720p and 30 FPS, it serves creators needing high-throughput interactive editing.
Releases foundation-model weights for interactive world simulation, supporting high-speed video generation and agent behavior planning for researchers and developers exploring embodied interaction.
Lightweight multimodal model family covering image and video understanding alongside text-to-image generation, with released weights supporting research, fine-tuning, and deployment by developers.
Community-maintained open-source text-to-speech model for expressive, long-form multi-speaker conversational speech, offering multiple weight sizes plus fine-tuning and inference code for researchers.
Releases a 20-billion-parameter end-to-end continuous speech synthesis foundation model with pretrained and distilled weights plus inference and fine-tuning code for researchers.
Fully open multimodal training framework unifying image, long-video, and spatial understanding, releasing models, datasets, and complete training pipelines for researchers building vision-language systems.
An open roadmap project for building world models to evaluate robot policies, sharing training and inference pipelines, model weights, datasets, and evaluation leaderboards.
Trains a world foundation model based on next-state prediction, learning unified latent world representations from video and language for downstream reasoning.
Multimodal model for embodied and edge vision delivering fine-grained open-world visual understanding, supporting object localization with abstention on missing targets for robotics developers.
Releases native language world-model weights spanning seven agent-interaction domains, together with cross-domain simulation benchmarks and a technical report for general-agent research.
Open-weight generalist vision-language model for medical text, 2D/3D imaging, and surgical video, emphasizing transparent evaluation and strong performance across multiple medical understanding benchmarks.
Apache-2.0 open-source unified image generation and editing model family offering base, high-speed, and editing variants, supporting high-quality Chinese and English text rendering with relatively little training data.
Family of self-supervised Vision Transformer backbones for dense spatial perception, offering pretrained weights from compact variants up to 1.1B parameters for robotics and vision researchers.
Open-weight mixture-of-experts video model series for embodied intelligence pretraining, providing scalable architectures and checkpoints for researchers building world models and robotic perception systems.
Vision-language-action foundation model weights for robotic manipulation, supporting generalized task execution across multiple robotic arm configurations for research and deployment.
Long-form music generation model using a hybrid hierarchical architecture, generating complete five-minute songs from lyrics and descriptions for musicians and creators, with released weights.
Full-stack open-source interactive world model for long-horizon video generation, supporting real-time camera control and prompt switching, with released weights and training and inference code.
Efficient mixture-of-experts foundation model weights from Tencent Hunyuan for real-world embodied agents, balancing physical understanding, reasoning performance, and inference efficiency.
Official inference repository for the aesthetic open-source Krea 2 image model, providing RAW base and Turbo distilled weights plus high-resolution generation examples.
A general-purpose interactive world model releasing weights that generate high-fidelity, explorable, and controllable dynamic world videos for researchers and creative developers.
A unified multimodal generative vision model releasing weights, datasets, and training and inference code to support researchers studying structured scene understanding and dense geometric prediction tasks.
Tencent Hunyuan's 295B MoE reasoning and agentic model weights release, emphasizing cost efficiency while matching larger flagship models on production tasks.
The official implementation of τ0-VLA, a hierarchical robot foundation model for long-horizon manipulation, providing weights, training, and deployment code with world-model-guided test-time reasoning.
Releases open weights and inference code for LavaSR, a lightweight high-speed speech restoration and enhancement model for developers cleaning noisy voice recordings.
Large diffusion-based multiplayer world model for Rocket League that generates four-view match video frame by frame in response to real-time player inputs, supporting interactive world-model research.
A series of large-scale diffusion language models from Ant Group's InclusionAI team, releasing multiple weight sizes with optimized parallel inference for researchers and developers.