中文

Models

73 projects · page 1 / 1

1-73 of 73
01
vibevoice
Microsoft's open-source frontier voice AI series providing streaming text-to-speech synthesis and recognition with a CPU-capable inference engine and demos for developers.
02
ChatTTS
Generative speech model designed for daily dialogue, supporting multiple speakers and fine-grained prosody control for conversational synthesis research and development.
03
kronos
Releases the Kronos family of foundation models for financial market candlestick sequences, including multi-scale pretrained weights and hierarchical discrete tokenizers for quantitative researchers and forecast developers.
04
voxcpm
Tokenizer-free multilingual text-to-speech model supporting thirty languages, creative voice design, high-fidelity cloning, and 48kHz output for natural speech generation tasks.
05
timesfm
Google Research pretrained time-series foundation model for multivariate forecasting, releasing weights and evaluation benchmarks for researchers and developers building prediction systems.
06
unlimited-ocr
Baidu-released open OCR model weights and inference examples for one-shot, long-horizon document parsing, designed for developers handling lengthy and structurally complex documents.
07
unilm
Microsoft's collection of large-scale self-supervised pretraining research spanning language, vision, speech, and multimodal foundation models, with code and weights for research reproduction.
08
Qwen
Collection of open-weight Tongyi Qianwen Qwen large language models, including base and chat variants with technical reports, intended for deployment, fine-tuning, and research by developers.
09
FinGPT
Open-source financial large language models with instruction-tuned weights, evaluation benchmarks, sentiment analysis, and robo-advisory forecasting capabilities for finance researchers and practitioners.
10
DeepResearch
Open 30B-parameter deep-research model from Tongyi focused on long-horizon information seeking with ReAct reasoning, achieving top results on intelligent search benchmarks.
11
Chinese-LLaMA-Alpaca
Provides open Chinese LLaMA and Alpaca model weights alongside pretraining and instruction-tuning scripts plus local quantized deployment solutions supporting Chinese LLM research and experimentation.
12
Janus
A family of autoregressive unified multimodal models that decouple visual encoding to support both multimodal understanding and text-to-image generation within one architecture.
13
lingbot-map
Feed-forward 3D foundation model reconstructing scenes from streaming data, released with weights, paper, and long-sequence inference implementation for real-time spatial understanding applications.
14
kittentts
Ultra-lightweight open-source text-to-speech model library under 25MB that runs on CPU, offering multiple voices and speed control for developers and applications.
15
openmythos
Reconstructs the Mythos architecture as an open theoretical model from published research, implementing recurrent-depth Transformers and sparse mixture-of-experts structures for study.
16
RWKV-LM
Official implementation of RWKV large language models with RNN architecture, providing training and inference code plus model weights that combine Transformer-level performance with linear-time inference efficiency.
17
CogVideo
Open text-and-image-to-video model series released by Zhipu. Supplies CogVideoX weights plus inference and fine-tuning code for researchers deploying video generation systems.
18
ace-step-1.5
Offers a powerful open-source local music-generation model supporting multilingual lyrics, style control and fast composition across Mac, AMD, Intel and CUDA devices for independent musicians and producers.
19
HRM
Lightweight 27M-parameter general reasoning model using hierarchical inference. Achieves near-perfect results on Sudoku, maze, and ARC benchmarks with only a thousand training examples.
20
MOSS
Open-source bilingual Chinese-English conversational LLM from Fudan University, providing model weights, dialogue data, and fine-tuning code with plugin-call support for researchers and developers.
21
needle
14MB foundation model for tiny devices including phones, wearables, smart homes, and robots, optimized for on-device tool calling and structured extraction tasks.
22
trellis.2
Provides a large-scale 3D generation model using native and compact structured latents, enabling researchers and artists to create high-fidelity 3D assets with material modeling from images.
23
MiniCPM
Lightweight LLM series optimized for edge deployment, offering 1B/2B weights, training recipes, and fine-tuning guides for developers building on-device applications.
24
ltx-2
Official Python package for the LTX-2 audio-video generative model, providing inference code and LoRA training scripts for researchers and developers working with synchronized media generation.
25
kimi-k3
Open-weight frontier multimodal foundation model from Moonshot AI, natively supporting long-context and agentic tasks for developers building advanced AI applications.
26
longcat-video
Meituan LongCat foundation video generation model unifying multiple tasks to produce minute-long videos efficiently, positioned against open-source and commercial high-end generators.
27
Qwen-Image
Releases a 20-billion-parameter open-source image foundation model excelling at complex text rendering and precise image editing, supporting both text-to-image generation and instruction-based editing tasks.
28
weathernext
Releases Google DeepMind's global medium-range weather and cyclone forecasting models with runnable code and pretrained weights for research and operational meteorological use.
29
glm-5
Offers Z.ai's open-source mixture-of-experts language model series designed for agentic workflows, providing long-context understanding and coding capabilities for developers building autonomous applications.
30
Chinese-LLaMA-Alpaca-2
Second-phase Chinese extension of LLaMA-2, releasing Chinese pretraining and instruction-tuned weights plus 64K long-context variants for deployment, fine-tuning and research.
31
LLM4Decompile
Open-source language models specialized for binary decompilation, offering multiple model sizes plus training scripts and benchmarks to convert x86_64 binaries back into readable C code for reverse engineers.
32
sensenova-u1
A family of unified native multimodal model weights supporting text-to-image generation, image editing, and high-resolution up to 4K output for developers and researchers.
33
neutts
An open-source on-device text-to-speech model family supporting instant voice cloning and real-time multilingual synthesis for developers building low-latency voice applications locally with efficient deployment.
34
dflash
Lightweight block-diffusion draft model for speculative decoding acceleration, providing weights across multiple model families along with inference and evaluation support for faster generation.
35
limix
Generalist foundation model weights for structured data, unifying classification, regression, and missing-value imputation tasks for practitioners working with tabular datasets in production.
36
qwen3.8
Qwen's open large language model series offering downloadable weights, benchmarks, and guides for inference and fine-tuning.
37
moss-tts
Open-source family of speech generation models for long-form speech, conversational dialogue, real-time streaming, and sound effects, with multiple weight sizes and companion inference code for researchers and developers.
38
gpa
Compact single-model release unifying speech recognition, text-to-speech synthesis, and voice conversion, enabling multiple audio tasks with one tiny general-purpose model.
39
tabfm
Pretrained tabular foundation model from Google Research for regression and classification on mixed-type tabular data, supporting zero-shot prediction for data scientists and practitioners.
40
moss-transcribe-diarize
Open-source 0.9-billion-parameter model for long-form transcription in more than fifty languages, featuring speaker diarization, timestamps, and acoustic event awareness for audio analysis pipelines.
41
joyai-echo
Open long-horizon audio-visual generation and omnimodal world models for persistent stories and interactive worlds, with inference code and checkpoints for research.
42
fireredaudio
A 9B general-purpose audio language model for ASR, understanding, voice cloning, TTS, and speech editing.
43
hunyuanocr
Tencent's lightweight end-to-end vision-language OCR model focused on faster, more accurate recognition of documents, tables, and scene text for developers.
44
joyai-video-edit
Real-time open-ended video editing model using autoregressive diffusion for causal frame-by-frame edits. Supporting live streams at 720p and 30 FPS, it serves creators needing high-throughput interactive editing.
45
lingbot-world-v2
Releases foundation-model weights for interactive world simulation, supporting high-speed video generation and agent behavior planning for researchers and developers exploring embodied interaction.
46
fireredtts3
A multilingual voice-cloning model covering 24 languages and dialects with voice design and speech editing.
47
wemm-embedding
A family of universal multimodal embedding models unifying text, image, video and interleaved inputs for state-of-the-art retrieval and understanding.
48
mage
Lightweight multimodal model family covering image and video understanding alongside text-to-image generation, with released weights supporting research, fine-tuning, and deployment by developers.
49
vibevoice
Community-maintained open-source text-to-speech model for expressive, long-form multi-speaker conversational speech, offering multiple weight sizes plus fine-tuning and inference code for researchers.
50
dots.tts
Releases a 20-billion-parameter end-to-end continuous speech synthesis foundation model with pretrained and distilled weights plus inference and fine-tuning code for researchers.
51
llava-onevision-2
Fully open multimodal training framework unifying image, long-video, and spatial understanding, releasing models, datasets, and complete training pipelines for researchers building vision-language systems.
52
giga-world-1
An open roadmap project for building world models to evaluate robot policies, sharing training and inference pipelines, model weights, datasets, and evaluation leaderboards.
53
orca
Trains a world foundation model based on next-state prediction, learning unified latent world representations from video and language for downstream reasoning.
54
vlx-seek
Multimodal model for embodied and edge vision delivering fine-grained open-world visual understanding, supporting object localization with abstention on missing targets for robotics developers.
55
qwen-agentworld
Releases native language world-model weights spanning seven agent-interaction domains, together with cross-domain simulation benchmarks and a technical report for general-agent research.
56
hulu-med
Open-weight generalist vision-language model for medical text, 2D/3D imaging, and surgical video, emphasizing transparent evaluation and strong performance across multiple medical understanding benchmarks.
57
boogu-image
Apache-2.0 open-source unified image generation and editing model family offering base, high-speed, and editing variants, supporting high-quality Chinese and English text rendering with relatively little training data.
58
lingbot-vision
Family of self-supervised Vision Transformer backbones for dense spatial perception, offering pretrained weights from compact variants up to 1.1B parameters for robotics and vision researchers.
59
lingbot-video
Open-weight mixture-of-experts video model series for embodied intelligence pretraining, providing scalable architectures and checkpoints for researchers building world models and robotic perception systems.
60
lingbot-vla-v2
Vision-language-action foundation model weights for robotic manipulation, supporting generalized task execution across multiple robotic arm configurations for research and deployment.
61
minimax-music3
Long-form music generation model using a hybrid hierarchical architecture, generating complete five-minute songs from lyrics and descriptions for musicians and creators, with released weights.
62
alayaworld
Full-stack open-source interactive world model for long-horizon video generation, supporting real-time camera control and prompt switching, with released weights and training and inference code.
63
hy-embodied
Efficient mixture-of-experts foundation model weights from Tencent Hunyuan for real-world embodied agents, balancing physical understanding, reasoning performance, and inference efficiency.
64
krea-2
Official inference repository for the aesthetic open-source Krea 2 image model, providing RAW base and Turbo distilled weights plus high-resolution generation examples.
65
dreamx-world
A general-purpose interactive world model releasing weights that generate high-fidelity, explorable, and controllable dynamic world videos for researchers and creative developers.
66
sensenova-vision
A unified multimodal generative vision model releasing weights, datasets, and training and inference code to support researchers studying structured scene understanding and dense geometric prediction tasks.
67
hy3
Tencent Hunyuan's 295B MoE reasoning and agentic model weights release, emphasizing cost efficiency while matching larger flagship models on production tasks.
68
tau-0-vla
The official implementation of τ0-VLA, a hierarchical robot foundation model for long-horizon manipulation, providing weights, training, and deployment code with world-model-guided test-time reasoning.
69
helixworld
A real-time interactive audio-visual world model that generates navigable video with spatial sound from an image and prompt.
70
lavasr
Releases open weights and inference code for LavaSR, a lightweight high-speed speech restoration and enhancement model for developers cleaning noisy voice recordings.
71
mira
Large diffusion-based multiplayer world model for Rocket League that generates four-view match video frame by frame in response to real-time player inputs, supporting interactive world-model research.
72
llada2.x
A series of large-scale diffusion language models from Ant Group's InclusionAI team, releasing multiple weight sizes with optimized parallel inference for researchers and developers.
73
anigen
Introduces a unified S-cubed field model that generates riggable, animatable 3D assets from a single image with skeletal skinning and motion driving.