A multimodal model which can process text, vision, and audio as both inputs and outputs.
Specs
Params73B total, 3B active
LicenseMIT
Resources
Adoption · Hugging Face
RAM score
Relative Adoption Metric not scored because required parameter, download, or API metadata is not cataloged. This is not a zero score.
Hugging Face Downloads
200
last 30d
31K
all time
HF Likes
211
Relative Adoption Metric not scored because required parameter, download, or API metadata is not cataloged. This is not a zero score.
Related Models
More from Meituan LongCat
Similar Models
Ming-Lite-Omni-1.5
Inclusion AI15 Jul 2025multimodalimage generationaudio generation
Ming-Lite-Omni
Inclusion AI28 May 2025multimodalimage generationaudio generation
HunyuanImage-3.0-Instruct
Tencent26 Jan 2026multimodalimage generation80B params
Emu3.5-Image
BAAI31 Oct 2025multimodalimage generation34.1B params
