A multimodal version of Qwen3 with a rather beefy 1.7B vision encoder and 1.2T tokens of continued, multimodal training, followed by an RL stage.
Specs
Params10B
LicenseApache-2.0
Capability · Artificial Analysis
AA Index
3.8
Adoption · Hugging Face
RAM score
Relative Adoption Metric not scored because required parameter, download, or API metadata is not cataloged. This is not a zero score.
Hugging Face Downloads
20.2K
last 30d
1.8M
all time
HF Likes
414
Relative Adoption Metric not scored because required parameter, download, or API metadata is not cataloged. This is not a zero score.
Related Models



