All models

Open-Reasoner-Zero-32B

byOpen-Reasoner-ZeroOpen-Reasoner-Zero· 18 Feb 2025
General purpose

An open replication of DeepSeek-R1-Zero-Qwen-32B, i.e., a fine-tuned model of Qwen2.5 32B with RL only, missing the SFT cold start. They share the used code and data, as well as a tech report. This work will likely be covered more in a future post, but you should think of it as the most robust RL-on-base-model report since the DeepSeek R1 release.

Specs
Params32B
LicenseMIT
Similarity · VAIL
VAIL
VAIL Fingerprint
0057:006f:0096:00d4:00fd:01b6:0501:5b2d

Explore other models with behavioral similarity to Open-Reasoner-Zero-32B.

Adoption · Hugging Face
RAM score
Relative Adoption Metric not scored because required parameter, download, or API metadata is not cataloged. This is not a zero score.
Hugging Face Downloads
101
last 30d
29K
all time
HF Likes
33

Relative Adoption Metric not scored because required parameter, download, or API metadata is not cataloged. This is not a zero score.

Related Models