All models

Virtuoso-Medium-v2

byArcee AIArcee AI· 27 Jan 2025
General purpose

A "true" distillation of DeepSeek V3 onto Qwen2.5-32B. Most models use supervised fine-tuning on outputs generated by the teacher, but this model is trained directly on V3's logits. For this to work, the model had to be re-trained on the V3 tokenizer, then distilled to finally use the original tokenizer and post-train the model.

Specs
Params32B
LicenseApache-2.0
Adoption · Hugging Face
RAM score
Relative Adoption Metric not scored because required parameter, download, or API metadata is not cataloged. This is not a zero score.
Hugging Face Downloads
95
last 30d
9.3K
all time
HF Likes
58

Relative Adoption Metric not scored because required parameter, download, or API metadata is not cataloged. This is not a zero score.

Related Models