We updated our OLMo 7B model with some small improvements to training data and annealing techniques. A pretty minor bump, but we have some exciting new models coming soon. An interesting thing I've learned about pretraining is how loss spikes often relate to "skipped tokens," making the models worse at a fixed compute budget. This is because the gradients is loss spikes get clipped and the data effectively does nothing.
Similarity · VAIL
VAIL Fingerprint
03e7:05e2:071f:0961:0c51:10ba:1fc9:6422
Explore other models with behavioral similarity to OLMo-1B-0724-hf.
Adoption · Hugging Face
RAM score
Relative Adoption Metric not applicable.
Hugging Face Downloads
5.6K
last 30d
950.3K
all time
HF Likes
24
Relative Adoption Metric not applicable.
Related Models
