I don’t understand their post on X. So they’re starting with DeepSeek-R1 as a starting point? Isn’t that circular? How did DeepSeek themselves produce DeepSeek-R1 then? I am not sure what the right terminology is but there’s a cost to producing that initial “base model” right? And without that, isn’t a lot of the expensive and difficult work being omitted?
No, the steps 1 vs 2+3 refer to different things, they do not depend on each other. They start with the distillation process (which is probably easier because it just requires synthetic data). Then they will try to recreate the R1 itself (first r1zero in step 2, and then the r1 in step 3), which is harder because it requires more training data and training in general. But in principle they do not need step 1 to go to step 2.
https://x.com/_lewtun/status/1883142636820676965
https://github.com/huggingface/open-r1
Hugging Face Journal Club - DeepSeek R1 https://www.youtube.com/watch?v=1xDVbu-WaFo