We’re really excited about self-distillation as a new paradigm for post-training.
Also check our work applying the same algorithm to offline data: self-distillation.github.io/SDFT
Here the baseline is SFT, not GRPO.
We show: Self-distillation enables continual learning.
self-distillation.github.io
SDFT: Self-Distillation Enables Continual Learning