Sharpen Without Search: On-Policy Distillation of Sequence-Level Power Distribution
This paper explores on-policy distillation using a power distribution to shift probability toward the most likely answers in language models.
This paper explores on-policy distillation using a power distribution to shift probability toward the most likely answers in language models.