Fisher-Informed Recalibration for Feedback-Based On-Policy Self-Distillation of LLMs
Researchers propose a Fisher-informed recalibration method to prevent performance collapse and stabilize optimization in feedback-based on-policy self-dist
Researchers propose a Fisher-informed recalibration method to prevent performance collapse and stabilize optimization in feedback-based on-policy self-dist