Distillation Defenses Easily Break After Reinforcement Learning
A study shows that existing defenses against model distillation attacks fail after reinforcement learning is applied to the distilled models.
A study shows that existing defenses against model distillation attacks fail after reinforcement learning is applied to the distilled models.