Evaluating Self-Repair Mechanisms in Language Model Ablations
Researchers argue that language model self-repair following component ablation is driven by pre-existing gains rather than dynamic compensation.
Researchers argue that language model self-repair following component ablation is driven by pre-existing gains rather than dynamic compensation.