Predicting Alignment Generalization with Value Representations
This paper examines how narrow post-training behaviors in LLMs influence generalization across unseen contexts and environments.
This paper examines how narrow post-training behaviors in LLMs influence generalization across unseen contexts and environments.