World Wires · story 31280 · corroborated · 1 source(s)

Predicting Alignment Generalization with Value Representations

This paper examines how narrow post-training behaviors in LLMs influence generalization across unseen contexts and environments.

Open in the desk

Coverage

What this site indexes