Decoupling Exploration from Optimization in RLVR
This paper examines reinforcement learning with verifiable rewards and the challenges of incorporating novelty incentives to discover new reasoning strateg
This paper examines reinforcement learning with verifiable rewards and the challenges of incorporating novelty incentives to discover new reasoning strateg