World Wires · story 28620 · corroborated · 1 source(s)

Q-Learning with Scalar Adjoint Matching

Researchers propose adjoint matching to fine-tune flow policies with off-policy reinforcement learning against learned value functions.

Open in the desk

Coverage

What this site indexes