World Wires · story 31263 · corroborated · 1 source(s)

Closing the Horizon Gap in Policy Optimization for Adversarial MDPs

Researchers close the horizon gap in regret bounds for policy optimization algorithms in adversarial Markov decision processes.

Open in the desk

Coverage

What this site indexes