Closing the Horizon Gap in Policy Optimization for Adversarial MDPs
Researchers close the horizon gap in regret bounds for policy optimization algorithms in adversarial Markov decision processes.
Researchers close the horizon gap in regret bounds for policy optimization algorithms in adversarial Markov decision processes.