Achieving an O(1/N) Optimality Gap in Average-Reward Weakly-Coupled MDPs
The study derives an O(1/N) optimality gap for average-reward weakly-coupled Markov decision processes with identical arms and budget constraints.
The study derives an O(1/N) optimality gap for average-reward weakly-coupled Markov decision processes with identical arms and budget constraints.