Gradient-Free versus Gradient-Based Risk Constraints in Reinforcement Learning: A Reservoir Governance Case Study

Read the full article

Discuss this preprint

Start a discussion What are Sciety discussions?

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

We ask whether the choice of reinforcement-learning optimizer matters for solving a CVaR-constrained reservoir-governance problem, and why. Training CEM (population-based), SAC, and PPO (gradient-based, with a Lagrangian expected-cost safety critic) on an identical simulator and evaluating all three on the full training objective, we find a sharp, reproducible pattern across three of four physical regimes: CEM achieves strictly lower realized tail risk than every SAC/PPO variant even its least conservative policy is safer than their most conservative ones – while SAC and PPO achieve strictly lower cost. A fourth regime, where management capacity saturates, is an instructive exception where CEM is both safest and cheapest. We explain this mechanistically: CEM evaluates the true, non-differentiable CVaR of the episode-maximum deficit directly, while gradient-based methods must substitute a differentiable expected-discounted-cost surrogate. We prove this substitution yields a discount-induced blindness result a per-step safety guarantee that decays geometrically with the time index and confirm it empirically via a discount-factor ablation and an extreme-budget stress test. An independent five-seed replication, using a scale-invariant dual-ascent correction, sharpens the risk ordering to statistical significance; a genuinely trained quantile/distributional safety critic, however, does not close the gap and in fact increases realized risk in every seed tested, pointing to Monte-Carlo quantile-estimation noise rather than La-grangian under-tuning as the remaining obstacle. Closed-loop existence, local-stability, and probabilistic practical-stability theorems complete the analysis, applicable regardless of optimizer.

Article activity feed