Standard reinforcement learning assumes an agent's actions fully determine how the environment changes. But in many real-world systems, outside events also shape the next state. A new paper on arXiv addresses this gap by studying Markov decision processes with continuous state and action spaces where transition dynamics are influenced by external events.

The authors develop algorithms with formal guarantees for this setting, and they analyze the sample complexity required to learn effectively. The abstract is brief, but it signals a move toward making reinforcement learning robust to disturbances that are not under the agent's control.

Because only this single source was available, there are no conflicting findings to compare. The paper's contribution appears to be a theoretical framework for a problem that is often handled only in ad hoc ways.