Chapter 8: Outlook And Open Problems
The deterministic stationary theory is now complete. Chapters 2—5 proved the prescribed-closure attractor theorem: for any Lipschitz closure map
, the enlarged-state actor-critic-law system defines a continuous semiflow on
with a compact global attractor. Chapter 6 supplied the genuine controlled-chain closure by proving that the frozen invariant-law map
exists uniquely and is Lipschitz under uniform exponential mixing. Chapter 7 then proved the finite-time tracking estimate for the exact fast-law system, turned it into upper semicontinuity of attractors as
, and showed that the abstract pathwise hypothesis follows from a concrete minorization condition.
The two running examples played complementary roles throughout that argument. The Chapter 0 system remained the hand-computable laboratory in which every variable, estimate, and asymptotic picture could be seen directly. The three-state routing chain from Section 2.8 showed that the same theorem applies in a model that already looks like an actual policy-routing problem: the stationary customer mix is explicit, the reference-state geometry is structural, and the fast-slow reduction has a clear operational reading. Concretely, the retail hub served as the reference state for minorization, the resulting contraction constants were fully explicit, and upper semicontinuity confirmed that fast customer-mix equilibration keeps the exact dynamics near the reduced attractor.
This chapter does not add theorems. It marks the next two directions in which the present theory extends, and closes the theoretical part of the notes with a summary of the logical structure.
8.1 Non-Autonomous Forcing
Everything proved in Chapters 2—7 belongs to an autonomous setting. The vector field depends on the current state
, but not on time itself. That is why the correct asymptotic object has been a global attractor for a semiflow.
The prescribed-closure result of Chapters 2—5 is the first step. The exact controlled-chain system analyzed in Chapters 6—7 is the next one: the closure map is no longer prescribed data but is generated by the chain itself. A new kind of problem begins when the environment has genuine external time dependence. The simplest way to see the difference is to modify the Chapter 0 model so that the reward of the preferred action oscillates seasonally:

Now even if we freeze
, the reward landscape still moves with time. The resulting actor-critic system is not an autonomous ODE on
. It is what the dynamical-systems literature calls a non-autonomous process: the solution operator now depends on the starting time as well as the elapsed time, not on elapsed time alone.
This is conceptually different from the enlarged-state closure story of the earlier chapters. In Chapters 2—7, the distribution variable
had to be added because the system was not closed on
. But after enlarging the state space, the system became autonomous. In a genuinely time-dependent environment, enlarging the state space does not remove the time dependence unless we also lift the external forcing into an autonomous drive system.
For the reinforcement-learning reader, this is the right language for settings in which the market, the reward table, or the user population shifts for reasons other than the agent’s current policy. For the dynamical-systems reader, the classical replacement objects are trajectory attractors, which compare entire forward trajectories rather than time-
snapshots (Chepyzhov-Vishik), or pullback attractors, which fix the present time and pull the initial condition back into the distant past (Kloeden-Rasmussen), depending on the structure of the forcing.
The autonomous stationary theory proved in Chapters 2—7 is already mathematically complete on its own, which is why the theory developed in these notes closes at this boundary. The non-autonomous extension is genuinely harder: the invariant-law graph
is no longer stationary, the reduction becomes a time-dependent graph transform, and the correct notion of asymptotic comparison must remember the forcing history. Each of those is a real new task that requires its own theorem.
8.2 Stochastic Perturbations
The deterministic mean dynamics studied throughout these notes arise by averaging out the stochastic noise in the underlying actor-critic recursion. The stochastic extension asks what happens when that noise is retained rather than averaged away.
In the Chapter 0 model, we may imagine the following discrete-time picture. At each step, the platform samples a state, chooses an action, observes a random reward, updates the critic with a noisy temporal-difference term, and updates the actor with a noisy policy-gradient term. The averaged drift of that recursion is the deterministic ODE we studied. The stochastic system itself is not governed by a single deterministic semiflow on
.
The deterministic theory still matters, because it identifies the skeleton that the stochastic system should track on finite time horizons and around which its long-run behavior should concentrate. In that sense, Theorem 7.3 is the right deterministic starting point for a stochastic perturbation theory: it tells us which reduced object the exact fast-law system follows before random sampling noise is added on top.
The order of ideas matters here. In these notes the dissipative mechanism comes from actor confinement, critic coercivity, simplex invariance, and fast-law contraction. Noise enters as a later perturbation of an already dissipative deterministic system, not as the source of the confinement.
The natural asymptotic objects at that stage are no longer deterministic global attractors alone. Depending on the formulation, we are led to invariant measures of the perturbed recursion, to random attractors that depend measurably on the noise sample, to concentration estimates that quantify how closely sample paths stay near the deterministic attractor, or to finite-time tracking results that compare each stochastic trajectory to its deterministic counterpart. In each case, the compact global attractor from Chapters 2—5 provides the deterministic skeleton around which the stochastic asymptotic theory is organized, and the answers rest on the deterministic geometry developed here.
8.3 Closing Perspective
The notes began with a simple observation: a typical actor-critic algorithm updates three coupled quantities at once, and the averaged dynamics of that update cannot be analyzed on the actor parameter space alone. Chapters 2—7 built the mathematical infrastructure needed to take that observation seriously.
The result is a self-contained set of theorems. Chapters 2—3 defined the enlarged state space and established local well-posedness. Chapter 4 proved the four a priori estimates–actor confinement, simplex invariance, critic coercivity, and compact absorption–that trap trajectories in a compact region. Chapter 5 turned those estimates into a global attractor. Chapter 6 supplied the genuine controlled-chain closure by identifying the Lipschitz invariant-law map. Chapter 7 proved that the exact fast-law system tracks the reduced invariant-law dynamics on finite time horizons, and that the exact attractors converge upper semicontinuously to the lifted reduced attractor as the chain mixes quickly: every point of the exact attractor ends up near the lifted reduced one, while the converse is not claimed.
A reader who has followed that arc now has something concrete: the ability to analyze RL mean dynamics as a bona fide dynamical system on the correct state space, with a rigorous attractor picture that does not depend on convergence to a single equilibrium. The Chapter 0 system and the three-state routing chain showed that the abstract theory specializes cleanly to hand-computable models and to models with a realistic operational reading.
What remains is extension rather than repair. Another slow-fast layer for the critic, genuine external time dependence, and stochastic perturbations of the deterministic skeleton are all natural next steps, and each rests on the geometry developed here.
A reader who has worked through these chapters is equipped to read the relevant literature with a clear picture of what the deterministic autonomous skeleton provides and where the new difficulties begin.
Before turning to those extensions, Chapters 9—11 demonstrate the proved theory on two substantive domain models. Chapter 9 develops a systematic protocol for translating domain problems into the Chapter 2 template. Chapter 10 applies that protocol to a content recommendation platform and shows that filter bubbles emerge as competing boundary attractors of the coupled actor-critic-law system. Chapter 11 extends the three-state routing chain to a five-state hub-and-spoke network and identifies hub-heavy versus spoke-heavy traffic configurations as competing attractors. Both chapters work entirely within the autonomous deterministic framework of Chapters 2—7.
Exercises
Exercise 8.1 (Extend: a second fast scale). Write the two-parameter system

State the reduced system one would expect when both
and
are small. Which extra stability input, beyond the hypotheses of Chapter 7, would be needed to justify that further reduction?
Exercise 8.2 (Connect: why Level 3 is genuinely harder). Using the Chapter 0 example with a time-dependent reward
from Section 8.1, identify what changes in the asymptotic objects and in the comparison argument when the reward becomes time-dependent. In particular, explain why the invariant-law graph
is no longer stationary and what replaces the global attractor as the asymptotic target.