References

Benaim, M. (1999). Dynamics of stochastic approximation algorithms. In Seminaire de Probabilites XXXIII, Lecture Notes in Mathematics 1709, pp. 1—68. Springer.

Benaim, M., Hofbauer, J., and Sorin, S. (2005). Stochastic approximations and differential inclusions. SIAM Journal on Control and Optimization, 44(1):328—348.

Benveniste, A., Metivier, M., and Priouret, P. (1990). Adaptive Algorithms and Stochastic Approximations. Applications of Mathematics, vol. 22. Springer.

Bertsekas, D. P. and Tsitsiklis, J. N. (1996). Neuro-Dynamic Programming. Athena Scientific.

Bhatnagar, S., Sutton, R. S., Ghavamzadeh, M., and Lee, M. (2009). Natural actor-critic algorithms. Automatica, 45(11):2471—2482.

Borkar, V. S. (1997). Stochastic approximation with two time scales. Systems & Control Letters, 29(5):291—294.

Borkar, V. S. (2008). Stochastic Approximation: A Dynamical Systems Viewpoint. Cambridge University Press.

Borkar, V. S. and Meyn, S. P. (2000). The O.D.E. method for convergence of stochastic approximation and reinforcement learning. SIAM Journal on Control and Optimization, 38(2):447—469.

Chepyzhov, V. V. and Vishik, M. I. (2002). Attractors for Equations of Mathematical Physics. American Mathematical Society Colloquium Publications, vol. 49.

Coddington, E. A. and Levinson, N. (1955). Theory of Ordinary Differential Equations. McGraw-Hill. Reprinted by Krieger, 1984.

Engelking, R. (1989). General Topology. Revised and completed edition. Heldermann Verlag.

Hale, J. K. (1988). Asymptotic Behavior of Dissipative Systems. Mathematical Surveys and Monographs, vol. 25. American Mathematical Society.

Kloeden, P. E. and Rasmussen, M. (2011). Nonautonomous Dynamical Systems. American Mathematical Society Mathematical Surveys and Monographs, vol. 176.

Konda, V. R. and Tsitsiklis, J. N. (2000). Actor-critic algorithms. In Advances in Neural Information Processing Systems 12, pp. 1008—1014. MIT Press.

Kushner, H. J. and Yin, G. G. (2003). Stochastic Approximation and Recursive Algorithms and Applications. 2nd ed. Springer.

Kuznetsov, Y. A. (2004). Elements of Applied Bifurcation Theory. 3rd ed. Applied Mathematical Sciences, vol. 112. Springer.

Ljung, L. (1977). Analysis of recursive stochastic algorithms. IEEE Transactions on Automatic Control, 22(4):551—575.

Meyn, S. P. and Tweedie, R. L. (2009). Markov Chains and Stochastic Stability. 2nd ed. Cambridge University Press.

Munkres, J. R. (2000). Topology. 2nd ed. Prentice Hall.

Norris, J. R. (1997). Markov Chains. Cambridge Series in Statistical and Probabilistic Mathematics. Cambridge University Press.

Pavliotis, G. A. and Stuart, A. M. (2008). Multiscale Methods: Averaging and Homogenization. Springer.

Prytula, V. (2026). Global attractors and fast-slow reduction for finite-state actor-critic mean dynamics. Preprint. Referred to throughout these notes as “the short note.” arXiv:2604.13259.

Puterman, M. L. (1994). Markov Decision Processes: Discrete Stochastic Dynamic Programming. Wiley.

Robbins, H. and Monro, S. (1951). A stochastic approximation method. Annals of Mathematical Statistics, 22(3):400—407.

Robinson, J. C. (2001). Infinite-Dimensional Dynamical Systems: An Introduction to Dissipative Parabolic PDEs and the Theory of Global Attractors. Cambridge University Press.

Strogatz, S. H. (2015). Nonlinear Dynamics and Chaos: With Applications to Physics, Biology, Chemistry, and Engineering. 2nd ed. Westview Press.

Sutton, R. S., McAllester, D., Singh, S., and Mansour, Y. (2000). Policy gradient methods for reinforcement learning with function approximation. In Advances in Neural Information Processing Systems 12, pp. 1057—1063. MIT Press.

Sutton, R. S. and Barto, A. G. (2018). Reinforcement Learning: An Introduction. 2nd ed. MIT Press.

Temam, R. (1997). Infinite-Dimensional Dynamical Systems in Mechanics and Physics. 2nd ed. Applied Mathematical Sciences, vol. 68. Springer.

Tsitsiklis, J. N. and Van Roy, B. (1997). An analysis of temporal-difference learning with function approximation. IEEE Transactions on Automatic Control, 42(5):674—690.

Williams, R. J. (1992). Simple statistical gradient-following algorithms for connectionist reinforcement learning. Machine Learning, 8(3—4):229—256.