References

Abdulla, Mohammed Shahid, and Shalabh Bhatnagar. 2007. “Reinforcement Learning Based Algorithms for Average Cost Markov Decision Processes.” Discrete Event Dynamic Systems 17 (1): 23–52.
Abounadi, Jinane, Dimitri P Bertsekas, and Vivek Borkar. 2002. “Stochastic Approximation for Nonexpansive Maps: Application to q-Learning Algorithms.” SIAM Journal on Control and Optimization 41 (1): 1–22.
Acerbi, C. 2002. “Spectral Measures of Risk: A Coherent Representation of Subjective Risk Aversion.” Journal of Banking & Finance 26 (7): 1505–18.
A.Dvoretzky. 1956. “On Stochastic Approximation.” Proc. Third Berkeley Symp. Math. Stat. And Prob. 1: 39–55.
Agarwal, Alekh, Ofer Dekel, and Lin Xiao. 2010. “Optimal Algorithms for Online Convex Optimization with Multi-Point Bandit Feedback.” In COLT, 28–40.
Alzantot, Moustafa, Yash Sharma, Supriyo Chakraborty, Huan Zhang, Cho-Jui Hsieh, and Mani B. Srivastava. 2019. “GenAttack: Practical Black-Box Attacks with Gradient-Free Optimization.” In Proceedings of the Genetic and Evolutionary Computation Conference, 1111–19. GECCO ’19. New York, NY, USA: Association for Computing Machinery. https://doi.org/10.1145/3321707.3321749.
Anandkumar, Animashree, and Rong Ge. 2016. “Efficient Approaches for Escaping Higher Order Saddle Points in Non-Convex Optimization.” In Conference on Learning Theory, 81–102. PMLR.
Andronov, Aleksandr Aleksandrovich, Aleksandr Adol’fovich Vitt, and Semen Emmanuilovich Khaikin. 2013. Theory of Oscillators: Adiwes International Series in Physics. Vol. 4. Elsevier.
Arnold, Vladimir I. 1992. Ordinary Differential Equations. Springer Science & Business Media.
Artzner, P., F. Delbaen, J. Eber, and D. Heath. 1999. “Coherent measures of risk.” Mathematical Finance 9 (3): 203–28.
Asmussen, S., and P. W. Glynn. 2007. Stochastic Simulation: Algorithms and Analysis. Springer.
Asmussen, Søren, and Peter W Glynn. 2007. Stochastic Simulation: Algorithms and Analysis. Vol. 57. Springer.
Aubin, J., and A. Cellina. 1984. Differential Inclusions: Set-Valued Maps and Viability Theory. Springer.
Aubin, J., and H. Frankowska. 1990. Set-Valued Analysis. Birkhauser.
Balasubramanian, K., and S. Ghadimi. 2022b. “Zeroth-Order Nonconvex Stochastic Optimization: Handling Constraints, High Dimensionality, and Saddle Points.” Foundations of Computational Mathematics 22 (1): 35–76.
———. 2022a. “Zeroth-Order Nonconvex Stochastic Optimization: Handling Constraints, High Dimensionality, and Saddle Points.” Foundations of Computational Mathematics 22 (1): 35–76.
Barakat, A., P. Bianchi, W. Hachem, and S. Schechtman. 2021. “Stochastic optimization with momentum: Convergence, fluctuations, and traps avoidance.” Electronic Journal of Statistics 15 (2): 3892–3947. https://doi.org/10.1214/21-EJS1880.
Bardou, O., N. Frikha, and G. Pages. 2009. “Computing VaR and CVaR using stochastic approximation and adaptive unconstrained importance sampling.” Monte Carlo Methods and Applications 15 (3): 173–210.
Benaïm, M. 1996. “A Dynamical System Approach to Stochastic Approximations.” SIAM J. Control Optim. 34 (2): 437–72.
———. 1999. “Dynamics of Stochastic Approximation Algorithms.” Seminaire De Probabilities (Strasbourg) 1709: 1–68.
Benaïm, M., and M. W. Hirsch. 1996. “Asymptotic Pseudotrajectories and Chain Recurrent Flows, with Applications.” J. Dynam. Differential Equations 8: 141–76.
Benaïm, M., J. Hofbauer, and S. Sorin. 2005. “Stochastic Approximations and Differential Inclusions.” SIAM Journal on Control and Optimization, 328–48.
———. 2012. “Perturbations of Set-Valued Dynamical Systems, with Applications to Game Theory.” Dynamic Games and Applications 2 (2): 195–205.
Berahas, Albert S., Liyuan Cao, Krzysztof Choromanski, and Katya Scheinberg. 2022. “A Theoretical and Empirical Comparison of Gradient Approximations in Derivative-Free Optimization.” Foundations of Computational Mathematics 22 (2): 507–60.
Bertsekas, D. P. 2012. Dynamic Programming and Optimal Control, Vol.II. Athena Scientific.
Bertsekas, D. P., and J. N. Tsitsiklis. 1996. Neuro-Dynamic Programming. Athena Scientific.
Bertsekas, Dimitri. 2019. Reinforcement Learning and Optimal Control. Vol. 1. Athena Scientific.
Bertsekas, Dimitri P. 1999. Nonlinear Programming. 2nd ed. Belmont, MA: Athena Scientific.
Bertsekas, DP, and JN Tsitsiklis. 1989. Parallel and Distributed Computation. Prentice Hall Inc.
Bhagoji, Arjun Nitin, Warren He, Bo Li, and Dawn Song. 2018. “Practical Black-Box Attacks on Deep Neural Networks Using Efficient Query Mechanisms.” In Computer Vision – ECCV 2018, edited by Vittorio Ferrari, Martial Hebert, Cristian Sminchisescu, and Yair Weiss, 158–74. Cham: Springer International Publishing.
Bhatnagar, S. 2005. “Adaptive Multivariate Three-Timescale Stochastic Approximation Algorithms for Simulation Based Optimization.” ACM Transactions on Modeling and Computer Simulation 15 (1): 74–107.
———. 2007. “Adaptive Newton-Based Smoothed Functional Algorithms for Simulation Optimization.” ACM Transactions on Modeling and Computer Simulation 18 (1): 2:1–35.
Bhatnagar, Shalabh. 2010. “An Actor–Critic Algorithm with Function Approximation for Discounted Cost Constrained Markov Decision Processes.” Systems & Control Letters 59 (12): 760–66.
———. 2023. “The Reinforce Policy Gradient Algorithm Revisited.” arXiv Preprint arXiv:2310.05000.
Bhatnagar, Shalabh, and Mohammed Shahid Abdulla. 2008. “Simulation-Based Optimization Algorithms for Finite-Horizon Markov Decision Processes.” Simulation 84 (12): 577–600.
Bhatnagar, Shalabh, and K Mohan Babu. 2008. “New Algorithms of the q-Learning Type.” Automatica 44 (4): 1111–19.
Bhatnagar, Shalabh, and Vivek S Borkar. 1998. “A Two Timescale Stochastic Approximation Scheme for Simulation-Based Parametric Optimization.” Probability in the Engineering and Informational Sciences 12 (4): 519–31.
———. 2003. “Multiscale Chaotic SPSA and Smoothed Functional Algorithms for Simulation Optimization.” Simulation 79 (10): 568–80.
Bhatnagar, Shalabh, Vivek S Borkar, Madhukar Akarapu, and Shie Mannor. 2006. “A Simulation-Based Algorithm for Ergodic Control of Markov Chains Conditioned on Rare Events.” Journal of Machine Learning Research 7 (10).
Bhatnagar, Shalabh, Michael C Fu, Steven I Marcus, I Wang, et al. 2003. “Two-Timescale Simultaneous Perturbation Stochastic Approximation Using Deterministic Perturbation Sequences.” ACM Transactions on Modeling and Computer Simulation 13 (2): 180–209.
Bhatnagar, Shalabh, N Hemachandra, and Vivek Kumar Mishra. 2011. “Stochastic Approximation Algorithms for Constrained Optimization via Simulation.” ACM Transactions on Modeling and Computer Simulation (TOMACS) 21 (3): 1–22.
Bhatnagar, Shalabh, and Shishir Kumar. 2004. “A Simultaneous Perturbation Stochastic Approximation-Based Actor-Critic Algorithm for Markov Decision Processes.” IEEE Transactions on Automatic Control 49 (4): 592–98.
Bhatnagar, Shalabh, and K Lakshmanan. 2012. “An Online Actor–Critic Algorithm with Function Approximation for Constrained Markov Decision Processes.” Journal of Optimization Theory and Applications 153: 688–708.
———. 2016. “Multiscale q-Learning with Linear Function Approximation.” Discrete Event Dynamic Systems 26: 477–509.
Bhatnagar, Shalabh, Vivek Kumar Mishra, and Nandyala Hemachandra. 2011. “Stochastic Algorithms for Discrete Parameter Simulation Optimization.” IEEE Transactions on Automation Science and Engineering 8 (4): 780–93.
Bhatnagar, Shalabh, and L. A. Prashanth. 2023. “Generalized Simultaneous Perturbation Stochastic Approximation with Reduced Estimator Bias.” In 2023 57th Annual Conference on Information Sciences and Systems (CISS), 1–6. IEEE.
Bhatnagar, S, H. L. Prasad, and L. A. Prashanth. 2013. Stochastic Recursive Algorithms for Optimization: Simultaneous Perturbation Methods (Lecture Notes in Control and Information Sciences). Vol. 434. Springer.
Bhatnagar, S., and L. A. Prashanth. 2015. “Simultaneous Perturbation Newton Algorithms for Simulation Optimization.” Journal of Optimization Theory and Applications 164 (2): 621–43.
Bhatnagar, S., R. S. Sutton, M. Ghavamzadeh, and M. Lee. 2009. “Natural Actor-Critic Algorithms.” Automatica 45 (11): 2471–82.
Bhavsar, N., and L. A. Prashanth. 2022. “Non-Asymptotic Bounds for Stochastic Optimization with Biased Noisy Gradient Oracles.” IEEE Transactions on Automatic Control, 1–1. https://doi.org/10.1109/TAC.2022.3159748.
Bhojanapalli, Srinadh, Behnam Neyshabur, and Nati Srebro. 2016. “Global Optimality of Local Search for Low Rank Matrix Recovery.” Advances in Neural Information Processing Systems 29.
Billingsley, Patrick. 2013. Convergence of Probability Measures. John Wiley & Sons.
———. 2017. Probability and Measure. John Wiley & Sons.
Borkar, V. S. 1995. Probability Theory: An Advanced Course. New York: Springer.
———. 2022. Stochastic Approximation: A Dynamical Systems Viewpoint, 2’nd Edition. Cambridge University Press.
Borkar, V. S., and S. P. Meyn. 1999. “The O.D.E. Method for Convergence of Stochastic Approximation and Reinforcement Learning.” SIAM J. Control Optim 38: 447–69.
———. 2000. “The O.D.E. Method for Convergence of Stochastic Approximation and Reinforcement Learning.” SIAM Journal of Control and Optimization 38 (2): 447–69.
Borkar, Vivek S. 2003. “Avoidance of Traps in Stochastic Approximation.” Systems & Control Letters 50 (1): 1–9.
Bottou, L., F. E. Curtis, and J. Nocedal. 2018. “Optimization Methods for Large-Scale Machine Learning.” SIAM Review 60 (2): 223–311.
Boyd, Stephen, and Lieven Vandenberghe. 2004. Convex Optimization. Cambridge university press.
Brandiere, Odile, and Marie Duflo. 1996. “Les Algorithmes Stochastiques Contournent-Ils Les Pièges?” In Annales de l’IHP Probabilités Et Statistiques, 32:395–427. 3.
Bunch, James R, and Beresford N Parlett. 1971. “Direct Methods for Solving Symmetric Indefinite Systems of Linear Equations.” SIAM Journal on Numerical Analysis 8 (4): 639–55.
Cai, HanQin, Daniel McKenzie, Wotao Yin, and Zhenliang Zhang. 2022. “Zeroth-Order Regularized Optimization (Zoro): Approximately Sparse Gradients and Adaptive Sampling.” SIAM Journal on Optimization 32 (2): 687–714.
Carmon, Yair, John Duchi, Oliver Hinder, and Aaron Sidford. 2016. “Accelerated Methods for Non-Convex Optimization.” arXiv Preprint arXiv:1611.00756.
Cassandras, Christos G, and Stéphane Lafortune. 2008. Introduction to Discrete Event Systems. Springer.
Chen, H. F., L. Guo, and A. J. Gao. 1987. “Convergence and robustness of the Robbins-Monro algorithm truncated at randomly varying bounds.” Stochastic Processes and Their Applications 27: 217–31.
Chen, Han-Fu, Tyrone E Duncan, and Bozenna Pasik-Duncan. 1999. “A Kiefer-Wolfowitz Algorithm with Randomized Differences.” IEEE Transactions on Automatic Control 44 (3): 442–53.
Chen, Jianbo, Michael I Jordan, and Martin J Wainwright. 2020. “HopSkipJumpAttack: A Query-Efficient Decision-Based Attack.” In IEEE Symposium on Security and Privacy (SP), 1277–94. IEEE.
Chen, Pin-Yu, Huan Zhang, Yash Sharma, Jinfeng Yi, and Cho-Jui Hsieh. 2017. “ZOO: Zeroth Order Optimization Based Black-Box Attacks to Deep Neural Networks Without Training Substitute Models.” In Proceedings of the 10th ACM Workshop on Artificial Intelligence and Security, 15–26. AISec ’17. New York, NY, USA: Association for Computing Machinery. https://doi.org/10.1145/3128572.3140448.
Chen, Shuhang, Adithya Devraj, Ana Busic, and Sean Meyn. 2020. “Explicit Mean-Square Error Bounds for Monte-Carlo and Linear Stochastic Approximation.” In International Conference on Artificial Intelligence and Statistics, 4173–83. PMLR.
Chin, D. C. 1997. “Comparative Study of Stochastic Algorithms for System Optimization Based on Gradient Approximations.” IEEE Transactions on Systems, Man, and Cybernetics, Part B: Cybernetics 27 (2): 244–49.
Choromanski, Krzysztof, Mark Rowland, Vikas Sindhwani, Richard Turner, and Adrian Weller. 2018. “Structured Evolution with Compact Architectures for Scalable Policy Optimization.” In Proceedings of the 35th International Conference on Machine Learning, edited by Jennifer Dy and Andreas Krause, 80:970–78. Proceedings of Machine Learning Research. PMLR.
Coddington, Earl A, Norman Levinson, and T Teichmann. 1956. “Theory of Ordinary Differential Equations.” American Institute of Physics.
Cover, Thomas M, and Joy A Thomas. 2012. Elements of Information Theory. John Wiley & Sons.
Dalal, G., B. Szorenyi, G. Thoppe, and S. Mannor. 2018. “Finite Sample Analysis of Two-Timescale Stochastic Approximation with Applications to Reinforcement Learning.” In Conference on Learning Theory, 1–35.
Dippon, Jürgen. 2003. “Accelerated Randomized Stochastic Optimization.” The Annals of Statistics 31 (4): 1260–81.
Dong, Yinpeng, Hang Su, Jun Zhu Wu, Ziwei Zhang, and Jiansheng Liu. 2020. “Improving Black-Box Adversarial Attacks with a Transfer-Based Prior.” In Advances in Neural Information Processing Systems (NeurIPS), 17631–41.
Duchi, John C., Peter L. Bartlett, and Martin J. Wainwright. 2012. “Randomized Smoothing for Stochastic Optimization.” SIAM Journal on Optimization 22 (2): 674–701.
Dunkel, J., and S. Weber. 2010. “Stochastic Root Finding and Efficient Estimation of Convex Risk Measures.” Operations Research 58 (5): 1505–21.
Durrett, Rick. 2019. Probability: Theory and Examples. Vol. 49. Cambridge university press.
Erdogdu, M. A. 2016. “Newton-Stein method: An optimization method for glms via stein’s lemma.” Journal of Machine Learning Research 17 (215): 1–52.
Fabian, V. 1968. “On Asymptotic Normality in Stochastic Approximation.” The Annals of Mathematical Statistics, 1327–32.
———. 1971. “Stochastic Approximation.” In Optimizing Methods in Statistics (Ed. J.j.rustagi), 439–70. New York: Academic Press.
Filippov, Aleksei Fedorovich. 2013. Differential Equations with Discontinuous Righthand Sides: Control Systems. Vol. 18. Springer Science & Business Media.
Flaxman, Abraham D, Adam Tauman Kalai, and H Brendan McMahan. 2005. “Online Convex Optimization in the Bandit Setting: Gradient Descent Without a Gradient.” In SODA, 385–94.
Föllmer, H., and A. Schied. 2002. “Convex Measures of Risk and Trading Constraints.” Finance and Stochastics 6 (4): 429–47.
Frikha, N., and S. Menozzi. 2012. “Concentration Bounds for Stochastic Approximations.” Electronic Communications in Probability 17: no. 47, 1–15.
Fu, M. C., ed. 2015. Handbook of Simulation Optimization. Springer.
Furmston, T., G. Lever, and D. Barber. 2016. “Approximate Newton Methods for Approximate Policy Search in Markov Decision Processes.” Journal of Machine Learning Research 17: 1–51.
Gadat, S., and I. Gavra. 2022. “Asymptotic Study of Stochastic Adaptive Algorithms in Non-Convex Landscape.” Journal of Machine Learning Research 23 (228): 1–54.
Gallager, Robert G. 2013. Stochastic Processes: Theory for Applications. Cambridge University Press.
Gasnikov, Alexander, Anton Novitskii, Vasilii Novitskii, Farshed Abdukhakimov, Dmitry Kamzolov, Aleksandr Beznosikov, Martin Takac, Pavel Dvurechensky, and Bin Gu. 2022. “The Power of First-Order Smooth Optimization for Black-Box Non-Smooth Problems.” In Proceedings of the 39th International Conference on Machine Learning, edited by Kamalika Chaudhuri, Stefanie Jegelka, Le Song, Csaba Szepesvari, Gang Niu, and Sivan Sabato, 162:7241–65. Proceedings of Machine Learning Research. PMLR.
Ge, R., F. Huang, C. Jin, and Y. Yuan. 2015. “Escaping from Saddle Points – Online Stochastic Gradient for Tensor Decomposition.” Conference of Learning Theory.
Ge, Rong, Chi Jin, and Yi Zheng. 2017. “No Spurious Local Minima in Nonconvex Low Rank Problems: A Unified Geometric Analysis.” In International Conference on Machine Learning, 1233–42. PMLR.
Ge, Rong, Jason D Lee, and Tengyu Ma. 2016. “Matrix Completion Has No Spurious Local Minimum.” Advances in Neural Information Processing Systems 29.
Gelfand, Saul B, and Sanjoy K Mitter. 1991. “Recursive Stochastic Algorithms for Global Optimization in r^d.” SIAM Journal on Control and Optimization 29 (5): 999–1018.
Gerencser, Laszlo, Stacy D Hill, and Zsuzsanna Vago. 1999. “Optimization over Discrete Sets via SPSA.” In Proceedings of the 31st Conference on Winter Simulation: Simulation—a Bridge to the Future-Volume 1, 466–70.
Ghadimi, S., and G. Lan. 2013. “Stochastic First-and Zeroth-Order Methods for Nonconvex Stochastic Programming.” SIAM Journal on Optimization 23 (4): 2341–68.
Ghoshdastidar, D., A. Dukkipati, and S. Bhatnagar. 2014a. “Newton-Based Stochastic Optimization Using q-Gaussian Smoothed Functional Algorithms.” Automatica 50 (10): 2606–14.
———. 2014b. “Smoothed Functional Algorithms for Stochastic Optimization Using q-Gaussian Distributions.” ACM Transactions on Modeling and Computer Simulation 26 (3): 17:1–26.
Giesecke, K., T. Schmidt, and S. Weber. 2008. “Measuring the Risk of Large Losses.” Journal of Investment Management, Fourth Quarter.
Grimmett, Geoffrey, and David Stirzaker. 2020. Probability and Random Processes. Oxford university press.
Hegde, V., A. S. Menon, L. A. Prashanth, and K. Jagannathan. 2021. “Online Estimation and Optimization of Utility-Based Shortfall Risk.” Papers 2111.08805. arXiv.org.
Hirsch, Morris W, Stephen Smale, and Robert L Devaney. 2013. Differential Equations, Dynamical Systems, and an Introduction to Chaos. Academic press.
Hu, X., L. A. Prashanth, A. György, and C. Szepesvári. 2016. “(Bandit) Convex Optimization with Biased Noisy Gradient Oracles.” In Artificial Intelligence and Statistics, 819–28.
Huang, Feihu, Lue Tao, and Songcan Chen. 2020. “Accelerated Stochastic Gradient-Free and Projection-Free Methods.” In Proceedings of the 37th International Conference on Machine Learning, edited by Hal Daumé III and Aarti Singh, 119:4519–30. Proceedings of Machine Learning Research. PMLR.
Hurley, M. 1995. “Chain Recurrence, Semiflows, and Gradients.” Journal of Dynamics and Differential Equations, 437–56.
Ilyas, Andrew, Logan Engstrom, Anish Athalye, and Jessy Lin. 2018. “Black-Box Adversarial Attacks with Limited Queries and Information.” In Proceedings of the 35th International Conference on Machine Learning, edited by Jennifer Dy and Andreas Krause, 80:2137–46. Proceedings of Machine Learning Research. PMLR.
Ilyas, Andrew, Logan Engstrom, and Aleksander Madry. 2019. “Prior Convictions: Black-Box Adversarial Attacks with Bandits and Priors.” In 7th International Conference on Learning Representations (ICLR).
Ince, Edward L. 1956. Ordinary Differential Equations. Courier Corporation.
Jain, P., D. Nagaraj, and P. Netrapalli. 2021. “Making the Last Iterate of SGD Information Theoretically Optimal.” SIAM Journal on Optimization 31 (2): 1108–30.
Ji, Kaiyi, Zhe Wang, Yi Zhou, and Yingbin Liang. 2019. “Improved Zeroth-Order Variance Reduced Algorithms and Analysis for Nonconvex Optimization.” In Proceedings of the 36th International Conference on Machine Learning, edited by Kamalika Chaudhuri and Ruslan Salakhutdinov, 97:3100–3109. Proceedings of Machine Learning Research. PMLR.
Jin, C., R. Ge, P. Netrapalli, S. M. Kakade, and M. I. Jordan. 2017. “How to Escape Saddle Points Efficiently.” ICML, 1724–32.
Jin, C., P. Netrapalli, R. Ge, S. M. Kakade, and M. I Jordan. 2021. “On Nonconvex Optimization for Machine Learning: Gradients, Stochasticity, and Saddle Points.” Journal of the ACM (JACM) 68 (2): 1–29.
Karandikar, Rajeeva Laxman, and Mathukumalli Vidyasagar. 2024. “Convergence Rates for Stochastic Approximation: Biased Noise with Unbounded Variance, and Applications.” Journal of Optimization Theory and Applications, 1–39.
Karmakar, Prasenjit, and Shalabh Bhatnagar. 2018. “Two Time-Scale Stochastic Approximation with Controlled Markov Noise and Off-Policy Temporal-Difference Learning.” Mathematics of Operations Research 43 (1): 130–51.
———. 2021. “Stochastic Approximation with Iterate-Dependent Markov Noise Under Verifiable Conditions in Compact State Space with the Stability of Iterates Not Ensured.” IEEE Transactions on Automatic Control 66 (12): 5941–54.
Katkovnik, V. Ya, and Yu Kulchitsky. 1972. “Convergence of a Class of Random Search Algorithms.” Automation Remote Control 8: 1321–26.
Kawaguchi, Kenji. 2016. “Deep Learning Without Poor Local Minima.” Advances in Neural Information Processing Systems 29.
Kiefer, J., and J. Wolfowitz. 1952. “Stochastic Estimation of the Maximum of a Regression Function.” Ann. Math. Statist. 23: 462–66.
Kirkpatrick, Scott, C Daniel Gelatt Jr, and Mario P Vecchi. 1983. “Optimization by Simulated Annealing.” Science 220 (4598): 671–80.
Konda, Vijay R, and John N Tsitsiklis. 2003. “On Actor-Critic Algorithms.” SIAM Journal on Control and Optimization 42 (4): 1143–66.
Kornowski, Guy, and Ohad Shamir. 2024. “An Algorithm with Optimal Dimension Dependence for Zero-Order Nonsmooth Nonconvex Stochastic Optimization.” Journal of Machine Learning Research 25 (122): 1–14.
Kozak, David, Cesare Molinari, Lorenzo Rosasco, Luis Tenorio, and Silvia Villa. 2023. “Zeroth-Order Optimization with Orthogonal Random Directions.” Mathematical Programming 199 (1): 1179–219.
Kulkarni, Vidyadhar G. 2016. Modeling and Analysis of Stochastic Systems. Chapman; Hall/CRC.
Kushner, H. J., and D. S. Clark. 1978. Stochastic Approximation Methods for Constrained and Unconstrained Systems. New York: Springer Verlag.
Kushner, H. J., and G. G. Yin. 2003. Stochastic Approximation Algorithms and Applications, 2’nd Ed. New York: Springer Verlag.
Lasalle, J. P., and S. Lefschetz. 1961. Stability by Liapunov’s Direct Method with Applications. New York: Academic Press.
Lei, J. 2020. “Convergence and concentration of empirical measures under Wasserstein distance in unbounded functional spaces.” Bernoulli 26 (1): 767–98.
Levin, David A, and Yuval Peres. 2017. Markov Chains and Mixing Times. Vol. 107. American Mathematical Soc.
Ljung, L. 1977. “Analysis of Recursive Stochastic Algorithms.” Automatic Control, IEEE Transactions on 22 (4): 551–75.
Malladi, Sadhika, Tianyu Gao, Eshaan Nichani, Alex Damian, Jason D Lee, Danqi Chen, and Sanjeev Arora. 2023. “Fine-Tuning Language Models with Just Forward Passes.” Advances in Neural Information Processing Systems 36: 53038–75.
Mania, Horia, Aurelia Guy, and Benjamin Recht. 2018. “Simple Random Search of Static Linear Policies Is Competitive for Reinforcement Learning.” Advances in Neural Information Processing Systems 31.
Maniyar, Mizhaan P., L. A. Prashanth, Akash Mondal, and Shalabh Bhatnagar. 2024. “A Cubic-Regularized Policy Newton Algorithm for Reinforcement Learning.” In Proceedings of the 27th International Conference on Artificial Intelligence and Statistics, edited by Sanjoy Dasgupta, Stephan Mandt, and Yingzhen Li, 238:4708–16. Proceedings of Machine Learning Research. PMLR.
Maryak, John L, and Daniel C Chin. 2001. “Global Random Optimization by Simultaneous Perturbation Stochastic Approximation.” In Proceedings of the 2001 American Control Conference.(cat. No. 01CH37148), 2:756–62. IEEE.
Meyn, Sean. 2022. Control Systems and Reinforcement Learning. Cambridge University Press.
Meyn, Sean P, and Richard L Tweedie. 2012. Markov Chains and Stochastic Stability. Springer Science & Business Media.
Mondal, A., L. A. Prashanth, and S. Bhatnagar. 2024. “Truncated Cauchy Random Perturbations for Smoothed Functional-Based Stochastic Optimization.” Automatica 162: 111528. https://doi.org/https://doi.org/10.1016/j.automatica.2024.111528.
Mou, Wenlong, Chris Junchi Li, Martin J Wainwright, Peter L Bartlett, and Michael I Jordan. 2020. “On Linear Stochastic Approximation: Fine-Grained Polyak-Ruppert and Non-Asymptotic Concentration.” In Conference on Learning Theory, 2947–97. PMLR.
Mukhoty, Bhaskar, Velibor Bojkovic, William de Vazelhes, Xiaohan Zhao, Giulia De Masi, Huan Xiong, and Bin Gu. 2023. “Direct Training of SNN Using Local Zeroth Order Method.” In Advances in Neural Information Processing Systems (NeurIPS).
Nandakumaran, A. K, P. S Datti, and R. K George. 2017. Ordinary Differential Equations: Principles and Applications. Cambridge University Press.
Nemirovski, Arkadi, Anatoli Juditsky, Guanghui Lan, and Alexander Shapiro. 2009. “Robust Stochastic Approximation Approach to Stochastic Programming.” SIAM Journal on Optimization 19 (4): 1574–1609.
Nesterov, Y., and B. T. Polyak. 2007. “Cubic Regularization of Newton Method and Its Global Performance.” Mathematical Programming 112: 159–81.
Nesterov, Y., and V. Spokoiny. 2017. “Random Gradient-Free Minimization of Convex Functions.” Foundations of Computational Mathematics 17 (2): 527–66.
Nesterov, Yurii, and Boris Polyak. 2006. “Cubic regularization of Newton method and its global performance.” Math. Program. 108 (August): 177–205.
Nocedal, Jorge, and Stephen J Wright. 1999. Numerical Optimization. Springer.
Norris, James R. 1998. Markov Chains. 2. Cambridge university press.
Pachal, Soumen, Shalabh Bhatnagar, and L. A. Prashanth. 2023. “Generalized Simultaneous Perturbation-Based Gradient Search with Reduced Estimator Bias.” arXiv Preprint arXiv:2212.10477.
Pemantle, R. 1990. “Non-Convergence to Unstable Points in Urn Models and Stochastic Approximations.” The Annals of Probability 18(2): 698–712.
Polyak, B. T., and A. B. Tsybakov. 1990. “Optimal Orders of Accuracy for Search Algorithms of Stochastic Optimization.” Problems in Information Transmission, 126–33.
Polyak, Boris T, and Anatoli B Juditsky. 1992. “Acceleration of Stochastic Approximation by Averaging.” SIAM Journal on Control and Optimization 30 (4): 838–55.
Powell, Warren B. 2021. “Reinforcement Learning and Stochastic Optimization.” John Wiley & Sons Hoboken, NJ.
Prashanth, L. A., and S. P. Bhat. 2022. “A Wasserstein Distance Approach for Concentration of Empirical Risk Estimates.” Journal of Machine Learning Research 23 (238): 1–61.
Prashanth, L. A., and S. Bhatnagar. 2012. “Threshold Tuning Using Stochastic Optimization for Graded Signal Control.” IEEE Transactions on Vehicular Technology 61 (9): 3865–80.
Prashanth, L. A., S. Bhatnagar, N. Bhavsar, M. Fu, and S. I. Marcus. 2020. “Random Directions Stochastic Approximation with Deterministic Perturbations.” IEEE Transactions on Automatic Control 65 (6): 2450–65.
Prashanth, L. A., Shalabh Bhatnagar, Michael C. Fu, and Steve I Marcus. 2017. “Adaptive System Optimization Using Random Directions Stochastic Approximation.” IEEE Transactions on Automatic Control 62 (5): 2223–38.
Prashanth, L. A., A. Chatterjee, and S. Bhatnagar. 2014. “Two Timescale Convergent q-Learning for Sleep-Scheduling in Wireless Sensor Networks.” Wireless Networks 20: 2589–2604.
Prashanth, L. A., and Mohammad Ghavamzadeh. 2016. “Variance-Constrained Actor-Critic Algorithms for Discounted and Average Reward MDPs.” Machine Learning 105: 367–417.
Prashanth, L. A., N. Korda, and R. Munos. 2021. “Concentration Bounds for Temporal Difference Learning with Linear Function Approximation: The Case of Batch Data and Uniform Sampling.” Mach. Learn. 110 (3): 559–618.
Ramaswamy, A., and S. Bhatnagar. 2016. “A Generalization of the Borkar-Meyn Theorem for Stochastic Recursive Inclusions.” Mathematics of Operations Research 42 (3): 648–61.
———. 2018. “Analysis of Gradient Descent Methods with Nondiminishing Bounded Errors.” IEEE Transactions on Automatic Control 63 (5): 1465–71.
———. 2019. “Stability of Stochastic Approximations with Controlled Markov Noise and Temporal Difference Learning.” IEEE Transactions on Automatic Control 64 (6): 2614–20.
———. 2021. “Analyzing Approximate Value Iteration Algorithms.” Mathematics of Operations Research. https://doi.org/10.1287/moor.2021.1202.
Ramaswamy, Arunselvan, and Shalabh Bhatnagar. 2016. “Stochastic Recursive Inclusion in Two Timescales with an Application to the Lagrangian Dual Problem.” Stochastics 88 (8): 1173–87.
Rando, Marco, Cesare Molinari, Lorenzo Rosasco, and Silvia Villa. 2023. “An Optimal Structured Zeroth-Order Algorithm for Non-Smooth Optimization.” In Advances in Neural Information Processing Systems, edited by A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, 36:36738–67. Curran Associates, Inc.
Rando, Marco, Cesare Molinari, Silvia Villa, and Lorenzo Rosasco. 2024. “Stochastic Zeroth Order Descent with Structured Directions.” Computational Optimization and Applications, October.
Rastogi, Pushpendre, Jingyi Zhu, and James C Spall. 2016. “Efficient Implementation of Enhanced Adaptive Simultaneous Perturbation Algorithms.” In 2016 Annual Conference on Information Science and Systems (CISS), 298–303. IEEE.
R.Blum, J. 1954. “Approximation Methods Which Converge with Probability One.” Annals of Mathematical Statistics, 382–86.
Robbins, H., and S. Monro. 1951. “A Stochastic Approximation Method.” Ann. Math. Statist. 22: 400–407.
Rockafellar, R. T., and S. Uryasev. 2000. “Optimization of conditional value-at-risk.” Journal of Risk 2: 21–42.
Royden, H. L., and P. M. Fitzpatrick. 2010. Real Analysis. 4th ed. Boston: Pearson.
Rubinstein, R. Y. 1981. Simulation and the Monte Carlo Method. New York: Wiley.
Ruppert, D. 1985. “A Newton-Raphson Version of the Multivariate Robbins-Monro Procedure.” Annals of Statistics 13: 236–45.
Saha, Ankan, and Ambuj Tewari. 2011. “Improved Regret Guarantees for Online Smooth Convex Optimization with Bandit Feedback.” In AISTATS, 636–42.
Salimans, Tim, Jonathan Ho, Xi Chen, Szymon Sidor, and Ilya Sutskever. 2017. “Evolution Strategies as a Scalable Alternative to Reinforcement Learning.” arXiv Preprint arXiv:1703.03864.
Serfling, Robert J. 2009. Approximation Theorems of Mathematical Statistics. Vol. 162. John Wiley & Sons.
Shamir, Ohad. 2017. “An Optimal Algorithm for Bandit and Zero-Order Convex Optimization with Two-Point Feedback.” Journal of Machine Learning Research 18 (52): 1–11.
Spall, J. C. 2000. “Adaptive Stochastic Approximation by the Simultaneous Perturbation Method.” IEEE Trans. Autom. Contr. 45: 1839–53.
———. 2005. Introduction to Stochastic Search and Optimization: Estimation, Simulation, and Control. Vol. 65. John Wiley & Sons.
Spall, James C. 1992. “Multivariate Stochastic Approximation Using a Simultaneous Perturbation Gradient Approximation.” IEEE Transactions on Automatic Control 37 (3): 332–41.
———. 1997. “A One-Measurement Form of Simultaneous Perturbation Stochastic Approximation.” Automatica 33 (1): 109–12.
———. 2009. “Feedback and Weighting Mechanisms for Improving Jacobian Estimates in the Adaptive Simultaneous Perturbation Algorithm.” IEEE Transactions on Automatic Control 54 (6): 1216–29.
Stein, C. 1972. “A Bound for the Error in the Normal Approximation to the Distribution of a Sum of Dependent Random Variables.” In Proceedings of the Sixth Berkeley Symposium on Mathematical Statistics and Probability, Volume 2: Probability Theory, 6:583–603. University of California Press.
———. 1981. “Estimation of the Mean of a Multivariate Normal Distribution.” The Annals of Statistics, 1135–51.
Styblinski, M. A., and T.-S. Tang. 1990. “Experiments in Nonconvex Optimization: Stochastic Approximation with Function Smoothing and Simulated Annealing.” Neural Networks 3: 467–83.
Sun, Ju, Qing Qu, and John Wright. 2016. “Complete Dictionary Recovery over the Sphere i: Overview and the Geometric Picture.” IEEE Transactions on Information Theory 63 (2): 853–84.
Sutton, R. S., and A. W. Barto. 2018. Reinforcement Learning, 2’nd Edition. MIT Press.
Sutton, Richard S, Hamid Reza Maei, Doina Precup, Shalabh Bhatnagar, David Silver, Csaba Szepesvári, and Eric Wiewiora. 2009. “Fast Gradient-Descent Methods for Temporal-Difference Learning with Linear Function Approximation.” In Proceedings of the 26th Annual International Conference on Machine Learning, 993–1000.
Sutton, Richard S, David McAllester, Satinder Singh, and Yishay Mansour. 1999. “Policy Gradient Methods for Reinforcement Learning with Function Approximation.” Advances in Neural Information Processing Systems 12.
Swain, J. J. 2017. “Simulation Software Survey-Simulation: New and Improved Reality Show.” OR/MS Today 44 (5): 38–49.
Tamar, Aviv, Yinlam Chow, Mohammad Ghavamzadeh, and Shie Mannor. 2015. “Policy Gradient for Coherent Risk Measures.” In Advances in Neural Information Processing Systems. Vol. 28. Curran Associates, Inc.
Tripuraneni, N., M. Stern, C. Jin, J. Regier, and M. I. Jordan. 2018. “Stochastic Cubic Regularization for Fast Nonconvex Optimization.” In NeurIPS. Vol. 31. Curran Associates, Inc.
Tropp, J. A. 2016. “The Expected Norm of a Sum of Independent Random Matrices: An Elementary Approach.” In High Dimensional Probability VII: The Cargèse Volume, 173–202.
Tsitsiklis, J. N., and B. Van Roy. 1997. “An Analysis of Temporal-Difference Learning with Function Approximation.” IEEE Transactions on Automatic Control 42 (5): 674–90.
Tsitsiklis, John N. 1994. “Asynchronous Stochastic Approximation and q-Learning.” Machine Learning 16 (3): 185–202.
Vijayan, Nithia, and L. A. Prashanth. 2021. “Smoothed Functional-Based Gradient Algorithms for Off-Policy Reinforcement Learning: A Non-Asymptotic Viewpoint.” Systems & Control Letters 155: 104988.
———. 2023. “A Policy Gradient Approach for Optimization of Smooth Risk Measures.” In Proceedings of the Thirty-Ninth Conference on Uncertainty in Artificial Intelligence, edited by Robin J. Evans and Ilya Shpitser, 216:2168–78. Proceedings of Machine Learning Research. PMLR.
Wang, Tianyu, and Yasong Feng. 2024. “Convergence Rates of Zeroth Order Gradient Descent for Lojasiewicz Functions.” INFORMS Journal on Computing.
Williams, Ronald J. 1992. “Simple Statistical Gradient-Following Algorithms for Connectionist Reinforcement Learning.” Machine Learning 8: 229–56.
Wright, S. J., and B. Recht. 2022. Optimization for Data Analysis. Cambridge University Press.
Yaji, V. G., and S. Bhatnagar. 2019. “Analysis of Stochastic Approximation Schemes with Set-Valued Maps in the Absence of a Stability Guarantee and Their Stabilization.” IEEE Transactions on Automatic Control 65 (3): 1100–1115.
Yaji, Vinayaka G, and Shalabh Bhatnagar. 2018. “Stochastic Recursive Inclusions with Non-Additive Iterate-Dependent Markov Noise.” Stochastics 90 (3): 330–63.
———. 2020. “Stochastic Recursive Inclusions in Two Timescales with Nonadditive Iterate-Dependent Markov Noise.” Mathematics of Operations Research 45 (4): 1405–44.
Yao, A. C. C. 1977. “Probabilistic Computations: Toward a Unified Measure of Complexity.” In FOCS, 222–27.
Zhang, Yimeng, Yuguang Yao, Jinghan Jia, Jinfeng Yi, Mingyi Hong, Shiyu Chang, and Sijia Liu. 2022. “How to Robustify Black-Box ML Models? A Zeroth-Order Optimization Perspective.” In International Conference on Learning Representations. https://openreview.net/forum?id=W9G_ImpHlQd.
Zhu, Jingyi, Long Wang, and James C Spall. 2019. “Efficient Implementation of Second-Order Stochastic Approximation Algorithms in High-Dimensional Problems.” IEEE Transactions on Neural Networks and Learning Systems 31 (8): 3087–99.
Zhu, X., and J. C. Spall. 2002. “A Modified Second-Order SPSA Optimization Algorithm for Finite Samples.” Int. J. Adapt. Control Signal Process. 16: 397–409.