References
Abdulla, Mohammed Shahid, and Shalabh Bhatnagar. 2007.
“Reinforcement Learning Based Algorithms for Average Cost Markov
Decision Processes.” Discrete Event Dynamic Systems 17
(1): 23–52.
Abounadi, Jinane, Dimitri P Bertsekas, and Vivek Borkar. 2002.
“Stochastic Approximation for Nonexpansive Maps: Application to
q-Learning Algorithms.” SIAM Journal on Control and
Optimization 41 (1): 1–22.
Acerbi, C. 2002. “Spectral Measures of Risk: A Coherent
Representation of Subjective Risk Aversion.” Journal of
Banking & Finance 26 (7): 1505–18.
A.Dvoretzky. 1956. “On Stochastic Approximation.” Proc.
Third Berkeley Symp. Math. Stat. And Prob. 1: 39–55.
Agarwal, Alekh, Ofer Dekel, and Lin Xiao. 2010. “Optimal
Algorithms for Online Convex Optimization with Multi-Point Bandit
Feedback.” In COLT, 28–40.
Alzantot, Moustafa, Yash Sharma, Supriyo Chakraborty, Huan Zhang,
Cho-Jui Hsieh, and Mani B. Srivastava. 2019. “GenAttack: Practical
Black-Box Attacks with Gradient-Free Optimization.” In
Proceedings of the Genetic and Evolutionary Computation
Conference, 1111–19. GECCO ’19. New York, NY, USA: Association for
Computing Machinery. https://doi.org/10.1145/3321707.3321749.
Anandkumar, Animashree, and Rong Ge. 2016. “Efficient Approaches
for Escaping Higher Order Saddle Points in Non-Convex
Optimization.” In Conference on Learning Theory, 81–102.
PMLR.
Andronov, Aleksandr Aleksandrovich, Aleksandr Adol’fovich Vitt, and
Semen Emmanuilovich Khaikin. 2013. Theory of Oscillators: Adiwes
International Series in Physics. Vol. 4. Elsevier.
Arnold, Vladimir I. 1992. Ordinary Differential Equations.
Springer Science & Business Media.
Artzner, P., F. Delbaen, J. Eber, and D. Heath. 1999. “Coherent measures of risk.”
Mathematical Finance 9 (3): 203–28.
Asmussen, S., and P. W. Glynn. 2007. Stochastic Simulation:
Algorithms and Analysis. Springer.
Asmussen, Søren, and Peter W Glynn. 2007. Stochastic Simulation:
Algorithms and Analysis. Vol. 57. Springer.
Aubin, J., and A. Cellina. 1984. Differential Inclusions: Set-Valued
Maps and Viability Theory. Springer.
Aubin, J., and H. Frankowska. 1990. Set-Valued Analysis.
Birkhauser.
Balasubramanian, K., and S. Ghadimi. 2022b. “Zeroth-Order
Nonconvex Stochastic Optimization: Handling Constraints, High
Dimensionality, and Saddle Points.” Foundations of
Computational Mathematics 22 (1): 35–76.
———. 2022a. “Zeroth-Order Nonconvex Stochastic Optimization:
Handling Constraints, High Dimensionality, and Saddle Points.”
Foundations of Computational Mathematics 22 (1): 35–76.
Barakat, A., P. Bianchi, W. Hachem, and S. Schechtman. 2021.
“Stochastic optimization with momentum:
Convergence, fluctuations, and traps avoidance.”
Electronic Journal of Statistics 15 (2): 3892–3947. https://doi.org/10.1214/21-EJS1880.
Bardou, O., N. Frikha, and G. Pages. 2009. “Computing VaR and CVaR using
stochastic approximation and adaptive unconstrained importance
sampling.” Monte Carlo Methods and Applications
15 (3): 173–210.
Benaïm, M. 1996. “A Dynamical System Approach to Stochastic
Approximations.” SIAM J. Control Optim. 34 (2): 437–72.
———. 1999. “Dynamics of Stochastic Approximation
Algorithms.” Seminaire De Probabilities (Strasbourg)
1709: 1–68.
Benaïm, M., and M. W. Hirsch. 1996. “Asymptotic Pseudotrajectories
and Chain Recurrent Flows, with Applications.” J. Dynam.
Differential Equations 8: 141–76.
Benaïm, M., J. Hofbauer, and S. Sorin. 2005. “Stochastic
Approximations and Differential Inclusions.” SIAM Journal on
Control and Optimization, 328–48.
———. 2012. “Perturbations of Set-Valued Dynamical Systems, with
Applications to Game Theory.” Dynamic Games and
Applications 2 (2): 195–205.
Berahas, Albert S., Liyuan Cao, Krzysztof Choromanski, and Katya
Scheinberg. 2022. “A Theoretical and Empirical Comparison of
Gradient Approximations in Derivative-Free Optimization.”
Foundations of Computational Mathematics 22 (2): 507–60.
Bertsekas, D. P. 2012. Dynamic Programming and Optimal Control,
Vol.II. Athena Scientific.
Bertsekas, D. P., and J. N. Tsitsiklis. 1996. Neuro-Dynamic
Programming. Athena Scientific.
Bertsekas, Dimitri. 2019. Reinforcement Learning and Optimal
Control. Vol. 1. Athena Scientific.
Bertsekas, Dimitri P. 1999. Nonlinear Programming. 2nd ed.
Belmont, MA: Athena Scientific.
Bertsekas, DP, and JN Tsitsiklis. 1989. Parallel and Distributed
Computation. Prentice Hall Inc.
Bhagoji, Arjun Nitin, Warren He, Bo Li, and Dawn Song. 2018.
“Practical Black-Box Attacks on Deep Neural Networks Using
Efficient Query Mechanisms.” In Computer Vision – ECCV
2018, edited by Vittorio Ferrari, Martial Hebert, Cristian
Sminchisescu, and Yair Weiss, 158–74. Cham: Springer International
Publishing.
Bhatnagar, S. 2005. “Adaptive Multivariate Three-Timescale
Stochastic Approximation Algorithms for Simulation Based
Optimization.” ACM Transactions on Modeling and Computer
Simulation 15 (1): 74–107.
———. 2007. “Adaptive Newton-Based Smoothed Functional
Algorithms for Simulation Optimization.” ACM Transactions on
Modeling and Computer Simulation 18 (1): 2:1–35.
Bhatnagar, Shalabh. 2010. “An Actor–Critic Algorithm with Function
Approximation for Discounted Cost Constrained Markov Decision
Processes.” Systems & Control Letters 59 (12):
760–66.
———. 2023. “The Reinforce Policy Gradient Algorithm
Revisited.” arXiv Preprint arXiv:2310.05000.
Bhatnagar, Shalabh, and Mohammed Shahid Abdulla. 2008.
“Simulation-Based Optimization Algorithms for Finite-Horizon
Markov Decision Processes.” Simulation 84 (12): 577–600.
Bhatnagar, Shalabh, and K Mohan Babu. 2008. “New Algorithms of the
q-Learning Type.” Automatica 44 (4): 1111–19.
Bhatnagar, Shalabh, and Vivek S Borkar. 1998. “A Two Timescale
Stochastic Approximation Scheme for Simulation-Based Parametric
Optimization.” Probability in the Engineering and
Informational Sciences 12 (4): 519–31.
———. 2003. “Multiscale Chaotic SPSA and Smoothed Functional
Algorithms for Simulation Optimization.” Simulation 79
(10): 568–80.
Bhatnagar, Shalabh, Vivek S Borkar, Madhukar Akarapu, and Shie Mannor.
2006. “A Simulation-Based Algorithm for Ergodic Control of Markov
Chains Conditioned on Rare Events.” Journal of Machine
Learning Research 7 (10).
Bhatnagar, Shalabh, Michael C Fu, Steven I Marcus, I Wang, et al. 2003.
“Two-Timescale Simultaneous Perturbation Stochastic Approximation
Using Deterministic Perturbation Sequences.” ACM Transactions
on Modeling and Computer Simulation 13 (2): 180–209.
Bhatnagar, Shalabh, N Hemachandra, and Vivek Kumar Mishra. 2011.
“Stochastic Approximation Algorithms for Constrained Optimization
via Simulation.” ACM Transactions on Modeling and Computer
Simulation (TOMACS) 21 (3): 1–22.
Bhatnagar, Shalabh, and Shishir Kumar. 2004. “A Simultaneous
Perturbation Stochastic Approximation-Based Actor-Critic Algorithm for
Markov Decision Processes.” IEEE Transactions on Automatic
Control 49 (4): 592–98.
Bhatnagar, Shalabh, and K Lakshmanan. 2012. “An Online
Actor–Critic Algorithm with Function Approximation for Constrained
Markov Decision Processes.” Journal of Optimization Theory
and Applications 153: 688–708.
———. 2016. “Multiscale q-Learning with Linear Function
Approximation.” Discrete Event Dynamic Systems 26:
477–509.
Bhatnagar, Shalabh, Vivek Kumar Mishra, and Nandyala Hemachandra. 2011.
“Stochastic Algorithms for Discrete Parameter Simulation
Optimization.” IEEE Transactions on Automation Science and
Engineering 8 (4): 780–93.
Bhatnagar, Shalabh, and L. A. Prashanth. 2023. “Generalized
Simultaneous Perturbation Stochastic Approximation with Reduced
Estimator Bias.” In 2023 57th Annual Conference on
Information Sciences and Systems (CISS), 1–6. IEEE.
Bhatnagar, S, H. L. Prasad, and L. A. Prashanth. 2013. Stochastic
Recursive Algorithms for Optimization: Simultaneous Perturbation Methods
(Lecture Notes in Control and Information Sciences). Vol. 434.
Springer.
Bhatnagar, S., and L. A. Prashanth. 2015. “Simultaneous
Perturbation Newton Algorithms for Simulation Optimization.”
Journal of Optimization Theory and Applications 164 (2):
621–43.
Bhatnagar, S., R. S. Sutton, M. Ghavamzadeh, and M. Lee. 2009.
“Natural Actor-Critic Algorithms.” Automatica 45
(11): 2471–82.
Bhavsar, N., and L. A. Prashanth. 2022. “Non-Asymptotic Bounds for
Stochastic Optimization with Biased Noisy Gradient Oracles.”
IEEE Transactions on Automatic Control, 1–1. https://doi.org/10.1109/TAC.2022.3159748.
Bhojanapalli, Srinadh, Behnam Neyshabur, and Nati Srebro. 2016.
“Global Optimality of Local Search for Low Rank Matrix
Recovery.” Advances in Neural Information Processing
Systems 29.
Billingsley, Patrick. 2013. Convergence of Probability
Measures. John Wiley & Sons.
———. 2017. Probability and Measure. John Wiley & Sons.
Borkar, V. S. 1995. Probability Theory: An Advanced Course. New
York: Springer.
———. 2022. Stochastic Approximation: A Dynamical Systems Viewpoint,
2’nd Edition. Cambridge University Press.
Borkar, V. S., and S. P. Meyn. 1999. “The O.D.E.
Method for Convergence of Stochastic Approximation and Reinforcement
Learning.” SIAM J. Control Optim 38: 447–69.
———. 2000. “The O.D.E. Method for Convergence of
Stochastic Approximation and Reinforcement Learning.” SIAM
Journal of Control and Optimization 38 (2): 447–69.
Borkar, Vivek S. 2003. “Avoidance of Traps in Stochastic
Approximation.” Systems & Control Letters 50 (1):
1–9.
Bottou, L., F. E. Curtis, and J. Nocedal. 2018. “Optimization
Methods for Large-Scale Machine Learning.” SIAM Review
60 (2): 223–311.
Boyd, Stephen, and Lieven Vandenberghe. 2004. Convex
Optimization. Cambridge university press.
Brandiere, Odile, and Marie Duflo. 1996. “Les Algorithmes
Stochastiques Contournent-Ils Les Pièges?” In
Annales de l’IHP Probabilités Et Statistiques,
32:395–427. 3.
Bunch, James R, and Beresford N Parlett. 1971. “Direct Methods for
Solving Symmetric Indefinite Systems of Linear Equations.”
SIAM Journal on Numerical Analysis 8 (4): 639–55.
Cai, HanQin, Daniel McKenzie, Wotao Yin, and Zhenliang Zhang. 2022.
“Zeroth-Order Regularized Optimization (Zoro): Approximately
Sparse Gradients and Adaptive Sampling.” SIAM Journal on
Optimization 32 (2): 687–714.
Carmon, Yair, John Duchi, Oliver Hinder, and Aaron Sidford. 2016.
“Accelerated Methods for Non-Convex Optimization.”
arXiv Preprint arXiv:1611.00756.
Cassandras, Christos G, and Stéphane Lafortune. 2008. Introduction
to Discrete Event Systems. Springer.
Chen, H. F., L. Guo, and A. J. Gao. 1987. “Convergence and robustness of the Robbins-Monro algorithm
truncated at randomly varying bounds.” Stochastic
Processes and Their Applications 27: 217–31.
Chen, Han-Fu, Tyrone E Duncan, and Bozenna Pasik-Duncan. 1999. “A
Kiefer-Wolfowitz Algorithm with Randomized Differences.” IEEE
Transactions on Automatic Control 44 (3): 442–53.
Chen, Jianbo, Michael I Jordan, and Martin J Wainwright. 2020.
“HopSkipJumpAttack: A Query-Efficient Decision-Based
Attack.” In IEEE Symposium on Security and Privacy
(SP), 1277–94. IEEE.
Chen, Pin-Yu, Huan Zhang, Yash Sharma, Jinfeng Yi, and Cho-Jui Hsieh.
2017. “ZOO: Zeroth Order Optimization Based Black-Box Attacks to
Deep Neural Networks Without Training Substitute Models.” In
Proceedings of the 10th ACM Workshop on Artificial Intelligence and
Security, 15–26. AISec ’17. New York, NY, USA: Association for
Computing Machinery. https://doi.org/10.1145/3128572.3140448.
Chen, Shuhang, Adithya Devraj, Ana Busic, and Sean Meyn. 2020.
“Explicit Mean-Square Error Bounds for Monte-Carlo and Linear
Stochastic Approximation.” In International Conference on
Artificial Intelligence and Statistics, 4173–83. PMLR.
Chin, D. C. 1997. “Comparative Study of Stochastic Algorithms for
System Optimization Based on Gradient Approximations.” IEEE
Transactions on Systems, Man, and Cybernetics, Part B: Cybernetics
27 (2): 244–49.
Choromanski, Krzysztof, Mark Rowland, Vikas Sindhwani, Richard Turner,
and Adrian Weller. 2018. “Structured Evolution with Compact
Architectures for Scalable Policy Optimization.” In
Proceedings of the 35th International Conference on Machine
Learning, edited by Jennifer Dy and Andreas Krause, 80:970–78.
Proceedings of Machine Learning Research. PMLR.
Coddington, Earl A, Norman Levinson, and T Teichmann. 1956.
“Theory of Ordinary Differential Equations.” American
Institute of Physics.
Cover, Thomas M, and Joy A Thomas. 2012. Elements of Information
Theory. John Wiley & Sons.
Dalal, G., B. Szorenyi, G. Thoppe, and S. Mannor. 2018. “Finite
Sample Analysis of Two-Timescale Stochastic Approximation with
Applications to Reinforcement Learning.” In Conference on
Learning Theory, 1–35.
Dippon, Jürgen. 2003. “Accelerated Randomized Stochastic
Optimization.” The Annals of Statistics 31 (4): 1260–81.
Dong, Yinpeng, Hang Su, Jun Zhu Wu, Ziwei Zhang, and Jiansheng Liu.
2020. “Improving Black-Box Adversarial Attacks with a
Transfer-Based Prior.” In Advances in Neural Information
Processing Systems (NeurIPS), 17631–41.
Duchi, John C., Peter L. Bartlett, and Martin J. Wainwright. 2012.
“Randomized Smoothing for Stochastic Optimization.”
SIAM Journal on Optimization 22 (2): 674–701.
Dunkel, J., and S. Weber. 2010. “Stochastic Root Finding and
Efficient Estimation of Convex Risk Measures.” Operations
Research 58 (5): 1505–21.
Durrett, Rick. 2019. Probability: Theory and Examples. Vol. 49.
Cambridge university press.
Erdogdu, M. A. 2016. “Newton-Stein method: An
optimization method for glms via stein’s lemma.”
Journal of Machine Learning Research 17 (215): 1–52.
Fabian, V. 1968. “On Asymptotic Normality in Stochastic
Approximation.” The Annals of Mathematical Statistics,
1327–32.
———. 1971. “Stochastic Approximation.” In Optimizing
Methods in Statistics (Ed. J.j.rustagi), 439–70. New York: Academic
Press.
Filippov, Aleksei Fedorovich. 2013. Differential Equations with
Discontinuous Righthand Sides: Control Systems. Vol. 18. Springer
Science & Business Media.
Flaxman, Abraham D, Adam Tauman Kalai, and H Brendan McMahan. 2005.
“Online Convex Optimization in the Bandit Setting: Gradient
Descent Without a Gradient.” In SODA, 385–94.
Föllmer, H., and A. Schied. 2002. “Convex Measures of Risk and
Trading Constraints.” Finance and Stochastics 6 (4):
429–47.
Frikha, N., and S. Menozzi. 2012. “Concentration Bounds for Stochastic
Approximations.” Electronic Communications in
Probability 17: no. 47, 1–15.
Fu, M. C., ed. 2015. Handbook of Simulation Optimization.
Springer.
Furmston, T., G. Lever, and D. Barber. 2016. “Approximate
Newton Methods for Approximate Policy Search in
Markov Decision Processes.” Journal of Machine
Learning Research 17: 1–51.
Gadat, S., and I. Gavra. 2022. “Asymptotic Study of Stochastic
Adaptive Algorithms in Non-Convex Landscape.” Journal of
Machine Learning Research 23 (228): 1–54.
Gallager, Robert G. 2013. Stochastic Processes: Theory for
Applications. Cambridge University Press.
Gasnikov, Alexander, Anton Novitskii, Vasilii Novitskii, Farshed
Abdukhakimov, Dmitry Kamzolov, Aleksandr Beznosikov, Martin Takac, Pavel
Dvurechensky, and Bin Gu. 2022. “The Power of First-Order Smooth
Optimization for Black-Box Non-Smooth Problems.” In
Proceedings of the 39th International Conference on Machine
Learning, edited by Kamalika Chaudhuri, Stefanie Jegelka, Le Song,
Csaba Szepesvari, Gang Niu, and Sivan Sabato, 162:7241–65. Proceedings
of Machine Learning Research. PMLR.
Ge, R., F. Huang, C. Jin, and Y. Yuan. 2015. “Escaping from Saddle
Points – Online Stochastic Gradient for Tensor Decomposition.”
Conference of Learning Theory.
Ge, Rong, Chi Jin, and Yi Zheng. 2017. “No Spurious Local Minima
in Nonconvex Low Rank Problems: A Unified Geometric Analysis.” In
International Conference on Machine Learning, 1233–42. PMLR.
Ge, Rong, Jason D Lee, and Tengyu Ma. 2016. “Matrix Completion Has
No Spurious Local Minimum.” Advances in Neural Information
Processing Systems 29.
Gelfand, Saul B, and Sanjoy K Mitter. 1991. “Recursive Stochastic
Algorithms for Global Optimization in r^d.” SIAM Journal on
Control and Optimization 29 (5): 999–1018.
Gerencser, Laszlo, Stacy D Hill, and Zsuzsanna Vago. 1999.
“Optimization over Discrete Sets via SPSA.” In
Proceedings of the 31st Conference on Winter Simulation:
Simulation—a Bridge to the Future-Volume 1, 466–70.
Ghadimi, S., and G. Lan. 2013. “Stochastic First-and Zeroth-Order
Methods for Nonconvex Stochastic Programming.” SIAM Journal
on Optimization 23 (4): 2341–68.
Ghoshdastidar, D., A. Dukkipati, and S. Bhatnagar. 2014a.
“Newton-Based Stochastic Optimization Using q-Gaussian Smoothed Functional Algorithms.”
Automatica 50 (10): 2606–14.
———. 2014b. “Smoothed Functional Algorithms for Stochastic
Optimization Using q-Gaussian Distributions.” ACM
Transactions on Modeling and Computer Simulation 26 (3): 17:1–26.
Giesecke, K., T. Schmidt, and S. Weber. 2008. “Measuring the Risk
of Large Losses.” Journal of Investment Management, Fourth
Quarter.
Grimmett, Geoffrey, and David Stirzaker. 2020. Probability and
Random Processes. Oxford university press.
Hegde, V., A. S. Menon, L. A. Prashanth, and K. Jagannathan. 2021.
“Online Estimation and Optimization of
Utility-Based Shortfall Risk.” Papers 2111.08805.
arXiv.org.
Hirsch, Morris W, Stephen Smale, and Robert L Devaney. 2013.
Differential Equations, Dynamical Systems, and an Introduction to
Chaos. Academic press.
Hu, X., L. A. Prashanth, A. György, and C. Szepesvári. 2016.
“(Bandit) Convex Optimization with Biased
Noisy Gradient Oracles.” In Artificial Intelligence
and Statistics, 819–28.
Huang, Feihu, Lue Tao, and Songcan Chen. 2020. “Accelerated
Stochastic Gradient-Free and Projection-Free Methods.” In
Proceedings of the 37th International Conference on Machine
Learning, edited by Hal Daumé III and Aarti Singh, 119:4519–30.
Proceedings of Machine Learning Research. PMLR.
Hurley, M. 1995. “Chain Recurrence, Semiflows, and
Gradients.” Journal of Dynamics and Differential
Equations, 437–56.
Ilyas, Andrew, Logan Engstrom, Anish Athalye, and Jessy Lin. 2018.
“Black-Box Adversarial Attacks with Limited Queries and
Information.” In Proceedings of the 35th International
Conference on Machine Learning, edited by Jennifer Dy and Andreas
Krause, 80:2137–46. Proceedings of Machine Learning Research. PMLR.
Ilyas, Andrew, Logan Engstrom, and Aleksander Madry. 2019. “Prior
Convictions: Black-Box Adversarial Attacks with Bandits and
Priors.” In 7th International Conference on Learning
Representations (ICLR).
Ince, Edward L. 1956. Ordinary Differential Equations. Courier
Corporation.
Jain, P., D. Nagaraj, and P. Netrapalli. 2021. “Making the Last
Iterate of SGD Information Theoretically Optimal.” SIAM
Journal on Optimization 31 (2): 1108–30.
Ji, Kaiyi, Zhe Wang, Yi Zhou, and Yingbin Liang. 2019. “Improved
Zeroth-Order Variance Reduced Algorithms and Analysis for Nonconvex
Optimization.” In Proceedings of the 36th International
Conference on Machine Learning, edited by Kamalika Chaudhuri and
Ruslan Salakhutdinov, 97:3100–3109. Proceedings of Machine Learning
Research. PMLR.
Jin, C., R. Ge, P. Netrapalli, S. M. Kakade, and M. I. Jordan. 2017.
“How to Escape Saddle Points Efficiently.” ICML,
1724–32.
Jin, C., P. Netrapalli, R. Ge, S. M. Kakade, and M. I Jordan. 2021.
“On Nonconvex Optimization for Machine Learning: Gradients,
Stochasticity, and Saddle Points.” Journal of the ACM
(JACM) 68 (2): 1–29.
Karandikar, Rajeeva Laxman, and Mathukumalli Vidyasagar. 2024.
“Convergence Rates for Stochastic Approximation: Biased Noise with
Unbounded Variance, and Applications.” Journal of
Optimization Theory and Applications, 1–39.
Karmakar, Prasenjit, and Shalabh Bhatnagar. 2018. “Two Time-Scale
Stochastic Approximation with Controlled Markov Noise and Off-Policy
Temporal-Difference Learning.” Mathematics of Operations
Research 43 (1): 130–51.
———. 2021. “Stochastic Approximation with Iterate-Dependent Markov
Noise Under Verifiable Conditions in Compact State Space with the
Stability of Iterates Not Ensured.” IEEE Transactions on
Automatic Control 66 (12): 5941–54.
Katkovnik, V. Ya, and Yu Kulchitsky. 1972. “Convergence of a Class
of Random Search Algorithms.” Automation Remote Control
8: 1321–26.
Kawaguchi, Kenji. 2016. “Deep Learning Without Poor Local
Minima.” Advances in Neural Information Processing
Systems 29.
Kiefer, J., and J. Wolfowitz. 1952. “Stochastic Estimation of the
Maximum of a Regression Function.” Ann. Math. Statist.
23: 462–66.
Kirkpatrick, Scott, C Daniel Gelatt Jr, and Mario P Vecchi. 1983.
“Optimization by Simulated Annealing.” Science 220
(4598): 671–80.
Konda, Vijay R, and John N Tsitsiklis. 2003. “On Actor-Critic
Algorithms.” SIAM Journal on Control and Optimization 42
(4): 1143–66.
Kornowski, Guy, and Ohad Shamir. 2024. “An Algorithm with Optimal
Dimension Dependence for Zero-Order Nonsmooth Nonconvex Stochastic
Optimization.” Journal of Machine Learning Research 25
(122): 1–14.
Kozak, David, Cesare Molinari, Lorenzo Rosasco, Luis Tenorio, and Silvia
Villa. 2023. “Zeroth-Order Optimization with Orthogonal Random
Directions.” Mathematical Programming 199 (1): 1179–219.
Kulkarni, Vidyadhar G. 2016. Modeling and Analysis of Stochastic
Systems. Chapman; Hall/CRC.
Kushner, H. J., and D. S. Clark. 1978. Stochastic Approximation
Methods for Constrained and Unconstrained Systems. New York:
Springer Verlag.
Kushner, H. J., and G. G. Yin. 2003. Stochastic Approximation
Algorithms and Applications, 2’nd Ed. New York: Springer Verlag.
Lasalle, J. P., and S. Lefschetz. 1961. Stability by Liapunov’s
Direct Method with Applications. New York: Academic Press.
Lei, J. 2020. “Convergence and concentration
of empirical measures under Wasserstein distance in unbounded functional
spaces.” Bernoulli 26 (1): 767–98.
Levin, David A, and Yuval Peres. 2017. Markov Chains and Mixing
Times. Vol. 107. American Mathematical Soc.
Ljung, L. 1977. “Analysis of Recursive Stochastic
Algorithms.” Automatic Control, IEEE Transactions on 22
(4): 551–75.
Malladi, Sadhika, Tianyu Gao, Eshaan Nichani, Alex Damian, Jason D Lee,
Danqi Chen, and Sanjeev Arora. 2023. “Fine-Tuning Language Models
with Just Forward Passes.” Advances in Neural Information
Processing Systems 36: 53038–75.
Mania, Horia, Aurelia Guy, and Benjamin Recht. 2018. “Simple
Random Search of Static Linear Policies Is Competitive for Reinforcement
Learning.” Advances in Neural Information Processing
Systems 31.
Maniyar, Mizhaan P., L. A. Prashanth, Akash Mondal, and Shalabh
Bhatnagar. 2024. “A Cubic-Regularized Policy Newton
Algorithm for Reinforcement Learning.” In Proceedings of the
27th International Conference on Artificial Intelligence and
Statistics, edited by Sanjoy Dasgupta, Stephan Mandt, and Yingzhen
Li, 238:4708–16. Proceedings of Machine Learning Research. PMLR.
Maryak, John L, and Daniel C Chin. 2001. “Global Random
Optimization by Simultaneous Perturbation Stochastic
Approximation.” In Proceedings of the 2001 American Control
Conference.(cat. No. 01CH37148), 2:756–62. IEEE.
Meyn, Sean. 2022. Control Systems and Reinforcement Learning.
Cambridge University Press.
Meyn, Sean P, and Richard L Tweedie. 2012. Markov Chains and
Stochastic Stability. Springer Science & Business Media.
Mondal, A., L. A. Prashanth, and S. Bhatnagar. 2024. “Truncated
Cauchy Random Perturbations for Smoothed Functional-Based Stochastic
Optimization.” Automatica 162: 111528.
https://doi.org/https://doi.org/10.1016/j.automatica.2024.111528.
Mou, Wenlong, Chris Junchi Li, Martin J Wainwright, Peter L Bartlett,
and Michael I Jordan. 2020. “On Linear Stochastic Approximation:
Fine-Grained Polyak-Ruppert and Non-Asymptotic Concentration.” In
Conference on Learning Theory, 2947–97. PMLR.
Mukhoty, Bhaskar, Velibor Bojkovic, William de Vazelhes, Xiaohan Zhao,
Giulia De Masi, Huan Xiong, and Bin Gu. 2023. “Direct Training of
SNN Using Local Zeroth Order Method.” In Advances in Neural
Information Processing Systems (NeurIPS).
Nandakumaran, A. K, P. S Datti, and R. K George. 2017. Ordinary
Differential Equations: Principles and Applications. Cambridge
University Press.
Nemirovski, Arkadi, Anatoli Juditsky, Guanghui Lan, and Alexander
Shapiro. 2009. “Robust Stochastic Approximation Approach to
Stochastic Programming.” SIAM Journal on Optimization 19
(4): 1574–1609.
Nesterov, Y., and B. T. Polyak. 2007. “Cubic Regularization of
Newton Method and Its Global Performance.” Mathematical
Programming 112: 159–81.
Nesterov, Y., and V. Spokoiny. 2017. “Random Gradient-Free
Minimization of Convex Functions.” Foundations of
Computational Mathematics 17 (2): 527–66.
Nesterov, Yurii, and Boris Polyak. 2006. “Cubic regularization of Newton method and its global
performance.” Math. Program. 108 (August):
177–205.
Nocedal, Jorge, and Stephen J Wright. 1999. Numerical
Optimization. Springer.
Norris, James R. 1998. Markov Chains. 2. Cambridge university
press.
Pachal, Soumen, Shalabh Bhatnagar, and L. A. Prashanth. 2023.
“Generalized Simultaneous Perturbation-Based Gradient Search with
Reduced Estimator Bias.” arXiv Preprint
arXiv:2212.10477.
Pemantle, R. 1990. “Non-Convergence to Unstable Points in Urn
Models and Stochastic Approximations.” The Annals of
Probability 18(2): 698–712.
Polyak, B. T., and A. B. Tsybakov. 1990. “Optimal Orders of
Accuracy for Search Algorithms of Stochastic Optimization.”
Problems in Information Transmission, 126–33.
Polyak, Boris T, and Anatoli B Juditsky. 1992. “Acceleration of
Stochastic Approximation by Averaging.” SIAM Journal on
Control and Optimization 30 (4): 838–55.
Powell, Warren B. 2021. “Reinforcement Learning and Stochastic
Optimization.” John Wiley & Sons Hoboken, NJ.
Prashanth, L. A., and S. P. Bhat. 2022. “A
Wasserstein Distance Approach for Concentration of Empirical Risk
Estimates.” Journal of Machine Learning Research
23 (238): 1–61.
Prashanth, L. A., and S. Bhatnagar. 2012. “Threshold Tuning Using Stochastic Optimization for Graded
Signal Control.” IEEE Transactions on Vehicular
Technology 61 (9): 3865–80.
Prashanth, L. A., S. Bhatnagar, N. Bhavsar, M. Fu, and S. I. Marcus.
2020. “Random Directions Stochastic Approximation with
Deterministic Perturbations.” IEEE Transactions on Automatic
Control 65 (6): 2450–65.
Prashanth, L. A., Shalabh Bhatnagar, Michael C. Fu, and Steve I Marcus.
2017. “Adaptive System Optimization Using Random Directions
Stochastic Approximation.” IEEE Transactions on Automatic
Control 62 (5): 2223–38.
Prashanth, L. A., A. Chatterjee, and S. Bhatnagar. 2014. “Two
Timescale Convergent q-Learning for Sleep-Scheduling in Wireless Sensor
Networks.” Wireless Networks 20: 2589–2604.
Prashanth, L. A., and Mohammad Ghavamzadeh. 2016.
“Variance-Constrained Actor-Critic Algorithms for Discounted and
Average Reward MDPs.” Machine Learning 105: 367–417.
Prashanth, L. A., N. Korda, and R. Munos. 2021. “Concentration
Bounds for Temporal Difference Learning with Linear Function
Approximation: The Case of Batch Data and Uniform Sampling.”
Mach. Learn. 110 (3): 559–618.
Ramaswamy, A., and S. Bhatnagar. 2016. “A Generalization of the
Borkar-Meyn Theorem for Stochastic Recursive
Inclusions.” Mathematics of Operations Research 42 (3):
648–61.
———. 2018. “Analysis of Gradient Descent Methods with
Nondiminishing Bounded Errors.” IEEE Transactions on
Automatic Control 63 (5): 1465–71.
———. 2019. “Stability of Stochastic Approximations with Controlled
Markov Noise and Temporal Difference Learning.”
IEEE Transactions on Automatic Control 64 (6): 2614–20.
———. 2021. “Analyzing Approximate Value Iteration
Algorithms.” Mathematics of Operations Research. https://doi.org/10.1287/moor.2021.1202.
Ramaswamy, Arunselvan, and Shalabh Bhatnagar. 2016. “Stochastic
Recursive Inclusion in Two Timescales with an Application to the
Lagrangian Dual Problem.” Stochastics 88 (8): 1173–87.
Rando, Marco, Cesare Molinari, Lorenzo Rosasco, and Silvia Villa. 2023.
“An Optimal Structured Zeroth-Order Algorithm for Non-Smooth
Optimization.” In Advances in Neural Information Processing
Systems, edited by A. Oh, T. Naumann, A. Globerson, K. Saenko, M.
Hardt, and S. Levine, 36:36738–67. Curran Associates, Inc.
Rando, Marco, Cesare Molinari, Silvia Villa, and Lorenzo Rosasco. 2024.
“Stochastic Zeroth Order Descent with Structured
Directions.” Computational Optimization and
Applications, October.
Rastogi, Pushpendre, Jingyi Zhu, and James C Spall. 2016.
“Efficient Implementation of Enhanced Adaptive Simultaneous
Perturbation Algorithms.” In 2016 Annual Conference on
Information Science and Systems (CISS), 298–303. IEEE.
R.Blum, J. 1954. “Approximation Methods Which Converge with
Probability One.” Annals of Mathematical Statistics,
382–86.
Robbins, H., and S. Monro. 1951. “A Stochastic Approximation
Method.” Ann. Math. Statist. 22: 400–407.
Rockafellar, R. T., and S. Uryasev. 2000. “Optimization of conditional value-at-risk.”
Journal of Risk 2: 21–42.
Royden, H. L., and P. M. Fitzpatrick. 2010. Real Analysis. 4th
ed. Boston: Pearson.
Rubinstein, R. Y. 1981. Simulation and the Monte Carlo Method.
New York: Wiley.
Ruppert, D. 1985. “A Newton-Raphson Version of the
Multivariate Robbins-Monro Procedure.” Annals of
Statistics 13: 236–45.
Saha, Ankan, and Ambuj Tewari. 2011. “Improved Regret Guarantees
for Online Smooth Convex Optimization with Bandit Feedback.” In
AISTATS, 636–42.
Salimans, Tim, Jonathan Ho, Xi Chen, Szymon Sidor, and Ilya Sutskever.
2017. “Evolution Strategies as a Scalable Alternative to
Reinforcement Learning.” arXiv Preprint
arXiv:1703.03864.
Serfling, Robert J. 2009. Approximation Theorems of Mathematical
Statistics. Vol. 162. John Wiley & Sons.
Shamir, Ohad. 2017. “An Optimal Algorithm for Bandit and
Zero-Order Convex Optimization with Two-Point Feedback.”
Journal of Machine Learning Research 18 (52): 1–11.
Spall, J. C. 2000. “Adaptive Stochastic Approximation by the
Simultaneous Perturbation Method.” IEEE Trans. Autom.
Contr. 45: 1839–53.
———. 2005. Introduction to Stochastic Search
and Optimization: Estimation, Simulation, and Control. Vol.
65. John Wiley & Sons.
Spall, James C. 1992. “Multivariate Stochastic Approximation Using
a Simultaneous Perturbation Gradient Approximation.” IEEE
Transactions on Automatic Control 37 (3): 332–41.
———. 1997. “A One-Measurement Form of Simultaneous Perturbation
Stochastic Approximation.” Automatica 33 (1): 109–12.
———. 2009. “Feedback and Weighting Mechanisms for Improving
Jacobian Estimates in the Adaptive Simultaneous Perturbation
Algorithm.” IEEE Transactions on Automatic Control 54
(6): 1216–29.
Stein, C. 1972. “A Bound for the Error in the Normal Approximation
to the Distribution of a Sum of Dependent Random Variables.” In
Proceedings of the Sixth Berkeley Symposium on Mathematical
Statistics and Probability, Volume 2: Probability Theory,
6:583–603. University of California Press.
———. 1981. “Estimation of the Mean of a Multivariate Normal
Distribution.” The Annals of Statistics, 1135–51.
Styblinski, M. A., and T.-S. Tang. 1990. “Experiments in Nonconvex
Optimization: Stochastic Approximation with Function Smoothing and
Simulated Annealing.” Neural Networks 3: 467–83.
Sun, Ju, Qing Qu, and John Wright. 2016. “Complete Dictionary
Recovery over the Sphere i: Overview and the Geometric Picture.”
IEEE Transactions on Information Theory 63 (2): 853–84.
Sutton, R. S., and A. W. Barto. 2018. Reinforcement Learning, 2’nd
Edition. MIT Press.
Sutton, Richard S, Hamid Reza Maei, Doina Precup, Shalabh Bhatnagar,
David Silver, Csaba Szepesvári, and Eric Wiewiora. 2009. “Fast
Gradient-Descent Methods for Temporal-Difference Learning with Linear
Function Approximation.” In Proceedings of the 26th Annual
International Conference on Machine Learning, 993–1000.
Sutton, Richard S, David McAllester, Satinder Singh, and Yishay Mansour.
1999. “Policy Gradient Methods for Reinforcement Learning with
Function Approximation.” Advances in Neural Information
Processing Systems 12.
Swain, J. J. 2017. “Simulation Software Survey-Simulation: New and
Improved Reality Show.” OR/MS Today 44 (5): 38–49.
Tamar, Aviv, Yinlam Chow, Mohammad Ghavamzadeh, and Shie Mannor. 2015.
“Policy Gradient for Coherent Risk Measures.” In
Advances in Neural Information Processing Systems. Vol. 28.
Curran Associates, Inc.
Tripuraneni, N., M. Stern, C. Jin, J. Regier, and M. I. Jordan. 2018.
“Stochastic Cubic Regularization for Fast Nonconvex
Optimization.” In NeurIPS. Vol. 31. Curran Associates,
Inc.
Tropp, J. A. 2016. “The Expected Norm of a
Sum of Independent Random Matrices: An Elementary
Approach.” In High Dimensional Probability VII: The
Cargèse Volume, 173–202.
Tsitsiklis, J. N., and B. Van Roy. 1997. “An Analysis of
Temporal-Difference Learning with Function Approximation.”
IEEE Transactions on Automatic Control 42 (5):
674–90.
Tsitsiklis, John N. 1994. “Asynchronous Stochastic Approximation
and q-Learning.” Machine Learning 16 (3): 185–202.
Vijayan, Nithia, and L. A. Prashanth. 2021. “Smoothed
Functional-Based Gradient Algorithms for Off-Policy Reinforcement
Learning: A Non-Asymptotic Viewpoint.” Systems & Control
Letters 155: 104988.
———. 2023. “A Policy Gradient Approach for Optimization of Smooth
Risk Measures.” In Proceedings of the Thirty-Ninth Conference
on Uncertainty in Artificial Intelligence, edited by Robin J. Evans
and Ilya Shpitser, 216:2168–78. Proceedings of Machine Learning
Research. PMLR.
Wang, Tianyu, and Yasong Feng. 2024. “Convergence Rates of Zeroth
Order Gradient Descent for Lojasiewicz Functions.”
INFORMS Journal on Computing.
Williams, Ronald J. 1992. “Simple Statistical Gradient-Following
Algorithms for Connectionist Reinforcement Learning.” Machine
Learning 8: 229–56.
Wright, S. J., and B. Recht. 2022. Optimization for Data
Analysis. Cambridge University Press.
Yaji, V. G., and S. Bhatnagar. 2019. “Analysis of Stochastic
Approximation Schemes with Set-Valued Maps in the Absence of a Stability
Guarantee and Their Stabilization.” IEEE Transactions on
Automatic Control 65 (3): 1100–1115.
Yaji, Vinayaka G, and Shalabh Bhatnagar. 2018. “Stochastic
Recursive Inclusions with Non-Additive Iterate-Dependent Markov
Noise.” Stochastics 90 (3): 330–63.
———. 2020. “Stochastic Recursive Inclusions in Two Timescales with
Nonadditive Iterate-Dependent Markov Noise.” Mathematics of
Operations Research 45 (4): 1405–44.
Yao, A. C. C. 1977. “Probabilistic Computations: Toward a Unified
Measure of Complexity.” In FOCS, 222–27.
Zhang, Yimeng, Yuguang Yao, Jinghan Jia, Jinfeng Yi, Mingyi Hong, Shiyu
Chang, and Sijia Liu. 2022. “How to Robustify
Black-Box ML Models? A Zeroth-Order Optimization
Perspective.” In International Conference on Learning
Representations. https://openreview.net/forum?id=W9G_ImpHlQd.
Zhu, Jingyi, Long Wang, and James C Spall. 2019. “Efficient
Implementation of Second-Order Stochastic Approximation Algorithms in
High-Dimensional Problems.” IEEE Transactions on Neural
Networks and Learning Systems 31 (8): 3087–99.
Zhu, X., and J. C. Spall. 2002. “A Modified Second-Order
SPSA Optimization Algorithm for Finite Samples.”
Int. J. Adapt. Control Signal Process. 16: 397–409.