Latest book

Gradient-based algorithms for zeroth-order optimization

Foundations and Trends in Optimization, Vol. 8: No. 1–3, pp 1-332, 2025

Research Interests

Reinforcement Learning, Zeroth-order Optimization, Multi-armed Bandits

News

May-2026: Serving as an area chair for NeurIPS-2026.
Feb-2026: Serving as an area chair for ICML-2026 and RLC-2026.
Jan-2026: Teaching a course on topics in RL. For details, click here.
Nov-2025: A paper entitled Policy Newton methods for Distortion Riskmetrics accepted for publication in AAAI (2026).
Jul-2025: Back at IITM after visiting C-MinDS at IITB from Aug-2024 to Jul-2025.
May-2025: A book entitled ‘Gradient-based algorithms for zeroth-order optimization’ published. See book page for the details.
May-2025: A paper entitled ‘Finite Time Analysis of Temporal Difference Learning for Mean-Variance in a Discounted MDP’ accepted for publication in Reinforcement Learning Conference (RLC). Click here for the arxiv report.
Mar-2025: Invited talk on ‘Distorted bandits or How I learned to be risk-seeking without regretting it’ at National Conference on Communications (NCC-2025) held at IIT Delhi. Click here for the slides.
Jan-2025: Invited talk on ‘Reinforcement Learning and Bandit Algorithms for Distortion Riskmetrics’ at Reinforcement learning workshop held at IISc, Bengaluru. Click here for the video.
Jan-2025: A paper entitled Generalized Simultaneous Perturbation-based Gradient Search with Reduced Estimator Bias accepted for publication in IEEE Transactions on Automatic Control and another paper entitled ‘‘Risk-sensitive Bandits: Arm Mixture Optimality and Regret-efficient Algorithms’’ accepted for publication in AISTATS.
Jan-2025: Teaching a course on stochastic optimization. For details, click here.
Jun-2024: A paper entitled Optimization of utility-based shortfall risk: A non-asymptotic viewpoint accepted to IEEE Conference on Decision and Control (CDC).
Jun-2024: A paper entitled Online Estimation and Optimization of Utility-Based Shortfall Risk accepted for publication in Mathematics of Operations Research.
May-2024: Two papers accepted to ICML, see here.
Feb-2024: Invited talk on ‘A cubic-regularized policy Newton for reinforcement learning’ at Reinforcement learning workshop held at IISc, Bengaluru. Click here for the video.
Jan-2024: A paper entitled A Cubic-regularized Policy Newton Algorithm for Reinforcement Learning accepted for publication in AISTATS.
Aug-2023: Teaching a course on operating systems. For details, click here.
Jul-2023: Invited talk on ‘Finite time analysis of temporal difference learning with linear function approximation’ at Data science: Probabilistic and optimization methods held at International Centre for Theoretical Sciences, Bengaluru. Click here for the video.
Jan-2023: A paper entitled A policy gradient approach for optimization of smooth risk measures accepted for publication in UAI.
Feb-2023: Invited talk on ‘Finite time analysis of temporal difference learning with linear function approximation: Tail averaging and regularisation’ at Networks Seminar Series held (in-person) at Indian Institute of Science. Click here for the video.
Feb-2023: Tutorial on risk-sensitive reinforcement learning at AAAI-2023. Click here for details.
Jan-2023: A paper entitled Finite time analysis of temporal difference learning with linear function approximation: Tail averaging and regularisation accepted for publication in AISTATS.
Jan-2023: Teaching a course on stochastic optimization. For details, click here.
Jan-2023: Invited talk on ‘A Wasserstein distance approach for concentration of empirical risk estimates’ at Information Theory and Data Science Workshop held (in-person) at National University of Singapore.
Aug-2022: A paper entitled A Wasserstein distance approach for concentration of empirical risk estimates accepted for publication in Journal of Machine Learning Research.
Jul-2022: Teaching a course on programming and data structures. For details, click here.
Jul-2022: Tutorial on Risk-Aware Multi-armed Bandits at SPCOM 2022. Slides here.
Jun-2022: A monograph entitled Risk-Sensitive Reinforcement Learning via Policy Gradient Search published by Foundations and Trends in Machine Learning.
Apr-2022: A survey article entitled A Survey of Risk-Aware Multi-Armed Bandits accepted at IJCAI-2022.
Feb-2022: Invited talk on ‘Concentration of risk measures: A Wasserstein distance approach’ at ‘IITB Workshop on Stochastic Models’.
Jan-2022: Teaching a course on object oriented analysis using C++. For details, click here.

Prospective interns

I do not have open positions and you are encouraged to apply directly for the IITM summer fellowship programme. Details here