Немонотонное поведение и тяжелые хвосты в методах первого порядка тема диссертации и автореферата по ВАК РФ 00.00.00, кандидат наук Данилова Марина Юрьевна
- Специальность ВАК РФ00.00.00
- Количество страниц 216
Оглавление диссертации кандидат наук Данилова Марина Юрьевна
Contents
List of Figures vi
List of Tables ix
1 Introduction
1.1 Deterministic First-Order Methods
1.2 Stochastic First-Order Methods
1.3 Aim of the work
1.4 Proposition for The Defense
1.5 Scientific Novelty
1.6 Theoretical and Practical Importance
1.7 Presentations and Validation of Research Results
1.8 Publications
1.9 Personal Contribution
1.10 Thesis Structure
2 Non-monotone Behavior of the Heavy Ball Method
2.1 Introduction
2.2 Peak effect for second-order difference equations
2.3 Analysis of the Heavy Ball method
2.3.1 The Heavy Ball method
2.3.2 Convergence analysis
2.3.3 Peak effect
2.4 Lyapunov function for the Heavy ball method
2.4.1 Construction of the Lyapunov function
2.4.2 Global convergence
2.4.3 Adaptive algorithm
2.5 Conclusion
3 Averaged Heavy-Ball Method
3.1 Introduction
3.1.1 Preliminaries
3.1.2 Related work
3.2 Maximal Deviations on Quadratic Problems
3.2.1 Heavy-Ball Method
3.2.2 Averaged Heavy-Ball method
3.2.3 Maximal Deviation of AHB for Arbitrary Initialization
3.3 Convergence Guarantees for Non-Quadratics
3.3.1 Weighted Averaged Heavy-Ball Method
3.3.2 Restarted Averaged Heavy-Ball Method
3.4 Numerical Experiments
3.4.1 Quadratic Functions
3.4.2 Logistic Regression with ^-Regularization
3.5 Conclusion
4 On the Convergence Analysis of Aggregated Heavy-Ball Method
4.1 Introduction
4.1.1 Motivational Example
4.1.2 Our Contributions
4.1.3 Technical Preliminaries
4.1.4 Related Work
4.2 Analysis of Aggregated Heavy-Ball Method
4.2.1 Non-Convex Case
4.2.2 Convex and Strongly-Convex Cases
4.3 Numerical Experiments
4.4 Conclusion
5 Stochastic Optimization with Heavy-Tailed Noise via Accelerated Gradient Clipping
5.1 Introduction
5.1.1 Preliminaries
5.1.2 Simple Motivational Example: Convergence in Expectation and Clipping
5.1.3 Related Work
5.1.4 Our Contributions
5.1.5 Chapter Organization
5.2 Accelerated SGD with Clipping
5.2.1 Convex Case
5.2.2 Strongly Convex Case
5.3 SGD with Clipping
5.3.1 Convex Case
5.3.2 Strongly Convex Case
5.4 Numerical Experiments
5.5 Discussion
5.6 Extra Experiments
5.6.1 Detailed Description of Experiments from Section
5.6.2 Additional Details for Experiments with Logistic Regression
6 Near-Optimal High Probability Complexity Bounds for Non-Smooth Stochas-
tic Convex Optimization with Heavy-Tailed Noise
6.1 Introduction
6.1.1 Preliminaries
6.1.2 Contributions
6.1.3 Related Work
6.1.4 Chapter Organization
6.2 Clipped Stochastic Similar Triangles Method
6.3 SGD with Clipping
6.4 Clipped Similar Triangles Method: Missing Details and Proofs
6.4.1 Convergence in the convex case
6.4.2 Convergence in the Strongly Convex Case
6.5 SGD with Clipping: Missing Details and Proofs
6.5.1 Convex Case
6.5.2 Strongly Convex Case
6.6 Numerical experiments
6.7 Additional experimental details
6.7.1 Main experiment hyper-parameters
6.7.2 On the relation between stepsize parameter a and batchsize
6.7.3 Evolution of the noise distribution
References
A Basic Facts, Technical Lemmas, and Auxiliary Results
A.1 Basic Inequalities
A.2 Identities and Inequalities Involving Random Variables
A.3 Auxiliary Results
A.3.1 Bernstein Inequality
A.3.2 About the Sum of i.i.d. Random Variables with Heavy Tails
A.3.3 (<,MU)-Hölder continuous gradient
A.4 Technical Results
Рекомендованный список диссертаций по специальности «Другие cпециальности», 00.00.00 шифр ВАК
Децентрализованная оптимизация на меняющихся со временем сетях / Decentralized optimization over time-varying networks2023 год, кандидат наук Рогозин Александр Викторович
Безградиентные методы выпуклой оптимизации в условиях шума / Gradient-Free Methods for Convex Optimization under Noise Conditions2025 год, кандидат наук Лобанов Александр Владимирович
Безградиентные методы решения седловых задач и не только / Gradient-Free Methods for Saddle-Point Problems and Beyond2023 год, кандидат наук Безносиков Александр Николаевич
Численные методы решения негладких задач выпуклой оптимизации с функциональными ограничениями / Numerical Methods for Non-Smooth Convex Optimization Problems with Functional Constraints2020 год, кандидат наук Алкуса Мохаммад
Оптимизация функционалов предыскажения сигнала по типу Виннера-Гаммерштейна, для устранения интермодуляционных компонент, возникающих при усилении мощности2024 год, кандидат наук Масловский Александр Юрьевич
Список литературы диссертационного исследования кандидат наук Данилова Марина Юрьевна, 2022 год
References
[1] Ildar Radikovich Abdrakhmanov, Evgenii Alekseevich Kanin, Sergei Andreevich Boronin, Evgeny Vladimirovich Burnaev, and Andrei Aleksandrovich Osiptsov. Development of deep transformer-based models for long-term prediction of transient production of oil wells. In SPE Russian Petroleum Technology Conference. OnePetro, 2021.
[2] Amir Beck and Marc Teboulle. A fast iterative shrinkage-thresholding algorithm for linear inverse problems. SIAM journal on imaging sciences, 2(1):183-202, 2009.
[3] George Bennett. Probability inequalities for the sum of independent random variables. Journal of the American Statistical Association, 57(297):33-45, 1962.
[4] Aleksandr Alekseevich Borovkov and Konstantin Aleksandrovich Borovkov. On probabilities of large deviations for random walks. i. regularly varying distribution tails. Theory of Probability & Its Applications, 46(2):193-213, 2002.
[5] Irving W Burr. Cumulative frequency functions. The Annals of mathematical statistics, 13(2):215-232, 1942.
[6] Chih-Chung Chang and Chih-Jen Lin. Libsvm: a library for support vector machines. ACM transactions on intelligent systems and technology (TIST), 2(3):1-27, 2011.
[7] Caroline Chaux, Patrick L Combettes, Jean-Christophe Pesquet, and Valérie R Wajs. A variational formulation for frame-based inverse problems. Inverse Problems, 23(4):1495-1518, jun 2007.
[8] Marina Danilova. On the convergence analysis of aggregated heavy-ball method. In Panos Pardalos, Michael Khachay, and Vladimir Mazalov, editors, Mathematical Optimization Theory and Operations Research, pages 3-17, Cham, 2022. Springer International Publishing.
[9] Marina Danilova, Pavel Dvurechensky, Alexander Gasnikov, Eduard Gorbunov, Sergey Guminov, Dmitry Kamzolov, and Innokentiy Shibaev. Recent theoretical advances in non-convex optimization. arXiv preprint arXiv:2012.06188, 2020.
[10] Marina Danilova, Anastasiia Kulakova, and Boris Polyak. Non-monotone behavior of the heavy ball method. In International Conference on Difference Equations and Applications, pages 213-230. Springer, 2018.
[11] MY Danilova, GS Malinovskiy, et al. Averaged heavy-ball method. Computer research and modeling, 14(2):277-308, 2022.
[12] Damek Davis and Dmitriy Drusvyatskiy. Stochastic model-based minimization of weakly convex functions. SIAM Journal on Optimization, 29(1):207-239, 2019.
[13] Damek Davis, Dmitriy Drusvyatskiy, Lin Xiao, and Junyu Zhang. From low probability to high confidence in stochastic convex optimization. arXiv preprint arXiv:1907.13307, 2019.
[14] Damek Davis, Dmitriy Drusvyatskiy, Lin Xiao, and Junyu Zhang. From low probability to high confidence in stochastic convex optimization. Journal of Machine Learning Research, 22(49):1-38, 2021.
[15] Aaron Defazio. Momentum via primal averaging: Theoretical insights and learning rate schedules for non-convex optimization. arXiv preprint arXiv:2010.00406, 2020.
[16] Aaron Defazio. Understanding the role of momentum in non-convex optimization: Practical insights from a lyapunov analysis. arXiv preprint arXiv:2010.00406, 2020.
[17] Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018.
[18] Olivier Devolder. Exactness, inexactness and stochasticity in first-order methods for large-scale convex optimization. PhD thesis, PhD thesis, 2013.
[19] Olivier Devolder et al. Stochastic first order methods in smooth convex optimization. Technical report, CORE, 2011.
[20] Olivier Devolder, François Glineur, and Yurii Nesterov. First-order methods of smooth convex optimization with inexact oracle. Mathematical Programming, 146(1):37-75, 2014.
[21] Pavel Dvurechenskii, Darina Dvinskikh, Alexander Gasnikov, Cesar Uribe, and Angelia Nedich. Decentralize and randomize: Faster algorithm for wasserstein barycenters. In Advances in Neural Information Processing Systems, pages 10760-10770, 2018.
[22] Pavel Dvurechensky and Alexander Gasnikov. Stochastic intermediate gradient method for convex problems with stochastic inexact oracle. Journal of Optimization Theory and Applications, 171(1):121-145, 2016.
[23] Kacha Dzhaparidze and JH Van Zanten. On bernstein-type inequalities for martingales. Stochastic processes and their applications, 93(1):109-117, 2001.
[24] Saber Elaydi. An Introduction to Difference Equations. Springer, 2005.
[25] O Bousquet A Elisseeff and Olivier Bousquet. Stability and generalization. Journal of Machine Learning Research, 2:499-526, 2002.
[26] David A Freedman et al. On tail probabilities for martingales. the Annals of Probability, 3(1):100-118, 1975.
[27] Alexander Gasnikov. Universal gradient descent. arXiv preprint arXiv:1711.00394, 2017.
[28] Alexander Gasnikov, Pavel Dvurechensky, and Yurii Nesterov. Stochastic gradient methods with inexact oracle. arXiv preprint arXiv:1411.4218, 2014.
[29] Alexander Gasnikov and Yurii Nesterov. Universal fast gradient method for stochastic composit optimization problems. arXiv:1604.05275, 2016.
[30] Alexander Vladimirovich Gasnikov, Yu E Nesterov, and Vladimir Grigor'evich Spokoiny. On the efficiency of a randomized mirror descent algorithm in online optimization problems. Computational Mathematics and Mathematical Physics, 55(4):580-596, 2015.
[31] Jonas Gehring, Michael Auli, David Grangier, Denis Yarats, and Yann N Dauphin. Convolutional sequence to sequence learning. In Proceedings of the 34th International Conference on Machine Learning-Volume 70, pages 1243-1252. JMLR. org, 2017.
[32] Euhanna Ghadimi, Hamid Reza Feyzmahdavian, and Mikael Johansson. Global convergence of the heavy-ball method for convex optimization. In 2015 European control conference (ECC), pages 310-315. IEEE, 2015.
[33] Saeed Ghadimi and Guanghui Lan. Optimal stochastic approximation algorithms for strongly convex stochastic composite optimization i: A generic algorithmic framework. SIAM Journal on Optimization, 22(4):1469-1492, 2012.
[34] Saeed Ghadimi and Guanghui Lan. Optimal stochastic approximation algorithms for strongly convex stochastic composite optimization, ii: shrinking procedures and optimal algorithms. SIAM Journal on Optimization, 23(4):2061-2089, 2013.
[35] Saeed Ghadimi and Guanghui Lan. Stochastic first-and zeroth-order methods for nonconvex stochastic programming. SIAM Journal on Optimization, 23(4):2341-2368, 2013.
[36] Pontus Giselsson and Stephen Boyd. Monotonicity and restart in fast gradient methods. In 53rd IEEE Conference on Decision and Control, pages 5058-5063. IEEE, 2014.
[37] Ian Goodfellow, Yoshua Bengio, and Aaron Courville. Deep Learning. MIT Press, 2016. http://www.deeplearningbook.org.
[38] Eduard Gorbunov, Adel Bibi, Ozan Sener, El Houcine Bergou, and Peter Richtarik. A stochastic derivative free optimization method with momentum. In International Conference on Learning Representations, 2020.
[39] Eduard Gorbunov, Marina Danilova, and Alexander Gasnikov. Stochastic optimization with heavy-tailed noise via accelerated gradient clipping. In H. Larochelle, M. Ranzato, R. Hadsell, M. F. Balcan, and H. Lin, editors, Advances in Neural Information Processing Systems, volume 33, pages 15042-15053. Curran Associates, Inc., 2020.
[40] Eduard Gorbunov, Marina Danilova, Innokentiy Shibaev, Pavel Dvurechensky, and Alexander Gasnikov. Near-optimal high probability complexity bounds for non-smooth stochastic optimization with heavy-tailed noise. arXiv preprint arXiv:2106.05958, 2021.
[41] Eduard Gorbunov, Darina Dvinskikh, and Alexander Gasnikov. Optimal decentralized distributed algorithms for stochastic convex optimization. arXiv preprint arXiv:1911.07363, 2019.
[42] Eduard Gorbunov, Pavel Dvurechensky, and Alexander Gasnikov. An accelerated method
for derivative-free smooth stochastic convex optimization. arXiv preprint arXiv:1802.09022, 2018.
[43] Eduard Gorbunov, Filip Hanzely, and Peter Richtárik. A unified theory of sgd: Variance reduction, sampling, quantization and coordinate descent. arXiv preprint arXiv:1905.11261, 2019.
[44] Robert Mansel Gower, Nicolas Loizou, Xun Qian, Alibek Sailanbayev, Egor Shulgin, and Peter Richtárik. Sgd: General analysis and improved rates. In International Conference on Machine Learning, pages 5200-5209, 2019.
[45] Vincent Guigues, Anatoli Juditsky, and Arkadi Nemirovski. Non-asymptotic confidence bounds for the optimal value of a stochastic program. Optimization Methods and Software, 32(5):1033-1058, 2017.
[46] Julia Gusak, Daria Cherniuk, Alena Shilova, Alexander Katrutsa, Daniel Bershatsky, Xunyi Zhao, Lionel Eyraud-Dubois, Oleg Shlyazhko, Denis Dimitrov, Ivan Oseledets, et al. Survey on large scale neural network training. arXiv preprint arXiv:2202.104 35, 2022.
[47] Cristóbal Guzmán and Arkadi Nemirovski. On lower complexity bounds for large-scale smooth convex optimization. Journal of Complexity, 31(1):1-14, 2015.
[48] Moritz Hardt, Benjamin Recht, and Yoram Singer. Train faster, generalize better: Stability of stochastic gradient descent. arXiv preprint arXiv:1509.01240, 2015.
[49] Elad Hazan and Satyen Kale. Beyond the regret minimization barrier: optimal algorithms for stochastic strongly-convex optimization. The Journal of Machine Learning Research, 15(1):2489-2512, 2014.
[50] Elad Hazan, Kfir Levy, and Shai Shalev-Shwartz. Beyond convexity: Stochastic quasi-convex optimization. In Advances in Neural Information Processing Systems, pages 1594-1602, 2015.
[51] Magnus R Hestenes and Eduard Stiefel. Methods of conjugate gradients for solving. Journal of research of the National Bureau of Standards, 49(6):409, 1952.
[52] Daniel Hsu and Sivan Sabato. Loss minimization and parameter estimation with heavy tails. The Journal of Machine Learning Research, 17(1):543-582, 2016.
[53] Bin Hu and Laurent Lessard. Dissipativity theory for nesterov's accelerated method. In International Conference on Machine Learning, pages 1549-1557. PMLR, 2017.
[54] Chi Jin, Praneeth Netrapalli, Rong Ge, Sham M Kakade, and Michael I Jordan. A short note on concentration inequalities for random vectors with subgaussian norm. arXiv preprint arXiv:1902.03736, 2019.
[55] Anatoli Juditsky, Arkadi Nemirovski, et al. First order methods for nonsmooth convex
large-scale optimization, i: general purpose methods. Optimization for Machine Learning, pages 121-148, 2011.
[56] Anatoli Juditsky and Yuri Nesterov. Deterministic and stochastic primal-dual subgradient algorithms for uniformly convex minimization. Stochastic Systems, 4(1):44-80, 2014.
[57] Sham M Kakade and Ambuj Tewari. On the generalization ability of online strongly convex programming algorithms. In Advances in Neural Information Processing Systems, pages 801-808, 2009.
[58] Ahmed Khaled and Peter Richtárik. Better theory for sgd in the nonconvex world. arXiv preprint arXiv:2002.03329, 2020.
[59] Rahul Kidambi, Praneeth Netrapalli, Prateek Jain, and Sham Kakade. On the insufficiency of existing momentum schemes for stochastic optimization. In 2018 Information Theory and Applications Workshop (ITA), pages 1-9. IEEE, 2018.
[60] Donghwan Kim and Jeffrey A Fessler. Adaptive restart of the optimized gradient method for convex optimization. Journal of Optimization Theory and Applications, 178(1):240-263, 2018.
[61] Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014.
[62] Alex Krizhevsky, Vinod Nair, and Geoffrey Hinton. Cifar-10 and cifar-100 datasets. URl: https://www. cs. toronto. edu/kriz/cifar. html, 6, 2009.
[63] Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. Imagenet classification with deep convolutional neural networks. In Advances in neural information processing systems, pages 1097-1105, 2012.
[64] Guanghui Lan. An optimal method for stochastic composite optimization. Mathematical Programming, 133(1-2):365-397, 2012.
[65] Laurent Lessard, Benjamin Recht, and Andrew Packard. Analysis and design of optimization algorithms via integral quadratic constraints. SIAM Journal on Optimization, 26(1):57-95, 2016.
[66] Kfir Y Levy. The power of normalization: Faster evasion of saddle points. arXiv preprint arXiv:1611.04831, 2016.
[67] Nicolas Loizou and Peter Richtárik. Momentum and stochastic momentum for stochastic gradient, newton, proximal point and subspace descent methods. Computational Optimization and Applications, 77(3):653-710, 2020.
[68] James Lucas, Shengyang Sun, Richard Zemel, and Roger Grosse. Aggregated momentum: Stability through passive damping. In International Conference on Learning Representations, 2019.
[69] Vien V Mai and Mikael Johansson. Stability and convergence of stochastic gradient clipping: Beyond lipschitz continuity and smoothness. arXiv preprint arXiv:2102.06489, 2021.
[70] Horia Mania, Xinghao Pan, Dimitris Papailiopoulos, Benjamin Recht, Kannan Ram-chandran, and Michael I Jordan. Perturbed iterate analysis for asynchronous stochastic optimization. SIAM Journal on Optimization, 27(4):2202-2229, 2017.
[71] Bernard Martinet. Régularisation d'inéquations variationnelles par approximations successives. rev. française informat. Recherche Opérationnelle, 4:154-158, 1970.
[72] Bernard Martinet. Détermination approchée d'un point fixe d'une application pseudocontractante. CR Acad. Sci. Paris, 274(2):163-165, 1972.
[73] Michael P McLaughlin. A compendium of common probability distributions. Michael P. McLaughlin, 2001.
[74] Aditya Krishna Menon, Ankit Singh Rawat, Sashank J Reddi, and Sanjiv Kumar. Can gradient clipping mitigate label noise? In International Conference on Learning Representations, 2020.
[75] Stephen Merity, Nitish Shirish Keskar, and Richard Socher. Regularizing and optimizing lstm language models. arXiv preprint arXiv:1708.02182, 2017.
[76] Tomás Mikolov. Statistical language models based on neural networks. Presentation at Google, Mountain View, 2nd April, 80, 2012.
[77] Konstantin Mishchenko, Eduard Gorbunov, Martin Takác, and Peter Richtárik. Distributed learning with compressed gradient differences. arXiv preprint arXiv:1901.09269, 2019.
[78] Hesameddin Mohammadi, Samantha Samuelson, and Mihailo R Jovanovic. Transient growth of accelerated first-order methods for strongly convex optimization problems. arXiv preprint arXiv:2103.08017, 2021.
[79] Eric Moulines and Francis R Bach. Non-asymptotic analysis of stochastic approximation algorithms for machine learning. In Advances in Neural Information Processing Systems, pages 451-459, 2011.
[80] Aleksandr Viktorovich Nazin, AS Nemirovsky, Aleksandr Borisovich Tsybakov, and AB Ju-ditsky. Algorithms of robust stochastic optimization based on mirror descent method. Automation and Remote Control, 80(9):1607-1627, 2019.
[81] Deanna Needell, Nathan Srebro, and Rachel Ward. Stochastic gradient descent, weighted sampling, and the randomized kaczmarz algorithm. Mathematical Programming, 155(1-2):549-573, 2016.
[82] Arkadi Nemirovski, Anatoli Juditsky, Guanghui Lan, and Alexander Shapiro. Robust
stochastic approximation approach to stochastic programming. SIAM Journal on optimization, 19(4):1574-1609, 2009.
[83] Arkadi S Nemirovski and David Berkovich Yudin. Cesari convergence of the gradient method of approximating saddle points of convex-concave functions. In Doklady Akademii Nauk, volume 239, pages 1056-1059. Russian Academy of Sciences, 1978.
[84] A.S. Nemirovsky and D.B. Yudin. Problem Complexity and Method Efficiency in Optimization. J. Wiley & Sons, New York, 1983.
[85] Yu Nesterov. Smooth minimization of non-smooth functions. Mathematical programming, 103(1):127-152, 2005.
[86] Yu Nesterov. Universal gradient methods for convex optimization problems. Mathematical Programming, 152(1):381-404, 2015.
[87] Yu Nesterov and J-Ph Vial. Confidence level solutions for stochastic programming. Automatical, 44(6):1559-1568, 2008.
[88] Yurii Nesterov. A method for unconstrained convex minimization problem with the rate of convergence O(1/k2). In Doklady an ussr, volume 269, pages 543-547, 1983.
[89] Yurii Nesterov. Introductory lectures on convex optimization: A basic course, volume 87. Springer Science & Business Media, 2003.
[90] Yurii Nesterov. Lectures on convex optimization, volume 137. Springer, 2018.
[91] Lam Nguyen, Phuong Ha Nguyen, Marten Dijk, Peter Richtarik, Katya Scheinberg, and Martin Takac. Sgd and hogwild! convergence without the bounded gradients assumption. In International Conference on Machine Learning, pages 3750-3758, 2018.
[92] Mykola Novik. torch-optimizer - collection of optimization algorithms for PyTorch. github repository, 2020.
[93] Brendan O'donoghue and Emmanuel Candes. Adaptive restart for accelerated gradient schemes. Foundations of computational mathematics, 15(3):715-732, 2015.
[94] Razvan Pascanu, Tomas Mikolov, and Yoshua Bengio. On the difficulty of training recurrent neural networks. In International conference on machine learning, pages 1310-1318, 2013.
[95] Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. Pytorch: An imperative style, high-performance deep learning library. In H. Wallach, H. Larochelle, A. Beygelzimer, F. d'Alché-Buc, E. Fox, and R. Garnett, editors, Advances in Neural Information Processing Systems 32, pages 8024-8035. Curran Associates, Inc., 2019.
[96] Matthew E Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. Deep contextualized word representations. arXiv preprint arXiv:1802.05365, 2018.
[97] IB Petrov and AI Lobanov. Lectures in computational mathematics. M: The Internet University of Information Technology, 2006.
[98] Boris Polyak. Introduction to Optimization. New York, Optimization Software, 1987.
[99] Boris Polyak and Pavel Shcherbakov. Lyapunov functions: An optimization theory perspective. IFAC-PapersOnLine, 50(1):7456-7461, 2017.
[100] Boris T Polyak. Gradient methods for the minimisation of functionals. USSR Computational Mathematics and Mathematical Physics, 3(4):864-878, 1963.
[101] Boris T Polyak. Some methods of speeding up the convergence of iteration methods. USSR Computational Mathematics and Mathematical Physics, 4(5):1-17, 1964.
[102] Boris T Polyak, Pavel S Shcherbakov, and Georgi Smirnov. Peak effects in stable linear difference equations. Journal of Difference Equations and Applications, 24(9):1488-1502, 2018.
[103] Alexander Rakhlin, Ohad Shamir, and Karthik Sridharan. Making gradient descent optimal for strongly convex stochastic optimization. arXiv preprint arXiv:1109.5647, 2011.
[104] Peter Richtarik, Igor Sokolov, and Ilyas Fatkhullin. Ef21: A new, simpler, theoretically better, and practically faster error feedback. arXiv preprint arXiv:2106.05203, 2021.
[105] Herbert Robbins and Sutton Monro. A stochastic approximation method. The annals of mathematical statistics, pages 400-407, 1951.
[106] R Tyrrell Rockafellar. Monotone operators and the proximal point algorithm. SIAM journal on control and optimization, 14(5):877-898, 1976.
[107] Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei. ImageNet Large Scale Visual Recognition Challenge. International Journal of Computer Vision (IJCV), 115(3):211-252, 2015.
[108] Shai Shalev-Shwartz and Shai Ben-David. Understanding machine learning: From theory to algorithms. Cambridge university press, 2014.
[109] Shai Shalev-Shwartz, Ohad Shamir, Nathan Srebro, and Karthik Sridharan. Stochastic convex optimization. In COLT, 2009.
[110] Shai Shalev-Shwartz, Yoram Singer, Nathan Srebro, and Andrew Cotter. Pegasos: Primal estimated sub-gradient solver for svm. Mathematical programming, 127(1):3-30, 2011.
[111] Alexander Shapiro, Darinka Dentcheva, and Andrzej Ruszczynski. Lectures on stochastic programming: modeling and theory. SIAM, 2014.
[112] DN Shiyan and AV Kolnogorov. Simulation of the mirror descent algorithm on distributions with different variances. In Journal of Physics: Conference Series, volume 1658, page 012051. IOP Publishing, 2020.
[113] Umut §im§ekli, Mert Gurbuzbalaban, Thanh Huy Nguyen, Gaël Richard, and Levent Sagun. On the heavy-tailed theory of stochastic gradient descent for deep neural networks. arXiv preprint arXiv:1912.00018, 2019.
[114] Umut Simsekli, Levent Sagun, and Mert Gurbuzbalaban. A tail-index analysis of stochastic gradient noise in deep neural networks. arXiv preprint arXiv:1901.06053, 2019.
[115] Vladimir Spokoiny et al. Parametric estimation. finite sample theory. The Annals of Statistics, 40(6):2877-2909, 2012.
[116] Ilya Sutskever, James Martens, George Dahl, and Geoffrey Hinton. On the importance of initialization and momentum in deep learning. In International conference on machine learning, pages 1139-1147. PMLR, 2013.
[117] Adrien Taylor and Francis Bach. Stochastic first-order methods: non-asymptotic and computer-aided analyses via potential functions. In Conference on Learning Theory, pages 2934-2992. PMLR, 2019.
[118] Adrien Taylor, Bryan Van Scoy, and Laurent Lessard. Lyapunov functions for first-order methods: Tight automated convergence guarantees. In International Conference on Machine Learning, pages 4897-4906. PMLR, 2018.
[119] Adrien B Taylor, Julien M Hendrickx, and François Glineur. Performance estimation toolbox (pesto): automated worst-case analysis of first-order optimization methods. In 2017 IEEE 56th Annual Conference on Decision and Control (CDC), pages 1278-1283. IEEE, 2017.
[120] Ilnura Usmanova. Robust solutions to stochastic optimization problems. Master Thesis (MSIAM); Institut Polytechnique de Grenoble ENSIMAG, Laboratoire Jean Kuntzmann, 2017.
[121] Bryan Van Scoy, Randy A Freeman, and Kevin M Lynch. The fastest known globally convergent first-order method for minimizing strongly convex functions. IEEE Control Systems Letters, 2(1):49-54, 2017.
[122] Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, £ukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Advances in neural information processing systems, pages 5998-6008, 2017.
[123] Alex Warstadt, Amanpreet Singh, and Samuel R Bowman. Neural network acceptability judgments. arXiv preprint arXiv:1805.12471, 2018.
[124] Waloddi Weibull. A statistical distribution function of wide applicability. Journal of Applied Mechanics, 18:293-297, 1951.
[125] Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander M. Rush. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 38-45, Online, October 2020. Association for Computational Linguistics.
[126] Tianbao Yang, Qihang Lin, and Zhe Li. Unified convergence analysis of stochastic momentum methods for convex and non-convex optimization. arXiv preprint arXiv:1604.03257, 2016.
[127] Hao Yu, Rong Jin, and Sen Yang. On the linear speedup analysis of communication efficient momentum sgd for distributed non-convex optimization. In International Conference on Machine Learning, pages 7184-7193. PMLR, 2019.
[128] Jingzhao Zhang, Tianxing He, Suvrit Sra, and Ali Jadbabaie. Why gradient clipping accelerates training: A theoretical justification for adaptivity. In International Conference on Learning Representations, 2020.
[129] Jingzhao Zhang, Tianxing He, Suvrit Sra, and Ali Jadbabaie. Why gradient clipping accelerates training: A theoretical justification for adaptivity. In International Conference on Learning Representations, 2020.
[130] Jingzhao Zhang, Sai Praneeth Karimireddy, Andreas Veit, Seungyeon Kim, Sashank Reddi, Sanjiv Kumar, and Suvrit Sra. Why are adaptive methods good for attention models? In H. Larochelle, M. Ranzato, R. Hadsell, M. F. Balcan, and H. Lin, editors, Advances in Neural Information Processing Systems, volume 33, pages 15383-15393. Curran Associates, Inc., 2020.
[131] Jingzhao Zhang, Sai Praneeth Karimireddy, Andreas Veit, Seungyeon Kim, Sashank J Reddi, Sanjiv Kumar, and Suvrit Sra. Why adam beats sgd for attention models. arXiv preprint arXiv:1912.03194, 2019.
Обратите внимание, представленные выше научные тексты размещены для ознакомления и получены посредством распознавания оригинальных текстов диссертаций (OCR). В связи с чем, в них могут содержаться ошибки, связанные с несовершенством алгоритмов распознавания. В PDF файлах диссертаций и авторефератов, которые мы доставляем, подобных ошибок нет.