Adam: A Method for Stochastic Optimization
Adam: метод стохастической оптимизации
2014-12-22
SCID: 54.1/x68mz84h
Discuss with AI
AdaMaxAdam optimization algorithmadaptive moment estimationonline convex optimizationstochastic optimization
Figures from the paper
Abstract (AI)
We introduce Adam, an algorithm for first-order gradient-based optimization of stochastic objective functions, based on adaptive estimates of lower-order moments. The method is straightforward to implement, is computationally efficient, has little memory requirements, is invariant to diagonal rescaling of the gradients, and is well suited for problems that are large in terms of data and/or parameters. The method is also appropriate for non-stationary objectives and problems with very noisy and/or sparse gradients. The hyper-parameters have intuitive interpretations and typically require little tuning. Some connections to related algorithms, on which Adam was inspired, are discussed. We also analyze the theoretical convergence properties of the algorithm and provide a regret bound on the convergence rate that is comparable to the best known results under the online convex optimization framework. Empirical results demonstrate that Adam works well in practice and compares favorably to other stochastic optimization methods. Finally, we discuss AdaMax, a variant of Adam based on the infinity norm.
Key Findings
1
Adam is designed to handle non-stationary objectives, noisy gradients, and sparse gradients, while requiring relatively little hyperparameter tuning.
2
Adam is introduced as a first-order stochastic optimization algorithm using adaptive estimates of lower-order gradient moments.
3
Empirical results show that Adam performs well and compares favorably with other stochastic optimization methods; the paper also introduces the AdaMax variant based on the infinity norm.
4
The method is computationally efficient, memory-efficient, invariant to diagonal gradient rescaling, and suitable for large-scale problems.
5
Theoretical analysis provides a regret-bound convergence rate comparable to the best known results for online convex optimization.
Research Object
Adam stochastic optimization algorithm
Research Subject
the algorithm’s adaptive moment estimation, convergence properties, computational performance, and practical effectiveness for stochastic objective functions
Publication Details
Publication Date
2014-12-22
Journal
Publisher
ISSN
Cited by
84624
Open access PDF
Access Type
Author Information
Download PDF
Subscribe to digest
References available in scid.ai1
Cited by20
Swin Transformer: Hierarchical Vision Transformer using Shifted Windows2021
Photo-Realistic Single Image Super-Resolution Using a Generative Adversarial Network2017
Beyond a Gaussian Denoiser: Residual Learning of Deep CNN for Image Denoising2017
Least Squares Generative Adversarial Networks2017
A Survey of Convolutional Neural Networks: Analysis, Applications, and Prospects2021
Neural Architectures for Named Entity Recognition2016
Convolutional 2D Knowledge Graph Embeddings2018
Video Swin Transformer2022
Review of Deep Learning Algorithms and Architectures2019
Tacotron: Towards End-to-End Speech Synthesis2017
Automated Machine Learning2019
BioGPT: generative pre-trained transformer for biomedical text generation and mining2022
AST: Audio Spectrogram Transformer2021
Automated Breast Ultrasound Lesions Detection Using Convolutional Neural Networks2017
Leveraging Passage Retrieval with Generative Models for Open Domain Question Answering2021
End-to-end Optimized Image Compression2016
HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million\n Narrated Video Clips2019
Training Spiking Neural Networks Using Lessons From Deep Learning2023
Improving Deep Learning with Generic Data Augmentation2018
A Novel Cascade Binary Tagging Framework for Relational Triple Extraction2020