![]() |
|
Mathematical Foundations of Deep Learning [Ye] - Printable Version +- MKLab (https://mklab.gr) +-- Forum: [INDEX] (https://mklab.gr/forumdisplay.php?fid=1) +--- Forum: ARTFICIAL INTELLIGENCE (AI) (https://mklab.gr/forumdisplay.php?fid=5) +---- Forum: BOOKS (https://mklab.gr/forumdisplay.php?fid=32) +----- Forum: FREE EBOOKS (https://mklab.gr/forumdisplay.php?fid=111) +----- Thread: Mathematical Foundations of Deep Learning [Ye] (/showthread.php?tid=1712) |
Mathematical Foundations of Deep Learning [Ye] - mklabgr - 08-20-2026 Mathematical Foundations of Deep Learning Author: Xiaojing Ye arXiv: 2603.18387 Submitted: 19 March 2026 Type: Draft book; final version published by Chapman & Hall/CRC in 2026. This work develops a rigorous mathematical framework for understanding modern deep learning. Rather than treating neural networks mainly as engineering tools, Ye presents them as mathematical objects: neural networks are viewed as classes of function approximators, while training becomes a large-scale, generally non-convex optimization problem. The book connects deep learning with approximation theory, functional analysis, probability, statistics, optimization and dynamical systems, with particular emphasis on explaining why neural networks can approximate complex functions, how they are trained, and what theoretical principles govern their behavior. A major early result discussed is the Universal Approximation Theorem, followed by network architectures, activation functions, automatic differentiation and deterministic and stochastic optimization methods. A distinctive feature is that the mathematical treatment extends well beyond standard supervised neural networks. Deep networks are linked to optimal control theory, including Euler–Lagrange equations, Hamiltonian systems, the Pontryagin Maximum Principle and Hamilton–Jacobi–Bellman equations. This leads naturally to Neural ODEs and then to reinforcement learning, where Markov decision processes, Bellman equations and policy-improvement methods are developed within essentially the same mathematical framework. The final part turns to modern generative AI, covering variational autoencoders, GANs, diffusion models and flow matching, together with their interpretation through probability distributions, stochastic differential equations and probability-density control. The central message is that seemingly different areas of modern AI share a surprisingly unified mathematical structure. Approximation explains what neural networks can represent; optimization explains how their parameters are learned; control theory and dynamic programming explain sequential decision-making; and probability, differential equations and transport ideas explain many modern generative models. The intended audience is advanced undergraduate or graduate students and researchers with a solid background in calculus, linear algebra and probability, with real analysis being useful. The work is therefore particularly valuable as a bridge between traditional mathematics and contemporary AI, rather than as a practical programming manual. Key takeaways
arXiv paper |