Linearized attention
Without the softmax, regrouping one product avoids the \(n\times n\) matrix: cost \(O(n^2d)\to O(nd^2)\).
Qatar Computing Research Institute (QCRI), Hamad Bin Khalifa University (HBKU)Department of Computer Science and Engineering, University of Minnesota
This survey examines modern large language models from a numerical linear algebra perspective. It has two goals: to explain the key concepts behind LLMs in a form accessible to specialists in numerical methods, and to show where numerical linear algebra appears throughout modern LLM design, training, and analysis. It is written for both numerical analysts and machine learning researchers.
Six selected connections between familiar LLM mechanisms and results from numerical linear algebra and related mathematics.
Without the softmax, regrouping one product avoids the \(n\times n\) matrix: cost \(O(n^2d)\to O(nd^2)\).
Each output row is a similarity-weighted average of value vectors; the attention matrix is row-stochastic.
Truncated SVD gives the best rank-\(r\) approximation; LoRA learns a low-rank update with \(r\ll d\).
Distance preservation and near-orthogonality in high dimensions; intrinsic dimension of hidden states.
The reverse-mode chain rule, layer by layer: a sequence of transposed-Jacobian products.
A Newton–Schulz polynomial iteration approximates the polar factor \(\mathbf{U}\mathbf{V}^\top\) of the momentum matrix.
Six selected connections, not an exhaustive list.
A decoder-only Transformer with each topic pinned to the component it acts on, tagged with its section.
Click the figure to open it at full size.
Additional resources and updates will be added as they become available.
@article{baggag2026nla,
title = {The Numerical Linear Algebra of Large Language Models},
author = {Baggag, Abdelkader and Saad, Yousef},
journal = {arXiv preprint arXiv:2610.04631},
year = {2026}
}