The Numerical Linear Algebra of Large Language Models — Abdelkader Baggag and Yousef Saad

The Numerical Linear Algebra of Large Language Models by Abdelkader Baggag and Yousef Saad: a Transformer pipeline from tokens to next token, annotated with six numerical linear algebra connections — linearized attention, kernel regression, low-rank structure, representation geometry, backpropagation and the Muon optimizer.

Qatar Computing Research Institute (QCRI), Hamad Bin Khalifa University (HBKU)Department of Computer Science and Engineering, University of Minnesota

Overview

This survey examines modern large language models from a numerical linear algebra perspective. It has two goals: to explain the key concepts behind LLMs in a form accessible to specialists in numerical methods, and to show where numerical linear algebra appears throughout modern LLM design, training, and analysis. It is written for both numerical analysts and machine learning researchers.

Six Numerical Linear Algebra Connections

Six selected connections between familiar LLM mechanisms and results from numerical linear algebra and related mathematics.

Linearized attention

Cayley (1858) · associativity of matrix products
\((\mathbf{Q}\mathbf{K}^\top)\mathbf{V}=\mathbf{Q}(\mathbf{K}^\top\mathbf{V})\)

Without the softmax, regrouping one product avoids the \(n\times n\) matrix: cost \(O(n^2d)\to O(nd^2)\).

§4.3–4.4

Kernel regression view

Nadaraya–Watson (1964)
\(\operatorname{softmax}\!\left(\mathbf{Q}\mathbf{K}^\top/\sqrt{d_k}\right)\mathbf{V}\)

Each output row is a similarity-weighted average of value vectors; the attention matrix is row-stochastic.

§2.8, §3.4

Low-rank structure: LoRA & compression

Eckart–Young (1936)
\(\mathbf{W}_r=\mathbf{U}_r\boldsymbol{\Sigma}_r\mathbf{V}_r^\top,\quad \mathbf{W}=\mathbf{W}_0+\mathbf{A}\mathbf{B}\)

Truncated SVD gives the best rank-\(r\) approximation; LoRA learns a low-rank update with \(r\ll d\).

§4.1–4.2

Representation geometry

Johnson–Lindenstrauss (1984)
\(\|\phi(\mathbf{u})-\phi(\mathbf{v})\|\approx\|\mathbf{u}-\mathbf{v}\|\)

Distance preservation and near-orthogonality in high dimensions; intrinsic dimension of hidden states.

§3.14–3.15

Backpropagation

Linnainmaa (1970) · reverse-mode differentiation
\(\mathbf{J}_\ell=\partial\mathbf{s}_\ell/\partial\mathbf{s}_{\ell-1},\quad \boldsymbol{\delta}_{\ell-1}=\mathbf{J}_\ell^\top\boldsymbol{\delta}_\ell\)

The reverse-mode chain rule, layer by layer: a sequence of transposed-Jacobian products.

§2.6

Muon optimizer

Schulz (1933) · Newton–Schulz
\(\mathbf{X}_{k+1}=a\mathbf{X}_k+b(\mathbf{X}_k\mathbf{X}_k^\top)\mathbf{X}_k+c(\mathbf{X}_k\mathbf{X}_k^\top)^2\mathbf{X}_k\)

A Newton–Schulz polynomial iteration approximates the polar factor \(\mathbf{U}\mathbf{V}^\top\) of the momentum matrix.

§5.5

Six selected connections, not an exhaustive list.

Technical Overview

A decoder-only Transformer with each topic pinned to the component it acts on, tagged with its section.

Technical overview of a decoder-only Pre-LN Transformer: attention, MLP/MoE, LayerNorm and residual stream, each annotated with the matrix operation it uses, plus low-rank approximation, linearized attention, geometry and interpretability, and matrix-aware optimizers.

Click the figure to open it at full size.

What the Paper Covers

Foundations

§2
  • AI and deep-learning foundations
  • Multilayer perceptrons and loss functions
  • Stochastic optimization
  • Computational graphs and backpropagation
  • Matrix and tensor computations

LLMs and Transformers

§3.1–3.13
  • Language modeling, tokens, and embeddings
  • Self-attention and multi-head attention
  • Positional representations
  • Feed-forward layers, normalization, and residual connections
  • Decoder-only LLMs and Mixture-of-Experts
  • Training dynamics

Geometry & Interpretability

§3.14–3.16
  • High-dimensional representation geometry
  • Johnson–Lindenstrauss and near-orthogonality
  • Intrinsic dimension
  • Interpretability through linear algebra
  • Scaling laws and emergent phenomena

Spectral & efficient methods

§4
  • Model compression
  • Low-rank structure and LoRA
  • Linearized and efficient attention
  • Randomization and random projections

Advanced optimization

§5
  • Natural gradients
  • Fisher and Gauss–Newton structure
  • K-FAC
  • Shampoo
  • Muon and Newton–Schulz iterations

Resources

Media & Figures

Additional resources and updates will be added as they become available.

Citation

@article{baggag2026nla,
  title   = {The Numerical Linear Algebra of Large Language Models},
  author  = {Baggag, Abdelkader and Saad, Yousef},
  journal = {arXiv preprint arXiv:2610.04631},
  year    = {2026}
}

Updates

Version 1 posted on arXiv (2610.04631).