Project
The Theory of Large Language Models
A Bayesian account of how large language models learn: in-context learning as inference on a latent manifold, and the geometry of attention.
Related Publications
Siddhartha Dalal, Vishal Misra, Abhay Parekh
Bayesian Wind Tunnels for Model Selection CoRR, January 2026
Naman Aggarwal, Siddhartha Dalal, Vishal Misra
The Bayesian Geometry of Transformer Attention CoRR, January 2025
Naman Aggarwal, Siddhartha Dalal, Vishal Misra
Gradient Dynamics of Attention: How Cross-Entropy Sculpts Bayesian Manifolds CoRR, January 2025
Naman Aggarwal, Siddhartha Dalal, Vishal Misra
Geometric Scaling of Bayesian Inference in LLMs CoRR, January 2025
Siddhartha Dalal, Vishal Misra
The Matrix: A Bayesian learning model for LLMs CoRR, February 2024