Notes
129 notes, all linked to each other. Search by title, summary, or tag.
Start here
AI & ML
Activation Functions
Activation functions introduce nonlinearity after the linear transformation in each neuron.
Actor-Critic Methods
Actor-critic combines a policy (actor) with a value function (critic) for stable, low-variance policy gradient learning.
Advantage Function
The advantage function measures how much better an action is compared to the average action in that state:
Autoregressive Generation
Autoregressive generation produces output one token at a time, feeding each generated token back as input for the next step.
Backpropagation
Backpropagation computes the gradient of the loss with respect to every parameter in the network by applying the Chain Rule (Multivariable) on the computation graph.
Batch Normalization
Batch normalization normalizes activations within each mini-batch to stabilize and accelerate training.
Bellman Equations
The Bellman equations express the recursive relationship in Value Functions: the value of a state equals the immediate reward plus the discounted value of the next state.
Bias-Variance Tradeoff
The expected test error of a model decomposes into three terms:
Causal Masking
Causal masking restricts Self-Attention so that each position can only attend to itself and earlier positions — never future tokens.
Chain-of-Thought Prompting
Chain-of-thought (CoT) prompting improves LLM reasoning by eliciting intermediate steps before the final answer.
Classical Machine Learning
Classical machine learning studies how models generalize from data without relying on large neural networks. The useful mental model is: choose a hypothesis class, define a loss…
Convolutional Neural Networks
CNNs exploit spatial structure through three key ideas: local connectivity, parameter sharing, and translation equivariance.
Cross-Validation
Cross-validation provides honest estimates of model performance on unseen data.
Curse of Dimensionality
As dimensionality increases, data becomes exponentially sparse. This breaks intuitions from low-dimensional spaces and has deep consequences for ML.
Deep Q-Network
DQN extends Q-Learning to high-dimensional state spaces by approximating with a neural network.
Direct Preference Optimization
DPO (Rafailov et al., 2023) trains directly on preference pairs without a separate reward model or RL loop.
Dropout
Dropout randomly sets each neuron's output to zero with probability during training:
Dynamic Programming in RL
Dynamic programming solves MDPs exactly when the model ( , ) is known. Foundation for understanding all RL algorithms.
Embeddings
An embedding maps discrete objects (words, tokens, users, items) into a continuous vector space where geometric relationships encode semantic relationships.
Evaluation Metrics
Classification metrics:
Inductive Bias
An inductive bias is an assumption baked into a model's architecture that constrains what functions it can learn. It encodes prior knowledge about the problem structure.
K-Nearest Neighbors
KNN is a non-parametric algorithm: it stores all training data and classifies new points by majority vote among the nearest neighbors.
KL Penalty
The KL penalty in RLHF constrains the trained policy from drifting too far from the reference (SFT) model.
Layer Normalization
Layer normalization normalizes across the feature dimension for each individual sample:
Linear Regression
Linear regression models the relationship (or with bias absorbed).
LLMs as RL Agents
The framing of LLM generation as a reinforcement learning problem reveals deep connections between language modeling and sequential decision-making.
Logistic Regression
Logistic regression is a linear classifier that models the probability of class membership:
LoRA
LoRA (Low-Rank Adaptation) makes fine-tuning large models practical by training only small low-rank matrices.
Loss Functions
The loss function measures how wrong the model's predictions are. Training minimizes it via Backpropagation.
LSTM and GRU
LSTMs and GRUs solve the vanishing gradient problem of Recurrent Neural Networks through gating mechanisms.
Markov Decision Process
An MDP is the formal framework for sequential decision-making under uncertainty.
Monte Carlo Methods in RL
Monte Carlo (MC) methods learn value functions from complete episodes of experience — no model needed.
Multi-Head Attention
Multi-head attention runs Self-Attention multiple times in parallel, each head learning different relationships.
Multimodality
Multimodal models extend a language model to consume (and sometimes produce) images, audio, or video by mapping each modality into the same embedding space the LLM already…
Neural Networks
Neural networks learn composed functions from data. The useful mental model is: layers apply differentiable transformations, nonlinearities make the composition expressive, and…
Policy Gradient Theorem
The policy gradient theorem gives the gradient of the expected return with respect to policy parameters — enabling direct optimization of the policy.
Positional Encoding
Transformers have no built-in notion of order — Self-Attention is permutation-equivariant. Positional encodings inject position information.
Pretraining
Pretraining teaches an LLM to predict the next token on massive text corpora — the foundational stage of the LLM pipeline.
Proximal Policy Optimization (PPO)
PPO is the default policy gradient algorithm for practical RL, including RLHF.
Q-Learning
Q-learning is an off-policy TD control algorithm that directly learns the optimal action-value function .
Random Forest
Random forest is an ensemble of decision trees that reduces variance through bagging and feature randomization.
ReAct
ReAct (Yao et al., 2023) interleaves reasoning and acting, enabling LLMs to use external tools.
Recurrent Neural Networks
RNNs process sequences by maintaining a hidden state that is updated at each time step:
Regularization
Regularization adds a penalty to the loss function to prevent overfitting by constraining model complexity.
REINFORCE
REINFORCE is the simplest policy gradient algorithm. It uses complete episode returns to estimate the gradient.
Reinforcement Learning
Reinforcement learning studies agents that learn by acting in an environment and receiving reward. The useful mental model is: an MDP defines the interaction, value functions…
Residual Connections
A residual (skip) connection adds the input of a layer to its output:
RLHF Pipeline
RLHF (Reinforcement Learning from Human Feedback) aligns LLMs with human preferences by training against a learned reward model.
Sampling and Decoding Strategies
At each step, an autoregressive language model outputs logits over the vocabulary. The decoding strategy determines how to pick the next token from this distribution.
SARSA
SARSA is an on-policy TD control algorithm. The name comes from the tuple .
Scaling Laws
Scaling laws describe how LLM performance (loss) improves predictably as you increase model size, data, and compute.
Self-Attention
Self-attention is the core mechanism of the transformer. It computes relationships between all pairs of positions in a sequence.
Self-Improvement in LLMs
Self-improvement methods allow LLMs to improve their own capabilities by generating and learning from their own outputs.
Softmax
Softmax converts a vector of raw scores (logits) into a probability distribution:
Supervised Fine-Tuning
SFT fine-tunes a pretrained LLM on instruction-response pairs, teaching it to be a helpful assistant instead of a text completer.
Temporal Difference Learning
TD learning combines the model-free nature of Monte Carlo with the bootstrapping of dynamic programming.
Test-Time Compute
Test-time compute refers to strategies that spend more computation at inference to improve output quality, rather than scaling the model itself.
The Artificial Neuron
The fundamental unit of a neural network:
Tokenization
Tokenization converts raw text into the integer sequences that transformers process.
Transformers and LLMs
Transformers and LLMs model token sequences by repeatedly mixing contextual information across positions. The useful mental model is: tokenization creates discrete symbols…
Value Functions
Value functions estimate how good it is to be in a state (or to take an action in a state) under a given policy’
Weight Initialization
Wrong initialization causes vanishing or exploding gradients before training even begins.
Mathematics
Adam Optimizer
Adam (Adaptive Moment Estimation) combines Momentum with per-parameter adaptive learning rates:
Bayes' Theorem
Components: - — posterior: updated belief about after observing - — likelihood: probability of evidence given - — prior: belief about before seeing evidence - — marginal…
Bayesian vs Frequentist Inference
Bayesian and frequentist inference differ in how they interpret probability and unknown parameters.
Bootstrap and Resampling
The bootstrap estimates uncertainty by resampling from the observed data.
Calculus and Optimization
Calculus studies local change; optimization uses local change to choose better parameters. The useful mental model is: derivatives describe sensitivity, Taylor expansions…
Central Limit Theorem
The Central Limit Theorem says that sums and averages of many independent random variables become approximately normal under broad conditions.
Chain Rule (Multivariable)
If , the chain rule gives:
Closure
A set is closed under an operation if applying the operation to elements of always produces an element of :
Computation Graphs
A computation graph is a directed acyclic graph (DAG) where nodes are operations and edges carry values. It makes the Chain Rule (Multivariable) systematic.
Confidence Intervals
A confidence interval is a procedure that produces a range of plausible parameter values from data.
Convexity
A function is convex if a line segment between any two points on its graph lies above the graph:
Correlation and Covariance
Covariance measures how two variables vary together.
Cross Product
The cross product is defined only in and and produces a vector (unlike the Dot Product which produces a scalar):
Determinant and Inverse
The determinant of a square matrix is a scalar that captures how the transformation scales volume:
Dot Product
The dot product (inner product) of two vectors can be understood through two equivalent views:
Eigendecomposition
An eigenvector of matrix is a nonzero vector whose direction is unchanged by the transformation: , where is the eigenvalue.
Entropy and Cross-Entropy
Entropy measures the average surprise (information content) of a distribution:
Estimators
An estimator is a rule for using data to estimate an unknown population parameter.
Expectation and Variance
Expectation (mean): the average value of a random variable. - Discrete: - Continuous: - Linearity: (always, even if dependent)
Experimental Design
Experimental design is about collecting data so comparisons support valid causal or statistical conclusions.
Gaussian Elimination
Gaussian elimination transforms a matrix into row echelon form (REF) using elementary row operations to solve linear systems .
Gradient
The gradient of a scalar function is the vector of all partial derivatives:
Hessian Matrix
The Hessian of a scalar function is the matrix of second partial derivatives:
Hypothesis Testing
A hypothesis test asks whether observed data is surprising under a null hypothesis.
Jacobian Matrix
The Jacobian of a vector-valued function is the matrix of all partial derivatives:
Key Probability Distributions
Discrete:
KL Divergence
KL divergence measures how one probability distribution diverges from a reference distribution :
Law of Large Numbers
The Law of Large Numbers says that the sample average converges to the expected value as sample size grows.
Learning Rate Schedules
The learning rate controls step size in gradient descent. Too large → divergence; too small → slow convergence. Schedules vary during training.
Linear Algebra
Linear algebra studies vector spaces and linear maps between them. The useful mental model is: vectors live in spaces, bases give coordinates, matrices represent transformations…
Linear Transformations
A linear transformation satisfies . Every linear transformation can be represented as multiplication by a matrix , and every matrix defines one.
Matrix Multiplication
Matrix multiplication can be understood through three equivalent views:
Maximum A Posteriori Estimation
MAP estimation adds a prior to Maximum Likelihood Estimation :
Maximum Likelihood Estimation
MLE finds the parameters that maximize the probability of the observed data:
Momentum
Momentum accelerates Stochastic Gradient Descent by accumulating a velocity vector in directions of persistent gradient:
Norms and Distance Metrics
A norm measures the "size" of a vector. Norms underlie nearly every loss function, regularizer, and similarity measure in ML.
Orthogonality and Projections
Two vectors are orthogonal if . An orthonormal set has all vectors mutually orthogonal with unit length.
Population vs Sample
A population is the full data-generating group you care about. A sample is the observed subset used to infer properties of that population.
Positive Definite Matrices
A symmetric matrix is positive definite (PD) if for all nonzero . Positive semi-definite (PSD) allows .
Power Iteration
Power iteration is a simple algorithm to find the dominant eigenvalue (largest in absolute value) and its eigenvector.
Principal Component Analysis (PCA)
PCA finds the directions of maximum variance in data and projects onto them for dimensionality reduction.
Probability
Probability studies uncertainty before observing data. The useful mental model is: random variables turn outcomes into quantities, distributions assign mass or density…
Random Variables
A random variable is a function from outcomes to numbers, equipped with a probability distribution describing the likelihood of each value.
Rank and Null Space
For a matrix ( ):
Sampling Distributions
A sampling distribution is the distribution of a statistic across repeated samples from the same population.
Singular Value Decomposition (SVD)
Every matrix (any shape) can be decomposed as :
Statistical Power
Statistical power is the probability that a test correctly detects a real effect.
Statistical Significance vs Practical Significance
Statistical significance asks whether an observed effect is unlikely under a null hypothesis.
Statistics Fundamentals
Statistics studies inference from finite data. The useful mental model is: probability goes from model to possible data; statistics goes from observed data back to uncertain…
Stochastic Gradient Descent
SGD and its variants are the workhorses of neural network training.
Taylor Expansion
Taylor expansion approximates a function near a point using its derivatives:
Vector Spaces and Basis
A vector space over is a set of vectors closed under addition and scalar multiplication. is the canonical example.
Computer Science
Approximate Nearest Neighbor Search
Exact nearest neighbor search in high dimensions is per query (brute-force scan). For millions of vectors, this is too slow. Approximate methods trade a small accuracy loss for…
Big-O and Complexity Analysis
Big-O notation describes how an algorithm's time or space scales with input size , ignoring constants and lower-order terms.
Computational Complexity of Attention
Standard Self-Attention has time and memory in sequence length . This is the fundamental bottleneck of transformers.
Computer Science Foundations
Computer science foundations study the cost and structure of computation. The useful mental model is: algorithms transform inputs into outputs, data structures control access…
Distributed Training Strategies
When a model or dataset is too large for a single GPU, training must be distributed across multiple devices.
Dynamic Programming
Dynamic programming (DP) solves problems with optimal substructure (optimal solution built from optimal sub-solutions) and overlapping subproblems (same subproblems recur) by…
Floating Point and Quantization
Numbers in hardware have finite precision. Choosing the right format trades off range, precision, memory, and speed.
GPU Architecture and CUDA
GPUs achieve massive parallelism through thousands of simple cores executing the same instruction on different data (SIMT — Single Instruction, Multiple Threads).
Graphs and Traversals
A graph consists of vertices and edges. Directed graphs (digraphs) have ordered edges. A DAG (directed acyclic graph) has no cycles.
Hash Tables
A hash table maps keys to values via a hash function , giving average-case lookup, insert, and delete.
Memory Hierarchy and IO-Awareness
Modern hardware is memory-bound, not compute-bound for most ML operations. Understanding the memory hierarchy is the key to writing fast code.
P vs NP and Intractability
P = problems solvable in polynomial time. NP = problems whose solutions are verifiable in polynomial time. The question is open, but widely believed to be .
Randomized Algorithms
Randomized algorithms use random choices to achieve better average-case performance, simpler implementations, or solutions to problems where deterministic approaches are…
Sorting and Selection
Sorting arranges elements in order. Selection finds the -th smallest (or largest) element without fully sorting.
Systems and Scaling
Systems and scaling study how model training and inference behave on real hardware. The useful mental model is: performance is limited by compute, memory, communication, and…