Loss Functions

1 min read

The loss function measures how wrong the model's predictions are. Training minimizes it via Backpropagation.

Regression:

Mean Squared Error (MSE): L=1n∑i(yi−y^i)2L = \frac{1}{n}\sum_i(y_i - \hat{y}_i)^2

  • Equivalent to MLE assuming Gaussian noise
  • Penalizes large errors quadratically → sensitive to outliers
  • Gradient: ∂L∂y^i=2n(y^i−yi)\frac{\partial L}{\partial \hat{y}_i} = \frac{2}{n}(\hat{y}_i - y_i)

Classification:

Binary cross-entropy: L=−1n∑i[yilog⁡y^i+(1−yi)log⁡(1−y^i)]L = -\frac{1}{n}\sum_i [y_i\log\hat{y}_i + (1-y_i)\log(1-\hat{y}_i)]

  • Equivalent to negative log-likelihood of Bernoulli distribution
  • Used with sigmoid output

Categorical cross-entropy: L=−1n∑i∑kyiklog⁡y^ikL = -\frac{1}{n}\sum_i\sum_k y_{ik}\log\hat{y}_{ik}

  • Equivalent to negative log-likelihood of categorical distribution
  • Used with softmax output
  • One-hot yy: simplifies to L=−1n∑ilog⁡y^i,ciL = -\frac{1}{n}\sum_i \log\hat{y}_{i,c_i} where cic_i is the true class

Connection: minimizing cross-entropy = minimizing KL Divergence from the true distribution (since H(p,q)=H(p)+DKL(p∥q)H(p,q) = H(p) + D_\text{KL}(p \| q) and H(p)H(p) is constant).

In LLM pretraining: cross-entropy on next-token prediction is the standard objective → Pretraining.

See also: Maximum Likelihood Estimation, Logistic Regression

Linked from