
Lecture 11: Fundamental Concepts in Machine Learning
08 October 2026

Both curves are averages of the next election’s vote share within bins of today’s margin. The right one fits this sample far better. Which one would you use to predict next year’s elections?
In-sample fit is not the goal. A prediction model is judged on data it has never seen.
Additional derivations and code follow the main lesson.
Run the examples from the lectures/ directory.
| Object | Role |
|---|---|
lee08 (from RDHonest) |
6,558 U.S. House elections: the Democratic margin of victory and the Democratic vote share at the next election |
Everything else in this lecture is simulated inside the deck, so no file is read from data/.
Fitting this sample is not the same as predicting the next one
Fitting past data too well usually makes a model predict new data worse. This is overfitting, and avoiding it is one of the main concerns of machine learning.
Y_i = \underbrace{X_i'\beta}_{\text{signal}} + \underbrace{\varepsilon_i}_{\text{noise}}, \qquad \mathbf{E}[\varepsilon_i \mid X_i] = 0.
## [1] 0.00 0.33 0.68 1.00

In-sample fit measures complexity as much as it measures signal.
Place the observations into d bins by the value of X_i: (t_0, t_1],\ (t_1, t_2],\ \dots,\ (t_{d-1}, t_d].
The prediction at x is the average outcome of the observations in the same bin as x: \hat f(x) = \frac{1}{n_j}\sum_{i:\ t_{j-1} < X_i \le t_j} Y_i \qquad \text{for } t_{j-1} < x \le t_j, \quad n_j = \#\{i: t_{j-1} < X_i \le t_j\}.
Three implementation choices for the election data:
nbins bins on each side of zero, 2 \times nbins in total.This is the cell-mean idea from the regression lecture, with cells defined by bins of a continuous X.
bin_cutpoints <- function(x_data, nbins) {
x_neg <- sort(x_data[x_data < 0]); x_neg[length(x_neg)] <- 0 # make 0 a cutpoint
x_pos <- sort(x_data[x_data > 0])
unique(c(-Inf, x_neg[ceiling((1:nbins) * length(x_neg) / nbins)], # quantile cutpoints,
x_pos[ceiling((1:nbins) * length(x_pos) / nbins)])) # no empty bins
}
binned_predictor <- function(x_data, y_data, new_x, nbins) {
cutpoints <- bin_cutpoints(x_data, nbins)
means_by_bin <- tapply(y_data, cut(x_data, breaks = cutpoints), mean)
means_by_bin[cut(new_x, breaks = cutpoints)] # look up the bin of each new x
}
binned_predictor(lee08$margin, lee08$voteshare, new_x = c(-30, -1, 1, 30), nbins = 5)## (-30.2,-19.5] (-9.46,0] (0,12.6] (27.8,43.7]
## 36.30102 43.33206 56.30646 67.91751
cut() assigns each value to a bin (t_{j-1}, t_j], tapply() averages Y within bins, and indexing by the bins of new_x looks up the predictions.unique() drops duplicate cutpoints when many elections share a margin (the uncontested races at \pm 100), so the number of bins formed can be a little below 2 \times nbins.xgrid <- seq(-100, 100, by = 0.1)
fit10 <- binned_predictor(lee08$margin, lee08$voteshare, new_x = xgrid, nbins = 5)
ggplot() +
geom_point(data = lee08, aes(margin, voteshare), colour = "grey70", alpha = 0.4, size = 0.6) +
geom_line(aes(xgrid, fit10), colour = smu_red, linewidth = 1) +
labs(x = "Democratic margin at t", y = "Predicted vote share at t + 1")
Each flat segment is a bin mean. Bins are narrow where elections are dense, near zero, and wide in the tails.

The same function with nbins = 1, 5, 10, 100, 1,000 and 2,500. At 5,000 bins almost every bin holds one election, and the “prediction” is the raw Y_i of whichever election is nearest.
In OLS, complexity is the number of coefficients. For the binned mean it is the number of bins, and the two are the same thing: the binned mean is a regression on bin indicators.
## [1] 1.563194e-13
| nbins | bins | in_sample_r2 |
|---|---|---|
| 1 | 2 | 0.517 |
| 5 | 10 | 0.673 |
| 10 | 20 | 0.678 |
| 100 | 200 | 0.692 |
| 1000 | 2000 | 0.761 |
| 2500 | 5000 | 0.878 |
Measure prediction error on data the model never saw
The in-sample deviance for least squares, fitted on i = 1, \dots, n: \text{dev}_{IS}(\hat\beta) = \sum_{i=1}^n (Y_i - X_i'\hat\beta)^2.
Suppose m new observations arrive, labelled n+1, \dots, n+m. The out-of-sample deviance is \text{dev}_{OOS}(\hat\beta) = \sum_{i=n+1}^{n+m} (Y_i - X_i'\hat\beta)^2.
R^2_{OOS} = 1 - \frac{\text{dev}_{OOS}(\hat\beta)}{\text{dev}_{OOS}(\hat\beta_{\text{null}})}, \qquad \text{dev}_{OOS}(\hat\beta_{\text{null}}) = \sum_{i=n+1}^{n+m} (Y_i - \bar Y)^2,
where \bar Y is the mean of the estimation sample, the prediction of the no-covariate model.
set.seed(0)
n <- nrow(lee08)
est <- sample.int(n, size = round(2 * n / 3)) # 4,372 elections to fit, 2,186 to test
x_est <- lee08$margin[est]; y_est <- lee08$voteshare[est]
x_test <- lee08$margin[-est]; y_test <- lee08$voteshare[-est]
r2_split <- function(nbins) {
fit_in <- binned_predictor(x_est, y_est, new_x = x_est, nbins)
fit_out <- binned_predictor(x_est, y_est, new_x = x_test, nbins)
tibble(bins = 2 * nbins,
in_sample = 1 - sum((y_est - fit_in)^2) / sum((y_est - mean(y_est))^2),
out_of_sample = 1 - sum((y_test - fit_out)^2) / sum((y_test - mean(y_est))^2))
}
r2_split(nbins = 5)## # A tibble: 1 × 3
## bins in_sample out_of_sample
## <dbl> <dbl> <dbl>
## 1 10 0.662 0.693
| bins | in_sample | out_of_sample |
|---|---|---|
| 10 | 0.662 | 0.693 |
| 20 | 0.668 | 0.697 |
| 200 | 0.688 | 0.693 |
| 2000 | 0.796 | 0.587 |
| 5000 | 0.935 | 0.448 |

oos_r2 <- function(seed, nbins = 100) {
set.seed(seed)
est <- sample.int(n, size = round(2 * n / 3))
fit <- binned_predictor(lee08$margin[est], lee08$voteshare[est], lee08$margin[-est], nbins)
1 - sum((lee08$voteshare[-est] - fit)^2) /
sum((lee08$voteshare[-est] - mean(lee08$voteshare[est]))^2)
}
round(sapply(1:10, oos_r2), 3)## [1] 0.677 0.690 0.673 0.651 0.716 0.658 0.679 0.686 0.672 0.660
Step 1. Split the data at random into K folds of (almost) equal size.
Step 2. For each k = 1, \dots, K, fit the model on all folds except the kth, giving a predictor \hat f_{-k}, and score it on fold k: \text{dev}_{OOS,k} = \sum_{i \in \text{fold } k} \big(Y_i - \hat f_{-k}(X_i)\big)^2.
Step 3. Average the K out-of-sample deviances: \text{dev}_{CV} = \frac{1}{K}\sum_{k=1}^K \text{dev}_{OOS,k}.
Every observation is predicted exactly once, by a model that did not see it. We report the result per observation, as a mean squared prediction error.

## foldid
## 1 2 3 4 5 6 7 8 9 10
## 656 656 656 656 656 656 656 656 656 654
rep(1:K, each = ceiling(n / K)) writes the labels 1, 1, \dots, 1, 2, 2, \dots, K: ten blocks of 656, slightly more than n = 6{,}558 labels.sample(1:n) is a random permutation of 1, \dots, n. Indexing by it shuffles the labels, keeps n of them and drops the rest.nbins_vec <- 1:250 # 2 to 500 bins: more was clearly worse in the single split
cv_mat <- matrix(NA, nrow = K, ncol = length(nbins_vec)) # row k: fold k held out
for (i in seq_along(nbins_vec)) {
for (k in 1:K) {
train <- which(foldid != k)
fit_k <- binned_predictor(lee08$margin[train], lee08$voteshare[train],
new_x = lee08$margin[-train], nbins = nbins_vec[i])
cv_mat[k, i] <- mean((lee08$voteshare[-train] - fit_k)^2)
}
}
cv_error <- colMeans(cv_mat) # K-fold CV estimate for each candidate
cv_se <- apply(cv_mat, 2, sd) / sqrt(K) # standard error of that average across folds
nbins_cv <- nbins_vec[which.min(cv_error)]
c(bins_chosen = 2 * nbins_cv, cv_error = min(cv_error), se = cv_se[nbins_cv])## bins_chosen cv_error se
## 46.000000 184.421240 4.345252
cv_mat[k, i] is the mean squared prediction error on fold k for candidate i. The column mean is \text{dev}_{CV} per observation.

Refitted on the full sample with the chosen 23 bins on each side of zero, using the same plotting code as the ten-bin slide. The staircase is fine where elections are dense, and the jump at zero survives.
## min_rule one_se_rule
## 46 12
cv.glmnet() reports both choices next lecture, as lambda.min and lambda.1se.| K | Training size | Properties |
|---|---|---|
| 2 | n/2 | Cheap; pessimistic, since each model sees half the data |
| 5 or 10 | 0.8\,n or 0.9\,n | The standard choices |
| n (leave-one-out) | n - 1 | Closest to the full-sample model; n fits, and the n errors are highly correlated |
Why cross-validation works, and why complexity cuts both ways
Posit Y = f(X) + \varepsilon with \mathbf{E}[\varepsilon \mid X] = 0. Linear regression is the case f(x) = x'\beta; the binned mean, trees and neural networks estimate more flexible f.
Let \hat f be an estimator computed from the sample (X_1, Y_1), \dots, (X_n, Y_n), and (X_{n+1}, Y_{n+1}) a new draw from the same population. The population mean squared error is
\text{MSE} = \mathbf{E}\big[(Y_{n+1} - \hat f(X_{n+1}))^2\big].
The expectation averages over two sources of randomness:
For m new observations, \mathbf{E}\left[\frac{1}{m}\,\text{dev}_{OOS}\right] = \frac{1}{m}\sum_{i=n+1}^{n+m} \mathbf{E}\big[(Y_i - \hat f(X_i))^2\big] = \text{MSE}.
\begin{aligned} \mathbf{E}\big[(Y_{n+1} - \hat f(X_{n+1}))^2\big] &= \mathbf{E}\big[(\varepsilon_{n+1} + f(X_{n+1}) - \hat f(X_{n+1}))^2\big] \\ &= \underbrace{\mathbf{E}[\varepsilon_{n+1}^2]}_{\text{irreducible error}} + \underbrace{\mathbf{E}\big[(\hat f(X_{n+1}) - f(X_{n+1}))^2\big]}_{\text{estimation error}} + 2\,\mathbf{E}\big[\varepsilon_{n+1}\big(f(X_{n+1}) - \hat f(X_{n+1})\big)\big]. \end{aligned}
The cross term is zero: \hat f depends on the sample only, and \varepsilon_{n+1} is a fresh mean-zero draw, so that \mathbf{E}[\varepsilon_{n+1} \mid \text{sample}, X_{n+1}] = 0.
Fix x and add and subtract \mathbf{E}[\hat f(x)], the average of \hat f(x) over samples:
\begin{aligned} \mathbf{E}\big[(\hat f(x) - f(x))^2\big] &= \mathbf{E}\Big[\big(\hat f(x) - \mathbf{E}[\hat f(x)] + \mathbf{E}[\hat f(x)] - f(x)\big)^2\Big] \\ &= \underbrace{\mathbf{E}\Big[\big(\hat f(x) - \mathbf{E}[\hat f(x)]\big)^2\Big]}_{V(x):\ \text{variance}} + \underbrace{\big(\mathbf{E}[\hat f(x)] - f(x)\big)^2}_{B(x)^2:\ \text{squared bias}}, \end{aligned}
since the cross term is 2\,\big(\mathbf{E}[\hat f(x)] - f(x)\big)\,\mathbf{E}\big[\hat f(x) - \mathbf{E}[\hat f(x)]\big] = 0.
Averaging over the distribution of the new X_{n+1}, \text{MSE} = \underbrace{\sigma^2}_{\text{irreducible}} + \underbrace{\mathbf{E}[V(X_{n+1})]}_{\text{variance}} + \underbrace{\mathbf{E}[B(X_{n+1})^2]}_{\text{squared bias}}.
| Component | Where it comes from | Can we reduce it? |
|---|---|---|
| Irreducible error | Noise in Y given X | No. Only better predictors X help. |
| Variance | The particular sample we drew | Yes: a simpler model, or more data |
| Squared bias | The model’s inability to represent f | Yes: a more flexible model |
No need to memorize the algebra. The point is that the two reducible components move in opposite directions when we change the complexity of the model.
Next lecture adds a knob that runs the other way, a penalty \lambda on the size of the coefficients:
\text{penalization } \lambda \uparrow \;\Longrightarrow\; \text{complexity} \downarrow \;\Longrightarrow\; \text{bias} \uparrow,\ \text{variance} \downarrow.
In the extreme \lambda \to \infty, \hat\beta \to 0 and we predict Y_{n+1} = 0 without looking at the data: zero variance, large bias (unless f = 0).
nbins is simply the number of bins, and the binned mean approximates a line with a staircase.
R <- 300
grid <- seq(0.025, 0.975, length.out = 191) # where the fits are evaluated
bins_sim <- c(1, 2, 3, 5, 10, 20, 50, 100)
fits <- array(NA, dim = c(R, length(grid), length(bins_sim)))
for (r in 1:R) {
x <- runif(n_sim); y <- 1 + x + 0.2 * rnorm(n_sim) # a fresh sample each time
for (j in seq_along(bins_sim))
fits[r, , j] <- binned_predictor(x, y, new_x = grid, nbins = bins_sim[j])
}
bias_var <- map_dfr(seq_along(bins_sim), function(j) {
avg_fit <- colMeans(fits[, , j]) # E[f_hat(x)] at each grid point
tibble(bins = bins_sim[j],
bias2 = mean((avg_fit - (1 + grid))^2), # B(x)^2 averaged over the grid
variance = mean(apply(fits[, , j], 2, var))) # V(x) averaged over the grid
}) %>% mutate(mse = bias2 + variance + 0.2^2)fits[r, , j] holds sample r’s fit with bins_sim[j] bins along the grid. The mean over r approximates \mathbf{E}[\hat f(x)] and the variance over r approximates V(x).
Grey: fits from 20 different samples. Red: the average fit over all 300. Blue: the truth. With 3 bins the average is a staircase far from the line and the grey fits lie on top of each other: bias without variance. With 50 bins the average sits on the line but the individual fits scatter: variance without bias.

| bins | bias2 | variance | mse |
|---|---|---|---|
| 1 | 0.0760 | 0.0002 | 0.1162 |
| 2 | 0.0160 | 0.0034 | 0.0595 |
| 3 | 0.0062 | 0.0028 | 0.0489 |
| 5 | 0.0016 | 0.0021 | 0.0436 |
| 10 | 0.0001 | 0.0015 | 0.0416 |
| 20 | 0.0000 | 0.0018 | 0.0418 |
| 50 | 0.0000 | 0.0041 | 0.0441 |
| 100 | 0.0000 | 0.0081 | 0.0481 |
The MSE is smallest at 10 bins, 0.0416, just above the floor of 0.04 set by the noise. Fewer bins pay in bias, more bins pay in variance.
## bias2 variance
## 4.848163e-07 1.529540e-04
A formula in place of the cross-validation loop
\begin{aligned} \text{dev}_{IS} &= \sum_{i=1}^n (Y_i - \hat f(X_i))^2 = \sum_{i=1}^n \big(Y_i - f(X_i) + f(X_i) - \hat f(X_i)\big)^2 \\ &= \underbrace{\sum_{i=1}^n (Y_i - f(X_i))^2}_{\text{noise}} + \underbrace{\sum_{i=1}^n (\hat f(X_i) - f(X_i))^2}_{\text{estimation error}} - 2 \underbrace{\sum_{i=1}^n (Y_i - f(X_i))\big(\hat f(X_i) - f(X_i)\big)}_{\text{the overfitting term}}. \end{aligned}
For least squares, define the degrees of freedom of an estimator as df = \frac{1}{\sigma^2}\,\mathbf{E}\left[\sum_{i=1}^n (Y_i - f(X_i))\big(\hat f(X_i) - f(X_i)\big)\right], \qquad \sigma^2 = \mathbf{E}[\varepsilon_i^2], so the overfitting term is 2\,\sigma^2 df on average.
| Estimator | df |
|---|---|
| OLS, logit, other maximum likelihood estimators | Number of parameters |
| Binned mean | Number of bins |
| Lasso | Number of nonzero coefficients (a deep result; next lecture) |
| Ridge | A formula in \lambda and the design matrix |
summary() of an lm reports residual degrees of freedom, n - df, not df.Adding the estimated overfitting term to the deviance gives Mallows’ C_p: C_p = \frac{1}{\hat\sigma^2}\sum_{i=1}^n (Y_i - \hat f(X_i))^2 + 2\,df, with \hat\sigma^2 taken from a low-bias model (many bins, or all available regressors).
For likelihood models, Akaike’s information criterion has the same shape, \text{AIC}(\hat f) = \text{dev}_{IS}(\hat f) + 2\,df, and reduces to C_p for a Gaussian likelihood with \sigma^2 treated as known. The corrected AIC inflates the penalty when df is not small relative to n, which also accounts for estimating \sigma^2: \text{AICc}(\hat f) = \text{dev}_{IS}(\hat f) + 2\,df\,\frac{n}{n - df - 1}.
Smaller is better. Packages differ in constants and normalizations, so compare AIC values only within one routine.
dev_df <- function(nbins) {
fit <- binned_predictor(lee08$margin, lee08$voteshare, new_x = lee08$margin, nbins)
tibble(bins = 2 * nbins, dev = sum((lee08$voteshare - fit)^2),
df = length(bin_cutpoints(lee08$margin, nbins)) - 1) # bins actually formed
}
ic <- map_dfr(nbins_vec, dev_df)
sigma2_hat <- with(ic[250, ], dev / (n - df)) # from the 500-bin, low-bias model
ic <- ic %>% mutate(cp = dev / sigma2_hat + 2 * df,
cp_as_mse = cp * sigma2_hat / n) # same ranking, on the CV error's scale
c(sigma2_hat = sigma2_hat, bins_cp = ic$bins[which.min(ic$cp)], bins_cv = 2 * nbins_cv)## sigma2_hat bins_cp bins_cv
## 184.5481 44.0000 46.0000
df counts the bins actually formed, since ties in the margin merge some cutpoints.
| Cross-validation | AIC and C_p | |
|---|---|---|
| Estimates | Population MSE of a model fitted on n(K-1)/K observations | Expected error of predicting a new Y at the sample’s X_i (appendix) |
| Needs | Only the ability to refit the model | A df formula; for the theory, homoskedastic (and often normal) errors |
| Cost | K fits per candidate | One fit per candidate |
| Scope | Any predictor, any loss function | Likelihood-based models with a known df |
Both estimate out-of-sample performance from a single sample, and both are used the same way: compute the criterion along the candidates and take the minimum. AIC is a computational shortcut to CV, with a slightly different target and stronger assumptions.
In-sample fit rewards complexity. R^2 rises with every added parameter, even pure noise, and a model that tracks \varepsilon_i predicts nothing.
Out-of-sample error is the criterion. Hold data out, predict it, and score the predictions. K-fold cross-validation does this once for every observation and averages the result.
Cross-validation estimates the population MSE. The MSE is irreducible error plus variance plus squared bias, and complexity trades the last two against each other.
Information criteria are the shortcut. AIC and C_p add 2\,df to the in-sample deviance to undo its optimism, and they track the CV curve closely at a fraction of the cost.
Choose the tuning parameter for the target. CV optimizes average prediction. A local parameter such as the incumbency jump needs its own criterion.
RDHonest (Kolesár): the lee08 data set and the regression discontinuity estimator used for comparison, from Armstrong and Kolesár (2020), Quantitative Economics 11, 1-39.Optional material for reference and practice
cv_choice <- function(seed) {
set.seed(seed)
foldid <- rep(1:K, each = ceiling(n / K))[sample(1:n)]
err <- sapply(nbins_vec, function(nb) mean(sapply(1:K, function(k) {
train <- which(foldid != k)
fit_k <- binned_predictor(lee08$margin[train], lee08$voteshare[train],
new_x = lee08$margin[-train], nbins = nb)
mean((lee08$voteshare[-train] - fit_k)^2)
})))
c(seed = seed, bins = 2 * nbins_vec[which.min(err)], min_error = min(err))
}
t(sapply(1:3, cv_choice))## seed bins min_error
## [1,] 1 46 184.1956
## [2,] 2 42 184.1855
## [3,] 3 44 184.0691
The minimum moves between 42 and 46 bins, and the flat valley from about 30 to 130 bins is there in every case. Report the region, not only the point.
Leaving out observation i changes only its own bin’s mean. If bin j holds n_j observations, \hat f_{-i}(X_i) = \frac{n_j\,\hat f(X_i) - Y_i}{n_j - 1} \qquad\Longrightarrow\qquad Y_i - \hat f_{-i}(X_i) = \frac{Y_i - \hat f(X_i)}{1 - 1/n_j}.
loocv <- sapply(nbins_vec, function(nb) {
bin <- cut(lee08$margin, breaks = bin_cutpoints(lee08$margin, nb))
n_bin <- table(bin)[bin] # size of each observation's bin
fit <- tapply(lee08$voteshare, bin, mean)[bin]
mean(((lee08$voteshare - fit) / (1 - 1 / n_bin))^2)
})
c(bins_loocv = 2 * nbins_vec[which.min(loocv)], loocv_error = min(loocv, na.rm = TRUE))## bins_loocv loocv_error
## 44.0000 184.3498
The same identity holds for any linear smoother with 1/n_j replaced by the leverage h_{ii}, the diagonal of the hat matrix for OLS. Candidates with a single-observation bin make the formula undefined and are skipped.
## estimate std.error bandwidth eff.obs
## I(margin>0) 5.849736 1.365882 7.715099 764.5629
RDHonest() fits a local linear regression on each side of the cutoff within a bandwidth of about 7.7 points of margin, chosen for estimating the jump rather than for global fit, and reports a bias-aware confidence interval.Stein’s lemma. If \varepsilon_i \sim N(0, \sigma^2) independently and \hat f_i = \hat f(X_i) is a smooth function of Y = (Y_1, \dots, Y_n), \mathbf{E}\left[\sum_{i=1}^n \varepsilon_i\,(\hat f_i - f_i)\right] = \sigma^2\, \mathbf{E}\left[\sum_{i=1}^n \frac{\partial \hat f_i}{\partial Y_i}\right], \qquad\text{so}\qquad df = \mathbf{E}\left[\sum_{i=1}^n \frac{\partial \hat f_i}{\partial Y_i}\right].
AIC() on the binned regression## AIC_R by_hand
## 52830.86 52830.86
## # A tibble: 1 × 3
## bins df aic
## <dbl> <dbl> <dbl>
## 1 44 42 34218.
lm, R’s AIC() uses the Gaussian log-likelihood with \sigma^2 estimated: n \log(\text{dev}/n) + 2(df + 1) plus a constant, where the +1 counts \sigma^2 as a parameter.Run once in your R environment if needed:
RDHonest provides the lee08 data set and the regression discontinuity estimator used for comparison.lee08, so no files are read from lectures/data/ and the deck runs anywhere once the packages are installed.Which of the two models on the first slide would you trust next November? The 10-bin one: its cross-validated error is far lower, and the 5,000-bin model is reproducing noise. Cross-validation would nudge you to about 46 bins, and an information criterion to about the same.