
Lecture 13: Trees and Random Forests
08 October 2026

tree() on motorcycle crashes and TV shows.Additional derivations and code follow the main lesson.
Run the examples from the lectures/ directory.
| File | Role |
|---|---|
MASS::mcycle |
Helmet acceleration after a simulated motorcycle crash, 133 readings |
data/nbc_showdetails.csv, data/nbc_demographics.csv |
40 TV shows: genre, ratings, engagement and audience shares |
data/prostate.csv |
97 prostate cancer patients |
data/CAhousing.csv |
20,640 California census block groups |
A tree is a binned sample mean whose bins are chosen by the data
In lecture 11 we predicted Y from a scalar X by
The number of bins was the complexity parameter, and cross-validation chose it.
Two things were unsatisfying.
With c = 4 cutpoints per covariate, the grid of bins grows as 5^d:
| Covariates d | Bins 5^d | Observations per bin when n = 20{,}640 |
|---|---|---|
| 1 | 5 | 4,128 |
| 2 | 25 | 826 |
| 5 | 3,125 | 7 |
| 9 | 1,953,125 | 0.01 |
Most bins are empty long before d reaches ten. We need bins that are adaptive: narrow only along the covariates and in the regions where Y actually changes, and wide elsewhere.
The lasso solved an analogous problem for linear models by selecting covariates. A regression tree is a way of selecting bins.
A decision tree is a nested sequence of if-else statements that takes an input and returns a conclusion. Here the input is the weather and the conclusion is whether to carry an umbrella.



Send the n observations down the tree. In each leaf, output a prediction from the observations that landed there:
A new X_{n+1} is dropped down the same tree; its leaf gives the prediction.
A regression tree is the binned sample mean with the bins given by the leaves. The tree decides which covariate to cut, where, and how finely.
The same construction handles classification, so the family is called CART, classification and regression trees. We say “regression tree” for both.


A deviance to minimize and a greedy search
Which tree? As with every estimator this semester, choose the one that minimizes a deviance, a loss function summed over the sample.
Minimizing the deviance over all trees is infeasible: the number of trees grows exponentially with the number of splits.
Instead, recursive binary splitting builds the tree one split at a time. For the first split, consider every covariate j and every observed value s of X_{ij}:
\text{left} = \{\ell : X_{\ell j} \le s\}, \qquad \text{right} = \{\ell : X_{\ell j} > s\},
predict by the mean (or the class shares) within each half, and compute the resulting deviance
\sum_{i \in \text{left}} (Y_i - \bar Y_{\text{left}})^2 + \sum_{i \in \text{right}} (Y_i - \bar Y_{\text{right}})^2.
Pick the (j, s) with the smallest deviance. Then repeat inside each child, using only its observations, and keep going.
mcycle records helmet acceleration (in g) against time (ms) after a simulated impact. Scan every candidate cutpoint s on times and record the deviance of the two-leaf tree.
split_deviance <- function(s) {
left <- mcycle$accel[mcycle$times <= s]
right <- mcycle$accel[mcycle$times > s]
sum((left - mean(left))^2) + sum((right - mean(right))^2)
}
cutpoints <- head(sort(unique(mcycle$times)), -1) # every observed time but the last
dev_path <- tibble(s = cutpoints, deviance = sapply(cutpoints, split_deviance))
dev_path |> slice_min(deviance, n = 1)## # A tibble: 1 × 2
## s deviance
## <dbl> <dbl>
## 1 27.2 200123.
## [1] 308222.7
The best single cut is between 27.2 and 27.6 ms and removes a third of the deviance.

Both halves are still far from flat. Recursive binary splitting now repeats the scan inside each half.
In tree() the stopping rules are
| Argument | Rule | Default |
|---|---|---|
mincut |
each child must contain at least this many observations | 5 |
minsize |
a node smaller than this is not split | 10 |
mindev |
a split must cut the deviance by at least mindev times the root deviance |
0.01 |
tree() on the motorcycle dataThe syntax is that of lm() and glm(). plot() draws the dendrogram and text() labels it.

Branch length is proportional to the deviance removed by the split, so the first cut at 27.4 ms is the long one.
## node), split, n, deviance, yval
## * denotes terminal node
##
## 1) root 133 308200.0 -25.550
## 2) times < 27.4 84 160500.0 -47.320
## 4) times < 16.5 43 18020.0 -16.480
## 8) times < 15.1 28 724.1 -4.357 *
## 9) times > 15.1 15 5494.0 -39.120 *
## 5) times > 16.5 41 58660.0 -79.660
## 10) times < 24.4 27 17040.0 -98.940
## 20) times < 19.5 15 9045.0 -86.310 *
## 21) times > 19.5 12 2616.0 -114.700 *
## 11) times > 24.4 14 12240.0 -42.490 *
## 3) times > 27.4 49 39670.0 11.780
## 6) times < 35 16 13300.0 29.290
## 12) times < 29.8 6 3900.0 10.250 *
## 13) times > 29.8 10 5919.0 40.720 *
## 7) times > 35 33 19080.0 3.291 *
grid <- data.frame(times = seq(0, 60, length.out = 1000))
grid$fit <- predict(collision, newdata = grid)
ggplot(mcycle, aes(times, accel)) + geom_point(alpha = 0.5) +
geom_line(data = grid, aes(times, fit), colour = smu_red, linewidth = 1.2) +
labs(x = "time since impact (ms)", y = "acceleration (g)")
The leaves are narrow where the acceleration swings and wide where it is flat, exactly the adaptivity that quantile bins lacked.
## [1] 40 56
## TERRITORY.EAST.CENTRAL TERRITORY.NORTHEAST COUNTY.SIZE.A COUNTY.SIZE.B
## Living with Ed 6 19 49 31
## Monarch Cove 10 14 35 32
## Top Chef 8 24 48 31
## Iron Chef America 13 25 47 30
## WIRED.CABLE.W.O.PAY DBS.OWNER
## Living with Ed 44 20
## Monarch Cove 40 29
## Top Chef 34 23
## Iron Chef America 30 26
Each row of demos is a show and each column is the percentage of its audience in a demographic cell: region, county size, cable subscription, household size, and so on. We will predict the show’s genre.
Make the outcome a factor, so that tree() fits class shares rather than a mean. With 40 shows we allow leaves of a single observation.
##
## Drama/Adventure Reality Situation Comedy
## 19 17 4
genre ~ . offers all 56 audience shares as candidate split variables; the tree will use four of them.Leaves are labelled with the most likely genre; text(genretree, label = "yprob") prints the three class shares instead. Basic-cable viewers watch reality shows, VCR owners watch drama.
## node), split, n, deviance, yval, (yprob)
## * denotes terminal node
##
## 1) root 40 75.800 Drama ( 0.47500 0.42500 0.10000 )
## 2) WIRED.CABLE.W.O.PAY < 28.6651 22 33.420 Drama ( 0.72727 0.09091 0.18182 )
## 4) VCR.OWNER < 83.749 5 6.730 Comedy ( 0.00000 0.40000 0.60000 ) *
## 5) VCR.OWNER > 83.749 17 7.606 Drama ( 0.94118 0.00000 0.05882 )
## 10) TERRITORY.EAST.CENTRAL < 16.4555 16 0.000 Drama ( 1.00000 0.00000 0.00000 ) *
## 11) TERRITORY.EAST.CENTRAL > 16.4555 1 0.000 Comedy ( 0.00000 0.00000 1.00000 ) *
## 3) WIRED.CABLE.W.O.PAY > 28.6651 18 16.220 Reality ( 0.16667 0.83333 0.00000 )
## 6) BLACK < 17.2017 15 0.000 Reality ( 0.00000 1.00000 0.00000 ) *
## 7) BLACK > 17.2017 3 0.000 Drama ( 1.00000 0.00000 0.00000 ) *
yprob lists the estimated class shares in the order of levels(genre): drama, reality, comedy. * marks a leaf.## Drama Reality Comedy
## Living with Ed 0 1 0
## Monarch Cove 1 0 0
## Top Chef 0 1 0
## [1] Reality Drama Reality Reality Reality Reality Reality Reality
## Levels: Drama Reality Comedy
For a new show, predict() returns its leaf’s class shares; type = "class" picks the largest.
Projected engagement (PE, 0 to 100) measures how much a focus group recalls of a pilot. Predict it from ratings (GRP) and genre.
## node), split, n, deviance, yval
## * denotes terminal node
##
## 1) root 40 5646.00 72.68
## 2) GRP < 223.05 7 1513.00 56.64 *
## 3) GRP > 223.05 33 1949.00 76.09
## 6) reality < 0.5 22 622.50 78.85
## 12) comedy < 0.5 19 513.80 78.01
## 24) GRP < 1545.15 12 250.20 75.98 *
## 25) GRP > 1545.15 7 129.80 81.48 *
## 13) comedy > 0.5 3 10.11 84.17 *
## 7) reality > 0.5 11 823.40 70.57
## 14) GRP < 433.85 3 145.00 63.13 *
## 15) GRP > 433.85 8 450.40 73.35 *

The first split separates the seven low-rated shows; above that, each genre gets its own concave-looking step function.
mincut = 5. Which cutpoints are admissible?mindev = 0.01 and a root deviance of 308,000, a node with deviance 19,000 has a best split that removes 950. Is it split?Grow a big tree, then cut it back
With mincut = 1 and mindev = 0, recursive splitting runs until every leaf holds one observation: the in-sample deviance is zero and the predictor is the nearest-neighbour step function, maximally overfit.
The obvious remedy is to treat mincut and mindev as complexity parameters and cross-validate them. Two things go wrong.
mincut alone forces leaves of similar size everywhere, which is the equal-count binning we wanted to escape.mindev stops when the best next split helps little. But the greedy algorithm lacks foresight: a split that is nearly useless on its own can be the one that makes the next split decisive.
Rather than stop early, overgrow the tree (mindev = 0, small mincut) and then prune it back:
Pruning does not suffer from the lack of foresight. Breiman, Friedman, Olshen and Stone (1984) show that the sequence contains, for each size, the subtree with the smallest deviance among all subtrees of that size.
The sequence is a set of candidate models indexed by the number of leaves, so cross-validation can choose among them, as it chose the number of bins and the penalty \lambda.
prune.tree(tree, best = m) returns the subtree with m leaves. The genre tree has five, so the sequence runs 5, 4, 3, 2, 1.

The first snip removes TERRITORY.EAST.CENTRAL < 16.4555, a split that separated one comedy from sixteen dramas.
Called without best, prune.tree() returns the whole sequence: sizes, in-sample deviances and the threshold k at which each subtree becomes optimal.
## size deviance k
## 1 5 6.73 -Inf
## 2 4 14.34 7.61
## 3 3 30.56 16.22
## 4 2 49.64 19.08
## 5 1 75.80 26.16
Each k is the deviance increase per leaf removed by the next snip: snipping node 5 costs 7.61 for one leaf, snipping node 3 costs 16.22, and so on. Larger k means a smaller tree.
Penalize the number of leaves, as the lasso penalized the coefficients. For \lambda \ge 0, among subtrees T of the big tree T_0 minimize
\sum_{i=1}^n \big(Y_i - \hat{\mathbf{E}}_T[Y \mid X_i]\big)^2 + \lambda \cdot |T|, \qquad |T| = \text{number of leaves of } T.
The minimizer for each \lambda is a member of the weakest-link sequence, and the k values are the \lambda’s at which the solution changes.
Choosing \lambda by K-fold cross-validation (James et al., Algorithm 8.1):
Ninety-seven patients; predict log cancer volume lcavol from age, log benign hyperplasia lbph, log capsular penetration lcp, Gleason score and log PSA lpsa.
##
## Regression tree:
## tree(formula = lcavol ~ ., data = prostate, mincut = 1, mindev = 0)
## Variables actually used in tree construction:
## [1] "lcp" "lpsa" "age" "lbph"
## Number of terminal nodes: 20
## Residual mean deviance: 0.2673 = 20.58 / 77
## Distribution of residuals:
## Min. 1st Qu. Median Mean 3rd Qu. Max.
## -1.38900 -0.27080 0.08845 0.00000 0.27050 1.36600
Twenty leaves for 97 patients, a few observations per leaf, and a residual deviance a sixth of the total. The printed tree runs to 39 lines.
cv.tree()cv.tree() implements the cross-validation of the previous slide and plot() draws the summed out-of-fold deviance against size (bottom axis) and k (top axis).
## [,1] [,2] [,3] [,4] [,5] [,6] [,7] [,8] [,9] [,10] [,11] [,12] [,13] [,14] [,15]
## size 20.0 19.0 18.0 17.0 15.0 14.0 13.0 12.0 11.0 8 7 6 5.0 4.0 3.0
## dev 72.1 72.1 72.1 72.1 72.1 72.1 72.1 72.1 73.4 74 72 73 71.2 71.2 69.8
## [,16] [,17]
## size 2.0 1.0
## dev 82.6 134.6
## [1] 3
cv.tree()cv.tree() grows each fold’s tree with the default stopping rules, whatever options built the tree you pass it. On the prostate data a default tree has about nine leaves, so a fold tree “pruned” at a small k is just the whole fold tree, evaluated again and again: that is the flat segment.
Warning
cv.tree() never evaluates subtrees larger than a default-grown tree. With mindev = 0 trees of hundreds of leaves, its curve goes flat at a few dozen leaves and the “best” size it reports is wrong.
The fix is to write the cross-validation loop ourselves, growing every fold’s tree with the same options as the full tree.
Ten lines reproduce Algorithm 8.1 and pass the growing options to every fold through ...:
cv_prune <- function(formula, data, K = 10, ...) {
full <- tree(formula, data = data, ...)
path <- prune.tree(full) # sizes and k of the full tree
fold <- sample(rep(1:K, length.out = nrow(data)))
dev <- 0
for (i in 1:K) {
t_i <- tree(formula, data = data[fold != i, ], ...) # same stopping rules as the full tree
dev <- dev + prune.tree(t_i, newdata = data[fold == i, ], k = path$k)$dev
}
data.frame(size = path$size, k = path$k, dev = dev)
}prune.tree(t_i, newdata = ..., k = path$k) prunes the fold tree at every k of the full tree’s path and evaluates each subtree on the held-out fold.
## [1] 3

Now the curve varies all the way up, and the minimum is still at three leaves. Here the two procedures agree; on the California tree they will not.

Three leaves summarize 97 patients: capsular penetration first, then PSA among the low-penetration patients. Age, hyperplasia and the Gleason score never enter the pruned tree, though all but the Gleason score appeared in the overgrown one.
k = 19.08 for the two-leaf subtree mean?cv.tree() and the curve is flat from 13 leaves up. What happened?Average many noisy trees instead of choosing one
In lecture 10 the bootstrap measured sampling uncertainty:
Bagging (Breiman, 1996) uses the same resamples to build a new estimator, their average:
\hat\beta_{\text{BAG}} = \frac{1}{B}\sum_{b=1}^B \hat\beta^*_b, \qquad \hat f_{\text{BAG}}(x) = \frac{1}{B}\sum_{b=1}^B \hat f^*_b(x).
For a regression tree, \hat f^*_b is the tree grown on resample b, and the bagged predictor is the average of B step functions.
newtimes <- data.frame(times = seq(0, 60, length.out = 1000))
B <- 100
set.seed(0)
boot_fits <- matrix(NA, B, nrow(newtimes))
for (b in 1:B) {
boot_sample <- mcycle[sample.int(nrow(mcycle), replace = TRUE), ]
boot_fits[b, ] <- predict(tree(accel ~ times, data = boot_sample), newdata = newtimes)
}
forest <- colMeans(boot_fits)Each resample grows its own tree, with its own cutpoints. The bagged predictor averages the 100 step functions pointwise.

A random forest (Breiman, 2001) is bagging applied to trees with one more source of randomness.
Forests do well without much tuning:
ranger’s default minimum node size is 5 for regression, 1 for classification), because overfitting a single tree is harmless once it is averaged;Adapted from Hastie, Tibshirani and Friedman (2009), Algorithm 15.1.
Prediction at x:
Median house value for 20,640 census block groups (the book says tracts; the count says block groups), with location, income, age and household composition. We make per-household variables and take logs.
CAhousing <- read.csv("data/CAhousing.csv")
CAhousing$AveBedrms <- CAhousing$totalBedrooms / CAhousing$households
CAhousing$AveRooms <- CAhousing$totalRooms / CAhousing$households
CAhousing$AveOccupancy <- CAhousing$population / CAhousing$households
logMedVal <- log(CAhousing$medianHouseValue)
CAhousing <- CAhousing[, -c(4, 5, 9)] # drop the room totals and the value in levels
CAhousing$logMedVal <- logMedVal
names(CAhousing)## [1] "longitude" "latitude" "housingMedianAge" "population"
## [5] "households" "medianIncome" "AveBedrms" "AveRooms"
## [9] "AveOccupancy" "logMedVal"
Nine covariates, so a forest’s default is m = \lfloor\sqrt{9}\rfloor = 3 candidates per split.

Value is a function of where you are: the Bay Area and the coast from Santa Barbara to San Diego are expensive, the Central Valley is not. A tree finds such pockets with a few splits on longitude and latitude; a linear model in longitude and latitude cannot.
ranger()## Ranger result
##
## Call:
## ranger(logMedVal ~ ., data = CAhousing, num.trees = 200, min.node.size = 5, importance = "impurity")
##
## Type: Regression
## Number of trees: 200
## Sample size: 20640
## Number of independent variables: 9
## Mtry: 3
## Target node size: 5
## Variable importance mode: impurity
## Splitrule: variance
## OOB prediction error (MSE): 0.05255512
## R squared (OOB): 0.8377498

Income and location dominate on both measures; the permutation measure, computed out of bag, puts latitude level with income. Household counts matter little once the per-household ratios are in.
Lasso. Standardize, interact everything with longitude and latitude (31 columns), and cross-validate.
## [1] 21
Pruned tree. Overgrow with leaves of at least ten block groups, then cross-validate the pruning path.

Income first and income again at both ends; latitude for the poorest block groups and the per-household ratios in the middle. The chosen tree continues for another 390 leaves, most of them splits on longitude and latitude.
Ten random splits: fit on 5,000 block groups, predict the other 15,640, compute the RMSE of log value. This is the code behind the opening figure.
set.seed(0)
compare <- tibble(rep = integer(), LASSO = numeric(), CART = numeric(), RF = numeric())
for (i in 1:10) {
train <- sample(1:nrow(CAhousing), 5000)
lin <- cv.glmnet(x = XXca[train, ], y = logMedVal[train], alpha = 1, lambda.min.ratio = 1e-4)
yhat_lin <- drop(predict(lin, XXca[-train, ]))
cvrt <- cv_prune(logMedVal ~ ., CAhousing[train, ], mindev = 0, mincut = 10)
rt <- tree(logMedVal ~ ., data = CAhousing[train, ], mindev = 0, mincut = 10)
yhat_rt <- predict(prune.tree(rt, best = cvrt$size[which.min(cvrt$dev)]), newdata = CAhousing[-train, ])
rf <- ranger(logMedVal ~ ., data = CAhousing[train, ], num.trees = 200, min.node.size = 5)
yhat_rf <- predict(rf, data = CAhousing[-train, ])$predictions
compare <- add_row(compare, rep = i, LASSO = rmse(logMedVal[-train], yhat_lin),
CART = rmse(logMedVal[-train], yhat_rt), RF = rmse(logMedVal[-train], yhat_rf))
}## # A tibble: 1 × 6
## LASSO_median LASSO_mean CART_median CART_mean RF_median RF_mean
## <dbl> <dbl> <dbl> <dbl> <dbl> <dbl>
## 1 0.775 0.819 0.316 0.314 0.253 0.253

## # A tibble: 3 × 4
## rep LASSO worst_lasso worst_occupancy
## <int> <dbl> <dbl> <dbl>
## 1 9 1.58 170. 1243.
## 2 10 1.44 145. 1243.
## 3 5 1.02 94.6 1243.
Could the lasso find cutpoints? Yes, with cleverly defined indicator variables (Wang, Sharpnack, Smola and Tibshirani, 2016, do this for trend filtering), but trees do it automatically.
mtry matter?set.seed(0)
MSE_mtry <- matrix(NA, 10, 9)
for (i in 1:10) {
train <- sample(1:nrow(CAhousing), 5000)
for (j in 1:9) {
rf <- ranger(logMedVal ~ ., data = CAhousing[train, ], num.trees = 200, min.node.size = 5, mtry = j)
MSE_mtry[i, j] <- rmse(logMedVal[-train], predict(rf, data = CAhousing[-train, ])$predictions)
}
}
par(mar = c(4, 4, 0.5, 0.5)); boxplot(MSE_mtry, col = "dodgerblue", xlab = "mtry", ylab = "out-of-sample RMSE")
mtry experimentmtry = 1 (a random covariate at every split) is clearly worse and mtry = 2 slightly so. From 3 to 9 the curve is flat within the noise of the ten splits, with the minimum at 4.mtry, including plain bagging at mtry = 9, beats the pruned tree and the lasso.mtry change, and why can a smaller mtry lower the error of the average even though it raises the error of each tree?A tree is an adaptive binned sample mean. Recursive binary splits choose which covariate to cut, where, and how finely; the leaves are the bins and the fit is a step function.
Growing is greedy, so do not stop early. Each split minimizes the deviance one step ahead. A near-useless split can precede a decisive one, so overgrow the tree and prune it back.
Prune by weakest link, choose by cross-validation. Cost-complexity pruning yields a nested path indexed by the number of leaves; cv.tree() evaluates it, but only up to the size of a default tree, so write the loop yourself for big trees.
Forests average many overgrown trees. Bagging removes the tree’s variance; random covariate subsets decorrelate the trees; the out-of-bag error comes free. The defaults usually work, and on the California data the forest beat both the pruned tree and the lasso in every split.
Trees are robust to the scale of X, linear models are not. A single extreme covariate value sank the lasso and left the trees untouched.
ranger package; the tree package by Ripley.Optional material for reference and practice
Let B tree predictions at a point be identically distributed with variance \sigma^2 and pairwise correlation \rho. The variance of their average is
\operatorname{Var}\Big(\frac{1}{B}\sum_{b=1}^B T_b\Big) = \rho\sigma^2 + \frac{1-\rho}{B}\sigma^2.

The second term vanishes as B grows; the first, \rho\sigma^2, is a floor that more trees cannot breach. The random covariate subset lowers \rho and so lowers the floor: that is the whole point of mtry (Hastie, Tibshirani and Friedman, section 15.2).
A bootstrap sample draws n times with replacement. The chance that a given observation is never drawn is
\Big(1 - \frac{1}{n}\Big)^n \;\longrightarrow\; e^{-1} \approx 0.368.
## exact limit
## 0.3678705 0.3678794
## [1] 0.3704942
Each tree sees about 63 percent of the observations and leaves 37 percent out, so with B = 200 every observation is out of bag for about 74 trees. The out-of-bag prediction for observation i averages those trees, and the resulting error is the OOB prediction error that ranger() printed (MSE 0.0526, R^2 0.838).
For a subtree T of the big tree T_0 with leaves m = 1, \dots, |T|, write the deviance R(T) = \sum_m \sum_{i \in R_m} (Y_i - \bar Y_m)^2 and the cost-complexity criterion R_\alpha(T) = R(T) + \alpha |T|.
The k column of prune.tree() is (\alpha_1, \alpha_2, \dots): in the genre path, g = 7.606 for node 5 (deviance 7.606, two pure children, one leaf removed), then 16.22 for node 3, then (33.42 - 14.34)/1 = 19.08 for node 2, then (75.80 - 49.64)/1 = 26.16 for the root.
For a node with class shares p_1, \dots, p_K, three measures of how mixed it is:
| Measure | Formula | In tree() |
|---|---|---|
| Misclassification error | 1 - \max_k p_k | not used for splitting |
| Gini impurity | \sum_k p_k(1 - p_k) | split = "gini" |
| Entropy (the deviance is 2n \times entropy) | -\sum_k p_k \log p_k | default |

Gini and entropy are strictly concave, so they reward a split that purifies the children even when the majority class does not change; misclassification error does not, which is why trees are not grown with it.
A tree with leaves R_1, \dots, R_M is the regression
\hat f(x) = \sum_{m=1}^M \hat c_m \, \mathbf{1}\{x \in R_m\}, \qquad \hat c_m = \text{mean of } Y_i \text{ in } R_m,
a linear model in M indicator variables whose definition is chosen by the data. Two consequences:
The lasso can mimic a tree if the analyst supplies the indicators: Wang et al. (2016) penalize differences of neighbouring fitted values on a grid, which selects cutpoints. Trees find the cutpoints themselves.
cv.tree() against the written-out loop on the California tree
## cv.tree cv_prune
## 13 397
The package curve is flat at 3,094 from 13 leaves to 1,596; the honest curve reaches 1,606 at 397 leaves, a very different tree.
Run once in your R environment if needed:
Data files, all read from lectures/data/:
nbc_showdetails.csv and nbc_demographics.csv, the 40 TV shows and their audience shares;prostate.csv, the 97 prostate cancer patients;CAhousing.csv, the 20,640 California block groups.MASS::mcycle ships with R. The notes use the same packages; randomForest is the older alternative to ranger.
Three predictors of California house prices: the lasso extrapolated off a single block group, the pruned tree needed cross-validation to find its 160 leaves, and the random forest, two hundred overgrown trees averaged, won every split with the defaults. Trees pick the bins; forests average the trees.