ECON 4354 / 6354 Computing for Economics

R Programming (Continued)

Zhan Gao

27 August 2026

Overview

  1. Functions
  2. Packages
  3. Control Flow
  4. Iteration
  5. Statistics
  6. R Markdown and Quarto

Functions


Why write functions?

In R, we do things using functions, as soon as we use it beyond a calculator.

  • Modularity: break a complex problem into small, manageable pieces. During development you focus on one function at a time instead of the whole project.
  • Scope management: functions provide local scope. Objects created inside a function stay inside it; only inputs and outputs touch the global environment.
  • Reusability, the DRY principle: “Don’t Repeat Yourself”. When something changes, you edit one definition instead of the same code in ten places.

Calling functions

We have already seen and used a multitude of functions in R.

  • Some come pre-packaged with base R (e.g. mean())
  • Others come from external packages (e.g. mvtnorm::rmvnorm())

Regardless of where they come from, functions in R all adopt the same basic syntax:

function_name(ARGUMENTS)

Writing your own functions

Most of the time we rely on functions that other people have written for us. But you can, and should, write your own too, with the generic function() function.1

function(ARGUMENTS) {
  OPERATIONS
  return(VALUE)
}

1 Yes, it is a function that lets you write functions. Very meta.

Name your functions

Writing anonymous functions like the above is possible and reasonably common. But we typically write functions because we want to reuse code, and for that it makes sense to name them.1

my_func =
  function(ARGUMENTS) {
    OPERATIONS
    return(VALUE)
  }

Short functions need neither the curly brackets nor an explicit return object, so you can write them on a single line:

my_short_func = function(ARGUMENTS) OPERATION

1 Remember: “In R, everything is an object and everything has a name.”

Name your functions (cont.)

Tip

Try to give your functions short, pithy names that are informative to both you and anyone else reading your code. This is harder than it sounds, but will pay off down the road.

Two things we still owe you, and which the next few slides deliver:

  • when the curly brackets and the explicit return() actually matter;
  • what happens to the objects created inside the brackets.

Simple example: addition

# Built-in function
sum(c(3, 4))
## [1] 7
# User-defined function
add_func <- function(x, y) {
  total <- x + y ## Create an intermediary object (that will be returned)
  return(total) ## The value(s) or object(s) that we want returned.
}
add_func(3, 4)
## [1] 7

Return explicitly

Caution

R’s default behaviour is to automatically return the final object that you created within the function. However, this won’t always be the case. Get into the habit of assigning the return object(s) explicitly.

An explicit return() is also how we hand back more than one object:

add_func <- function(x, y) {
  total <- x + y
  return(list(sum = total, first_element = x, second_element = y))
}
add_func(3, 4)
## $sum
## [1] 7
## 
## $first_element
## [1] 3
## 
## $second_element
## [1] 4

Return explicitly (cont.)

  • Multiple return objects have to be combined in a list.
  • Naming the list elements, i.e. “sum”, “first_element” and “second_element”, is optional, but helpful for users of our function.

For this simple example we could have written everything on a single line:

my_short_add_func <- function(x, y) x + y
my_short_add_func(3, 4)
## [1] 7

Specifying default argument values

You can also assign default argument values. You have encountered examples of this already.1

add_func <- function(x = 1, y = 1) {
  total <- x + y
  return(list(total, first_element = x, second_element = y))
}

1 E.g. type ?rnorm and see that it provides a default mean and standard deviation of 0 and 1, respectively.

Specifying default argument values (cont.)

add_func()
## [[1]]
## [1] 2
## 
## $first_element
## [1] 1
## 
## $second_element
## [1] 1
add_func(3)
## [[1]]
## [1] 4
## 
## $first_element
## [1] 3
## 
## $second_element
## [1] 1
add_func(3, 4)
## [[1]]
## [1] 7
## 
## $first_element
## [1] 3
## 
## $second_element
## [1] 4

Lexical scoping

None of the intermediate objects created inside the functions above (total, etc.) made their way into our global environment. Confirm this in the “Environment” pane of RStudio.

R has a set of lexical scoping rules governing where it stores and evaluates the values of different objects.

  • Functions operate in a quasi-sandboxed environment.
  • They do not return or use objects in the global environment unless forced to, e.g. by a return() command.
  • A function only looks to outside environments, e.g. a level “up”, if it does not see the object named within itself.

Example: OLS estimation

Consider a simple linear regression model with a constant term and one regressor:

y_i = \beta_0 + \beta_1 x_i + \varepsilon_i

  • y_i: the dependent variable
  • x_i: the independent variable
  • \beta_0: the intercept, \beta_1: the slope coefficient
  • \varepsilon_i: the error term for observation i

Example: OLS estimation (cont.)

In matrix notation:

\mathbf{y} = \mathbf{X}\boldsymbol{\beta} + \boldsymbol{\varepsilon}

  • \mathbf{y} is an n \times 1 vector of dependent variable observations
  • \mathbf{X} is an n \times 2 matrix: the first column all ones (for the intercept), the second the x_i values
  • \boldsymbol{\beta} = \begin{bmatrix} \beta_0 \\ \beta_1 \end{bmatrix} is a 2 \times 1 vector of parameters
  • \boldsymbol{\varepsilon} is an n \times 1 vector of error terms

The OLS estimator is given by

\hat{\boldsymbol{\beta}} = (\mathbf{X}' \mathbf{X})^{-1} \mathbf{X}'\mathbf{y}

To conduct OLS estimation in R, we literally translate this mathematical expression into code.

Step 1: simulate the data

We need data Y and X to run OLS, so we simulate an artificial dataset.

# simulate data
set.seed(111) # can be removed to allow the result to change

# set the parameters
n <- 100
b0 <- matrix(1, nrow = 2)

# generate the data
e <- rnorm(n)
X <- cbind(1, rnorm(n))
Y <- X %*% b0 + e

Step 1: simulate the data (cont.)

# Snapshot of X (first 10 rows):
print(head(X, 10))
##       [,1]       [,2]
##  [1,]    1  0.5996197
##  [2,]    1 -1.1603295
##  [3,]    1  0.4390934
##  [4,]    1  0.2048537
##  [5,]    1 -0.6991813
##  [6,]    1 -0.9266257
##  [7,]    1 -1.0134824
##  [8,]    1  0.6049871
##  [9,]    1  1.7344001
## [10,]    1 -0.3498505
# Snapshot of Y (first 10 elements):
print(head(Y, 10))
##             [,1]
##  [1,]  1.8348405
##  [2,] -0.4910654
##  [3,]  1.1274696
##  [4,] -1.0974919
##  [5,]  0.1299426
##  [6,]  0.2136525
##  [7,] -1.5109090
##  [8,]  0.5947986
##  [9,]  1.7859245
## [10,]  0.1561873

Step 2: translate the formula to code

\hat{\boldsymbol{\beta}} = (\mathbf{X}' \mathbf{X})^{-1} \mathbf{X}'\mathbf{y} becomes one line of R.

# OLS estimation
bhat <- solve(t(X) %*% X, t(X) %*% Y)
print(bhat)
##           [,1]
## [1,] 0.9861773
## [2,] 0.9404956
class(bhat)
## [1] "matrix" "array"

solve(A, b) solves Ax = b, which is numerically better behaved than forming the inverse with solve(A) %*% b.

Step 2: wrap it in a function

# User-defined function
ols_est <- function(X, Y) {
  bhat <- solve(t(X) %*% X, t(X) %*% Y)
  return(bhat)
}
bhat_2 <- ols_est(X, Y)
print(bhat_2)
##           [,1]
## [1,] 0.9861773
## [2,] 0.9404956
class(bhat_2)
## [1] "matrix" "array"

Step 2: or use a built-in function

# Use built-in functions
bhat_3 <- lsfit(X, Y, intercept = FALSE)$coefficients
print(bhat_3)
##        X1        X2 
## 0.9861773 0.9404956
class(bhat_3)
## [1] "numeric"

Same numbers, three routes. Note that the returned class differs.

Step 3: plot the regression

Scatter points, the fitted regression line (black) and the true coefficient line (red).

plot(y = Y, x = X[, 2], xlab = "X", ylab = "Y", main = "regression")
abline(a = bhat[1], b = bhat[2])
abline(a = b0[1], b = b0[2], col = "#CC0035")
abline(h = 0, lty = 2)
abline(v = 0, lty = 2)

Step 4: hypothesis testing

In econometrics we are often interested in hypothesis testing, and the t-statistic is widely used. To test H_0: \beta_1 = 1, again a translation:

t = \frac{\hat{\beta}_1 - \beta_{01}}{ \hat{\sigma}_{\hat{\beta}_1} } = \frac{\hat{\beta}_1 - \beta_{01}}{ \sqrt{ \left[ (X'X)^{-1} \hat{\sigma}^2 \right]_{22} } },

where [\cdot]_{22} is the (2,2)-element of a matrix.

Step 4: hypothesis testing (cont.)

# calculate the t-value
bhat2 <- bhat[2] # the parameter we want to test
e_hat <- Y - X %*% bhat
sigma_hat_square <- sum(e_hat^2) / (n - 2)
Sigma_B <- solve(t(X) %*% X) * sigma_hat_square
t_value_2 <- (bhat2 - b0[2]) / sqrt(Sigma_B[2, 2])
print(t_value_2)
## [1] -0.5615293

Exercise (in Homework 1): can you write a function with both \hat{\beta} and the t-value as outputs?

Accessing function source code

Looking inside a function is not only important for debugging, it is also a great way to pick up programming tips and tricks.

  • For simple functions: type the function name into the console without parentheses and let R print the object, e.g. dplyr::bind_rows.
  • This breaks down for functions with dispatch methods (S3/S4 classes) or compiled code underneath (C, Fortran).

You can still view the source, it just takes more legwork:

Packages


R’s package ecosystem

A base R installation is relatively compact. R’s true power lies in its extensive ecosystem of add-on packages, most of them hosted on CRAN, the Comprehensive R Archive Network.

A common practice in the statistical community: researchers publish an R package alongside the academic paper.

  • New statistical methods become immediately accessible to practitioners.
  • A virtuous cycle: researchers gain visibility, users get cutting-edge tools within months, sometimes weeks, of publication.

Install and load packages

A package is installed by install.packages("package_name") and invoked by library(package_name).

To install ggplot2, either:

  • Console: enter install.packages("ggplot2").
  • RStudio: click the “Packages” tab in the bottom-right pane, then “Install”, and search for the package.

Install and load packages (cont.)

Loading packages

Once installed, load a package into your R session with library().

library(ggplot2)

Note

Notice that you do not need quotes around the package name any more.

Reason: R now recognizes the package as a defined object with a given name. (“Everything in R is an object and everything has a name.”)

Shortcuts and namespaces

Tip

A convenient way to combine installation and loading is the pacman package’s p_load() function. pacman::p_load(ggplot2) first checks whether it needs to install the package, then loads it. Clever.

We can also run a function from an installed package without loading it, using PACKAGE::package_function():

stats::sd(1:10)
## [1] 3.02765

Packages beyond CRAN

Many authors distribute packages through GitHub or their personal websites, especially for cutting-edge research or development versions.

  • Install them with devtools or remotes.
  • Follow the instructions on the project’s repository, e.g. RDHonest.
if (!requireNamespace("remotes")) {
  install.packages("remotes")
}
remotes::install_github("kolesarm/RDHonest")

Build your own packages

Why build a package?

  • Share and reuse code across projects
  • Enforce conventions and structure
  • Improve reproducibility and collaboration

Core tools: RStudio Projects, usethis, devtools, roxygen2, testthat, knitr, styler.

The standard reference is Wickham and Bryan’s R Packages.

One-page workflow

  1. create_package() → open project
  2. use_git() → commit
  3. use_r() → write function(s) in R/
  4. roxygen comments → document()
  5. load_all() to try code
  1. use_testthat() → write tests → test()
  2. use_package() for deps (avoid library() in code)
  3. (Optional) use_vignette(), use_data()
  4. use_readme_rmd() → build_readme()
  5. check() → install() → build() → (optional) use_github()

Literate programming

A modern approach to developing R packages is through literate programming.

  • The litr package, by Jacob Bien, lets you write a complete R package inside a single R Markdown document.
  • The package is created when you knit the Rmd file.
  • For larger packages, you can write a bookdown that defines the package.

Interested? The videos show how it works in a few minutes.

Control Flow


Controlling the flow

Now that we have a good sense of the basic function syntax, it is time to learn control flow: controlling the order, or “flow”, of the statements and operations that our functions evaluate.

This is common to all programming languages.

  • if is used for choice.
  • for and while are used for loops.

if

The basic syntax for an if statement:

if (CONDITION) {
  OPERATION
}
square_root <- function(x) {
  if (x < 0) {
    message("x is negative")
    return(NA)
  }
  return(sqrt(x))
}
square_root(9)
## [1] 3
square_root(-9)
## x is negative
## [1] NA

if ... else

We can extend this with an else statement to handle alternative conditions:

square_root <- function(x) {
  if (x < 0) {
    message("x is negative")
    return(NA)
  } else {
    return(sqrt(x))
  }
}

ifelse()

For simple conditional operations, ifelse() provides a more concise syntax:

ifelse(CONDITION, DO IF TRUE, DO IF FALSE)

In our example, ifelse() gives a simpler version of square_root():

square_root <- function(x) {
  ifelse(x < 0, NA, sqrt(x))
}

A “gotcha” with ifelse()

Warning

The base R ifelse() function normally works great, and I use it all the time. But there are a couple of “gotcha” cases you should be aware of.

Consider this silly function, designed to return either today’s date or the day before:

print(Sys.Date())
## [1] "2026-08-27"
print(Sys.Date() - 1)
## [1] "2026-08-26"
today <- function(...) ifelse(..., Sys.Date(), Sys.Date() - 1)
today(TRUE)
## [1] 20692

A number, not a date!

A “gotcha” with ifelse() (cont.)

Why: ifelse() automatically converts date objects to numeric, as a way to get around some other type conversion strictures.

  • Confirm it for yourself by converting back the other way: as.Date(today(TRUE), origin = "1970-01-01").

Aside: the “dot-dot-dot” argument (...) used above is a convenient shortcut that allows users to enter unspecified arguments into a function.

  • Beyond the scope of this lecture, but an incredibly useful and flexible programming strategy.
  • See the relevant section of Advanced R.

Safer alternatives

To guard against this behaviour, and to add some other optimisations, both the tidyverse (through dplyr) and data.table offer their own versions of ifelse.

First, dplyr::if_else():

today2 <- function(...) {
  dplyr::if_else(..., Sys.Date(), Sys.Date() - 1)
}
today2(TRUE)
## [1] "2026-08-27"

Second, data.table::fifelse():

today3 <- function(...) {
  data.table::fifelse(..., Sys.Date(), Sys.Date() - 1)
}
today3(TRUE)
## [1] "2026-08-27"

Nested ifelse()

As you may have guessed, it is certainly possible to write nested ifelse() statements:

ifelse(CONDITION1, DO IF TRUE, ifelse(CONDITION2, DO IF TRUE, ifelse(...)))

or

if (CONDITION1) {
  DO IF TRUE
} else if (CONDITION2) {
  DO IF TRUE
} else {
  DO IF FALSE
}

However, these nested statements quickly become difficult to read and troubleshoot.

case when

A better solution was originally developed in SQL with the CASE WHEN statement. Both dplyr (case_when()) and data.table (fcase()) provide implementations in R.

x <- 1:10
## dplyr::case_when()
case_when(
  x <= 3 ~ "small",
  x <= 7 ~ "medium",
  TRUE ~ "big" ## Default value. Could also write `x > 7 ~ "big"` here.
)
##  [1] "small"  "small"  "small"  "medium" "medium" "medium" "medium" "big"   
##  [9] "big"    "big"
## data.table::fcase()
fcase(
  x <= 3, "small",
  x <= 7, "medium",
  default = "big" ## Default value. Could also write `x > 7, "big"` here.
)
##  [1] "small"  "small"  "small"  "medium" "medium" "medium" "medium" "big"   
##  [9] "big"    "big"

Iteration


Iteration

Alongside control flow, the most important early programming skill to master is iteration.

In particular, we want to write functions that can iterate, or map, over a set of inputs.

for loops

By far the most common way to iterate, across programming languages, is the for loop.

for (i in 1:3) {
  x <- rnorm(i)
  print(x)
}
## [1] -0.2986997
## [1] -0.5935356  0.8267038
## [1] -2.1614064 -0.9210425  0.5171631

Minimize the work inside a loop

When we use loops, we should minimize the tasks within the loop.

  • Pre-specify the dimension and allocate memory for variables outside the loop.
  • In cases where we want to “grow” an object via a for loop, we first have to create an empty (or NULL) object, and R re-allocates memory at every iteration.

Let us measure how much this actually costs.

Example: empirical coverage of a confidence interval

CI <- function(x) {
  # construct confidence interval
  # x is a vector of random variables
  n <- length(x)
  mu <- mean(x)
  sig <- sd(x)
  upper <- mu + 1.96 / sqrt(n) * sig
  lower <- mu - 1.96 / sqrt(n) * sig
  return(list(lower = lower, upper = upper))
}

Rep <- 100000
sample_size <- 1000
mu <- 2

Option 1: append a new outcome after each loop

pts0 <- Sys.time() # check time
for (i in 1:Rep) {
  x <- rpois(sample_size, mu)
  bounds <- CI(x)
  out_i <- ((bounds$lower <= mu) & (mu <= bounds$upper))
  if (i == 1) {
    out <- out_i
  } else {
    out <- c(out, out_i)
  }
}
mean(out)
## [1] 0.95023
cat("Takes", Sys.time() - pts0, "seconds\n")
## Takes 7.249099 seconds

Option 2: initialize the result vector

# Initialize the result vector
out <- rep(0, Rep)
pts0 <- Sys.time() # check time
for (i in 1:Rep) {
  x <- rpois(sample_size, mu)
  bounds <- CI(x)
  out[i] <- ((bounds$lower <= mu) & (mu <= bounds$upper))
}
mean(out)
## [1] 0.94893
cat("Takes", Sys.time() - pts0, "seconds\n")
## Takes 2.395516 seconds

Identical answer, and the empirical coverage sits close to the nominal 95%. Pre-allocating is the cheaper route.

Vectorization

Before we write any loops, we should ask ourselves: “Do I need to iterate at all?”

R is vectorized. Often you can apply a function to every element of a vector at once, rather than one at a time.

a_vec <- c(1, 2, 3)
b_vec <- c(4, 5, 6)

c_vec <- rep(NA, length(a_vec))
for (i in 1:length(a_vec)) {
  c_vec[i] <- a_vec[i] * b_vec[i]
}
c_vec

… or just a_vec * b_vec.

Why vectorize?

  • In R and other high-level languages, for loops can be slow. When you loop through many elements and perform an expensive operation on each one, your code may take a long time to execute.
  • Many operations typically implemented with for-loops can be rewritten using vectorized operations or matrix computations.
  • Whenever possible, leverage R’s built-in vectorization capabilities for better performance.

Example: 2-D random walk

Borrowed from Ross Ihaka’s online note.

  • A 2-D discrete random walk:
    • Start at the point (0, 0)
    • For t = 1, 2, \cdots, take a unit step in a randomly chosen direction
      • N, S, E, W

Naive implementation with a for loop

rw2d_loop <- function(n) {
  xpos <- rep(0, n)
  ypos <- rep(0, n)
  xdir <- c(TRUE, FALSE)
  step_size <- c(1, -1)
  for (i in 2:n) {
    if (sample(xdir, 1)) {
      xpos[i] <- xpos[i - 1] + sample(step_size, 1)
      ypos[i] <- ypos[i - 1]
    } else {
      xpos[i] <- xpos[i - 1]
      ypos[i] <- ypos[i - 1] + sample(step_size, 1)
    }
  }
  return(data.frame(x = xpos, y = ypos))
}

Vectorization without loops

rw2d_vec <- function(n) {
  xsteps <- c(-1, 1, 0, 0)
  ysteps <- c(0, 0, -1, 1)
  dir <- sample(1:4, n - 1, replace = TRUE)
  xpos <- c(0, cumsum(xsteps[dir]))
  ypos <- c(0, cumsum(ysteps[dir]))

  return(data.frame(x = xpos, y = ypos))
}

The trick: draw all n-1 directions in one call, then cumsum() the steps.

Comparison

n <- 100000
t0 <- Sys.time()
df <- rw2d_loop(n)
cat("Naive implementation: ", difftime(Sys.time(), t0, units = "secs"))
## Naive implementation:  0.329658
t0 <- Sys.time()
df2 <- rw2d_vec(n)
cat("Vectorized implementation: ", difftime(Sys.time(), t0, units = "secs"))
## Vectorized implementation:  0.001796007

Two orders of magnitude, for the same random walk.

Functional programming

Loops can run very slowly, and an inconspicuous for loop has been known to bring an entire analysis crashing to its knees.1

The bigger problem with for loops, however, is that they deviate from the norms and best practices of functional programming.

What is functional programming?

Functional programming (FP) is arguably the most important thing you can take away from today’s lecture. Hadley Wickham puts the core idea like this in Advanced R:

R, at its heart, is a functional programming (FP) language. […] R has what’s known as first class functions.

You can do anything with a function that you can do with a vector:

  • assign it to a variable,
  • store it in a list,
  • pass it as an argument to another function,
  • create it inside a function,
  • return it as the result of a function.

Hadley explains it better

Why not for loops?

Summary: for loops emphasise the objects we are working with (say, a vector of numbers) rather than the operations we want to apply to them (get the mean, or the median, or whatever).

  • This is inefficient: it requires us to write out the for loop by hand rather than getting an R function to create the loop for us.
  • As a corollary, for loops pollute our global environment with counting variables. Look at your “Environment” pane: i is sitting there, equal to the last value of its loop.

Creating those auxiliary variables is almost certainly not an intended outcome, and they can cause errors when we inadvertently refer to a similarly-named variable elsewhere in our script. So we best remove them as soon as we are finished.

rm(i)

Why not for loops? (cont.)

Another annoyance arrived in cases where we want to “grow” an object as we iterate over it. To do that with a for loop, we had to create an empty object first.

FP lets us avoid the explicit loop construct and its associated downsides. In practice, there are two ways to implement FP in R:

  1. The *apply family of functions in base R.
  2. The map*() family of functions from purrr.

Let’s explore these in more depth.

1) lapply()

Its syntax closely mimics the syntax of a basic for-loop.

# for(i in 1:10) print(LETTERS[i]) ## Our original for loop (for comparison)
lapply(1:10, function(i) LETTERS[i])
## [[1]]
## [1] "A"
## 
## [[2]]
## [1] "B"
## 
## [[3]]
## [1] "C"
## 
## [[4]]
## [1] "D"
## 
## [[5]]
## [1] "E"
## 
## [[6]]
## [1] "F"
## 
## [[7]]
## [1] "G"
## 
## [[8]]
## [1] "H"
## 
## [[9]]
## [1] "I"
## 
## [[10]]
## [1] "J"

Three things to notice

  1. No i in your global environment. Because of R’s lexical scoping rules, any object created and invoked by a function is evaluated in a sandbox outside your global environment.
  2. The syntax barely changed when switching from for() to lapply(). The essential structure is the same: first the iteration list (1:10), then the desired function or operation (LETTERS[i]).
  3. The returned object is a list. lapply() accepts vectors, data frames and lists as arguments, but always returns a list, with one element per iteration of the loop. (So now you know where the “l” in “lapply” comes from.)

Getting output that is not a list

Several options exist.1 The one I use most commonly is to bind the list elements into a single data frame with dplyr::bind_rows() or data.table::rbindlist().

lapply(1:10, function(i) {
  df <- tibble(num = i, let = LETTERS[i])
  return(df)
}) %>%
  bind_rows()
## # A tibble: 10 × 2
##      num let  
##    <int> <chr>
##  1     1 A    
##  2     2 B    
##  3     3 C    
##  4     4 D    
##  5     5 E    
##  6     6 F    
##  7     7 G    
##  8     8 H    
##  9     9 I    
## 10    10 J

1 E.g. pipe the output to unlist() if you want a vector, or use sapply(), which is covered next.

Why lapply() anyway?

The default list-return behaviour may not sound ideal at first, but I use lapply() more frequently than any of the other apply family members.

  • My functions normally return multiple objects of different type, which makes a list the only sensible format;
  • or they return a single data frame, which is where dplyr::bind_rows() and data.table::rbindlist() come in.

The *apply family: sapply()

sapply() stands for “simplify apply”. It is essentially a wrapper around lapply that tries to return simplified output that matches the input type. Feed it a vector and it will try to return a vector.

sapply(1:10, function(i) LETTERS[i])
##  [1] "A" "B" "C" "D" "E" "F" "G" "H" "I" "J"

The *apply family: apply()

apply() applies functions over array margins: the rows or columns of matrices and data frames.

# Create a simple matrix
mat <- matrix(1:12, nrow = 3, ncol = 4)
print(mat)
##      [,1] [,2] [,3] [,4]
## [1,]    1    4    7   10
## [2,]    2    5    8   11
## [3,]    3    6    9   12
apply(mat, 2, mean) ## apply mean over columns
## [1]  2  5  8 11
apply(mat, 1, mean) ## apply mean over rows
## [1] 5.5 6.5 7.5

For a concise overview of the different *apply() functions, see this blog post by Neil Saunders.

2) The purrr package

purrr’s map() is the tidyverse’s lapply(). Same syntax, same list output:

map(1:10, function(i) { ## only need to swap `lapply` for `map`
  df <- tibble(num = i, let = LETTERS[i])
  return(df)
})
## [[1]]
## # A tibble: 1 × 2
##     num let  
##   <int> <chr>
## 1     1 A    
## 
## [[2]]
## # A tibble: 1 × 2
##     num let  
##   <int> <chr>
## 1     2 B    
## 
## [[3]]
## # A tibble: 1 × 2
##     num let  
##   <int> <chr>
## 1     3 C    
## 
## [[4]]
## # A tibble: 1 × 2
##     num let  
##   <int> <chr>
## 1     4 D    
## 
## [[5]]
## # A tibble: 1 × 2
##     num let  
##   <int> <chr>
## 1     5 E    
## 
## [[6]]
## # A tibble: 1 × 2
##     num let  
##   <int> <chr>
## 1     6 F    
## 
## [[7]]
## # A tibble: 1 × 2
##     num let  
##   <int> <chr>
## 1     7 G    
## 
## [[8]]
## # A tibble: 1 × 2
##     num let  
##   <int> <chr>
## 1     8 H    
## 
## [[9]]
## # A tibble: 1 × 2
##     num let  
##   <int> <chr>
## 1     9 I    
## 
## [[10]]
## # A tibble: 1 × 2
##     num let  
##   <int> <chr>
## 1    10 J

map_df(): pick your output type

map() comes with its own variants for returning objects of a desired type. purrr::map_df() returns a data frame:

map_df(1:10, function(i) { ## don't need bind_rows with `map_df`
  df <- tibble(num = i, let = LETTERS[i])
  return(df)
})
## # A tibble: 10 × 2
##      num let  
##    <int> <chr>
##  1     1 A    
##  2     2 B    
##  3     3 C    
##  4     4 D    
##  5     5 E    
##  6     6 F    
##  7     7 G    
##  8     8 H    
##  9     9 I    
## 10    10 J

More efficient, i.e. less typing, than the lapply() version: no extra row-binding step at the end. Jenny Bryan’s purrr tutorial goes much deeper.

Create and iterate over named functions

We can split the function and the iteration (and binding) into separate steps. This is generally a good idea, since you typically create named functions with the goal of reusing them.

## Create a named function
num_to_alpha <-
  function(i) {
    df <- tibble(num = i, let = LETTERS[i])
    return(df)
  }

Create and iterate over named functions (cont.)

Now we can easily iterate over our function using different input values.

lapply(1:10, num_to_alpha) %>%
  bind_rows()
## # A tibble: 10 × 2
##      num let  
##    <int> <chr>
##  1     1 A    
##  2     2 B    
##  3     3 C    
##  4     4 D    
##  5     5 E    
##  6     6 F    
##  7     7 G    
##  8     8 H    
##  9     9 I    
## 10    10 J
map_df(c(1, 5, 26, 3), num_to_alpha)
## # A tibble: 4 × 2
##     num let  
##   <dbl> <chr>
## 1     1 A    
## 2     5 E    
## 3    26 Z    
## 4     3 C

R Markdown and Quarto


R Markdown

A document format that combines the power of R programming with the simplicity of Markdown syntax. Code, results, and narrative text coexist seamlessly.

R Markdown provides an authoring framework for data science. You can use a single R Markdown file to both

  • save and execute code
  • generate high quality reports that can be shared with an audience

Key features of R Markdown

  • Reproducible research: your analysis and results are automatically updated when you change your code or data.
  • Multiple output formats: HTML, PDF, Word documents, presentations and more, all from the same source.
  • Code integration: embed R code chunks that execute and display results inline with your text.
  • Version control friendly: the plain text format works well with Git and other version control systems.

Basic structure

An R Markdown document consists of three main components:

  1. YAML header: metadata and output options, enclosed by ---
  2. Markdown text: regular text formatted with Markdown syntax
  3. Code chunks: R code blocks enclosed by triple backticks with {r}

This deck, and the long-form note it mirrors, are themselves such documents.

Code chunks

Code chunks are the heart of R Markdown. They allow you to:

  • execute R code and display results;
  • control output with chunk options, e.g. echo=FALSE, eval=FALSE;
  • create plots, tables, and other visualizations;
  • cache results for faster compilation.

Getting started

Quarto

Quarto (.qmd) represents the next generation of R Markdown, building upon its foundation while introducing significant improvements.

  • Most R Markdown syntax remains compatible.
  • Enhanced features and better multi-language support, beyond just R.

For comprehensive documentation and tutorials, visit the Quarto website.