::: {.callout-smu-blue}
- Questions 1–3 are **conceptual**. Answer them in your own words, in complete sentences. Question 1 also asks for a short plot.
- Question 4 is an **empirical** exercise. Use Quarto/Knitr chunks for all work, and discuss what you find. A number without a sentence next to it earns no credit.
- The questions are taken from Diebold, *Forecasting*, Sections 2.12 (Problems 1, 5a, 6c) and 3.5 (Problem 1).
- Turn in the rendered HTML **and** your `.qmd` source via Canvas.
:::

## 0. Setup

```{r setup, message=FALSE, warning=FALSE}
library(tidyverse)

# Question 4 will also use some of the following. Load what you need.
# library(lmtest)       # dwtest(), bgtest(), bptest(), resettest(), coeftest()
# library(sandwich)     # NeweyWest(), vcovHC()
# library(tseries)      # jarque.bera.test()
# library(strucchange)  # sctest(), efp() for structural stability
```


## 1. Properties of Loss Functions

Recall that we write a forecast error as $e = y - \hat{y}$, and that a loss function $L(e)$ maps that error into the cost it imposes on the decision maker.

For each of the following candidate loss functions, state whether it meets the criteria introduced in the text. If it does, state whether it is symmetric or asymmetric, and explain what kind of forecaster would want to use it. If it does not, say exactly which criterion fails and why.

a. $L(e) = e^2 + e$

b. $L(e) = e^4 + 2e^2$

c. $L(e) = 3e^2 + 1$

d. $L(e) = \begin{cases} \sqrt{e} & \text{if } e > 0 \\[4pt] |e| & \text{if } e \le 0 \end{cases}$


### Plot the four candidates

Plot all four functions on the same grid of forecast errors, say $e \in [-2, 2]$, so that you can *see* which criteria fail. A single faceted `ggplot` is the cleanest way to do this.

```{r q1-plot}
# TODO: build a data frame with a column of error values e, a column for the
# value of each candidate loss function, and plot the four panels.
```

### Which one would you use?

Of the candidates that *are* valid loss functions, which would you choose to evaluate forecasts of a firm's quarterly sales? Why?


## 2. Univariate and Multivariate Information Sets

Which of the following modeling situations involve univariate information sets? Which involve multivariate information sets? In each case write down the information set $\Omega_T$ explicitly, and justify your classification.

i. Using a stock's price history to forecast its price over the next week.

ii. Using a stock's price history *and* volatility history to forecast its price over the next week.

iii. Using a stock's price history and volatility history to forecast its price *and* volatility over the next week.

Then answer the following: two of these three situations share exactly the same information set. Which two, and what is it that actually distinguishes them?


## 3. Assessing a Forecasting Situation

You work for D&D, a major Los Angeles advertising firm, and you must create an ad for a client's product. The ad must be targeted toward teenagers, because they constitute the primary market for the product. You must (somehow) find out what kids currently think is "cool," incorporate that information into your ad, and make your client's product attractive to the new generation. If your hunch is right, your firm basks in glory, and you can expect multiple future clients from this one advertisement. If you miss, however, and the kids don't respond to the ad, then your client's sales fall and the client may reduce or even close its account with you.

Discuss this scenario under each of the following headings. Devote a short paragraph to each.

a. **The decision environment.** What decision does the forecast guide, who makes it, and when?

b. **The object to be forecast.** Is it an event outcome, an event timing, or a time series? Is it quantitative or qualitative?

c. **The forecast statement.** Point, interval, density, or probability forecast?

d. **The forecast horizon.** What sets it, and why does that matter here?

e. **The loss function.** Is it symmetric? Is it a function of the forecast error alone? What does its shape imply about the forecast you should report?

f. **The information set.** What is in it, and what is conspicuously *not* in it?

g. **Methods and complexity.** What simple approaches would you try, and would anything be gained from a complex one?


## 4. Regression, Regression Diagnostics, and Regression Graphics in Action

At the end of each quarter, you forecast a series $y$ for the next quarter. You do this using a regression model that relates the current value of $y$ to the lagged value of a single predictor $x$. That is, you regress
$$y_t \rightarrow c, \; x_{t-1}.$$

In this exercise, $y_t$ is the annualized growth rate of U.S. real GDP in quarter $t$, and $x_t$ is the **term spread** in quarter $t$: the yield on the 10-year Treasury bond minus the yield on the 3-month Treasury bill. The term spread is one of the oldest and best-known leading indicators of U.S. economic activity, so it is a natural candidate for $x$.

### The data

```{r q4-data, message=FALSE}
dat <- read_csv("data/hw2_gdp_spread.csv")
glimpse(dat)
```

The file `data/hw2_gdp_spread.csv` holds 262 quarterly observations, 1959Q1 through 2024Q2:

| Variable | Description |
|:---|:---|
| `date` | first day of the quarter |
| `quarter` | the quarter as a label, e.g. `2008 Q4` |
| `gdp` | U.S. real GDP, billions of chained 2017 dollars (FRED series `GDPC1`) |
| `gdp_growth` | **this is $y_t$**: annualized quarterly growth of real GDP in percent, $400 \times \Delta \log(\text{gdp}_t)$ |
| `spread` | **this is $x_t$**: term spread in percentage points, quarterly average of the monthly 10-year Treasury yield minus the 3-month bill rate (FRED-MD series `GS10` and `TB3MS`) |

Your first job is to build `XLAG1`, the one-quarter lag of `spread`.

```{r q4-lag}
# TODO: create y (gdp_growth) and xlag1 (the one-quarter lag of spread), and
# drop the row that has no lagged value. How many observations are left?
```

### (a) Why a lagged right-hand-side variable?

Why might you include a *lagged*, rather than current, right-hand-side variable? Give at least two distinct reasons, and be precise about what is and is not in the information set at the moment the forecast is made.

### (b) Graph Y vs. XLAG1

Graph `y` against `xlag1` and discuss. A time series plot of each variable is also worth a look. What does the scatterplot suggest about the strength of the relationship, and is there anything in it that worries you?

```{r q4-graph}
# TODO: scatterplot of y vs. xlag1 with a fitted line, plus time plots of both
# series.
```

### (c) Regress Y on XLAG1

Regress `y` on `xlag1` and discuss, including whatever regression diagnostics you deem relevant.

```{r q4-reg}
# TODO: fit the regression, report the output, and interpret it: the sign and
# size of the slope, its statistical significance, the R-squared, the SER.
```

Interpret the estimated slope in economic terms: if the term spread widens by one percentage point this quarter, what does the model say about next quarter's GDP growth? Is the $R^2$ disappointing? Should it be?

### (d) Variations on the theme

Consider as many variations as you deem relevant on the general theme. At a minimum, address each of the following. Support every answer with output, a graph, or a test, and say what you conclude.

i. Does it appear necessary to include an intercept in the regression?

ii. Does the functional form appear adequate? Might the relationship be nonlinear?

iii. Do the regression residuals seem completely random? If not, do they appear serially correlated, heteroskedastic, or something else?

iv. Are there any outliers? If so, does the estimated model appear robust to their presence?

v. Do the regression disturbances appear normally distributed?

vi. How might you assess whether the estimated model is structurally stable?

```{r q4-diagnostics}
# TODO: work through (i)-(vi).
```

::: {.callout-tip}
Useful functions: `lmtest::dwtest()` and `lmtest::bgtest()` for serial correlation; `feasts::ACF()` for the residual autocorrelation function; `lmtest::bptest()` for heteroskedasticity; `lmtest::resettest()` for functional form; `tseries::jarque.bera.test()` and `qqnorm()` for normality; `strucchange::sctest()` for a Chow test and `strucchange::efp()` for recursive residuals.

Order matters. Look at the residual plot *first*: it will tell you which formal test to reach for, and it will sometimes tell you that a test statistic is not to be trusted.
:::

### (e) What would you actually forecast with?

Having worked through (a)–(d), write a short paragraph: would you use this regression to forecast next quarter's GDP growth? If not, what would you change? Be concrete.
