library(tidyverse)
# Question 4 will also use some of the following. Load what you need.
# library(lmtest) # dwtest(), bgtest(), bptest(), resettest(), coeftest()
# library(sandwich) # NeweyWest(), vcovHC()
# library(tseries) # jarque.bera.test()
# library(strucchange) # sctest(), efp() for structural stabilityECON 4354/6354 — Homework 2
- Questions 1–3 are conceptual. Answer them in your own words, in complete sentences. Question 1 also asks for a short plot.
- Question 4 is an empirical exercise. Use Quarto/Knitr chunks for all work, and discuss what you find. A number without a sentence next to it earns no credit.
- The questions are taken from Diebold, Forecasting, Sections 2.12 (Problems 1, 5a, 6c) and 3.5 (Problem 1).
- Turn in the rendered HTML and your
.qmdsource via Canvas.
0. Setup
1. Properties of Loss Functions
Recall that we write a forecast error as \(e = y - \hat{y}\), and that a loss function \(L(e)\) maps that error into the cost it imposes on the decision maker.
For each of the following candidate loss functions, state whether it meets the criteria introduced in the text. If it does, state whether it is symmetric or asymmetric, and explain what kind of forecaster would want to use it. If it does not, say exactly which criterion fails and why.
\(L(e) = e^2 + e\)
\(L(e) = e^4 + 2e^2\)
\(L(e) = 3e^2 + 1\)
\(L(e) = \begin{cases} \sqrt{e} & \text{if } e > 0 \\[4pt] |e| & \text{if } e \le 0 \end{cases}\)
Plot the four candidates
Plot all four functions on the same grid of forecast errors, say \(e \in [-2, 2]\), so that you can see which criteria fail. A single faceted ggplot is the cleanest way to do this.
# TODO: build a data frame with a column of error values e, a column for the
# value of each candidate loss function, and plot the four panels.Which one would you use?
Of the candidates that are valid loss functions, which would you choose to evaluate forecasts of a firm’s quarterly sales? Why?
2. Univariate and Multivariate Information Sets
Which of the following modeling situations involve univariate information sets? Which involve multivariate information sets? In each case write down the information set \(\Omega_T\) explicitly, and justify your classification.
Using a stock’s price history to forecast its price over the next week.
Using a stock’s price history and volatility history to forecast its price over the next week.
Using a stock’s price history and volatility history to forecast its price and volatility over the next week.
Then answer the following: two of these three situations share exactly the same information set. Which two, and what is it that actually distinguishes them?
3. Assessing a Forecasting Situation
You work for D&D, a major Los Angeles advertising firm, and you must create an ad for a client’s product. The ad must be targeted toward teenagers, because they constitute the primary market for the product. You must (somehow) find out what kids currently think is “cool,” incorporate that information into your ad, and make your client’s product attractive to the new generation. If your hunch is right, your firm basks in glory, and you can expect multiple future clients from this one advertisement. If you miss, however, and the kids don’t respond to the ad, then your client’s sales fall and the client may reduce or even close its account with you.
Discuss this scenario under each of the following headings. Devote a short paragraph to each.
The decision environment. What decision does the forecast guide, who makes it, and when?
The object to be forecast. Is it an event outcome, an event timing, or a time series? Is it quantitative or qualitative?
The forecast statement. Point, interval, density, or probability forecast?
The forecast horizon. What sets it, and why does that matter here?
The loss function. Is it symmetric? Is it a function of the forecast error alone? What does its shape imply about the forecast you should report?
The information set. What is in it, and what is conspicuously not in it?
Methods and complexity. What simple approaches would you try, and would anything be gained from a complex one?
4. Regression, Regression Diagnostics, and Regression Graphics in Action
At the end of each quarter, you forecast a series \(y\) for the next quarter. You do this using a regression model that relates the current value of \(y\) to the lagged value of a single predictor \(x\). That is, you regress \[y_t \rightarrow c, \; x_{t-1}.\]
In this exercise, \(y_t\) is the annualized growth rate of U.S. real GDP in quarter \(t\), and \(x_t\) is the term spread in quarter \(t\): the yield on the 10-year Treasury bond minus the yield on the 3-month Treasury bill. The term spread is one of the oldest and best-known leading indicators of U.S. economic activity, so it is a natural candidate for \(x\).
The data
dat <- read_csv("data/hw2_gdp_spread.csv")
glimpse(dat)Rows: 262
Columns: 5
$ date <date> 1959-01-01, 1959-04-01, 1959-07-01, 1959-10-01, 1960-01-01…
$ quarter <chr> "1959 Q1", "1959 Q2", "1959 Q3", "1959 Q4", "1960 Q1", "196…
$ gdp <dbl> 3352.129, 3427.667, 3430.057, 3439.832, 3517.181, 3498.246,…
$ gdp_growth <dbl> 7.6009, 8.9137, 0.2788, 1.1383, 8.8949, -2.1592, 1.9549, -5…
$ spread <dbl> 1.2167, 1.2567, 0.9633, 0.3533, 0.6133, 1.2667, 1.4733, 1.5…
The file data/hw2_gdp_spread.csv holds 262 quarterly observations, 1959Q1 through 2024Q2:
| Variable | Description |
|---|---|
date |
first day of the quarter |
quarter |
the quarter as a label, e.g. 2008 Q4 |
gdp |
U.S. real GDP, billions of chained 2017 dollars (FRED series GDPC1) |
gdp_growth |
this is \(y_t\): annualized quarterly growth of real GDP in percent, \(400 \times \Delta \log(\text{gdp}_t)\) |
spread |
this is \(x_t\): term spread in percentage points, quarterly average of the monthly 10-year Treasury yield minus the 3-month bill rate (FRED-MD series GS10 and TB3MS) |
Your first job is to build XLAG1, the one-quarter lag of spread.
# TODO: create y (gdp_growth) and xlag1 (the one-quarter lag of spread), and
# drop the row that has no lagged value. How many observations are left?(a) Why a lagged right-hand-side variable?
Why might you include a lagged, rather than current, right-hand-side variable? Give at least two distinct reasons, and be precise about what is and is not in the information set at the moment the forecast is made.
(b) Graph Y vs. XLAG1
Graph y against xlag1 and discuss. A time series plot of each variable is also worth a look. What does the scatterplot suggest about the strength of the relationship, and is there anything in it that worries you?
# TODO: scatterplot of y vs. xlag1 with a fitted line, plus time plots of both
# series.(c) Regress Y on XLAG1
Regress y on xlag1 and discuss, including whatever regression diagnostics you deem relevant.
# TODO: fit the regression, report the output, and interpret it: the sign and
# size of the slope, its statistical significance, the R-squared, the SER.Interpret the estimated slope in economic terms: if the term spread widens by one percentage point this quarter, what does the model say about next quarter’s GDP growth? Is the \(R^2\) disappointing? Should it be?
(d) Variations on the theme
Consider as many variations as you deem relevant on the general theme. At a minimum, address each of the following. Support every answer with output, a graph, or a test, and say what you conclude.
Does it appear necessary to include an intercept in the regression?
Does the functional form appear adequate? Might the relationship be nonlinear?
Do the regression residuals seem completely random? If not, do they appear serially correlated, heteroskedastic, or something else?
Are there any outliers? If so, does the estimated model appear robust to their presence?
Do the regression disturbances appear normally distributed?
How might you assess whether the estimated model is structurally stable?
# TODO: work through (i)-(vi).Useful functions: lmtest::dwtest() and lmtest::bgtest() for serial correlation; feasts::ACF() for the residual autocorrelation function; lmtest::bptest() for heteroskedasticity; lmtest::resettest() for functional form; tseries::jarque.bera.test() and qqnorm() for normality; strucchange::sctest() for a Chow test and strucchange::efp() for recursive residuals.
Order matters. Look at the residual plot first: it will tell you which formal test to reach for, and it will sometimes tell you that a test statistic is not to be trusted.
(e) What would you actually forecast with?
Having worked through (a)–(d), write a short paragraph: would you use this regression to forecast next quarter’s GDP growth? If not, what would you change? Be concrete.