Once you know the parameters of interest and the meaning of the parameters of interest in terms of the research question you have, you can proceed with determining how to estimate or test claims about the parameters of interest. When we say “estimate”, we want to guess the values of the parameters of interest (or functions of these parameters of interest). This activity is called estimation. When we say “test claims”, we actually claim that the parameters of interest (or functions thereof) take some specific values. These specific values usually come from theory or from a wide accumulation of past experiences. This activity is called hypothesis testing.
Whether you estimate or test claims, you always have parameters in mind and your goal is to use a procedure with “desirable” properties. Unfortunately, any procedure used to estimate or test claims about parameters is subject to sampling variation. This means that your procedure produces a set of possible values of estimates. Some of these values are more likely than others to occur. Therefore, we can imagine the sampling distribution of your procedure.
When you estimate parameters, you would definitely want a procedure which allows you to get as close as you could to the actual unknown values of these parameters, especially when the sample size is very large. This desirable property is called consistency. In other words, the impact of the sampling variation becomes practically negligible if the sample size is large.
Note
To know what consistency means from a more formal point of view and to see a connection with what you learned before in ECON221, read the beginning part of Section C-3a in page 721.
You may see references to unbiasedness, but this is less of a priority nowadays. If you want to understand it, think of it as the center of the sampling distribution of your procedure being equal to the parameter of interest no matter how large or small your sample size is. Unfortunately, there are only very few procedures where unbiasedness can be proved. Even if we can prove unbiasedness, the conditions required are too strong.
Although consistency is desirable property of a procedure, we are unable to use it to describe how far off our procedure is in producing results close to the actual unknown values of the parameters. We need a way to describe the sampling distribution of the estimates more directly. Thus, a desirable property is to know the shape of this distribution (and name if possible) when the sample size is large. One version of this desirable property is called asymptotic normality.
If asymptotic normality can be demonstrated for a procedure, then it allows you to quantify the uncertainty associated with a procedure due to sampling variation. It also allows you to state with a high level of probability that your procedure is able to “capture” the parameters, even if you will never know their actual unknown values. This is called confidence. Finally, it allows you to test claims because you can decide how far is an observed estimate of an unknown parameter from a claim about this unknown parameter.
Note
To know what asymptotic normality means from a more formal point of view and to see a connection with what you learned before in ECON221, read the beginning part of Section C-3b.
The bottom line is that, at the very least, you should check whether the procedure you are using is consistent for the parameters of interest and whether the sampling distribution of the procedure could be known or approximated very well in large samples. If the approximation is via a standard normal distribution, then you are in the most typical of situations. If not, then you are in a relatively special situation which requires further study.
3.2 Key results for OLS
The good news is that, under some key assumptions, the procedure you know as OLS has many the two desirable properties of consistency and asymptotic normality.
Assumptions MLR.1 to MLR.4 (found in page 153 in Wooldridge) are enough to guarantee consistency for the \(\beta\)’s found in the conditional expectation of \(Y\) given the \(X\)’s.
Note
Refer to Theorem 5.1 in page 164 of Wooldridge.
It is important to note that the \(\beta\)’s the theorem is referring to are the \(\beta\)’s found in \[\mathbb{E}\left(Y|X_1,\ldots, X_k\right)=\beta_0+\beta_1X_1+\ldots+\beta_kX_k.\]
The key driving assumptions for the result are really MLR.1 and MLR.4. If you lose these assumptions, we have no guarantee of consistency of OLS right away.
It is possible to weaken MLR.2 (to cover non-random samples such as time series data, clustered data, etc.). It is possible to weaken MLR.3 because what is important is whether we can learn something unique at the population level.
Note
Chapter 5 of Wooldridge also allows a further weakening of MLR.4 in the case of random sampling (when MLR.2 holds).
Start from Assumption MLR.4’ in page 166 and read until the end of this page.
OLS is still consistent, but it is consistent for the \(\beta\)’s in the best linear approximation of the conditional expectation of \(Y\) given the \(X\)’s. In other words, when we apply OLS to a situation where MLR.1 and MLR.4 are not necessarily satisfied, the best we can obtain from OLS is that it becomes consistent for the \(\beta\)’s in the best linear approximation of the conditional expectation of \(Y\) given the \(X\)’s, whatever form this conditional expectation would take. This is an extremely subtle point which is sometimes lost among practitioners. The issue is whether the \(\beta\)’s in the best linear approximation to the conditional expectation actually contains answers to your research question.
The zero covariance conditions are precisely the conditions which give meaning to the \(\beta\)’s written earlier. A a result, if the zero covariance conditions found in Assumption MLR.4’ are not satisfied, then it is not surprising that OLS will fail to be consistent for those \(\beta\)’s.
3.3 The key expression for evaluating procedures
Note
Refer to page 165 of Wooldridge, specifically Equation 5.2.
To obtain this expression start from a simple linear regression model and express the regression slope \(\widehat{\beta}_1\) in a form which will lead to Equation (5.2). This is achieved by substituting the model into the regression slope and simplifying.
To evaluate procedures, we need to somehow express any estimator into a sum of two major parts: the actual unknown value the estimator is trying to recover and a portion that is due to model errors and sampling variation. Typically, you would need assumptions about the model and the origins of your measurements. In the context of OLS, it would give rise to Equation 5.2.
When sample size is large, the sampling variation would tend to stabilize leading to something similar to Equation 5.3. From there, you will find out what assumptions are needed to recover the actual unknown value any estimator is trying to recover.
Equation 5.2 is also the starting point for demonstrating asymptotic normality like in Theorem 5.2 of Wooldridge in page 170. Demonstrating asymptotic normality takes much more work.
To demonstrate all of these, the easiest case is regression with only an intercept. By now, you should know that in this case, \(\widehat{\beta}_0\) is nothing but the sample average of \(Y\). Here the linear regression model is \(Y=\beta_0+u\) and \(\mathbb{E}\left(u\right)=0\).
When you draw a sample from a distribution which obeys the previous linear regression model, then you would observe \(Y_1,Y_2,\ldots, Y_n\) and that \(Y_i=\beta_0+U_i\) for every \(i\). The sample average of \(Y\) could then be written as \[\widehat{\beta}_0=\overline{Y}=\beta_0+\frac{1}{n}\sum_{i=1}^n U_i.\]
When \(n\) is large enough, the sample average of the \(u_i\)’s will be very close (but not equal) to \(\mathbb{E}\left(U\right)\). So when \(n\) is large enough, \(\widehat{\beta}_0\) becomes hard to distinguish from \(\beta_0\). In this sense, we have learned \(\beta_0\) even if we do not actually know its actual value, under the assumptions we laid out.
When we look at the impact of sampling variation and its impact on \(\widehat{\beta}_0\), we need to determine its sampling distribution. For our purposes, it is enough to get a sense of the spread of this sampling distribution. Specifically, we wil calculate the variance (and as a consequence, the standard deviation) of the procedure \(\widehat{\beta}_0\).
If we calculate \(\mathsf{Var}\left(\widehat{\beta}_0\right)\), we will obtain from PROPERTY.VAR4 of page 700 that \[\mathsf{Var}\left(\widehat{\beta}_0\right)=\frac{1}{n^2}\mathsf{Var}\left(U_1\right)+\frac{1}{n^2}\mathsf{Var}\left(U_2\right)+\ldots+ \frac{1}{n^2}\mathsf{Var}\left(uU_n\right).\] But this result only works for pairwise uncorrelated random variables.
If we make an assumption similar to MLR.5 that all the \(U_i\)’s have a common variance (call this common variance \(\sigma^2\)), then \[\mathsf{Var}\left(\widehat{\beta}_0\right)=\frac{\sigma^2}{n},\] which should be familiar from ECON221.
If the assumption that all the \(U_i\)’s have a common variance is not true, then the earlier calculation is wrong and any software which relies on such a calculation may produce incorrect results. If the assumption of pairwise uncorrelated random variables is not true, then the calculation which relies on PROPERTY.VAR4 will also be incorrect. Specifically, we have to account for multiple covariances, i.e.,
Taking square roots of the expressions for the variances will give you the standard deviation of the sampling distribution of \(\widehat{\beta}_0\). Unfortunately, all these expressions involve unknown quantities which have to be estimated.
Thankfully, modern software allows easy estimation of these quantities and they are then called standard errors. If MLR.1 and MLR.4 (or MLR.4’) are not appropriate for your application, then using OLS will not even make sense and computing these standard errors will be nonsensical.
There are many types of standard errors:
If MLR.1 to MLR.5 are satisfied, then we have the usual standard errors or the homoscedasticity-only standard errors. They are available by default in lm().
If you lose MLR.5, then you have to use heteroscedasticity-consistent (HC) standard errors. There are more details in Chapter 8 of Wooldridge. These are not available by default in lm(). You need additional R packages for this.
If you lose MLR.2 and MLR.5, then you have to use heteroscedasticity and autocorrelation consistent (HAC) standard errors. There are more details in Chapter 12. These are not available by default in lm(). You need additional R packages for this.
3.4 Some demonstrations in R
3.4.1 Replicating Example 8.1
Note
Read Examples 7.6 and 8.1 found in pages 228 to 229 and 265 to 266 of Wooldridge, respectively.
library(wooldridge)options(digits=3)wage1$marrmale <- (wage1$married==1)*(wage1$female==0)wage1$marrfem <- (wage1$married==1)*(wage1$female==1)wage1$singfem <- (wage1$married==0)*(wage1$female==1)example8.1<-lm(log(wage) ~ marrmale + marrfem + singfem + educ + exper +I(exper^2) + tenure +I(tenure^2), data = wage1)# Compare standard errorslibrary(sandwich)cbind("Coef"=coef(example8.1), "Default SE"=sqrt(diag(vcov(example8.1))), "HC0 SE"=sqrt(diag(vcovHC(example8.1, type ="HC0"))), "HC3 SE"=sqrt(diag(vcovHC(example8.1, type ="HC3"))))
Read Example 8.2 found in pages 266 to 267 of Wooldridge.
library(wooldridge)options(digits=3)# Spring semester results OLSexample8.2<-lm(cumgpa ~ sat + hsperc + tothrs + female + black + white, data = gpa3, subset = (gpa3$term ==2))coef(example8.2)
(Intercept) sat hsperc tothrs female black
1.47006 0.00114 -0.00857 0.00250 0.30343 -0.12828
white
-0.05872
# Compare standard errorslibrary(sandwich)cbind("Coef"=coef(example8.2), "Default SE"=sqrt(diag(vcov(example8.2))), "HC0 SE"=sqrt(diag(vcovHC(example8.2, type ="HC0"))), "HC3 SE"=sqrt(diag(vcovHC(example8.2, type ="HC3"))))
Coef Default SE HC0 SE HC3 SE
(Intercept) 1.47006 0.229803 0.218560 0.229380
sat 0.00114 0.000179 0.000190 0.000195
hsperc -0.00857 0.001240 0.001404 0.001444
tothrs 0.00250 0.000731 0.000734 0.000749
female 0.30343 0.059020 0.058570 0.060040
black -0.12828 0.147370 0.118095 0.128188
white -0.05872 0.140990 0.110322 0.120435
# Test given null hypothesislibrary(car)
Loading required package: carData
# if MLR.5 is satisfiedlinearHypothesis(example8.2, c("black", "white"), rhs =0, test ="F", vcov. =vcov(example8.2))
Linear hypothesis test:
black = 0
white = 0
Model 1: restricted model
Model 2: cumgpa ~ sat + hsperc + tothrs + female + black + white
Note: Coefficient covariance matrix supplied.
Res.Df Df F Pr(>F)
1 361
2 359 2 0.68 0.51
# Iif MLR.5 is not satisfiedlinearHypothesis(example8.2, c("black", "white"), rhs =0, test ="F", vcov. =vcovHC(example8.2, type="HC0"))
Linear hypothesis test:
black = 0
white = 0
Model 1: restricted model
Model 2: cumgpa ~ sat + hsperc + tothrs + female + black + white
Note: Coefficient covariance matrix supplied.
Res.Df Df F Pr(>F)
1 361
2 359 2 0.75 0.47
# HC3 default in vcovHC will not match what you see in WooldridgelinearHypothesis(example8.2, c("black", "white"), rhs =0, test ="F", vcov. =vcovHC(example8.2))
Linear hypothesis test:
black = 0
white = 0
Model 1: restricted model
Model 2: cumgpa ~ sat + hsperc + tothrs + female + black + white
Note: Coefficient covariance matrix supplied.
Res.Df Df F Pr(>F)
1 361
2 359 2 0.67 0.51
# Chi-squared versionlinearHypothesis(example8.2, c("black", "white"), rhs =0, test ="Chisq", vcov. =vcovHC(example8.2, type ="HC0"))
Linear hypothesis test:
black = 0
white = 0
Model 1: restricted model
Model 2: cumgpa ~ sat + hsperc + tothrs + female + black + white
Note: Coefficient covariance matrix supplied.
Res.Df Df Chisq Pr(>Chisq)
1 361
2 359 2 1.5 0.47
3.4.3 Working on Example 8.3
Note
Read Example 8.3 found in page 268 of Wooldridge.
library(wooldridge)options(digits=3)example8.3<-lm(narr86 ~ pcnv + avgsen +I(avgsen^2) + ptime86 + qemp86 + inc86 + black + hispan, data = crime1) # comparing SEslibrary(sandwich)cbind("Coef"=coef(example8.3), "Default SE"=sqrt(diag(vcov(example8.3))), "HC0 SE"=sqrt(diag(vcovHC(example8.3, type ="HC0"))), "HC3 SE"=sqrt(diag(vcovHC(example8.3, type ="HC3"))))
# Test given null hypothesislibrary(car)# if MLR.5 is satisfiedlinearHypothesis(example8.3, c("avgsen", "I(avgsen^2)"), rhs =0, test ="F", vcov. =vcov(example8.3))
Linear hypothesis test:
avgsen = 0
I(avgsen^2) = 0
Model 1: restricted model
Model 2: narr86 ~ pcnv + avgsen + I(avgsen^2) + ptime86 + qemp86 + inc86 +
black + hispan
Note: Coefficient covariance matrix supplied.
Res.Df Df F Pr(>F)
1 2718
2 2716 2 1.73 0.18
# Iif MLR.5 is not satisfiedlinearHypothesis(example8.3, c("avgsen", "I(avgsen^2)"), rhs =0, test ="F", vcov. =vcovHC(example8.3, type="HC0"))