Showing posts with label Normal distribution. Show all posts
Showing posts with label Normal distribution. Show all posts

Monday, April 29, 2019

Recursions for the Moments of Some Continuous Distributions

This post follows on from my recent one, Recursions for the Moments of Some Discrete Distributions. I'm going to assume that you've read the previous post, so this one will be shorter. 

What I'll be discussing here are some useful recursion formulae for computing the moments of a number of continuous distributions that are widely used in econometrics. The coverage won't be exhaustive, by any means. I provide some motivation for looking at formulae such as these in the previous post, so I won't repeat it here. 

When we deal with the Normal distribution, below, we'll make explicit use of Stein's Lemma. Several of the other results are derived (behind the scenes) by using a very similar approach. So, let's begin by stating this Lemma.

Stein's Lemma (Stein, 1973):


"If  X ~ N[θ , σ2], and if g(.) is a differentiable function such that E|g'(X)| is finite, then 

                            E[g(X)(X - θ)] = σ2 E[g'(X)]."

It's worth noting that although this lemma relates to a single Normal random variable, in the bivariate Normal case the lemma generalizes to:


"If  X and Y follow a bivariate Normal distribution, and if g(.) is a differentiable function such that E|g'(Y)| is finite, then 

                            Cov.[g(Y ), X] = Cov.(X , Y) E[g'(Y)]."

In this latter form, the lemma is useful in asset pricing models.

There are extensions of Stein's Lemma to a broader class univariate and multivariate distributions. For example, see Alghalith (undated), and Landsman et al. (2013), and the references in those papers. Generally, if a distribution belongs to an exponential family, then recursions for its moments can be obtained quite easily.

Now, let's get down to business............


Tuesday, January 1, 2019

New Year Reading Suggestions for 2019

With a new year upon us, it's time to keep up with new developments -
  • Basu, D., 2018. Can we determine the direction of omitted variable bias of OLS estimators? Working Paper 2018-16, Department of Economics, University of Massachusetts, Amherst.
  • Jiang, B., Y. Lu, & J. Y. Park, 2018. Testing for stationarity at high frequency. Working Paper 2018-9, Department of Economics, University of Sydney. 
  • Psaradakis, Z. & M. Vavra, 2018. Normality tests for dependent data: Large-sample and bootstrap approaches. Communications in Statistics - Simulation and Computation, online.
  • Spanos, A., 2018. Near-collinearity in linear regression revisited: The numerical vs. the statistical perspective. Communications in Statistics - Theory and Methods, online.
  • Thorsrud, L. A., 2018. Words are the new numbers: A newsy coincident index of the business cycle. Journal of Business Economics and Statistics, online. (Working Paper version.)
  • Zhang, J., 2018. The mean relative entropy: An invariant measure of estimation error. American Statistician, online.
© 2019, David E. Giles

Sunday, September 10, 2017

Econometrics Reading List for September

A little belatedly, here is my September reading list:
  • Benjamin, D. J. et al., 2017. Redefine statistical significance. Pre-print.
  • Jiang, B., G. Athanasopoulos, R. J. Hyndman, A. Panagiotelis, and F. Vahid, 2017. Macroeconomic forecasting for Australia using a large number of predictors. Working Paper 2/17, Department of Econometrics and Business Statistics, Monash University.
  • Knaeble, D. and S. Dutter, 2017. Reversals of least-square estimates and model-invariant estimations for directions of unique effects. The American Statistician, 71, 97-105.
  • Moiseev, N. A., 2017. Forecasting time series of economic processes by model averaging across data frames of various lengths. Journal of Statistical Computation and Simulation, 87, 3111-3131.
  • Stewart, K. G., 2017. Normalized CES supply systems: Replication of Klump, McAdam and Willman (2007). Journal of Applied Econometrics, in press.
  • Tsai, A. C., M. Liou, M. Simak, and P. E. Cheng, 2017. On hyperbolic transformations to normality. Computational Statistics and Data Analysis, 115, 250-266,


© 2017, David E. Giles

Saturday, June 3, 2017

June Reading List

Here are some suggestions for you:
  • Ai, C. and E. C. Norton, 2003. Interaction terms in logit and probit models. Economics Letters, 80, 123-129.
  • Hirschberg, J. and J. Lye, 2017. Inverting the indirect - the ellipse and the Boomerang: Visualizing the confidence intervals of the structural coefficient from two-stage least squares. Journal of Econometrics, in press.
  • Kim, I. and S. Park, 2017. Likelihood ratio tests for multivariate normality. Communications in Statistics - Theory and Methods, in press.
  • Knotek, E. S. and S. Zaman, 2017. Financial nowcasts and their usefulness in macroeconomic forecasting. Working Paper 17-02, Federal Reserve Bank of Cleveland.
  • Marczak, M. and V. Goméz, 2017. Monthly US business cycle indicators: A new multivariate approach based on a band-pass filter. Empirical Economics, 52, 1379-1408.
  • Sherwood, C. and D. W. Kwak, 2017. New insights into an old problem - enhancing student learning outcomes in an introductory statistics course. Applied Economics, in press.
© 2017, David E. Giles

Tuesday, February 2, 2016

February Reading List

Here's a suggested reading list for February:
  • Casey, G. and M. Klemp, 2016. Instrumental variables in the long run. MPRA Paper No. 68696.
  • Coglianese, J., L. W. Davis, L. Kilian, and J. H. Stock, 2016. Anticipation, tax avoidance, and the price elasticity of gasoline demand. Journal of Applied Econometrics, in press.
  • Falorsi, S., A. Naccarato, and A. Pierini, 2015. Using Google trend data to predict the Italian unemployment rate. Working Paper No. 203, Dipartimento di Economia, Università degli studi Roma Tre.
  • Harris, D., S. J. Leybourne, and A. M. Robert, 2016. Test of the co-integration rank in VAR models in the presence of a possible break in trend at an unknown point. Working Paper No. 5 01-2016, Essex Finance Centre, Essex Business School, University of Essex.
  • Inoue, A. and G. Solon, 2010. Two-sample instrumental variables estimators. Review of Economics and Statistics, 93, 557-561.
  • Kim, N., 2016. A robustified Jarque-Bera test for multivariate normality. Economics Letters, in press.

© 2016, David E. Giles

Saturday, January 16, 2016

Why Does "Pi" Appear in the Normal Density

Every now and then a student will ask me why the formula for the density of a Normal random variable includes the constant, π, or more correctly (2π)-½.

The answer is that this term ensures that the density function is "proper" - that is, the integral of the function over the full real line takes the value "1". The area under the density, or "total probability", is "1".

Some students are happy with this (partial) answer, but others want to see a proof. Fair enough!

However, there's a trick to proving that this integral (area) is "1" in value. Let's take a look at it.

Friday, October 2, 2015

Illustrating Spurious Regressions

I've talked a bit about spurious regressions a bit in some earlier posts (here and here). I was updating an example for my time-series course the other day, and I thought that some readers might find it useful.

Let's begin by reviewing what is usually meant when we talk about a "spurious regression".

In short, it arises when we have several non-stationary time-series variables, which are not cointegrated, and we regress one of these variables on the others.

In general, the result that we get are nonsensical, and the problem is only worsened if we increase the sample size. This phenomenon was observed by Granger and Newbold (1974), and others, and Phillips (1986) developed the asymptotic theory that he then used to prove that in a spurious regression the Durbin-Watson statistic converges in probability to zero; the OLS parameter estimators and R2 converge to non-standard limiting distributions; and the t-ratios and F-statistic diverge in distribution, as T ↑ ∞ .

Let's look at some of these results associated with spurious regressions. We'll do so by means of a simple simulation experiment.

Tuesday, August 25, 2015

The Distribution of a Ratio of Correlated Normals

Suppose that the random variables X1 and X2 are jointly distributed as bivariate Normal, with means of θ1 and θ2, variances of σ12 and σ22 respectively, and a correlation coefficient of ρ.

In this post we're going to be looking at the distribution of the ratio, W = (X1 / X2).

You probably know that if X1 and X2 are independent standard normal variables, then W follows a Cauchy distribution. This will emerge as a special case in what follows.

The more general case that we're concerned with is of interest to econometricians for several reasons.

Thursday, July 9, 2015

'Student', on Kurtosis

W. S. Gosset (Student) provided this useful aid to help us remember the difference between platykurtic and leptokurtic distributions:


('Student', 1927. Errors of routine analysis. Biometrika, 19, 151-164. See p. 160.)

Here, β2 is the fourth standardized moment of the distribution about its mean. The Normal distribution has β2 = 3.

The appropriate definition of "kurtosis" for uni-modal distributions has been the subject of considerable discussion in the statistical literature. Should it be based on the characteristics of the tail of the distribution; the shape of the density around its mode; or both? 

© 2015, David E. Giles

Sunday, February 15, 2015

Testing for Multivariate Normality

The assumption that multivariate data are (multivariate) normally distributed is central to many statistical techniques. The need to test the validity of this assumption is of paramount importance, and a number of tests are available.

A recently released R package, MVN, by Korkmaz et al. (2014) brings together several of these procedures in a friendly and accessible way. Included are the tests proposed by Mardia, Henze-Zirkler, and Royston, as well as a number of useful graphical procedures.

If for some inexplicable reason you're not a user of R, the authors have thoughtfully created a web-based application just for you!


Reference

Korkmaz, S., D. Goksuluk, and G. Zarasiz
, 2014. An R package for assessing multivariate normality. The R Journal, 6/2, 151-162.


© 2015, David E. Giles

Thursday, December 11, 2014

Two Non-Problems!

I just love Dick Startz's "byline" on the EViews 9 Beta Forum: 

"Non-normality and collinearity are NOT problems!"

Why do I like it so much? Regarding "normality", see here, and here. As for "collinearity": see here, here, here, and  here

© 2014, David E. Giles

Thursday, December 4, 2014

More on Prediction From Log-Linear Regressions

My therapy sessions are actually going quite well. I'm down to just one meeting with Jane a week, now. Yes, there are still far too many log-linear regressions being bandied around, but I'm learning to cope with it!

Last year, in an attempt to be helpful to those poor souls I had a post about forecasting from models with a log-transformed dependent variable. I felt decidedly better after that, so I thought I follow up with another good deed.

Let's see if it helps some more:

Tuesday, November 11, 2014

Normality Testing & Non-Stationary Data

Bob Jensen emailed me about my recent post about the way in which the Jarque-Bera test can be impacted when temporally aggregated data are used. Apparently he publicized my post on the listserv for Accounting Educators in the U.S.. He also drew my attention to a paper from Two former presidents of the AAA: "Some Methodological Deficiencies in Empirical Research Articles in Accounting", by Thomas R. Dyckman and Stephen A. Zeff, Accounting Horizons, September 2014, 28 (3), 695-712. (Here.) 

Bob commented that an even more important issue might be that our data may be non-stationary. Indeed, this is always something that should concern us, and regular readers of this blog will know that non-stationary data, cointegration, and the like have been the subject of a lot of my posts.

In fact, the impact of unit roots on the Jarque-Bera test was mentioned in this old post about "spurious regressions". There, I mentioned a paper of mine (Giles, 2007) in which I proved that:

Friday, November 7, 2014

The Econometrics of Temporal Aggregation V - Testing for Normality

This post is one of a sequence of posts, the earlier members of which can be found here, here, here, and here. These posts are based on Giles (2014).

Some of the standard tests that we perform in econometrics can be affected by the level of aggregation of the data. Here, I'm concerned only with time-series data, and with temporal aggregation. I'm going to show you some preliminary results from work that I have in progress with Ryan Godwin. Although these results relate to just one test, our work covers a range of testing problems.

I'm not supplying the EViews program code that was used to obtain the results below - at least, not for now. That's because what I'm reporting is based on work in progress. Sorry!

As in the earlier posts, let's suppose that the aggregation is over "m" high-frequency periods. A lower case symbol will represent a high-frequency observation on a variable of interest; and an upper-case symbol will denote the aggregated series.

So,
               Yt = yt + yt - 1 + ......+ yt - m + 1 .

If we're aggregating monthly (flow) data to quarterly data, then m = 3. In the case of aggregation from quarterly to annual data, m = 4, etc.

Now, let's investigate how such aggregation affects the performance of the well-known Jarque-Bera (1987) (J-B) test for the normality of the errors in a regression model. I've discussed some of the limitations of this test in an earlier post, and you might find it helpful to look at that post (and this one) at this point. However, the J-B test is very widely used by econometricians, and it warrants some further consideration.

Consider the following small Monte Carlo experiment.

Monday, November 3, 2014

Central and Non-Central Distributions

Let's imagine that you're teaching an econometrics class that features hypothesis testing. It may be an elementary introduction to the topic itself; or it may be a more detailed discussion of a particular testing problem. We're not talking here about a course on Bayesian econometrics, so in all likelihood you'll be following the "classical" Neyman-Pearson paradigm.

You set up the null and alternative hypotheses. You introduce the idea of a test statistic, and hopefully, you explain why we try to find one that's "pivotal". You talk about Type I and Type II errors; and the trade-off between the probabilities of these errors occurring. 

You might talk about the idea of assigning a significance level for the test in advance of implementing it; or you might talk about p-values. In either case, you have to emphasize to the classt that in order to apply the test itself, you have to know the sampling distribution of your test statistic for the situation where the null hypothesis is true.

Why is this?

Monday, April 21, 2014

More On the Limitations of the Jarque-Bera Test

Testing the validity of the assumption, that the errors in a regression model are normally distributed, is a standard pastime in econometrics. We use this assumption when we construct standard confidence intervals  for, or test hypotheses about, the parameters of our models. In a post some time ago I pointed out that this assumption is actually is sufficient, but not necessary, for the validity of these inferences.

More recently, here and here, I discussed some aspects of the normality test that most econometricians use - the asymptotically valid test of Jarque and Bera (1987). Let's refer to this as the JB test. In the first of those posts I made brief mention of the finite-sample properties of the JB test, and I concluded:
"However, more recent evidence suggests that the power of the J-B test can be quite low in small samples, for a number of important alternative hypotheses - e.g., see Thadewald and Buning (2004). I'll address this aspect of the J-B test more fully in a later post."
The main results obtained by Thadewald and Buning are summed up in the abstract to their paper .............

Sunday, March 9, 2014

Testing for Multivariate Normality

In a recent post I commented on the connection between the multivariate normal distribution and marginal distributions that are normal. Specifically, the latter do not necessarily imply the former.

So, let's think about this in terms of testing for normality.

Suppose that we have several variables which we think may have a joint distribution that's normal. We could test each of the variables for normality, separately, perhaps using the Jarque-Bera LM test. If the null hypothesis of normality was rejected for one or more of the variables, this could be taken as evidence against multivariate normality. However, if normality couldn't be rejected for any of the variables, this wouldn't tell us anything about their joint distribution.

What we need is a test for multivariate normality itself. Let's see what's available.

Saturday, February 15, 2014

Some Things You Should Know About the Jarque-Bera Test

What test do you usually use if you want to test if the errors of your regression model are normally distributed? I bet it's the Jarque-Bera (1982, 1987) test. After all, it's a standard feature in pretty well every econometrics package. And with very good reason.

However, there some things relating to this test that you may not have learned in your econometrics courses. Let's take a look at them.

Friday, December 27, 2013

Unbiased Estimation of a Standard Deviation

Frequently, we're interested in using sample data to obtain an unbiased estimator of a population variance. We do this by using the sample variance, with the appropriate correction for the degrees of freedom. Similarly, in the context of a linear regression model, we use the sum of the squared OLS residuals, divided by the degrees of freedom, to get an unbiased estimator of the variance of the model's error term.

But what if we want an unbiased estimator of the population standard deviation, rather than the variance?

Thursday, December 19, 2013

Maximum Likelihood Estimation in EViews

This post is all about estimating regression models by the method of Maximum Likelihood, using EViews. It's based on a lab. class from one of my grad. econometrics courses.

We don't go through all of the material below in class - PART 3 is left as an exercise for the students to pursue in their own time.

The data and the EViews workfile can be found on the data page and the code page for this blog.

The purpose of this lab. exercise is to help the students to learn how to use EViews to estimate the parameters of a regression model by Maximum Likelihood, when the model is of some non-standard type. Specifically, find lout how to estimate models of types that are not “built in” as a standard option in EViews. This involves setting up the log-likelihood function for the model, based on the assumption of independent observations; and then maximizing this function numerically with respect to the unknown parameters. 

First, to introduce the concepts and commands that are involved, we consider the standard  linear multiple regression model with normal errors, for which we know that the MLE of the coefficient vector is just the same as the OLS estimator. This will give us a “bench-mark” against which to check our understanding of what is going on. Then we can move on to some more general models.