Showing posts with label Count data. Show all posts
Showing posts with label Count data. Show all posts

Sunday, October 27, 2019

Reporting an R-Squared Measure for Count Data Models

This post was prompted by an email query that I received some time ago from a reader of this blog. I thought that a more "expansive" response might be of interest to other readers............

In spite of its many limitations, it's standard practice to include the value of the coefficient of determination (R2) - or its "adjusted" counterpart - when reporting the results of a least squares regression. Personally, I think that R2 is one of the least important statistics to include in our results, but we all do it. (See this previous post.)

If the regression model in question is linear (in the parameters) and includes an intercept, and if the parameters are estimated by Ordinary Least Squares (OLS), then R2 has a number of well-known properties. These include:
  1. 0 ≤ R2 ≤ 1.
  2. The value of R2 cannot decrease if we add regressors to the model.
  3. The value of R2 is the same, whether we define this measure as the ratio of the "explained sum of squares" to the "total sum of squares" (RE2); or as one minus the ratio of the "residual sum of squares" to the "total sum of squares" (RR2).
  4. There is a correspondence between R2 and a significance test on all slope parameters; and there is a correspondence between changes in (the adjusted) R2 as regressors are added, and significance tests on the added regressors' coefficients.   (See here and here.)
  5. R2 has an interpretation in terms of information content of the data.  
  6. R2 is the square of the (Pearson) correlation (RC2) between actual and "fitted" values of the model's dependent variable. 
However, as soon as we're dealing with a model that excludes an intercept or is non-linear in the parameters, or we use an estimator other than OLS, none of the above properties are guaranteed.

Sunday, April 2, 2017

Read Some Econometrics this Month!

There are no April Fool's tricks in the following list of suggestions. 馃槓
© 2017, David E. Giles

Friday, January 22, 2016

Modelling With the Generalized Hermite Distribution

"Count" data occur frequently in economics. These are simply data where the observations are integer-valued - usually 0, 1, 2, ....... . However, the range of values may be truncated (e.g., 1, 2, 3, ....).

To model data of this form we typically resort to distributions such as the Poisson, negative binomial, or variations of these. These variations may account for truncation or censoring of the data, or the over-representation of certain count values (e.g., the "zero-inflated" Poisson distribution).

Covariates (explanatory variables) can be included into the model by making the mean of the distribution a function of these variables. After all, that's exactly what we do in a linear regression model.

If the "count" data form a time-series, then there are other issues that have to be taken into account.

However, the discrete distributions that we typically use have a number of limitations. The fact that the Poisson distribution is, of necessity, "equi-dispersed" (its variance equals its mean) is a big limitation. This leads us to consider distributions such as the negative binomial, in which he variance exceeds the mean. This enables us to model "over-dispersed" data, which are encountered frequently in practice.

The standard distributions are also limited in terms of what they can model in terms of distributional shapes. In particular, there are limitations on modal values in the data.

For instance, in the case of the Poisson distribution, these limitations are the following. If the parameter (位) of the Poisson distribution is an integer, then there are two adjacent modes with equal modal height, at x = 位 and x = 位-1. If lambda is non-integer, then there is a single mode at int(位), the integer part of 位.

In the case of the negative binomial distribution, there is a single mode.

This suggests that standard discrete distributions of the type that we typically use to mode l"count" data will not be very satisfactory if our data exhibit multi-modality.

We need to look to alternative distributions.

Here's an example of what I mean.

In an earlier post, I discussed some of my work involving the use of the so-called Hermite distribution, introduced by Kemp and Kemp (1965). As an example, I showed the distribution of data relating to the number of financial crises in various countries, as reproduced here:

You can see that, apart from being multi-modal, this empirical distribution is over-dispersed (its variance is approximately twice its mean).

In Giles (2010) I used the Hermite distribution, and various covariates, to model these data using maximum likelihood estimation.

The Hermite distribution can be generalized in various ways. Recently, Mori帽a et al. (2015) have released a terrific R package, called hermite, that makes it really easy to model "count data" using the Generalized Hermite distribution. We now have a convenient way of dealing with data that exhibit both over-dispersion and multi-modality.

I strongly recommend this new addition to R.


References

Giles, D. E., 2010. Hermite regression analysis of multi-modal count data. Economics Bulletin, 30(4), 2936–2945.

Kemp, C. D. and A. W. Kemp, 1965. Some properties of the ‘Hermite’ distribution. Biometrika, 52, 381-394.

Mori帽a, D,, M. Higueras, P. Puig, and M. Oliveira, 2015. Generalized Hermite distribution modelling with the R package hermite. The R Journal, 7(2), 263-274.  


© 2016, David E. Giles

Wednesday, April 1, 2015

April Reading

April 1 already - time to update your reading list. Here are some suggestions:


© 2015, David E. Giles

Monday, December 1, 2014

Here's Your Reading List!

As we count the year down, there's always time for more reading!
  • Birg, L. and A. Goeddeke, 2014. Christmas economics - A sleigh ride. Discussion Paper No. 220, CEGE, University of Gottingen.
  • Geraci, A., D. Fabbri, and C. Monfardini, 2014. Testing exogeneity of multinomial regressors in count data models: Does two stage residual inclusion work? Working Paper 14/03, Health, Econometrics and Data Group, University of York.
  • Li, Y. and D. E. Giles, 2014. Modelling volatility spillover effects between developed stock markets and Asian emerging stock markets. International Journal of Finance and Economics, in press.
  • Ma, J. and M. Wohar, 2014. Expected returns and expected dividend growth: Time to rethink an established literature. Applied Economics, 46, 2462-2476. 
  • Qin, D., 2014. Resurgence of instrument variable estimation and fallacy of endogeneity. Economics Discussion Papers No. 2014-42, Kiel Institute for the World Economy. 
  • Romano, J. P. and M. Wolf, 2014. Resurrecting weighted least squares. Working Paper No. 172, Department of Economics, University of Zurich.
  • Tchatoka, F.D., 2014. Specification tests with weak and invalid instruments. Working Paper No. 2014-05, School of Economics, University of Adelaide.

© 2014, David E. Giles

Friday, August 29, 2014

September Reading List

In North America, Labo(u)r Day weekend is upon us. The end of summer. Back to school. Last chance to get some pre-class reading done!

  • Blackburn, M. L., 2014. The relative performance of Poisson and negative binomial regression estimators. Oxford Bulletin of Economics and Statistics, in press.
  • Giannone, D., M. Lenza, and G. E. Primiceri, 2014. Prior selection for vector autoregressions. Review of Economics and Statistics, in press.
  • Gulesserian, S. G. and M. Kejriwal, 2014. On the power of bootstrap tests for stationarity: A Monte Carlo comparison. Empirical Economics, 46, 973-998.
  • Elliot, G. and A. Timmerman, 2008. Economic forecasting. Journal of Economic Literature, 46, 3-56.
  • Kiviet, J. F., 1986, On the rigour of some misspecification tests for modelling dynamic relationships. Review of Economic Studies, 53, 241-261.
  • Otto, G. D. and G. M. Voss, 2014. Flexible inflation forecast targeting: Evidence from Canada. Canadian Journal of Economics, 47, 398-421. 

© 2014, David E. Giles

Monday, February 10, 2014

Modelling Olympic Medal Wins

With the Sochi Winter Olympics now well underway, I was reminded of the empirical literature that has attempted to model the number of medals that different countries win.

Back in 2006, one of our students, Glen Roberts, wrote an excellent paper, titled "Accounting for Achievement in Athens: A Count Data Analysis of National Olympic Performance", on this topic. Glen's work was based on a term project that he undertook for my ECON 546, "Themes in Econometrics" course. This is an elective course for M.A. students, and it emphasises the thematic content of econometric methods - MLE, IV/GMM, Bayesian inference, etc.

Glen found that the empirical analyses of Olympic medal wins largely ignored the "count data" aspect of the problem. You can find plenty of references in Glen's paper. He then set about rectifying this situation, as the abstract to his paper describes:

Saturday, January 11, 2014

Reading for the New Year

Back to work, and back to reading:
  • Basturk, N., C. Cakmakli, S. P. Ceyhan, and H. K. van Dijk, 2013. Historical developments in Bayesian econometrics after Cowles Foundation monographs 10,14. Discussion Paper 13-191/III, Tinbergen Institute.
  • Bedrick, E. J., 2013. Two useful reformulations of the hazard ratio. American Statistician, in press.
  • Nawata, K. and M. McAleer, 2013. The maximum number of parameters for the Hausman test when the estimators are from different sets of equations.  Discussion Paper 13-197/III, Tinbergen Institute.
  • Shahbaz, M, S. Nasreen, C. H. Ling, and R. Sbia, 2013. Causality between trade openness and energy consumption: What causes what  high, middle and low income countries. MPRA Paper No. 50832. 
  • Tibshirani, R., 2011. Regression shrinkage and selection via the lasso: A retrospective. Journal of the Royal Statistical Society, B, 73, 273-282.
  • Zamani, H. and N. Ismail, 2014. Functional form for the zero-inflated generalized Poisson regression model. Communications in Statistics - Theory and Methods, in press.


© 2014, David E. Giles

Friday, July 19, 2013

Some Current Projects

I often get emails asking me what research projects I'm working on. Generally, I have several projects underway at any given time - usually at various stages of development or completion. In that respect I guess I'm pretty typical.

I also tend to have a mixture of theoretical and applied projects, some econometric and some essentially statistical in nature. I find that this provides some continuity in my work. It's not easy to focus on just one or two research projects all of the time, especially if they're not progressing as well as you'd like them to!

So, what am I up to right now? Here are some of the papers/projects that I'm working on:

Friday, July 5, 2013

Paper With Jacob Schwartz

It was nice to get the final "acceptance" yesterday for a paper co-authored with former grad. student, Jacob Schwartz.

The paper, titled "Bias-Reduced Maximum Likelihood Estimation of the Zero-Inflated Poisson Distribution", and with Jacob as lead author, will appear in Communications in Statistics - Theory & Methods. You can download a copy of the paper from here.

Jacob has been in the Ph.D. program at UBC for a while now. It seems quieter around the computing lab. without him!


© 2013, David E. Giles

Monday, June 3, 2013

Last Week's Reading

There are some great econometrics papers out there, just waiting to be read. I need more hours in the day!

Some of the papers I enjoyed reading last week were:
  • Barsoum, F. and S. Stankiewicz, 2013. Forecasting GDP growth using mixed-frequency models with switching regimes. Department of Economics, University of Konstanz, Working Paper 2013-10.
  • Castle, J. L., M. P. Clements, and D. F. Hendry, 2013. Forecasting by factors, by variables, by both or neither? Journal of Econometrics, in press.
  • Chiu, C. W., B. Eraker, A. T. Foerster, T. B. Kim, and H. D. Seoane, 2012. Estimating VAR's sampled at mixed or irregular spaced frequencies: A Bayesian approach. Federal Reserve Bank of Kansas City, Research Working Paper 11-11 (revised, December 2012).
  • Dufour, J-M. and J. Wilde, 2013. Weak identification in probit models with endogenous covariates.
  • Lim, H. K., J. Song, and B. C. Jung, 2013. Score tests for zero-inflation and overdispersion in two-level count data. Computational Statistics and Data Analysis, 61, 67-82.
  • Millimet, D. L. and I. K. McDonough, 2013. Dynamic panel data models with irregular spacing: With applications to early childhood development. IZA Discussion Paper 7359.
  • Pesaran, H. H., A. Pick, and M. Pranovich, 2013. Optimal forecasts in the presence of structural breaks. Journal of Econometrics, in press.

© 2013, David E. Giles

Monday, March 4, 2013

Measuring the Quality of an Estimator


In which, with almost no symbols, I encourage students and practitioners to question what they've been taught............

When it comes to introducing our students to the notion of the "quality" of an estimator, most of us begin by observing that estimators are functions of the random sample data, and hence they are "statistics" in the literal sense. As such, estimators have a probability distribution. We give this distribution a special name - the "sampling distribution" of the estimator in question.

It's understandable that students sometimes find the concept of the sampling distribution a little tricky when they first encounter it. After all, it's based on a "thought game" of sorts. We have to consider the idea of repeatedly drawing samples of a fixed size, for ever, constructing the statistic in question, and then keeping track of all of the possible values that the statistic can take, together with the relative frequency of occurrence for each value. A Monte Carlo experiment is the obvious way to introduce students to this concept.


Friday, August 24, 2012

Analysing Olympic Medal Data

So, the London Olympics are over - with the Paralympics still to come, of course. Sports, and events such as the Olympic Games, generate lots of lovely data. It's also usually "hard" data. So, there's a cottage industry out there comprised of statisticians of all shapes and forms who love to work sports data.

The American Statistical Association has a Section for Statistics in Sport, publishes the Journal of Quantitative Analysis in Sports, and provides access to some interesting sports data-sets.

Friday, April 13, 2012

Count Data & the Hermite Distribution

One of the limitations of the usual discrete distributions that we use when modeling "count data" is that they can't allow for multi-modality (except in a trivial manner). So, there's no use in trying to model multi-modal data using a Poisson regression  model, or a Negative Binomial regression model, for example.

However, such data occur frequently in practice. So, what options are open to us?