Showing posts with label Big data. Show all posts
Showing posts with label Big data. Show all posts

Wednesday, October 30, 2019

Everything's Significant When You Have Lots of Data

Well........, not really!

It might seem that way on the face of it, but that's because you're probably using a totally inappropriate measure of what's (statistically) significant, and what's not.

I talked a bit about this issue in a previous post, where I said:
"Granger (1998, 2003) has reminded us that if the sample size is sufficiently large, then it's virtually impossible not to reject almost any hypothesis. So, if the sample is very large and the p-values associated with the estimated coefficients in a regression model are of the order of, say, 0.10 or even 0.05, then this really bad news. Much, much, smaller p-values are needed before we get all excited about 'statistically significant' results when the sample size is in the thousands, or even bigger."
This general point, namely that our chosen significance level should be decreased as the sample size grows, is pretty well understood by most statisticians and econometricians. (For example, see Good, 1982.) However, it's usually ignored by the authors of empirical economics studies based on samples of thousands (or more) observations. Moreover, a lot of practitioners seem to be unsure of just how much they should revise their significance levels (or re-interpret their p-values) in such circumstances.

There's really no excuse for this, because there are some well-established guidelines to help us. In fact, as we'll see, some of them have been around since at least the 1970's.

Let's take a quick look at this, because it's something that all students need to be made aware of as we work more and more with "big data". Students certainly won't gain this awareness by looking at the  interpretation of the results in the vast majority of empirical economics papers that use even sort-of-large samples!

Sunday, January 13, 2019

Machine Learning & Econometrics

What is Machine Learning (ML), and how does it differ from Statistics (and hence, implicitly, from Econometrics)?

Those are big questions, but I think that they're ones that econometricians should be thinking about. And if I were starting out in Econometrics today, I'd take a long, hard look at what's going on in ML.

Here's a very rough answer - it comes from a post by Larry Wasserman on his (now defunct) blog, Normal Deviate:
"The short answer is: None. They are both concerned with the same question: how do we learn from data?
But a more nuanced view reveals that there are differences due to historical and sociological reasons.......... 
If I had to summarize the main difference between the two fields I would say: 
Statistics emphasizes formal statistical inference (confidence intervals, hypothesis tests, optimal estimators) in low dimensional problems. 
Machine Learning emphasizes high dimensional prediction problems. 
But this is a gross over-simplification. Perhaps it is better to list some topics that receive more attention from one field rather than the other. For example: 
Statistics: survival analysis, spatial analysis, multiple testing, minimax theory, deconvolution, semiparametric inference, bootstrapping, time series.
Machine Learning: online learning, semisupervised learning, manifold learning, active learning, boosting. 
But the differences become blurrier all the time........ 
There are also differences in terminology. Here are some examples:
Statistics       Machine Learning
———————————–
Estimation        Learning
Classifier          Hypothesis
Data point         Example/Instance
Regression        Supervised Learning
Classification    Supervised Learning
Covariate          Feature
Response          Label 
Overall, the the two fields are blending together more and more and I think this is a good thing."
As I said, this is only a rough answer - and it's by no means a comprehensive one.

For an econometrician's perspective on all of this you can't do better that to take a look at Frank Dielbold's blog, No Hesitations. If you follow up on his posts with the label "Machine Learning" - and I suggest that you do - then you'll find 36 of them (at the time of writing).

If (legitimately) free books are your thing, then you'll find some great suggestions for reading more about the Machine Learning / Data Science field(s) on the KDnuggets website - specifically, here in 2017 and here in 2018.

Finally, I was pleased that the recent ASSA Meetings (ASSA2019) included an important contribution by Susan Athey (Stanford), titled "The Impact of Machine Learning on Econometrics and Economics". The title page for Susan's presentation contains three important links to other papers and a webcast.

Have fun!

© 2019, David E. Giles

Thursday, November 22, 2018

A New Canadian Macroeconomic Database

Anyone who's undertaken empirical macroeconomic research relating to Canada will know that there are some serious data challenges that have to be surmounted.

In particular, getting access to long-term, continuous, time series isn't as easy as you might expect.

Statistics Canada has been criticized frequently over the years by researchers who find that crucial economic series are suddenly "discontinued", or are re-defined in ways that make it extremely difficult to splice the pieces together into one meaningful time-series.

In recognition of these issues, a number of efforts have been made to provide Canadian economic data in forms that researchers need. These include, for instance, Boivin et al. (2010), Bedock and Stevanovic (2107), and Stephen Gordon's on-going "Project Link".

Thanks to Olivier Fortin-Gagnon, Maxime Leroux, Dalibor Stevanovic, &and Stéphane Suprenant we now have an impressive addition to the available long-term Canadian time-series data. Their 2018 working paper, "A Large Canadian Database for Macroeconomic Analysis", discusses their new database and illustrates its usefulness in a variety of ways.

Here's the abstract:
"This paper describes a large-scale Canadian macroeconomic database in monthly frequency. The dataset contains hundreds of Canadian and provincial economic indicators observed from 1981. It is designed to be updated regularly through (the) StatCan database and is publicly available. It relieves users to deal with data changes and methodological revisions. We show five useful features of the dataset for macroeconomic research. First, the factor structure explains a sizeable part of variation in Canadian and provincial aggregate series. Second, the dataset is useful to capture turning points of the Canadian business cycle. Third, the dataset has substantial predictive power when forecasting key macroeconomic indicators. Fourth, the panel can be used to construct measures of macroeconomic uncertainty. Fifth, the dataset can serve for structural analysis through the factor-augmented VAR model."
Note - these are monthly data! And they're freely available. Although the paper doesn't appear to provide the source for accessing the data, Dalibor kindly pointed out to me that there's a download link here, on his webpage. This link will give you the data in spreadsheet form, together with all of the necessary background information.

The only slight concern that I have about this resource - and I don't want to sound ungrateful - is the issue of the updating of the data over time. You'll note from the abstract that the database "...... is designed to be updated regularly through (the) StatCan database....". Given my comments (above) about some of the issues that we've all faced for a very long time when it comes to StatCan data, I  know that updating this new database on a regular basis is going to be a bit of a challenge.

Added 8 March 2019: I'm glad to learn that new update of the database is now available here.

However, let's not let this concern detract from the considerable benefits that we'll all derive from having access to this rich set of Canadian macroeconomic time-series.

Thanks, again, to the authors for constructing this database, and for making it freely available!

References

Bedock, N. & D. Stevanovic, 2017. An empirical study of credit shock transmission in a small open economy. Canadian Journal of Economics, 50, 541–570.

Boivin, J., M. Giannoni, & D. Stevanovic, 2010. Monetary transmission in a small open economy: more data, fewer puzzles. Technical report, Columbia Business School, Columbia University.

Fortin-Gagnon, O., M. Leroux, D. Stevanovic, & S. Suprenant, 2018. A large Canadian database for macroeconomic analysis. CIRANO Working Paper 2018s-25.

Gordon, S., 2018. Project Link - Piecing together Canadian economic history. Département d'économique, Université Laval.

© 2018, David E. Giles

Monday, October 1, 2018

Essential Fall Reading

  • Buono, D., G. Kapetanios, M. Marcellino, G. Mazzi, & F. Papailias, 2018. Big data econometrics - Now casting and early estimates. Working paper N. 82, Baffi Carefin Centre for Applied Research on International Markets, Banking, Finance, and Regulation, Bocconi University.
  • Fair, R. C., 2018. Information content of DSGE forecasts. Mimeo
  • Lewbel, A., 2018. The identification zoo - Meanings of Identification. Forthcoming, Journal of Economic Literature.
  • Pretis, F., J. J. Reade, & G. Sucarrat, 2018. Automated general-to-specific (GETS) regression modeling and indicator saturation for outliers and structural breaks. Journal of Statistical Software, 86, 3.
  • Woodruff, R. S., 1971. A simple method for approximating the variance of a complicated estimate. Journal of the American Statistical Association, 66, 411-414.
  • Zhang, R. & N. H. Chan, 2018. Portmanteau-type tests for unit-root and cointegration. Journal of Econometrics, in press.
© 2018, David E. Giles

Tuesday, January 2, 2018

Econometrics Reading for the New Year

Another year, and lots of exciting reading!
  • Davidson, R. & V. Zinde-Walsh, 2017. Advances in specification testing. Canadian Journal of Economics, online.
  • Dias, G. F. & G. Kapetanios, 2018. Estimation and forecasting in vector autoregressive moving average models for rich datasets. Journal of Econometrics, 202, 75-91.  
  • González-Estrada, E. & J. A. Villaseñor, 2017. An R package for testing goodness of fit: goft. Journal of Statistical Computation and Simulation, 88, 726-751.
  • Hajria, R. B., S. Khardani, & H. Raïssi, 2017. Testing the lag length of vector autoregressive models:  A power comparison between portmanteau and Lagrange multiplier tests. Working Paper 2017-03, Escuela de Negocios y EconomÍa. Pontificia Universidad Católica de ValaparaÍso.
  • McNown, R., C. Y. Sam, & S. K. Goh, 2018. Bootstrapping the autoregressive distributed lag test for cointegration. Applied Economics, 50, 1509-1521.
  • Pesaran, M. H. & R. P. Smith, 2017. Posterior means and precisions of the coefficients in linear models with highly collinear regressors. Working Paper BCAM 1707, Birkbeck, University of London.
  • Yavuz, F. V. & M. D. Ward, 2017. Fostering undergraduate data science. American Statistician, online. 

© 2018, David E. Giles

Tuesday, January 17, 2017

Royal Economic Society Webcasts on Econometrics

The Royal Economic Society has recently released videos of interviews with three leading econometricans, recorded during the Society's 2016 Meeting. These are: 

Webcasts of Special (Econometrics) Sessions at RES Meetings between 2011 and 2016 are also available for viewing - here.     
© 2017, David E. Giles

Sunday, November 15, 2015

November Reading

Somewhat belatedly, here is some suggested reading for this month:
  • Al-Sadoon, M. M., 2015. Testing subspace Granger causality. Barcelona GSE Working Paper Series, Working Paper nº 850.
  • Droumaguet, M., A. Warne, & T. Wozniak, 2015. Granger causality and regime influence in Bayesian Markov-switching VAR's. Department of Economics, University of Melbourne. 
  • Foroni, C., P. Guerin, & M. Marcellino, 2015. Using low frequency information for predicting high frequency variables. Working Paper 13/2015, Norges Bank.
  • Hastie, T., R. Tibshirani, & J. Friedman, 2009. The Elements of Statistical Learning: Data Mining, Inference, and Prediction (2nd. ed.). Springer, New York. (Legitimate download.) 
  • Hesterberg, T. C., 2015. What teachers should know about the bootstrap: Resampling in the undergraduate statistics curriculum. American Statistician, in press. 
  • Quineche. R. & G. Rodríguez, 2015. Data-dependent methods for the lag selection in unit root tests with structural change. Documento de Trabajo No. 404, Departmento de Economía, Pontificia Universidad Católica del Perú.


© 2015, David E. Giles

Tuesday, September 1, 2015

September Reading List

  • Abeln, B. and J. P. A. M. Jacobs, 2015. Seasonal adjustment with and without revisions: A comparison of X-13ARIMA-SEATS and CAMPLET. CAMA Working Paper 25/2015, Crawford School of Public Policy, Australian National University.
  • Chan, J. C. C. and A. L. Grant, 2015. A Bayesian model comparison for trend-cycle decompositions of output. CAMA Working Paper 31/2015, Crawford School of Public Policy, Australian National University.
  • Chen, K. and K-S. Chan, 2015. A note on rank reduction in sparse multivariate regression. Journal of Statistical Theory and Practice, in press.
  • Fan, Y., S. Pastorello, and E. Renault, 2015. Maximization by parts in extremum estimation. Econometrics Journal, 18, 147-171.
  • Horowitz, J., 2014. Variable selection and estimation in high-dimensional models. Cemmap Working Paper CWP35/15, Institute of Fiscal Studies, Department of Economics, University College London.
  • Larson, W., 2015. Forecasting an aggregate in the presence of structural breaks in the disaggregates. RPF Working Paper No. 2015-002, Research Program on Forecasting, Center of Economic Research, George Washington University.


© 2015, David E. Giles

Saturday, April 25, 2015

Introductory Statistics for Data Science

The latest issue of Chance contains a very timely article by Nicholas Horton, Benjamin Baumer, and Hadley Wickham. It's titled, "Setting the Stage for Data Science: Integration of Data Management Skills in Introductory and Second Courses in Statistics".

Ask yourself - "Is the traditional way that we teach introductory and second-level statistics courses really suited for preparing students for future work in modern data science?"

More specifically, do our undergraduate courses provide the data-related skills that are increasingly needed? The same question could be asked of undergraduate training in econometrics.

Horton et al. itemize five things which, in their opinion, deserve more attention in this context:

Tuesday, August 19, 2014

David Mimno on "Data Carpentry"

There's a post on David Mimno's blog  today titled, "Data Carpentry".

I like it a lot, because it emphasises just how much effort, time and creativity can be required in order to get one's data in order before we can get on with the fun stuff - estimating models, testing hypotheses, making forecasts, and so on. I know that this was something that I didn't fully appreciate when I was starting my career. And when I did get the message, I found it rather irksome!

However, the message isn't going to change, so we just have to live with it, and accept the realities of working with "real" data.

In his post, David explains why he doesn't like the oft-used term"data cleaning" (which makes us sound like "data janitors"), and why he prefers the term "data carpentry". Certainly, the latter has more constructive overtones.

As he says:
"To me these imply that there is some kind of pure or clean data buried in a thin layer of non-clean data, and that one need only hose the dataset off to reveal the hard porcelain underneath the muck. In reality, the process is more like deciding how to cut into a piece of material, or how much to plane down a surface. It’s not that there’s any real distinction between good and bad, it’s more that some parts are softer or knottier than others. Judgement is critical.
The scale of data work is more like woodworking, as well. Sometimes you may have a whole tree in front of you, and only need a single board. There’s nothing wrong with the rest of it, you just don’t need it right now."
A nice post, and a very nice "take" on a crucial part of the work that we do.


© 2014, David E. Giles

Wednesday, April 30, 2014

HDDA Workshop, 2015

The Fourth International Workshop on the Perspectives on High-Dimensional Data Analysis  is going to be held here at the University of Victoria next summer. The workshop will bring together researchers involved in statistics for high-dimensional data, with researchers involved with topological methods for data analysis and visualization. It will be an "Applied Topology - Applied Statistics" event.

The link for the third such workshop, held in 2013, is here.

I'm on the organising committee, so watch this blog for further developments.


© 2014, David E. Giles

Sunday, April 13, 2014

Open Science Through R

There's so much being written about R these days, and justifiably so. If you use R for your econometrics, you should also keep in mind that its applicability is far wider than statistical analysis. 

A big HT to the folks at Quandl for leading me to a nice overview of the way in which R is enabling some big changes in the way in which scientific research is being conducted more generally. The article in question is by Tina Amirtha, "How the Rise of the "R" Language is Bringing Open Source to Science", which you'll find here.

If you think that R is just about statistics, and you can't see the point of investing some time (not money) in getting on board, then read Tina's piece. 

You'll change your mind if you consider yourself a survivor.



© 2014, David E. Giles

Monday, March 24, 2014

Thumbs Up; Thumbs Down

People say and do the darnedest things! 

I'll let you assign your own "thumbs up" and "thumbs down" to the following gems. I imagine you can guess where I stand on each of them!

'But which is a bigger menace to society, laziness about data or laziness about theory? Theory-laziness is seductive because it's easy - mining for correlations isn't very mentally taxing. But data-laziness is seductive because it's hard - the more complicated and intricate a theory you make, the smarter it makes you feel, even if the theory sucks. 
 In the past, data-laziness was probably more of a threat to humanity. Since systematic data was scarce, people had a tendency to sit around and daydream about how stuff might work. But now that Big Data is getting bigger and computing power is cheap, theory-laziness seems to be becoming more of a menace. The lure of Big Data is that we can get all our ideas from mining for patterns, but A) we get a lot of false patterns that way, and B) the patterns insidiously and subtly suggest interpretations for themselves, and those interpretations are often wrong.'
(Noah Smith in his post, Which is Better, Data or Theory?)

'........ which raises the question "who should be teaching students econometrics?" Should it be someone like ****, who is basically an applied micro guy, or should it be an econometric theorist?' 
(Frances Woolley, commenting on her own post)

'Developing statistical methods is hard and often frustrating work. One of the under appreciated rules in statistical methods development is what I call the 80/20 rule (maybe could even by the 90/10 rule). The basic idea is that the first reasonable thing you can do to a set of data often is 80% of the way to the optimal solution. Everything after that is working on getting the last 20%.'
(Jeff Leek, on the Simply Statistics blog)

'The micro stuff that people like myself and most of us do has contributed tremendously and continues to contribute. Our thoughts have had enormous influence. It just happens that macroeconomics, firstly, has been done terribly and, secondly, in terms of academic macroeconomics, these guys are absolutely useless, most of them. Ask your brother-in-law. I’m sure he thinks, as do 90% of us, that most of what the macro guys do in academia is just worthless rubbish. Worthless, useless, uninteresting rubbish, catering to a very few people in their own little cliques.'
(Chris Auld, reputedly quoting someone else, in a blog post from 2011)

'The combination of some data and an aching desire for an answer does not ensure that a reasonable answer can be extracted from a given body of data.'
(John Tukey)
'So, we produce our papers, as if on a relentless production line. We cannot wait for inspiration; we must maintain our output. To do our jobs successfully, we need to acquire a fundamental academic skill that the scholars of old generally did not possess; modern academics must be able to keep writing and publishing even when they have nothing to say. ....'
(Michael Billig, as quoted by Timothy Taylor)

So, thumbs up, and thumbs down. Or, from the sublime to the ridiculous - take your pick.
Boy - it was hard to resist giving my reaction  to some of these!
© 2014, David E. Giles

Sunday, March 23, 2014

Data Transfer Advice From Francis Smart

I always enjoy reading the posts by Francis Smart on his Econometrics by Simulation blog. A couple of days ago he wrote a nice piece titled, "It is Time for RData Files to Become the Standard for Data Transfer". 

Francis made some very good points about the handling of large amounts of data, and he provided some convincing examples regarding the compression rates and opening times for RData files as compared with other options. The comments to his post are also very relevant.

If you're using or exchanging large data files, you'll find Francis's post most helpful.


© 2014, David E. Giles

Sunday, February 9, 2014

The Statsguys on Data Analytics

It's good to see that more and more students of econometrics are taking an interest in "Data Analytics" / "Big Data" /"Data Science" literature. As I've commented previously, there's a lot that we can all learn from each other. Moreover, many of "boundaries" are very soft, and are more perceived than real.

So, I was delighted to see the arrival of The Statsguys, last month. (Hat-tip to the team at Quandl for alerting me to this.

Saturday, January 11, 2014

Reading for the New Year

Back to work, and back to reading:
  • Basturk, N., C. Cakmakli, S. P. Ceyhan, and H. K. van Dijk, 2013. Historical developments in Bayesian econometrics after Cowles Foundation monographs 10,14. Discussion Paper 13-191/III, Tinbergen Institute.
  • Bedrick, E. J., 2013. Two useful reformulations of the hazard ratio. American Statistician, in press.
  • Nawata, K. and M. McAleer, 2013. The maximum number of parameters for the Hausman test when the estimators are from different sets of equations.  Discussion Paper 13-197/III, Tinbergen Institute.
  • Shahbaz, M, S. Nasreen, C. H. Ling, and R. Sbia, 2013. Causality between trade openness and energy consumption: What causes what  high, middle and low income countries. MPRA Paper No. 50832. 
  • Tibshirani, R., 2011. Regression shrinkage and selection via the lasso: A retrospective. Journal of the Royal Statistical Society, B, 73, 273-282.
  • Zamani, H. and N. Ismail, 2014. Functional form for the zero-inflated generalized Poisson regression model. Communications in Statistics - Theory and Methods, in press.


© 2014, David E. Giles

Tuesday, December 31, 2013

My Top 5 For 2013

Everyone seems to be doing it at this time of the year. So, here are the five most popular new posts on this blog in 2013:
  1. Econometrics and "Big Data"
  2. Ten Things for Applied Econometricians to Keep in Mind
  3. ARDL Models - Part II - Bounds Tests
  4. The Bootstrap - A Non-Technical Introduction
  5. ARDL Models - Part I

Thanks for reading, and for your comments.

Happy New Year!


© 2013, David E. Giles

Monday, December 30, 2013

A Cautionary Bedtime Story

Once upon a time, when all the world and you and I were young and beautiful, there lived in the ancient town of Metrika a young boy by the name of Joe.

Saturday, December 28, 2013

Statistical Significance - Again

With all of this emphasis on "Big Data", I was pleased to see this post on the Big Data Econometrics blog, today.

When you have a sample that runs to the thousands (billions?), the conventional significance levels of 10%, 5%, 1% are completely inappropriate. You need to be thinking in terms of tiny significance levels.

I discussed this in some detail back in April of 2011, in a post titled, "Drawing Inferences From Very Large Data-Sets". If you're of those (many) applied researchers who uses large cross-sections of data, and then sprinkles the results tables with asterisks to signal "significance" at the 5%, 10% levels, etc., then I urge you read that earlier post.

It's sad to encounter so many papers and seminar presentations in which the results, in reality, are totally insignificant!


© 2013, David E. Giles

Thursday, December 12, 2013

Time for Some More Reading!

With the weekend upon us once again, it's time to settle down with the papers - the econometrics research papers, that is. Here are my latest picks:
  • Cook, S., D. Watson, and L. Parker, 2014. New evidence on the importance of gender and asymmetry in the crime-unemployment relationship. Applied Economics, 46, 119-126.
  • Fan, J., F. Han, and H. Liu, 2013. Challenges of big data analysis. Mimeo.
  • Hashmi, A. R., 2014. Competition and innovation: The inverted-U relationship revisited. Review of Economics and Statistics, in press.
  • Juselius, K., N. F. Moller, and F. Tarp, 2104. The long-run impact of foreign aid in 36 African countries: Insights from multivariate time series analysis. Oxford Bulletin of Economics and Statistics, in press.
  • Li, R., D. K. J. Lin, and B. Li, 2013. Statistical inference in massive data sets. Applied Stochastic Models in Business and Industry, 29, 399-409.
  • Sanderson, E. and F. Windmeijer, 2013. A weak instrument F-test in linear IV models with multiple endogenous variables. CEMMAP Working Paper CWP58/13, The Institute for Fiscal Studies.

© 2013, David E. Giles