Saturday, September 24, 2016

The Atlantic weighs in on reproducible research

DTLR has dwelled on the reproducibility crisis since its birth.  Most of the discussions I have cited are from the scientific literature, but some have appeared in venues intended for a general audience.  One of the best I've seen has just appeared in the Atlantic, in a post by Ed Yong.  It appropriately places the focus on the incentive system for scientists.  A one sentence summary is provided by its quotation of Richard Horton, editor of the Lancet, who said:  "No one is incentivized to be right.  Instead, scientists are incentivized to be productive."

Please take a look at Yong's post.


Sunday, August 14, 2016

Book Review: Stephen Stigler's "The Seven Pillars of Statistical Wisdom"



The Seven Pillars of Statistical Wisdom, by Stephen M. Stigler (Harvard University Press, Cambridge, Mass., 2016).

The book presents seven principles that the author believes support the core of statistics as a unified science of data, “the original and still preeminent data science” (p. 195).  It is intended for both professional statisticians and the “interested layperson,” though I suspect the latter would struggle a bit, as the author does not shy from formulae, calculations, and even name-drops of advanced statistical methods and concepts.  The author is a distinguished professor of statistics at the University of Chicago, and a leading historian of the field.  Each of the seven main chapters discusses one of the “pillars,” illustrated with historical examples (as opposed to contemporary ones) and often accompanied by discussions of the pitfalls involved with each principle.  The author states that “I will try to convince you that each of these was revolutionary when introduced, and each remains a deep and important conceptual advance” (p. 2).

The first principle is titled “Aggregation” or “the combination of observations,” of which the arithmetic mean is the chief example discussed.  The author implies that the method of least squares, and more general smoothing methods, also falls under aggregation, broadly understood.  The concept of aggregation is radical because the principle implies that individual observations can be discarded in favor of sufficient statistics.  Prior to a general acceptance of averages, scientists would often simply choose the “best” of a set of observations, or perhaps take a midrange (average of the highest and lowest values).  The concept offers other dangers, as the author illustrates with Quetelet’s notion of the Average Man.

The second principle is titled “Information:  its measurement and rate of change”, which focuses on the Central Limit Theorem and the root-N law (which roughly states that the gain in precision of an estimate increases with only the square root of the amount of data used to calculate it).  The author acknowledges the contrast of statisticians’ usage of the term “Information” (specifically, Fisher Information) with its more general use in signal processing and information theory (specifically, Shannon information).  Again, pitfalls are discussed, including a case where randomly selecting one of two data points is better than using their average.  This is a case where cannonballs of two calibers are reported by different spies.  A cannon whose caliber equals their average would not exist.  “The measurement of information clearly required attention to the goal of the investigation” (p. 59).  (In my view, one could write an entire chapter on that last point, and it would be more important than most of the 7 principles selected by the author for this book.)

The third principle is titled “Likelihood:  Calibration on a Probability Scale”.  Here the concept of a statistical significance test is introduced, along with p-values, Bayesian induction, and the theory of maximum likelihood.  The fourth principle is titled “Intercomparison:  within-sample variation as the standard.”  He illustrates it with Student’s distribution and t-test, and the analysis of variance.  Pitfalls are illustrated with an example of data dredging in the hands of economist William Stanley Jevons.  The author acknowledges further pitfalls, “for the lack of appeal to an outside standard can remove our conclusions from all relevance” (p. 198). (In my view this concern is understated:  statisticians are fond of standardizing data, but this prevents multiple data sets from being compared using an external standard.  Dimensional analysis offers an alternate approach.)

The fifth principle is titled “Regression:  Multivariate Analysis, Bayesian Inference, and Causal Inference”.  This principle warrants the longest chapter of the book, and begins by focusing on regression to the mean, a discovery made by Francis Galton.  This discovery resolved a paradox Galton had noticed in Darwin’s theory of evolution:  if each generation produced heritable variation of traits to its offspring, why was the aggregate variation in those traits stable over time?  Later in the chapter, Stein’s paradox is discussed, and shrinkage estimation is presented as a version of regression.  The correlation-causation fallacy is also discussed, including spurious correlation and Austin Bradford Hill’s principles for epidemiological inference.  The chapter also covers multivariate analysis, Bayesian statistics, and path analysis-- a real hodge podge.

The sixth principle is “Design:  Experimental Planning and the Role of Randomization”.  Fisher’s demolishment of one-factor-at-a-time experimentation is discussed, as is Pierce’s innovation of using randomization in experimental psychology studies, and later Neyman’s discussion of random sampling in social science.  The chapter ends with a brief discussion of clinical trials and a lengthier discussion of the French lottery of the 18th and 19th centuries.

The seventh principle is titled “Residual”, by which the author means both residual analysis, a commonly used approach for statistical model criticism, as well as formal model comparison for nested models, using a significance test.  The author also detours into the history of data graphics.  The chapter is marred by infelicities in its history of physics and astronomy.  At one point the author states that “We are still looking for that [lumineferous] aether” (p. 172).  Rest assured, most physicists are not worried about that.  The author then describes Laplace’s approach to resolving an apparent discrepancy in the orbits of Jupiter and Saturn; Laplace was able to show that the motions could be explained using a mutual 3-body problem with the sun.  Using an exaggeration worthy of our current Presidential candidates, the author observes that “A residual analysis had saved the solar system.”  Finally, in the Conclusion, the author speculates about the possibility of an (as yet unknown) eighth pillar to accommodate the era of big data.

At this point, readers should be warned that I have an unconventional and dissident view of statistical ideology.  For instance, where the author states about statistical significance tests, “misleading uses have been paraded as if they were evidence to damn the entire enterprise rather than the particular use” (p. 197), I would number myself among those who would damn the entire enterprise.  (This is a topic of current controversy, as evidenced by Wasserstein, 2016.)  There is some value in distilling the ideas of statistics into a set of principles; similar exercises are commonly embarked on, and Kass et. al (2016) is another example published in the same year.  Were I to write such an account, it would differ from both Stigler’s and others’, and present my own statistical ideology.  This will have to wait for another day.  Suffice it to say that my selection of pillars would differ, and any discussion I'd offer of Stigler's would dwell far more on the pitfalls and hazards than he has.

In my view, this book’s chapter on “Design” is the best (except for the digression on the French lottery), while the topics discussed in the other chapters are so fraught with difficulties that the concepts described might be as potentially harmful as they are helpful to the serious data analyst.  I found the book disappointing and less enlightening than I had hoped. While not as bad as Salsburg's The Lady Tasting Tea, I would find it difficult to recommend this book to readers of any level of statistical sophistication.


References


Kass, R.E., et al., 2016:  Ten simple rules for effective statistical practice.  PLoS Computational Biology, vol. 12 (6), e1004961.

Wasserstein, R. L. (ed.), 2016:  ASA statement on statistical significance and P-values.  The American Statistician, vol. 70, pp. 129-133.

Friday, July 22, 2016

A less dismal science

This blog rarely strays into the behavioral sciences, and for good reasons.  Some of these reasons are outlined in an article in last week's The Economist, in a special insert, "The World If".  This particular piece ponders the scenario, What if economists reformed themselves?  One of the criticisms identified in the article is "model mania"; the author writes, "problems arise when they mistake the map for the territory."  Frankly, I think this is a criticism that applies more broadly, to any area of mathematical modeling where the contact between model and reality is very loose or non-existent.  This occurs when mathematical models are not validated by comparison with actual data; the ultimate validation regime is to predict new phenomena or future data, and compare such predictions with experimental or observational data.  Theories of physics are usually test driven in this way, as are engineering models, and many of those in data science.  Such validation is often lacking in both economics and inferential (as opposed to predictive) statistical modeling in general.  The author of the Economist piece recommends that economists repeat the mantra, "My model is a model, not the model."  DTLR advises all other users of mathematical and statistical models to do the same.

For further reading, see The Financial Modelers' Manifesto by Paul Wilmott and Emanuel Derman (2009).






Exploratory or confirmatory?

In last week's issue of Science, outgoing editor Marcia McNutt was interviewed (Shell, 2016) on the occasion of beginning a term as President of the National Academy of Sciences.  I am going to reproduce a lengthy quote from the interview.

At Science, the paradigm is changing.  We're talking about asking authors, 'Is this hypothesis testing or exploratory?'  An exploratory study explores new questions rather than tests an existing hypothesis.  But scientists have felt that they had to disguise an exploratory study as hypothesis testing, and that is totally dishonest.  I have no problem with true exploratory science.  That is what I did most of my career.  But it is important that scientists call it as such and not try to pass it off as something else.  If the result is important and exciting, we want to publish exploratory studies, but at the same time make clear that they are generally statistically underpowered, and need to be reproduced.

Bravo, Dr. McNutt!  DTLR agrees completely with the sentiment here.  It matters because the statistical dressing that accompanies much scientific research is usually only appropriate for confirmatory studies, or those that McNutt calls hypothesis testing, rather than hypothesis finding (exploratory).  It is rare to find the editor of a major scientific journal express this view in such a crisp, precise manner.   DTLR hopes that her successor, and other editors and referees of scientific journals, follow the lead set by McNutt.  DTLR also recommends all readers of this blog to take a look at Tukey (1980).

Reference


Ellen Ruppel Shell, 2016:  Hurdling obstacles:  Meet Marcia McNutt, scientist, administrator, editor, and now National Academy of Sciences president.  Science, vol. 353, pp. 116-119.

John W. Tukey, 1980:  We need both exploratory and confirmatory.  The American Statistician, vol. 34, pp. 23-25.

Friday, June 17, 2016

Randomized clinical trials, defended

Medical blogger Vinay Prasad has posted a vigorous defense of randomized clinical trials, responding to a recent paper in the New England Journal of Medicine.  It's worth a look.

H/T:  In the Pipeline by Derek Lowe (discussion here).

Wednesday, May 25, 2016

Nature keeps the heat up on reproducible research

This week's issue of Nature has a good article by Monya Baker on a wide-ranging survey of scientists about reproducible research, and a related editorial.  DTLR is most encouraged by the final table in Baker's article, the ratings of factors most likely to improve reproducibility.  "More robust experimental design" received the most combined "likely" and "very likely" ratings.  I think that this is the right answer.  Also highly ranked were "better mentoring/supervision" and "better understanding of statistics".  This latter one is a tough call, as statisticians themselves seem not to have reached a consensus on how to move forward, as evidenced in the extensive Discussion items published along with the American Statistical Association's Statement on Statistical Significance and P-valuesposted in early March.

DTLR expresses thanks to Nature for keeping the drums beating on reproducible research.  The issue is very visible right now, and the community should strike while the iron is hot, in terms of reforming the infrastructure of our community (laboratory practices, publication standards, and incentives for grant funding, promotion, and tenure).  Mis-aligned incentives are ultimately the cause of non-reproducibility, though methodological issues (poor study design and execution, inappropriate use of statistical methods, etc.) are key enablers.