Sunday, December 29, 2013

The replication myth?

Earlier this month, Scientific American blogger (and neuroscience graduate student) Jared Horvath posted a rebuttal to the Economist cover story on “How Science Goes Wrong”, the latter of which I discussed in a previous post. Let's examine Horvath's post in detail; as we'll see, there is not much to like.

Horvath asserts that “[U]nreliable research and irreproducible data have been the status quo since the inception of modern science. Far from being ruinous, this unique feature of research is integral to the evolution of science.” Horvath presents several historical examples of great scientists publishing findings based on data that later investigators have not been able to reproduce. Three cases are described in detail (Galileo, Dalton, and Millikan) and others are mentioned (Mendel, Darwin, and Einstein). Nonetheless, he agrees that their contributions were of tremendous value to science. “Their work, if ultimately invalid, proved useful.” Horvath concludes that “If replication were the gold standard of scientific progress, we would still be banging our heads against our benches trying to arrive at the precise values that Galileo reported. Clearly this isn't the case.”

I believe Horvath's reading of history is deeply flawed. First of all, my take-home from his case studies is that it is very easy for scientists to fool themselves, and sometimes they are lucky enough to be right. Second, the “irreproducible” findings were only useful because they were consistent with the findings of others, which in the aggregate pointed the way to correct theories. In modern times, the development of scientific theory would be greatly expedited by publishing more defensible results in the first place.

Horvath goes on to describe the serendipitous route to the discovery of Viagra by Pfizer scientists, stating that it illustrates “the true path by which science evolves.” Although the original goal of the drug had failed, the data revealed other, unexpected applications. “Had the initial researchers been able to massage their data to a point where they were able to publish results that were later found to be irreproducible, this would not have changed the utility of a sub-set of their results for the field of male potency.”

Again, I disagree with Horvath's interpretation of events. The point of clinical research is not to massage the data until it is reproducible. This in fact was not the problem with the drug, rather its failure for the original study objectives was the problem. If the study design and execution are sound, there is no fiddling with the data needed at all, and this is precisely how the Pfizer scientists salvaged the drug. It is a non-sequitur to jump from the nonlinear path of drug development to somehow giving a blessing to non-reproducible research.

Horvath goes on to criticize the conventional portrait of the scientific method as progressing in discrete, cumulative steps. “In reality, science progresses in subtle degrees, half-truths and chance. An article that is 100 percent valid has never been published. While direct replication may be a myth, there may be information or bits of data that are useful among the noise.” Once again, I find these arguments non-sequiturs. The orthodox portrayal of the scientific method has been criticized by countless philosophers of science. However, this issue is completely disconnected from that of non-reproducible research. If we were to bless the half-baked publication of results, as Horvath seems to do, we would also give blessing to the kind of work discussed by Glenn Begley. Begley once cornered an oncologist whose work he (Begley) could not reproduce. Upon questioning, the oncologist admitted, “We did this experiment a dozen times, got this answer once, and that's the one we decided to publish” (Couzin-Frankel, 2013).

Horvath goes on to celebrate what we can learn from failure, and that “with enough time all scientific wells run dry.” Learning from failure is certainly a good thing, as discussed by Couzin-Frankel (2013). However, the right way to do that is to learn from failures resulting from well designed, conducted, and reported studies, not from the garbage that the Economist article was addressing. In the latter case, I believe “garbage in, garbage out” is the lesson. Horvath seems to be arguing that “garbage in, gospel out” (the ironic motto of this blog)!

Do all scientific wells run dry? I think not. The classical mechanics of Galileo and Newton still provides the framework for the study of classical fluid dynamics, a lively discipline that thrives to this day. The concepts of Darwinian evolution continue to influence how we fight infectious disease, for instance, by updating the influenza vaccine on an annual basis.

I'm all for disclosing the messy nature of scientific progress, learning from failure (and publishing the results), and so on. However, all of these things can only be done when critical thinking (including statistical thinking) are constantly at work throughout the process. We should all be seeking to make research reproducible not by “massaging data” but by thinking critically, using good study design and execution, employing sound data collection, and providing full disclosure of methods and results. It seems to me that Horvath has not really understood the reasons for non-reproducible research that have been put forward by Ioannidis (2005) and others.

I hate to end the year on a sour note, but this is likely to be my last post for 2013. Nonetheless the raw material for this blog is considerable, and hopefully I will have more to post in the new year. Thanks for reading!

References


Jennifer Couzin-Frankel, 2013:  The power of negative thinking.  Science, 342:  68-69.

John P.A. Ioannidis, 2005:  Why most published research findings are false.  PLoS Medicine, 2 (8),  e124:  696-701.

Wednesday, December 25, 2013

Some notes on high energy physics

DTLR has been on hiatus for the last month or so, due to travel and other commitments.  I will slowly try to get back into the groove of things in the coming weeks.

A couple items of interest related to high energy physics have appeared in the interim.  First, there is a very candid Guardian interview with 2013 Physics Nobel Laureate, Peter Higgs, by Decca Aitkenhead.  I found a few things striking in the interview.  First, he did not seem particularly bothered by the fact that only two of the six or so theorists involved in the prediction of the Higgs particle were awarded the Nobel Prize.  I was bothered, as described in earlier posts here and here.  Second, the article states that Higgs has not written many papers over the course of his career, and was considered an embarrassment to his department for his lack of productivity.  Again he does not seem to be bothered by this.  I do sympathize with his complaint that in today's research environment, he might not have had the time or space for the deep thinking required for formulating his theory.  In fact he says he might not be able to get a job in today's environment.  This should lead us to reflect on the degraded condition of scientific research infrastructure in our own time.

On the other hand, the lack of productivity would certainly bother me if it were to happen to me.  I would have either changed fields or careers, or somehow found some other way to contribute to society.  Perhaps Higgs continued at least to teach before his retirement?  Richard Hamming (1997) was unsympathetic to great scientists whose productivity slowed when they were given the opportunity to work without constraints, as at the Institute for Advanced Study in Princeton, NJ.  Hamming, like the Nobel laureate economist, James M. Buchanan (1994), most admired hard workers who were driven to be productive.  (Hamming names John Tukey as a prototype of the successful hard-working scientist.)

Third, the story of the rejection of Higgs' initial manuscript is perhaps a salutary one, as it forced him to rewrite it and explicitly identify the new particle.  Evidently it made the importance of the paper more obvious and led to publication.  This is a (perhaps rare) example of peer review doing what it was intended to do.

Finally, it is stated that Higgs turned down a knighthood, but was tricked into accepting another national honor.  I agree wholeheartedly with his explanation:  "I'm rather cynical about the way the honours system is used, frankly. A whole lot of the honours system is used for political purposes by the government in power."  I also concur with his discomfort with the name "God Particle" that has been used to describe the Higgs boson.

Looking forward now, the incoming Fermilab director, Nigel Lockyer, has an interesting perspective on the future of big high energy physics experiments, published in Nature earlier this month.  The lack of resources for such big projects is enforcing global cooperation.  For instance, he says that the world really needs a long baseline neutrino experiment, but that the world can only afford to pay for one.  With three candidate sites, physicists from all nations need to work together whichever one ends up getting funded.  It is interesting to note that he says that talk of brain drains and gains has been replaced with 'brain circulation.'  All of this sits well with me.

References


James M. Buchanan, 1994:  Ethics and Economic Progress.    University of Oklahoma Press.  See Chapter 1 in particular.

Richard W. Hamming, 1997:  The Art of Doing Science and Engineering:  Learning to Learn.  Gordon and Breach.  See Chapter 30 in particular.




Tuesday, November 19, 2013

Follow up on John Bohannon's 'Open access sting'

Last month, I made note of the 'open access sting' carried out by John Bohannon and published in Science.  Last week, this post at the Scholarly Kitchen features an interview with Bohannon in which he is given an opportunity to answer his critics.  I find all of his answers persuasive.  Bohannon has adequately defended his work, in my view.

However, I still sympathize with the critics who would have liked a control group of non-open-access publishers as well as those who believe the peer review system, writ large, is broken.  Bohannon makes clear that addressing both of these issues was out of scope for his project.  That's fair.  Many of us, however, are indeed concerned with these broader matters.  I don't think it would be necessary to carry out a 'sting' to demonstrate that non-open-access publishers, including top-notch ones like Science, have a track record of publishing substandard or even deeply flawed research.  I wrote at length about this earlier.

H/T:  In the Pipeline

Friday, November 15, 2013

bioRxiv goes live

In an earlier post I mentioned the forthcoming preprint server for the life sciences, bioRxiv.  According to Nature, the site has now launched.  See the write-up by Ewen Callaway here.



The STEM-Crisis Myth

The Chronicle of Higher Education this week has an article by Michael Anft examining the "STEM-Crisis Myth", or "The much-hyped shortage of science and tech graduates."  Politicians, business and academic leaders, and professional science societies are constantly warning of a shortage of American college students who want to study science and engineering.  The actual data backing up such claims, the Chronicle shows, is at best contestable.  The article tries to air both sides of the debate.  The president of Arizona State University is quoted as stating, "there's too much reliance on anecdotes" about the alleged poor job market for science graduates.

The Chronicle cites data alleging that about half of STEM majors leave their field within 10 years, and that 1 in 5 American scientists contemplate leaving the country.  There are dueling studies, however.  My personal experience is more consistent with the allegation that the STEM-jobs-crisis is a myth.  The industry in which I once worked has experienced a massive downsizing of its workforce in the last decade or so, including in my field.  More senior scientists have had trouble finding new positions that make full use of their talents.  Meanwhile, young graduates have trouble finding jobs, and many post-docs have been trapped in academic limbo with too many chasing too few faculty and industry positions.  Sequestration and the instability of the federal budget threatens federal funding for science across academia as well as the national labs.  Put simply, there isn't much money available for basic and applied research in the public and private sectors these days.

The 'myth' has been discussed in other venues long before this, and I am glad that the Chronicle has decided to cover it.  DTLR encourages skepticism about the alleged shortage of scientists in the labor market.  Those who have been talking up the shortage have a vested interest in increasing enrollments in universities, membership in scientific societies, and expanding the labor pool of science graduates in order to push down labor costs.  This includes academic, business, and professional society leaders across the sciences, as well as politicians.  The rhetoric is highly self-serving, because it promotes their own vested interests at the expense of the young people to whom they are serving up deceptive statements.  DTLR believes that any professional society, university, or business leader is committing fraud when they recruit youngsters into science and engineering with the promise of a bountiful job market when they graduate.  Many of them realize this, because they take an alternate tack by arguing that STEM training is a good foundation for any career, as the Chronicle notes, and the article ends by promoting the liberal arts idea of a "broader education" including the sciences and humanities.  I'm not sure there is much data to support these views.

Read the article and decide for yourself.

Reference


Michael Anft, 2013:  The STEM-Crisis Myth.  Chronicle of Higher Education, LX (11):  A30-A33 (Nov. 15, 2013).






Sunday, November 3, 2013

Non-reproducible research in the news

An epidemic of non-reproducible research in the life and behavioral sciences has been revealed in recent years. Much of the spadework has been done by John Ioannidis and collaborators, discussed earlier on DTLR. Well known biopharmaceutical industry reports from Bayer (Prinz, et al., 2011) and Amgen (Begley & Ellis, 2012) provide further confirmation. 

Glenn Begley, one of the co-authors of these papers, was interviewed for a story by Jennifer Couzin-Frankel in the recent Science special issue on Communication in Science, discussed on DTLR last month. Couzin-Frankel (2013) discusses Begley's failed attempts to reproduce the results published by a prominent oncologist in Cancer Cell. At a 2011 conference, Begley invited the author to breakfast and inquired about his team's inability to reproduce the results from the paper. According to Begley, the oncologist replied, “We did this experiment a dozen times, got this answer once, and that's the one we decided to publish.” Begley couldn't believe what he'd heard.

Indeed, I am simultaneously shocked but not surprised. Shocked, because it displays an utter lack of critical thinking on the oncologist's part. Not surprised, because in my experience critical thinking is rarely formally taught to scientific researchers, and the incentive system for scientists rewards such lax behavior. The oncologist may have forgotten why he got into science and medicine to begin with. The pressures of a career in academic medicine may have corrupted his integrity, but the work of Ioannidis and others alluded to above shows that this phenomenon is pretty common.

The rest of Couzin-Frankel's article discusses how clinical studies often get published even when the primary objective of the study has failed. Usually (but not always) the authors are up front about the failure, but try to spin the results positively in various ways. For instance, by making enough unplanned post hoc statistical comparisons, inevitably they'll find one that achieves (nominal) statistical significance, and they'll use that to justify the publication. Evidently journals allow this to occur, resulting in tremendous bias in what gets published. These are examples of selective reporting (cherry-picking) and exaggeration that result in misleading interpretations. This is not how science ought to be done.

Couzin-Frankel's article ends with a discussion of journals dedicated to publishing negative results, as well as recent efforts by mainstream medical journals to allow publishing negative studies.

Non-reproducible research has also gotten the attention of The Economist, which ran a cover story and editorial on it a few weeks ago. As additional evidence they cite the following statistic: “In 2000-2010 roughly 80,000 patients took part in clinical trials based on research that was later retracted because of mistakes or improprieties.” Thus there are real consequences. Patients are needlessly exposed to clinical trials that may have negligible scientific value; their altruism is being abused. This should be a worldwide scandal, and I congratulate The Economist for shining a harsh light on the problem.

The Economist points out that much of this research is publicly funded, and hence a scientific scandal becomes a political and financial one. “When an official at America's National Institutes of Health (NIH) reckons, despairingly, that researchers would find it hard to reproduce at least three-quarters of all published biomedical findings, the public part of the process seems to have failed.” They then discuss the journal PLoS One, which publishes papers without regard to novelty and significance, but only for methodological soundness. “Remarkably, almost half the submissions to PLoS One are rejected for failing to clear that seemingly low bar.” Among the statistical issues the article discusses are multiplicity, blinding, and overfitting.

The Economist points discusses the main reasons for these problems: scarcity of funding for science, which leads to hyper-competition; the incentive system that rewards non-reproducible research and punishes those interested in reproducibility; incompetent peer review; and statistical malpractice. The suggest a number of solutions: raising publication standards, particularly on statistical matters; making study protocols publicly available prior to running a trial; making trial data publicly available; and making funding available for attempt to reproduce work, not just publish new work.

The Economist's article has generated a certain amount of controversy, but I think it gets it mostly right.  I would have formulated the statistical discussion differently, and I think the article misses the chance to point out more fundamental statistical problems.  I also don't give much weight to the comments by Harry Collins about "tacit knowledge".  A truly robust scientific result should be reproducible under slightly varying conditions.

References


Begley, C.G., and Ellis, L.M. (2012): Drug development: raise standards for preclinical cancer research. Nature, 483: 531-533.

Jennifer Couzin-Frankel, 2013: The power of negative thinking. Science, 342: 68-69.

Prinz, F., Schlange, T., and Asadullah, K. (2011): Believe it or not: how much can we rely on published data on potential drug targets? Nature Reviews Drug Discovery, 10: 712.


Data management plan to be included in open access mandate

This month's APS News has a front-page report by Michael Lucibella, “Open Access Mandate will Include Raw Data.” The story focuses on the forthcoming mandate from the U.S. Office of Science and Technology Policy (OSTP), regarding open access to journal papers derived from federally funded research, one year after publication. Lucibella says that although no official statement has been made, it is expected that the mandate would include a data management plan, to make data sets generated by public funds available to the public as well. The story quotes the OSTP memo stating that “scientific data resulting from unclassified research supported wholly or in part by Federal funding should be stored and publicly accessible to search, retrieve, and analyze.” Specifics are “just starting to take shape.” The story goes on to outline various challenges to such a mandate.

One valuable feature of the mandate is that “Data points that have been expunged from the final analysis will likely have to be included, the idea being that scientists can evaluate why those points were eliminated.” In principle this is a good thing, but it will be nearly impossible to enforce. Also, there may be some subjectivity involved, as data that are clearly from documented technical errors should probably not be included (in my view); transcription errors should be corrected before posting. Also, I would like meta data to be included along with the raw data files.

The story states that computer codes would not be included in the mandate, “though talks are continuing over this point.” The story quotes statistician Victoria Stodden, who expresses concern about the omission of computer codes, which will obstruct reproducible research. I share Stodden's concern and I hope the mandate will include computer codes.

Modulo the concern about computer programs, DTLR endorses both the mandate to make journal articles public after one year, as well as the mandate to make the data publicly available.  I was a co-author on two publications where we provided supplemental information that included data sets and computer scripts.  However, I've co-authored nearly 20 refereed papers in total, and obviously most of them did not include such supplemental information.  As a result, all these years later, it is impossible for me to reproduce any of that work.  (Caveat:  some of this research was not financed by public funds; nonetheless I believe the principle should apply to all published research.)  I wish such a mandate had been in place at the beginning of my career, so that all of my published work could be reproducible.  With job changes and so on, I've long lost track of data sets and computer codes that were employed in doing the work reported in those papers.

Reference