Sunday, February 23, 2014

A call for more reproducible research in drug discovery/development

Phase III drug trials are typically randomized, blinded, controlled clinical trials.  However, last month in JAMA, Djulbegovic et al. (2013) argued "more than 80% of phase 1 studies and more than 50% of phase 2 studies are currently nonrandomized."  They argue that these early phase studies should all be randomized, and that even preclinical studies in animals and cell cultures should also be randomized.  (Randomization is probably even less prevalent in preclinical research than in the early phase studies discussed in the quote.)  In other words, the authors advocate study designs that encourage reproducibility across the spectrum of clinical and preclinical research.

They argue that non-randomized studies can easily lead to incorrect decisions, both pro and con.  Thus randomized studies are more efficient and provide stronger backing for decision making.  (They also make the case that randomized studies are more ethical; I'm not sure I find their reasoning here as compelling.)  They hope that use of more rigorous study designs across the drug development arena could be one way to address the industry's infamously high failure rate.

Here is a key passage from the article, describing the literature in preclinical research.

This literature yields an excess of statistically significant findings that cannot be eventually replicated, let alone translated into clinical successes.  For preclinical research conducted by the industry, routine adoption of rigorous randomized designs should be straightforward--no company wants to spend millions of dollars for the clinical testing of useless treatments.  In fact, industry researchers have taken the lead in raising the concerns about the reproducibility of preclinical research and suggesting partial solutions.  For preclinical research conducted by non industry researchers, similar rigorous practices can also be routinely adopted and requested.  Funders and journals can specify that they will sponsor and publish animal studies only if they fulfill rigorous randomization criteria.  Justified exceptions to this rule are likely to be rare.
(I have not included the footnotes; see the original.)  The major lesson for me in the above passage is that the main contribution of statistics to such studies is in the design, not the analysis.  "Statistically significant" findings by no means guarantee that the study has a chance of being reproduced, whereas good study design principles would greatly enhance the likelihood that such studies are reproducible.  Unfortunately much of the teaching and practice of statistics, both by statisticians and non-statisticians, tends to emphasize the mathematical/calculational side, rather than the study design side.

I'd like to see a devil's advocate's response to all this.  I find the authors' views compelling, and have difficulty imagining the grounds for which one might disagree.

Reference


Benjamin Djulbegovich, Iztok Hozo, and John P. A. Ioannidis, 2013:  Improving the drug development process:  more not less randomized trials.  Journal of the American Medical Association, 311 (4):  355-356.


Saturday, February 1, 2014

Responses to "When Mice Mislead"

This past week's issue of Science (the Jan. 24, 2014 issue) has two letters to the editor, responding to a report last November, "When Mice Mislead" by Jennifer Couzin-Frankel, which I discussed in an earlier post.  The first letter, by Richard Traystman and Paco Herson, points to earlier findings, similar to those reported by Couzin-Frankel, in the stroke research community.  Most importantly, they assert that "It is unlikely that poor methods used in animal studies account for all the negative clincial trials that have been performed based on preclinical studies.  After all, some investigators do perform appropriate experiments, and even those studies rarely lead to positive clinical trials."  The authors point to the fact that mouse studies are usually done with healthy young mice, whereas human subjects in neuroprotective drug clinical trials are often older and have many co-morbidities.  They propose that aged mice with comorbid diseases be used in stroke trials, as a better animal model of human disease.

The second letter is from statistician Gary Churchill.  He zeroes in on one key question:  "Was the result replicated in more than one genetic background?"  He goes on to identify two "root causes" for nonreproducible research:

Science today is driven by an incentive system that often rewards precedence and impact over quality of the work.  Statistical training of scientists often emphasizes analytical techniques over experimental design and quantitative reasoning.  These are systemic problems that will not change without substantial effort
Meanwhile, Churchill endorses the message of Couzin-Frankel's article with his maxim:  "Be wise, randomize."

I think that both of these letters add value to the original piece by Couzin-Frankel. In particular, Churchill's second "root cause" is particularly interesting, as both statisticians and lay scientists or mathematicians who teach statistics are all guilty of overemphasizing methodology, modeling, and inference at the expense of study design and critical thinking. 

References

Jennifer Couzin-Frankel, 2013: When mice mislead. Science, 342: 922-925.

Richard J. Traystman and Paco S. Herson, 2014:  Misleading results:  translational challenges.  Science, 343:  369-370.

Gary Churchill, 2014:  Misleading results:  don't blame the mice.  Science, 343, 370.


The value and place of prespecifying data analysis plans

DTLR does not usually stray into the social sciences, but a paper in Science last month (Miguel, et al., 2013) provides another opportunity to dwell on reproducible research. Prospective, designed experiments are becoming more common in the social and behavioral sciences, particularly in economics and program evaluation. However, as in the natural sciences, “Commentators point to a dysfunctional reward structure in which statistically significant, novel, and theoretically tidy results are published more easily than null, replication, or perplexing results.” Reporting standards in social science journals are similarly lax as those in biology journals, and “researchers have incentives to analyze and present data to make them more 'publishable,' even at the expense of accuracy.” Examples of poor practices include the publication of positive results which form a subset of a larger study with mixed or null results, as well as presenting exploratory findings dressed up as confirmatory results.

The authors propose that three core practices be emphasized: disclosure, registration and preanalysis plans, and open data and materials. These concepts are familiar to those who work in clinical trials. However, the authors believe that the situation can be improved yet further than in the medical trial model. In the latter, the “dominant role” of government regulatory agencies “arguably slows adoption of innovative statistical methods.” The authors are also resistant to a “one-size-fits-all” approach for trial registration, preferring a method-specific approach. They foresee some convergence between methods used in behavioral research with those in medical trials, particularly in the neuroscience arena.

Near the end of the paper, there is a particularly eloquent passage that I'd like to quote in full.
The most common objection to the move toward greater research transparency pertains to preregistration. Concerned that preregistration implies a rejection of exploratory research, some worry that it will stifle creativity and serendipitous discovery. We disagree.
Scientific inquiry requires imaginative exploration. Many important findings originate as unexpected discoveries. But findings from such inductive analysis are necessarily more tentative because of the greater flexibility of methods and tests and, hence, the greater opportunity for the outcome to obtain by chance. The purpose of prespecification is not to disparage exploratory analysis but to free it from the tradition of being portrayed as formal hypothesis testing.
The above two paragraphs can easily carry over into all experimental research, not just those in the social and behavioral sciences.


Reference

 
E. Miguel, et al., 2014: Promoting transparency in social science research. Science, 343: 30-31.

Monday, January 27, 2014

NIH ponders reproducible research, redux

Last August I wrote about the NIH's considerations of measures to encourage reproducible research.  Today, the NIH's leaders Francis Collins and Lawrence Tabak have announced what they've been up to.  Their note quite appropriately emphasizes study design and the reporting of details of experimental design.  They are basically running a range of pilot projects and plan to decide which measures to adopt permanently at the end of this year.  Although this is a very cautious and gradualist approach, I applaud their intention to act on this issue.  And I agree with them that the whole scientific community needs to take ownership of reproducible research.  I look forward to reading about their findings and their decisions in 4Q14.

DTLR is delighted to endorse the NIH's actions to support reproducible research.

Reference


Francis S. Collins and Lawrence Tabak, 2014:  NIH plans to enhance reproducibility.  Nature, 505:  612-613.

Friday, January 24, 2014

Science Magazine embraces reproducibility (finally)

Last week's issue of Science carried an editorial titled simply, "Reproducibility" (McNutt, 2014).  The Editor-in-Chief, Marcia McNutt, announced that Science would be adopting recommendations of the U.S. National Institute of Neurological Disorders and Stroke (NINDS) to encourage transparency and reproducible research in preclinical studies (Landis, et al., 2012).  In addition, the journal would spend the next six months gathering examples of "excellence in transparency" in order to further develop guidelines to incentive reproducible research.  Finally Science plans to add more statisticians to its editorial board, to involve them in manuscript review.  Science follows its sister journal, Science Translational Medicine, which has already adopted the NINDS guidelines, as well as Nature which published the NINDS guidelines and posted its own, discussed here. Although Science has certainly dragged its feet on this issue, in comparison to Nature, they should be commended for finally joining the bandwagon.

The NINDS paper (Landis, et al., 2012) is a small masterpiece.  Not only does it provide a concise and thoughtful list of issues that preclinical investigators should consider in the design, execution, and analysis of preclinical studies, the paper also provides a useful literature review documenting both the problems and the proposed solutions and their outcomes, particularly in comparison to the clinical trials literature.  I would also like to point readers to a more specific set of guidelines for preclinical imaging studies (Stout, et al., 2013).  (This blog, after all, is named after an imaging method.)

References


S. C. Landis, et al., 2012:  A call for transparent reporting to optimize the predictive value of preclinical research.  Nature, 490:  187-191.

M. McNutt, 2014:  Reproducibility.  Science, 343:  229.

D. Stout, et al., 2013:  Guidance for methods descriptions used in preclinical imaging papers.  Molecular Imaging, 2013:  1-15.



Monday, January 20, 2014

The New York Times on nonreproducible research

Today's New York Times has a new column, Raw Data, by George Johnson.  Appropriately enough, the inaugural column is about non-reproducible research.  It covers mostly the same ground that the Economist did in its special issue, "How Science Goes Wrong" last October.  We've talked about many of these issues on this blog, DTLR, over the past half year as well.  It is gratifying to see that the venerable newspaper deems the topic fit to write about.

There is one topic in Johnson's column that we haven't examined here on DTLR yet, and that is the opinion piece by Mina Bissell writing in Nature last November.  Her thoughtful piece will be the subject of a future post. 

Hurricane Sandy, climate change, and the limits of data science

Climate science is a very touchy subject, because unfortunately nearly all discussion of it becomes quickly entwined with politics. A reminder of this was recently discussed by Kerr (2013). In the President's State of the Union address a year ago, Mr. Obama said “We can choose to believe that Superstorm Sandy, and the most severe drought in decades, and the worst wildfires some states have ever seen were all just a freak coincidence. Or we can choose to believe in the overwhelming judgment of science—and act before it's too late.”

However, Kerr (2013) notes that “there is little or no evidence that global warming steered Sandy into New Jersey or made the storm any stronger. And scientists haven't even tried yet to link climate change with particular fires.” Kerr also points to a Republican Congressman's equally incorrect claim that “Extreme weather isn't linked to climate change.” Kerr states that several heat waves have indeed been “securely linked” to global warming. Kerr says that “Links between extreme weather and climate change are not only often scientifically suspect, they may also be a risky strategy to take climate change seriously.” After all, climate by definition is a statistical average of weather, which is what we experience on a day-to-day basis. The President, alas, was wrong.

Last March, I attended a lecture by Dr. Richard A. Anthes, an eminent meteorologist and president emeritus of the University Corporation for Atmospheric Research, and a former president of the American Meteorological Society. He pointed out that the track of Hurricane Sandy, with its left turn towards New Jersey, had been predicted days in advance by the ECMWF forecast model (European Center for Medium-Range Weather Forecasts). The U.S. forecast models were not able to give as early a warning, due to technical limitations of the computers and computer models (see my earlier post).

Let's use this example in a thought experiment on how data could be used to make a hurricane forecast. A purely empirical approach (whether by conventional statistical methods or by data mining/machine learning) would likely have failed: never before had a hurricane approached New Jersey from the east in late October. The ECMWF model uses data too, but not for statistical forecasting. It uses data as initial conditions to simulate the atmosphere using partial differential equations that incorporate subject matter knowledge of the physics and chemistry of the atmosphere and ocean. By doing so, it forecast the formation of the storm before it actually formed, as well as its subsequent track. The successful forecast of the ECMWF model was interpreted correctly by political authorities and heeded by the public, saving tens of thousands of lives. If you want an example of science at its best, here it is.

Make no mistake: meteorology as a science does make very judicious use of statistical and monte carlo methods. The atmospheric sciences, however, are driven primarily by methods based on subject matter knowledge, not strictly empirical methods such as those used by statisticians and data scientists (“superficial statistics” in the devastating words of Salby, 2012). Note that both approaches are equally data hungry.

Let's return to climate now. The whole point of climate change is that data in the future will not be like data from the past. As Showstack (2013) reports, Kathryn Sullivan, acting administrator of the National Oceanic and Atmospheric Administration (NOAA) gave the keynote address at the National Research Council's Board on Earth Sciences and Resources in November. She said, “The past is no longer prologue when it comes to the risks we bear at any given place on this planet. The statistical pattern of our past cannot be relied upon fully to tell us what our future will be.” Isn't this a conundrum for a statistician or a data scientist?  What good is the training data when we know it will not be informative about data from the future? 

In my view, the solution to all this is to use first principles modeling, as climate scientists do. As with meteorology, in climatology knowledge of the physics and chemistry of the atmosphere, embodied in the partial differential equations of climate modeling, is preferred to the methods of statistics and data science for predicting both the weather and the climate.

References

 

 

Richard A. Kerr, 2013: In the hot seat. Science, 342: 688-689.

Murray L. Salby, 2012: Physics of the Atmosphere and Climate. Cambridge University Press, p. xvi.

Randy Showstack, 2013: Earth sciences and societal needs explored at National Research Council meeting. EOS, Transactions of the American Geophysical Union, 94 (48): 457-459.