Showing posts with label reproducible research. Show all posts
Showing posts with label reproducible research. Show all posts

Sunday, March 29, 2020

In praise of high journal acceptance rates?

A Penn State professor of astronomy and astrophysics, Jason Wright, published a commentary in the February 2020 issue of Physics Today, titled "High journal acceptance rates are good for science".  He notes that the journals he publishes in have an 85% acceptance rate.  He says that "Therefore nearly all significant astronomical results submitted to those journals that are not obviously fatally flawed are likely to be published....it is the sign of a healthy culture of science, and astronomy is better for it."  He writes about how this state of affairs helps to avoid publication bias, because paper rejection is often influenced by reasons other than quality:  "scientific taste, politics, professional advantage, and science's inherent conservatism."  He even writes that "it's important that scientists be allowed to be wrong in the literature, as long as they have made no errors."  He concludes that "referees best serve when they act not as gatekeepers but as editorial consultants and independent voices that offer construcive criticism that improves submitted papers."

DTLR finds much merit in Wright's arguments.  However, I do think referees and editors have to be gatekeepers when it comes to one point that Wright alludes to, but does not emphasize.  I and many others have seen methodologically unsound research published, even in prestigious journals with low acceptance rates, in the life and social sciences.  A methdological critic would have been able to reject the paper before seeing any of the data.  This is a collective failing of authors, referees, and editors, who are often themselves untutored in even basic principles of research design and execution. There has been much discussion of these phenomena, for instance, in a special issue of the Lancet in 2014, and in closely related discussions of reproducible research (for instance).  Sadly, that this has continued to be a recognized problem for almost 15 years is a sign of an unhealthy culture of science.

One proposed remedy has much appeal:  "Results-Blind Manuscript Evaluation" (RBME), initially proposed by Joseph Locascio over 20 years ago (eg, Locascio, 2019).  This involves a two-stage manuscript review, where a mansucript is first evaluated on the basis of the Introduction and Methods sections, without knowledge of the results.  Methodologically unsound research can be rejected out of hand at this stage; if not, the manuscript moves to the second stage where its entirety is reviewed, though the decision to accept or reject may still not be based on the results, but only on the soundness of execution, analysis, and presentation thereof.  See the cited paper by Locascio for details.  RBME would not solve all the problems, but would go a long way in changing the incentives.

DTLR believes that a scientific journal should combine the insights of Wright and Locascio in its efforts to fight publication bias, while ensuring that research with a chance of contributing to, rather than misleading, the work of others, sees the light of day.

References


J. J. Locascio, 2019:  The impact of results blind science publishing on statistical consultation and collaboration.  The American Statistician, 73 sup 1:  346-351.

J. Wright, 2020:  High journal acceptance rates are good for science.  Physics Today, 73 (2):  10-11.

Wednesday, September 11, 2019

Science Magazine embraces replication (and gets its nose bloodied)

Science has put its money where its mouth is.  In this week's issue, the editor Jeremy Berg writes about "Replication Challenges".  They have published not one, but three NIH-funded replication studies of an earlier paper.  That is the good news.  The bad news is that a sister journal published a 10-patient human clinical trial, also based on the original paper, in parallel with the 3 replication studies.  Like all 3 replication studies, the human study was not able to reproduce the effect claimed in the original paper.

Can you see why journals may be reluctant to press forward with publishing replication studies?  They can certainly get their noses bloodied.  DTLR commends the editor of Science for restating his commitment to publishing replication studies.


Monday, November 12, 2018

Scientific American on "How to Fix Science"

The October 2018 issue of Scientific American features a set of articles under the umbrella, "How to Fix Science."  It begins with a prologue describing "antiscience currents" and arguing that one of the ways to "fight back" is to "tackle its own problems" and "shore up their enterprise from the inside."  The articles in the feature are as follows.

  • "Rethink funding" by John. P. A. Ioannidis.
  • "Make research reproducible" by Shannon Palus.
  • "End harassment" by Clara Moskowitz.
  • "Help young scientists" by Rebecca Boyle.
  • "Break down silos" by Graham A. J. Worthy and Cherie L. Yestrebsky.
The middle three articles are by science journalists, while the first and last are by academic scientists. Readers of DTLR will recognize that the first two articles discuss topics that have been kicked around on this blog over the years.  DTLR readers will also notice that I have been castigating Dr. Rush Holt, CEO of AAAS, recently, for avoiding these very issues.  Holt believes science is just fine the way it is.  DTLR has previously made a similar argument as the prologue of the Scientific American feature:  we need to clean up our own house before we can, with any good conscience, ask the public to continue financially supporting our endeavors.  Thus, I view the Scientific American piece as a rebuke to the Rush Holt model for responding to the "significant erosion of trust in science over recent years" (a phrase used by Worthy & Yestrebsky).

The leading piece by Ioannidis is a good starting point for this conversation.  Perverse incentives stimulate perverse outcomes.  His piece not only touches on funding issues, but hiring, promotion, and tenure decisions. While I disagree with many of the details in his proposed solutions, I believe the spirit of his piece is absolutely on target.  We need to fix the incentive systems in the scientific enterprise, and this is fundamental infrastructure.  Unfortunately the elites who thrive in the current scientific infrastructure have a vested interest in maintaining the status quo, and would oppose all Ioannidis' proposed reforms.


The article by Palus is a decent introduction to the non-reproducibility issue for the general reader, but lacks the space to delve into the methodological details of good experimental design, experimental quality control, and adequate reporting that are discussed, for instance, in Richard Harris' recent book, Rigor Mortis.

Moskowitz states that sexual harassment is more prevalent in academia than in any other arena except the military.  I am not surprised by this.  Having worked in various sectors of the economy - academia, industry, and government - my personal preference for a setting conducive to scientific research is the private sector, not academia.  In fact, the academic tenure system protects 'assholes' including those who practice harassment.

Boyle's article consists of comments by various young scientists on the following issues:  moving, money, culture, family, industry vs. academia, getting jobs/fellowships/into school, and representation and inequality.  I agree with the writer as well as Ioannidis that the treatment of young scientists in today's science ecosystem is appalling, and needs to be drastically reformed.

Finally, the article on interdisciplinary research is really another reflection of the perverse incentive systems for academic science.

I mentioned that I agreed with the spirit, but not many of the details of Ioannidis' piece.  Readers, what problems do you see in the incentive structure for scientists, and what solutions would you propose?



Tuesday, February 7, 2017

Videotaping experiments?

Today's issue of Nature has an interesting commentary by Timothy D. Clark, advocating videotaping of experiments.  His motivation is to combat scientific fraud.  However, there are broader reasons to consider the idea.  Clark alludes to some, and here is another.  Back when Google Glass came out, I read about a scientist who recorded the execution of her lab protocol simply as a means of documenting what was done.  This can be an extremely valuable supplement to a written protocol, as a video recording is more likely to capture "folk knowledge" within a lab, that nobody thinks to write down.  In other words, such practices could enhance reproducibility.

Of course, videotaping experiments is a long tradition in fluid dynamics, dating back to Henri Benard's films of what we now call von Karman vortex streets, in the first decades of the 20th century.  Today, image analysis methods are often used to extract data from moving images of fluids.

The other point made by Clark's post is that the burden of proof for allegations of misconduct is on the accuser, rather than the accused.  He makes a point about the trust-based nature of the scientific enterprise.  Shifting the burden of proof to the authors not only reduces the likelihood of fraud, but enhances the likelihood of reproducibility.    DTLR endorses Clark's proposal.


Sunday, January 1, 2017

Happy New Year

DTLR expects the new year to bring more scrutiny to non-reproducible research, and poor study design, conduct, analysis, and reporting.  Here is an example from last month in Nature.


Wednesday, November 30, 2016

"To err is human, but so often?" - David Freedman

Nature's editorial this week discusses the unleashing of the "statcheck" computer program on psychology journal articles.  Evidently it is an automated mechanism to detect errors in the calculation of p-values reported in published papers.

While I do not object to anything they said in the editorial, my concern is that there is very little penalty for carelessness in scientific research.  P-values are actually the least of my concerns; of greater concern are errors or even sub-optimal practices in the design, execution, and reporting of research.  Statistical inference is of no value if these other issues are present, and even when not present, statistical inference remains of incredibly limited value compared to a descriptive presentation of the data.  There are several reasons for this, such as:
  • Statistical inference presumes some kind of generalization, usually to a larger, stable population of which the data in the study can be thought of as representative.  This is rarely justified.
  • The statistical analysis adds information to the data  in the form of an assumed probability model.  This model's assumptions may well influence the outcome more than the data does.
  • Statistical inference is an inherently confirmatory activity, while most research is exploratory.  Statistical models in this context are overfitted to the data, and the generalization implied by statistical inference is invalid.
Nonetheless, sloppy calculation is a sign of carelessness, and for this reason the "statcheck" episode has certainly done a service, if it dis-incentivizes future carelessness.  On the other hand, legitimate criticisms of "statcheck's" own error rate have been raised.  I see this as the needed back-and-forth in the ongoing discussion of reproducible research, and accusations of "harassment" on the part of "statcheck's" creators are over-sensitive and unwarranted.



Saturday, September 24, 2016

The Atlantic weighs in on reproducible research

DTLR has dwelled on the reproducibility crisis since its birth.  Most of the discussions I have cited are from the scientific literature, but some have appeared in venues intended for a general audience.  One of the best I've seen has just appeared in the Atlantic, in a post by Ed Yong.  It appropriately places the focus on the incentive system for scientists.  A one sentence summary is provided by its quotation of Richard Horton, editor of the Lancet, who said:  "No one is incentivized to be right.  Instead, scientists are incentivized to be productive."

Please take a look at Yong's post.


Friday, July 22, 2016

Exploratory or confirmatory?

In last week's issue of Science, outgoing editor Marcia McNutt was interviewed (Shell, 2016) on the occasion of beginning a term as President of the National Academy of Sciences.  I am going to reproduce a lengthy quote from the interview.

At Science, the paradigm is changing.  We're talking about asking authors, 'Is this hypothesis testing or exploratory?'  An exploratory study explores new questions rather than tests an existing hypothesis.  But scientists have felt that they had to disguise an exploratory study as hypothesis testing, and that is totally dishonest.  I have no problem with true exploratory science.  That is what I did most of my career.  But it is important that scientists call it as such and not try to pass it off as something else.  If the result is important and exciting, we want to publish exploratory studies, but at the same time make clear that they are generally statistically underpowered, and need to be reproduced.

Bravo, Dr. McNutt!  DTLR agrees completely with the sentiment here.  It matters because the statistical dressing that accompanies much scientific research is usually only appropriate for confirmatory studies, or those that McNutt calls hypothesis testing, rather than hypothesis finding (exploratory).  It is rare to find the editor of a major scientific journal express this view in such a crisp, precise manner.   DTLR hopes that her successor, and other editors and referees of scientific journals, follow the lead set by McNutt.  DTLR also recommends all readers of this blog to take a look at Tukey (1980).

Reference


Ellen Ruppel Shell, 2016:  Hurdling obstacles:  Meet Marcia McNutt, scientist, administrator, editor, and now National Academy of Sciences president.  Science, vol. 353, pp. 116-119.

John W. Tukey, 1980:  We need both exploratory and confirmatory.  The American Statistician, vol. 34, pp. 23-25.

Wednesday, May 25, 2016

Nature keeps the heat up on reproducible research

This week's issue of Nature has a good article by Monya Baker on a wide-ranging survey of scientists about reproducible research, and a related editorial.  DTLR is most encouraged by the final table in Baker's article, the ratings of factors most likely to improve reproducibility.  "More robust experimental design" received the most combined "likely" and "very likely" ratings.  I think that this is the right answer.  Also highly ranked were "better mentoring/supervision" and "better understanding of statistics".  This latter one is a tough call, as statisticians themselves seem not to have reached a consensus on how to move forward, as evidenced in the extensive Discussion items published along with the American Statistical Association's Statement on Statistical Significance and P-valuesposted in early March.

DTLR expresses thanks to Nature for keeping the drums beating on reproducible research.  The issue is very visible right now, and the community should strike while the iron is hot, in terms of reforming the infrastructure of our community (laboratory practices, publication standards, and incentives for grant funding, promotion, and tenure).  Mis-aligned incentives are ultimately the cause of non-reproducibility, though methodological issues (poor study design and execution, inappropriate use of statistical methods, etc.) are key enablers.



Wednesday, September 2, 2015

The role of institutions in promoting reproducible research

Nature published a superb commentary yesterday by Glenn Begley, Alistair Buchan, and Ulrich Dirnagl, advocating reforms among institutions to help reduce irreproducible research.  I have little to add except unabashed praise.  DTLR endorses the views expressed in the commentary.




Sunday, August 9, 2015

Self-correction and an open research culture in science

Back in the June 26 issue of Science, there was a pair of commentaries by Alberts et al. (2015) and Nosek et al. (2015) entitled, respectively, "Self-correction in science at work" and "Promoting an open research culture."  Both articles are reactions to concerns about non-reproducible research.  The first one focuses on incentives, investigations of misconduct, and ethics education.  The second article deals with transparency, and introduces a set of guidelines called TOP (Transparency and Openness Promotion).  It discusses various standards for transparency, with different levels.  While I do not think this pair of articles captures the whole scope of the irreproducibility crisis, they do provide much food for thought on how we should address it.  Thus I recommend both articles to DTLR readers.

One passage in Nosek, et al. (2015), discussing the preregistration of studies and analysis plans, caught my eye.  "Preregistration of analysis plans certify the distinction between confirmatory and exploratory research, or what is also called hypothesis testing versus hypothesis generating research.  Making transparent the distinction between confirmatory and exploratory methods can enhance reproducibility."  I fully endorse this perspective.  I have met scientists who fail to understand the point made here so succinctly, and it really deserves emphasis.


References


B. Alberts, et al., 2015:  Self-correction in science at work.   Science, 348:  1420-1422.

B. A. Nosek, et al., 2015:  Promoting an open research culture.  Science, 348:  1422-1425.

Thursday, June 4, 2015

The grubby details of reproducing research

This week in Nature, Richard van Noorden writes of some preliminary findings of the Reproducibility Initiative:  Cancer Biology, presented at a conference in Brazil.  It is an interesting report, though I feel somewhat uncomfortable about the meta-analysis type statistical significance measure that they plan to calculate.  Nonetheless, the effort seems a worthy one, and DTLR awaits the final results.


Tuesday, May 19, 2015

More pieces of the nonreproducibility puzzle

A couple of recent news features have shed light on pieces of the reproducible research puzzle. Back in February, Jill Neimark in Science wrote about contaminated cell lines.  And just this week, Monya Baker in Nature wrote about batch-to-batch variation and non-specificity of antibodies.  Cells and antibodies are workhorses of modern biological research, and growing attention is needed to these potential sources of error.  I commend both of these articles to DTLR readers.

References


M. Baker, 2015:  Blame it on the antibodies.  Nature, 521:  274-276.

J. Neimark, 2015:  Line of  attack.  Science, 347:  938-940.


Wednesday, April 15, 2015

Scientific software reproducibility

Writing in Nature, Erika Check Hayden describes steps being taken by the journal Nature Biotechnology to improve the reproducibility of computational results.  She cites cases where error-prone software led to the publication of incorrect results in good journals.  Lack of software documentation seems to be one of the major issues; software testing and reproducibility were also mentioned.  The article also delves into the challenges faced by such a policy, such as finding qualified reviewers, and the danger of public shaming. 

DTLR feels that such challenges should not be allowed to block efforts to reform publication standards.  I endorse the stance taken by Nature Biotechnology, which can be found here

Sunday, April 12, 2015

Reproducible research in the Chronicle of Higher Education

Writing for the Chronicle of Higher Education last month, Paul Voosen covers the National Institutes of Health's work in encouraging reproducible research.  (The article is behind a pay-wall, so I have not linked it here.)   The NIH seems to be getting its act together.  The article points to universities, however, as the weak link.  "Indeed, more than any part of the scientific system, the universities have been ignoring the replication crisis," the article states, attributing the thought to Glenn Begley, author of the well known Amgen study of non reproducible research (Begley & Ellis, 2012).  This has the ring of truth to me.  As I've stated before, I commend the NIH for finally acknowledging the issue and taking steps to remedy it, some of which are detailed in the brief article.  The NIH and other funding agencies, as well as prominent journals, must necessarily take a top-down approach.  However this needs to be complemented by a bottom-up approach, the embracing of a cultural change by scientists in the trenches.  Such a change will be opposed by those who benefit under the status quo, as Arturo Casadevall suggests in the article.  The tight and thus highly competitive funding environment exasperates the problem, as his colleague Ferric Fang notes in the article. I would argue that scientists need to remember why they became scientists, and they need to be angry about the non-reproducible research that pollutes the academic literature.  However, this is a tough thing to say when labs have to fight for funding survival.  Anger and idealism don't do much good if your lab has gone out of business and you failed to achieve tenure.  It's a conundrum.

I highly recommend Voosen's article to DTLR readers.


Reference


C. G. Begley and L. M. Ellis, 2012:   Raise standards for preclinical cancer research.  Nature, 483:  531-533.

Paul Voosen, 2015:  Amid a sea of false findings, the NIH tries reform.  Chronicle of Higher Education, March 20, 2015, page A12.

Thursday, January 29, 2015

More on Reproducible research

Earlier this week, Joel Achenbach covered reproducible research in a Washington Post article.  This follows hot on the heels of the Science News series mentioned in my last post.  This is a topic that has received much discussion within the scientific literature, such as in Nature and Science, and bled into the popular press, for instance, with an Economist cover story in the fall of 2013, and a National Public Radio piece last fall by Richard Harris.  Achenbach's article doesn't have anything particularly new for those who have been following this thread over the last few years (including readers of this blog).  However, this topic deserves attention from major news organizations such as the Post and the Economist.  The taxpayers, after all, are bankrolling much of scientific research, and deserve to be kept in the loop on how their money is spent, or mis-spent as the case may be.

I also want to call readers' attention to an opinion piece last fall by John Ioannidis (2014).  It is a forward looking piece on how to make research more reproducible.  Much of the paper is focused on the infrastructure of the scientific community, including the incentive systems.  Ultimately this is indeed where change must occur.  He also has a list of "Some research practices that may help increase the proportion of true research findings."  Some of these are not explained in detail in this paper.  Third to last on his list is "Improvement of study design standards," an issue I feel is paramount.  Unfortunately Ioannidis does not go into great detail on this particular point, though it could deserve a paper all of its own.

My feeling is that, despite all the attention, reproducible research is not yet a big deal in the scientific community.  Scientists, and those who fund them, aren't angry enough yet to push for serious changes.  Until that happens, DTLR will not rest in promoting reproducible research practices.

Reference


John P. A. Ioannidis, 2014:  How to make more published research true.  PLoS Medicine, 11 (10):  e1001747.

Tuesday, January 27, 2015

Science news on reproducible research

Science News offers an excellent two-part examination of the woes of reproducible research by Tina Hesman Saey.  Here is part one and here is part two.  Highly recommended.

Thursday, November 6, 2014

Journals unite for reproducibility: DTLR is pleased!

Today Nature and Science posted joint editorials (here and here) endorsing a Proposed Principles and Guidelines for Reporting Preclinical Research, posted at the U.S. National Institutes of Health.  It is the product of a June, 2014, workshop sponsored by NIH and the two journals, and endorsed by over 30 other biomedical journals.  It provides a bare minimum list of criteria for good reporting practices of animal experiments in biomedical research.  The list is less detailed than one given by Landis et al. (2012) which they cite.

DTLR joins in endorsing the proposed principles and congratulates all the participants for taking a major step forward in promoting reproducible research.  It reflects a discipline-wide concern about the metastasis of non reproducible research across the spectrum of journals, as abundant evidence has made clear in recent years.  The proposed principles emphasize study design issues such as randomization, blinding, sample size, and appropriate replication.  Appropriately, such issues receive more space than analysis.  (The use of the term "inclusion/exclusion criteria" is a little confusing here - in clinical research, this refers to patient enrolment criteria; but the authors here seem to use it to refer to selective reporting and data omission, certainly an important issue, but one I would have found other language to describe.)  The sharing of data sets and the full disclosure of biological reagents are also welcome features.  In general, materials and methods sections of papers really should be expanded such that an independent laboratory could reproduce the experiment and expect to obtain similar results.

DTLR does not believe that the proposed principles go far enough, however.  For instance, the section on statistics requires a disclosure of "the statistical test used"; the fact that the emphasis here is on a statistical test rather than an estimation procedure is a serious oversight, in my view.  The reporting of confidence intervals instead of tests provides a sense of magnitude and direction that is lacking in a p-value, allowing evaluation of both clinical and statistical significance.  A statistical test outcome only communicates statistical significance. A confidence interval implicitly reports a test result when the confidence limits are compared with zero (for a conventional null hypothesis test).  Other types of estimation (tolerance intervals, prediction intervals) may be more appropriate in some situations.

DTLR is glad to see a growing consensus in the scientific community that nonreproducible research is corrosive and must be reduced.  This is a terrific step forward, but just one step.  Now, individual laboratories must take these guidelines to heart and use them to improve the design, execution, analysis, and reporting of their studies. 

References


Nature, vol. 515, p. 7 (2014).

Science, vol. 346, p. 679 (2014).

S. C. Landis, et al., 2012:  A call for transparent reporting to optimize the predictive value of preclinical research.  Nature, 490:  187-191.

Wednesday, May 28, 2014

Welcome indeed, Scientific Data

The Nature journals have recently launched a new journal, Scientific Data, for the publication of data sets with detailed descriptions.  Readers should take a look at the journal's website and an editorial in the main journal, Nature, here.  This is a welcome experiment in improving the infrastructure of science (including providing a new incentive for sharing data) and promoting reproducible research.