Showing posts with label biology. Show all posts
Showing posts with label biology. Show all posts

Wednesday, December 29, 2021

"New data from old poop"

Don't you want to learn about changes to our microbiome over time, by examining paleofeces?  You do.  See the commentary by Andrew Curry from earlier this year.



Tuesday, May 19, 2015

More pieces of the nonreproducibility puzzle

A couple of recent news features have shed light on pieces of the reproducible research puzzle. Back in February, Jill Neimark in Science wrote about contaminated cell lines.  And just this week, Monya Baker in Nature wrote about batch-to-batch variation and non-specificity of antibodies.  Cells and antibodies are workhorses of modern biological research, and growing attention is needed to these potential sources of error.  I commend both of these articles to DTLR readers.

References


M. Baker, 2015:  Blame it on the antibodies.  Nature, 521:  274-276.

J. Neimark, 2015:  Line of  attack.  Science, 347:  938-940.


Sunday, March 30, 2014

Surface tension and biology

This blog has tended to focus on methodological issues in science and medicine, but occasionally I do want to lavish praise for substantive work.  Here I'd like to call readers' attention to the delightful article in Science a couple weeks ago by Elizabeth Pennisi, "Water's Tough Skin."  It is a feature article describing a number of ways that surface tension is important in biology, including for plants, animals, and microbes.  A number of scientists, engineers, and mathematicians are interviewed about their work.  This is the kind of article that reminds us of why we became interested in science, engineering, and medicine in the first place.  I won't review the content here, but I commend it to readers to enjoy for themselves.

Reference


Elizabeth Pennisi, 2014:  Water's tough skin.  Surface tension is a force to be reckoned with, especially if you are small.  Science, 343:  1194-1197.



Saturday, February 1, 2014

Responses to "When Mice Mislead"

This past week's issue of Science (the Jan. 24, 2014 issue) has two letters to the editor, responding to a report last November, "When Mice Mislead" by Jennifer Couzin-Frankel, which I discussed in an earlier post.  The first letter, by Richard Traystman and Paco Herson, points to earlier findings, similar to those reported by Couzin-Frankel, in the stroke research community.  Most importantly, they assert that "It is unlikely that poor methods used in animal studies account for all the negative clincial trials that have been performed based on preclinical studies.  After all, some investigators do perform appropriate experiments, and even those studies rarely lead to positive clinical trials."  The authors point to the fact that mouse studies are usually done with healthy young mice, whereas human subjects in neuroprotective drug clinical trials are often older and have many co-morbidities.  They propose that aged mice with comorbid diseases be used in stroke trials, as a better animal model of human disease.

The second letter is from statistician Gary Churchill.  He zeroes in on one key question:  "Was the result replicated in more than one genetic background?"  He goes on to identify two "root causes" for nonreproducible research:

Science today is driven by an incentive system that often rewards precedence and impact over quality of the work.  Statistical training of scientists often emphasizes analytical techniques over experimental design and quantitative reasoning.  These are systemic problems that will not change without substantial effort
Meanwhile, Churchill endorses the message of Couzin-Frankel's article with his maxim:  "Be wise, randomize."

I think that both of these letters add value to the original piece by Couzin-Frankel. In particular, Churchill's second "root cause" is particularly interesting, as both statisticians and lay scientists or mathematicians who teach statistics are all guilty of overemphasizing methodology, modeling, and inference at the expense of study design and critical thinking. 

References

Jennifer Couzin-Frankel, 2013: When mice mislead. Science, 342: 922-925.

Richard J. Traystman and Paco S. Herson, 2014:  Misleading results:  translational challenges.  Science, 343:  369-370.

Gary Churchill, 2014:  Misleading results:  don't blame the mice.  Science, 343, 370.


Monday, January 20, 2014

When mice mislead

Two months ago, the Nov. 22, 2013, issue of Science announced the detection of high energy neutrinos from beyond the solar system. This of course is a major achievement for physicists. However, my attention was drawn to another story in the same issue, “When mice mislead” (Couzin-Frankel, 2013). I regard it as one of the most important works of science journalism of the year just ended.

The article documents the following “bad habits” in studies of laboratory animals, such as mice, that can lead to non-reproducible results and misleading conclusions. The bad habits discussed include:
  • Removing data, such as from animals enrolled in a study but removed from the analysis for any number of reasons.
  • Lack of randomization and blinding.
  • No attention paid to inclusion/exclusion criteria for enrolling animals.
  • Different experimental conditions for different groups of animals.
  • Sample sizes too small to lead to definitive results, since researchers have very good reasons (ethical and financial) to minimize the number of animals used for research.
  • Publication bias, along the lines of Ioannidis (2005).
A good example is discussed by Lisa Bero, interviewed in the article. About scientists and their mentors, Bero states that “Their idea of randomization is, you stick your hand in the cage and whichever one comes up to you, you grab. That is not a random way to select an animal.” Couzin-Frankel goes on to say that “Some animals might be fearful, or biters, or they might just be curled up in the corner, asleep. None will be chosen. And there, bias begins.”

All of these bad habits are ones that have been largely eliminated from randomized clinical trials. Lab animal studies are traditionally pursued with far less rigor than clinical studies, but the article suggests that the good habits that dominate clinical trials could really clean up preclinical research if they were to be adopted widely there too. In other words, it's time to raise the level of the game in lab animal studies, and practically it wouldn't take much additional effort to do so. In fact, this very article will help those of us who try to push back on bad habits. We now have a convenient summary of the findings that such bad habits really matter, and should be avoided.

Surprisingly, Lisa Bero found that industry funded research is less likely to endorse a drug that research funded through other means, “maybe because companies don't want to pour millions of dollars into testing a treatment in people that's unlikely to help them.” This certainly has the ring of truth; however, I think even within industrial labs, the good habits of clinical trials are not always pervasive among users of lab animals.

Joseph Bass is also interviewed with a less pessimistic view. He believes there are substantive reasons why many mouse studies fail to reproduce, such as the temperature that mice are housed at, or variation in their response with age.  However, he seems too optimistic.  Nonreproducible research is more pervasive than I would like, and while fitness for purpose as a decision criterion should always over-rule "one size fits all" rules and checklists, I think such rules and checklists will do more good than harm at this point in the history of science.

Couzin-Frankel (2013) refers to the ongoing effort by the NIH to draft rules for research it funds, to encourage openness and reproducibility, as well as the checklist for biology research promulgated by Nature last year (discussed here). She even says that Science is considering a similar policy! An NIH official is quoted as saying “Sometimes the fundamentals get pushed aside—the basics of experimental design, the basics of statistics.” Amen!  This quote summarizes the problem in a nutshell.

I have been critical of Science as a follower, not a leader, on reproducible research. However, their publication of Couzin-Frankel's report goes a long way to earning forgiveness.

References


Jennifer Couzin-Frankel, 2013: When mice mislead. Science, 342: 922-925.

John P.A. Ioannidis, 2005: Why most published research findings are false. PLoS Medicine, 2 (8), e124: 696-701.



Friday, November 15, 2013

bioRxiv goes live

In an earlier post I mentioned the forthcoming preprint server for the life sciences, bioRxiv.  According to Nature, the site has now launched.  See the write-up by Ewen Callaway here.



Sunday, October 27, 2013

Software validation in computational biology

Last month in Nature, there was a brief article by Erika Check Hayden about an experiment in peer review of scientific software being carried out by the new Mozilla Science Lab. Nine papers published in PLoS Computational Biology, selected by its editors, would have their code subject to a peer review by software engineers. The experiment and its motivation are described in the article; I also recommend reading the user comments posted at the end. (See also the earlier pieces by Zeeya Merali and Nick Barnes, published together in Nature in 2010.)  Apparently there has been some controversy, as scientists are understandably nervous about having their work subjected to a new form of review. However, scientists are not well trained in software development concepts such as version control, validation, and verification, and the code they write may become difficult to maintain, or even worse, produce undetected errors that have worked their way into published research.

The Mozilla Science Lab was introduced this past summer, and is led by Kaitlin Thaney. It sponsors Greg Wilson's Software Carpentry; I strongly recommend having a look at the latter's website. I've read Wilson's essays in Computing in Science and Engineering and other publications over the years, and have been sympathetic to his views. I've heard rumors that the some of the code at Fermilab is spaghetti code, with bits and pieces of it written by many hands over many decades. Such an unwieldy mass of legacy code is almost impossible to maintain. I was told about one bug whose fix generated another, more serious bug that was impossible to debug. It was decided to restore the original bug and leave it in the code!

I am fortunate in that one of my formative experiences was an internship with a small company that, as a matter of survival, implemented a fairly disciplined software construction methodology, based in part on Steve McConnell's Code Complete. Because the company was small and had a certain rate of turnover, all of their software had to be highly maintainable, assuming the original coder was no longer employed at the firm. It was a point of pride there that you wouldn't be able to tell who wrote a piece of code found in the software they developed, without looking at the header (which had version control data), for we all conformed to the same software style.

DTLR endorses the Mozilla experiment in peer review of software. I hope we learn a lot from their experiment, even if it is deemed to be a failure in the end. In a letter to the editor, Alden and Read (2013) state that software quality should be built in from the beginning, before any data are taken, and not “inspected in” at the peer review stage. They are of course right, but to protect the rest of the community I do think software peer review is a concept that should at least be explored.

References


Nick Barnes, 2010: Publish your computer code: it is good enough. Nature, 467: 753.

Zeeya Merali, 2010: Computational science:...error. Nature, 467, 775-777.

Erika Check Hayden, 2013: Mozilla plan seeks to debug scientific code. Nature, 501: 472.

Kieran Alden and Mark Read, 2013: Scientific software needs quality control. Nature, 502: 448.


Biology's dry future

A few weeks ago, Science magazine featured a very interesting story by Robert F. Service titled “Biology's Dry Future.” The subtitle tells us, “The explosion of publicly available databases housing sequences, structures, and images allows life scientists to make fundamental discoveries without ever getting their hands 'wet' at the lab bench.” The story highlights two quotes from interviews. The first is by Atul Butte of Stanford University School of Medicine: “I'm like a kid in a candy store. There is so much we can do.” The second is by David Heckerman of Microsoft Research: “You basically don't need a wet lab to explore biology.”

The title of the story is not quite accurate. There will always need to be wet lab biologists to do experiments and generate data. What is novel here is the new breed of biologist who works on data generated by other labs, but need not have a lab themselves. This may be new to biology, but physicists have long had a split between experimentalists and theorists, recently joined by computationalists.

Of particular interest to DTLR are the three “growing pains” mentioned in the article: data access, data standardization, and genetic privacy. I will focus on the first two here. Regarding data access:

In many cases, researchers who have spent their careers generating powerful data sets are reluctant to share. They may be hoping to mine it themselves before others make discoveries based on their work. Or the data may be raw and in need of further analyses or annotation. “These are really hard problems,” Butte says. “We need better systems to reward people that share their data.”

DTLR endorses that last sentence. First of all, anyone who makes the effort to generate a good data set should make the effort to document and annotate it for use. Even if the data are never shared, pretending that it might be shared one day instills the necessary discipline for documentation and annotation. Moreover, if the work is publicly funded, then in my view the social contract requires that the data be made available to the broader scientific community at some point, perhaps after an appropriate time period of exclusive use, say, no more than two years. (This is about the time needed for a grad student or post-doc to squeeze at least one paper out of results.) The new Nature online journal for data sets would provide an excellent venue to generate a peer reviewed publication for the data set alone, rather than discoveries that can be made with it. Bear in mind that in physics, Nobel prizes are awarded to both theorists and experimentalists. Biology as a discipline should adopt a similar cultural mindset to reward both wet bench and dry bench biologists.

Regarding standardization:

Not only do research groups file their data using different software tools and file formats, but also in many cases the design of the experiments—and therefore precisely what is being measured—can differ. Butte and others argue that dealing with multiple file formats is somewhat cumbersome but that the problem is surmountable. But it can be harder to account for differences in experimental design when comparing large data sets.

DTLR could not have said it better. The core problem here is experimental design, and it will always be a limiting factor for dry lab biologists trying to combine data from more than one experiment. A similar problem exists in clinical medicine, under the term 'meta-analysis', and I'm not sure there are really good solutions there either. The best approach, in my view, is to take any findings based on multiple data sets as tentative, exploratory, and hypothesis-generating, rather than definitive. The findings should then be confirmed (or refuted) in a new experiment. This is where the dry lab biologist might have to return to the bench.

Finally, DTLR cautions that dry lab biologists should still spend some time in the lab, at least while in training. There is no substitute for bench time for getting a feel for how sloppy and imprecise experimental data can be, and where the pitfalls and potential systematic and random errors may arise from. It is too easy for a dry bench scientist to take data found in a database at face value. Spending time at the bench will provide a needed reality check.

Reference


Robert F. Service, 2013: Biology's dry future. Science, 342: 186-189.