Reader Comments

Post a new comment on this article

Incorrect data analysis

Posted by janhul on 07 May 2007 at 07:45 GMT

Apparently, the paper exhibits at least four flaws, which render the authors’ conclusions invalid.

First, in the microarray study, M-values were calculated to estimate differential expression due to treatment. Data from each group of 8 birds (same generation and population) were collapsed, because the observations were not matched individually, but groupwise. In this way, all information on the relative expression of genes in individual birds is lost and the variation between experimental units (the basis of statistical inference) cannot be estimated. In the analysis that follows, the observations are pairwise M-values from the different spots, but there is only one group of animals, i.e. essentially one single observation, behind all data-points. Instead of using M-values, it would probably be wise to keep the raw values of relative expression as the outcome trait, thus utilizing information from each bird.

Second, correlation analysis, as in the study of associations between M-values in parents and offspring, requires that observations are independent. It is very reasonable to assume that there is some dependence between M-values of different spots, i.e. that the expressions of different genes are related to each other. Hence, the observed association probably reflects the degree of dependence among spots more than anything else.

Third, sub-samples of the 9,033 spots were selected randomly to calculate a mean correlation coefficient and a standard error of that mean. However, the sub-samples are not independent either, primarily because of dependence among spots, but also because some sub-samples even share the same spots. The number of sub-samples is of course arbitrary; nevertheless, it is crucial when calculating standard errors and drawing further inferences. Therefore, the sub-sampling approach does not solve anything – it only makes things worse.

Fourth, learning ability was tested and the cumulative proportion of successful birds at each test round was compared between treatment and control groups by chi-square analysis. No adjustment was made for multiple comparisons at different tests. The graphs in Figure 1 allow us to reconstruct raw data with a fair degree of accuracy and re-run the analysis with such an adjustment. Alternatively, considering the time-to-event type of data, it seems more straightforward to test the effect of treatment using the log-rank test. With either analytical approach (both were tried by me), there is no significant difference between treated and control White Leghorn offspring (P>>0.05), as the authors claim. I did not check the learning ability results of the parental generation.

Cheers,
Jan Hultgren

RE: Incorrect data analysis

perje replied to janhul on 24 May 2007 at 07:58 GMT

We thank Jan Hultgren for his important and stimulating comments on our paper, and would like to offer the following responses:

1. I am afraid we do not fully understand the objection here. Underlying the M-values are of course data on the relative expression level of each spot from each individual. For every individual (4 per treatment group), one microarray (plus technical replicates) were run. The M-value, by definition, estimates the difference in average expression between two treatments or groups of animals. The statistical analysis developed for microarray analysis is based on the variation between experimental units (biological replicates), and we therefore continue to believe that the analysis we have done is the most appropriate one.

2. Hultgren is correct in stating that the spots on one micro-array are inter-dependent, since regulatory networks will affect many genes in the same way. However, the main point of the correlation analysis is that there is a dependency of M-values (which, as already said, estimate the average differential expression) between generations, something which is not biologically obvious. The important message here is that this dependency was only seen in the population of birds where we simultaneously saw a phenotypic effect.

3. The randomization procedure used for calculating an average coefficient of correlation does of course not render independent datasets, as Hultgren rightly points out. However, we think this was the best available type of analysis for our purpose. We could have simply calculated a correlation coefficient based on all 9033 pairs of M-values from the two generations, but this would be a biased analysis with an inflated p-value due to the facts that there were so many data-points, and that most of them had small absolute M-values. We could also have decided to look at only the most differentially expressed genes (those shown in the fig), but again this would have rendered a biased r. We therefore used a randomization procedure and picked random subsets of M-values as described in the paper. Even if the randomly generated datasets are not independent, they do provide a clear indication that the M-values of the two generations are correlated. However, we agree that the SD and the actual p-value of the t-test must be treated with care.

4. As we do say in the paper, there was no overall effect of treatment on the average (or median) number of tests needed to reach the learning criterium in the offspring, which is probably why Hultgren does not find any significant result in his analysis. However, by visual inspection, it is obvious that the learning curves appear quite different between Leghorn offspring of stressed and control birds. After 5,6 and 7 rounds of testing, the differences between treatment groups were visually striking, and we therefore decided to test the differences in proportion of birds having solved the task at the points which appeared different. We did not apply a correction for multiple testing, because we only tested three data-points in the offspring, and in the parental generation, the differences were already shown by an ANOVA. We admit that the effects are small in the offspring, but still clearly discernible.

Cheers,
Per Jensen, on behalf on the authors.

RE: RE: Incorrect data analysis

ErikiLund replied to perje on 13 Nov 2007 at 21:04 GMT

I have read this paper with great interest, and we actually used it in a seminar about evolutionary biology at the Department of Ecology in Lund. The students and myself found the paper interesting and the results suggestive. Hopefully, this paper will stimulate further research in this area and this study is a first step.

However, one of the students pointed out that the experiment was not replicated, i. e. all individual parents subject to stress-treatment were in one batch. If this criticism is correct, I believe it is a serious statistical problem with the current study. This problem could easily have been avoided if the the treatments had been replicated, but according to my reading of the paper, they have not. It is possible that this is a misunderstanding on my part, but I do not think so, at least when reading through the paper.