Skip to main content
Advertisement

< Back to Article

A complete statistical model for calibration of RNA-seq counts using external spike-ins and maximum likelihood theory

Fig 1

Diagrammatic summarization of the approach.

Asterix (*) on a variable or constant quantity means its is known by design or can be measured/calculated. Double asterisk (**) on a variable means its value is an estimate defined in this study. Index i = 1…s, denotes a spike-in, while i = s + 1…s + q, cellular RNA. For clarity in the diagram, these indices have been given the notation ERCC1…s and mRNAs + 1…s+q. The remaining mathematical notations in this figure follows exactly that of Table 1. (A) A fixed amount of spike-in RNA is added in fresh lysates from m cells in r repeats. The quantity of added spike-ins is known, and we want to calculate the quantity of endogenous mRNAs. (B) RNA is extracted from the lysates, RNA-seq libraries are prepared using a multi-step protocol, sequenced, aligned and count tables are constructed for spike-ins (i) and cellular RNA (ii). We use the spike-in count table together with the vector of spike-in abundance to estimate the library calibration factor ν, which is in turn applied for the estimation of nominal abundance of endogenous RNA in the sample. The mathematical definition of relative yield (α), and nominal abundance (z) are also shown. Note that the definition of z as a function of α cannot be estimated (**) as neither αmRNAs+1, nor nmRNAs+1, j are known.

Fig 1

doi: https://doi.org/10.1371/journal.pcbi.1006794.g001