Skip to main content
Advertisement

< Back to Article

Table 1.

Patient data.

More »

Table 1 Expand

Fig 1.

De novo assembly of var gene sequence.

(A) Expression profiles for the ItG subclone E8B. The assembled transcripts were annotated with their closest BLAST match to the IT4 (a clone of the ItG isolate) sequences from the database of [14]. The expression levels in RPKM are then compared to RPKM levels of reads annotated directly to the whole gene DNA sequences of [14] and to those obtained using qPCR. (B) Var gene chromosomal arrangements, group A and B var genes are present in subtelomeric clusters, group C var genes are present in chromosome internal var gene clusters. The different resolutions of var sequences investigated in this manuscript are illustrated. Var gene transcripts are obtained by de novo assembly of transcriptome data. Domain regions are then identified within these transcripts along with smaller subdomain segments and homology blocks. The number and order of both domains and segments varies between var genes. ATS, acidic terminal sequence; BLAST, basic local alignment search tool; CIDR, cysteine-rich interdomain region; DBL, Duffy binding-like; NTS, N-terminal sequence; qPCR, quantitative PCR; RPKM, Reads Per Kilobase of transcript per Million mapped reads; TM, transmembrane.

More »

Fig 1 Expand

Fig 2.

Genome-wide analysis of RNAseq data using 3D7 annotation.

(A) Estimated stage proportions for each sample. The mixture model was constrained to require that each sample be made up of a combination of ring, early trophozoite, late trophozoite, schizont, and gametocyte stages. Consequently, the columns in this barplot must add to 1 for each sample. A small bias towards the early trophozoite appears in the nonsevere malaria samples. Sample SFC21 also appears to be an outlier due to its higher proportion of late-stage and gametocyte parasites, a finding which was confirmed by microscopy. Plotted proportions are available in S2 Data. (B) A PCA plot of read counts normalised for library size (read counts are available in S2 Data). Samples are coloured by phenotype, red for severe and blue for nonsevere. Some separation by disease severity phenotype is evident; however, staging effects are apparent as is seen in the outlying position of sample SFC21, which has been identified as having more late-stage and gametocyte parasites. (C) A PCA plot of read counts normalised for library size, staging effects, and other unwanted batch effects using the novel mixture model along with 3 unwanted factors of variation estimated by RUV4 (normalised read counts are available in S2 Data). Sample SFC21 has been appropriately dealt with and a better separation of the samples by disease phenotype can be observed. PC, principal component; PCA, principal component analysis; RUV, Remove Unwanted Variation.

More »

Fig 2 Expand

Fig 3.

Gene sets enriched in deregulated genes in severe malaria.

(A) Summary of highly ranked GO and KEGG gene annotation pathways that included significantly deregulated genes in severe malaria. Only gene sets that contained more than 1 deregulated gene are shown; deregulated gene set data available in S3 Data, deregulated genes available in S1 Data. (B) The glycolysis pathway in P. falciparum in severe malaria. Fold-change in gene expression in severe malaria relative to uncomplicated malaria (x) and p-value for the fold-change are indicated beside genes. Genes that were significantly (adjusted p < 0.1) down-regulated in severe malaria are indicated in red. (C) LC-MS metabolomic analysis of plasma samples from patients with severe and uncomplicated malaria. Ion counts for metabolites commonly affected by malaria are presented; data available in S4 Data. adj-p, adjusted p; GO, Gene Ontology; KEGG, Kyoto Encyclopedia of Genes and Genomes; LC-MS, liquid chromatography–mass spectrometry; logFC, log fold-change; uncompl, uncomplicated.

More »

Fig 3 Expand

Fig 4.

Analysis of RNAseq data at the level of var gene transcripts: Combined assembly.

(A) Expression levels of transcripts from the combined sample assembly found to be up-regulated in severe disease. Samples and clusters have been grouped using complete linkage hierarchical clustering (raw read counts available in S6 Data). (B) Expression levels of transcripts from the combined sample assembly found to be up-regulated in severe disease; values for all samples and the IQR and median are indicated. RPKM is reads per kb of transcript per million reads mapped to total var transcripts (RPKM available in S6 Data). IQR, interquartile range; RPKM, Reads Per Kilobase of transcript per Million mapped reads.

More »

Fig 4 Expand

Fig 5.

Analysis of RNAseq data at the level of var gene transcripts: Separate assembly.

(A) Expression levels of clusters identified by Corset found to be up-regulated in severe disease. Samples and clusters have been grouped using complete linkage hierarchical clustering. Raw read counts are available in S7 data. (B) Expression levels of clusters identified by Corset found to be up-regulated in severe disease. Values for all samples and the IQR and median are indicated and are available in S7 Data. RPKM is reads per kb of transcript per million reads mapped to total var transcripts. IQR, interquartile range; RPKM, Reads Per Kilobase of transcript per Million mapped reads.

More »

Fig 5 Expand

Fig 6.

Summary of PfEMP1 transcripts, domains, and segments that were up-regulated in severe malaria.

Sequences up-regulated in severe malaria are organised in columns for each analysis method separated by grey bars. Multiple domains found in the same single transcripts from the combined or separate assemblies are on a single row. Closely related sequences found in multiple analyses are colour coded for each of the major domain types and are grouped together across analyses by unbroken horizontal lines. Domains and/or segments that clustered together by expression profile in multiple individuals within a single analysis are also grouped by unbroken horizontal lines. Grey shaded sequences at the bottom of the diagram are unrelated to each other. For example, in the case of DC4, 2 transcripts from the combined assembly were amongst the closest BLAST hits to the DC4-like transcripts from the CORSET cluster of the separate assembly; 6 domains and 5 blocks identified by HMM in the separate assembly are found in DC4 domains; and clusters for 1 domain and 4 segments identified by hierarchical analysis contained DC4 domain sequences, including those from the DC4-like transcripts from the CORSET cluster of the separate assembly. aCombined assembly transcripts up-regulated in severe malaria were all adjusted p < 0.05 except for domains marked b (adjusted p < 0.153). Domains HMM and blocks HMM were identified using the HMM of [14]. Domains and segments %ID were identified using the novel hierarchical approach developed for this study. cNon–DC8-like DBLδ1 and non–DC4-like DBLβ3 that clustered by expression profile in the same patients with a highly conserved CIDRβ1. A dashed line separates DBLβ12 from DC8 because DC8 typically contain DBLβ12, but these DBLβ12 formed a phylogenetic cluster with non-DC8 DBLβ12. Dashed lines separate putative DC9 components because transcripts containing all components were not up-regulated in the combined assembly or the Corset analysis, but the clusters from which the up-regulated segments were identified contained multiple transcripts carrying the DC9 domains. ATS, acidic terminal sequence; CIDR, cysteine-rich interdomain region; DBL, Duffy binding-like; DC, domain cassette; HMM, Hidden Markov Model; PfEMP1, Plasmodium falciparum Erythrocyte Membrane Protein 1; TM, transmembrane.

More »

Fig 6 Expand

Fig 7.

Analysis of RNAseq data via de novo assembly at the level of var gene domains.

(A) Expression levels of domain subfamilies from [14] found to be up-regulated in severe disease as identified using HMMER3 models. These models were built from the domain sequences of [14]. Samples and clusters have been grouped using complete linkage hierarchical clustering. (B) PCA plot of read counts that align to domain regions of the de novo–assembled transcripts identified using HMMER3 models. There is less separation by phenotypes in this plot than was observed at the whole-transcript–and all-gene–analysis levels. Read count data for Fig 7 is available in S8 Data. (C) An example of the hierarchical clustering tree. Colours represent significance, with red indicating a significant difference in expression after multiple testing correction and blue indicating not significant. Nodes are coloured grey if there is insufficient evidence for them to be considered in the testing either because they have less than 5 samples present or they are marked by DESeq2’s prefilter step. At the 60% identity level, cluster 670_X0.6 becomes significant. This significance is then obscured at the 50% identity level, demonstrating the importance of considering different levels of the hierarchy. (D) Clustering the domain level counts at 50% sequence identity rather than using the previous classifications of [14] improves the grouping of severe samples. At 50% identity, the severe samples are grouped more closely together, suggesting that they have more in common than the nonsevere samples; transformed read count data available in S9 Data. PCA, principal component analysis; RNAseq, RNA sequencing.

More »

Fig 7 Expand

Fig 8.

Analysis of RNAseq data via de novo assembly at the level of var gene domains: Hierarchical analysis.

(A) Expression levels of the domain clusters identified using the hierarchical approach. Samples and domains are grouped using complete linkage hierarchical clustering. The colourings on the left indicate notable groups identified using the hierarchical cutting algorithm of [95]. The clusters are also annotated with the domain model of [14] that they most closely resemble. Raw read counts are available in S9 Data. (B) Expression levels of the domain clusters identified using the hierarchical approach, values for all samples, and the IQR and median are indicated. RPKM is reads per kb of domain per million reads mapped to total var domains (RPKM data available in S9 Data). IQR, interquartile range; RPKM, Reads Per Kilobase of transcript per Million mapped reads.

More »

Fig 8 Expand

Fig 9.

Analysis of RNAseq data via de novo assembly at the level of var gene segments.

Expression levels of novel conserved segment clusters found to be up-regulated in severe disease. Samples and segment clusters have been grouped using complete linkage hierarchical clustering. The raw read counts that were transformed for this figure are available in S12 Data.

More »

Fig 9 Expand