Figure 1.
Systematic mining of publicly available biomedical text combines literature mining and the genome analysis of associate genes and phenotypes.
Words describing infectious disease related syndromes, names of pathogens and names of individual genes obtained from [32], [33] and CMR database are retrieved from PubMed abstracts [31]. Syndrome-pathogen and microbial genus-syndrome word pairs with a point-wise mutual information (PMI) scores, respectively, greater than 20 and 5 suggestive of informative associations which are visualized as ‘heat maps’. (A) Initial representation of raw counts of PMI scores for all pairs of syndrome-genus name of a pathogen. (B) Preferential co-occurrence of specific syndromes and pathogen names is emphasized by hierarchical clustering. (C) ‘Doppler’ map of syndrome-pathogen associations revealing clusters of increased interest, which are judged by the rapid growth in the number of publications. An example of highly clustered pathogen Mycobacterium is then explored further by building the heat map of associations between syndromes and the list of individual Mycobacterium species, then by zooming into the heat map of associations between the list of syndromes and individual genes of Mycobacterium tuberculosis H37Rv. Individual genes not mentioned in the searched collection of abstracts were not included in the heatmap. Clusters of associated syndromes and pathogens include many previously known relationships.
Figure 2.
Associations between individual Mycobacterium species and infectious disease related syndromes.
Figure 3.
Examples of species (A) and genome level (B) views of the tuberculosis landscapes with the selection of associations that suggest new hypotheses for testing. Peaks represent counts of co-occurrences of concepts. The table provides examples of associations between individual genes of M. tuberculosis H37Rv and syndromes (re)discovered in the experiment.
Figure 4.
Syndromic signatures quantify differences between pathogens.
Radar chart scale reflects the frequencies of co-occurrence of individual pathogen genus names and infectious disease syndromes.
Figure 5.
Clinically relevant reclassification of pathogens.
Maximum parsimony tree represents the distance matrix of associated syndromic signatures.
Figure 6.
Associational network representation of syndrome-pathogen relationships.
(a) Hypothetical network of syndromes and pathogens with edges representing the number of co-occurences in the text (mixed syndrome-pathogen network). (b) Hypothetical network of syndromes linking syndromes that co-occur with the same pathogens (syndrome only network); the length of the edge is inversely proportional to the number of shared pathogens. (c) Hypothetical network of pathogens linking pathogens that co-occur with the same syndromes (pathogen only network); the length of the edge is inversely proportional to the number of shared syndromes. Syndrome-pathogen (d), syndrome (e) and pathogen (f) networks built from text mining experiments; the size of a node represents the number of citations. The most highly connected nodes are in the middle. The edge per node distribution for each network is shown in respective graphs (g, h, i). Specific subsets of networks and lists of the most connected nodes can be found in SOM.
Figure 7.
Associational networks of syndromes and pathogens.
Examples of associational networks of syndromes (A), pathogens (B) and genes of an individual species (C). Node size depics the number of citations. Edge distance between two nodes is inversely proportional to the number of pathogens (A), syndromes (B and C) shared by concepts of nodes. The thickness of an edge is proportional to the normalised number of co-occurrences. Minimum number of co-citations with other pathogens for each entity in the network A to be included in this network was 15. The individual species gene associational networks (B) are presented using M.tuberculosis H37Rv genome as an example. Gene names and gene locus numbers are presented.