Figure 1.
Spurious alignments found by BLAST.
This is the output of a blastp search with a reversed protein (B6D5L7_PERAZ) against the nr database at NCBI.
Figure 2.
A NUMT in the X chromosome of C. elegans.
This shows an alignment between the X chromosome (upper) and the mitochondrial chromosome (lower). Lowercase red letters were masked by tantan. The blue arrowheads indicate the first unit of an inexact tandem repeat.
Table 1.
Example of an enlarged score matrix for gentle masking.
Table 2.
NUMTs found with gentle or harsh masking.
Figure 3.
A metagenomic DNA read aligned to a bacterial genome.
The upper sequence is the DNA read “1_lane2_104963”; the lower sequence is from the genome “A1-86”. Lowercase red letters were masked by tantan.
Figure 4.
Alignments between a protein and the human genome.
This shows two local alignments between a protein (Q494U1, upper sequence) and human chromosome 1 (lower sequence). Lowercase red letters were masked by tantan. The upper alignment was found with gentle masking, but not with harsh masking.
Figure 5.
Alignments of reversed sequences, with gentle masking.
This shows alignments between: (A) the C. elegans genome and the reversed P. pacificus genome; (B) the A. thaliana genome and the reversed P. patens genome; (C) vertebrate proteins and reversed plant proteins; (D) the human genome and the reversed opossum genome; (E) the P. falciparum genome and the reversed D. discoideum genome; (F) the P. falciparum genome and the reversed human genome. The colors indicate alignments after: masking both sets of sequences (solid red); masking the first-named set only (dotted magenta); masking the second-named set only (dashed blue); shuffling the letters in each set (dashed brown). The black lines indicate the expected number of alignments for random sequences.
Figure 6.
Alignments of reversed sequences, using the HOXD70 scoring scheme.
Alignments between: (A) the C. elegans genome and the reversed P. pacificus genome; (B) the A. thaliana genome and the reversed P. patens genome; (C) the human genome and the reversed opossum genome. The colors indicate alignments after: masking both sets of sequences (solid red); masking the first-named set only (dotted magenta); masking the second-named set only (dashed blue); shuffling the letters in each set (dashed brown). The black lines indicate the expected number of alignments for random sequences.
Figure 7.
Alignments between DNA sequences and reversed protein sequences, with gentle masking.
This shows alignments between: (A) the C. elegans genome and reversed plant proteins; (B) the P. falciparum genome and reversed vertebrate proteins. The colors indicate alignments after: masking the proteins, and the DNA at the protein level (solid red); masking the proteins, and the DNA at the DNA level (solid blue); masking the proteins only (dashed red); masking the DNA only, at the DNA level (dashed blue); shuffling the letters in each set (dashed brown). The black lines indicate the expected number of alignments for random sequences.
Figure 8.
Alignment problem using a mask score of 0.
This kind of nonsensical alignment may occur if masked letters (lowercase red) always receive a score of 0.
Figure 9.
Alignments using BLOSUM80 or BLOSUM62.
(A) An alignment that is significant when scored with blosum80 but not blosum62. (B) Two local alignments between a protein (Q9BZQ4, upper sequence) and human chromosome 1 (lower sequence). The blue arrowheads indicate spurious over-extension of the second alignment. This over-extension occurs when the blosum62 matrix is used for alignment, but not when blosum80 is used. (Actually, it is conceivable that the extension correctly indicates homology: the genomic segment marked by blue arrowheads could be paralogous to the genomic segment in the first alignment.)