Fig 1.
Diagrammatic illustration of the key steps of Seek & Blastn (S&B).
S&B extracts facts to check from published text (nucleotide sequences with associated targeting/ non-targeting status), and then performs blastn analyses and fact checking.
Fig 2.
Example of Seek & Blastn (S&B) output for a retracted Corpus P paper (Ref. [22]).
Columns shown from left to right are: “Tested file”; “Nearest dist”, which provides intertextual distance analysis results [9]; “Genes”, which provides gene and species identifiers extracted from the text; “Cont. CL”, which provides identifiers that correspond to contaminated or misidentified cell lines; and “Sequences”, which lists all nucleotide sequences that were extracted, and their corresponding blastn results. The tested publication forms part of the reference cohort [16], and its closest match is the same publication within the reference cohort. The tested pdf included 6 nucleotide sequences, which were correctly extracted and identified by S&B. Two extracted sequences were recognized as a previously reported sequence (SeqA) [16], and a mismatch was detected between the claimed non-targeting status and the blastn identity, as shown in red hypertext.
Table 1.
Descriptions of Corpus P and Corpus U analysed by Seek & Blastn.
Table 2.
Seek & Blastn nucleotide sequence and associated status extraction (targeting versus non-targeting) from Corpus P and Corpus U publications.
Fig 3.
The proportions of Seek & Blastn (S&B) status predictions for 304 correctly extracted Corpus P sequences that were either confirmed or refuted by manual analyses.
Predictions were classified as either true negative, false negative, true positive or false positive outcomes. The numbers of targeting and non-targeting sequences for each of the 4 possible outcomes are listed separately. Where sequences were correctly flagged by S&B as true positives, “Targeting” and “Non-targeting” refer to the incorrect claimed status in the relevant publication.
Fig 4.
The proportions of Seek & Blastn (S&B) status predictions for 1066 correctly extracted Corpus U sequences that were either confirmed or refuted by manual analyses.
Predictions were classified as either true negative, false negative, true positive or false positive outcomes. The numbers of targeting and non-targeting sequences for each of the 4 possible outcomes are listed separately. Where sequences were correctly flagged by S&B as true positives, “Targeting” and “Non-targeting” refer to the incorrect claimed status in the relevant publication.
Table 3.
Numbers and proportions of Corpus P and Corpus U papers that were correctly or incorrectly flagged by Seek & Blastn (S&B).
Table 4.
Corpus P and Corpus U papers with apparent nucleotide sequence typographic versus identity errors.
Fig 5.
The three Seek & Blastn automata (A1, A2, A3) and associated stacks (StatStk, NucStack, AllStack).
Circles represent states and arrows represent state transitions. The upper part of the label on an arrow specifies the property of the scanned word (W) that causes the transition. The lower part of the label described other actions triggered by the transition. The A1 automaton (shown at left) builds nucleotide sequence encounters in a sentence, extracting a nucleotide sequence (Seq) and building the sentence nucleotide stack (@NucStk). The A2 automaton (shown at centre) tracks the different targeting/ non-targeting status encounters when reading a sentence, which may be unknown (S = ʽU’), targeting (S = ʽT’) or non-targeting (S = ʽNT’). The automaton A3 (shown at right) tracks the use of the word “respectively” in the sentence. The ending state of this automaton is used to determine which stacks need to be used to decide the targeting/ non-targeting status of each nucleotide sequence.