Skip to main content
Advertisement

< Back to Article

Figure 1.

The DNA barcoding workflow.

Students participating in the Barcoding Life's Matrix program are engaged in a complete and integrated analytic workflow that includes the following core activities: collecting targeted marine specimens and processing/curating corresponding tissue; recording specimen data and collection event details using mobile computing technology; uploading specimen data to BOLD-SDP, an online workbench and data repository that was specifically designed for the educational community (see Figure 3 for additional details); generating CO1 amplicons using accessible and widely applied molecular-biology-based laboratory techniques (e.g., gDNA extraction, PCR amplification, agarose gel electrophoresis, silica spin-column purification of PCR products); assembling contigs and editing nucleotide sequence data using a suite of Internet-based bioinformatics tools and software; and uploading raw trace files and edited nucleotide sequence data to BOLD-SDP.

More »

Figure 1 Expand

Figure 2.

Acquisition and management of field data.

Specimen data and collection event details are acquired by students in the field using the DNA Barcoding Assistant, a free smartphone utility application. The date and time, elevation/depth, and GPS coordinates are automatically captured by the app when a new specimen record is created (1). A unique specimen identifier can be entered into the record manually using the smartphone keypad, or by scanning a barcode symbol affixed to a specimen storage container using the barcode reader function of the app (2). Currently supported barcodes include DataMatrix, EAN/UPC product codes, QR codes, Code 128, Code 39, Code 93, DataBar (RSS), and Interleaved 2 of 5. A digital photo of the specimen is captured using the smartphone camera (3) and linked to the specimen record. Additional information, including provisional genus and species names (4), collection site designation (5), and field notes (including the name of the collector) can also be added to the record manually by using the smartphone keypad (6). Completed records are transferred from the app to a computer, directly or via email (7). Once data are inspected for accuracy, they can be submitted to BOLD-SDP through the student data management console (8).

More »

Figure 2 Expand

Figure 3.

Barcode of Life Data Systems Student Data Portal (BOLD-SDP).

BOLD-SDP is a public-access data repository, data management tool, and analytical workbench for educational users that consists of five customized consoles. The Explore console of BOLD-SDP provides students and teachers with a gateway to the BOLD Taxonomy Browser (where they can determine the barcoding status of specific taxa) and the BOLD Public Data Portal (where they can access and download data generated by professional iBOL scientists and researchers). Through this console, students and teachers can also access the Integrated Taxonomic Information System and US National Center for Biotechnology Information Taxonomy Database (where they can obtain the scientific name and accepted taxonomy of a specimen from its common name; not shown in figure). The Teacher Admin console was designed for educators to register their class, compile a roster of student contributors, create a destination folder for specimen and sequence data, monitor student progress, and inspect student-generated data for accuracy. The Data Management console of BOLD-SDP permits student users to upload specimen data and collection event details, images of specimens and lab results, forward and reverse trace files generated from their amplicons, and edited nucleotide sequence data. The Sequence Analysis console of BOLD-SDP contains an integrated suite of analytical tools that enables students to (1) visualize the relatedness of specimens by building a genetic distance-based phenogram or tree (Taxon ID Tree), (2) compare sequence data obtained from their specimens against barcode records contained in the BOLD data repository (to confirm their specimen identifications; ID Engine), and (3) calculate the differences among nucleotide sequences generated from their specimens (from species to class levels; Distance Summary). The Submission console of BOLD-SDP allows project leaders and members of the scientific community to vet student-generated barcode records before moving them to the BOLD researcher workbench and submitting corresponding nucleotide sequences for publication in GenBank/INSDC (refer to Figure 4 for additional details).

More »

Figure 3 Expand

Figure 4.

Meeting the Barcode Data Standard.

Reference barcode sequences are linked to collateral data associated with the source specimen and collection event within biphasic records. Given their importance in ensuring the accuracy of species identifications through BOLD, reference barcode records are subject to a variety of formal data standards established by the scientific community. Through the Submission console of BOLD-SDP, project leaders and researchers review student-generated barcode records for their compliance with current data standards. Required data elements minimally include a species name assigned by an expert taxonomist (or a provisional name), a unique specimen identifier, information related to the voucher specimen (including the name of the institution storing the voucher), a collection record (e.g., collector, collection date, collection location, and geospatial coordinates), a CO1 sequence (for animals) of at least 500 nucleotides with fewer than 1% ambiguous base calls (Ns), the sequence of PCR primers used to generate the CO1 amplicon, and trace files. Student-generated records that satisfy these criteria are moved from BOLD-SDP to the BOLD researcher workbench and published in INSDC with the BARCODE designation.

More »

Figure 4 Expand

Figure 5.

Reference DNA barcode records generated by project participants.

(A) As of October 2012, students and teachers have generated and submitted complete reference DNA barcode records from 716 unique individuals representing four animal phyla, eight orders, 18 families, 26 genera (not shown), and 53 species (not shown). 716 records are currently published in GenBank (see Text S2 for the accession numbers corresponding to each published record). GenBank accession numbers, specimen and collection data, nucleotide sequences, trace files, and primer details are also available within the Barcoding Life's Matrix project folder, which is accessible through the BOLD Public Data Portal (http://www.boldsystems.org/index.php/Public_BINSearch?searchtype=records; use search term “BLM”). (B) Quality statistics for edited and unedited CO1 nucleotide sequence data. CO1 sequences edited by project participants from raw trace files contain no ambiguous base calls (Ns), stop codons, contaminating sequences, or insertions or deletions. Each amplicon was sequenced bidirectionally to yield at least one forward and one reverse trace file for each barcode record. Of the 1,444 trace files generated, 94.46% are categorized as high quality (mean Phred quality score >40 [28]), 4.43% are categorized as medium quality (mean Phred quality score = 30–40), and 1.11% are categorized as low quality (mean Phred quality score <30). (C) Nucleotide length distribution of CO1 sequences generated for each specimen. All 716 sequences generated by project participants exceeded the minimum barcode length of 500 nucleotides, with a minimum sequence length of 502 nucleotides, a maximum sequence length of 1,152 nucleotides, and a mean sequence length of 720 nucleotides.

More »

Figure 5 Expand