Figure 1.
Embedding a watermark into a functional gene.
A) The 5′ end of HIV gag ORF optimized for Homo sapiens, HIVgag (optimized) with its amino acid and nucleotide sequence, and original codon ranking. The modified codon ranking and altered nucleotide sequence of HIVgag (message) shown below give the desired binary message in blue, spelling “GENE”. B) The ASCII symbols spelling “GENE” convert to the binary digits 100111, 100101, 101110, 100101, based on C) the modified ASCII table. A total of 64 typographic characters (Char) were chosen from the print characters 32 to 95 of the standard ASCII decimal code (ASCII Dec). Subtracting 32 from each value gave numbers ranging from 0 to 63 (Minus 32), which were converted into a 6-bit binary code (Binary). D) The sorted human codon usage table was used to incorporate this bitstream into the modified HIVgag (message) sequence depicted above. Only amino acids with ≥4 alternative codons were changed (red letters in HIVgag sequence at the top). Binary 1 represents codons ranking 1, 3 or 5 (odd); binary 0 is for codons ranking 2, 4 or 6 (even). To secure binary 0 at nucleotide position 43 the leucine codon ranking 4 was chosen since the 2nd best codon would have created an undesirable SacI restriction site (GAGCTC). Embedding the four-letter text message required 12 silent substitutions (shaded grey in A) in the watermarked DNA sequence.
Figure 2.
Protein expression and functionality.
A) Western blots of HIV gag protein without [opt] and with [msg] message expressed in HEK293 cells. Equal amounts of protein from 5 independent transfections were analyzed from cell supernatants (top panels) and cell lysates (bottom panels), with actin detection serving as a control. B) Western blot of HEK293 lysates expressing GFP from optimized genes without [opt] and with [msg] embedded watermarks, including the appended human codon ranking sequence [msg+cut] (see Fig. 3), an encrypted watermark [msg enc] (see Table 2) and with a longer embedded message [msg long], also employing amino acids with 2 or 3 alternative codons (CDEFHIKNQY). Protein expression was quantified by densitometry. Results are derived from five independent experiments. C) Fluorescence microscopy images of GFP transfected into tobacco leaves show no visible differences in cellular location and only little variation in abundance. D) GST-T7 RNA polymerase [opt] and [msg] expressed in triplicates in E. coli was analyzed by SDS-PAGE and Coomassie staining (top panel) or Western blotting using a specific T7 RNA polymerase antibody (α-T7RNAP; lower panel). Equal amounts of purified T7 RNA polymerase were used for in vitro transcription. Synthesized RNA detected with a molecular beacon in real-time directly revealed almost identical RNA polymerase activities mediated by T7 [opt] (blue line) and T7 [msg] (orange line).
Table 1.
Summary of the genes tested. Wildtype gene sequences optimized for expression in different organisms1 were used for embedding various messages.
Table 2.
Vigenère polyalphabetic substitution.
Figure 3.
Procedure for storing the key to any possible codon usage table ranking order within a string of 35 nt (▴) and vice versa (Δ). In this example the sequential arrangement of the human codon usage is specified in the table columns “AA” and “Order”. For example, the most frequently used alanine codon in the human genome is GCC, followed by GCT, GCA and the least frequent GCG. Applying this sequential frequency arrangement to the alphabetically ordered codons results in GCA(3), GCC(1), GCG(4), GCT(2) or - in short - 3142. In total, there are 24 possible combinations of 4 alternative codons (#0 = 1234 … #23 = 4321) and the showcase order 3142 is at position #13 in this list. Thus, from the lookup table for four alternative codons (4 Alt. Codons), this order can be represented by the binary identifier 01101 ( = 13). When performed for each amino acid, a binary string of 70 digits defines the frequency arrangement of all 64 codons for that specific codon usage table, here listed for H. sapiens. A simple binary-to-nucleotide translation table (A = 00, C = 01, G = 10, T = 11) can then represent this binary string in a sequence of 35 nucleotides.
Figure 4.
Influence of gene labeling in HIV-1 gag on viral infectivity and replication.
A) Viral infectivity was assessed after transiently infecting HEK293T cells with the indicated viral plasmids, harvesting cells 72 h post transfection and using equal amounts of capsid-normalized virus particles to infect CD4-positive TZM-bl indicator cells. Luciferase activity (RLU) was measured in cells lysed 48 h post infection. The control represents uninfected cells. Error bars indicate the standard deviations from quadruplet infection. B) To monitor and maintain viral replication over a period of 20 weeks, CEM cells were infected in duplicate with virus-containing supernatants. Capsid protein (p24) amounts in culture supernatants were quantified by ELISA at the indicated time points. The control represents supernatant from an uninfected cell culture. Error bars indicate standard deviations from duplicate infections.