Skip to main content
Advertisement

< Back to Article

Designing diverse and high-performance proteins with a large language model in the loop

Fig 2

Fitness score statistics for BADASS optimization of alpha-amylase (AMY_BACSU).

BADASS was run for 140 iterations with a batch size of 1,000. The plot shows fitness score averaged over sampled sequences per iteration, with the shaded area representing scores within of the average, with the standard deviation of batch scores. Horizontal lines denote the set points and that govern the transitions between cooling and heating phases (Fig 1C). Vertical lines mark iterations where phase transitions occur as a moving average of fitness scores crosses the thresholds. The optimization was performed using the unsupervised ESM2 model and the semi-supervised Seq2Fitness model. Fitness scores were standardized as described in the methods. Runs included either exactly 6 mutations per variant (A, C) or an even mix of 2 to 6 mutations (B, D).

Fig 2

doi: https://doi.org/10.1371/journal.pcbi.1013119.g002