Workflow
Conditioned generation and constraints
Structure prediction and scoring
Experimental assays and iteration
Example
Example: De Novo Enzyme Design
Specify catalytic constraints.
Generate and rank sequences.
Assay and optimize variants.
Frequently asked questions
Generalization?
Generalization in AI-driven protein design relies heavily on the quality and diversity of the training data, as well as the inductive bias embedded within the model’s architecture. Models trained on a broad range of protein families are more likely to generalize effectively across different sequences and functional requirements.
Stability?
Stability is assessed through energy models, such as Rosetta or CHARMM, which predict the thermodynamic stability of the designed proteins. These predictions are then validated experimentally using techniques like differential scanning calorimetry (DSC) to confirm the protein’s resistance to unfolding under various conditions.
Function?
The AI is conditioned on active sites and motifs known to be associated with specific functions, guiding sequence selection towards regions likely to exhibit the desired catalytic or binding activity. This approach leverages existing knowledge of protein function while allowing for novel combinations within a designed sequence.
Screening?
High-throughput readouts are crucial for rapidly screening large numbers of generated sequences, often utilizing techniques like fluorescence polarization or surface plasmon resonance. These assays provide quantitative data on protein activity, allowing the AI to quickly identify and prioritize the most promising candidates for further investigation.
Safety?
Risk assessment and filters are incorporated throughout the design process to mitigate potential safety concerns associated with novel proteins. These measures include evaluating predicted toxicity, immunogenicity, and potential off-target effects before experimental validation.
IP?
Intellectual property considerations center on determining whether a newly designed protein sequence is sufficiently novel compared to naturally occurring sequences. Patentability hinges on demonstrating that the design represents a genuinely inventive step, rather than simply mimicking existing proteins.
Computing?
Significant GPU budgets and optimized pipelines are essential for efficiently processing the computational demands of AI-driven protein design, particularly when generating and scoring vast numbers of sequences. Parallelization strategies are often employed to accelerate these calculations.
Data?
Curated assays and labels – including detailed structural information, activity data, and stability measurements – provide the critical feedback loop necessary for training and refining AI models in protein design. High-quality data ensures that the model learns effectively from experimental results.
Clinical?
Developability criteria are considered early on, focusing on factors such as solubility, aggregation propensity, and potential for large-scale production – all crucial aspects for translating protein designs into therapeutic applications or industrial processes.
Outlook?
The future of protein design with AI lies in the development of autonomous design-build-test loops, where AI systems iteratively refine protein sequences based on continuous experimental feedback, accelerating the discovery and optimization of novel proteins for a wide range of applications.
Try it live
Everything above runs in your browser — open AI Protein Design: Diffusion Folding Simulator and change the parameters while it is running. Nothing is installed, nothing is uploaded, the whole model lives in one tab.
▶ Open AI Protein Design: Diffusion Folding Simulator simulation