Data Collection Modeling General Biology Encoding Blog
The Vision for Prism
4 min read
Words by:

Stephanie WankowiczScientific Program Director
Predicting function and how it changes under perturbation is a central goal of biology.
Structural biology has pursued this goal by obtaining static structures. Those structures have provided key insights into enzyme mechanisms, protein folds, and evolutionary relationships. The goals of static structural biology organized seventy years of method development, instrumentation, and data infrastructure. Building off of this work, the first generation of biomolecular artificial intelligence (AI) solved the sequence-to-structure problem. Accurate static structure prediction is now available for nearly all sequences, and generative models have inverted the same mapping into a design direction. This generation demonstrated what becomes possible when biological data is combined with powerful AI tools.
However, these algorithms have limitations. While they can predict form, they often fail to predict function. For example, they often cannot explain mutation pathogenicity or drug selectivity. The design direction falls short in a similar way. While we can design binders, we cannot predict how a binder changes function beyond blocking a binding site, nor can we predict how selective it will be.
The limitation of these algorithms is representational. Function depends on the distribution of accessible conformational states. However, structural data that drove the first generation of AI tools collapses this population average to a single point. No amount of scaling over single points will recover a distribution.
Prism is transforming structural biology by shifting the baseline from static snapshots to dynamic ensembles. We are rebuilding the field's foundational layer around conformational ensembles, grounding molecular understanding in how macromolecules behave, and building toward predicting those ensembles. Function emerges from the ensemble of conformations a protein adopts and from the motions between them. How that motion drives biology remains uncharted. Making that link actionable will transform how we discover drugs, design enzymes and synthetic systems, interpret disease, and understand human health.
_________
To drive the next phase of structural biology, we can learn from and build on the field's strong foundation. First, the Protein Data Bank imposed a common format, a common deposition expectation, and a common access point on a heterogeneous experimental community. Second, methods in crystallography, nuclear magnetic resonance spectroscopy, and cryogenic electron microscopy were sustained and developed, each expanding the range of tractable systems while feeding standardized output back into the same archive. Third, computational algorithms were developed to support each new form of experimental data and were paired directly with the PDB. These advances were mutually dependent. The field succeeded because its stages were integrated, and that integration is the foundation Prism builds on.
That same integration is why we aim to approach this problem collaboratively. Partial change, or only attacking part of this problem – just modeling or just data processing will likely fall short or take too much time. New data is wasted if the processing removes motion. Better models are useless if the machine learning algorithms cannot represent that motion. Progress is bounded by the weakest stage, and information is lost at the points where stages meet. Transforming structural biology requires coordinated change at every stage of the pipeline and at every interface between them. Prism therefore treats data collection, modeling, representation, validation, and interpretation as a single program of work rather than as separable problems.
_________
This is an era of algorithmic abundance. Empowering those algorithms requires the right data, models, and encodings. Much of the technology and infrastructure needed to measure conformational ensembles does not yet exist. Prism is building it. We aim to democratize access to structural biology pipelines from data collection through modeling to infrastructure. Researchers will be able to generate the volume and variety of measurements required to connect ensembles to function.
Currently, Prism has four technological focus areas: data collection & processing, modeling, infrastructure, and biological impact. The four areas work together, and each advances the others. Together they build the foundation needed to measure, predict, and act on protein motion.
Stay tuned here to find out what’s next.
Share article: