Biology Blog
How Silent Code Updates Can Alter Protein Predictions
6 min read
Words by:

Feng YuScientist
Over the past few years, we have witnessed an explosion of protein prediction models such as AlphaFold3, OpenFold3, Protenix, Boltz, and Chai. These structure predictions serve not only as standalone tools but also as critical inputs to downstream algorithms, from generative protein design engines to molecular docking pipelines.
However, these foundational models are revised and improved over time to improve performance or fix bugs. While open-source tools are tracked on GitHub or other platforms, the rapid development cycle alongside the dependencies of other downstream uses of these tools raises a critical question for the scientific community: do these routine version updates silently introduce significant changes to our predictions?
Mismatched Structure Predictions
Sampleworks’ core mission is to use and improve generative biomolecular structure predictors to model experimental data describing protein conformational ensembles. Sampleworks is a highly modular framework (Chrispens et al. 2026) that lets us choose the best predictors and methods for modeling conformational ensembles. To track progress towards that goal, we rigorously track the versions of code and model weights we use, operating under the assumption that locked weights yield locked results.
Recently, we used sampleworks to examine results published by the Bronstein team in their paper, Inference-time optimization for experiment-grounded protein ensemble generation. We evaluated predictions for PDB entry 5G51 chain A, a single-chain virus VP3 protein that forms complexes. The paper used AlphaFold3 (AF3)(Abramson et al. 2024) as the base model, whereas ours utilized Protenix (Protenix team et al. 2026). While we anticipated minor variations due to algorithmic differences between the base models, the results revealed a stark divergence.
The sampleworks Protenix baseline prediction (without any experimental guidance) yielded an essentially unstructured, "spaghetti"-like conformation (Figure 1). Therefore, we ran the exact same protein sequence through the official online Protenix server. We configured the server to use the original v0.5.0 weights (the ones used in sampleworks) alongside its v2.0 model code. This produced a well-constructed model of 5G51, raising the question of why we saw such divergent structure predictions between the server and sampleworks.

Figure 1: 5G51 protein structure predictions. Our local Protenix version (yellow) poorly matches the real structure (dark blue). However, the updated Protenix server (orange) matches both the real structure and AlphaFold 3 (pink) closely, showing the huge difference caused by the software update.
We then tested a standalone Protenix 0.7.3 installation with weight version 0.5.0 to verify our results, which were nearly identical to the spaghetti model we got from running the structure prediction sampleworks, confirming that the sampleworks implementation and configuration were not causing the divergent structure predictions, but rather the problem lay in a hidden discrepancy between the version we were running and the version deployed on the server.
After reviewing the Protenix repository release notes, we found that over a several month development window, the maintainers pushed two major releases. Crucially, during one of these updates,reworked its multi-sequence alignment (MSA) handling code, removing components previously derived from OpenFold (Ahdritz, G et al. 2024) with a custom-built alternative. While our thorough review confirmed this architectural shift was documented in the repository's GitHub release notes and in the paper's supplemental material, it was omitted from the main text of their preprint (Proteinix team et al. 2026).
Because the online server runs Protenix 2.0, its updated pipeline produced different MSA results for a given sequence than older versions. We had executed our local pipeline using an older codebase (Protenix 0.7.3) while rigidly pinning the model weights (version 0.5.0) to enforce consistency. Because the preprint did not explicitly mention the module replacement in the main text, we assumed that pinned weights guarantee reproducibility. We therefore hypothesized that this underlying difference in MSA generation caused the mismatched results between our local pipeline and the Protenix server, despite both using identical model weights.
To test this hypothesis, we first predicted the same protein chain with Boltz2, which obtains its alignments from the ColabFold MMseqs2 service (Mirdita, M. 2022 et al.). Compared with Protenix v0.7.3, which derives its alignment features from ColabFold MMseqs2 output using OpenFold-derived parsing and pairing code, Boltz obtains paired and unpaired alignments directly from the ColabFold MMseqs2 service and ingests them through its own pipeline. Boltz2 produced a more accurate prediction to the AlphaFold3 baseline, indicating the chain is tractable given a ColabFold-derived alignment. Supplying that alignment to our local Protenix pipeline — with the codebase (v0.7.3) and weights (v0.5.0) unchanged — recovered AF3-level accuracy (Fig. 2). Because only the input alignment varied, this indicates that the alignment supplied to the model, rather than the pinned weights, accounts for the mismatch between our local predictions and the server.

Figure 2: CA RMSD of 5G51 predictions across model configurations. Swapping the default Protenix MSA module for ColabFold MSA reduces median RMSD from 12.14 Å to 7.74 Å, rescuing accuracy to match the AlphaFold3 baseline.
Uncharted Upgrades
One might logically ask, "If the newer version yields a more accurate prediction, why is this an issue? Isn't improvement always the goal?" The danger lies not in the improvement itself, but in the integrity of comparative research.
Consider a hypothetical scenario: a research team submits a paper claiming their novel computational method—built on top of Protenix—significantly improves the Root Mean Square Deviation (RMSD) compared to historical benchmarks. They might document that they used the exact same model weights as previous studies. However, if their environment automatically pulled the updated underlying codebase through a routine package update, the final results would show a massive performance leap, that should be credited to Proteinx updates, not the model built on top of it.
Researchers routinely compare their novel methods against established models. To do so, they rely on a dated, fixed
toml file for established models, creating a sense of stability. However, a dated toml file will always point to a specific version in history, which may keep the weights the same but have hidden code improvements. On the other hand, new models usually use the latest or newer version of similar software. When underlying data processing pipelines or alignment modules evolve without explicit, highly visible acknowledgment of their impact on final outputs, the bedrock of scientific reproducibility fractures.This oversight points to a much broader vulnerability in how we conceptualize model architectures versus their surrounding frameworks. Historically, we have operated under the flawed assumption that feature-processing pipelines—the scaffolding around the model—remain static or behave predictably across systems. This same dynamic is currently playing out across the broader artificial intelligence landscape, particularly in agentic AI. In that domain, many celebrated leaps in performance are not driven by fundamental improvements to the core Large Language Models (LLMs) themselves, but rather by highly optimized "harnesses"—the sophisticated wrappers, tool integrations, and processing frameworks enveloping the models. Whether in protein structure prediction or language generation, when we fail to distinguish between the raw capability of a core architecture and the heavy lifting done by its surrounding framework, we fundamentally misattribute the source of innovation.
As the field continues to lean heavily on these powerful predictive tools, we must elevate our standards for computational vigilance. True open science demands more than just sharing model weights; it requires absolute transparency and rigid control over the precise configuration of the entire predictive pipeline. Otherwise, we risk measuring the ghosts of software updates rather than genuine scientific progress.
Share article: