{"publication":{"abstract":"Proteins are dynamic ensembles of interconverting conformations. This ensemble, rather than any single structure, governs functions such as catalysis, mutational effects, and molecular recognition. Predicting protein conformational ensembles is a central goal of structural biology, but progress is constrained by a shortage of training data. Current ensemble predictors rely on molecular dynamics simulations, which are limited in number and constrained by force-field accuracy and accessible timescales. Experimental structural data is an untapped source of alternative data for training ensemble predictors of structure. X-ray crystallography and cryo-EM measure many copies of a protein and average over space and time, so the data encodes an ensemble of states. Conventional refinement collapses this signal into a single set of coordinates. Most Protein Data Bank (PDB) depositions, therefore, report a single averaged structure, leaving the underlying heterogeneity unmodeled and hidden in the experimental data. Here, we applied qFit to recover this latent heterogeneous signal at scale. Starting from structures resolved to better than 2 Å with deposited structure factors, we re-refined them and ran qFit multiconformer modeling. This produced over 60,000 completed multiconformer models, the largest dataset of experimentally derived ensemble protein structural models to date, spanning a broad range of sequence and structural diversity. 83.7% of structures have a lower $R_\\mathrm{free}$ with qFit multiconformer models compared to the deposited, re-refined model. These models identified widespread side-chain heterogeneity absent from the deposited models. This resource aims to help address the data bottleneck in ensemble prediction and reframes the experimental data deposited in the PDB as a source of ensemble information.","body":"# Background\n\nStatic structural prediction has been transformed by machine-learning predictors [](https://doi.org/10.1038/s41586-021-03819-2) [](https://doi.org/10.1038/s41586-024-07487-w) [](https://doi.org/10.1101/2025.06.14.659707) [](https://doi.org/10.64898/2026.02.05.703733) [](https://doi.org/10.1126/science.abj8754) [](https://doi.org/10.1101/2024.10.10.615955) [](https://doi.org/10.1101/2025.01.08.631967). However, proteins are not static. They exist as dynamic ensembles of interconverting conformations [](https://doi.org/10.1021/cr040423+). This ensemble, not any single structure, drives biological function, governing catalysis, mutational effects, and molecular recognition [](https://doi.org/10.1146/annurev-biophys-061824-104900) [](https://doi.org/10.1038/nchembio.232).\n\nPredicting an ensemble is significantly harder than predicting a single structure, as the target is a distribution of weighted states rather than a single set of coordinates [](https://doi.org/10.1038/s41592-026-03084-z). Beyond the significantly more difficult learning problem, the training data needed to learn that distribution is scarce. Current ensemble structural predictors are trained on and benchmarked against molecular dynamics (MD) data [](https://doi.org/10.1126/science.adv9817) [](https://doi.org/10.1016/j.sbi.2026.103251) [](https://doi.org/10.48550/ARXIV.2402.04845) [](https://doi.org/10.1101/2025.06.14.659707). However, simulations are constrained by force field accuracy, accessible timescales, and the diversity of the training set [](https://doi.org/10.1016/j.neuron.2018.08.011) [](https://doi.org/10.1073/pnas.1800690115) [](https://doi.org/10.1002/prot.26409). Accurately predicting conformational ensembles will require far more training data.\n\nOne untapped resource for training data is the experimental data used to derive static structures. X-ray crystallography and cryo-electron microscopy (cryo-EM) measure tens of thousands to billions of copies of the protein. These measurements are then ensemble-averaged over both space and time, so the experimental data carries information about an ensemble of states. While crystal packing and cryogenic temperatures restrict the accessible conformational space, they do not eliminate heterogeneity [](https://doi.org/10.1126/science.1218231). Conventional modeling and refinement nonetheless collapse this signal into a single, approximate structure. As a result, most Protein Data Bank (PDB) depositions report a single, averaged set of coordinates [](https://doi.org/10.1016/j.str.2021.04.010) [](https://doi.org/10.1038/nsmb0306-184) [](https://doi.org/10.1038/s41592-022-01760-4). These are the models that helped unlock static structure prediction. The unmodeled heterogeneity in the experimental data is itself a resource for ensembles. Recovering it would substantially increase the available ensemble training data.\n\nHowever, identifying and modeling the underlying conformational heterogeneity in experimental data is difficult. It is typically done by hand, identifying often sub-Angstrom differences, using visualization tools such as Coot [](https://doi.org/10.1107/s0907444904019158). This is complicated by low density and noise from many sources, including crystal imperfections, radiation damage, and poor initial modeling [](https://doi.org/10.1107/s1399004715006045) [](https://doi.org/10.1107/s0907444909047337) [](https://doi.org/10.1126/science.1218231). Manual modeling would be incredibly challenging to scale across the PDB. To make multiconformer modeling routine and impartial, we previously developed qFit [](https://doi.org/10.1371/journal.pcbi.1004507) [](https://doi.org/10.1002/pro.4001) [](https://doi.org/10.1107/s0907444909030613). qFit takes a refined single-conformer structure and a high-resolution X-ray or cryo-EM real-space map as input. It then uses optimization algorithms to identify alternative protein [](https://doi.org/10.1107/s0907444909030613) [](https://doi.org/10.1371/journal.pcbi.1004507) or ligand [](https://doi.org/10.1002/pro.4001) conformations, supplementing manual modeling.\n\nqFit produces a multiconformer model [](https://doi.org/10.7554/elife.90606). A multiconformer model represents conformational heterogeneity within a single set of coordinates. It encodes alternative states locally as altloc-labeled conformations and assigns each a weight in the population proportional to its occupancy. Altlocs are added only where the density requires it, so an ordered region can be modeled by a single conformation, whereas a more heterogeneous area can be modeled by multiple conformations. In contrast, an ensemble model represents heterogeneity as multiple complete copies of the system. A region that adopts several conformations appears as several slightly different versions of itself, one per copy, and the population is read from the spread of the copies rather than from any single one. The number of copies is set by the refinement protocol rather than by the local complexity of the density, so even a well-ordered region is duplicated across every model [](https://doi.org/10.1016/j.sbi.2014.07.005).\n\nHere, we ran qFit on almost 80,000 deposited PDB structures with resolution better than 2 Å and deposited structure factors. In most structures, qFit improves the fit to experimental data and identifies additional conformational heterogeneity latent in the deposited data. This produced the largest dataset of multiconformer protein structures assembled to date, helping address the data bottleneck that constrains ensemble prediction. Beyond ensemble prediction, it also provides a rich resource for systematic studies of how conformational heterogeneity affects allostery, macromolecular interactions, and mutations [](https://doi.org/10.1101/2025.11.25.690589) [](https://doi.org/10.7554/elife.111298.1) [](https://doi.org/10.1107/s2059798318017941).\n\n# Data Generation\n\nThe initial dataset includes all structures from PDBRedo with a resolution of 2 Å or better (downloaded December 2024; n=80,876). Structures composed solely of nucleic acids were removed (n=1043). We then re-refined all models using _phenix.refine_ with the parameters shown below [](https://doi.org/10.1107/s0907444912001308). This re-refined structure was used as the ‘deposited’ comparison. We then generated a composite omit map, which was used to run qFit (version 2025.3) with default parameters [](https://doi.org/10.7554/elife.90606). Finally, we ran the post-qFit refinement script using _phenix.refine._ The final refinement cycle is outlined below and matches the re-refined model. All code and parameters needed to reproduce this pipeline are available in the qFit repository: <https://github.com/ExcitedStates/qfit-3.0>. All re-refined models (PDB and CIF), qFit models (PDB and CIF), and FASTA files are included in the Zenodo deposition ([https://zenodo.org/records/20801853](https://zenodo.org/records/20801853)).\n\n```bash\nphenix.refine \\\n\"${pdb}.pdb\" \\\n\"${pdb}.mtz\" \\\n\"$ligand_cif\" \\\nrefinement.main.number_of_macro_cycles=5 \\\nrefinement.main.nqh_flips=True \\\nrefinement.refine.adp.individual.isotropic=all \\\nrefinement.output.write_maps=False \\\nrefinement.hydrogens.refine=riding \\\nrefinement.main.ordered_solvent=True \\\nrefinement.target_weights.optimize_xyz_weight=true \\\nrefinement.target_weights.optimize_adp_weight=true \\\nordered_solvent.mode=every_macro_cycle\n```\n\n# Dataset Description\n\nOf the 79,833 structures that entered the pipeline, 60,549 (85.9%) produced completed models. Among the completed structures, the median resolution was 1.70 Å (range: 0.5–2.0 Å). The median number of residues was 298 (range: 3–2,338)([Figure 1](#figure1)). The distribution of chain count per structure (range:1-36). The re-refined models had a median $R_\\mathrm{free}$ of 0.198 (range 0.06–0.71) and a median $R_\\mathrm{work}$ of 0.17 (range 0.06–0.68). $R_\\mathrm{free}$ rises modestly as resolution worsens (slope 0.062, R² = 0.22)([Figure 2](#figure2)). The median $R_\\mathrm{free}$ of about 0.2 corresponds to a residual disagreement of roughly 20% between the experimental data and the model. This residual partially reflects the limitation of reducing ensemble-averaged data to a discrete model rather than experimental error [](https://doi.org/10.1016/j.cell.2024.01.003) [](https://doi.org/10.1111/febs.12922). All models are located in the Zenodo deposition (https://zenodo.org/records/20801853).\n\n::::::figure{#figure1 align=\"center\" type=\"image\" label=\"Figure 1\"}\n\n:::::image{width=\"100%\" src=\"https://thestacks-01.s3.amazonaws.com/publications/qfit-at-scale/media_48b74625_f3b4690c8a1c\"}\n:::::\n\n:::::figcaption\n**Figure 1.** Overview of the qFit multiconformer model dataset (n = 60,549). A. Distribution of resolutions; median resolution 1.70 Å (range 0.50–2.00 Å). B. Distribution of total residues per structure; median 298 residues (range 3–2,338). C. Distribution of chain count per structure (range:1-36).\n:::::\n\n::::::\n\n::::::figure{#figure2 align=\"center\" type=\"image\" label=\"Figure 2\"}\n\n:::::image{width=\"100%\" src=\"https://thestacks-01.s3.amazonaws.com/publications/qfit-at-scale/media_f494fc13_f3b4690c8a1c\"}\n:::::\n\n:::::figcaption\n**Figure 2.** Data Description of qFit multiconformer models (n=60,549). A. Distribution of $R_\\mathrm{work}$. B. Distribution of $R_\\mathrm{free}$. C. Relationship between resolution and $R_\\mathrm{free}$ (Slope: 0.062, R²: 0.22).\n:::::\n\n::::::\n\nTo assess the dataset's diversity, we clustered the entries at the sequence and structural levels. Sequence clustering with MMseqs2 [](https://doi.org/10.1038/nbt.3988) yielded 16,133 clusters at 30% sequence identity, with many being singletons, and 25,608 clusters at 90% sequence identity. Structural clustering with Foldseek produced 6,590 clusters [](https://doi.org/10.1038/s41587-023-01773-0).\n\n# Multiconformer models improve fit to experimental data\n\nTo assess whether multiconformer modeling improves the fit to the data, we compared the $R_\\mathrm{free}$ of each qFit multiconformer model with that of the corresponding re-refined PDB-Redo model. Across 60,549 models, qFit lowered $R_\\mathrm{free}$ by a mean of −0.009 and a median of −0.01 ([Figure 3](#figure3)A/B). Individual outcomes vary widely, ranging from a 0.320 decrease to a 0.310 increase, with a standard deviation of 0.026. However, 54,274 (89.7%) structures have a lower $R_\\mathrm{free}$ with qFit compared to the deposited, re-refined model. While the average improvement in $R_\\mathrm{free}$ is small in magnitude, it represents a real improvement in fit to the data through identifying and modeling conformational heterogeneity (see below). Further, the improvement in $R_\\mathrm{free}$ is partly obscured by poor solvent fitting around heterogeneous models [](https://doi.org/10.82153/2ME0-HD96). Of note, there is a subset of structures (1.3%; n=802) in which qFit yields significantly worse $R_\\mathrm{free}$, defined as an increase in $R_\\mathrm{free}$ by over 0.05. Examining these structures, we could not find a pattern that explained this increase in $R_\\mathrm{free}$; this remains the subject of ongoing work.\n\n::::::figure{#figure3 align=\"center\" type=\"image\" label=\"Figure 3\"}\n\n:::::image{width=\"100%\" src=\"https://thestacks-01.s3.amazonaws.com/publications/qfit-figure-update/media_bfea6b75_0e214fd16a5b\"}\n:::::\n\n:::::figcaption\n**Figure 3.** Comparison of $R_\\mathrm{free}$ between deposited and qFit models. A lower $R_\\mathrm{free}$ indicates a better fit to the experimental data. A. Scatterplot of deposited $R_\\mathrm{free}$ versus qFit $R_\\mathrm{free}$. Points below the diagonal indicate the qFit model fits the data better. B. Histogram of the per-structure $R_\\mathrm{free}$ difference (deposited minus qFit).\n:::::\n\n::::::\n\n# Multiconformer models have significantly more conformational heterogeneity compared to static structures\n\nWe quantified the additional heterogeneity in qFit models relative to deposited models using root-mean-square fluctuation (RMSF) and the number of alternate conformations (altlocs). RMSF measures the spatial spread of atomic positions across the modeled conformers and captures the magnitude of discrete displacement that qFit introduces through multiconformer modeling. The number of altlocs is the count of discrete alternate conformations assigned per residue and provides a direct measure of how many distinct conformational states qFit resolves from the density.\n\nThe majority of residues had one altloc (90.4%) across deposited and qFit models ([Figure 4](#figure4)A). qFit increased the altloc count for 9.4% of residues. Among residues that gained conformations, a single additional altloc was most common at 4.9%, followed by two at 3.6% and three at 0.8%. Alanine, glycine, and proline had the fewest additional altlocs, while leucine and cysteine had the most ([Supplementary Figure 1](#sfig1)). Only 0.2% of residues also had a reduction in the number of altlocs modeled.\n\nAcross the dataset, qFit increased RMSF by a mean of 0.13 Å relative to the deposited model, with a median difference of 0, consistent with the majority of residues not having an alternative conformer modeled ([Figure 4](#figure4)B). However, a small subset of residues shows a large increase in RMSF. Unsurprisingly, the longer the amino acid, the larger the qFit RMSF, with lysine, arginine, glutamate, and glutamine showing the largest changes in RMSF ([Supplementary Figure 2](#sfig2)). As with the removal of altlocs, a minority of residues show reduced RMSF.\n\n::::::figure{#figure4 align=\"center\" type=\"image\" label=\"Figure 4\"}\n\n:::::image{width=\"100%\" src=\"https://thestacks-01.s3.amazonaws.com/publications/qfit-figure-update/media_ee3beb24_7a6caafb7032\"}\n:::::\n\n:::::figcaption\n**Figure 4.** Per-residue changes in altloc count and RMSF between qFit and deposited models. A. Distribution of the per-residue difference in altloc count across all qFit models relative to deposited models. The y-axis shows the number of residues on a log scale. 9.4% of residues gained an altloc in the qFit model relative to the deposited model. B. Distribution of the per-residue difference in RMSF across all qFit models relative to deposited models. The y-axis shows the number of residues on a log scale. qFit increased RMSF by a mean of 0.13 Å relative to the deposited model. The median difference was 0 Å.\n:::::\n\n::::::\n\n# Hidden heterogeneity in a high-quality structure\n\nStructures with excellent refinement statistics can still contain conformational heterogeneity that the deposited model does not capture. For example, the crystal structure of death-associated protein kinase 1 (DAPK1) in complex with resveratrol (PDB 7CCU; [Figure 5](#figure5)A). This structure has a resolution of 1.65 Å, an $R_\\mathrm{free}$ of 0.189, and an $R_\\mathrm{work}$ of 0.172, about median values for our dataset. By these metrics, the model is of relatively high quality. However, the deposited model has no alternative conformations modeled. After applying the qFit pipeline, we lowered the $R_\\mathrm{free}$ to 0.1812 and the $R_\\mathrm{work}$ to 0.159, indicating a small but meaningful improvement in the fit to the scattering factors. Across the structure, we found that about 60% (164/277) of residues had at least one altloc, and 31% (85/277) had more than 2 altlocs. When removing altlocs that exist within the same rotamer well, we still see that 52% (144/277) and 17.3% (48/277) of residues had one or multiple altlocs. Figure 5B-D shows a visualization of altlocs identified in Phe240 ([Figure 5](#figure5)B), Asp220 ([Figure 5](#figure5)C), and Glu118 ([Figure 5](#figure5)D). All of these had a single conformation modeled in the deposited structure, but clearly support multiple conformations.\n\n::::::figure{#figure5 align=\"center\" type=\"image\" label=\"Figure 5\"}\n\n:::::image{width=\"100%\" src=\"https://thestacks-01.s3.amazonaws.com/publications/qfit-at-scale/media_32e538a2_745556479dbf\"}\n:::::\n\n:::::figcaption\n**Figure 5.** DAPK1 in complex with resveratrol (PDB 7CCU). The deposited model is shown in green and the qFit multiconformer model in blue. Electron density is contoured at 0.5σ (mesh). (A) Deposited model in electron density. (B) Phe240. (C) Asp220. (D) Glu118.\n:::::\n\n::::::\n\n# Related datasets of conformational heterogeneity\n\nSeveral resources catalog the deposited alternative conformations. PDBFlex characterizes flexibility by comparing different PDB structures of the same protein. The Alternate Location Server maintains a list of PDB structures containing alternate locations, and a recent survey has assembled a custom dataset of alternately modeled backbone segments [](https://doi.org/10.1093/nar/gkv1316). However, most deposited structures have no alternative conformations, and the difficulty of modeling them often leads to human bias, with some parts of the protein more likely to be modeled as multiconformers while other regions remain single-conformer models [](https://doi.org/10.1101/2024.12.14.628518).\n\nBeyond experimentally derived datasets, most datasets used for training ensemble predictors are MD datasets. The most widely used resources are ATLAS and mdCATH, which provide all-atom trajectories for a broad range of folded domains, spanning 1,390 protein chains in ATLAS and 5,398 CATH domains in mdCATH [](https://doi.org/10.1093/nar/gkad1084) [](https://doi.org/10.1038/s41597-024-04140-z). Other resources include MISATO, which simulates approximately 20,000 protein-ligand complexes drawn from PDBbind and focuses on common drug targets [](https://doi.org/10.1038/s43588-024-00627-2). GPCRmd provides MD on GPCRs, with most systems belonging to class A [](https://doi.org/10.1038/s41592-020-0884-y). IDRome covers many disordered regions of the human proteome using coarse-grained simulations [](https://doi.org/10.1002/pro.5172).\n\nThere are a few key differences between the derivatives of these MD datasets and our qFit multiconformer model dataset. MD produces an explicit, time-ordered trajectory. It resolves motions from femtosecond-scale bond vibrations upward but is limited by the simulation length and force-field accuracy. For example, the ATLAS database, which has been used by many ensemble prediction algorithms, provides three 100 ns trajectories per protein [](https://doi.org/10.1093/nar/gkad1084). Many functionally important motions, including loop rearrangements and allosteric transitions, occur on microsecond to millisecond timescales and are undersampled at these lengths [](https://doi.org/10.1038/nature06522). In contrast, X-ray crystallographic measurements are inherently time-agnostic, capturing conformational changes regardless of the timescale, provided they occur within the crystal. A single X-ray experimental dataset reports the equilibrium distribution of conformations populated across the crystal, averaged over time and over all unit cells. However, there is no time information, and conformations are biased by the crystalline environment. Further, this data remains latent in the experimental data until modeled.\n\n# Impact\n\nX-ray crystallography and cryo-EM data encode rich ensemble information that conventional modeling pipelines largely discard. Beyond the impact of understanding individual proteins' functions, this leaves valuable training data for ensemble methods unused. Here, we applied qFit across high-resolution structures and experimental data in the PDB to recover this heterogeneity. We show that qFit recovers the widespread conformational heterogeneity present in high-resolution X-ray structure factors but absent from the deposited models [](https://doi.org/10.7554/elife.90606) [](https://doi.org/10.1002/pro.4001). The resulting multiconformer models more faithfully represent the true underlying ensemble, as shown by improved $R_\\mathrm{free}$ values. This yields the largest dataset of multiconformer protein models to date. It offers experimentally derived protein ensemble information at scale, a significant addition to the largely MD-derived datasets used for ensemble training [](https://doi.org/10.1093/nar/gkad1084) [](https://doi.org/10.1038/s41597-024-04140-z) [](https://doi.org/10.1038/s43588-024-00627-2) [](https://doi.org/10.1038/s41592-020-0884-y). This foundational resource opens new avenues for studying protein ensembles and advancing computational methods for their prediction [](https://doi.org/10.1038/s41592-026-03084-z).\n\nBeyond ensemble training,  it is possible that this resource may help the accuracy gap between prediction methods and experiments may result from an incomplete consideration of ensembles [](https://doi.org/10.1038/s41592-022-01760-4). All public predictors are trained on single-conformer models [](https://doi.org/10.1038/s41586-024-07487-w) [](https://doi.org/10.1101/2025.06.14.659707) [](https://doi.org/10.1101/2024.10.10.615955) [](https://doi.org/10.1101/2025.01.08.631967). When a residue populates multiple states, modeling it as a single conformer introduces strain and other local geometric artifacts [](https://doi.org/10.7554/elife.90606). Predictors trained on these structures may learn the distortions along with the correct geometry. Beyond incorrect geometry, it is possible that ensemble models will help with “the last angstrom problem”, the gap between the accurate backbones that current predictors produce and the sub-angstrom side-chain and bond geometry needed for ligand docking and mechanistic interpretation [](https://doi.org/10.1101/2025.06.30.662466). The last angstrom problem may be due to the precision with which the algorithms are built and to underlying uncertainty, but it is also possible that the poor precision reflects the statistical distribution of structural models. Additionally, this data can also help with the analysis and benchmarking of MD data. \n\nMore broadly, the dataset supports a shift toward dynamic structural biology. It reframes the PDB as a source of ensemble information. In the vast majority of cases, experimental structural biology data are too rich to be described by a single conformation. Making heterogeneous multiconformer models available at scale lowers the barrier to integrating dynamics into routine structural analysis. This opens up rich possibilities for systematic study across a range of biological phenomena through the lens of conformational ensembles.\n\nSeveral dataset limitations warrant consideration. First, structures are restricted to those resolved to better than 2 Å with available structure factors, limiting our dataset and hiding heterogeneity in lower-resolution structures. Second, the dataset inherits conformational constraints imposed by crystal packing and cryogenic temperatures in most structures. Third, the conformational heterogeneity identified here arises predominantly from side-chain rather than backbone movements. Beyond these specific constraints, it is important to recognize that structures are models rather than experimental reality [](https://doi.org/10.1016/j.cell.2024.01.003). Our multiconformer models improve fit to the experimental data, as shown by improved $R_\\mathrm{free}$ values, but bridging the remaining gap between deposited models and the full information content of experimental observations remains an important challenge for the field.\n\nFuture work will focus on improving these algorithmic approaches and integrating them with guidance methods [](https://doi.org/10.82153/JKXJ-TW08) [](https://doi.org/10.48550/ARXIV.2602.24007) [](https://doi.org/10.48550/ARXIV.2506.04490) [](https://doi.org/10.1038/s41592-026-03047-4). This helps establish a self-reinforcing loop ([Figure 6](#figure6)). Training data derived from qFit can improve the ability to predict a larger portion of the conformational ensemble. Using these improved structure predictors, we can then use them alongside guidance frameworks, enabling the discovery of richer heterogeneity, which in turn generates higher-quality ensemble training data for subsequent prediction algorithms.\n\n::::::figure{#figure6 align=\"center\" type=\"image\" label=\"Figure 6\"}\n\n:::::image{width=\"100%\" src=\"https://thestacks-01.s3.amazonaws.com/publications/qfit-at-scale/media_73316a4e_745556479dbf\"}\n:::::\n\n:::::figcaption\n**Figure 6.** A self-reinforcing loop for conformational ensemble prediction. Training data derived from qFit improves structure predictors' ability to capture conformational ensembles (modeling heterogeneity). These improved predictors, combined with guidance frameworks, enable discovery of richer heterogeneity. This richer heterogeneity, in turn, generates higher-quality training data for the next generation of predictive algorithms.\n:::::\n\n::::::\n\nRealizing this loop also depends on how we encode this information. The experimental data capture conformational heterogeneity spanning length scales from single atoms to loop movements. There is also considerable compositional heterogeneity within experimental data. While qFit and a growing number of approaches can model this complex heterogeneity landscape, the current PDBx/mmCIF format does not support its description or encoding well [](https://doi.org/10.1107/s2052252524005098). This means that even when heterogeneity is modeled, much of it cannot be communicated through a model. This limits the ensemble information available to describe biological phenomena and the prediction methods used for training.\n\nFinally, closing the loop will also require new infrastructure, including tools to share, query, and compare models, built around a living database that treats experimental data as ground truth and continuously incorporates modeling improvements as they emerge. The circular prediction model we present above depends on capturing this information at each pass ([Figure 6](#figure6)). As ensemble-based methods become more accurate and computationally accessible, a unified and continually updated repository would let those improvements propagate easily to downstream users, enabling new biological insights rather than locking them behind the original depositors' modeling choices.\n\n::::::div{.bordered.lead data-internal=\"acknowledgements\"}\n### Acknowledgements\n\nThis material is based upon work supported by the Defense Advanced Research Projects Agency under this Agreement (SAW received research support). The views and conclusions contained in this document are those of the authors and should not be interpreted as representing the official policies, either expressed or implied, of the U.S. Government. This work leveraged the Research Software Engineer Services provided by the Vanderbilt Advanced Computing Center for Research and Education (ACCRE), operated by and for Vanderbilt faculty.\n::::::\n\n::::::details\n\n:::::summary\n# Supplementary Figures\n:::::\n\n:::::figure{#sfig1 align=\"center\" type=\"image\" unnumbered=\"true\" label}\n\n::::image{width=\"75%\" src=\"https://thestacks-01.s3.amazonaws.com/publications/qfit-at-scale/media_924f40fb_f3b4690c8a1c\"}\n::::\n\n::::figcaption\nSupplementary Figure 1. Per-residue altloc count differences between qFit multiconformer models and deposited Final models (qFit minus Final), shown separately for each canonical amino acid type. Positive values indicate additional modeled conformers in qFit. Across all residue types, the median difference is zero, with about 89-92 percent of residues unchanged. Mean increases are modest and consistent across types, ranging from 0.11 for GLY and ALA to 0.16 for LEU and CYS. Among residues that change, gains predominate: increases of one or two conformers are most common.\n::::\n\n:::::\n\n:::::figure{#sfig2 align=\"center\" type=\"image\" unnumbered=\"true\" label}\n\n::::image{width=\"75%\" src=\"https://thestacks-01.s3.amazonaws.com/publications/qfit-at-scale/media_a3556b06_745556479dbf\"}\n::::\n\n::::figcaption\nSupplementary Figure 2. Per-residue RMSF differences between qFit multiconformer models and deposited Final models (qFit minus Final), shown separately for each canonical amino acid type. Positive values indicate greater modeled conformational heterogeneity in qFit. For every residue type, the median difference is at or near zero while the mean is positive, reflecting a right-skewed distribution in which a subset of residues gains substantial flexibility rather than a uniform shift. Mean increases are largest for long charged side chains (LYS 0.23 Å, ARG 0.20 Å) and smallest for small or conformationally restricted residues (GLY 0.05 Å, ALA 0.06 Å, PRO 0.08 Å). Per-residue counts range from 6,233 (CYS) to 39,107 (LEU).\n::::\n\n:::::\n\n::::::\n\n::::::bibtex\n@article{jumper2021, author = {Jumper, John and Evans, Richard and Pritzel, Alexander and Green, Tim and Figurnov, Michael and Ronneberger, Olaf and Tunyasuvunakool, Kathryn and Bates, Russ and Žídek, Augustin and Potapenko, Anna and Bridgland, Alex and Meyer, Clemens and Kohl, Simon A. A. and Ballard, Andrew J. and Cowie, Andrew and Romera-Paredes, Bernardino and Nikolov, Stanislav and Jain, Rishub and Adler, Jonas and Back, Trevor and Petersen, Stig and Reiman, David and Clancy, Ellen and Zielinski, Michal and Steinegger, Martin and Pacholska, Michalina and Berghammer, Tamas and Bodenstein, Sebastian and Silver, David and Vinyals, Oriol and Senior, Andrew W. and Kavukcuoglu, Koray and Kohli, Pushmeet and Hassabis, Demis}, title = {Highly accurate protein structure prediction with AlphaFold}, journal = {Nature}, volume = {596}, number = {7873}, pages = {583–589}, publisher = {Springer Science and Business Media LLC}, year = {2021}, month = {July}, doi = {10.1038/s41586-021-03819-2}, url = {https://doi.org/10.1038/s41586-021-03819-2} }\n\n@article{abramson2024, author = {Abramson, Josh and Adler, Jonas and Dunger, Jack and Evans, Richard and Green, Tim and Pritzel, Alexander and Ronneberger, Olaf and Willmore, Lindsay and Ballard, Andrew J. and Bambrick, Joshua and Bodenstein, Sebastian W. and Evans, David A. and Hung, Chia-Chun and O’Neill, Michael and Reiman, David and Tunyasuvunakool, Kathryn and Wu, Zachary and Žemgulytė, Akvilė and Arvaniti, Eirini and Beattie, Charles and Bertolli, Ottavia and Bridgland, Alex and Cherepanov, Alexey and Congreve, Miles and Cowen-Rivers, Alexander I. and Cowie, Andrew and Figurnov, Michael and Fuchs, Fabian B. and Gladman, Hannah and Jain, Rishub and Khan, Yousuf A. and Low, Caroline M. R. and Perlin, Kuba and Potapenko, Anna and Savy, Pascal and Singh, Sukhdeep and Stecula, Adrian and Thillaisundaram, Ashok and Tong, Catherine and Yakneen, Sergei and Zhong, Ellen D. and Zielinski, Michal and Žídek, Augustin and Bapst, Victor and Kohli, Pushmeet and Jaderberg, Max and Hassabis, Demis and Jumper, John M.}, title = {Accurate structure prediction of biomolecular interactions with AlphaFold 3}, journal = {Nature}, volume = {630}, number = {8016}, pages = {493–500}, publisher = {Springer Science and Business Media LLC}, year = {2024}, month = {May}, doi = {10.1038/s41586-024-07487-w}, url = {https://doi.org/10.1038/s41586-024-07487-w} }\n\n@article{passaro2025, author = {Passaro, Saro and Corso, Gabriele and Wohlwend, Jeremy and Reveiz, Mateo and Thaler, Stephan and Somnath, Vignesh Ram and Getz, Noah and Portnoi, Tally and Roy, Julien and Stark, Hannes and Kwabi-Addo, David and Beaini, Dominique and Jaakkola, Tommi and Barzilay, Regina}, title = {Boltz-2: Towards Accurate and Efficient Binding Affinity Prediction}, publisher = {openRxiv}, year = {2025}, month = {June}, doi = {10.1101/2025.06.14.659707}, url = {https://doi.org/10.1101/2025.06.14.659707} }\n\n@article{protenixteam2026, author = {{Protenix Team} and Zhang, Yuxuan and Gong, Chengyue and Zhang, Hanyu and Ma, Wenzhi and Liu, Zhenyu and Chen, Xinshi and Guan, Jiaqi and Wang, Lan and Yang, Yanping and Xia, Yu and Xiao, Wenzhi}, title = {Protenix-v1: Toward High-Accuracy Open-Source Biomolecular Structure Prediction}, publisher = {openRxiv}, year = {2026}, month = {Feb}, doi = {10.64898/2026.02.05.703733}, url = {https://doi.org/10.64898/2026.02.05.703733} }\n\n@article{baek2021, author = {Baek, Minkyung and DiMaio, Frank and Anishchenko, Ivan and Dauparas, Justas and Ovchinnikov, Sergey and Lee, Gyu Rie and Wang, Jue and Cong, Qian and Kinch, Lisa N. and Schaeffer, R. Dustin and Millán, Claudia and Park, Hahnbeom and Adams, Carson and Glassman, Caleb R. and DeGiovanni, Andy and Pereira, Jose H. and Rodrigues, Andria V. and van Dijk, Alberdina A. and Ebrecht, Ana C. and Opperman, Diederik J. and Sagmeister, Theo and Buhlheller, Christoph and Pavkov-Keller, Tea and Rathinaswamy, Manoj K. and Dalwadi, Udit and Yip, Calvin K. and Burke, John E. and Garcia, K. Christopher and Grishin, Nick V. and Adams, Paul D. and Read, Randy J. and Baker, David}, title = {Accurate prediction of protein structures and interactions using a three-track neural network}, journal = {Science}, volume = {373}, number = {6557}, pages = {871–876}, publisher = {American Association for the Advancement of Science (AAAS)}, year = {2021}, month = {Aug}, doi = {10.1126/science.abj8754}, url = {https://doi.org/10.1126/science.abj8754} }\n\n@article{chaidiscovery2024, author = {{Chai Discovery} and Boitreaud, Jacques and Dent, Jack and McPartlon, Matthew and Meier, Joshua and Reis, Vinicius and Rogozhnikov, Alex and Wu, Kevin}, title = {Chai-1: Decoding the molecular interactions of life}, publisher = {openRxiv}, year = {2024}, month = {Oct}, doi = {10.1101/2024.10.10.615955}, url = {https://doi.org/10.1101/2024.10.10.615955} }\n\n@article{bytedanceamlai4sciencetea2025, author = {{ByteDance AML AI4Science Team} and Chen, Xinshi and Zhang, Yuxuan and Lu, Chan and Ma, Wenzhi and Guan, Jiaqi and Gong, Chengyue and Yang, Jincai and Zhang, Hanyu and Zhang, Ke and Wu, Shenghao and Zhou, Kuangqi and Yang, Yanping and Liu, Zhenyu and Wang, Lan and Shi, Bo and Shi, Shaochen and Xiao, Wenzhi}, title = {Protenix - Advancing Structure Prediction Through a Comprehensive AlphaFold3 Reproduction}, publisher = {openRxiv}, year = {2025}, month = {Jan}, doi = {10.1101/2025.01.08.631967}, url = {https://doi.org/10.1101/2025.01.08.631967} }\n\n@article{hilser2006, author = {Hilser, Vincent J. and García-Moreno E., Bertrand and Oas, Terrence G. and Kapp, Greg and Whitten, Steven T.}, title = {A Statistical Thermodynamic Model of the Protein Ensemble}, journal = {Chemical Reviews}, volume = {106}, number = {5}, pages = {1545–1558}, publisher = {American Chemical Society (ACS)}, year = {2006}, month = {Mar}, doi = {10.1021/cr040423+}, url = {https://doi.org/10.1021/cr040423+} }\n\n@article{hilser2025, author = {Hilser, Vincent J. and Wrabl, James O. and Millard, Charles E.F. and Schmitz, Anna and Brantley, Sarah J. and Pearce, Marie and Rehfus, Joe and Russo, Miranda M. and Voortman-Sheetz, Keila}, title = {Statistical Thermodynamics of the Protein Ensemble: Mediating Function and Evolution}, journal = {Annual Review of Biophysics}, volume = {54}, number = {1}, pages = {227–247}, publisher = {Annual Reviews}, year = {2025}, month = {May}, doi = {10.1146/annurev-biophys-061824-104900}, url = {https://doi.org/10.1146/annurev-biophys-061824-104900} }\n\n@article{boehr2009, author = {Boehr, David D and Nussinov, Ruth and Wright, Peter E}, title = {The role of dynamic conformational ensembles in biomolecular recognition}, journal = {Nature Chemical Biology}, volume = {5}, number = {11}, pages = {789–796}, publisher = {Springer Science and Business Media LLC}, year = {2009}, month = {Oct}, doi = {10.1038/nchembio.232}, url = {https://doi.org/10.1038/nchembio.232} }\n\n@article{wankowicz2026, author = {Wankowicz, Stephanie A. and Bonomi, Massimiliano}, title = {From possibility to precision in macromolecular ensemble prediction}, journal = {Nature Methods}, volume = {23}, number = {6}, pages = {1100–1108}, publisher = {Springer Science and Business Media LLC}, year = {2026}, month = {May}, doi = {10.1038/s41592-026-03084-z}, url = {https://doi.org/10.1038/s41592-026-03084-z} }\n\n@article{lewis2025, author = {Lewis, Sarah and Hempel, Tim and Jiménez-Luna, José and Gastegger, Michael and Xie, Yu and Foong, Andrew Y. K. and Satorras, Victor García and Abdin, Osama and Veeling, Bastiaan S. and Zaporozhets, Iryna and Chen, Yaoyi and Yang, Soojung and Foster, Adam E. and Schneuing, Arne and Nigam, Jigyasa and Barbero, Federico and Stimper, Vincent and Campbell, Andrew and Yim, Jason and Lienen, Marten and Shi, Yu and Zheng, Shuxin and Schulz, Hannes and Munir, Usman and Sordillo, Roberto and Tomioka, Ryota and Clementi, Cecilia and Noé, Frank}, title = {Scalable emulation of protein equilibrium ensembles with generative deep learning}, journal = {Science}, volume = {389}, number = {6761}, publisher = {American Association for the Advancement of Science (AAAS)}, year = {2025}, month = {Aug}, doi = {10.1126/science.adv9817}, url = {https://doi.org/10.1126/science.adv9817} }\n\n@article{jing2026, author = {Jing, Bowen and Berger, Bonnie and Jaakkola, Tommi}, title = {AI-based methods for simulating, sampling, and predicting protein ensembles}, journal = {Current Opinion in Structural Biology}, volume = {98}, pages = {103251}, publisher = {Elsevier BV}, year = {2026}, month = {June}, doi = {10.1016/j.sbi.2026.103251}, url = {https://doi.org/10.1016/j.sbi.2026.103251} }\n\n@misc{jing2024, author = {Jing, Bowen and Berger, Bonnie and Jaakkola, Tommi}, title = {AlphaFold Meets Flow Matching for Generating Protein Ensembles}, publisher = {arXiv}, year = {2024}, doi = {10.48550/ARXIV.2402.04845}, url = {https://doi.org/10.48550/ARXIV.2402.04845} }\n\n@article{hollingsworth2018, author = {Hollingsworth, Scott A. and Dror, Ron O.}, title = {Molecular Dynamics Simulation for All}, journal = {Neuron}, volume = {99}, number = {6}, pages = {1129–1143}, publisher = {Elsevier BV}, year = {2018}, month = {Sept}, doi = {10.1016/j.neuron.2018.08.011}, url = {https://doi.org/10.1016/j.neuron.2018.08.011} }\n\n@article{robustelli2018, author = {Robustelli, Paul and Piana, Stefano and Shaw, David E.}, title = {Developing a molecular dynamics force field for both folded and disordered protein states}, journal = {Proceedings of the National Academy of Sciences}, volume = {115}, number = {21}, publisher = {National Academy of Sciences}, year = {2018}, month = {May}, doi = {10.1073/pnas.1800690115}, url = {https://doi.org/10.1073/pnas.1800690115} }\n\n@article{pedersen2022, author = {Pedersen, Kasper B. and Flores‐Canales, Jose C. and Schiøtt, Birgit}, title = {Predicting molecular properties of α‐synuclein using force fields for intrinsically disordered proteins}, journal = {Proteins: Structure, Function, and Bioinformatics}, volume = {91}, number = {1}, pages = {47–61}, publisher = {Wiley}, year = {2022}, month = {Aug}, doi = {10.1002/prot.26409}, url = {https://doi.org/10.1002/prot.26409} }\n\n@article{karplus2012, author = {Karplus, P. Andrew and Diederichs, Kay}, title = {Linking Crystallographic Model and Data Quality}, journal = {Science}, volume = {336}, number = {6084}, pages = {1030–1033}, publisher = {American Association for the Advancement of Science (AAAS)}, year = {2012}, month = {May}, doi = {10.1126/science.1218231}, url = {https://doi.org/10.1126/science.1218231} }\n\n@article{burley2021, author = {Burley, Stephen K. and Berman, Helen M.}, title = {Open-access data: A cornerstone for artificial intelligence approaches to protein structure prediction}, journal = {Structure}, volume = {29}, number = {6}, pages = {515–520}, publisher = {Elsevier BV}, year = {2021}, month = {June}, doi = {10.1016/j.str.2021.04.010}, url = {https://doi.org/10.1016/j.str.2021.04.010} }\n\n@article{furnham2006, author = {Furnham, Nicholas and Blundell, Tom L and DePristo, Mark A and Terwilliger, Thomas C}, title = {Is one solution good enough?}, journal = {Nature Structural & Molecular Biology}, volume = {13}, number = {3}, pages = {184–185}, publisher = {Springer Science and Business Media LLC}, year = {2006}, month = {Mar}, doi = {10.1038/nsmb0306-184}, url = {https://doi.org/10.1038/nsmb0306-184} }\n\n@article{lane2023, author = {Lane, Thomas J.}, title = {Protein structure prediction has reached the single-structure frontier}, journal = {Nature Methods}, volume = {20}, number = {2}, pages = {170–173}, publisher = {Springer Science and Business Media LLC}, year = {2023}, month = {Jan}, doi = {10.1038/s41592-022-01760-4}, url = {https://doi.org/10.1038/s41592-022-01760-4} }\n\n@article{emsley2004, author = {Emsley, Paul and Cowtan, Kevin}, title = {Coot: model-building tools for molecular graphics}, journal = {Acta Crystallographica Section D Biological Crystallography}, volume = {60}, number = {12}, pages = {2126–2132}, publisher = {International Union of Crystallography (IUCr)}, year = {2004}, month = {Nov}, doi = {10.1107/s0907444904019158}, url = {https://doi.org/10.1107/s0907444904019158} }\n\n@article{weichenberger2015, author = {Weichenberger, Christian X. and Afonine, Pavel V. and Kantardjieff, Katherine and Rupp, Bernhard}, title = {The solvent component of macromolecular crystals}, journal = {Acta Crystallographica Section D Biological Crystallography}, volume = {71}, number = {5}, pages = {1023–1038}, publisher = {International Union of Crystallography (IUCr)}, year = {2015}, month = {Apr}, doi = {10.1107/s1399004715006045}, url = {https://doi.org/10.1107/s1399004715006045} }\n\n@article{kabsch2010, author = {Kabsch, Wolfgang}, title = {XDS}, journal = {Acta Crystallographica Section D Biological Crystallography}, volume = {66}, number = {2}, pages = {125–132}, publisher = {International Union of Crystallography (IUCr)}, year = {2010}, month = {Jan}, doi = {10.1107/s0907444909047337}, url = {https://doi.org/10.1107/s0907444909047337} }\n\n@article{keedy2015, author = {Keedy, Daniel A. and Fraser, James S. and van den Bedem, Henry}, title = {Exposing Hidden Alternative Backbone Conformations in X-ray Crystallography Using qFit}, journal = {PLOS Computational Biology}, volume = {11}, number = {10}, pages = {e1004507}, publisher = {Public Library of Science (PLoS)}, year = {2015}, month = {Oct}, doi = {10.1371/journal.pcbi.1004507}, url = {https://doi.org/10.1371/journal.pcbi.1004507} }\n\n@article{riley2020, author = {Riley, Blake T. and Wankowicz, Stephanie A. and de Oliveira, Saulo H. P. and van Zundert, Gydo C. P. and Hogan, Daniel W. and Fraser, James S. and Keedy, Daniel A. and van den Bedem, Henry}, title = {qFit 3: Protein and ligand multiconformer modeling for X‐ray crystallographic and single‐particle cryo‐EM density maps}, journal = {Protein Science}, volume = {30}, number = {1}, pages = {270–285}, publisher = {Wiley}, year = {2020}, month = {Nov}, doi = {10.1002/pro.4001}, url = {https://doi.org/10.1002/pro.4001} }\n\n@article{vandenbedem2009, author = {van den Bedem, Henry and Dhanik, Ankur and Latombe, Jean-Claude and Deacon, Ashley M.}, title = {Modeling discrete heterogeneity in X-ray diffraction data by fitting multi-conformers}, journal = {Acta Crystallographica Section D Biological Crystallography}, volume = {65}, number = {10}, pages = {1107–1117}, publisher = {International Union of Crystallography (IUCr)}, year = {2009}, month = {Sept}, doi = {10.1107/s0907444909030613}, url = {https://doi.org/10.1107/s0907444909030613} }\n\n@article{wankowicz2024, author = {Wankowicz, Stephanie A and Ravikumar, Ashraya and Sharma, Shivani and Riley, Blake and Raju, Akshay and Hogan, Daniel W and Flowers, Jessica and van den Bedem, Henry and Keedy, Daniel A and Fraser, James S}, title = {Automated multiconformer model building for X-ray crystallography and cryo-EM}, journal = {eLife}, volume = {12}, publisher = {eLife Sciences Publications, Ltd}, year = {2024}, month = {June}, doi = {10.7554/elife.90606}, url = {https://doi.org/10.7554/elife.90606} }\n\n@article{woldeyes2014, author = {Woldeyes, Rahel A and Sivak, David A and Fraser, James S}, title = {E pluribus unum, no more: from one crystal, many conformations}, journal = {Current Opinion in Structural Biology}, volume = {28}, pages = {56–62}, publisher = {Elsevier BV}, year = {2014}, month = {Oct}, doi = {10.1016/j.sbi.2014.07.005}, url = {https://doi.org/10.1016/j.sbi.2014.07.005} }\n\n@article{seo2025, author = {Seo, Louella and Farran, Ian and Aslam, Ahmed and Li, Xinyun and Jaishankar, Priyadarshini and Ashworth, Alan and Fraser, James S. and Renslo, Adam R. and Wankowicz, Stephanie A.}, title = {Crystallographic Ensembles Reveal the Structural Basis of Binding Entropy in SARS-CoV2 Macrodomain}, publisher = {openRxiv}, year = {2025}, month = {Nov}, doi = {10.1101/2025.11.25.690589}, url = {https://doi.org/10.1101/2025.11.25.690589} }\n\n@article{miller2026, author = {Miller, Charlotte A and Wankowicz, Stephanie A}, title = {Binding Entropy Can Be Predicted by Crystallographic Ensembles}, publisher = {eLife Sciences Publications, Ltd}, year = {2026}, month = {May}, doi = {10.7554/elife.111298.1}, url = {https://doi.org/10.7554/elife.111298.1} }\n\n@article{keedy2019, author = {Keedy, Daniel A.}, title = {Journey to the center of the protein: allostery from multitemperature multiconformer X-ray crystallography}, journal = {Acta Crystallographica Section D Structural Biology}, volume = {75}, number = {2}, pages = {123–137}, publisher = {International Union of Crystallography (IUCr)}, year = {2019}, month = {Jan}, doi = {10.1107/s2059798318017941}, url = {https://doi.org/10.1107/s2059798318017941} }\n\n@article{afonine2012, author = {Afonine, Pavel V. and Grosse-Kunstleve, Ralf W. and Echols, Nathaniel and Headd, Jeffrey J. and Moriarty, Nigel W. and Mustyakimov, Marat and Terwilliger, Thomas C. and Urzhumtsev, Alexandre and Zwart, Peter H. and Adams, Paul D.}, title = {Towards automated crystallographic structure refinement with phenix.refine}, journal = {Acta Crystallographica Section D Biological Crystallography}, volume = {68}, number = {4}, pages = {352–367}, publisher = {International Union of Crystallography (IUCr)}, year = {2012}, month = {Mar}, doi = {10.1107/s0907444912001308}, url = {https://doi.org/10.1107/s0907444912001308} }\n\n@article{fraser2024, author = {Fraser, James S. and Murcko, Mark A.}, title = {Structure is beauty, but not always truth}, journal = {Cell}, volume = {187}, number = {3}, pages = {517–520}, publisher = {Elsevier BV}, year = {2024}, month = {Feb}, doi = {10.1016/j.cell.2024.01.003}, url = {https://doi.org/10.1016/j.cell.2024.01.003} }\n\n@article{holton2014, author = {Holton, James M. and Classen, Scott and Frankel, Kenneth A. and Tainer, John A.}, title = {The R‐factor gap in macromolecular crystallography: an untapped potential for insights on accurate structures}, journal = {The FEBS Journal}, volume = {281}, number = {18}, pages = {4046–4060}, publisher = {Wiley}, year = {2014}, month = {Sept}, doi = {10.1111/febs.12922}, url = {https://doi.org/10.1111/febs.12922} }\n\n@article{vankempen2023, author = {van Kempen, Michel and Kim, Stephanie S. and Tumescheit, Charlotte and Mirdita, Milot and Lee, Jeongjae and Gilchrist, Cameron L. M. and Söding, Johannes and Steinegger, Martin}, title = {Fast and accurate protein structure search with Foldseek}, journal = {Nature Biotechnology}, volume = {42}, number = {2}, pages = {243–246}, publisher = {Springer Science and Business Media LLC}, year = {2023}, month = {May}, doi = {10.1038/s41587-023-01773-0}, url = {https://doi.org/10.1038/s41587-023-01773-0} }\n\n@misc{holton2026, author = {Holton, James and Wankowicz, Stephanie A}, title = {Supplying a user-defined bulk solvent map for refinement}, publisher = {Radial}, year = {2026}, doi = {10.82153/2ME0-HD96}, url = {https://doi.org/10.82153/2ME0-HD96} }\n\n@article{hrabe2015, author = {Hrabe, Thomas and Li, Zhanwen and Sedova, Mayya and Rotkiewicz, Piotr and Jaroszewski, Lukasz and Godzik, Adam}, title = {PDBFlex: exploring flexibility in protein structures}, journal = {Nucleic Acids Research}, volume = {44}, number = {D1}, pages = {D423–D428}, publisher = {Oxford University Press (OUP)}, year = {2015}, month = {Nov}, doi = {10.1093/nar/gkv1316}, url = {https://doi.org/10.1093/nar/gkv1316} }\n\n@article{wankowicz2024a, author = {Wankowicz, Stephanie A.}, title = {Modeling Bias Toward Binding Sites in PDB Structural Models}, publisher = {openRxiv}, year = {2024}, month = {Dec}, doi = {10.1101/2024.12.14.628518}, url = {https://doi.org/10.1101/2024.12.14.628518} }\n\n@article{vandermeersche2023, author = {Vander Meersche, Yann and Cretin, Gabriel and Gheeraert, Aria and Gelly, Jean-Christophe and Galochkina, Tatiana}, title = {ATLAS: protein flexibility description from atomistic molecular dynamics simulations}, journal = {Nucleic Acids Research}, volume = {52}, number = {D1}, pages = {D384–D392}, publisher = {Oxford University Press (OUP)}, year = {2023}, month = {Nov}, doi = {10.1093/nar/gkad1084}, url = {https://doi.org/10.1093/nar/gkad1084} }\n\n@article{mirarchi2024, author = {Mirarchi, Antonio and Giorgino, Toni and De Fabritiis, Gianni}, title = {mdCATH: A Large-Scale MD Dataset for Data-Driven Computational Biophysics}, journal = {Scientific Data}, volume = {11}, number = {1}, publisher = {Springer Science and Business Media LLC}, year = {2024}, month = {Nov}, doi = {10.1038/s41597-024-04140-z}, url = {https://doi.org/10.1038/s41597-024-04140-z} }\n\n@article{siebenmorgen2024, author = {Siebenmorgen, Till and Menezes, Filipe and Benassou, Sabrina and Merdivan, Erinc and Didi, Kieran and Mourão, André Santos Dias and Kitel, Radosław and Liò, Pietro and Kesselheim, Stefan and Piraud, Marie and Theis, Fabian J. and Sattler, Michael and Popowicz, Grzegorz M.}, title = {MISATO: machine learning dataset of protein–ligand complexes for structure-based drug discovery}, journal = {Nature Computational Science}, volume = {4}, number = {5}, pages = {367–378}, publisher = {Springer Science and Business Media LLC}, year = {2024}, month = {May}, doi = {10.1038/s43588-024-00627-2}, url = {https://doi.org/10.1038/s43588-024-00627-2} }\n\n@article{rodriguezespigares2020, author = {Rodríguez-Espigares, Ismael and Torrens-Fontanals, Mariona and Tiemann, Johanna K. S. and Aranda-García, David and Ramírez-Anguita, Juan Manuel and Stepniewski, Tomasz Maciej and Worp, Nathalie and Varela-Rial, Alejandro and Morales-Pastor, Adrián and Medel-Lacruz, Brian and Pándy-Szekeres, Gáspár and Mayol, Eduardo and Giorgino, Toni and Carlsson, Jens and Deupi, Xavier and Filipek, Slawomir and Filizola, Marta and Gómez-Tamayo, José Carlos and Gonzalez, Angel and Gutiérrez-de-Terán, Hugo and Jiménez-Rosés, Mireia and Jespers, Willem and Kapla, Jon and Khelashvili, George and Kolb, Peter and Latek, Dorota and Marti-Solano, Maria and Matricon, Pierre and Matsoukas, Minos-Timotheos and Miszta, Przemyslaw and Olivella, Mireia and Perez-Benito, Laura and Provasi, Davide and Ríos, Santiago and R. Torrecillas, Iván and Sallander, Jessica and Sztyler, Agnieszka and Vasile, Silvana and Weinstein, Harel and Zachariae, Ulrich and Hildebrand, Peter W. and De Fabritiis, Gianni and Sanz, Ferran and Gloriam, David E. and Cordomi, Arnau and Guixà-González, Ramon and Selent, Jana}, title = {GPCRmd uncovers the dynamics of the 3D-GPCRome}, journal = {Nature Methods}, volume = {17}, number = {8}, pages = {777–787}, publisher = {Springer Science and Business Media LLC}, year = {2020}, month = {July}, doi = {10.1038/s41592-020-0884-y}, url = {https://doi.org/10.1038/s41592-020-0884-y} }\n\n@article{cao2024, author = {Cao, Fan and von Bülow, Sören and Tesei, Giulio and Lindorff‐Larsen, Kresten}, title = {A coarse‐grained model for disordered and multi‐domain proteins}, journal = {Protein Science}, volume = {33}, number = {11}, publisher = {Wiley}, year = {2024}, month = {Oct}, doi = {10.1002/pro.5172}, url = {https://doi.org/10.1002/pro.5172} }\n\n@article{henzlerwildman2007, author = {Henzler-Wildman, Katherine and Kern, Dorothee}, title = {Dynamic personalities of proteins}, journal = {Nature}, volume = {450}, number = {7172}, pages = {964–972}, publisher = {Springer Science and Business Media LLC}, year = {2007}, month = {Dec}, doi = {10.1038/nature06522}, url = {https://doi.org/10.1038/nature06522} }\n\n@article{lyu2025, author = {Lyu, Ningyi and Du, Siyuan and Shao, Qianzhen and Yang, Zhongyue and Ma, Jianpeng and Herschlag, Daniel}, title = {Physics-Grounded Evaluation to Guide Accurate Biomolecular Prediction}, publisher = {openRxiv}, year = {2025}, month = {July}, doi = {10.1101/2025.06.30.662466}, url = {https://doi.org/10.1101/2025.06.30.662466} }\n\n@misc{chrispens2026, author = {Chrispens, Karson and Collins, Marcus and Fraser, James S. and Mai, Doris and van den Bedem, Henry and Wankowicz, Stephanie A}, title = {sampleworks: A Modular Platform for Experimentally Guided Biomolecular Ensemble Generation}, publisher = {Radial}, year = {2026}, doi = {10.82153/JKXJ-TW08}, url = {https://doi.org/10.82153/JKXJ-TW08} }\n\n@misc{maddipatla2026, author = {Maddipatla, Advaith and Rzayev, Anar and Pegoraro, Marco and Pacesa, Martin and Schanda, Paul and Marx, Ailie and Vedula, Sanketh and Bronstein, Alex M.}, title = {Inference-time optimization for experiment-grounded protein ensemble generation}, publisher = {arXiv}, year = {2026}, doi = {10.48550/ARXIV.2602.24007}, url = {https://doi.org/10.48550/ARXIV.2602.24007} }\n\n@misc{raghu2025, author = {Raghu, Rishwanth and Levy, Axel and Wetzstein, Gordon and Zhong, Ellen D.}, title = {Multiscale guidance of protein structure prediction with heterogeneous cryo-EM data}, publisher = {arXiv}, year = {2025}, doi = {10.48550/ARXIV.2506.04490}, url = {https://doi.org/10.48550/ARXIV.2506.04490} }\n\n@article{fadini2026, author = {Fadini, Alisia and Li, Minhuan and McCoy, Airlie J. and Banjara, Suresh and Okumura, Hiroki and Napier, Eve and Fontana, Pietro and Khan, Amir R. and Jovine, Luca and Terwilliger, Thomas C. and Read, Randy J. and Hekstra, Doeke R. and AlQuraishi, Mohammed}, title = {AlphaFold as a prior: experimental structure determination conditioned on a pretrained neural network}, journal = {Nature Methods}, volume = {23}, number = {4}, pages = {785–795}, publisher = {Springer Science and Business Media LLC}, year = {2026}, month = {Apr}, doi = {10.1038/s41592-026-03047-4}, url = {https://doi.org/10.1038/s41592-026-03047-4} }\n\n@article{wankowicz2024b, author = {Wankowicz, Stephanie A. and Fraser, James S.}, title = {Comprehensive encoding of conformational and compositional protein structural ensembles through the mmCIF data structure}, journal = {IUCrJ}, volume = {11}, number = {4}, pages = {494–501}, publisher = {International Union of Crystallography (IUCr)}, year = {2024}, month = {June}, doi = {10.1107/s2052252524005098}, url = {https://doi.org/10.1107/s2052252524005098} }\n\n@article{steinegger2017, author = {Steinegger, Martin and Söding, Johannes}, title = {MMseqs2 enables sensitive protein sequence searching for the analysis of massive data sets}, journal = {Nature Biotechnology}, volume = {35}, number = {11}, pages = {1026–1028}, publisher = {Springer Science and Business Media LLC}, year = {2017}, month = {Oct}, doi = {10.1038/nbt.3988}, url = {https://doi.org/10.1038/nbt.3988} }\n::::::","contributors":[{"user_id":5245,"role":"Writing","first_name":"Stephanie A","last_name":"Wankowicz","avatar":"avatars/users/user_5245_1781698403.jpg","in_byline":1,"priority":null,"affiliations":[{"org_id":7,"name":"Radial","slug":"radial","ror_id":"050rbg919","avatar":"avatars/orgs/org_7_1780396782.jpg"},{"org_id":12,"name":"Vanderbilt University","slug":"vanderbilt","ror_id":"02vm5rt34","avatar":null}]},{"user_id":5245,"role":"Investigation","first_name":"Stephanie A","last_name":"Wankowicz","avatar":"avatars/users/user_5245_1781698403.jpg","in_byline":1,"priority":null,"affiliations":[{"org_id":7,"name":"Radial","slug":"radial","ror_id":"050rbg919","avatar":"avatars/orgs/org_7_1780396782.jpg"},{"org_id":12,"name":"Vanderbilt University","slug":"vanderbilt","ror_id":"02vm5rt34","avatar":null}]},{"user_id":5245,"role":"Data curation","first_name":"Stephanie A","last_name":"Wankowicz","avatar":"avatars/users/user_5245_1781698403.jpg","in_byline":0,"priority":null,"affiliations":[{"org_id":7,"name":"Radial","slug":"radial","ror_id":"050rbg919","avatar":"avatars/orgs/org_7_1780396782.jpg"},{"org_id":12,"name":"Vanderbilt University","slug":"vanderbilt","ror_id":"02vm5rt34","avatar":null}]},{"user_id":5245,"role":"Formal analysis","first_name":"Stephanie A","last_name":"Wankowicz","avatar":"avatars/users/user_5245_1781698403.jpg","in_byline":1,"priority":null,"affiliations":[{"org_id":7,"name":"Radial","slug":"radial","ror_id":"050rbg919","avatar":"avatars/orgs/org_7_1780396782.jpg"},{"org_id":12,"name":"Vanderbilt University","slug":"vanderbilt","ror_id":"02vm5rt34","avatar":null}]},{"user_id":5245,"role":"Conceptualization","first_name":"Stephanie A","last_name":"Wankowicz","avatar":"avatars/users/user_5245_1781698403.jpg","in_byline":1,"priority":null,"affiliations":[{"org_id":7,"name":"Radial","slug":"radial","ror_id":"050rbg919","avatar":"avatars/orgs/org_7_1780396782.jpg"},{"org_id":12,"name":"Vanderbilt University","slug":"vanderbilt","ror_id":"02vm5rt34","avatar":null}]},{"user_id":5245,"role":"Methodology","first_name":"Stephanie A","last_name":"Wankowicz","avatar":"avatars/users/user_5245_1781698403.jpg","in_byline":1,"priority":null,"affiliations":[{"org_id":7,"name":"Radial","slug":"radial","ror_id":"050rbg919","avatar":"avatars/orgs/org_7_1780396782.jpg"},{"org_id":12,"name":"Vanderbilt University","slug":"vanderbilt","ror_id":"02vm5rt34","avatar":null}]},{"user_id":5245,"role":"Software","first_name":"Stephanie A","last_name":"Wankowicz","avatar":"avatars/users/user_5245_1781698403.jpg","in_byline":1,"priority":null,"affiliations":[{"org_id":7,"name":"Radial","slug":"radial","ror_id":"050rbg919","avatar":"avatars/orgs/org_7_1780396782.jpg"},{"org_id":12,"name":"Vanderbilt University","slug":"vanderbilt","ror_id":"02vm5rt34","avatar":null}]}],"created_at":"2026-08-10T22:33:50.000Z","doi":"10.82153/pff0-ck46","version_doi":"10.82153/fjj4-gv26","feedback_form_embed_src":null,"id":175,"license":"CC BY","linked_assets":[{"type":"Code","url":"https://github.com/ExcitedStates/qfit-3.0","name":""},{"type":"Data","url":"https://zenodo.org/records/20801853","name":"Data"}],"orgs":[{"org_id":7,"priority":null,"slug":"radial","name":"Radial","avatar":"avatars/orgs/org_7_1780396782.jpg"}],"pdf_url":"https://thestacks-01.s3.us-west-2.amazonaws.com/publications/qfit-at-scale/qfit-at-scale_v1.pdf","references":[{"url":"https://doi.org/10.1038/s41586-021-03819-2","text":"Jumper J, Evans R, Pritzel A, Green T, Figurnov M, Ronneberger O, Tunyasuvunakool K, Bates R, Žídek A, Potapenko A, Bridgland A, Meyer C, Kohl SAA, Ballard AJ, Cowie A, Romera-Paredes B, Nikolov S, Jain R, Adler J, Back T, Petersen S, Reiman D, Clancy E, Zielinski M, Steinegger M, Pacholska M, Berghammer T, Bodenstein S, Silver D, Vinyals O, Senior AW, Kavukcuoglu K, Kohli P, Hassabis D. (2021). Highly accurate protein structure prediction with AlphaFold."},{"url":"https://doi.org/10.1038/s41586-024-07487-w","text":"Abramson J, Adler J, Dunger J, Evans R, Green T, Pritzel A, Ronneberger O, Willmore L, Ballard AJ, Bambrick J, Bodenstein SW, Evans DA, Hung C-C, O’Neill M, Reiman D, Tunyasuvunakool K, Wu Z, Žemgulytė A, Arvaniti E, Beattie C, Bertolli O, Bridgland A, Cherepanov A, Congreve M, Cowen-Rivers AI, Cowie A, Figurnov M, Fuchs FB, Gladman H, Jain R, Khan YA, Low CMR, Perlin K, Potapenko A, Savy P, Singh S, Stecula A, Thillaisundaram A, Tong C, Yakneen S, Zhong ED, Zielinski M, Žídek A, Bapst V, Kohli P, Jaderberg M, Hassabis D, Jumper JM. (2024). Accurate structure prediction of biomolecular interactions with AlphaFold 3."},{"url":"https://doi.org/10.1101/2025.06.14.659707","text":"Passaro S, Corso G, Wohlwend J, Reveiz M, Thaler S, Somnath VR, Getz N, Portnoi T, Roy J, Stark H, Kwabi-Addo D, Beaini D, Jaakkola T, Barzilay R. (2025). Boltz-2: Towards Accurate and Efficient Binding Affinity Prediction."},{"url":"https://doi.org/10.64898/2026.02.05.703733","text":"Team P, Zhang Y, Gong C, Zhang H, Ma W, Liu Z, Chen X, Guan J, Wang L, Yang Y, Xia Y, Xiao W. (2026). Protenix-v1: Toward High-Accuracy Open-Source Biomolecular Structure Prediction."},{"url":"https://doi.org/10.1126/science.abj8754","text":"Baek M, DiMaio F, Anishchenko I, Dauparas J, Ovchinnikov S, Lee GR, Wang J, Cong Q, Kinch LN, Schaeffer RD, Millán C, Park H, Adams C, Glassman CR, DeGiovanni A, Pereira JH, Rodrigues AV, van Dijk AA, Ebrecht AC, Opperman DJ, Sagmeister T, Buhlheller C, Pavkov-Keller T, Rathinaswamy MK, Dalwadi U, Yip CK, Burke JE, Garcia KC, Grishin NV, Adams PD, Read RJ, Baker D. (2021). Accurate prediction of protein structures and interactions using a three-track neural network."},{"url":"https://doi.org/10.1101/2024.10.10.615955","text":"Discovery C, Boitreaud J, Dent J, McPartlon M, Meier J, Reis V, Rogozhnikov A, Wu K. (2024). Chai-1: Decoding the molecular interactions of life."},{"url":"https://doi.org/10.1101/2025.01.08.631967","text":"Team BAA, Chen X, Zhang Y, Lu C, Ma W, Guan J, Gong C, Yang J, Zhang H, Zhang K, Wu S, Zhou K, Yang Y, Liu Z, Wang L, Shi B, Shi S, Xiao W. (2025). Protenix - Advancing Structure Prediction Through a Comprehensive AlphaFold3 Reproduction."},{"url":"https://doi.org/10.1021/cr040423+","text":"Hilser VJ, García-Moreno E. B, Oas TG, Kapp G, Whitten ST. (2006). A Statistical Thermodynamic Model of the Protein Ensemble."},{"url":"https://doi.org/10.1146/annurev-biophys-061824-104900","text":"Hilser VJ, Wrabl JO, Millard CE, Schmitz A, Brantley SJ, Pearce M, Rehfus J, Russo MM, Voortman-Sheetz K. (2025). Statistical Thermodynamics of the Protein Ensemble: Mediating Function and Evolution."},{"url":"https://doi.org/10.1038/nchembio.232","text":"Boehr DD, Nussinov R, Wright PE. (2009). The role of dynamic conformational ensembles in biomolecular recognition."},{"url":"https://doi.org/10.1038/s41592-026-03084-z","text":"Wankowicz SA, Bonomi M. (2026). From possibility to precision in macromolecular ensemble prediction."},{"url":"https://doi.org/10.1126/science.adv9817","text":"Lewis S, Hempel T, Jiménez-Luna J, Gastegger M, Xie Y, Foong AYK, Satorras VG, Abdin O, Veeling BS, Zaporozhets I, Chen Y, Yang S, Foster AE, Schneuing A, Nigam J, Barbero F, Stimper V, Campbell A, Yim J, Lienen M, Shi Y, Zheng S, Schulz H, Munir U, Sordillo R, Tomioka R, Clementi C, Noé F. (2025). Scalable emulation of protein equilibrium ensembles with generative deep learning."},{"url":"https://doi.org/10.1016/j.sbi.2026.103251","text":"Jing B, Berger B, Jaakkola T. (2026). AI-based methods for simulating, sampling, and predicting protein ensembles."},{"url":"https://doi.org/10.48550/arxiv.2402.04845","text":"Jing B, Berger B, Jaakkola T. (2024). AlphaFold Meets Flow Matching for Generating Protein Ensembles."},{"url":"https://doi.org/10.1016/j.neuron.2018.08.011","text":"Hollingsworth SA, Dror RO. (2018). Molecular Dynamics Simulation for All."},{"url":"https://doi.org/10.1073/pnas.1800690115","text":"Robustelli P, Piana S, Shaw DE. (2018). Developing a molecular dynamics force field for both folded and disordered protein states."},{"url":"https://doi.org/10.1002/prot.26409","text":"Pedersen KB, Flores‐Canales JC, Schiøtt B. (2022). Predicting molecular properties of <scp>α‐synuclein</scp> using force fields for intrinsically disordered proteins."},{"url":"https://doi.org/10.1126/science.1218231","text":"Karplus PA, Diederichs K. (2012). Linking Crystallographic Model and Data Quality."},{"url":"https://doi.org/10.1016/j.str.2021.04.010","text":"Burley SK, Berman HM. (2021). Open-access data: A cornerstone for artificial intelligence approaches to protein structure prediction."},{"url":"https://doi.org/10.1038/nsmb0306-184","text":"Furnham N, Blundell TL, DePristo MA, Terwilliger TC. (2006). Is one solution good enough?."},{"url":"https://doi.org/10.1038/s41592-022-01760-4","text":"Lane TJ. (2023). Protein structure prediction has reached the single-structure frontier."},{"url":"https://doi.org/10.1107/s0907444904019158","text":"Emsley P, Cowtan K. (2004). Coot: model-building tools for molecular graphics."},{"url":"https://doi.org/10.1107/s1399004715006045","text":"Weichenberger CX, Afonine PV, Kantardjieff K, Rupp B. (2015). The solvent component of macromolecular crystals."},{"url":"https://doi.org/10.1107/s0907444909047337","text":"Kabsch W. (2010). XDS."},{"url":"https://doi.org/10.1371/journal.pcbi.1004507","text":"Keedy DA, Fraser JS, van den Bedem H. (2015). Exposing Hidden Alternative Backbone Conformations in X-ray Crystallography Using qFit."},{"url":"https://doi.org/10.1002/pro.4001","text":"Riley BT, Wankowicz SA, de Oliveira SHP, van Zundert GCP, Hogan DW, Fraser JS, Keedy DA, van den Bedem H. (2020). <scp>qFit</scp> 3: Protein and ligand multiconformer modeling for X‐ray crystallographic and single‐particle <scp>cryo‐EM</scp> density maps."},{"url":"https://doi.org/10.1107/s0907444909030613","text":"van den Bedem H, Dhanik A, Latombe J, Deacon AM. (2009). Modeling discrete heterogeneity in X-ray diffraction data by fitting multi-conformers."},{"url":"https://doi.org/10.7554/elife.90606","text":"Wankowicz SA, Ravikumar A, Sharma S, Riley B, Raju A, Hogan DW, Flowers J, van den Bedem H, Keedy DA, Fraser JS. (2024). Automated multiconformer model building for X-ray crystallography and cryo-EM."},{"url":"https://doi.org/10.1016/j.sbi.2014.07.005","text":"Woldeyes RA, Sivak DA, Fraser JS. (2014). E pluribus unum, no more: from one crystal, many conformations."},{"url":"https://doi.org/10.1101/2025.11.25.690589","text":"Seo L, Farran I, Aslam A, Li X, Jaishankar P, Ashworth A, Fraser JS, Renslo AR, Wankowicz SA. (2025). Crystallographic Ensembles Reveal the Structural Basis of Binding Entropy in SARS-CoV2 Macrodomain."},{"url":"https://doi.org/10.7554/elife.111298.1","text":"Miller CA, Wankowicz SA. (2026). Binding Entropy Can Be Predicted by Crystallographic Ensembles."},{"url":"https://doi.org/10.1107/s2059798318017941","text":"Keedy DA. (2019). Journey to the center of the protein: allostery from multitemperature multiconformer X-ray crystallography."},{"url":"https://doi.org/10.1107/s0907444912001308","text":"Afonine PV, Grosse-Kunstleve RW, Echols N, Headd JJ, Moriarty NW, Mustyakimov M, Terwilliger TC, Urzhumtsev A, Zwart PH, Adams PD. (2012). Towards automated crystallographic structure refinement with phenix.refine."},{"url":"https://doi.org/10.1016/j.cell.2024.01.003","text":"Fraser JS, Murcko MA. (2024). Structure is beauty, but not always truth."},{"url":"https://doi.org/10.1111/febs.12922","text":"Holton JM, Classen S, Frankel KA, Tainer JA. (2014). The R‐factor gap in macromolecular crystallography: an untapped potential for insights on accurate structures."},{"url":"https://doi.org/10.1038/nbt.3988","text":"Steinegger M, Söding J. (2017). MMseqs2 enables sensitive protein sequence searching for the analysis of massive data sets."},{"url":"https://doi.org/10.1038/s41587-023-01773-0","text":"van Kempen M, Kim SS, Tumescheit C, Mirdita M, Lee J, Gilchrist CLM, Söding J, Steinegger M. (2023). Fast and accurate protein structure search with Foldseek."},{"url":"https://doi.org/10.82153/2me0-hd96","text":"Holton J, Wankowicz SA. (2026). Supplying a user-defined bulk solvent map for refinement."},{"url":"https://doi.org/10.1093/nar/gkv1316","text":"Hrabe T, Li Z, Sedova M, Rotkiewicz P, Jaroszewski L, Godzik A. (2015). PDBFlex: exploring flexibility in protein structures."},{"url":"https://doi.org/10.1101/2024.12.14.628518","text":"Wankowicz SA. (2024). Modeling Bias Toward Binding Sites in PDB Structural Models."},{"url":"https://doi.org/10.1093/nar/gkad1084","text":"Vander Meersche Y, Cretin G, Gheeraert A, Gelly J, Galochkina T. (2023). ATLAS: protein flexibility description from atomistic molecular dynamics simulations."},{"url":"https://doi.org/10.1038/s41597-024-04140-z","text":"Mirarchi A, Giorgino T, De Fabritiis G. (2024). mdCATH: A Large-Scale MD Dataset for Data-Driven Computational Biophysics."},{"url":"https://doi.org/10.1038/s43588-024-00627-2","text":"Siebenmorgen T, Menezes F, Benassou S, Merdivan E, Didi K, Mourão ASD, Kitel R, Liò P, Kesselheim S, Piraud M, Theis FJ, Sattler M, Popowicz GM. (2024). MISATO: machine learning dataset of protein–ligand complexes for structure-based drug discovery."},{"url":"https://doi.org/10.1038/s41592-020-0884-y","text":"Rodríguez-Espigares I, Torrens-Fontanals M, Tiemann JKS, Aranda-García D, Ramírez-Anguita JM, Stepniewski TM, Worp N, Varela-Rial A, Morales-Pastor A, Medel-Lacruz B, Pándy-Szekeres G, Mayol E, Giorgino T, Carlsson J, Deupi X, Filipek S, Filizola M, Gómez-Tamayo JC, Gonzalez A, Gutiérrez-de-Terán H, Jiménez-Rosés M, Jespers W, Kapla J, Khelashvili G, Kolb P, Latek D, Marti-Solano M, Matricon P, Matsoukas M, Miszta P, Olivella M, Perez-Benito L, Provasi D, Ríos S, R. Torrecillas I, Sallander J, Sztyler A, Vasile S, Weinstein H, Zachariae U, Hildebrand PW, De Fabritiis G, Sanz F, Gloriam DE, Cordomi A, Guixà-González R, Selent J. (2020). GPCRmd uncovers the dynamics of the 3D-GPCRome."},{"url":"https://doi.org/10.1002/pro.5172","text":"Cao F, von Bülow S, Tesei G, Lindorff‐Larsen K. (2024). A coarse‐grained model for disordered and multi‐domain proteins."},{"url":"https://doi.org/10.1038/nature06522","text":"Henzler-Wildman K, Kern D. (2007). Dynamic personalities of proteins."},{"url":"https://doi.org/10.1101/2025.06.30.662466","text":"Lyu N, Du S, Shao Q, Yang Z, Ma J, Herschlag D. (2025). Physics-Grounded Evaluation to Guide Accurate Biomolecular Prediction."},{"url":"https://doi.org/10.82153/jkxj-tw08","text":"Chrispens K, Collins M, Fraser JS, Mai D, van den Bedem H, Wankowicz SA. (2026). _sampleworks:_ A Modular Platform for Experimentally Guided Biomolecular Ensemble Generation."},{"url":"https://doi.org/10.48550/arxiv.2602.24007","text":"Maddipatla A, Rzayev A, Pegoraro M, Pacesa M, Schanda P, Marx A, Vedula S, Bronstein AM. (2026). Inference-time optimization for experiment-grounded protein ensemble generation."},{"url":"https://doi.org/10.48550/arxiv.2506.04490","text":"Raghu R, Levy A, Wetzstein G, Zhong ED. (2025). Multiscale guidance of protein structure prediction with heterogeneous cryo-EM data."},{"url":"https://doi.org/10.1038/s41592-026-03047-4","text":"Fadini A, Li M, McCoy AJ, Banjara S, Okumura H, Napier E, Fontana P, Khan AR, Jovine L, Terwilliger TC, Read RJ, Hekstra DR, AlQuraishi M. (2026). AlphaFold as a prior: experimental structure determination conditioned on a pretrained neural network."},{"url":"https://doi.org/10.1107/s2052252524005098","text":"Wankowicz SA, Fraser JS. (2024). Comprehensive encoding of conformational and compositional protein structural ensembles through the mmCIF data structure."}],"slug":"qfit-at-scale","social_posts_count":null,"social_posts_embed_src":null,"state":"PUBLISHED","subtitle":"An Ensemble Dataset of over 60,000 Structures","tags":["result"],"title":"Recovering Conformational Heterogeneity from the Protein Data Bank at Scale","version_desc":null,"version_number":1,"versions":[{"id":483,"version_number":1,"version_desc":null,"doi":"10.82153/fjj4-gv26","created_at":"2026-08-10T22:33:50.000Z"}]}}