Overlay databank unlocks data-driven analyses of biomolecules for all
Full text
This is a self-archived version of an original article. This version may differ from the original in pagination and typographic details. Author(s): Title: Year: Version: Copyright: Rights: Rights url: Please cite the original version: CC BY 4.0 https://creativecommons.org/licenses/by/4.0/ Overlay databank unlocks data-driven analyses of biomolecules for all © The Author(s) 2024 Published version Kiirikki, Anne M.; Antila, Hanne S.; Bort, Lara S.; Buslaev, Pavel; Favela-Rosales, Fernando; Ferreira, Tiago Mendes; Fuchs, Patrick F. J.; Garcia-Fandino, Rebeca; Gushchin, Ivan; Kav, Batuhan; Kučerka, Norbert; Kula, Patrik; Kurki, Milla; Kuzmin, Alexander; Lalitha, Anusha; Lolicato, Fabio; Madsen, Jesper J.; Miettinen, Markus S.; Mingham, Cedric; Monticelli, Luca; Nencini, Ricky; Nesterenko, Alexey M.; Piggot, Thomas J.; Piñeiro, Ángel; Reuter, Nathalie; Samantray, Suman; SuárezLestón, Fabián; Talandashti, Reza; Ollila, O. H. Samuli Kiirikki, A. M., Antila, H. S., Bort, L. S., Buslaev, P., Favela-Rosales, F., Ferreira, T. M., Fuchs, P. F. J., Garcia-Fandino, R., Gushchin, I., Kav, B., Kučerka, N., Kula, P., Kurki, M., Kuzmin, A., Lalitha, A., Lolicato, F., Madsen, J. J., Miettinen, M. S., Mingham, C., . . . Ollila, O. H. S. (2024). Overlay databank unlocks data-driven analyses of biomolecules for all. Nature Communications, 15, Article 1136. https://doi.org/10.1038/s41467-024-45189-z 2024
Article https://doi.org/10.1038/s41467-024-45189-z Overlay databank unlocks data-driven analyses of biomolecules for all Anne M. Kiirikki 1 ,HanneS.Antila 2,3 ,LaraS.Bort 2,4 , Pavel Buslaev 5 , Fernando Favela-Rosales 6 , Tiago Mendes Ferreira 7 , Patrick F. J. Fuchs 8,9 , Rebeca Garcia-Fandino 10 , Ivan Gushchin, Batuhan Kav 11,12 ,NorbertKučerka 13 , Patrik Kula 14 ,MillaKurki 15 ,AlexanderKuzmin ,AnushaLalitha 16 , Fabio Lolicato 17,18 ,JesperJ.Madsen 19,20 ,MarkusS.Miettinen 2,21,22 , Cedric Mingham 23 , Luca Monticelli 24,25 , Ricky Nencini 1,26 , Alexey M. Nesterenko 21,22 , Thomas J. Piggot 27 , Ángel Piñeiro 28 , Nathalie Reuter 21,22 ,SumanSamantray 11,29 ,FabiánSuárez-Lestón 10,28,30 , Reza Talandashti 21,22 &O.H.SamuliOllila 1,31 Tools based on artificial intelligence (AI) are currently revolutionising many fields, yet their applications are often limited by the lack of suitable training data in programmatically accessible format. Here we propose an effective solution to make data scattered in various locations and formats accessible for data-driven and machine learning applications using the overlay databank format. To demonstrate the practical relevance of such approach, we present the NMRlipids Databank—a community-driven, open-for-all database featuring programmatic access to quality-evaluated atom-resolution molecular dynamics simulations of cellular membranes. Cellular membrane lipid composition is implicated in diseases and controls major biological functions, but membranes are difficult to study experimentally due to their intrinsic disorder and complex phase behaviour. While MD simulations have been useful in understanding membrane systems, they require significant computational resources and often suffer from inaccuracies in model parameters. Here, we demonstrate how programmable interface for flexible implementation of datadriven and machine learning applications, and rapid access to simulation data through a graphical user interface, unlock possibilities beyond current MD simulation and experimental studies to understand cellular membranes. The proposed overlay databank concept can be further applied to other biomolecules, as well as in other fields where similar barriers hinder the AI revolution. Tools based on artificial intelligence (AI) are currently revolutionising many fields, yet their success relies on training data that is available in programmatically accessible and standardised format1. For example in structural biology, Protein Data Bank2has enabled revolutionary tools that can predict protein structures with unprecedented accuracy3. However, the lack of smart guidelines and community-consensus bestpractices for data sharing are limiting the development of such databanks and AI applications in many fields1. This is the case in biomolecular modelling, where vast amount of data is already available for proteins, lipids, nucleic acids, and carbohydrates, but tools that enable applications of these data for data-driven applications are not yet available4. Here we propose a cost-effective solution to make data Received: 2 June 2023 Accepted: 17 January 2024 Check for updates A full list of affiliations appears at the end of the paper. e-mail: samuli.ollila@helsinki.fi Nature Communications | (2024) 15:1136 1 1234567890():,; 1234567890():,;
scattered in various locations and formats accessible for data-driven and machine learning (ML) applications: an overlay databank. We demonstrate the practical relevance of such approach for understanding cellular membranes by incorporating available lipid bilayer simulations into the NMRlipids Databank. Importantly, the basic principles of an overlay databank can be applied to simulation data of any other biomolecules (that are becoming available in increasing amounts4), as well as to any other field where the development of AIbased tools is limited by the lack of data access and community bestpractices—such as the assignment of NMR spectra5. Incentives to share data can be further accelerated by combining overlay databanks with an open collaboration approach (see ref. 6). Cellular membranes contain hundreds of different types of lipid molecules that regulate the membrane properties, morphology, and biological functions7–9. Membrane lipid composition is implicated in diseases, such as cancer and neurodegenerative disorders, and therapeutics that affect membrane compositions are emerging10. However, biomembranes are often difficult to study experimentally, because they are complex mixtures of proteins and lipids in disordered fluid state with complicated phase behaviour at biological conditions. Molecular dynamics (MD) simulations can be used to model biomembranes in detail, but the computational cost for simulating all possible biological membrane compositions would be formidable. For those reasons, data-driven and machine-learning-based models that predict biomembrane properties will benefitwiderangeoffields covering academia and industry, from cell membrane biology to lipid nanoparticle formulations. Here we present the NMRlipids Databank—acommunity-driven, open-for-all database featuring programmatic access to atomresolution MD simulations of lipid bilayers. To demonstrate its advantages over existing approaches, we build a ML model that predicts membrane properties from its lipid composition and show how informationonrarephenomenathatarebeyondthescopeofstandard MD simulation investigations can be gleaned from the Databank. In addition, we demonstrate the immediate relevance of NMRlipids Databank in extending the scope of MD simulations to new fields: Using a data-driven approach, we are able to analyse how anisotropic diffusion of water depends on membrane properties; this benefits understanding in magnetic resonance imaging (MRI)11 and pharmacokinetics12, where MD simulations have, until now, been rarely applied. Furthermore, the Databank performs automatic quality evaluation of membrane simulations, which facilitates the selection of best-performing models for each given application and accelerates the development of simulation parameters and methodology. Notably, the overlay-databank and open-collaboration approaches would facilitate the collecting of community-contributed data and providing programmatic access to them also in fields other than membrane simulations. For example, substantial biomolecular MD simulation data are already available but in scattered locations and formats4. The technical advances presented here will immediately benefit making these data programmatically accessible. Furthermore, the overlay databank configuration will benefit a broad range of fields where access to training data is a bottleneck for building AI-based tools. Results NMRlipids overlay databank delivers access to MD simulations of membranes composed of the biologically most abundant lipids NMRlipids Databank is a community-driven catalogue containing atomistic MD simulations of biologically relevant lipid membranes emerging from the NMRlipids open collaboration6,13–16. It has been designed to improve the Findability, Accessibility, Interoperability, and Reuse17 of MD simulation data, most importantly the output trajectories and necessary information to their reuse. The NMRlipids Databank is constructed using the NMRlipids project protocol, in which all the content is openly accessible throughout the project6. Currently, the NMRlipids Databank contains 765 simulation trajectories with the total length of approximately 0.4 ms. Singlecomponent lipid membranes and binary mixtures are currently most abundant in the NMRlipids Databank, yet mixtures with up to five lipid types are available. For available mixtures, see Fig. 1E. The distribution of lipids among the available simulations, shown in Fig. 1B, roughly resembles the biological relative abundance of different lipid types, with phosphatidylcholine (PC) being the most common followed by cholesterol, phosphatidylethanolamine (PE), phosphatidylserine (PS), phosphatidylglycerol (PG), phosphatidylinositol (PI), and other lipids, depending on organism and organelle7. Abbreviations and full names of all lipids present in the Databank are listed in the NMRlipids Databank documentation18.Forcefields used in simulations cover all the essential parameter sets commonly used in lipid simulations, see Fig. 1C and Supplementary Table 1, including also united atom and polarisable force fields. Therefore, the averages calculated over the Databank can be considered as mean predictions from available lipid models (average over force field parameters) for an average cell membrane (average over lipid compositions). The overlay structure of the NMRlipids Databank, illustrated in Fig. 1A, is designed to enable efficient upcycling of MD simulations for data-driven and ML applications with minimal investment on new infrastructure. Raw simulation data in the Data layer can be stored in any publicly available location with long term stability and with permanent links to the data, such as digital object identifiers (DOIs), such as Zenodo (zenodo.org). The Databank layer (github.com/NMRlipids/ Databank) is the core of the Databank containing all the relevant information about the simulations: links to the raw data, relevant metadata describing the systems, universal naming conventions for lipids and their atoms, quality evaluation of simulations against experimental data, and the computer programmes to create the entries and to analyse the five basic properties extracted from all simulations (area per lipid, C–H bond order parameters, X-ray scattering form factors, membrane thickness, and equilibration times of principal components). Also the values for these five basic properties are stored in the Databank layer. The Application layer is composed of repositories and tools that read information from the Databank layer for further analyses. Because the Application layer does not interfere with the Databank layer, it can befreely extended by anyone for a wide range of purposes. This is demonstrated here with two examples: the NMRlipids Databank graphical user interface (NMRlipids Databank-GUI) at databank. nmrlipids.fiand a repository exemplifying novel analyses utilising NMRlipids Databank as discussed below (github.com/NMRlipids/ DataBankManuscript). A more detailed description of the NMRlipids Databank structure is available in the Supplementary Information. NMRlipids Databank-GUI: graphical access to the MD simulation data NMRlipids Databank-GUI, available at databank.nmrlipids.fi,provides easy access to the NMRlipids Databank content through a graphical user interface (GUI). Simulations can be searched based on their molecular composition, force field, temperature, membrane properties, and quality; the search results are ranked based on the simulation quality as evaluated against experimental data when available. Membranes can be visualised, and properties between different simulations and experiments compared. The NMRlipids Databank-GUI enables rapid surveying of what simulation data is available, selection of the best available simulations for specific systems based on ranking lists, and comparisons of basic properties between different types of membranes. Notably, the GUI enables these operations to be performed by scientists with a wide range of backgrounds—including those who do not necessarily have programming expertise or other means to access MD simulation data. Article https://doi.org/10.1038/s41467-024-45189-z Nature Communications | (2024) 15:1136 2
NMRlipids Databank-API: programmatic access to the MD simulation data The NMRlipids Databank-API provides programmatic access to all simulation data in the NMRlipids Databank through application programming interface (API). This enables wide range of novel datadriven applications—from construction of ML models that predict membrane properties, to automatic analysis of virtually any property across all simulations in the Databank. The flowchart in Fig. 1D illustrates the practical implementation of such an analysis. After cloning the Databank repository to a local computer, raw data of each simulation can be accessed and analyzed with the help of functions delivered by the NMRlipids Databank-API. Analyses over simulations with different naming conventions can be automatically performed with the help of mapping files that associate the specific naming conventions in each simulation with the universal molecule and atom names used by the Databank. Finally, the analysis results can be stored using the same structure as in the Databank layer. Documentation of the NMRlipids Databank-API and template for new user-defined analyses with further instructions are available at the NMRlipids Databank documentation19. While analysis codes and results for basic membrane properties are included in the Databank layer, unlimited further analyses canbeimplementedbyanyoneinseparaterepositoriesinthe Application layer.WhenApplication layer repositories are organised by mimicking the Databank layer structure, they can be accessed programmatically and further analyzed using the tools in the NMRlipids Databank-API by implementing the flowchart demonstrated in Supplementary Fig. 2. Novel analyses that demonstrate the power of NMRlipids Databank in selecting the best simulation models, analysing rare phenomena, and extending MD simulations to new fields are implemented in an Application-layer repository located at github.com/NMRlipids/ DataBankManuscript. The related codes are listed in Supplementary Table 2. Fig. 1 | Overview of the NMRlipids Databank. A Schematic presentation of the overlay structure used in the NMRlipids Databank. A more detailed structure of the Databank layer is shown in Supplementary Fig. 1. BDistribution of lipids present in the trajectories of the Databank. ‘Others’lists lipids occurring in six or fewer simulations. CDistribution of force fields in the simulations in the Databank. References for each force field are given in Supplementary Table 1. DFlowchart for performing an analysis of properties through all MD simulations in the NMRlipids Databank using the API. ECurrently available lipid mixtures in the NMRlipids Databank. Colorbar shows the number of available simulations with the darkest green indicating three or more. Article https://doi.org/10.1038/s41467-024-45189-z Nature Communications | (2024) 15:1136 3
Selecting simulation parameters using NMRlipids Databank: Best models for most abundant neutral membrane lipids MD simulations have been particularly useful in understanding membrane systems, although their accuracy has often been compromised by artefacts such as the quality of model parameters20,21.Presently,the accuracy of models is becoming increasingly important as researches are progressing from simulations of individualmolecules to simulating whole organelles or even cells using interdisciplinary approaches21–23. Such systems exhibit intricate emergent behaviour making inaccuracies more difficult to detect, and accumulation of even modest errors may have a dramatic impact on the conclusions drawn. To minimise the detrimental consequences of artificial MD simulation results for their applications, the quality of lipid bilayer MD simulations has to be carefully assessed20. This can be done, for example, against the C–H bond order parameters from NMR spectroscopy16,24 and the form factors from X-ray scattering13, although it requires comparisons between large number of simulations, which is laborious even with collaborative approaches6,14–16. Here we streamlined this process by defining quantitative quality measures for conformational ensembles of individual lipid molecules and membrane dimensions using C–H bond order parameters from NMR and X-ray scattering form factors13. These measures enable automatic ranking of lipid bilayer simulations based on their quality against experiments. Qualities of order parameters were evaluated by first calculating the probabilities for each C–H bond order parameter to locate within experimental error, and then averaging the possibilities over different lipid segments (Phg,Psn1,Psn2,andPtotal). Qualities against X-ray scattering experiments (FF q ) were estimated as the difference in the experimental and simulated locations of the first form factor minimum. These measures are good proxies for membrane properties because they correlate with the membrane lateral packing and thickness (Fig. 2G, Supplementary Figs. 3 and 4). Ergodicity of conformational sampling of lipids was estimated by calculating τ rel ,the convergence time of the slowest principal component divided by the simulation length. Figure 2demonstrates how the automatic simulation-quality evaluation and the NMRlipids Databank-API enable rapid selection of the best models for membrane simulations. Figure 2A illustrates that predictions for the lateral packing of membranes composed of two most biologically-abundant neutral membrane lipids, POPC and POPE7, diverge between different force fields. To find the most realistic parameters to simulate membranes with these lipids, we first ranked all simulations based on order parameter quality (Supplementary Fig. 5), then picked force fields that occur in Fig. 2A (that is: force fields for both POPC and POPE in the Databank), and then ranked them according to the quality of the sn-1 chain of POPC (Fig. 2B) and of POPE (Fig. 2C). Simulations with τ rel clearly above one (larger than 1.3) were discarded in this analysis. Because the average sn-1 chain order parameter is a good proxy for themembranepacking(Fig.2G), rankings in Fig. 2BandCcanbe used to select the simulations giving the most realistic results in Fig. 2A. Based on this, Lipid17 and Slipids simulations are most realistic for a POPC membrane, while CHARMM36 and GROMOSCKP simulations predict overly packed bilayers (overestimated order in Supplementary Fig. 6). For POPE, on the other hand, GROMOS-CKP and Slipids are most realistic, while CHARMM36 and Lipid17 predict too packed membranes. In conclusion, the quality evaluation based on the NMRlipids Databank suggests that the Slipids parameters are the best currently available choice for simulations with PC and PE lipids, at least for applications where membrane packing is relevant. Also direct comparisons with the experimental data for the most relevant simulations are shown in Fig. 2D–F and Supplementary Fig. 6A. Figure 2D shows the overall highest-ranked simulation, POPC bilayer with OPLS3e parameters, for the reference. Using NMRlipids Databank as a training set for machine learning applications: Predicting multi-component membrane properties To demonstrate the usage of the NMRlipids Databank to construct ML models that predict membrane properties, we trained a model that predicts area per lipids and thicknesses of membranes with diverse compositions. First we used randomly selected 80% of the Databank simulations to choose the hyperparameters for, and optimise, a set of ML models with the goal of predicting the area per lipid from the membrane composition. After the parameter optimisation, we tested predictions from different ML models for the area per lipid both against the remaining 20% of the Databank and against areas per lipid reported from simulations of membranes containing mixtures of POPC, POPE, POPS, PI, sphingomyelin lipids, and cholesterol25–27 that are not included in the Databank. Essential differences between models were not observed when predicting area per lipids of 20% of simulations selected as the test set, but linear regression and Ridge models gave the best correlations with the literature data (predictions from the linear regression modelareshowninFig.3and from other models in Supplementary Fig. 7). The linear regression model was selected for further studies due to its simplicity. To demonstrate the usefulness of the constructed models for understanding multi-component membrane properties, we predicted later packing (areas per lipid) and membrane thicknesses of common biological membranes based on their lipid compositions reported in the literature, see Table 1. The model predicts substantial 50% difference in area per lipid between most densely (influenza virus) and loosely (mitochondria) packed membranes. Difference of 0.8 nm in thickness is predicted between the thinnest (bacterial) and thickest (plasma) membranes. Such differences are expected to effect on many biologically relevant functions of membranes, such as permeation, cholesterol flip-flops (see next sections), and interactions with proteins28, demonstrating that NMRlipids Databank can be used to give valuable insights on biologically relevant properties of complex biological membranes. Most importantly, the delivered programmatic access to increasing amount of MD simulation data enables training of ML models that predict various membrane properties for all. Detecting rare phenomena using NMRlipids databank: cholesterol flip-flops Lipid flip-flops from one bilayer leaflet to another play an important role in lipid trafficking and regulating membrane properties7.Phospholipid flip-flop events are rare when not facilitated by proteins, occurring spontaneously on the timescale of hours or days, while cholesterol, diacylglycerol, and ceramide flip-flop much more often. Still, the reported timescales range from minutes to submillisecods7,29–31. These timescales were previously accessible only by coarse-grained simulations or free energy calculations30,andatomistic simulations reporting cholesterol flip-flop events have been published only recently31–33. The atomistic studies report an increase in cholesterol flip-flop rates with increasing acyl chain unsaturation level and decreasing cholesterol concentration31,32,buttheamountofdatain these individual studies was not sufficient to systematically assess correlations between cholesterol flip-flop rates and membrane properties. Here, we demonstrate that the NMRlipids Databank-API makes analyses of such rare phenomena accessible for all by enabling access to a large amount of MD simulation data as illustrated in Fig. 1.Thisis particularly useful for scientists in various fields of science and industry who lack access to the computational resources or the expertise to produce the large amounts of MD simulation data required for such analyses. Using the general workflow depicted in Fig. 1D, we first calculated the flip-flop rates from all the simulations available in the NMRlipids Databank. Flip-flops were observed for cholesterol, DCHOL Article https://doi.org/10.1038/s41467-024-45189-z Nature Communications | (2024) 15:1136 4
(18,19-di-nor-cholesterol), DOG (1,2-dioleoyl-sn-glycerol), and SDG (1stearoyl-2-docosahexaenoyl-sn-glycerol). The observed cholesterol flip-flop rates, ranging between 0.001–1.6 μs−1with the mean of 0.16 μs−1and median of 0.07 μs−1, are in line with the previously reported values from atomistic MD simulations31–33.Theflip-flop rate of DCHOL, 0.2 μs−1, was close to the average value of cholesterol, while the average rates for diacylglycerols DOG (0.4 μs−1)andSDG(0.5μs−1) were higher than for cholesterol. Flip-flops were not observed for other lipids, giving the upper limits for PC-lipid flip-flop rate as 9×10 −6μs−1and for ceramide (N-palmitoyl-D-erythro-sphingosine) as 0.002 μs−1. Thus, the available data in the NMRlipids Databank suggest that the lipid flip-flop rate decreases in the order: diacylglycerols > cholesterol > other lipids including ceramides. However, the amount of data for diacylgycerols (8 simulations with the Lipid17 force field) and ceramide (3 simulations with CHARMM36) is less than that for cholesterol (83 simulations); thus we cannot fully exclude the effect of force field or composition on this comparison. Nevertheless, we used the general workflow depicted in Supplementary Fig. 2 to analyse how the flip-flop rates calculated from the NMRlipids Databank depend on membrane properties. Figure 4B–D show cholesterol flip-flopratesandtheirhistogramsasafunctionof membrane thickness, lateral density, and acyl chain order. The results Fig. 2 | Examples of data obtained using NMRlipids Databank. A Area per lipid of POPC and POPE lipid bilayers predicted by different force fields at 310 K in simulations that are available in the NMRlipids Databank. The data points from the bestperforming simulations, based on rankings in (B,C), are surrounded by black circles. BBest POPC simulations ranked based on the sn-1 acyl chain order parameter quality (Psn1). Also sn-2 acyl chain (Psn2), headgroup (Phg)andtotal(Ptotal)order parameter qualities, form factor quality (FF q ), and relative equilibration time for conformations (τ rel ) are shown. Note that the best possible order parameter quality is one, while the best possible form factor quality is zero. CBest POPE simulations ranked based on the sn-1 acyl chain order parameter quality. Direct comparison against experimental (NMR order parameters and X-ray scattering) data exemplified for a simulation with the best overall order parameter quality (D), the best quality for POPE lipid (E), and the headgroup quality for POPE (F). Error bars for simulations are standard error of the mean over different lipids (n= number of lipids in a simulation shown in B). Error bars for experiments are 0.0213.GScatter plots and Pearson correlation coefficients, r, for the membrane area per lipid, thickness, first minimum of X-ray scattering form factor and average order parameter of the sn-1 acyl chain extracted from the NMRlipids Databank. All correlation coefficients have p-value below 0.001 with two-sided test. For more correlations see Supplementary Fig. 3. Article https://doi.org/10.1038/s41467-024-45189-z Nature Communications | (2024) 15:1136 5
reveal a non-linear correlation between cholesterol flip-flop rate and membrane packing (depicted as area per lipid): Flip-flop rates increase by an order of magnitude when membrane packing density decreases, and a major jump is observed at low membrane packing. Such orderof-magnitude changes in cholesterol flip-flop rate with the membrane composition may have major implications in understanding lipid trafficking and membrane biochemistry31,33.Becausetheresultsfrom the NMRlipids Databank are averaged over a large range of membrane compositions and force fields, they show that the strong dependence of cholesterol flip-flop rate on membrane properties is not limited to the particular lipid compositions or force fields used in the previous studies31–33. Extending the scope of MD simulations to new fields using NMRlipids Databank: Water diffusion anisotropy in membrane systems The anisotropic diffusion of water and hydrophilic molecules in directions parallel and perpendicular to membranes is an important parameter in models describing the translocation of drugs through biological material, particularly in the skin12,34–36. Water anisotropic diffusion plays a role also in the signal formation in diffusion-tensor MRI imaging11. MD simulations are rarely used to analyze the anisotropic diffusion of water, since only a few membrane permeation events of water are typically observed in a single MD simulation trajectory37,38, thereby making the collection of a sufficient amount of data challenging. Here, we show that the API access to the data in NMRlipids Databank enables systematic analysis on how the anisotropic diffusion of water depends on membrane properties in multilamellar membrane systems, thereby extending the application of MD simulations to new fields. To this end, we first calculated the water permeability through membranes from all simulations in the NMRlipids Databank using the general workflow depicted in Fig. 1D. The resulting non-zero values range between 0.3 and 322 μm/s with the mean of 14 μm/s and median of 8 μm/s. These values agree with the previously reported simulation results37,38, but are on average larger than experimental values reported for PC lipids in the liquid crystalline phase, 0.19–0.33 μm/s39.Using the workflow depicted in Supplementary Fig. 2, we then plotted the observed permeabilities and their histogrammed values in Fig. 5B–Eas a function of temperature, membrane thickness, area per lipid, and acyl chain order. As expected, the permeability increases with the temperature, giving an average energy barrier of 14 ± 3 k B Tfor the water permeation from the Arrhenius plot in Fig. 5B. On the other hand, the water permeability on average decreases when membranes become more packed, that is, with decreasing area per lipid and increasing thickness and acyl chain order (Fig. 5C–E). Permeation of water through bilayers depends on membrane properties also according to previous studies, but there is no established consensus on whether the area per lipid40 or bilayer thickness41 is the main parameter determining the permeability. Our analysis over the NMRlipids Databank, containing significantly more data than what was available in previous studies, suggest non-linear dependencies on both of these parameters. Clear dependencies of permeability on hydration level or the fraction of charged lipids, cholesterol, or POPE in the membrane were not observed (Supplementary Fig. 8). To examine how water diffusion anisotropy depends on membrane properties in a multi-lamellar lipid bilayer system, we analyzed the water diffusion parallel to the membrane surface from all simulations in the NMRlipids Databank using the general workflows depicted in Fig. 1D and Supplementary Fig. 2.The parallel diffusion coefficient of water, D ∥ , decreases with reduced hydration and increases with the temperature, but dependencies on the membrane area per lipid, thickness, or fraction of charged lipids were not observed in Fig. 5and Supplementary Fig. 9. Simulation results are close to the experimental values with low hydration levels in Fig. 5F, but increase to approximately 50% higher than the experimental value for bulkwater diffusion value (3.1 × 10−9m2/s at 313 K42) with high hydration levels. This is not surprising as the most common water model used in membrane simulations, TIP3P, overestimates the bulk water diffusion43.Toestimate the diffusion anisotropy of water, D ⊥ /D ∥ , in multilamellar membrane system, the permeability coefficients of water through membranes were translated to perpendicular diffusion coefficients, D ⊥ , using the Tanner equation44,45. The resulting perpendicular diffusion coefficients are approximately five orders of magnitude smaller than the lateral diffusion coefficients of water (Fig. 5G, H), which is at the upper limit of anisotropy estimated from experimental data12.A significant increase in the diffusion anisotropy with membrane packing is observed, as D ⊥ /D ∥ deviates further from unity with decreasing area per lipid and increasing thickness in Fig. 5G, H. This follows from Table 1 | Areas per lipid (APL) and membrane thicknesses predicted by the linear regression model trained using the NMRlipids Databank for membrane compositions corresponding different biological membranes Membrane Composition [lipid(%-fraction)] APL[Å2] Thickness [nm2] Mitochondria POPC(37) : POPE(31) : PI(6) : CL(22) : CHOL(4)7,70,71 74 (66)a4.4 Bacterial POPC(20) : POPE(35) : POPG(35) : CL(5) : CHOL(5)72 61 (60)a4.2 ER POPC(54) : POPE(20) : POPS(4) : PI(11) : SM(4) : CHOL(8)7,70,71 57 4.6 Golgi POPC(36) : POPE(21) : POPS(6) : PI(12) : SM(7) : CHOL(18)7,70,71 52 4.8 Plasma POPC(23) : POPE(11) : POPS(8) : PI(7) : SM(17) : CHOL(34)7,70,71 45 5.0 Synaptic POPC(28) : POPE(20) : POPS(9) : PI(2) : SM(4) : CHOL(37)73 45 4.8 Influenza POPC(5) : POPE(32) : POPS(15) : SM(9) : CHOL(40)74,75 42 4.8 Systems are sorted in the order of increasing membrane packing. Compositions of different membranes are estimated based on references given in the composition column. aValue in parenthesis is the area per acyl chain taking into account that cardiolipin has four chains per molecule while other lipids have two. Fig. 3 | Predictions of areas per lipid (APL) of multi-component membranes composed of POPC, POPE, POPS, PI, sphingomyelin lipids, and cholesterol from linear regression model against literature data from simulations (green25, blue26,andred 27). Error bars are from the same publications as the values. Black line indicates x=y. Article https://doi.org/10.1038/s41467-024-45189-z Nature Communications | (2024) 15:1136 6
Fig. 5 | Quantification of water diffusion in NMRlipids Databank simulations. AWater diffusion, D ⊥ , and permeability, P, through membranes, and lateral diffusion along the membrane, D ∥ , illustrated in a multilamellar stack of lipid bilayers. B–EWater permeation through membranes analyzed from the Databank as a function of temperature, thickness, area per lipid, and acyl chain order. Inset in (B) shows the Arrhenius plot of permeation (lnðPÞvs. 1/T)thatgives14±3k B Tfor the average activation energy for water permeation through lipid bilayer. FLateral diffusion of water as a function of hydration level. Experimental points for DMPC bilayers at 313 K at different hydration levels are shown76.G,HDiffusion anisotropy of water as a function of thickness and area per lipid. Non-zero permeation and diffusion values from simulations are shown with blue dots. Histogrammed values are shown with black dots. For the mean value in each bin, average weighted with the simulation lengths was used, and error bars show the standard error of the mean. Only bins with more than one microsecond of data in total were used for water permeation. Fig. 4 | Quantification of cholesterol flip-flop events in NMRlipids Databank simulations. A Illustration of cholesterol flip-flop. B–DCholesterol flip-flops analyzed from the Databank as a function of membrane thickness, area per lipid, and acyl chain order. Values from simulations with non-zero flip-flop rates are shown with blue dots.Histogrammed values are shown with black dots. For the mean value in each bin, average weighted with the simulation lengths was used, and error bars show the standard error of the mean. Article https://doi.org/10.1038/s41467-024-45189-z Nature Communications | (2024) 15:1136 7
decreasing water permeability with membrane packing (Fig. 5C, D), while lateral diffusion remains approximately constant (Supplementary Fig. 9A, C). In summary, our results suggest that the bilayer packing has a substantial effect on anisotropic water diffusion in multi-membrane lipid systems. The several-fold larger anisotropy in membranes with higher lateral density is expected to play a role in pharmacokinetic models not only for water but also for other hydrophilic molecules12. Furthermore, the enhanced understanding of this anisotropy may help in developing new diffusion-tensor-based MRI imaging methods where signals originate from the anisotropic diffusion of water in biological matter11. Discussion Sharing of biomolecular MD simulation and other data is becoming increasingly important in the age of big data and AI1,4. Besides the data itself, also programmatic access is a necessary requirement for datadriven and ML applications. This is particularly challenging when fieldspecific smart guidelines and community-consensus best-practices have not yet been defined, which is the case in biomolecular simulations4. The NMRlipids Databank demonstrates how these issues can be solved by the overlay databank design, where the raw data are distributed to already publicly available decentralised locations, while the core of the databank is composed only of the metadata stored in a version-controlled git repository with an open-access license. On the other hand, the open-collaboration approach developed in the NMRlipids Project6creates incentives for sharing the data by offering authorship in published articles to the contributors. Advantages of such an approach are demonstrated here for membrane simulations, yet the concept can be applied to any other biomolecules, as well as in any other field where similar barriers hinder the AI revolution, such as the assignment of NMR spectra5. The NMRlipids Databank-API delivers programmatic access to MD simulation data that can be used as training set for diverse data-driven and ML applications that predict membrane properties. Such applications could be analogous to AlphaFold3and other tools46,47 that predict protein structures from their sequence using AI. This is demonstrated here by building ML models to predict multicomponent membrane properties. Furthermore, the analysis of cholesterol flip-flop events (Fig. 4) and water permeation through membranes (Fig. 5) demonstrate how a large amount of accessible simulation data in terms of quantity (e.g., simulation length and number of conformations) and content (e.g., lipid compositions and ion concentrations) enable analyses of rare phenomena that are beyond the current possibilities for a single research group. Such analyses also pave the way for applications of MD simulations in new fields, as demonstrated here by analysing an essential parameter in pharmacokinetic modelling and MRI imaging:11,12 the anisotropic diffusion of water in membrane systems (Fig. 5). These possibilities are particularly valuable for scientists who do not typically have access to large-scale MD simulation data. The focus of biomolecular simulations is moving from studies of individual molecules to larger complexes and even whole cells and organelles21–23. Simultaneously, machine-learning-based models for predicting the behaviour of biomolecules and automatic approaches to parametrise models are emerging3,20. The resources delivered by the NMRlipids Databank will support developments in both of these directions. Automatic quality evaluation and ranking of simulations against experimental data enable the selection of best simulations for specific applications without laborious manual force field evaluation. This also streamlines automatic parametrization procedures for atomistic and coarse grained simulations by, for example, pinpointing typical failures of force fields and highlighting points of improvement. Such practises for fostering the accuracy of simulations are becoming increasingly important as small errors accumulate when complexity and size of simulated systems are increasing. Examples of impact of NMRlipids and other overlay databanks in different disciplines are listed in Table 2, yet the scope of applications is expected to further widen with increasing amount of publicly shared data. Methods Structure of the databank The overlay structure designed for the NMRlipids Databank is composed of three layers (Fig. 1A). The Data layer contains raw data that can be distributed to publicly available servers such as Zenodo (zenodo.org). The core content of the Databank locates in the Databank layer, which is a git repository at github.com/NMRlipids/ Databank and is also permanently stored in a Zenodo repository (https://doi.org/10.5281/zenodo.7875567). The essential information of each simulation is stored in a human-and-machine-readable README.yaml file located in a subfolder of the /Data/Simulations folder in the Databank layer repository; each subfolder has a unique name constructed based on a hash code of the trajectory and topology files of each simulation. The README.yaml files in these folders contain access to all information that is needed for further analysis of simulations, such as links to the rawdata and associations with the universal molecule and atom names. The content of these files is described in detail in Supplementary Table 3 and in the NMRlipids Databank documentation nmrlipids.github.io. Results from analyses of basic membrane properties (area per lipid, thickness, C–H bond order Table 2 | Examples of impact of NMRlipids and other overlay databanks in different disciplines Target group Outcomes Practical examples Computational and data scientists Databanks with programmatic access, training sets for machine learning applications, automatic optimisation of simulation models Programmatic access to all publicly available MD simulation trajectories, machine learning models predicting properties of complex biomolecular assemblies (Table 1), optimising force field and model parameters from atomistic to continuum scale Biomedical scientists Applications of biomolecular modelling in new fields, coupling omics to biomolecular structure and dynamics Anisotropic water diffusion for pharmacokinetics and MRI imaging applications (Fig. 5), properties of complex cellular structures based on composition analyses from omics Material scientists Predictions of complex biomolecular material properties from data-driven and machine learning models Optimising composition of lipid nanoparticle formulations for desired properties, predictions of bioinspired material properties Biophysicists Novel analyses and predictions for properties of biomolecular assemblies and their compositions Analyses of rare phenomena (for example, Lipid flip-flops and water permeationinFigs.4and 5), data-driven analyzes and machine learning models predicting correlations between properties and compositions of biomolecular assemblies (Figs. 2and 5) Students and teachers Graphical and programmatic access to biomolecular structure and dynamics Illustration of disorder and dynamics in biomolecules, example data for bioinformatics and data-science courses Reviewers and publishers Tool to facilitate FAIRness of data, transparency and reproducibility All publicly available data can be included in an overlay databank irrespectively of its location Article https://doi.org/10.1038/s41467-024-45189-z Nature Communications | (2024) 15:1136 8
Zacatecas, Mexico. 7 NMR group - Institute for Physics, Martin Luther University Halle-Wittenberg, 06120 Halle (Saale), Germany. 8 Sorbonne Université, Ecole Normale Supérieure, PSL University, CNRS, LaboratoiredesBiomolécules(LBM), F-75005 Paris, France. 9 Université Paris Cité, F-75006 Paris, France. 10 Center for Research in Biological Chemistry and Molecular Materials (CiQUS), Universidade de Santiago de Compostela, E-15782 Santiago de Compostela, Spain. 11 Institute of Biological Information Processing: Structural Biochemistry (IBI-7), Forschungszentrum Jülich, 52428 Jülich, Germany. 12 ariadne.ai GmbH (Germany), Häusserstraße 3, 69115 Heidelberg, Germany. 13 Department of Physical Chemistry of Drugs, Faculty of Pharmacy, Comenius University Bratislava, 832 32 Bratislava, Slovakia. 14 Institute of Organic Chemistry and Biochemistry of the Czech Academy of Sciences, Flemingovo nám. 542/2, CZ-16610 Prague, Czech Republic. 15 School of Pharmacy, University of Eastern Finland, 70211 Kuopio, Finland. 16 Institut Charles Gerhardt Montpellier (UMR CNRS 5253), Université Montpellier, Place Eugène Bataillon, 34095 Montpellier, Cedex 05, France. 17 Heidelberg University Biochemistry Center, 69120 Heidelberg, Germany. 18 Department of Physics, University of Helsinki, FI-00014 Helsinki, Finland. 19 Department of Molecular Medicine, Morsani College of Medicine, University of South Florida, 33612 Tampa, FL, USA. 20 Center for Global Health and Infectious Diseases Research, Global and Planetary Health, College of Public Health, University of South Florida, 33612 Tampa, FL, USA. 21 Department of Chemistry, University of Bergen, 5007 Bergen, Norway. 22 Department of Informatics, Computational Biology Unit, University of Bergen, 5008 Bergen, Norway. 23 Hochschule Mannheim, University of Applied Sciences, 68163 Mannheim, Germany. 24 University of Lyon, CNRS, Molecular Microbiology and Structural Biochemistry (MMSB, UMR 5086), F-69007 Lyon, France. 25 Institut National de la Santé et de la Recherche Médicale (INSERM), Lyon, France. 26 Division of Pharmaceutical Biosciences, Faculty of Pharmacy, University of Helsinki, 00014 Helsinki, Finland. 27 Chemistry, University of Southampton, Highfield SO17 1BJ Southampton, UK. 28 Department of Applied Physics, Faculty of Physics, University of Santiago de Compostela, E-15782 Santiago de Compostela, Spain. 29 Institute of Biotechnology, RWTH Aachen University, Worringerweg 3, 52074 Aachen, Germany. 30 MD.USE Innovations S.L., Edificio Emprendia, 15782 Santiago de Compostela, Spain. 31 VTT Technical Research Centre of Finland, Espoo, Finland. e-mail: samuli.ollila@helsinki.fi Article https://doi.org/10.1038/s41467-024-45189-z Nature Communications | (2024) 15:1136 15