Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,773
datasets available to search
ShareScore release 0.7.1
Dataset results
1,773 results for “predictive modeling”
Code and data for Bayesian joint species distribution model selection for community-level prediction
<p>Code and data for reproducing the analysis in the manuscript "Bayesian joint species distribution model selection for community-level prediction." Provided data include percent cover observations for 39 modeled vascular plant species within boreal forest understory communities and environmental model covariates. R code is provided to generate model inputs, apply alternative models, generate out-of-sample predictions, and calculate associated community and species log scores and alternative model evaluation metrics. Further, R source code is provided to implement the multinomial joint species distribution model defined in the manuscript. Details on the data, its processing, and the alternative model definitions and structure can be found in the main text of the manuscript. Provided data are currently being used in ongoing analyses and coordination with authors may be warranted to avoid duplicate publication. Potential users are encouraged to consider collaboration with authors when useful and appropriate. Misinterpretation of data may occur if used outside the context of the original analysis. All data are made available in their current state. While significant efforts have been made to ensure data accuracy, complete accuracy cannot be guaranteed. Data may be updated periodically. It is the responsibility of the data user to check for updated versions of the data.</p>
Data from: Integrated species distribution models to account for sampling biases and improve range wide occurrence predictions
<p><strong><span>Aim</span></strong></p> <p><span>Species distribution models (SDMs) that integrate presence-only and presence-absence data offer a promising avenue to improve information on species' geographic distributions. The use of such 'integrated SDMs' on a species range-wide extent has been constrained by the often-limited presence-absence data and by the heterogeneous sampling of the presence-only data. Here, we evaluate integrated SDMs for studying species ranges with a novel expert range map-based evaluation. We build a new understanding about how integrated SDMs address issues of estimation accuracy and data deficiency and thereby offer advantages over traditional SDMs.</span></p> <p><strong><span>Location</span></strong></p> <p><span>South and Central America.</span></p> <p><strong><span>Time period</span></strong></p> <p><span>1979-2017.</span></p> <p><strong><span>Major taxa studied</span></strong></p> <p><span>Hummingbirds.</span></p> <p><strong><span>Methods</span></strong></p> <p><span>We build integrated SDMs by linking two observation models – one for each data type – to the same underlying spatial process.</span> <span>We validate SDMs with two schemes: i) cross-validation with presence-absence data and ii) comparison with respect to the species' whole range as defined with IUCN range maps. We also compare models relative to the estimated response curves and compute the association between the benefit of the data integration and the number of presence records in each data set.</span></p> <p><strong><span>Results</span></strong></p> <p><span>The integrated SDM accounting for the spatially varying sampling intensity of the presence-only data was one of the top-performing models in both model validation schemes. Presence-only data alleviated overly large niche estimates, and data integration was beneficial compared to modelling solely presence-only data for species that had few presence points when predicting the species' whole range. On the community level, integrated models improved the species richness prediction.</span></p> <p><strong><span>Main conclusions</span></strong></p> <p><span>Integrated SDMs combining presence-only and presence-absence data are successfully able to borrow strengths from both data types and offer improved predictions of species' ranges. Integrated SDMs can potentially alleviate the impacts of taxonomically and geographically uneven sampling and to leverage the detailed sampling information in presence-absence data.</span></p>
Modelling the carbon balance in bryophytes and lichens: Presentation of PoiCarb 1.0, a new model for explaining distribution patterns and predicting climate-change effects
<p><strong>Premise </strong></p> <p>Bryophytes and lichens have important functional roles in many ecosystems. Insight into how their CO<sub>2</sub> exchange responds to climatic conditions is essential for understanding current and predicting future productivity and biomass patterns, but responses are hard to quantify at time-scales beyond instantaneous measurements. We present PoiCarb 1.0, a model to study how CO<sub>2</sub> exchange rates of these poikilohydric organisms change through time as a function of weather conditions.</p> <p><strong>Methods</strong></p> <p>PoiCarb simulates diel fluctuations of CO<sub>2</sub> exchange and estimates long-term carbon balances, identifying optimal and limiting climatic patterns. Modelled processes are net photosynthesis, dark respiration, evaporation and water uptake. Measured CO<sub>2</sub>-exchange responses to light, temperature, atmospheric CO<sub>2</sub> concentration, and thallus water content (calculated in a separate module) are used to parameterise the model's carbon module. We validated the model by comparing modelled diel courses of net CO<sub>2</sub> exchange to such courses from field measurements on the tropical lichen <em>Crocodia aurata</em>. To demonstrate the model's usefulness, we simulated potential climate-change effects.</p> <p><strong>Results </strong></p> <p>Diel patterns were reproduced well and modelled and observed diel carbon balances were strongly positively correlated. Simulated warming effects via changes in metabolic rates were consistently negative, while effects via faster drying were variable, depending on the timing of hydration.</p> <p><strong>Conclusions</strong></p> <p>Being able to reproduce the weather-dependent variation in diel carbon balances is a clear improvement compared to simple extrapolations of short-term measurements or potential photosynthetic rates. Apart from predicting climate-change effects, future uses of PoiCarb include testing hypotheses about distribution patterns of poikilohydric organisms and guiding species' conservation.</p>
Data for: Implementing detailed nucleation predictions in the Earth system model EC-Earth3.3.4: sulfuric acid-ammonia nucleation
<p>Model dataset variables produced from the IFS and TM5 modules in EC-Earth3 version 3.3.4. which contains the control case and three experiments with the NPF lookup table. This paper is published at EGUshpere by journal: Geoscientific Model Development.</p> <p>The files contain:</p> <p>Compressed tar file of NetCDF data from IFS output for all four simulations. All IFS data have been averaged to monthly means from 6-hourly grib datasets. The post-process bash script which contains the function for the CDN and cloud effective radius weighted average towards cloud_time is found in the supplemented zendo link.</p> <p>NetCDF files from TM5 general output for each simulation. </p>
Task-driven neural network models predict neural dynamics of proprioception: Synthetic muscle spindle datasets
<p>#############</p> <p>Task-driven neural network models predict neural dynamics of proprioception, Cell 2024</p> <p>#############</p> <p>Authors: Marin Vargas, Alessandro (orcid=0000-0001-7073-4120) and Bisi, Axel (orcid=0009-0006-8602-7555) and Chiappa, Alberto Silvio (orcid=0009-0001-2764-6552) and Versteeg, Christopher (orcid=0000-0002-4269-5109) and Miller, Lee E. (orcid=0000-0001-8675-7140) and Mathis, Alexander (orcid=0000-0002-3777-2202)</p> <p>Affiliation: EPFL</p> <p>Date: January, 2024</p> <p>Link to the Cell article: </p> <p><a href="https://www.cell.com/cell/pdf/S0092-8674(24)00239-3.pdf">https://www.cell.com/cell/pdf/S0092-8674(24)00239-3.pdf</a></p> <p>--------------------------------</p> <p>Here we provide the synthetic spindle datasets of our article "Task-driven neural network models predict neural dynamics of proprioception". It contains the synthetic generated training dataset of simulated muscle spindles during arm passive movements generated with either character writing (PCR) or with 3D target reaching using reinforcement learning (RL).</p> <p>The overall structure of the data is:</p> <p>└── spindle_datasets<br> ├── pcr_dataset - Contains PCR synthetic training dataset<br> └── rl_dataset - Contains RL-generated synthetic training dataset</p> <p>The code to generate the PCR synthetic spindle dataset is available at: <a href="https://github.com/amathislab/Task-driven-Proprioception/tree/master/PCR-data-generation">https://github.com/amathislab/Task-driven-Proprioception/tree/master/PCR-data-generation</a></p> <p>The code to generate the RL-generated synthetic spindle dataset is available at: <a href="https://github.com/amathislab/Task-driven-Proprioception/tree/master/RL-data-generation">https://github.com/amathislab/Task-driven-Proprioception/tree/master/RL-data-generation</a></p> <p>--------------------------------</p> <p>The datasets, weights, activations and predictions are released with Creative Commons Attribution 4.0 license.</p> <p>The code is released under the MIT license, see <a href="https://github.com/amathislab/Task-driven-Proprioception">https://github.com/amathislab/Task-driven-Proprioception</a></p> <p>If you find our code, weights, predictions or ideas useful, please cite:</p> <p>@article{vargas2024task,<br> title={Task-driven neural network models predict neural dynamics of proprioception},<br> author={{Marin Vargas}, Alessandro and Bisi, Axel and Chiappa, Alberto S and Versteeg, Chris and Miller, Lee E and Mathis, Alexander},<br> journal={Cell},<br> year={2024},<br> publisher={Elsevier}<br>}</p>
Task-driven neural network models predict neural dynamics of proprioception: Neural network model weights
<p>#############</p> <p>Task-driven neural network models predict neural dynamics of proprioception, Cell 2024</p> <p>#############</p> <p>Authors: Marin Vargas, Alessandro (orcid=0000-0001-7073-4120) and Bisi, Axel (orcid=0009-0006-8602-7555) and Chiappa, Alberto Silvio (orcid=0009-0001-2764-6552) and Versteeg, Christopher (orcid=0000-0002-4269-5109) and Miller, Lee E. (orcid=0000-0001-8675-7140) and Mathis, Alexander (orcid=0000-0002-3777-2202)</p> <p>Affiliation: EPFL</p> <p>Date: January, 2024</p> <p>Link to the Cell article: </p> <p><a href="https://www.cell.com/cell/pdf/S0092-8674(24)00239-3.pdf">https://www.cell.com/cell/pdf/S0092-8674(24)00239-3.pdf</a></p> <p>--------------------------------</p> <p>Here we provide the trained model checkpoints for all tasks of our article "Task-driven neural network models predict neural dynamics of proprioception". It contains 300 temporal convolutional networks (TCNs) and 50 LSTM models trained on 16 tasks as well as the untrained initialization. </p> <p>The overall structure of the data is:</p> <p>└── models<br> ├── deepdraw_models - Contains networks hyperparameters<br> │ ├── template_models - Contains the default parameters<br> │ ├── torque - Contains network hyperparameters for the torque task<br> │ └── ... <br> ├── experiment_*** - Contains checkpoint of trained and untrained models <br> ├── ... <br> └── ... </p> <p>--------------------------------</p> <p>The checkpoints are stored in experiment folders (experiment_***) that follow this scheme:<br>- Task: shallow exp id, deep TCNs exp id, LSTM id.</p> <p>Experiment IDs for each task:</p> <p>- Untrained: 15, 115, 45<br>- Classification: 4015, 5015, 4045</p> <p>- Torque: 8015, 8030, 8045</p> <p>- Regress joint pos: 17016, 17031, 17046<br>- Regress joint vel: 17216, 17231, 17246<br>- Regress joint pos & vel:: 17416, 17431, 17446<br>- Regress joint pos & vel & acc:: 20516, 20531, 20546</p> <p>- Regress hand pos: 4016, 5016, 4046<br>- Regress hand vel: 17316, 17331, 17346<br>- Regress hand pos & vel: 17516, 17531, 17546<br>- Regress hand pos & vel & acc: 20416, 17831, 17846</p> <p>- Regress hand and elbow pos: 20016, 20031, 20046<br>- Regress hand and elbow vel: 20916, 20931, 20946<br>- Regress hand and elbow pos & vel: 20616, 20631, 20646<br>- Regress hand and elbow pos & vel & acc: 20816, 20831, 20846</p> <p>- Redundancy reduction - task transfer (AR): 10020, 10035, 10050<br>- Redundancy reduction - task transfer (HP): 10021, 10036, 10051<br>- Autoencoder 20716 & 20717, 20731 & 20732, X</p> <p>The code to load, evaluate and train the models is available at: <a href="https://github.com/amathislab/Task-driven-Proprioception/tree/master/nn-training">https://github.com/amathislab/Task-driven-Proprioception/tree/master/nn-training</a></p> <p>--------------------------------</p> <p>The datasets, weights, activations and predictions are released with Creative Commons Attribution 4.0 license.</p> <p>The code is released under the MIT license, see <a href="https://github.com/amathislab/Task-driven-Proprioception">https://github.com/amathislab/Task-driven-Proprioception</a></p> <p>If you find our code, weights, predictions or ideas useful, please cite:</p> <p>@article{vargas2024task,<br> title={Task-driven neural network models predict neural dynamics of proprioception},<br> author={{Marin Vargas}, Alessandro and Bisi, Axel and Chiappa, Alberto S and Versteeg, Chris and Miller, Lee E and Mathis, Alexander},<br> journal={Cell},<br> year={2024},<br> publisher={Elsevier}<br>}</p>
Main model fits and substitution rate predictions for: A quantitative genetic model of background selection in humans
<p>Across the human genome, there are large-scale fluctuations in genetic diversity caused by the indirect effects of selection. This can be thought of as a "linked selection signal" that reflects the impact of selection varying according to the placement of functional regions and recombination rates along the genome. Previous work has shown that negative selection against the steady influx of new deleterious mutations into conserved regions is the predominant mode of selection in humans. However, the theoretic model that underpins these results, classic Background Selection theory, is only applicable when new mutations are so deleterious that they cannot fix in the population. Here, we develop a statistical method based on a quantitative genetics view of the linked selection, which models the effects of weak draft created according to how polygenic additive fitness variance is distributed along the genome. We use a recent model that jointly predicts the equilibrium fitness variance and substitution rates due to both strong and weakly deleterious mutations, we estimate the distribution of fitness effects (DFE) and mutation rate across three human populations. While our model can accommodate weaker selection, we initially find evidence across three human populations of very strong selection against deleterious mutations consistent with previous work. However, the corollary predicted substitution rates for conserved regions are unreasonably low, and in disagreement with observed rates. We hypothesize this could be due to selected sites experiencing a further diminished population size due to selective interference. When we account for this in our method, we find evidence of weakly deleterious mutations in conserved regions which brings the predicted substitution rate into agreement with observations. However, these models lead to implausibly large mutation rate estimates. Overall, while our model of the genomic linked selection signal brings us a step towards uniting population and quantitative genetic selection models with the substitution process, our work suggests considerable uncertainty remains about the processes generating fitness variance in humans.</p>
Task-driven neural network models predict neural dynamics of proprioception: Experimental data, activations and predictions of neural network models
<p>#############</p> <p>Task-driven neural network models predict neural dynamics of proprioception, Cell 2024</p> <p>#############</p> <p>Authors: Marin Vargas, Alessandro (orcid=0000-0001-7073-4120) and Bisi, Axel (orcid=0009-0006-8602-7555) and Chiappa, Alberto Silvio (orcid=0009-0001-2764-6552) and Versteeg, Christopher (orcid=0000-0002-4269-5109) and Miller, Lee E. (orcid=0000-0001-8675-7140) and Mathis, Alexander (orcid=0000-0002-3777-2202)</p> <p>Affiliation: EPFL</p> <p>Date: January, 2024</p> <p>Link to the Cell article:</p> <p><a href="https://www.cell.com/cell/pdf/S0092-8674(24)00239-3.pdf">https://www.cell.com/cell/pdf/S0092-8674(24)00239-3.pdf</a></p> <p>--------------------------------</p> <p>Here we provide the neural data, activation and predictions for the best models and result dataframes of our article "Task-driven neural network models predict neural dynamics of proprioception".</p> <p>It contains the behavioral and neural experimental data (cuneate nucleus and somatosensory recordings from the Miller Lab, Northwestern University), the result dataframes for task-driven and untrained models, the activations and predictions for the *best models for all tasks* for active and passive movements and the predictions for linear models for active and passive movements. </p> <p>Note, the predictions of other models can be computed from the network weights that were deposited for all trained models. </p> <p>The overall structure of the data is:</p> <p>└── exp_analysis<br> ├── results - Contains the result dataframe of the predictions for all models, tasks and primates<br> ├── activations<br> │ ├── active - Contains activations related to active movements<br> │ └── passive - Contains activations related to passive movements<br> ├── predictions<br> │ ├── active - Contains predictions related to active movements<br> │ └── passive - Contains predictions related to passive movements<br> └── beh_exp_datasets<br> ├── matlab_data - Contains raw behavioral and neural data<br> ├── MonkeyAlignedDatasets_new - Contains padded test behavioral input for generating network activations<br> ├── MonkeyDatasets - Contains not aligned padded test behavioral input for generating network activations<br> ├── MonkeySpikeRegressDatasets - Contains datasets for training data-driven models<br> ├── MonkeySpikeRegressDatasets_new - Contains trial index for regression splits <br> └── new_beh_exp_dataframe - Contains pre-processed behavioral and neural data</p> <p>--------------------------------</p> <p>The activations and predictions for the best 3 models and for all tasks are stored in experiments folder (in .h5 format) that follows the same name convention of the checkpoints.</p> <p>The checkpoints are stored in experiment folders (experiment_***) that follow this scheme:<br>- Task: shallow exp id, deep TCNs exp id, LSTM id.</p> <p>Experiment IDs for each task:</p> <p>- Untrained: 15, 115, 45<br>- Classification: 4015, 5015, 4045</p> <p>- Torque: 8015, 8030, 8045</p> <p>- Regress joint pos: 17016, 17031, 17046<br>- Regress joint vel: 17216, 17231, 17246<br>- Regress joint pos & vel:: 17416, 17431, 17446<br>- Regress joint pos & vel & acc:: 20516, 20531, 20546</p> <p>- Regress hand pos: 4016, 5016, 4046<br>- Regress hand vel: 17316, 17331, 17346<br>- Regress hand pos & vel: 17516, 17531, 17546<br>- Regress hand pos & vel & acc: 20416, 17831, 17846</p> <p>- Regress hand and elbow pos: 20016, 20031, 20046<br>- Regress hand and elbow vel: 20916, 20931, 20946<br>- Regress hand and elbow pos & vel: 20616, 20631, 20646<br>- Regress hand and elbow pos & vel & acc: 20816, 20831, 20846</p> <p>- Redundancy reduction: 10020, 10035, 10050<br>- Autoencoder 20716 & 20717, 20731 & 20732, X</p> <p> </p> <p>The code to process the behavioral data is available at: <a href="https://github.com/amathislab/Task-driven-Proprioception/tree/master/exp_data_processing">https://github.com/amathislab/Task-driven-Proprioception/tree/master/exp_data_processing</a><br>The code to load and use the models to generate activations and predictions is available at: <a href="https://github.com/amathislab/Task-driven-Proprioception/tree/master/neural_prediction">https://github.com/amathislab/Task-driven-Proprioception/tree/master/neural_prediction</a></p> <p>To reproduce the results, it is possible to reproduce the main figures using the result dataframe. See our repository for more details. </p> <p>--------------------------------</p> <p>The datasets, weights, activations and predictions are released with Creative Commons Attribution 4.0 license.</p> <p>The code is released under the MIT license, see <a href="https://github.com/amathislab/Task-driven-Proprioception">https://github.com/amathislab/Task-driven-Proprioception</a></p> <p>If you find our code, weights, predictions or ideas useful, please cite:</p> <p>@article{vargas2024task,<br> title={Task-driven neural network models predict neural dynamics of proprioception},<br> author={{Marin Vargas}, Alessandro and Bisi, Axel and Chiappa, Alberto S and Versteeg, Chris and Miller, Lee E and Mathis, Alexander},<br> journal={Cell},<br> year={2024},<br> publisher={Elsevier}<br>}</p>
Train and Evaluation Code, Road Classification Models and Test set of the paper "Impact of Image Resolution and Image Overlap on the Prediction Performance of Convolutional Neural Networks Trained for Road Classification"
<p>This repository contains the Python scripts built for training and evaluation of the implementation, together with the test data and the resulting road classification models corresponding to the paper "Impact of Image Resolution and Image Overlap on the Prediction Performance of Convolutional Neural Networks Trained for Road Classification". The scripts make use of the Tensorflow with Keras framework and the additional required dependencies.</p> <p>The training and validation set is based on the binary SROADEX dataset (https://zenodo.org/records/6482346) that was re-split into tiles that feature the image resolutions (256 x 256, 512 x 512, and 1024 x 1024 pixels) and image overlaps (0% and 12.5%) considered in this study. The data have been generated using scripts developed in Python using Open Source libraries (GDAL/OGR and MapScript) for rasterization of vector cartography that represents the axes of the different types of roads (urban, interurban and rural). This binary road data contains information from 16 full orthoimages (28.5 km * 18.5 km) with spatial resolution of 0.5 m/pixel from the insular and peninsular Spanish territory. Due to the size on disk of approximately 546 gigabytes, this training and validation data is only available upon request from the corresponding author. The test set has been generated from a novel area of 28.5 km * 18.5 km and features binary road labels. The test sets are provided in the repository for each resolution (with no overlap), so that additional DL models can be evaluated on the same data and compared with the results achieved in this study.</p> <p>The structure of the information shared in this repository is as follows:<br>The scripts have been grouped by tile resolution (256, 512 and 1024). First, the test set and the evaluation script can be found. For each tile resolution, there are two subfolders (corresponding to the "no overlap" and "12.5% overlap"). In each case, the Python scripts for training the models in the three repetitions are shared, and the trained models (H5 format) are shared in compressed form. Finally, for each resolution we also share the testing dataset which consists of two folders.</p> <p>The material is distributed under a CC-BY 4.0 license.</p>
Extending Grime's CSR model to predict plant demographic responses across resource availability gradients: evidence from the Patagonian steppes
<p>Sexual reproduction, growth, and survival are crucial demographic strategies for plant population viability. Here, we propose a conceptual model predicting demographic responses of species based on their ecological strategy and the heterogeneity of environmental conditions within a biogeographical unit and then applied it to a case study from a 5-degree latitudinal gradient in the Patagonian steppes. We also aim to disentangle genetic from environmental effects on demographic responses. We performed <em>in-situ </em>and common garden experiments with two species from six local populations of the Occidental Phytogeographical District of the Patagonian steppes. Species differ in key ecological traits, and thus fit into Grime´s model for evolutionary strategies in plants: one as competitive species and the other as stress-tolerant species. We calculated population growth rate (λ) and performed elasticity analyses to compare the contribution of each demographic strategy to population fitness between species and among local populations distributed along 600 km latitudinal gradient with differences in mean annual precipitation (MAP). We highlight four results. First, the competitive species change from sexual reproduction to growth as MAP increases. Second, the stress-tolerant species relied on growth and survival along the MAP gradient. Third, interannual variation in resource availability modulated demographic responses for both strategies. Fourth, based on the comparison of the <em>in-situ</em> and common garden experiments, we submit that demographic responses were genetically driven. Our study shows that demographic responses can be roughly predicted by the ecological strategy across environmental gradients. We show that differences arise not only between species, but also were genetically driven differences within species among local populations. Scaling up plant-level responses to population-level dynamics allows for a process-based understanding of current and future biogeographical species organization. Furthermore, conservation and restoration efforts should be guided by demographic strategies underlying population viability.</p>
Morphing libraries, QSAR models, and compounds predicted to be active on the Glucocorticoid receptor (GR)
<p>This repository contains datasets and files related to the computational drug discovery project of the chemical space exploration of the Glucocorticoid receptor. The accompanying Python code is freely available in the GitHub repository (<a title="https://github.com/Iagea/GRML_analyses" href="https://github.com/Iagea/GRML_analyses" target="_blank" rel="noreferrer noopener">https://github.com/Iagea/GRML_analyses</a>).</p> <p><strong>Morphing Libraries:</strong></p> <ul> <li><strong>GRML_library.csv:</strong> The GRML library is the collection of 999,015 virtual compounds generated by Molpher [1-2] starting from GR ligands with unique Bemis-Murcko scaffolds collected from the ChEMBL17 and IMG libraries.</li> <li><strong>RML_library.csv:</strong> The RML library is the collection of 1,346,310 virtual compounds generated by Molpher starting from compounds with unique Bemis-Murcko scaffolds randomly selected from the ZINC database.</li> </ul> <p><strong>IMG library:</strong></p> <ul> <li><strong>IMG_non_proprietary.csv</strong>: The non-proprietary IMG library subset containing 12,956 compounds and their corresponding B-scores from the primary screen.</li> </ul> <p><strong>Molpher inputs:</strong></p> <ul> <li><strong>GR_inputs.csv</strong>: The GR inputs are the ligands used to create the GRML library, 204 compounds from ChEMBL17 (95 compounds) and the non-proprietary dataset from IMG (109 compounds).</li> <li><strong>Random_inputs.csv</strong>: The random inputs are 249 random ZINC compounds used to create the Random library.</li> </ul> <p><strong>Model's training sets:</strong></p> <ul> <li><strong>Model33_training_set.csv</strong>: Random forest classification model training set, it includes 865 compounds; known GR actives and inactives from ChEMBL33 (738 compounds) and non-proprietary active ligands from the IMG library (127 compounds).</li> <li><strong>Model17_training_set.csv</strong>: Random forest classification model training set, it includes 601 compounds; known GR actives and inactives from ChEMBL17 (474 compounds) and non-proprietary active ligands from the IMG library (127 compounds).</li> <li><strong>RFR_training_set.csv</strong>: Random forest regression model training set, it includes 89 compounds; known GR actives and inactives from ChEMBL33 that fit into the GR pharmacophore with the four features we describe in our paper.</li> </ul> <p><strong>Models:</strong></p> <ul> <li><strong>Model33.pkl:</strong> Python pickle file containing the trained Random forest classification models used along with Mondrian cross-conformal prediction to classify GR actives/inactives. This model was trained with ChEMBL33 and IMG libraries.</li> <li><strong>Model17.pkl</strong>: Python pickle file containing the trained Random forest classification models used along with Mondrian cross-conformal prediction to classify GR actives/inactives. This model was trained with ChEMBL17 and IMG libraries.</li> <li><strong>RFR_models.pkl</strong><em>: </em>Python pickle file containing the 100 trained random forest regression models used to rank the proposed active morphs. These models were trained with the RFR_training_set.csv.</li> </ul> <p><strong>Active predicted morphs:</strong></p> <ul> <li><strong>all_morphs_actives</strong><em><strong>_</strong></em><strong>predicted.xlsx:</strong> An Excel spreadsheet containing two sheets. 1) All 22,524 GRML active predicted morphs. 2) All 4,341 RML active predicted morphs. The QED, NIBR Severity Score, and Molskill Score are given for each morph.</li> </ul> <p><strong>Proposed GR active ligands:</strong></p> <ul> <li><strong>designed_ligands.xlsx</strong>: An Excel spreadsheet containing two sheets. 1) All 54 designed GR ligands with their QED, NIBR severity score, MolSkill score, consensus ranking from the 100 RFR models, and the result of the manual annotation and remarks, if available. 2) The structure of the 54 ligands based on their manual annotation and presence or not in ChEMBL33 database.</li> </ul> <p>Researchers and professionals in the field of drug discovery and cheminformatics may find these resources useful for further analysis and investigations.</p> <p><strong>Bibliography</strong></p> <p>[1] Hoksza, D., Škoda, P., Voršilák, M. <em>et al.</em> Molpher: a software framework for systematic chemical space exploration. <em>J Cheminform</em> <strong>6</strong>, 7 (2014). https://doi.org/10.1186/1758-2946-6-7</p> <p>[2] <a href="https://github.com/lich-uct/molpher-lib">https://github.com/lich-uct/molpher-lib</a></p>
JSON files containing parameters of training gene models for ab-initio prediction software
<p>These are the JSON files containing parameters of training gene models for ab-initio prediction software. These training datasets are Phytophthora specific and can be further utilized for the gene prediction and annotation of other related Phytophthora strains.</p>
DeepBacs – E. coli SIM prediction dataset and CARE model
<p>Training and test images of live, membrane-labeled <em>E. coli </em>cells for prediction of SIM super-resolution images from widefield images, as well as a trained CARE model.</p> <p>Additional information can be found on this <a href="https://github.com/HenriquesLab/DeepBacs/wiki">github wiki</a>.</p> <p>The example image shows a widefield fluorescence image and SIM reconstruction of FM5-95 labelled, live <em>E. coli </em>cells.</p> <p> </p> <p><strong>Training and test dataset</strong></p> <p><strong>Data type</strong>: Paired microscopy images (fluorescence) of low (widefield) and high resolution (SIM)</p> <p><strong>Microscopy data type</strong>: Fluorescence microscopy (FM5-95)</p> <p><strong>Microscope</strong>: GE HealthCare Deltavision OMX system (with temperature and humidity control, 37°C) equipped with an Olympus 60x 1.42NA Oil immersion objective and 2 PCO Edge 5.5 sCMOS cameras (one for DIC, one for fluorescence)</p> <p><strong>Cell type</strong>: <em>E. coli </em>DH5α grown under agarose pads</p> <p><strong>File format</strong>: .tif (16-bit for widefield images and 32-bit for SIM reconstructions)</p> <p><strong>Image size</strong>: 1024 x 1024 px² (40 nm/px)<br> <strong>Image preprocessing</strong>: <em>E. coli</em> widefield images were scaled with a factor of 2 to match the SIM reconstruction pixel size. </p> <p> </p> <p><strong>CARE model</strong></p> <p>The CARE 2D model was generated using the ZeroCostDL4Mic platform (Chamier et al., 2021). It was trained from scratch for 300 epochs on 5500 paired image patches (image dimensions: (1024 x 1024 px²), patch size: (80 x 80 px²), 100 patches/image) with a batch size of 8 and a laplace loss function, using the CARE 2D ZeroCostDL4Mic notebook (v 1.12). Key python packages used include tensorflow (v 0.1.12), Keras (v2.3.1), csbdeep (v 0.6.1), numpy (v1.19.5), cuda (v 10.1.243). The training was accelerated using a Tesla P100GPU and data was augmented by a factor of 4 using rotation and flipping.</p> <p>Model weights can be used with the ZeroCostDL4Mic CARE 2D notebook or the CSBDeep Fiji plugin.</p> <p> </p> <p><strong>Author(s)</strong>: Pedro Matos Pereira<sup>1,2</sup>, Mariana Pinho<sup>1,3</sup></p> <p><strong>Contact email</strong>: <a href="mailto:pmatos@itqb.unl.pt">pmatos@itqb.unl.pt</a> and <a href="mailto:mgpinho@itqb.unl.pt">mgpinho@itqb.unl.pt</a></p> <p> </p> <p><strong>Affiliation</strong>: </p> <p>1) Bacterial Cell Biology, Instituto de Tecnologia Química e Biológica António Xavier, Universidade Nova de Lisboa, Oeiras, Portugal</p> <p>2) ORCID: https://orcid.org/0000-0002-1426-9540</p> <p>3) ORCID: https://orcid.org/0000-0002-7132-8842</p>
PlasX model, Predicted plasmids, and Known plasmids
<p>PlasX model and all analyses of known and predicted plasmids</p>
Raw data of: "Controlling Hand Movements Relying on Tactile Illusions: A Model Predictive Control Framework"
<p>in Fig4_a.txt: raw the data for the plot of Fig4_a (x and y of the first simulated trajectory from trajectory 1 to 50)</p> <p>in Fig4_b.txt: raw the data for the plot of Fig4_b </p> <p>in Fig4_c.txt raw the data for the plot of Fig4_b. Each column corresponds to the optimal angle of the plate for each of the 50 trajectories simulated in Fig4_a</p>
Cellpose models for Label Prediction from Brightfield and Digital Phase Contrast images
<p><strong>Name: </strong>Cellpose models for Brightfield and Digital Phase Contrast images</p> <p><strong>Data type: </strong>Cellpose models trained via transfer learning from the ‘nuclei’ and ‘cyto2’ pretrained model with additional <strong>Training Dataset . Includes</strong> corresponding csv files with 'Quality Control' metrics(§) (model.zip).</p> <p><strong>Training Dataset: </strong>Light microscopy (Digital Phase Contrast or Brightfield) and automatic annotations (nuclei or cyto) (<a href="https://doi.org/10.5281/zenodo.6140064">https://doi.org/10.5281/zenodo.6140064</a>)</p> <p><strong>Training Procedure: </strong>The cellpose models were trained using cellpose version 1.0.0 with GPU support (NVIDIA GeForce K40) using default settings as per the <a href="https://cellpose.readthedocs.io/en/latest/train.html">Cellpose documentation</a> . Training was done using a <a href="https://datascience.ch/renku/">Renku </a>environment (<a href="https://github.com/BIOP/renku-templates/tree/main/VNC-Napari-Fiji-Omero-CUDA11.4-cellpose-omnipose">renku template</a>).</p> <p> </p> <p><strong>Command Line Execution for the different trained models</strong></p> <p><strong>nuclei_from_bf: </strong></p> <pre><code class="language-python">cellpose --train --dir 'data/train/' --test_dir 'data/test/' --pretrained_model nuclei --img_filter _bf --mask_filter _nuclei --chan 0 --chan2 0 --use_gpu --verbose</code></pre> <p><strong>cyto_from_bf</strong>:</p> <pre><code class="language-python">cellpose --train --dir 'data/train/' --test_dir 'data/test/' --pretrained_model cyto2 --img_filter _bf --mask_filter _cyto --chan 0 --chan2 0 --use_gpu --verbose</code></pre> <p> </p> <p><strong>nuclei_from_dpc:</strong></p> <pre><code class="language-python">cellpose --train --dir 'data/train/' --test_dir 'data/test/' --pretrained_model nuclei --img_filter _dpc --mask_filter _nuclei --chan 0 --chan2 0 --use_gpu --verbose</code></pre> <p><strong>cyto_from_dpc</strong>:</p> <pre><code>cellpose --train --dir 'data/train/' --test_dir 'data/test/' --pretrained_model cyto2 --img_filter _dpc --mask_filter _cyto --chan 0 --chan2 0 --use_gpu --verbose</code></pre> <p> </p> <p><strong>nuclei_from_sqrdpc</strong>:</p> <pre><code class="language-python">cellpose --train --dir 'data/train/' --test_dir 'data/test/' --pretrained_model nuclei --img_filter _sqrdpc --mask_filter _nuclei --chan 0 --chan2 0 --use_gpu --verbose</code></pre> <p><strong>cyto_from_sqrdpc</strong>:</p> <pre><code class="language-python">cellpose --train --dir 'data/train/' --test_dir 'data/test/' --pretrained_model cyto2 --img_filter _sqrdpc --mask_filter _cyto --chan 0 --chan2 0 --use_gpu --verbose</code></pre> <p> </p> <p><em><strong>NOTE </strong></em>(§): We provide a notebook for Quality Control, which is an adaptation of the <a href="https://colab.research.google.com/github/HenriquesLab/ZeroCostDL4Mic/blob/master/Colab_notebooks/Beta%20notebooks/Cellpose_2D_ZeroCostDL4Mic.ipynb">"Cellpose (2D and 3D)" notebook from ZeroCostDL4Mic</a> .</p> <p><em><strong>NOTE</strong></em>: This dataset used a training dataset from the Zenodo entry(<a href="https://doi.org/10.5281/zenodo.6140064">https://doi.org/10.5281/zenodo.6140064</a>) generated from the “HeLa “Kyoto” cells under the scope” dataset Zenodo entry(<a href="https://doi.org/10.5281/zenodo.6139958">https://doi.org/10.5281/zenodo.6139958</a>) in order to automatically generate the label images.</p> <p><strong><em>NOTE</em></strong>:<strong> </strong>Make sure that you delete the “_flow” images that are auto-computed when running the training. If you do not, then the flows from previous runs will be used for the new training, which might yield confusing results.</p> <p> </p>
Protein language model embeddings and predictions for the fly proteome (FlyBase)
<p>Residue and sequence embeddings of the fly (drosophila melanogaster) proteome (FlyBase for organism drosophila melanogaster, downloaded on 2022.03.01) computed using bio_embeddings (bioembeddings.com) using the ProtT5 embedder at full precision (https://www.biorxiv.org/content/10.1101/2020.07.12.199554v3). To open the embeddings file, please see <a href="https://github.com/sacdallago/bio_embeddings/blob/develop/notebooks/open_embedding_file.ipynb">this notebook</a>. The embeddings will be indexed by numbers according to the mapping file (mapping_file.csv) in this dataset. All following results will share the same mapping (for instance, to access the variation prediction results, by accessing index "0", you will query results for the sequence "FBpp0304622").</p> <p>Additionally:</p> <p>- Sequence-level predictions of subcellular localization in 10 classes using LA (https://www.biorxiv.org/content/10.1101/2021.04.25.441334v1)</p> <p>- Residue-level three state secondary structure prediction (alpha, sheet or other) using models reported in the ProtTrans paper (https://www.biorxiv.org/content/10.1101/2020.07.12.199554v3)</p> <p>- Residue-level prediction of conservation (in 9 states) and of variation effect (from 0 [no-effect] to 1 [effect]) using VESPAl (https://doi.org/10.1007/s00439-021-02411-y)</p> <p> </p> <p>Files included:</p> <p>- dmel-all-translation-r6.44.fasta --> FASTA-formatted sequences of drosophila melanogaster from FlyBase</p> <p>- mapping_file.csv --> A CSV file mapping the identifiers used in the following files (from 0 to 30737) to the identifiers in the FlyBase fasta file (dmel-all-translation-r6.44.fasta).</p> <p>- DSSP3_fly_ProtT5Sec.fasta --> Secondary structure predictions in three states for each residue of each protein in dmel-all-translation-r6.44.fasta. "H" stands for Helix; "E" stands for Sheet; "C" stands for Other.</p> <p>- subcell_fly_LA_ProtT5.csv --> Subcellular location (10 states) and memrane-boundness (2 states) for each protein in dmel-all-translation-r6.44.fasta</p> <p>- embeddings_file.h5 --> per-residue embeddings of sequences in dmel-all-translation-r6.44.fasta. Each dataset in the .h5 file represents a protein sequence and contains a matrix of length Lx1024, with L being the length of the protein sequence. Datasets are indexed using integers. The original sequence identifier (from the FASTA header) can be accessed through the "original_id" attribute. See https://docs.bioembeddings.com/v0.2.0/notebooks/open_embedding_file.html for information on how to open the file.</p> <p>- reduced_embeddings_file.h5 --> per-sequence embeddings of sequences in dmel-all-translation-r6.44.fasta (obtained by mean-pooling the residue-embeddings along the length dimension of the protein sequence). Each dataset in the .h5 file represents a protein sequence and contains a vector of size 1024 (meaning, each sequence has the same dimension).</p> <p>- conspred_probs.h5 --> per-sequence conservation probability (softmax) prediction of sequences in dmel-all-translation-r6.44.fasta in 9 classes. Each dataset in the .h5 file represents a protein sequence and contains a matrix of length 9xL, with L being the length of the protein sequence, and 9 being the predicted conservation class (index 0 = very variable; index 8 = very conserved)</p> <p>- vespal_SAVeffect_fly.zip --> zipped .h5 file of per-sequence variation predictions of sequences in dmel-all-translation-r6.44.fasta on a scale from 0 (neutral) to 1 (effect). -1 indicates WT substitution. Each dataset in the .h5 file represents a protein sequence and contains a matrix of length 20xL, with L being the length of the protein sequence, and 20 being the predicted variation score for each residue substitution (AAs in the following order: "<strong>ALGVSREDTIPKFQNYMHWC</strong>" . Meaning that index 0 = substitution of the residue to "A", index = 1 substitution to residue "L", aso.)</p>
Custom language model checkpoints used in "Testing the limits of natural language models for predicting human language judgments"
<p>Checkpoint files for an RNN, LSTM, BILSTM and n-gram models used the paper "Testing the limits of natural language models for predicting human language judgments"</p>
Supplementary Dataset for Deep learning based kcat prediction enables improved enzyme constrained model reconstruction
<p>This dataset is the supplementary dataset for the paper "<strong>Deep learning based <em>k</em><sub>cat</sub> prediction enables improved enzyme constrained model reconstruction</strong>". Protein sequence fasta files, deep learning predicted <em>k</em><sub>cat</sub> values, classcial-ecGEMs, DL-ecGEMs and <em>Posterior</em>-mean-ecGEMs for 343 yeast/fungi species are available in this dataset.This repository also contains the computed results for reproducing the figures as model_build_files . The scripts can be found in Github (https://github.com/SysBioChalmers/DLKcat)</p>
Development of a machine learning model to predict non- durable response to anti-TNF therapy in Crohn's disease using transcriptome imputed from genotypes
<p>This is the expression value predicted using PrediXcan version 7 to find a gene feature that can distinguish between patients with and without effect on infliximab.</p> <p>Among the various tissue models provided by PrediXcan v7, three models were selected and used: whole blood, Colon transverse, and terminal ileum of small intestine, and the predicted gene counts of each model were 6,294, 5,612 and 3,107.</p> <p>For each of the three models, predicted gene expression values and phenotype information per sample were submitted.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.