Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
13,064
datasets available to search
ShareScore release 0.9.0
Dataset results
13,064 results for “Prediction”
Mean NDVI Values (1982-2018) and Future Predictions Using CHELSA Bioclim Variables for Türkiye
<p>This dataset contains mean Normalized Difference Vegetation Index (NDVI) values from 1982 to 2018 and their future predictions based on CHELSA bioclimatic variables, specifically for the region of Türkiye. The data is provided in .asc format and includes both historical and projected NDVI values under different climate scenarios.</p> <h4>Contents:</h4> <ul> <li><strong>Historical NDVI Data (1982-2018)</strong>: Mean NDVI values derived from remote sensing data.</li> <li><strong>Future NDVI Predictions</strong>: NDVI projections for the periods 2011-2040, 2041-2070, and 2071-2100 under three Shared Socioeconomic Pathways (SSPs): SSP1-2.6, SSP3-7.0, and SSP5-8.5.</li> </ul> <h4>Methodology:</h4> <ol> <li><strong>Model Training</strong>: <ul> <li>A Random Forest Regressor was used to model the relationship between NDVI and the selected bioclim variables.</li> <li>The model achieved an R² of 0.9341, Mean Absolute Error of 0.0275, and Root Mean Squared Error of 0.0499.</li> </ul> </li> <li><strong>Future Predictions</strong>: <ul> <li>Future NDVI values were predicted using the trained model and future CHELSA bioclim projections.</li> <li>Predictions were made for three future periods (2011-2040, 2041-2070, 2071-2100) under three SSPs (SSP1-2.6, SSP3-7.0, SSP5-8.5).</li> </ul> </li> </ol> <h4>Data Specifications:</h4> <ul> <li><strong>Extent</strong>: Covers the geographical area of Türkiye and adjacents.</li> </ul> <h4>Sources:</h4> <ul> <li><strong>NDVI Data</strong>: <ul> <li>Ma, Z., Dong, C., Lin, K., Yan, Y., Luo, J., Jiang, D., & Chen, X. (2022). A Global 250-m Downscaled NDVI Product from 1982 to 2018. Remote Sensing, 14(15), 3639.</li> </ul> </li> <li><strong>CHELSA Bioclim Data</strong>: <ul> <li>Karger, D.N., Conrad, O., Böhner, J., Kawohl, T., Kreft, H., Soria-Auza, R.W., Zimmermann, N.E., Linder, P., Kessler, M. (2017). Climatologies at high resolution for the Earth land surface areas. Scientific Data. 4 170122. <a href="https://doi.org/10.1038/sdata.2017.122" target="_new" rel="noreferrer">https://doi.org/10.1038/sdata.2017.122</a></li> <li>Karger, D.N., Conrad, O., Böhner, J., Kawohl, T., Kreft, H., Soria-Auza, R.W., Zimmermann, N.E., Linder, H.P., Kessler, M. Data from: Climatologies at high resolution for the earth’s land surface areas. Dryad Digital Repository. <a href="http://dx.doi.org/doi:10.5061/dryad.kd1d4" target="_new" rel="noreferrer">http://dx.doi.org/doi:10.5061/dryad.kd1d4</a></li> </ul> </li> </ul>
Data for: Crall et al., Spatial fidelity of workers predicts collective response to disturbance in a social insect
<p>Dataset for: Crall et al., Spatial fidelity of workers predicts collective response to disturbance in a social insect, in final revision for Nature Communications.</p> <p>Includes two files - one behavioral data from uniquely identified worker bumblebees, and the second containing metadata for the colonies from which these data were generated (including experimental treatments, locations, sizes, etc.).</p>
Datasets for predicting TF binding using Virtual ChIP-seq
<p>This repository contains datasets necessary for using the Virtual ChIP-seq software.</p> <p>Virtual ChIP-seq requires the following datasets to predict transcription factor binding:</p> <ul> <li> <p>chipExpDir_AtoH_V1.0.0.tar.gz: Reference matrices of correlation between TF binding and gene expression for TFs starting with letters A-H.</p> </li> <li> <p>chipExpDir_ItoZ_V1.0.0.tar.gz: Reference matrices of correlation between TF binding and gene expression for TFs starting with letters I-Z.</p> </li> <li> <p>refTables_V1.1.0.tar.gz: PhastCons genomic conservation, FIMO PWM scores for JASPAR motifs, and ChIP-seq data of ENCODE and Cistrome database.</p> </li> <li> <p>hg38_chrsize.tsv: Length of chromosomes in hg38</p> </li> <li> <p>trainedModels_V1.0.0.tar.gz: Virtual ChIP-seq scikit-learn trained models saved in joblib format</p> </li> <li> <p><CellType>.tar.gz: Pre-calculated matrices suitable for training with other algorithms or re-training with Virtual ChIP-seq.</p> </li> </ul> <p>Some predictive features of TF binding are the same in each cell type and are stored together for simplicity in refTables_V1.0.0.tar.gz. You can use datasets from other cell types (named here as <CellType>.tar.gz) for the purpose of re-training the model. The <CellType>.tar.gz files contain pre-calculated predictive features of transcription factor binding in 4 chromosomes (5, 10, 15, 20).</p> <p>These features include:</p> <ul> <li> <p>PhastCons genomic conservation</p> </li> <li> <p>FIMO score for sequence motifs of TF in the JASPAR database</p> </li> <li> <p>Chromatin accessibility</p> </li> <li> <p>TF binding in ENCODE + Cistrome DB datasets</p> </li> <li> <p>Virtual ChIP-seq expression score</p> </li> </ul> <p> </p>
Predictive models for off-target binding profiles generation
<p>Models for predicting off-target binding, built with Conformal Prediction, and the <a href="http://cpsign-docs.genettasoft.com">CPSign software</a>. The dataset is part of an upcoming publication (Manuscript in preparation), which will provide more details.</p> <p>The dataset is a GZipped Tar archive, with the models as Java Archive (JAR) files. For every JAR-file, there is also a corresponding audit log, with the extension ".audit.json", produced by the workflow software (<a href="http://scipipe.org">SciPipe</a>) used to train the models. This audit file contains all the shell commands used in the workflow that produced the models.</p>
R-code for publication: Ensembles of Ensembles: Combining the Predictions from Multiple Machine Learning Methods
<p>This is the R-code as well as the underlying data needed to reproduce the results of the springer book chapter: "Ensembles of Ensembles: Combining the Predictions from Multiple Machine Learning Methods"</p> <p>For more information contact: dlieske@mta.ca</p> <p> </p>
Virtual ChIP-seq predictions of binding of 36 transcription factor in Roadmap Epigenomics Project tissues
<p>This dataset contains predictions of Virtual ChIP-seq for binding of 36 transcription factors in Roadmap Epigenomics dataset tissues with matched DNase-seq and RNA-seq data.</p> <p>Tarball contains subfolders for each of the 36 TFs where Virtual ChIP-seq median MCC in validation cell types was > 0.3.</p> <p>Each subfolder contains gzipped BED files. Each file is named as <Tissue>_<Age>_<TF>_<Accession>_Predictions.bed.gz. Columns correspond to Chromosome, Start, End, <Tissue>_<Age>_<TF>_<Accession>, Posterior probability</p> <p>You can use the posterior probabilities provided in Virchip_PosteriorCutoffs_V3.0.0.tsv. These are posterior probability cutoffs which maximized MCC in H1-hESC cell type, or are set to 0.4 if there was no ChIP-seq data of that TF in H1-hESC (0.4 is the mode of all optimal posterior probability cutoffs in H1-hESC).</p>
Global derived datasets for use in k-NN machine learning prediction of global seafloor total organic carbon
<p>This dataset includes 663 predictor grids used for k-NN global prediction of seafloor total organic carbon.</p> <p>663 predictor grids available in netCDF4 HDF5 file format. Grids are cell-centered sized 4320 x 2160. File names adhere to the naming conventions discussed below. The naming structure is partioned by underscores and periods in the following order: interface to which the gridded values refer to, quantity of values contained within the grid, units and reference values/units (e.g. meters below sea level), data source, statistic calculated (if applicable), grid pitch, and file extension.</p> <p>Possible interfaces from the top – down:</p> <p>SS – Sea surface – atmosphere interface (may also be average of the entire water column)</p> <p>SF – Seafloor – water interface (may also be denoted by GL)</p> <p>GL – Ground level (e.g. bottom of pure liquid, top of dirt)</p> <p>SC – Sediment – crust interface (e.g. sediment above, igneous/metamorphic below)</p> <p>CM – Crust – mantle interface (e.g. Mohorovicic discontinuity)</p> <p>Appropriate reference naming marker (bold), original data source, and date of last access:</p> <p><strong>Becker</strong></p> <p>Becker, J. J., Wood, W. T., & Martin, K. M. (2014). <em>Global crustal heat flow using random decision forest prediction</em>, Abstract NG31A-3788 presented at 2014 Fall Meeting, AGU, San Francisco, California, U.S.A. Last access: 06/23/2015.</p> <p><strong>CRUST1</strong> </p> <p>Pasyanos, M.E., Masters, G., Laske, G. & Ma, Z. (2012). <em>LITHO1.0 - An Updated Crust and Lithospheric Model of the Earth Developed Using Multiple Data Constraints</em>, Abstract T11D-09 presented at 2012 Fall Meeting, AGU, San Francisco, California, U.S.A. Last access: 07/01/2014.</p> <p><strong>CRUST1_NOAA</strong></p> <p> As the NOAA sediment thickness database is globally not complete, data gaps in the NOAA grid with this have been supplemented by the CRUST1 sediment thickness (see above citation).</p> <p>Whittaker, J., Goncharov, A., Williams, S., Müller, R. D., & Leitchenkov, G. (2013) Global sediment thickness dataset updated for the Australian-Antarctic Southern Ocean, <em>Geochemistry, Geophysics, Geosystems. </em>https://doi.org/10.1002/ggge.2018.<em> </em>Last access: 09/02/2018.</p> <p><strong>GVP</strong></p> <p>Global Volcanism Program (2013) Volcanoes of the World. In E. Venzke (ed.). (Vol. 4.7.3). Smithsonian Institution. https://doi.org/10.5479/si.GVP.VOTW4-2013. Last access: 09/22/2014.</p> <p><strong>ETOPO2v2</strong></p> <p>National Geophysical Data Center (2006). 2-minute Gridded Global Relief Data (ETOPO2) v2. National Geophysical Data Center, NOAA. DOI: 10.7289/V5J1012Q. Last access: 02/06/2013.</p> <p><strong>PLATES</strong></p> <p>Coffin, M.F., Gahagan, L.M., & Lawver, L.A. (1998). Present-day Plate Boundary Digital Data Compilation. University of Texas Institute for Geophysics Technical Report (No. 174, pp. 5). Last access: 09/15/2014.</p> <p><strong>ONRL</strong></p> <p>Ludwig,W., Amiotte-Suchet, P., & Probst, J. L. (2011). ISLSCP II Global River Fluxes of Carbon and Sediments to the Oceans. In F. G. Hall, G. Collatz, B. Meeson, S. Los, E. Brown de Colstoun, and D. Landis (Eds.), <em>ISLSCP Initiative II Collection</em>. Oak Ridge National Laboratory Distributed Active Archive Center, Oak Ridge, Tennessee, U.S.A. http://dx.doi.org/10.3334/ORNLDAAC/1028. Last Access: 02/15/2015.</p> <p><strong>Muller</strong></p> <p>Müller, R. D., Sdrolias, M., Gaina, C., & Roest, W. R. (2008). Age, spreading rates, and spreading asymmetry of the world’s ocean crust, <em>Geochemistry, Geophysics, Geosystems</em>, 9(4), Q04006. https://doi.org/10.1029/2007GC001743. Last accessed: 07/19/2011.</p> <p><strong>Woa13x</strong></p> <p>Boyer, T.P., Antonov, J. I., Baranova, O. K., Coleman, C., Garcia, H. E., Grodsky, A., et al. (2013) World Ocean Database 2013. In S. Levitus, A. Mishonov (Ed.), <em>NOAA Atlas NESDIS 72, Technical Ed</em>. Silver Spring, MD. http://doi.org/10.7289/V5NZ85MT. Last Access: 09/18/2014.</p> <p><strong>KIM</strong></p> <p>Kim, S.S. & Wessel, P. (2011). New global seamount census from the altimetry-derived gravity data, <em>Geophysical Journal International</em>, 186, 615-631. https://doi.org/10.1111/j.1365-246X.2011.05076.x. Last access: 09/22/2014.</p> <p><strong>HYCOM</strong></p> <p>The 1/12 deg global HYCOM+NCODA Ocean Reanalysis was funded by the U.S. Navy and the Modeling and Simulation Coordination Office. Computer time was made available by the DoD High Performance Computing Modernization Program. The output is publicly available at https://hycom.org/publications/acknowledgements/ocean-reanalysis-data.Last access: 03/19/2014.</p> <p><strong>NCEDC</strong></p> <p>NCEDC (2016). Northern California Earthquake Data Center. UC Berkeley Seismological Laboratory. Dataset. doi:10.7932/NCEDC. Last access: 09/21/2014.</p> <p><strong>Wei2010</strong></p> <p>Wei, C.-L., Rowe, G. T., Escobar-Briones, E., Boetius, A., Soltwedel, T., Caley, M. J., et al.(2010). Global patterns and predictions of seafloor biomass using random forests. <em>PLoS ONE</em>,5(12), e15323. https://doi.org/10.1371/journal.pone.0015323 Last access: 06/20/2016.</p> <p><strong>NGA_egm2008</strong></p> <p>Pavlis, N.K., Holmes, S. A., Kenyon, S. C., & Factor, J. K. (2008). <em>The</em> <em>EGM2008 Global Gravitational Model</em>, Abstract 2008AGUFM.G22A..01P presented at the 2008 General Assembly of the European Geosciences Union, Vienna, Austria. Last access: 07/10/2014.</p> <p><strong>WAVEWATCH3</strong></p> <p>The 1/12 deg global HYCOM+NCODA Ocean Reanalysis was funded by the U.S. Navy and the Modeling and Simulation Coordination Office. Computer time was made available by the DoD High Performance Computing Modernization Program. The output is publicly available at https://hycom.org/publications/acknowledgements/ocean-reanalysis-data. Last access: 03/19/2014.</p> <p>Updated global seafloor porosity grid using our k-nearest neighbors algorithm using 5 nearest neighbors. Observed data used for prediction from Martin et al. (2015). </p> <p>Martin, K. M., Wood, W. T., & Becker, J. J. (2015). A global prediction of seafloor sediment porosity using machine learning. <em>Geophysical Research Letters</em>, 42(24), 10640. https://doi.org/10.1002/2015GL065279</p> <p>Other grids which have been generated by empirical means are latitude (and derivatives), longitude (and derivatives), Coriolis, coast_is_1.0, and the random noise grids. </p> <p>Units referenced are as follows:</p> <p>KGM3 - kilogram per cubic meter<br> MS - meters per second<br> KM - kilometer<br> M_ASL - meters above sea level (i.e. meters referenced to sea level)<br> MWM2 - milliwatt per square meter<br> TGCYR - terragram of carbon per year<br> TGYR - terragram per year<br> MA - megaannum<br> M - meters<br> MGCM2 - milligram of carbon per square meter<br> DEG - degree<br> S - seconds</p> <p>Statistics grids are calculated within a given radius (e.g. 10km, 50km, 125km, 250km, 500km, 1000km) of the respective cell-centered value. The statistics grids include mean (.men), average absolute deviation from the mean (.aad), and the common logarithm (.log) of the absolute value of the mean (.mlg). Additionally, some grids are a weighted count for given radii (e.g. seamounts) where weight is a cosine taper from the center of the grid cell. </p> <p>The grid pitch for this dataset is uniformly at 5-arc minute denoted by “.5m”. Additionally, the extension used (netCDF4) is denoted by “.nc”.</p>
Ontology based text mining of gene-phenotype associations: application to candidate gene prediction
<p>Gene-phenotype associations play an important role in understanding<br> the disease mechanisms which is a requirement for treatment<br> development. A portion of gene-phenotype associations are observed<br> mainly experimentally and made publicly available through several<br> standard resources such as MGI. However, there is still a vast<br> amount of gene--phenotype associations buried in the biomedical<br> literature. Given the large amount of literature data, we need<br> automated text mining tools to alleviate the burden in manual<br> curation of gene-phenotype associations and to develop<br> comprehensive resources. We developed an ontology based<br> approach in combination with statistical methods to text mine<br> gene-phenotype associations from literature. Our method achieved<br> AUC values of 0.90 and 0.75 in recovering known gene-phenotype<br> associations from HPO and MGI respectively. We posit that candidate<br> genes and their relevant diseases should be expressed with similar<br> phenotypes in publications. Thus, we demonstrate the utility of our<br> approach by predicting disease candidate genes based on the semantic<br> similarities of phenotypes associated with genes and diseases. We evaluated our disease candidate prediction model on<br> the gene-disease associations from MGI. Our model achieved AUC<br> values of 0.90 and 0.87 on OMIM (human) and MGI (mouse) datasets of<br> gene-disease associations respectively. Our manual analysis on the<br> text mined data revealed that, our method can accurately extract<br> gene-phenotype associations which are not currently covered by the<br> existing public gene-phenotype resources. Overall, results indicate<br> that our method can precisely extract known as well as new<br> gene-phenotype associations from literature. This released dataset at Zenodo covers our gene-phenotype extracts from the literature. All the methods used to extract the data are available at https://github.com/bio-ontology-research-group/genepheno.</p>
Replication package for "An Empirical Assessment of Best-Answer Prediction Models in Technical Q&A Sites" (EMSE 2018)
<p>Replication package for the paper:</p> <blockquote> <p>F. Calefato, F. Lanubile, and N. Novielli (2018) “<a href="http://collab.di.uniba.it/fabio/wp-content/uploads/sites/5/2018/07/EMSE-D-17-00159_R3.compressed.pdf">An Empirical Assessment of Best-Answer Prediction Models in Technical Q&A Sites</a>.” Empirical Software Engineering Journal, DOI: <a href="http://dx.doi.org/10.1007/s10664-018-9642-5">10.1007/s10664-018-9642-5</a>.</p> </blockquote>
Prediction Experiment for Western Kho-Bwa language data: dataset
<p><strong>Prediction Experiment on Western Kho-Bwa languages</strong></p> <p><em>Timotheus A. Bodt (SOAS, London) and Johann-Mattis List (Max Planck Institute, Jena)</em></p> <p>This database includes all the sound files and the transcriptions of the prediction experiment for Western Kho-Bwa. This experiment was registered as:</p> <p>Bodt, Timotheus A., Nathan W. Hill and Johann-Mattis List. 2018. <em>Prediction experiment for missing words in Kho-Bwa language data. </em>Open Science Framework Preregistrations October 5. <a href="https://osf.io/evcbp/">https://osf.io/evcbp/</a> </p> <p>The data and code can be found on:</p> <p>Timotheus A. Bodt, Nathan W. Hill, & Johann-Mattis List. (2018, October 8). Prediction experiment for missing words in Kho-Bwa language data (Version v1.0.1). Zenodo. <a href="http://doi.org/10.5281/zenodo.1451176">http://doi.org/10.5281/zenodo.1451176</a></p> <p>A paper explaining the experiment is under review:</p> <p>Bodt, Timotheus A. and Johann-Mattis List. 2019 (under review). Testing the predictive force of the comparative method: An ongoing experiment on unattested words in Western Kho-Bwa languages. <em>Papers in Historical Phonology</em> Volume 1: 1–21.</p> <p>The results of the experiment will be presented at the International Conference on Historical Linguistics 24: 01-Jul-2019 - 05-Jul-2019, Canberra, Australia.</p> <p>The uncut sound files, cut sound files, original field notes and preliminary transcriptions have been saved as:</p> <p>Bodt, Timotheus Adrianus. (2019). <em>'Retrodiction' experiment Western Kho-Bwa languages: data [Data set]</em>. Zenodo. <a href="http://doi.org/10.5281/zenodo.2529727">http://doi.org/10.5281/zenodo.2529727</a></p> <p><strong>How to use these files?</strong></p> <ul> <li>Download the zip folder soundfiles_prediction_experiment.zip</li> <li>Extract the files in a separate folder</li> <li>Search for the required sound file(s)</li> </ul> <p>Searching sound files can best be done using the English CONCEPTS from the predictions_results.csv file. For example, searching for BACK will give all the sound files that contain the English gloss ‘back’ (including ‘backwards’, ‘back’ as body part, turn ‘back’ etc.).</p> <p>Another option is the select all the sound files of a given linguistic variety / doculect by searching for the original sound file number.</p> <p>I would advise against using a certain attested form in the predictions_results.csv file and search for that (e.g. p a ŋ + b u ‘chest’), because the cut sound files have been saved without spaces and morpheme breaks and because the actual transcriptions of the sound files may have changed after analysis, but were not updated in the name of the cut sound files.</p> <p>If you cannot find a certain sound file, then it may simply not have been recorded or not cut from the main sound file. If you are really interested, please mail me at <a href="mailto:timintibet@hotmail.com">timintibet@hotmail.com</a> and I will attempt to find it or record it.</p>
Consensus models to predict oral rat acute toxicity and validation on a dataset coming from the industrial context
<p>We report predictive models of acute oral systemic toxicity representing a follow-up of our previous work in the framework of the NICEATM project. It includes the update of original models through the addition of new data and an external validation of the models using a dataset relevant for the chemical industry context. A regression model for LD50 and classification model for toxicity classes according to the Global Harmonized System categories were prepared. ISIDA descriptors were used to encode molecular structures. Machine learning algorithms included Support Vector Machine (SVM), Random Forest (RF) and Naïve Bayesian. Selected individual models were combined in consensus.</p> <p>The different datasets were compared using the Generative Topographic Mapping approach. It appeared that the NICEATM datasets were lacking some relevant chemotypes for chemical industry. The new models trained on enlarged data sets have applicability domain (AD) sufficiently large to accommodate industrial compounds. The fraction of compounds inside the models’ AD increased from 58 % (NICEATM model) to 94 % (new model). Yet, the increase of training sets only slightly improved of the models’ prediction performance: RMSE values decreased from 0.56 to 0.47 and balanced accuracies increased from 0.69 to 0.71 for NICEATM and new models, respectively.</p>
Extremely Imbalanced Smell-based Defect Prediction
<p><strong>Abstract: </strong>In continuous integration/continuous delivery, one of the main requirements for high-speed delivery of software is to find bugs efficiently. For this reason, multiple solutions were introduced in the literature. For instance, defect prediction approaches based on bad code smells detected in modules from each version of the software. Nevertheless, these approaches do not consider the problem where there may exist an extremely higher percentage of non-defective modules compared to defective modules. Given that, each version of the software may only have a small number of defects. As a result, in this thesis, we introduce a new model with an autoencoder algorithm that uses design and implementation smells to detect defective modules. Therefore, we trained five autoencoders with distinct architectures. Ad- ditionally, for evaluation, we compared each model against autoencoders with the same architecture, trained with traditional object-oriented metrics and the combination of both. Our analysis did not show promising results, as the use of only smells and the combination of features did not provide an improve- ment compared with the use of metrics. However, we introduce a starting point for smell-based defect prediction in the context of dataset imbalance. Furthermore, we introduce a baseline for future work.</p> <p> </p> <p><strong>Dataset Description:</strong></p> <p>We provide three datasets. The first results from the extraction of traditional object-oriented metrics (metric.csv). The second results from the extraction of design and implementation smells (smell.csv). The third is the combination of all the features (metricsmell.csv). Moreover, these features were extracted from Designite and Bugsdorjar software archives.</p>
Time-resolved compound repositioning predictions on a text-mined knowledge network
<p><strong>gs_positives.csv</strong>: The re-processed version of DrugCentral indications, utilized as training and testing positives in the analysis.</p> <p><strong>top_5000_predictions.csv</strong>: The top 5000 drug-disease pairs, by probability, produced by this analysis pipeline.</p> <p><strong>file_info.txt</strong>: Information about the column headings in of the two files.</p> <p> </p>
Data Sets and Prediction Models Created using MLP, RNN and LSTM
<p>Binary Data Sets</p> <p>- 2018tbi219_shuffled3.csv and 2018tbi219_shuffled5.csv</p> <p>- these are stratified data sets that were able to produce prediction models with high prediction rates.</p> <p>Models.zip</p> <p>- this zip file contains the prediction models created using Keras Deep Learning Algorithms: MLP, RNN and LSTM</p> <p>Model Creation - Python Code Snippets</p> <p>- Code snippets for creating the prediction models</p>
Microclimate predicts frost-hardiness of alpine Arabidopsis thaliana populations better than elevation
<p>In mountain regions, topological differences on the micro-scale can strongly affect microclimate and may counteract the average effects of elevation, such as decreasing temperatures. While these interactions are well understood, their effect on plant adaptation is understudied.</p> <p> </p> <p>We investigated winter frost hardiness of Arabidopsis thaliana accessions originating from 13 sites along altitudinal gradients in the Southern Alps during three winters on an experimental field station on the Swabian Jura and compared levels of frost damage with the observed number of frost days and the lowest temperature in eight collection sites.</p> <p> </p> <p>We found that frost-hardiness increased with elevation in a log-linear fashion. This is consistent with adaptation to a higher frequency of frost conditions, but also indicates a decreasing rate of change in frost hardiness with increasing elevation. Moreover, the number of frost days measured with temperature loggers at the collection sites correlated much better with frost-hardiness than the elevation of collection sites, suggesting that populations were adapted to their local microclimate. Notably, the variance in frost days across sites increased exponentially with elevation. Together, our results suggest that strong microclimate heterogeneity of high alpine environments can preserve functional genetic diversity among small populations.</p> <p> </p> <p>Synthesis. Here we tested how plant populations differed in their adaptation to frost exposure along an elevation gradient and whether microsite temperatures improve the prediction of frost hardiness. We found that local temperatures, particularly the number of frost days, is a better predictor of the frost hardiness of plants than elevation. This reflects a substantial variance in frost frequency between sites at similar high elevations. We conclude that high mountain regions harbor microsites that differ in their local microclimate and thereby can preserve a high functional genetic diversity among them. Therefore, high mountain regions have the potential to function as a refugium in times of global change.</p>
QSPR models for bioconcentration factor (BCF): Are they able to predict data of industrial interest?
<p>This dataset is described and studied in the article </p> <p>"QSPR models for bioconcentration factor (BCF): Are they able to predict data of industrial interest?"</p> <p>published in <em>SAR and QSAR Environmental Research</em> (Taylor&Francis).</p> <p>Files description:</p> <p>SI_BCFtrainset.xlsx: a collection of 1129 chemical structures and CAS identifiers with their logBCF values extracted from various literature sources.</p> <p>SI_BCFtestset.xlsx: a collection of 204 chemical structures for which the logBCF is considered of lower reliability and used as an external test set.</p> <p>SI_FullDataset_rawdata.csv: the raw data composed of 15372 entries with the following columns: CASRN, Tissue, Duration [d], Test organism, Exposure type, Steady state, RESPONSE, RESPONSE UNIT, Media type, TakenFrom, TITLE, AUTHOR, YEAR, SOURCE, SMILES</p> <p>SI_ExcludedOutliers34.csv: 34 chemical structures that have been identified as suspicious during analysis.</p> <p> </p>
DeepPredSpeech: computational models of predictive speech coding based on deep learning
<p>This dataset contains all data, source code, pre-trained computational predictive models and experimental results related to: </p> <p>Hueber T., Tatulli E., Girin L., Schwatz, J-L "How predictive can be predictions in the neurocognitive processing of auditory and audiovisual speech? A deep learning study." (<a href="https://doi.org/10.1101/471581">biorXiv preprint https://doi.org/10.1101/471581</a>). </p> <ul> <li>Raw data are extracted from the publicly available database NTCD-TIMIT (10.5281/zenodo.260228). <ul> <li>Audio recordings are available in the audio_clean/ directory</li> <li>Post-processed lip image sequences are available in the lips_roi/ directory (67x67 pixels, 8bits, obtained by lossless inverse DCT-2D transform from the DCT feature available in the original repository of NTCD-TIMIT)</li> <li>Phonetic segmentation (extracted from NTCD-TIMIT original zenodo repository) is available in the HTK MLF file volunteer_labelfiles.mlf</li> </ul> </li> <li>Audio features (MFCC-spectrogram and log-spectrogram) are available in the mfcc_16k/ and fft_16k/ directories. </li> <li>Models (audio-only, video-only and audiovisual, based on deep feed-forward neural networks and/or convolutional neural network, in .h5 format, trained with Keras 2.0 toolkit) and data normalization parameters (in .dat scikit-learn format) are available in models_mfcc/ and models_logspectro/ directories</li> <li>Predicted and target (ground truth) MFCC-spectro (resp. log-spectro) for the test databases (1909 sentences), and for the different values of <span class="math-tex">\(\tau_p\)</span> or <span class="math-tex">\(\tau_f\)</span> are available in pred_testdb_mfccspectro/ (resp. pred_testdb_logspectro/) directory</li> </ul> <p>Source code for extracting audio features, training and evaluating the models is available on GitHub https://github.com/thueber/DeepPredSpeech/</p> <p>All directories have been zipped before upload.</p> <p>Feel free to contact me for more details.</p> <p>Thomas Hueber, Ph. D., CNRS research fellow, GIPSA-lab, Grenoble, France, thomas.hueber@gipsa-lab.fr </p>
Data from "Predicting Global Ground Geoelectric Field With Coupled Geospace and Three‐Dimensional Geomagnetic Induction Models"
<p>Data presented in http://dx.doi.org/10.1029/2018SW001859 excluding the first and last hours which were determined to contain partially unphysical results and probably should not be used.</p> <p>Each file contains the external ground magnetic field components calculated on 5x5 degree geographic grid and the results of induction modeling using 1d and 3d ground conductivity models: total ground magnetic field components, horizontal ground electric field components. Times are given in UTC.</p> <p>To reproduce Figure 8 in above reference use:</p> <p> </p> <p>import numpy<br> import matplotlib.pyplot<br> data = numpy.load('2006-12-14T22:58:00.npz') # or 2006-12-14T22_58_00.npz<br> matplotlib.pyplot.colorbar(<br> matplotlib.pyplot.imshow(<br> data['B_3D_north'],<br> cmap = matplotlib.pyplot.get_cmap('bwr'),<br> vmin = -800, vmax = 800,<br> extent = (-180, 180, -90, 90)<br> ),<br> format = '%.0f', fraction = 0.02, pad = 0.03<br> )<br> matplotlib.pyplot.show()</p> <p> </p> <p>and substitute B_3D_east, E_3D_north, etc. for the different panels.</p>
Spring Precipitation Amount and Timing Predict Restoration Success in a Semi-Arid Ecosystem Code and Data
<table> <tbody> <tr> <td>The data here is summary data compiled from all years of the project that lead to the publication Spring Precipitation Amount and Timing Predict Restoration Success in a Semi-Arid Ecosystem with the Journal of Applied Ecology and code to analyze these data. Our study was focused on the Northern Great Basin ecosystem. We conducted surveys at 48 sites over the course of five years (2016-2020). All were located on public lands managed by either the Bureau of Land Management, Idaho Department of Lands, or Oregon State Lands Department. We looked at the influence of management, biotic, abiotic and weather variables predicting seedling establishment success, 45 predictor variables in all. Machine learning techniques were used to select most important predictor variables to be used in future work predicting good seedling establishment windows. </td> </tr> </tbody> </table>
Spatiotemporal prediction of soil organic carbon density (SOCD) for pan-Europe (2000-2022) in 3D+T
<h2><strong>Sub-dataset: WPCT p025, 2016–2020</strong></h2> <h2>Disclaimer</h2> <p>This is the first release of pan-EU predictions of soil health indicators (the Soil Health Data Cube). Use for testing purposes only. A publication describing methods used has been submitted to PeerJ and is in review. Funded by the European Union. Views and opinions expressed are however those of the author(s) only and do not necessarily reflect those of the European Union or European Commision. Neither the European Union nor the granting authority can be held responsible for them. The data is provided "as is". AI4SoilHealth project consortium and its suppliers and licensors hereby disclaim all warranties of any kind, express or implied, including, without limitation, the warranties of merchantability, fitness for a particular purpose and non-infringement. Neither AI4SoilHealth project Consortium nor its suppliers and licensors, makes any warranty that the Website will be error free or that access thereto will be continuous or uninterrupted. You understand that you download from, or otherwise obtain content or services through, the Website at your own discretion and risk.</p> <h2>Description</h2> <p>This dataset covers pan-European areas, including Ukraine, the UK, and Turkey. This data cube could be used for applications such as soil property mapping and comprehensive soil health assessment across Europe. The dataset spans four depth ranges and multiple time periods, providing information for studies on soil organic carbon stock and dynamics.</p> <p>This dataset is part of the Spatiotemporal prediction of soil organic carbon density for Europe (2000-2022) in 3D+T dataset. Check the related identifiers section below to access other parts of the dataset.</p> <p>This data set includes:</p> <ul> <li><strong>Soil Organic Carbon Density (SOCD) (2000-2022, 4-year intervals):</strong><br> This data includes mean, p975, and p025 SOCD maps for four depth ranges (0-20cm, 20-50cm, 50-100cm, and 100-200cm) in kg/m<sup>3</sup> (scaled 10x). </li> <li><strong>Organic carbon content based on dry combustion weight percentage (WPCT) (2000-2022, 4-year intervals):</strong><br> This data includes mean, p975, and p025 WPCT maps for four depth ranges (0-20cm, 20-50cm, 50-100cm, and 100-200cm) in percentage. </li> </ul> <h3>Related identifiers</h3> <ul> <li><strong>SOCD mean:</strong><br> <a href="https://zenodo.org/records/13754343">2000-2004</a> <a href="https://zenodo.org/records/13771721">2004-2008</a> <a href="https://zenodo.org/records/13771841">2008-2012</a> <a href="https://zenodo.org/records/13771911">2012-2016</a> <a href="https://zenodo.org/records/13771967">2016-2020</a> <a href="https://zenodo.org/records/13772054">2020-2022</a> </li> <li><strong>SOCD p025:</strong><br> <a href="https://zenodo.org/records/13779539">2000-2004</a> <a href="https://zenodo.org/records/13774064">2004-2008</a> <a href="https://zenodo.org/records/13774089">2008-2012</a> <a href="https://zenodo.org/records/13774114">2012-2016</a> <a href="https://zenodo.org/records/13774167">2016-2020</a> <a href="https://zenodo.org/records/13774196">2020-2022</a> </li> <li><strong>SOCD p975:</strong><br> <a href="https://zenodo.org/records/13778472">2000-2004</a> <a href="https://zenodo.org/records/13773396">2004-2008</a> <a href="https://zenodo.org/records/13773765">2008-2012</a> <a href="https://zenodo.org/records/13773828">2012-2016</a> <a href="https://zenodo.org/records/13773953">2016-2020</a> <a href="https://zenodo.org/records/13774003">2020-2022</a> </li> <li><strong>WPCT mean:</strong><br> <a href="https://zenodo.org/records/13785010">2000-2004</a> <a href="https://zenodo.org/records/13785079">2004-2008</a> <a href="https://zenodo.org/records/13785170">2008-2012</a> <a href="https://zenodo.org/records/13785306">2012-2016</a> <a href="https://zenodo.org/records/13785419">2016-2020</a> <a href="https://zenodo.org/records/13785553">2020-2022</a> </li> <li><strong>WPCT p025:</strong><br> <a href="https://zenodo.org/records/13786250">2000-2004</a> <a href="https://zenodo.org/records/13786314">2004-2008</a> <a href="https://zenodo.org/records/13786449">2008-2012</a> <a href="https://zenodo.org/records/13786565">2012-2016</a> <a href="https://zenodo.org/records/13786594">2016-2020</a> <a href="https://zenodo.org/records/13786702">2020-2022</a> </li> <li><strong>WPCT p975:</strong><br> <a href="https://zenodo.org/records/13785722">2000-2004</a> <a href="https://zenodo.org/records/13785854">2004-2008</a> <a href="https://zenodo.org/records/13785917">2008-2012</a> <a href="https://zenodo.org/records/13786029">2012-2016</a> <a href="https://zenodo.org/records/13786119">2016-2020</a> <a href="https://zenodo.org/records/13786214">2020-2022</a> </li> </ul> <h3>Data Details</h3> <ul> <li><strong>Time period:</strong> 2000–2022, in 4-year intervals (last period covers 2020–2022).</li> <li><strong>Type of data:</strong> Spatiotemporal soil organic carbon data cube, with depth ranges and weighted percentage data for soil carbon assessments.</li> <li><strong>How the data was collected or derived:</strong> The data was derived using machine learning models.</li> <li><strong>Statistical methods used:</strong> Quantile Random Forest</li> <li><strong>Limitations or exclusions in the data:</strong> The dataset does not include data for Svalbard. </li> <li><strong>Coordinate reference system:</strong> EPSG:3035</li> <li><strong>Bounding box (Xmin, Ymin, Xmax, Ymax):</strong> (900,000, 899,000, 7,401,000, 5,501,000)</li> <li><strong>Spatial resolution:</strong> 30m</li> <li><strong>Image size:</strong> 216,700P x 153,400L</li> <li><strong>File format:</strong> Cloud Optimized Geotiff (COG) format.</li> </ul> <h3>Support</h3> <p>If you discover a bug, artifact, or inconsistency, or if you have a question please raise a GitHub issue: GitLab Issues (tbc)</p> <h3>Name convention</h3> <p>To ensure consistency and ease of use across and within the projects, we follow the standard Ai4SoilHealth and Open-Earth-Monitor file-naming convention. The convention works with 10 fields that describe important properties of the data. In this way users can search files, prepare data analysis etc, without needing to open files. The fields are:</p> <ol> <li><strong>generic variable name:</strong> oc = organic carbon</li> <li><strong>variable procedure combination:</strong> iso.10694.1995.mg.cm3 = ISO method 10694:1995, with values in mg/cm<sup>3</sup> for SOCD | iso.10694.1995.wpct = ISO method 10694:1995, with values in weighted percentage of organic carbon content.</li> <li><strong>Position in the probability distribution/variable type:</strong> m = mean | p975 = percentile 97.5 | p025 = percentile 2.5</li> <li><strong>Spatial support:</strong> 30m</li> <li><strong>Depth reference:</strong> b0cm..20cm = depth range from 0 to 20cm</li> <li><strong>Time reference begin time:</strong> 20000101 = 2000-01-01</li> <li><strong>Time reference end time:</strong> 20041231 = 2004-12-31</li> <li><strong>Bounding box:</strong> eu = pan-Europe</li> <li><strong>EPSG code:</strong> epsg.3035</li> <li><strong>Version code:</strong> v20240804 = version from 2024-08-04</li> </ol>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.