Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

558

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

558 results for “Training Data”

Learn how ShareScore rates datasets ↗
dryad32/100

Data from: Out of Africa: the slow train to Australasia

We used mitochondrial DNA (mtDNA) sequences to test biogeographic hypotheses for Patiriella exigua (Asterinidae), one of the world's most widespread coastal sea stars. This small intertidal species has an entirely benthic life history and yet occurs in southern temperate waters of the Atlantic, Indian, and Pacific oceans. Despite its abundance around southern Africa, southeastern Australia, and several oceanic islands, P. exigua is absent from the shores of Western Australia, New Zealand, and South America. Phylogenetic analysis of mtDNA sequences (cytochrome oxidase I, control region) indicates that South Africa houses an assemblage of P. exigua that is not monophyletic (P = 0.04), whereas Australian and Lord Howe Island specimens form an interior monophyletic group. The placement of the root in Africa and small genetic divergences between eastern African and Australian haplotypes strongly suggest Pleistocene dispersal eastward across the Indian Ocean. Dispersal was probably achieved by rafting on wood or macroalgae, which was facilitated by the West Wind Drift. Genetic data also support Pleistocene colonization of oceanic islands (Lord Howe Island, Amsterdam Island, St. Helena). Although many biogeographers have speculated about the role of long-distance rafting, this study is one of the first to provide convincing evidence. The marked phylogeographic structure evident across small geographic scales in Australia and South Africa indicates that gene flow among populations may be generally insufficient to prevent the local evolution of monophyly. We suggest that P. exigua may rely on passive mechanisms of dispersal.

opencc-zeroDec 2017View details →
dryad32/100

Data from: Linking habitat composition, local population densities and traffic characteristics to spatial patterns of ungulate-train collisions

1. Total length of railways worldwide exceeds 1 million kilometres and recent railway development directly impacts wildlife because of animal-train collisions. Few studies, however, have analysed factors driving ungulate-train collisions. 2. We analysed over 3500 ungulate-train collisions including roe deer, red deer, wild boar, and moose collected in 2012-2015 in Poland. We compared train traffic characteristics (e.g. traffic intensity, speed, rail curvature), land-use and habitat characteristics (e.g. share of forests and build-up areas) and local ungulate population densities at collision sites and random sites distributed along the rail network. 3. Forest coverage generally increased, while urban areas decreased ungulate collision risk. Local density of ungulate species was strongly positively related to the relative collision risk in all four ungulate species, but above certain densities, the risk levelled off for all four species. 4. Train speed and train traffic intensity were positively associated with elevated collision risk in all four species, but the latter in a non-linear manner reached an asymptote at the level of ca. 10 trains per day. Rail curvature also increased probability of collisions with roe deer and red deer and possibly also wild boar. 5. Mortality rate of ungulates on railways in Poland is estimated to be 0.13-0.42% of annual hunting bags of studied species assuming that only one individual is killed at each occasion and ignoring undetected collisions. These values are expected to increase in near future due to increasing train speed in Central European countries. 6. Synthesis and applications. Ungulate-train collisions spots are characterised by surrounding forest, rail curvature, high train speed, and a moderate to high train traffic intensity. To reduce collision risk in a cost-effective way, we suggest to prioritise mitigation actions at sections of the railway characterized by those factors, e.g. by fencing and various warning devices. Due to nonlinear correlation between collision risk and population density, reducing density of ungulates will most likely reduce collision risk only marginally, and only in regions of low population densities where collision risk is relatively low anyway.

opencc-zeroSep 2019View details →
dryad32/100

Data from: Bachelors level soil science training at land grant institutions in the USA and its territories

Concern over the status of soil science education in the USA has led to a number of publications in recent years that track trends in student enrollment and offer suggestions for attracting more students to soil science. However, there is little information about changes in the number of degree programs that prepare students for careers as soil scientists, and such changes are obviously an important measure of the status of our field. This study established criteria to identify bachelor's degree programs that prepare students for soil science careers and used websites at land grant colleges to review the degree offerings of these schools in USA states and territories to determine if they met the established criteria. Fifty-nine land grant colleges were identified that offer bachelor's degree programs that prepare students for soil science careers, with a total of 61 degree programs since two of the schools had two separate bachelor's degrees that met the established criteria. This study provides guidelines for conducting similar future studies and a baseline against which they can be compared to allow us to determine whether we are gaining or losing soil science programs at the land grant colleges over time.

opencc-zeroDec 2018View details →
zenodo32/100

Data used for training glioblastoma NF1 classifier

<p>All data is publicly available and downloaded from UCSC Xena<br /> https://genome-cancer.ucsc.edu/proj/site/xena/datapages/?cohort=TCGA%20Pan-Cancer</p> <p>Because the database is continously updated and to ensure reproducibility,&nbsp;access data from this cached download.</p> <p>RNAseq and Clincal data were downloaded on 8 March 2016<br /> Mutation data was downloaded on 12 June 2015</p>

opencc-zeroJun 2016View details →
zenodo32/100

Data and trained model for iPXRDnet

<p>This data set is a collection of data sets and model checkpoint in the iPXRDnet.</p> <p><br>Model checkpoint file(model.zip):<br>hmof-130T_Hydrogen: Model of adsorption prediction for H2 in the hMOF-130T database obtained by training<br>hmof-130T_CarbonDioxide: Model of adsorption prediction for CO2 in the hMOF-130T database obtained by training<br>hmof-130T_Nitrogen: Model of adsorption prediction for N2 in the hMOF-130T database obtained by training<br>hmof-130T_Methane: Model of adsorption prediction for CH4 in the hMOF-130T database obtained by training<br>hmof-300T: Adsorption prediction model in the hMOF-300T database obtained by training<br>Gas_Se: Separation selectivity prediction model obtained by training<br>Gas_SD: Self-diffusion coefficients prediction model obtained by training<br>MOD: Bulk modulus and shear modulus prediction model obtained by training<br>exAPMOF-1bar-ALM+PXRD: Experimental adsorption at 1 bar model of Anion-pillared MOFs obtained by training with PXRD and material ligands<br>exAPMOF-1bar-ALM: Experimental adsorption at 1 bar model of Anion-pillared MOFs obtained by training with material ligands only<br>exAPMOF-1bar-PXRD: Experimental adsorption at 1 bar model of Anion-pillared MOFs obtained by training with PXRD only<br>exAPMOF-ISO:Experimental adsorption isotherm model of Anion-pillared MOFs obtained by training<br>exAPMOF-1bar-NOacvPXRD: Experimental adsorption at 1 bar model of Anion-pillared MOFs obtained by training with PXRD data before activation only<br>exAPMOF-1bar-acvPXRD: Experimental adsorption at 1 bar model of Anion-pillared MOFs obtained by training with PXRD data after activation only</p> <p>Checkpoint file for model without co-learning strategy(model-No-co-learning.zip):<br>hmof-130T_Hydrogen: Model of adsorption prediction for H2 in the hMOF-130T database obtained by training<br>hmof-130T_CarbonDioxide: Model of adsorption prediction for CO2 in the hMOF-130T database obtained by training<br>hmof-130T_Nitrogen: Model of adsorption prediction for N2 in the hMOF-130T database obtained by training<br>hmof-130T_Methane: Model of adsorption prediction for CH4 in the hMOF-130T database obtained by training<br>hmof-130T-str: Structural characteristics prediction model in the hMOF-130T database obtained by training<br>hMOF-130T-GradCAM: Model for GradCAM in the hMOF-130T database obtained by training<br>hmof-300T: Adsorption prediction model in the hMOF-300T database obtained by training<br>hmof-300T-str: Structural characteristics prediction model in the hMOF-300T database obtained by training<br>Gas_Se: Separation selectivity prediction model obtained by training<br>Gas_SD: Self-diffusion coefficients prediction model obtained by training<br>MOD: Bulk modulus and shear modulus prediction model obtained by training</p> <p><br>Data sets file (data.zip):<br>hmof-xrd+str+ad :PXRD and gas adsorption and structural feature of hmof-300T database<br>hMOF-130T_ad_list_mof :Gas adsorption data of hmof-130T database&nbsp;<br>hMOF-130T_GAS_DICT :Gas descriptors data of hmof-130T database<br>hMOF-130T_STR_DICT :Structural feature data of hmof-130T database<br>hMOF-130T_PXRD_DICT :PXRD data of hmof-130T database<br>MOD_data :Bulk modulus and shear modulus data of Moghadam's MOFs<br>MOD_PXRD_dict : PXRD data of Moghadam's MOFs<br>GAS_SD-data : self-diffusion coefficients data in CoREMOF database<br>SE-CO2,N2_data:Separation selectivity ,PXRD and structural feature of CO2/N2 selectivity database<br>Sa_sp:Data set partitioning results of CO2/N2 selectivity database<br>gas_dict : gas descriptors data used in the self-diffusion coefficients database<br>PXRD_DICT : PXRD data after activation of MOFs in Anion-pillared MOFs' experimental database<br>xrd_noacv : PXRD data before activation of MOFs in Anion-pillared MOFs' experimental database<br>Smiles_ads : Smiles data of gas in Anion-pillared MOFs' experimental database<br>all_exAPMOF-1bar : Anion-pillared MOFs' experimental adsorption data under 298K and 1 bar.<br>all_exAPMOF-1bar-NOacv : Experimental adsorption data for anion-pillared MOFs with PXRD before activation under 298K and 1 bar.<br>exAPMOF_DICT : Anion-pillared MOFs' Smiles data of MOFs' ligands and descriptors of metal centers in the experimental database<br>all_exAPMOF-iso : Key library of MOF and gas combinations in Anion-pillared MOFs' experimental isotherm database.<br>exAPMOF_ISOdata: Anion-pillared MOFs' experimental adsorption isotherm data under 298K.</p> <p>&nbsp;</p> <p>New Data sets file (data_new.zip):</p> <p>Files starting with &lsquo;4gas_&rsquo;: Files are named in the form of &lsquo;4gas_Gas_Pressure&rsquo;, recording the adsorption amounts of 20,000 randomly selected structures from the hMOF-130T database at different pressures at 298K.<br>all_adinfo_list_robustness: Gas adsorption data file used for study of model robustness .<br>hmof-130T_Xnm: PXRD data of structures in the hmof-130T database at different crystal sizes.<br>hMOF-130T_ad_list_mof: Corrected Gas adsorption data of hmof-130T database . There are problems with the data in the data.zip file.<br>X_PXRD,AD: Adsorption amount and PXRD data of ionic MOFs (iMOFs), zeolitic imidazolate frameworks (ZIF) and Uio-66 series from experimental literature. Among them, the Uio-66 series uses CO2 as the gas, and other files contain adsorption data of different gases.</p> <p>&nbsp;</p>

opencc-by-4.0Sep 2024View details →
zenodo32/100

"An efficient ptychography reconstruction strategy through fine-tuning of large pre-trained deep learning model" train and test data

<ul><li>Model for the article "An efficient ptychography reconstruction strategy through fine-tuning of large pre-trained deep learning model".</li><li>The &nbsp;.pth file is the pre-trained PtyNet-S model and the fine-tuned PtyNet-B model.</li><li>Please contact panxy@ihep.ac.cn if you have any questions.</li></ul>

opencc-by-4.0Nov 2023View details →
zenodo32/100

MME-only models trained with clean data for JAMES paper "Machine-learned uncertainty quantification is not magic"

<p>This tar file contains all 100 trained models in the MME-only ensemble from Experiment 1 (i.e., those trained with clean data, not with lightly perturbed data). &nbsp;To read one of the models into Python, you can use the method neural_net.read_model in the ml4rt library.</p>

opencc-by-4.0Nov 2023View details →
zenodo32/100

MME-only models trained with lightly perturbed data for JAMES paper "Machine-learned uncertainty quantification is not magic"

<p>This tar file contains all 100 trained models in the MME-only ensemble from Experiment 2 (i.e., those trained with lightly perturbed data). &nbsp;To read one of the models into Python, you can use the method neural_net.read_model in the ml4rt library.</p>

opencc-by-4.0Nov 2023View details →
zenodo32/100

MME/CRPS models trained with clean data for JAMES paper "Machine-learned uncertainty quantification is not magic"

<p>This tar file contains all 100 trained models in the MME/CRPS ensemble from Experiment 1 (i.e., those trained with clean data, not with lightly perturbed data). &nbsp;To pare the ensemble down to 50 models, we randomly select 50. &nbsp;To read one of the models into Python, you can use the method neural_net.read_model in the ml4rt library.</p>

opencc-by-4.0Nov 2023View details →
zenodo32/100

MME/CRPS models trained with lightly perturbed data for JAMES paper "Machine-learned uncertainty quantification is not magic"

<p>This tar file contains all 100 trained models in the MME/CRPS ensemble from Experiment 2 (i.e., those trained with lightly perturbed data). &nbsp;To pare the ensemble down to 50 models, we randomly select 50. &nbsp;To read one of the models into Python, you can use the method neural_net.read_model in the ml4rt library.</p>

opencc-by-4.0Nov 2023View details →
zenodo32/100

Data and training checkpoints of BTO

<p>Bright field images and training checkpoints of freestanding BTO in various temperatures and chemical environments</p>

opencc-by-4.0Nov 2023View details →
zenodo32/100

Dataset for "Explainable Offline-Online Training of Neural Networks for Parameterizations: A 1D Gravity Wave-QBO Testbed in the Small-data Regime" by Pahlavan et al. (2023)

Open the record for dataset details and reuse information.

opencc-by-4.0Dec 2023View details →
zenodo32/100

Training data for the shared task Ideology and Power Identification in Parliamentary Debates (2024)

<p>This dataset contains a selection of speeches from <a href="https://www.clarin.eu/parlamint">ParlaMint</a> corpora (version 4.0) as the training set for &nbsp;the shared task on "<a href="https://touche.webis.de/clef24/touche24-web/ideology-and-power-identification-in-parliamentary-debates.html">Ideology and Power Identification in Parliamentary Debates</a>" in <a href="https://clef2024.imag.fr/">CLEF 2024</a>.</p> <p>All files are tab-separated text files with the following fields:</p> <ul> <li>"<em>id</em>" is a unique (arbitrary) ID for each text.</li> <li>"<em>speaker</em>" is a unique (arbitrary) ID for each speaker. There may be multiple speeches from the same speaker.</li> <li>"<em>sex</em>" is the (binary/biological) sex of the speaker. This information is collected from varying sources (typically data published by the respective parliament), and in some cases it may be unspecified or unknown.</li> <li>"<em>text</em>" is the transcribed text of the parliamentary speech. Real examples may include line breaks, and other special sequences escaped or quoted.</li> <li>"<em>text_en</em>" is an automatic English translation of the corresponding text. This field may be empty (obviously) &nbsp;for speeches in English, but the translations may be missing for a small number of non-English speeches as well.</li> <li>"<em>label</em>" is the binary/numeric label. For political orientation, 0 is left and 1 is right. For power identification 0 indicates coalition (or governing party) and 1 indicates opposition.</li> </ul> <p>File names indicate the task and the parliament. We provide data from&nbsp;the following national and regional parliaments.</p> <ul> <li>Austria (at)</li> <li>Bosnia and Herzegovina (ba)</li> <li>Belgium (be)</li> <li>Bulgaria (bg)</li> <li>Czechia (cz)</li> <li>Denmark (dk)</li> <li>Estonia (ee) [only political orientation]</li> <li>Spain (es)</li> <li>Catalonia (es-ct)</li> <li>Galicia (es-ga)</li> <li>Basque Country (es-pv) [only power]</li> <li>Finland (fi)</li> <li>France (fr)</li> <li>Great Britain (gb)</li> <li>Greece (gr)</li> <li>Croatia (hr)</li> <li>Hungary (hu)</li> <li>Iceland (is) [only political orientation]</li> <li>Italy (it)</li> <li>Latvia (lv)</li> <li>The Netherlands (nl)</li> <li>Norway (no) [only political orientation]</li> <li>Poland (pl)</li> <li>Portugal (pt)</li> <li>Serbia (rs)</li> <li>Sweden (se) [only political orientation]</li> <li>Slovenia (si)</li> <li>Turkey (tr)</li> <li>Ukraine (ua)</li> </ul> <p>The number of training instances and the class imbalance differs for each training set. We do not provide a fixed validation split. Please see the <a href="https://touche.webis.de/clef24/touche24-web/ideology-and-power-identification-in-parliamentary-debates.html">shared task website</a> for further description of the data set and the sampling process.</p>

opencc-by-4.0Jan 2024View details →
zenodo32/100

A sample of the training data used in the paper "A Hybrid Physics-AI (HyPhAI) approach for probability fields advection: Application to cloud cover nowcasting"

<p>Copyright (2024) EUMETSAT</p>

opencc-by-4.0Feb 2024View details →
zenodo32/100

Potential and training data for 'Structure-property relations of silicon oxycarbides studied using a machine learning interatomic potential'

<p>Fitted potential, training and testing data.</p>

opencc-by-4.0Nov 2023View details →
zenodo32/100

DIPROMATS 2024 - Shared Task 2: few-shot training data for narrative identification

<p>Narratives are causally connected sequences of events that are selected and evaluated as meaningful for a particular audience. They make sense of the world by identifying the significance of people, places, objects, and events in time. In international relations, international actors create strategic narratives to &ldquo;construct a shared meaning of the past, present, and future of international politics to shape the behavior of domestic and international actors&rdquo;</p> <p>DIPROMATS 2024 Task 2 is a multiclass multilabel classification problem. Given a series of predefined narratives of each international actor, systems must determine which narrative the tweets belong to. Systems will receive the description of each narrative and a few examples of tweets in both languages (English and Spanish) that belong to each of them (few-shot learning). A tweet may be associated with one, several or none of the narratives.</p> <p>These are the few-shot training datasets for Englsih and Spanish.</p> <p>These files don't contain the narratives description. You can find them in the testing dataset:</p> <p>Pe&ntilde;as, A., Fraile-Hern&aacute;ndez, J. M., Moral, P., Rodrigo, &Aacute;., Deriu, J., Sharma, R., Centeno, R., Rodr&iacute;guez-Garc&iacute;a, R., Giedemann, P., &amp; Reyes-Montesinos, J. (2024). DIPROMATS 2024 - Shared Task 2: testing data for narrative identification (1.0.0) [Data set]. Zenodo. <a href="https://doi.org/10.5281/zenodo.12663310" target="_blank" rel="noopener">https://doi.org/10.5281/zenodo.12663310</a></p>

opencc-by-4.0Mar 2024View details →
zenodo32/100

preprocessed training data

Open the record for dataset details and reuse information.

opencc-by-4.0Mar 2024View details →
zenodo32/100

$swint vision model training data

<p>Training language models to see using strings that represent pixelized images.</p>

opencc-by-4.0Mar 2024View details →
zenodo32/100

Source data for manuscript(De novo protein design with a denoising diffusion network independent of pre-trained structure prediction models)

<p>This respository contains the source data for figure and supplementary figure in manuscript(SCUBA-D).</p>

opencc-by-4.0Apr 2024View details →
zenodo32/100

Source data for manuscript(De novo protein design with a denoising diffusion network independent of pre-trained structure prediction models)

<p>This respository contains the source data for figure and supplementary figure in manuscript(SCUBA-D).</p>

opencc-by-4.0Apr 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record