Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

3,481

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

3,481 results for “data set”

Learn how ShareScore rates datasets ↗
zenodo32/100

Data set for the publication

<p>Data set for the publication</p><p>Identification and validation of novel snoRNA-based biomarkers for clear cell renal cell carcinoma from urine-derived extracellular vesicles</p><p>&nbsp;</p>

opencc-by-4.0Jul 2023View details →
zenodo32/100

[Data set] Alloy [FA,Cs]PbI3 Perovskite Surfaces, Stability and Tolerance to Defect Formation

<p>This is the data repository related to the simulations of alloy [FA,Cs]PbI<sub>3</sub> Perovskite Surfaces.</p> <p>The surfaces here (Slabs.tgz) included are the relaxed structure of the slab models obtained by making cuts in the (001) direction in the special quasi-random structure (SQS) bulk models of FA<sub>1-x</sub>Cs<sub>x</sub>PbI<sub>3</sub> with (x=0.25 and 0.5), and the supercell models of pure FAPbI<sub>3</sub> and CsPbI<sub>3</sub> perovskites.[1] Besides, we include the optimized structures of each neutral vacancy pair defects (Slabs-defects.tgz) of formamidinium iodide and/or cesium iodide&nbsp;created in the most stable alloy [FA,Cs]I-C-FA<sub>0.75</sub>Cs<sub>0.25</sub>PbI<sub>3</sub> surface (see details in the paper). The ionic relaxations were performed with VASP code (version 6.2.1), using the PBE exchange-correlation functional, including Van der Waals corrections using the Grimme method with zero-damping function.</p> <p>Finally, the data set includes the simulated ab initio molecular dynamics (AIMD) trajectories of the most stable alloy and pure slabs, [FA,Cs]I-C-FA<sub>0.75</sub>Cs<sub>0.25</sub>PbI<sub>3</sub> &nbsp;and FAI-PbI<sub>3</sub>), including neutral vacancy pair defects of formamidinium iodide on a surface (Slabs-defects-AIMD.tgz). The AIMD calculations were performed with the CP2K code (V7.1), evaluating the forces with the PBE functional with the Grimme correction scheme (DFT- D3, Zero&ndash;damped correction). The trajectory productions include up to 15 ps using the microcanonical ensemble with 0.5 fs of time-step, considering 5 ps of thermalization time. More details in the article support information.&nbsp;</p> <p>&nbsp;</p> <p>Reference:</p> <p>1. G. M. Dalpian, X. G. Zhao, L. Kazmerski, A. Zunger, <em>Chem. Mater.</em> <strong>31</strong>, 2497&ndash;2506 (2019).</p>

opencc-by-4.0Feb 2024View details →
zenodo32/100

Reduced Gun Violence Frame Corpus data set for the Text2Story 2024 article: "Evaluating the Ability of Computationally Extracted Narrative Maps to Encode Media Framing"

<p><strong>Title: </strong>Simplified Gun Violence Frame Corpus (GVFC) Subset</p> <p><strong>Description:</strong><br>This data set is a simplified subset of the Gun Violence Frame Corpus (GVFC) from Liu et al. (2019). The original GVFC consists of 1300 news articles in English from multiple U.S. based sources extracted during the year 2018, focusing on media frames commonly used when reporting the issue of Gun Violence. The original data set has 9 types of frames, including both issue-specific and generic frames. Due to high computational costs in our analysis methods, we decreased the data set size from 1300 articles to 131 articles using stratified sampling, maintaining the original distribution of the frame labels. We also manually searched for the original sources of each article based on its headline and added the missing temporal information and news source to the data set, as it was required by our algorithms.</p> <p>To further reduce the complexity of the framing model and account for the smaller data set size, we grouped the original nine frames into three higher-level frames:</p> <p>1. Frame 1: Political Issues - Combining the first, second, and third frames, which focus on political issues mostly related to gun control.<br>2. Frame 2: Public Services - Combining the fourth and fifth frames, which focus on mental healthcare issues, as well as school and public safety.<br>3. Frame 3: Cultural and Societal Issues - Combining the last four frames, which are oriented towards cultural or societal issues, including discussions around race and ethnicity, public opinion, and economic consequences.</p> <p>The resulting simplified data set contains 131 news articles, each labeled with one of the three higher-level frames, along with the necessary temporal information and news source for the narrative extraction process.</p> <p>If you use this data set, please make sure to cite the original GVFC paper and our workshop paper please.&nbsp;</p> <p><strong>References:</strong></p> <ol> <li>Liu et al. (2019) "Detecting Frames in News Headlines and Its Application to Analyzing News Framing Trends Surrounding US Gun Violence", 23rd Conference on Computational Natural Language Learning (CoNLL 2019).</li> <li>Concha Mac&iacute;as, Sebasti&aacute;n and Keith Norambuena, Brian (2024). "Evaluating the Ability of Computationally Extracted Narrative Maps to Encode Media Framing", Text2Story 2024 Workshop, ECIR 2024.</li> </ol>

opencc-by-4.0Mar 2024View details →
zenodo32/100

Data set for Control of Ge island coalescence for the formation of nanowires on silicon

<p>This document contains all the data and the analysis used in the manuscript titled " Control of Ge island coalescence for the formation of nanaowires on silicon"</p> <p><a title="Link to landing page via DOI" href="https://doi.org/10.1039/D3NH00573A">https://doi.org/10.1039/D3NH00573A</a></p>

opencc-by-4.0Mar 2024View details →
zenodo32/100

Data set for Coherent Hole Transport in Selective Area Grown Ge Nanowire Networks

<p>This document contains all the data and analysis used in the manuscript titled&nbsp;</p> <h1><span>Coherent Hole Transport in Selective Area Grown Ge Nanowire Networks</span></h1> <p><span><a title="DOI URL" href="https://doi.org/10.1021/acs.nanolett.2c00358">https://doi.org/10.1021/acs.nanolett.2c00358</a></span></p>

opencc-by-4.0Mar 2024View details →
zenodo32/100

A data set from a survey investigating the SMART approach to develop good cyber security metrics

Open the record for dataset details and reuse information.

opencc-by-4.0Mar 2024View details →
zenodo32/100

NGII Data Set for Black Ice Traffic Accident Prediction

<p><a href="../api/records/10863284/draft/files/NGII%20Data%20Set%20for%20Black%20Ice%20Traffic%20Accident%20Prediction.zip/content" target="_blank" rel="noopener noreferrer">Title: NGII Data Set for Black Ice Traffic Accident Prediction</a></p> <p>This dataset has been processed for scholarly purposes, utilizing data provided by the National Geographic Information Institute of Korea.</p> <p>&lt;Reference&gt;</p> <p>National Geographic Information Institute. (n.d.). National Land Information Platform. Retrieved from https://map.ngii.go.kr/ms/map/NlipMap.do</p>

opencc-by-4.0Mar 2024View details →
zenodo32/100

Comparative membrane proteomic analysis of Tritrichomonas foetus isolates (non filtered Data Sets)

<p>Tritrichomonas foetus is a flagellated and anaerobic parasite able to infect cattle and felines. Despite its prevalence, there is no effective standardized or legal treatment for T. foetus-infected cattle; the vaccination still has limited success in mitigating infections and reducing abortion risk; and nowadays, the diagnosis of T. foetus presents important limitations in terms of sensitivity and specificity in bovines. Here, we characterize the plasma membrane proteome of T. foetus and identify proteins that are represented in different isolates of this protozoan.&nbsp; Raw proteomics data sets from MALDI-TOF Mass Spectrometry presented here corresponds to six T. foetus isolates (Tf0-Tf5). For Tf2 isolate also five membrane fractions are presented (f1-f5).&nbsp;&nbsp;</p>

opencc-by-4.0Mar 2024View details →
zenodo32/100

Capturing Periodic I/O Using Frequency Techniques [Data Set]

<div>This file contains the data set from the paper: "Capturing Periodic I/O Using Frequency Techniques," which was accepted at the IPDPS 2024. <div>The Instructions are provided in the <a href="https://github.com/tuda-parallel/FTIO/tree/main/artifacts/ipdps24">FTIO GitHub</a>: https://github.com/tuda-parallel/FTIO/tree/main/artifacts/ipdps24</div> </div> <div>&nbsp;</div> <div>After extracting data.zip, the folder named&nbsp;<em>data</em> has the following structure:</div> <div> <pre>data ├── application_traces │&nbsp;&nbsp; ├── HACC-IO │&nbsp;&nbsp; ├── IOR │&nbsp;&nbsp; ├── LAMMPS │&nbsp;&nbsp; ├── NEK5000 │&nbsp;&nbsp; └── README.md ├── exps_with_synthetic_traces ├── iosets_ftio_experiments └── README.md</pre> </div> <div><br>The folder iosets_ftio_experiments and exps_with_synthetic_traces in data.zip are snapshots from the repositories: <ul> <li><a href="https://gitlab.inria.fr/hpc_io/iosets-ftio-experiments">https://gitlab.inria.fr/hpc_io/iosets-ftio-experiments</a></li> <li><a href="https://gitlab.inria.fr/hpc_io/ftio_paper_exps_with_synthetic_traces">https://gitlab.inria.fr/hpc_io/ftio_paper_exps_with_synthetic_traces</a></li> </ul> </div>

opencc-by-4.0Feb 2024View details →
zenodo32/100

NemNet try out data set

<p>This data set contains publicly available monitoring data of NEMNET</p> <p><a href="https://www.onderzoeksfaciliteiten.nl/node/3910">National Environmental Monitoring Network | Grootschalige wetenschappelijke infrastructuur (onderzoeksfaciliteiten.nl)</a></p>

opencc-by-4.0Mar 2024View details →
zenodo32/100

Prediction of individual disease progression including parameter uncertainty in rare neurodegenerative diseases: the example of Autosomal-Recessive Spastic Ataxia Charlevoix Saguenay (ARSACS) - code and data sets

<p>This repository contains the scripts for the paper in revision to the AAPS J: Prediction of individual disease progression including parameter uncertainty in rare neurodegenerative diseases: the example of Autosomal-Recessive Spastic Ataxia Charlevoix Saguenay (ARSACS)&nbsp;</p> <p>Authors: Niels Hendrickx, MSc, France Mentr&eacute;, MD, PhD, Andreas Trasch&uuml;tz, MD, PhD, Cynthia Gagnon, PhD, Rebecca Sch&uuml;le, MD, ARCA Study Group, EVIDENCE-RND consortium, Matthis Synofzik, MD, Emmanuelle Comets, PhD</p> <p>A simulated dataset (<strong>simulated_arsacs.csv</strong>) has been included in the repository to make the code executable as a standalone. Four main scripts have been provided in addition with the present Readme describing the files. The repository also includes 3 R objects and 2 folders which will be overwritten when the scripts are run, and are included as examples of the expected outputs. The main scripts are:</p> <p>- <strong>Script_imputation_selection.R</strong>: runs the covariate selection method. It uses a simulated dataset provided in the depot. The multiple imputation model is hardcoded as an input to the mice package to generate 10 imputed datasets, saved in current_directory/imputed_data_sets/df_arsacs_mi_i.csv. The script then runs the covariate selection method. The script prints out the list of selected covariates and returns a saemixObject containing the fit of the selected covariate model.<br>&nbsp;After the script executes, a list will be saved with the name of the selected covariates in the current directory (an example is included under the name "cov_matrix_model.RData" in the repository), the output of the selection, containing the whole history of runs will be saved under "final_covariate_model.RData", the list of selected covariate names will be saved under "list_covariates.RData".</p> <p>- <strong>source_mi.R</strong>: contains the functions used by Script_imputation_selection.R</p> <p>- <strong>script_bootstrap_indfit.R</strong>: This script loads "cov_matrix_model.RData" containing the matrix of covariate effects (used by saemix) and "list_covariates.RData", the list of covariates included, fits the model on the imputed data sets and computes its bootstrap distribution for each imputed data set (in the script, using only 20 samples for computation time, saved in current_directory/bootstrap/boot.arsacs.case.mi.i). It then computes the mean parameter and relative standard error of each parameter. It then computes the conditional distribution of each patient in each bootstrap samples and returns a data frame of individual predictions. The script will then plot 4 indivudal predictions.&nbsp;</p> <p>-<strong> source_bootstrap.R</strong>: contains the functions used by script_bootstrap_indfit.R</p> <p>Both scripts need the saemix package to run, which we haven&rsquo;t included in the repository as it is freely available on the CRAN (https://cran.r-project.org/web/packages/saemix/index.html). Additional libraries we make use of in the code (MICE, tidyverse, ggplot2) also need to be installed prior to execution.&nbsp;<br>The R code provided can be further customised to be adapted to different scenarios.</p> <p>For the code to run, it is preferable to unzip the whole folder and set the working directory to the source file location as the script uses the "bootstrap" and "imputed_data_sets" sub-folders</p> <p>To execute this code, assuming the required libraries are available in the local R installation, please open an R session and run:<br>source("Script_imputation_selection.R") # for the covariate selection method (runtime: 3h on a &nbsp;i7-8565U laptop)<br>source("script_bootstrap_indfit.R") # to obtain individual trajectories (runtime: 1h on a &nbsp;i7-8565U laptop)</p>

opencc-by-4.0Apr 2024View details →
zenodo32/100

Learning tasks and result data for the 2024 GECCO short paper Length-niching Selection and Spatial Crossover in Variable-length Evolutionary Rule Set Learning

<p>Learning tasks and result data for the 2024 GECCO short paper Length-niching Selection and Spatial Crossover in Variable-length Evolutionary Rule Set Learning by David P&auml;tzel, Richard Nordsieck and J&ouml;rg H&auml;hner.</p>

opencc-by-4.0Apr 2024View details →
zenodo32/100

Data Sets For: Heterogeneous Structure, Mechanisms of Counterion Exchange, and the Spacer Salt Effect in Complex Molten Salt Mixtures Including LaCl3

<p>Data Sets For: Heterogeneous Structure, Mechanisms of Counterion Exchange, and the Spacer Salt Effect in Complex Molten Salt Mixtures Including LaCl3</p>

opencc-by-4.0Mar 2024View details →
zenodo32/100

Points_Data_Set

Open the record for dataset details and reuse information.

opencc-by-4.0Apr 2024View details →
zenodo32/100

Data set for A Novel VNS-based Algorithm for SVC Allocation in the Brazilian Interconnected Power System

<p>This release includes the 107-bus version of the Brazilian Interconnected Power System (available <a href="https://www.sistemas-teste.com.br/">here</a>). The system consists of 107 buses, 104 lines, and 67 transformers distributed across three areas: South, Southeast, and Mato Grosso. This test system provides extensive applications for problems related to steady-state analysis.</p>

opencc-by-4.0Nov 2024View details →
zenodo32/100

Mnemosine Data-Set

<p>This repository contains the OPEN-DATA information for collection Mnemosine.</p>

opencc-by-4.0Jun 2022View details →
zenodo32/100

Open Data Set for the article Analysing the impact of renewables on Iberian wholesale electricity market prices using machine learning techniques. Green Finance, 2024, 6 (2), 363-382

<p>The datasets available for open access from the article &lsquo;<em>Ballester, C. and Furi&oacute;, D. (2024). Analysing the impact of renewables on Iberian wholesale electricity market prices using machine learning techniques. Green Finance, 6 (2), 363-382</em>&rsquo; are provided here. This study has been supported by funding from the Spanish Ministry of Science, Innovation, and Universities (Project PGC2018-093645-B-100).</p> <p>The data encompasses the price series of the components of the final wholesale prices, other than the day-ahead market price, at an hourly frequency, from January 2017 to December 2021. In particular, the daily average of the hourly price series of the intraday market, which captures the net effect of the six sessions of the intraday market on the final price, the daily average of the hourly net effect on the final price of the procedure to solve technical constraints, the daily average of the hourly costs resulting from ancillary services and deviation management, the daily average of the hourly costs related to capacity payments and the daily average of the hourly costs associated with the interruptibility service. In addition, we compute the daily average of the hourly series of bids (price and amount) individually submitted by market participants to buy or sell energy, distinguishing between matched and non-matched bids, both in the day-ahead market and in the first session of the intraday market. Other energy-related price series included in the analysis are: the Dutch TTF futures price, the API2 index for the coal price, the EUA futures price, the percentage of hours with 100% use from the France-Spain interconnection, the spread from the France-Spain interconnection, the percentage of water reserves in the reservoirs of the Iberian Peninsula.&nbsp;All shareable data are made available in accordance with open data principles to promote transparency and reproducibility in research. However, there is a specific dataset that we are not authorized to share publicly. Specifically, this includes the data corresponding to the Dutch TTF futures price and the API2 index. Consequently, this dataset is published under restricted access.</p>

opencc-by-4.0Nov 2024View details →
zenodo32/100

Open Data Set for the article Dynamic demand response to electricity prices: Evidence from the Spanish retail market. Utilities Policy, 88 (2024), 101763

<p><span>The datasets available for open access from the article &lsquo;Furi&oacute;, D. and Moreno-del-Castillo, J.<em> Dynamic demand response to electricity prices: Evidence from the Spanish retail market. Utilities Policy, 88 (2024), 101763</em>&rsquo; are provided here.</span></p> <p><span>The datasets comprise time series data on wholesale market global prices and quantity demanded by reference suppliers and competing retailers, along with day-ahead market prices for each of the 24 hours of the day spanning from January 1, 2007, to March 31, 2022. These series have been downloaded from the CNMC website. Additionally, primary temperature data were obtained from the Spanish Meteorological Agency&rsquo;s website. This dataset provided hourly weather information from a nationwide network of weather stations. Given the scope of the demand series data, which reflects the entire Spanish market, the temperature series were constructed to ensure national representativeness. These data were then aggregated to develop comprehensive national temperature series, aligning them with the demand series. <span>Finally, a dummy variable for each day t of the studied period was constructed to capture the business/non-business day effect on electricity demand. It takes the value of 1 if </span></span><span>𝑡</span><span> corresponds to a working day (non-holiday, Monday to Friday) and 0 otherwise. The national holidays considered are as follows: January 1<sup>st</sup>, January 6<sup>th</sup>, May 1<sup>st</sup>, August 15<sup>th</sup>, October 12<sup>th</sup>, November 1<sup>st</sup>, December 6<sup>th</sup>, December 8<sup>th</sup>, and December 25th. Additionally, the corresponding Good Friday for each year included in the study period was also considered.</span></p>

opencc-by-4.0Nov 2024View details →
zenodo32/100

Data set consisting of proteins with associated hydration sites calculated by WATsite

<p>The data set contains 3793 proteins. Each protein is provided as "protein_sasa.pdb", also specifying the SASA value per atom.&nbsp; For each protein, the hydration sites with associated occupancy and thermodynamic properties were calculated using WATsite and are provided in the files "watsite.csv".</p>

opencc-by-4.0Nov 2024View details →
zenodo32/100

E-commerce data set

Open the record for dataset details and reuse information.

opencc-by-4.0Nov 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record