Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
747
datasets available to search
ShareScore release 0.7.1
Dataset results
747 results for “Open Data”
Open Access in Ukraine: characteristics and evolution from 2012 to 2021: supplementary data
<p>This dataset includes metadata for publications authored by researchers affiliated with Ukrainian research institutions and universities, gathered from Scopus, WoS, and Dimensions as of November 2022. Information on the papers' open access status, OA subtypes, publisher details and repository names was obtained from Unpaywall.</p>
Data for: Characteristics of Selected Open Infrastructures, 2024 State of Open Infrastructure Report
<p>The State of Open Infrastructure report provides an annual snapshot of general characteristics for open infrastructures (OIs) listed in Invest in Open Infrastructure’s (IOI) open infrastructure selection tool, Infra Finder (https://infrafinder.investinopen.org/).</p> <p>The data were summarized and reported in the “2024 State of Open Infrastructure Report” section “Characteristics of selected open infrastructures.” The full report is available at https://doi.org/10.5281/zenodo.10934089.</p> <p>A readme, data dictionary, and additional metadata definition file are provided with the dataset with additional detail.</p>
Open Data in German Forest Information Systems: Towards an EU Forest Resilience Monitor (Original dataset on Forest Resilience Indicators and their compliance with Open data criteria)
<p>This dataset represents the original analysis on which my Master's Thesis in the pioneer master programme at the Universities of Münster, Tallinn (Taltech) and Leuven (KUL), titled "Open Data in German Forest Governance: Towards an EU Forest Resilience Monitor". </p> <ul> <li>The first sheet contains the coding on information systems on <strong>bird species occurrence</strong> and their compliance with Open data criteria, with justifications, links or further information speciefied in comments, where necessary. </li> <li>The second sheet contains the coding on information systems on <strong>tree species distribution</strong> and their compliance with Open data criteria, with justifications, links or further information speciefied in comments, where necessary. </li> <li>The third sheet contains the coding on information systems on <strong>soil water conditions</strong> and their compliance with Open data criteria, with justifications, links or further information speciefied in comments, where necessary. </li> <li>The fourth sheet contains the coding on information systems on <strong>canopy cover</strong> and their compliance with Open data criteria, with justifications, links or further information speciefied in comments, where necessary. </li> <li>The fifth sheet contains the coding on information systems on <strong>carbon sequestration</strong> and their compliance with Open data criteria, with justifications, links or further information speciefied in comments, where necessary. </li> <li>The sixth sheet contains the data on the <strong>individual scores per policy level/state per indicator group and the respective averages</strong>. More information on the operationaliation can be found in the methodology section of the thesis. </li> <li>The seventh sheet contains the data on the <strong>individual scores per policy level/state per indicator group and the respective averages, ranked from highest to lowest compliance</strong>. More information on the operationaliation can be found in the methodology section of the thesis. </li> <li>The last sheet gives <strong>information on the coding</strong>. </li> </ul>
A Year of Journal of Open Humanities Data
I asked myself, "What can I learn by applying distant reading computing techniques against a single year of content from the Journal of Open Humanities Data?" In a sentence, I learned a great deal about the Journal, and it very much lives up to is name.
Phenotype variation in Niphargus (Amphipoda: Niphargidae): possible explanations and open challenges: data and R code
<p>Data and R code for performing the analyses of phylogenetic signal presented in the manuscript titled "Phenotype variation in Niphargus (Amphipoda: Niphargidae): possible explanations and open challenges. Data contains phylogenetic tree (Delić et al., 2023) and functional trait data in the RDS format (Premate & Fišer, 2024). The R code is available in the html format.</p> <p>References/data sources:</p> <p>Delić, T., Borko, S., Premate, E., Rexhepi, B., Alther, R., Knuesel, M., ... & Altermatt, F. (2023). Evolutionary origin of morphologically cryptic species imprints co-occurrence and sympatry patterns. <em>bioRxiv</em>, 2023-09.</p> <p>Premate, E., & Fišer, C. (2024). Functional trait dataset of European groundwater Amphipoda: Niphargidae and Typhlogammaridae. <em>Scientific Data</em>, <em>11</em>(1), 188.</p>
Books per Publisher: Data from the Directory of Open Access Books (DOAB)
<p>Quantitative information on publishers and published books from the Directory of Open Access Books (DOAB).</p>
Belmont Forum Open Data Survey 2014
<p>Belmont Forum Open Data Survey, September -- November 2014, by the Open Data working group of the E-Infrastructure and Data Management Collaborative Research Action (http://www.bfe-inf.org/).</p> <p>For the publication of the survey data free text answers have either be lightly edited (question 4, other), or are provided in separate files (questions 8, 15, 16, 17). </p> <p>Publication of the results is under preparation (manuscript submitted). </p>
Categories of Datasets Offered by Current Open Data Portals
<p>This dataset lists categories of open datasets from 40 European open data catalogs. The 40 European data catalogues were taken from four countries (France, Germany, Spain, and the United Kingdom) at a rate of 10 per country. These categories are an indicator of the topics of interests in the respective countries (back in 2016), the types of questions data publishers assume users will ask, and ultimately the types of questions citizens can ask.</p>
ASN 2016 evaluation with open data
<p>This publication contains several datasets that have been used in the paper "Crowdsourcing open citations with CROCI – An analysis of the current status of open citations, and a proposal" submitted to the <a href="https://www.issi2019.org/">17th International Conference on Scientometrics and Bibliometrics (ISSI 2019)</a>, available at <a href="https://arxiv.org/abs/1902.03287">https://arxiv.org/abs/1902.03287</a>.</p> <p>Additional information about the analyses described in the paper, including the code and the data we have used to compute all the figures, is available at <a href="https://github.com/sosgang/asn2016-issi2019">https://github.com/sosgang/asn2016-issi2019</a>. The datasets contain the following information.</p> <p><strong>[asncv | dblp | dblp_asncv]-data.csv:</strong> these CSV files contains all the final data used for the experiment with three conditions as described in the paper. The columns of the CSV file are the following ones (each row represents a particular candidate's application):</p> <ul> <li><em>level:</em> the professor level (1: Full Professor, 2: Associate Professor) to which the candidate wants to get the habilitation;</li> <li><em>session_date:</em> the date of the session for submitting the application;</li> <li><em>journal_number_open:</em> the number of journal article published by the candidate according to the open data available;</li> <li><em>citation_number_open:</em> the number of citations received by candidate's articles according to the open data available;</li> <li><em>h_index_open:</em> the h-index of the candidate according to the open data available;</li> <li><em>journal_number_real:</em> the number of journal article published by the candidate according to the official ASN 2016 data;</li> <li><em>citation_number_real:</em> the number of citations received by candidate's articles according to the official ASN 2016 data;</li> <li><em>h_index_real:</em> the h-index of the candidate according to the the official ASN 2016 data.</li> </ul> <p>All the other CSV files are intermediated data used to calculate the aforementioned one.</p>
Open trade data
<p>Trade-related datasets with a CC-BY or more permissive license</p>
Data supplement to: "Refractory depression - Mechanisms & Efficacy of Radically Open Dialectical Behaviour Therapy (RefraMED): findings of randomised trial on benefits and harms"
<p>Data and code to support primary analyses reported in "Refractory depression – mechanisms and efficacy of radically open dialectical behaviour therapy (RefraMED): findings of a randomised trial on benefits and harms". British Journal of Psychiatry. doi: 10.1192/bjp.2019.53</p>
Open access data for 2-year percutaneous osseointegrated implants - A sheep model
<p>Percutaneous osseointegrated <strong>(OI)</strong> devices for amputees are metallic endoprostheses, surgically implanted into the residual bone that protrude through the skin, allowing attachment of an exoprosthetic. In contrast to standard socket-type systems, these percutaneous OI devices can provide an improved prosthetics attachment platform. However, bone adaptations, which include atrophy and/or hypertrophy along the extent of the host bone-endoprosthetic interface, are known clinical outcomes and are dependent upon the load transfer region of the device to the host bone. The goal of this study was to determine if a percutaneous OI device, designed with a porous coated distal region and a collar, could promote and maintain stable bone attachment. A total of eight, 18 to 24-month old, mixed-breed sheep were surgically implanted with a percutaneous OI device. For 24-months, animals were allowed to bear weight as tolerated and monitored for signs of bone remodelling. At necropsy, the endoprosthesis and the surrounding tissues were harvested, radiographically imaged, and histomorphometrically analyzed to determine the periprosthetic bone adaptation in five animals. Bone growth into the porous coating was achieved in all five animals. Serial radiographic data showed stress-shielding related bone adaptation based on the placement of the endoprosthetic stem. When collar placement achieved end-bearing against the transected bone, distal bone conservation/hypertrophy was observed. The results supported the use of distally porous coated percutaneous OI devices for distal load-transfer and host bone maintenance.</p>
Source data belonging to "Visualisation of dCas9 target search in vivo using an open-microscopy framework"
<p>Source data corresponding to "Visualisation of dCas9 target search <em>in vivo</em> using an open-microscopy framework". Contains pTarget and pNonTarget raw datasets, as well as all localization data, cell UV intensity data, cell outline data, and analysed diffusion coefficient lists.</p>
Open Research Data in Medicine - Polish scientists' attitudes towards data sharing
<p>The survey on the attitudes and beliefs of research staff has been carried out at selected Polish medical universities. The purpose of the questionnaire was to collect respondents' opinions on opening research data created during their scientific work. The research was aimed at preparing the necessary educational, technical and legal support for scientists after launching the Polish Medical Platform, when Polish scientists will be asked to deposit their data in local repositories.</p>
6TiSCH Open Data Action Datasets
<p>This repository contains the datasets of the 6TiSCH Open Data Action experiment of Fed4FIRE. The datasets cover the execution of three application-level test scenarios, building-automation, home-automation and industrial-monitoring, on two testbeds, w-iLab.t in Ghent and OpenTestbed in Paris. The reference firmware image used was OpenWSN. The data format is documented at https://benchmark.6tis.ch/.</p>
Tutorial Weather Data Cutouts for PyPSA-Eur: An Open Optimisation Model of the European Transmission System
<p><strong>PyPSA-Eur</strong> is an open model dataset of the European power system at the transmission network level that covers the full ENTSO-E area. It can be built using the code provided at <a href="https://github.com/PyPSA/PyPSA-eur">https://github.com/PyPSA/PyPSA-eur</a>.</p> <p><strong>It contains</strong> alternating current lines at and above 220 kV voltage level and all high voltage direct current lines, substations, an open database of conventional power plants, time series for electrical demand and variable renewable generator availability, and geographic potentials for the expansion of wind and solar power.</p> <p><strong>Not all data dependencies</strong> are shipped with the <a href="https://github.com/PyPSA/PyPSA-eur">code repository</a>, since git is not suited for handling large changing files. Instead we provide separate <strong>data bundles and cutouts</strong> to be downloaded and extracted as noted in the <a href="https://pypsa-eur.readthedocs.io/en/latest/installation.html">documentation</a>.</p> <p>The provided lightweight <strong>cutouts </strong>are spatiotemporal subsets of the German weather data from the <a href="https://software.ecmwf.int/wiki/display/CKB/ERA5+data+documentation">ECMWF ERA5</a> reanalysis dataset for March 2013 to be used for the <a href="https://pypsa-eur.readthedocs.io/en/latest/tutorial.html">PyPSA-Eur tutorial</a>. They have been prepared by and are for use with the <a href="https://github.com/PyPSA/atlite">atlite</a> tool (<a href="https://atlite.readthedocs.io/">https://atlite.readthedocs.io/</a>).</p> <p><strong>ECMWF ERA5</strong></p> <ul> <li><strong>Source: </strong><a href="https://cds.climate.copernicus.eu/cdsapp#!/dataset/reanalysis-era5-single-levels?tab=overview">https://cds.climate.copernicus.eu/cdsapp#!/dataset/reanalysis-era5-single-levels?tab=overview</a></li> <li><strong>Terms of Use: </strong><a href="https://cds.climate.copernicus.eu/api/v2/terms/static/20180314_Copernicus_License_V1.1.pdf">https://cds.climate.copernicus.eu/api/v2/terms/static/20180314_Copernicus_License_V1.1.pdf</a></li> </ul>
The Politecnico di Torino rolling bearing test rig: description of the open-access data for vibration monitoring and diagnostics
<p>Accelerometric measurements from the rolling bearing test rig of the Dynamic and Identification Research Group (DIRG), Department of Mechanical and Aerospace Engineering, Politecnico di Torino.</p> <p>Goals:</p> <p> • Vibration Monitoring, Bearing Diagnostics, Damage detection, Damage localization, Damage classification, Damage assessment.</p> <p>Features:</p> <p> • high-speed spindle driving a hollow shaft supported by a couple of identical roller bearings B1 and B3. B1 is the bearing under analysis and features various damages.</p> <p> • two damage types (indentations on a roller and on the inner ring) and severities (0, 150, 250, 450 µm).</p> <p> • a central, larger roller bearing (B2) is loaded through a sledge generating a controlled radial force measured by a load cell.</p> <p> • lubrication is obtained by oil injection into the hollow shaft.</p> <p> • a K-type thermocouple is used to monitor the temperature (manually recorded).</p> <p> • two triaxial accelerometers are mounted on the supports of bearings B1 and B2.</p> <p>Dataset:</p> <p> • stationary acquisitions at different speed & load combinations (speed: 0, 100, 200, 300, 400, 500 Hz; load: 0, 1000, 1400, 1800 N).</p> <p> • endurance acquisitions of the bearing featuring the 450µm roller indentation. Monitoring of the damage evolution for about 230 hours under the same speed and load condition.</p> <p> </p> <p>The extended description of the dataset can be found in the attached pdf "Description and analysis of open access data" or in:</p> <p>A.P. Daga, A. Fasana, S. Marchesiello, L. Garibaldi, The Politecnico di Torino rolling bearing test rig: Description and analysis of open access data, Mechanical Systems and Signal Processing 120 (2019) 252–273. doi:10.1016/j.ymssp.2018.10.010.</p>
Galvanising the Open Access Community: A Study on the Impact of Plan S - Data and Code
<div> <div> <div> <div> <p>This repository contains the datasets and code underpinning <em>Chapter 3 "Counterfactual Impact Evaluation of Plan S"</em> of the report <em>"Galvanising the Open Access Community: A Study on the Impact of Plan S"</em> commissioned by the cOAlition S to scidecode science consulting.</p> <p>Two categories of files are part of this repository:</p> <p><strong>1. Datasets <br></strong></p> <p>The 21 CSV source files contain the subsets of publications funded by the funding agencies that are part of this study. These files have been provided by <em>OA.Works</em>, with whom scidecode has collaborated for the data collection process. Data sources and collection and processing workflows applied by <em>OA.Works</em> are described on their website and specifically at <a href="https://about.oa.report/docs/data">https://about.oa.report/docs/data</a>.</p> <p>The file "plan_s.dta" is the aggregated data file stored in the format ".dta", which can be accessed with <em>STATA </em>by default or with plenty of programming languages using the respective packages, e.g., <em>R</em> or <em>Python</em>. </p> <p><strong>2. Code files</strong></p> <p>The associated code files that have been used to process the data files are:</p> <pre> - data_prep_and_analysis_script.do<br> - coef_plots_script.R</pre> <p>The first file has been used to process the CSV data files above for data preparation and analysis purposes. Here, data aggregation and data preprocessing is executed. Furthermore, all statistical regressions for the ounterfactual impact evaluation are listed in this code file. The second code file "coef_plots_script.R" uses the computed results of the counterfactual impact evaluation to create the final graphic plots using the <em>ggplot2 </em>package.</p> <p>The first ".do" file has to be run in STATA, the second one (".R") requires the use of an integrated development environment for R. </p> Further Information are avilable in the final report and via the followng URLs:<br> <pre><a href="https://www.coalition-s.org/">https://www.coalition-s.org/</a> <a href="https://scidecode.com/">https://scidecode.com/</a> <a href="https://oa.works/">https://oa.works/</a> <a href="https://openalex.org/">https://openalex.org/</a><br><a href="https://sites.google.com/view/wbschmal">https://sites.google.com/view/wbschmal</a> </pre> </div> </div> </div> </div>
Enhancing multi-mode transport emission inventories: combining open-source data with traditional approaches
<p>The primary goal of this dataset is to enhance the spatial and temporal distribution of emissions from civil aviation (NFR1.A.3.a), road transport (NFR1.A.3.b), railways (NFR1.A.3.c), and military aviation (NFR1.A.5), using Portugal as case study. For more information, please refer to the published article “Enhancing multi-mode transport emission inventories: combining open-source data with traditional approaches” (<a title="Persistent link using digital object identifier" href="https://doi.org/10.1016/j.uclim.2024.102097" target="_blank" rel="noreferrer noopener"><span><span>https://doi.org/10.1016/j.uclim.2024.102097</span></span></a>). This dataset contains the following folders and files:</p> <p><strong>1. Spatial_Location</strong></p> <p> 1.1. NFR1_A_3_a.gdb: Geodatabase containing the locations of Portuguese airports and aerodromes.</p> <p> 1.2. NFR1_A_3_b.gdb: Geodatabase containing the locations of Portuguese roads.</p> <p> 1.3 NFR1_A_3_c.gdb: Geodatabase containing non-electrified Portuguese railways and train station locations.</p> <p> 1.4 NFR1_A_5.gdb: Geodatabase containing the locations of Portuguese military airport facilities.</p> <p><strong>2. Temporal_Profiles</strong></p> <p><em> 2.1. Daily</em></p> <p> 2.1.1. Daily_NFR1_A_3_a.csv: This csv file contains the daily movements profiles of civil aviation sites in Portugal.</p> <p><em> 2.2. Hourly</em></p> <p> 2.2.1. Hourly_NFR1_A_3_b.txt: This txt file contains the hourly road traffic volume profiles for the road transport activities in Portugal at different locations (BigAir column).</p> <p> 2.2.2. Hourly_NFR1_A_3_c.txt: This text file contains the hourly railway profile in Portugal, categorized by line and train station.</p> <p><strong>3. Emission_Factors</strong></p> <p> 3.1. EF_NFR1_A_3_a.xlsx: This Excel file contains emission factors for civil aviation activities, categorized by technology, flight phase, fuel, and pollutant. Additionally, it includes information about engines and aircraft.</p> <p> 3.2. EF_NFR1_A_3_b.xlsx: This Excel file contains emission factors for road transport activities, categorized by vehicle type, technology, fuel, abatement, and pollutant. Emission factors for road resuspension are not provided because the papers using this dataset are still under review.</p> <p> 3.3. EF_NFR1_A_3_c.xlsx: This Excel file contains emission factors for railways activities, categorized by technology, fuel, and pollutant.</p> <p> 3.4. EF_NFR1_A_5.xlsx: This Excel file contains emission factors for military aviation activities, categorized by fuel, and pollutant.</p> <p><strong>4. Other_Info</strong></p> <p> 4.1 NFR1_A_3_b: This folder contains information organized by road segments, including fuel consumption (in the “FuelConsumption” folder), hourly meteorology (in the “Meteorology” folder), population data (in the “Population” folder), daily traffic volume (in the “TrafficVolume” folder), vehicle categories (in the “VehicleCategory” folder), and vehicle classes (in the “VehicleClasses” folder). Additionally, it includes the link between road traffic volume measurement points and the Portuguese road network (in the “sensorsVSroads” folder)</p>
Data curation materials in "Daily life in the Open Biologist's second job, as a Data Curator"
<p>This is the supplementary material accompanying the manuscript "Daily life in the Open Biologist’s second job, as a Data Curator", published in <a href="https://doi.org/10.12688/wellcomeopenres.22899.1">Wellcome Open Research</a>. </p> <p>It contains:</p> <p><strong>- Python_scripts.zip</strong>: Python scripts used for data cleaning and organization:</p> <p> -add_headers.py: adds specified headers automatically to a list of csv files, creating new output files containing a "_with_headers" suffix.</p> <p> -count_NaN_values.py: counts the total number of rows containing null values in a csv file and prints the location of null values in the (row, column) format.</p> <p> -remove_rowsNaN_file.py: removes rows containing null values in a single csv file and saves the modified file with a "_dropNaN" suffix.</p> <p> -remove_rowsNaN_list.py: removes rows containing null values in list of csv files and saves the modified files with a "_dropNaN" suffix.</p> <p><strong>- README_template.txt</strong>: a template for a README file to be used to describe and accompany a dataset. </p> <p><strong>- template_for_source_data_information.xlsx</strong>: a spreadsheet to help manuscript authors to keep track of data used for each figure (e.g., information about data location and links to dataset description).</p> <p><strong>- Supplementary_Figure_1.tif</strong>: Example of a dataset shared by us on Zenodo. The elements that make the dataset FAIR are indicated by the respective letters. Findability (F) is achieved by the dataset unique and persistent identifier (DOI), as well as by the related identifiers for the publication and dataset on GitHub. Additionally, the dataset is described with rich metadata, (e.g., keywords). Accessibility (A) is achieved by the ease of visualization and downloading using a standardised communications protocol (https). Also, the metadata are publicly accessible and licensed under the public domain. Interoperability (I) is achieved by the open formats used (CSV; R), and metadata are harvestable using the Open Archives Initiative Protocol for Metadata Harvesting (OAI-PMH), a low-barrier mechanism for repository interoperability. Reusability (R) is achieved by the complete description of the data with metadata in README files and links to the related publication (which contains more detailed information, as well as links to protocols on protocols.io). The dataset has a clear and accessible data usage license (CC-BY 4.0).</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.