Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
339
datasets available to search
ShareScore release 0.9.0
Dataset results
339 results for “fairness”
EOSC Task Force on FAIR Metrics and Data Quality: FAIR Evaluation community survey 2023
<p>The EOSC-A FAIR Metrics and Data Quality Task Force (TF) supported the European Open Science Cloud Association (EOSC-A) by providing strategic directions on FAIRness (Findable, Accessible, Interoperable, and Reusable) and data quality. The Task Force conducted a survey using the <a href="https://ec.europa.eu/eusurvey/">EUsurvey tool</a> between 15.11.2022 and 18.01.2023, targeting both developers and users of FAIR assessment tools. The survey aimed at supporting the harmonisation of FAIR assessments, in terms of what it evaluated and how, across existing (and future) tools and services, as well as explore if and how a community-driven governance on these FAIR assessments would look like. The survey received 78 responses, mainly from academia, representing various domains and organisational roles. This is the anonymised survey dataset in csv format; most open-ended answers have been dropped. The codebook contains variable names, labels, and frequencies.</p>
Dataset Dental research data availability and quality according to FAIR principles
<p>This dataset contains open access publications in EPMC dental journals from 2016 to 2021 and 500 non-open access dental publications. We evaluated the level of compliance with the FAIR principles. The original dataset and codebook are attached. </p>
FAIR Data Package of a Tribological Showcase Pin-on-Disk Experiment
<p>To assess the feasibility of producing FAIR data via the integration of a controlled vocabulary, an ontology, and an ELN, this dataset demonstrates the implementation of a tribological experiment while accounting for as many details as possible. The showcase experiment had a lubricated pin-on-disk arrangement, ran at 15 N normal load and a velocity range of 20 to 170 mm/s. With this dataset, we hope to provide a possible blueprint for FAIR data publication in experimental tribology.</p> <p><a href="http://www.nature.com/articles/s41597-022-01429-9">https://www.nature.com/articles/s41597-022-01429-9</a> - Garabedian, N.T., Schreiber, P.J., Brandt, N., Greiner, C., et al.</p> <p>Quick start with the dataset in README.txt (<em>included in the newest version of the dataset</em>)</p> <p>Abstract: Generating FAIR research data in experimental tribology. Sci Data 9, 315 (2022). Digital solutions for the generation of FAIR (Findable, Accessible, Interoperable and Reusable) data and metadata in experimental tribology are currently lacking, despite the looming challenge of integrating cutting-edge data science techniques – a promising scientific route for any field that often relies on phenomenology and empiricism. Additionally, the broad interdisciplinarity of tribology is probably a main contributing factor for the lack of community-wide data and metadata standards, and the heavy reliance on custom workflows and equipment. This paper, first, outlines a sample framework for scalable generation of FAIR data, and second, delivers a showcase FAIR data package for a pin-on-disk tribological experiment. The resulting curated data, consisting of 2,008 key-value pairs and 1,696 logical axioms, is the result of (1) the close collaboration with developers of a virtual research environment, (2) crowd-sourced controlled vocabulary, (3) ontology building and (4) numerous – seemingly – small-scale digital tools. Thereby, this paper demonstrates a collection of scalable non-intrusive techniques that extend the life, reliability and reusability of experimental tribological data beyond typical publication practices.</p> <p><a href="http://youtu.be/xwCpRDnPFvs">https://youtu.be/xwCpRDnPFvs</a> - Generating FAIR Research Data in Experimental Tribology - Get Scientific Results Ready for ML</p> <p><a href="https://doi.org/10.5281/zenodo.5720626">https://doi.org/10.5281/zenodo.5720626</a> - FAIR Data Package of a Tribological Showcase Pin-on-Disk Experiment</p> <p><a href="https://doi.org/10.5281/zenodo.5720198">https://doi.org/10.5281/zenodo.5720198</a> or <a href="https://github.com/nick-garabedian/TriboDataFAIR-Ontology">https://github.com/nick-garabedian/TriboDataFAIR-Ontology</a> or <a href="https://fairsharing.org/3597">https://fairsharing.org/3597</a> - TriboDataFAIR Ontology</p> <p><a href="https://doi.org/10.5281/zenodo.5720218">https://doi.org/10.5281/zenodo.5720218</a> or <a href="https://github.com/nick-garabedian/SurfTheOWL">https://github.com/nick-garabedian/SurfTheOWL</a> - SurfTheOWL</p> <p><a href="https://kadi4mat.iam-cms.kit.edu/">https://kadi4mat.iam-cms.kit.edu/</a> - Kadi4Mat Virtual Research Environment and Electronic Lab Notebook </p>
Four Essential Components for FAIR Data: Capability & Category-Specific Requirements
<p>Adapted from Bailo (2019) and Peng (2023), this diagram illustrates FAIR requirements specific to data, metadata, and infrastructure, aligned with the definitions of individual FAIR principles. It highlights the critical role of enterprise capabilities—including processes, systems, standards, tools, and skills—in supporting FAIR data. These four components are essential for systematically enhancing the overall FAIRness of an organization's scientific data collection</p> <p> </p>
[Supplementary Information] Can LCA be FAIR? – Assessing the status quo and opportunities for FAIR data sharing
<p>This is the supplementary information related to a the manuscript - 'Can LCA be FAIR?' - Assessing the status quo and opportunities for FAIR data sharing. The purpose of this study is to assess the status quo of data sharing in LCA in relation to the FAIR data principles (Findability, Accessibility, Interoperability and Re-use).</p><p>The supplementary information consists of three files:</p><p><strong>SI 1</strong> - How the life cycle inventory is shared in relation to the FAIR data principles in 25 peer reviewed LCA journal articles between 2018 -2022.</p><p><strong>SI 2</strong> - Review of ten data management plans of EU Horizon Europe projects in relation to LCA to assess the recommendations on the implementation of FAIR principles.</p>
Data from: A FAIR and modular image-based workflow for knowledge discovery in the emerging field of imageomics
<p>Data and results from the Imageomics Workflow. These include data files from the Fish-AIR repository (https://fishair.org/) for purposes of reproducibility and outputs from the application-specific imageomics workflow contained in the Minnow_Segmented_Traits repository (https://github.com/hdr-bgnn/Minnow_Segmented_Traits).</p> <p>Fish-AIR:<br> This is the dataset downloaded from Fish-AIR, filtering for Cyprinidae and the Great Lakes Invasive Network (GLIN) from the Illinois Natural History Survey (INHS) dataset. These files contain information about fish images, fish image quality, and path for downloading the images. The data download ARK ID is dtspz368c00q. (2023-04-05). The following files are unaltered from the Fish-AIR download. We use the following files:</p> <p>extendedImageMetadata.csv: A CSV file containing information about each image file. It has the following columns: ARKID, fileNameAsDelivered, format, createDate, metadataDate, size, width, height, license, publisher, ownerInstitutionCode. Column definitions are defined https://fishair.org/vocabulary.html and the persistent column identifiers are in the meta.xml file.</p> <p>imageQualityMetadata.csv: A CSV file containing information about the quality of each image. It has the following columns: ARKID, license, publisher, ownerInstitutionCode, createDate, metadataDate, specimenQuantity, containsScaleBar, containsLabel, accessionNumberValidity, containsBarcode, containsColorBar, nonSpecimenObjects, partsOverlapping, specimenAngle, specimenView, specimenCurved, partsMissing, allPartsVisible, partsFolded, brightness, <br> uniformBackground, onFocus, colorIssue, quality, resourceCreationTechnique. Column definitions are defined https://fishair.org/vocabulary.html and the persistent column identifiers are in the meta.xml file.</p> <p>multimedia.csv: A CSV file containing information about image downloads. It has the following columns: ARKID, parentARKID, accessURI, createDate, modifyDate, fileNameAsDelivered, format, scientificName, genus, family, batchARKID, batchName, license, source, ownerInstitutionCode. Column definitions are defined https://fishair.org/vocabulary.html and the persistent column identifiers are in the meta.xml file.</p> <p>meta.xml: A XML file with the metadata about the column indices and URIs for each file contained in the original downloaded zip file. This file is used in the fish-air.R script to extract the indices for column headers.</p> <p>The outputs from the Minnow_Segmented_Traits workflow are:</p> <p>sampling.df.seg.csv: Table with tallies of the sampling of image data per species during the data cleaning and data analysis. This is used in Table S1 in Balk et al. </p> <p>presence.absence.matrix.csv: The Presence-Absence matrix from segmentation, not cleaned. This is the result of the combined outputs from the presence.json files created by the rule “create_morphological_analysis”. The cleaned version of this matrix is shown as Table S3 in Balk et al.</p> <p>heatmap.avg.blob.png and heatmap.sd.blob.png: Heatmaps of average area of biggest blob per trait (heatmap.avg.blob.png) and standard deviation of area of biggest blob per trait (heatmap.sd.blob.png). These images are also in Figure S3 of Balk et al.</p> <p>minnow.filtered.from.iqm.csv: Filtered fish image data set after filtering (see methods in Balk et al. for filter categories).</p> <p>burress.minnow.sp.filtered.from.iqm.csv: Fish image data set after filtering and selecting species from Burress et al. 2017.</p>
MARCSI - Inventory of Marine Citizen Science Initiatives and the FAIRness of the data they produce
<p>Inventory (data set) of Marine Citizen Science Intiatives collected and described in the publication entitled "Past and present marine citizen science around the globe: a cumulative inventory of initiatives and data produced" co-authored by Uta Wehn, Ane Bilbao, Luke Somerwill, Torsten Linders, Joan Maso, Stephen Parkinson, Christina Semasingha,<sup> </sup>Sasha Woods.</p>
OpenAIRE and FAIR Data Expert Group survey about Horizon 2020 template for Data Management Plans
<p>This dataset is published in 2017 by the OpenAIRE project and the FAIR Data Expert Group.</p> <p>It contains two survey data files, two pdf-files summarising the results in a report and an infographic, and a Readme.txt file.</p> <p>The OpenAIRE project supports the open science ambitions of the European Commission. The project and in particular the Research Data Management team provide support, training and information on the Open Research Data Pilot. In this context, a survey was carried out to collect feedback on the Horizon 2020 template for Data Management Plans (DMPs). The team collaborated with the FAIR data expert group, which is providing recommendations to the European Commission on turning FAIR data into reality. One of the specific tasks of the Expert Group is contributing to an evaluation of the Horizon 2020 approach to DMPs, including future revisions of the template and the development of additional sector/ discipline-specific guidance. The aim of the survey was to collect experiences of researchers and DMP reviewers with the DMP template and guidelines on FAIR data management in Horizon 2020. The survey assesses the usefulness of the guidelines and any aspects that are confusing and unclear to determine what improvements can be made.</p> <p>Feedback was sought from both researchers and research support staff. The survey was initially scheduled to run from 22 May to 21 June 2017. Several organisations were asked to help announce the survey, including OpenAIRE’s National Open Access Desks, the FAIR data expert group, FOSTER, LIBER, and the RDA Interest Group on Active DMPs. When the first survey responses showed only a small share of researchers, more stakeholders were contacted to specifically target this community. The European Research Area was approached, whose project officers circulated the survey call among award holders of EC projects. Early-career researchers were also informed through the YEAR network and EURODOC. This resulted in an extension of the survey to 21 July 2017.</p> <p>At the close of the survey on 21 July 2017, a total number of 289 responses were reached. 50% of the respondents indicated that they were researchers, and 60% that they were (also) research support staff. OpenAIRE and the FAIR data expert group are very pleased with this balanced outcome and would like to thank all colleagues and organisations who promoted the survey, as well as everyone who took part in it.</p> <p> </p>
Dissecting the FAIR Guiding Principles - Key Categories, Core Concepts, Focus Elements, and Harmonized Indicators
<p>A comprehensive workbook created to facilitate and document the process of decomposing the FAIR Guiding Principles and mapping them to key categories, requirements, core concepts, focus elements, and harmonized indicators. It also contains a complete list of the indicators.</p>
Example files to A FAIR archive based on the CERIF model
<p>An archival structure based on the CERIF model is proposed. The archive tree is represented by cfProjects and the archived objects by cfResult* entities with their descriptive metadata given in attached CERIF entities. Archival preservation metadata is stored in the Premis format inside attached cfMeasurment entities. An example in which EPrints repository items are transferred to the archive is presented. When CERIF is employed in relevant archive processes, a FAIR compliant archive is easier to achieve. </p>
Stocktaking GO FAIR Discovery IN - Use cases, infrastructure
<p>In order to build a better ecosystem for data discovery tools the Data Discovery Implementation Group of GO Fair (https://www.go-fair.org/implementation-networks/overview/discovery) collected use cases between 2019 and 2020 from a variety of sources. We also detail the ‘Actors’ for these use cases and the ‘Source’ providing links, whenever possible. Since we found over a hundred individual use cases, we decided to cluster them to provide a better overview. The clustering, as well as the results of a small survey among data infrastructure specialists to find how they rate the importance of the clusters are detailed in the documentation to this dataset, a draft of which can currently be found <a href="https://docs.google.com/document/d/1sq78eCFYgmWcMFYcbNonA2KrkUO1qdGr7d49tHuRcRM/edit?usp=sharing">here</a>. The code and data to produce the figures in the documentation are available as R code in the GO_FAIR_Discovery_Use_case-master.zip file. The use cases themselves are available as Excel sheet and csv. </p>
Atom probe tomography nomad-FAIR demonstrator dataset R76-23219-v01.epos.apth5
<p>This is the dataset of an atom probe tomography experiment which is provided open source for testing the possibility of implementing an open source encyclopedia for experimental materials science datasets, including techniques to begin with such as Scanning Transmission Electron Microscopy (STEM), Multidimensional Photo Emission Spectroscopy (MPES), and Atom Probe Tomography (APT) / Field Ion Microscopy (FIM).</p> <p><strong>This repository serves three aims:</strong></p> <p>1. The dataset is of scientific interest. Specifically, it captures the result of a cutting-edge APT experiment detailed exemplarily in DOI: 10.1038/s41467-018-03115-0 (Fig. 6a "Se+Na2Se treatment") by Torsten Schwarz and coworkers.</p> <p>2. The dataset contributes to tests of an extension to "The NOMAD Laboratory" (https://nomad-coe.eu/): nomad-FAIR. Specifically, to test various aspects of an automatized metadata parsing and processing pipeline to enable the extraction of domain-specific JSON metadata files into a NOMAD-conformant JSON file, ultimately aiming for searchable and repurposable dataset documentation. This serves two purposes: on the one hand to contextualize each dataset within NOMAD. On the other hand to serve as a starting point to parse potential interesting content from the heavy data HDF5 file to reduce unnecessary file access.<br> The implementation of nomad-FAIR is coordinated by Markus Scheidgen.<br> The APT domain-specific parser is developed by Markus Kühbach.</p> <p>3. The dataset constitutes further a test of an open format specification for storing atom probe tomography data using the Hierarchical Data Format (HDF5). This is a recent initiative of the International Field Emission Society's (IFES) atom probe tomography technical committee. In this repository it is detailed an exemplar proposal of how to store acquisition-side relevant results and context of an APT experiment into a HDF5 file and complementary metadata files such as JSON. Implementation of this HDF5-based storage solution for APT data is lead by Markus Kühbach.</p> <p><br> <strong>The organization of this repository with respect to above aims is as follows:</strong></p> <p>-The original EPOS file of the measured is contained in the compressed *.epos.tar.gz archive.</p> <p>-The *.apth5 file is a transcoded version of the EPOS file. Therein, x,y,z data columns are stripped.</p> <p>-The correspondingly named *.json file is the file which nomad-FAIR parses metadata from.</p> <p>-Other files constitute logs of the transcoding process.</p> <p><br> <strong>Funding:</strong><br> The work was partially supported by BiGmax, the Max Planck Society's Research Network on Big-Data-Driven Materials-Science.</p>
FAIR Data: just data done right
<p>An aphorism about FAIR Data, in graphical form, inspired by "Sticker open science: just science done right": Melanie Imming, & Jon Tennant. (2018). Sticker open science: just science done right (ENG). Zenodo. https://doi.org/10.5281/zenodo.1285575 <br> <br> . </p> <p> </p>
Figure data and code used in Technical comment on "Fairness considerations in global mitigation investments"
<p>The package contains the data and code to create the figure in the associated technical comment in Science published at <a href="https://www.science.org/doi/10.1126/science.adg5893">https://www.science.org/doi/10.1126/science.adg5893</a></p>
FAIR-CHO Citation Model Zotero Group Library Bibliography. Supplementary material
<p>This is a selected bibliography created during the project <em>A FAIR-enabling citation model for Cultural Heritage Objects</em>.</p> <p>This bibliography has been set up via a Zotero Library Group, by organizing it into subject folders representative of the project content. An initial list of descriptors was also defined to 'semantically' label the bibliographic references as they were collected.</p> <p>The dataset represents the Zotero Library on 1st August 2023.</p> <p>The dataset is published in .csv and .ris formats.</p> <p>See also on Zotero Groups: <a href="https://www.zotero.org/groups/4883319/cho_citation_model/library">https://www.zotero.org/groups/4883319/cho_citation_model/library</a>.</p>
Data: Algorithms for new types of fair stable matchings
<p>This data corresponds to the data and experiments described in Section 5 of<br> the following paper:</p> <p>Algorithms for new types of fair stable matchings<br> Authors: Frances Cooper and David Manlove</p> <ul> <li>The paper is located at: <a href="https://arxiv.org/abs/2001.10875">https://arxiv.org/abs/2001.10875</a></li> <li>The software is located at: <a href="https://zenodo.org/record/3630383">https://zenodo.org/record/3630383</a></li> <li>The data is located at: <a href="https://zenodo.org/record/3630349">https://zenodo.org/record/3630349</a></li> </ul> <p>See the README for more information.</p>
Dataset do DH2020 [The Lusophone Digital Humanities and What they (we) are doing from the South: textual corpus analysis and FAIR principles to tackle Hegemony]
<p>Planilha de dados recuperados do Google Scholar utilizado na análise e apresentação da pesquisa empírica intitulada - <strong>The Lusophone Digital Humanities and What they (we) are doing from the South: textual corpus analysis and FAIR principles to tackle Hegemony </strong>- no evento <strong>DH2020 Ottawa</strong>: <a href="https://hcommons.org/deposits/item/hc:32051/">https://hcommons.org/deposits/item/hc:32051/</a></p>
Collected recommendations and requirements for FAIR-enabling services
<p>Within FAIRsFAIR task 2.4, we carried out a structured literature review to extract requirements, recommendations and other desiderata for FAIR-enabling services. This document contains the full list of excerpts, structured and annotated. This work has been used as input for the basic framework on FAIRness of services developed by FAIRsFAIR task 2.4 (see https://doi.org/10.5281/zenodo.4292599).</p>
Analysed data from interviews on FAIR-enabling services
<p>Within FAIRsFAIR task 2.4, we carried out five semi-structured interviews with data service owners to understand how services currently support the FAIR principles, what are transferable insights and recommendations, and what are common challenges and pitfalls. This document collects the insights captured from the interviews. This work has been used as input for the basic framework on FAIRness of services developed by FAIRsFAIR task 2.4 (see https://doi.org/10.5281/zenodo.4292599).</p> <p> </p>
Supplementary data files for manuscript titled "From spreadsheet lab data templates to knowledge graphs: A FAIR data journey in the domain of AMR research"
<div>This data repository contains all the necessary supplementary files for the manuscript titled "<strong>From spreadsheet lab data templates to knowledge graphs: A FAIR data journey in the domain of AMR research.</strong>"</div> <div> </div> <div>The repository is a copy of the <a href="https://github.com/IMI-COMBINE/template2graphs">GitHub page</a> with the source code used to generate the graph and additional files required for the Lab Data Template.</div> <div> </div> <div>Below we provide a brief overview of the data files in the `additional folder` and their underlying purpose:</div> <div> <ul> <li>The <strong>Data Survey</strong> collects relevant project and data set information to set up a Data Management Plan. It can serve as an input for Lab Data Template development.</li> <li>The <strong>Lab Data Templates</strong> facilitate the collection of AMR research data (in vivo and in vitro) in several sub-tables. The Excel format is compatible with upload procedures into the data repository 'grit' and serves as input for a knowledge graph workflow.</li> <li>The <strong>Data dictionary</strong> is connected to the Lab Data Templates and ensures harmonized data entries. In addition, the dictionaries collect metadata beyond the content of the Lab Data Template (e.g. bacterial strain information or compound information) and link to ontologies where possible.</li> <li>The <strong>FAIR assessments</strong> have been used as a primer for improving the template. This report is generated using the FAIR-DSM model.</li> </ul> </div> <div>The templates have been used during the IMI2 GNA NOW project to collect information and have been improved according to FAIR standards in collaboration with the IMI FAIRplus project ("post FAIRification").</div>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.