Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
410
datasets available to search
ShareScore release 0.9.0
Dataset results
410 results for “Data Repositories”
LocoD Data Repository
<p>We publish this repository to offer a common data set to conduct comparisons between different methods. Furthermore, if access to equipment, facilities, and/or research participants is not possible, then this repository facilitates testing of the preliminary algorithms. The recorded signal from EMG, IMU, and pressure sensor, along with important information such as tags and recording properties, has been saved in a structure.</p> <p>Our Data Repository consists of data from 8 Female and 7 male subjects and none of them had prior experience with LocoD.</p> <p>The recorded data corresponds to one recording per participant digitalized at 2 kHz. Data includes 8 EMG (Delsys) channels, 3 IMUs (Delsys), and one pressure sensor (Delsys).</p> <p>Data consists of 30 trials of our circuit. Our circuit contains terrains for walking, stair ascent, stair descent, ramp ascent, and ramp descent. Data were tagged when the subjects started a terrain manually by an operator.</p> <p>SEMG electrodes were placed on the semitendinosus, biceps femoris, tensor fasciae latae, rectus femoris, vastus lateralis, vastus medialis, and gracilise. These muscles were found using palpation by an experienced physiotherapist and were selected based on a literature search for the most common muscle signals used to control lower limb prosthetics.</p> <p>IMUs were placed above the knee, below the knee, and on the foot to get all the joint orientations during different movements.</p> <p>A Pressure sensor was built into an insole used by each research participant.</p> <p>Participants were instructed to enter each terrain, such as stairs or ramps with their sensorized legs. These different locomotion modes were selected as they are the most common movements in daily life.</p>
Fermi-LAT data for crab-multi-instrument-systematics repository
Open the record for dataset details and reuse information.
Data Repository for Combined Experimental and Computational Study of the Reactivity of the Methanimine Radical Cation (H2CNH·+) and Its Isomer Aminomethylene (HCNH2·+) with Propene (CH3CHCH2)
Open the record for dataset details and reuse information.
Data Repository: Direct electron beam writing of silver using a β-diketonate precursor: first insights
<h2>Summary</h2> <p>The data is contained in a single zip file with the two main folders: "Tungsten_SEM_deposition" and "FESEM_deposition". </p> <p>The folder "Tungsten_SEM_deposition" contains all used data from deposition experiments in the Hitachi S3600 tungsten filament scanning electron microscope (SEM) using the precursor (hfac)AgPMe3 with the home-built gas-injection system.</p> <p>The folder "FESEM_deposition" contains all used data from deposition in the Zeiss Crossbeam 340 KMAT using the field emitter scanning electron microscope (FESEM) capability of the dual beam instrument using the precursor (hfac)AgPMe3 with the commercial gas-injection system (Kleindiek).</p> <p>In each folder all raw SEM images related to these experiments are provided. In addition all pattern files and the most important data on the microstructural characterization using transmission electron microscopy (TEM) and energy-dispersive X-ray (EDX) spectroscopy are provided and indicated be the corresponding folder names.</p> <h3><br>Folder structure: "Tungsten_SEM_deposition"</h3> <p>1) KH157_new_Si_Ag_hfacAgTMP_deposition: images of the deposition experiment taken in the tungesten filament microscope</p> <p>2) KH157_Si_Ag_hfacAgTMP_HRSEM: high-resolution images taken in the field emitter scanning electron microscope Hitachi S-4800</p> <p>3) KH157_new_Si_Ag_hfacAgTMP_EDX_10kV: elemental analyses done in the field emitter microscope Hitachi S-4800 using an EDAX Genesis 4000 detector and 10 kV acceleration voltage</p> <p>4) KH157_new_Si_Ag_hfacAgTMP_EDX 2022-09 15mm: elemental analyses done in the field emitter microscope Hitachi S-4800 using an EDAX Genesis 4000 detector using 5 kV and 7 kV acceleration voltage</p> <p>5) KH157 - Pillar structure in cross-section: imaging and cross-sectioning done in a Tescan Lyra dual beam instrument plus elemental analysis done in a Tescan Mira FESEM equipped with an EDAX EDX system</p> <p>6) KH157_Si_Ag_hfacAgTMP_TEM: data related to transmission electron microscopy studies done in a ThermoFischer Themis 200 G3 microscope</p> <p>7) pattern_Katja_Hoeflich_Aug2022: pattern and design files used for automation of the patterning with the Xenos patterning software</p> <h3><br>Folder structure: "FESEM_deposition"</h3> <p>1) FESEM_deposition_Si: images of the deposition experiment in the field emitter dual beam instrument</p> <p>2) FESEM_deposition_TEM_grid: data related to the tranmission electron microscopy studies done in a ThermoFischer Themis 200 G3 microscope for deposition directly onto a TEM grid</p> <p>3) pattern: pattern files for patterning carried out using the SmartFIB software</p>
FAIRness of Repositories & Their Data: A Report from LIBER's Research Data Management Working Group
<p>Data repositories play a crucial role in the evolution of Open Science. The FAIR Data Principles establish how to make data Findable, Accessible, Interoperable and Reusable (Wilkinson et al., 2016). The FAIR principles are as follows: </p> <p><strong>To Be Findable</strong></p> <ul> <li>F1. (meta)data are assigned a globally unique and eternally persistent identifier.</li> <li>F2. data are described with rich metadata.</li> <li>F3. (meta)data are registered or indexed in a searchable resource.</li> <li>F4. metadata specify the data identifier.</li> </ul> <p><strong>To Be Accessible:</strong></p> <ul> <li>A1 (meta)data are retrievable by their identifier using a standardized communications protocol.</li> <li>A1.1 the protocol is open, free, and universally implementable.</li> <li>A1.2 the protocol allows for an authentication and authorization procedure, where necessary.</li> <li>A2 metadata are accessible, even when the data are no longer available.</li> </ul> <p><strong>To Be Interoperable</strong></p> <ul> <li>I1. (meta)data use a formal, accessible, shared, and broadly applicable language for knowledge representation.</li> <li>I2. (meta)data use vocabularies that follow FAIR principles.</li> <li>I3. (meta)data include qualified references to other (meta)data.</li> </ul> <p><strong>To Be Reusable</strong></p> <ul> <li>R1. meta(data) have a plurality of accurate and relevant attributes.</li> <li>R1.1. (meta)data are released with a clear and accessible data usage license.</li> <li>R1.2. (meta)data are associated with their provenance.</li> <li>R1.3. (meta)data meet domain-relevant community standards. </li> </ul> <p><strong>Methodology</strong></p> <p>Based on the FAIR Data Principles, two questionnaires were created. The first (hereafter #Q1 - see Appendix #1) targeted repository managers and/or librarians and consisted of 40 questions. The second (hereafter #Q2 - see Appendix #2) targeted technical staff responsible for repository development and maintenance and consisted of 25 questions. </p> <p>Members of LIBER’s <a href="https://libereurope.eu/strategy/research-infrastructures/rdm/">Research Data Management (RDM) Working Group</a> circulated the questionnaires between December 2018 and February 2019. Responses were collected from managers and/or librarians of 29 repositories for the first (#Q1) questionnaire. </p> <p>In addition, technical staff responsible for the development and maintenance of 14 repositories (Table 1) responded to the second (#Q2) questionnaire. In 11 cases, repositories filled out both #Q1 and #Q2. </p> <p>In this report, the responses for both questionnaires have been merged and analyzed to gain a comprehensive picture about FAIRness at the level of repositories and their data.<br> </p>
Data repository Coen Prins bachelorthesis 2019
<p>This zipfile contains all the raw data used during my bachelor thesis regarding the Suppression and induction of early plant defenses by <em>Tetranychus urticae.</em>It also includes the statistical methods used to analyse the data. </p> <p>Within the folder are txt files that explain how to interpret the data & statistical methods </p> <p> </p> <p> </p>
CARE for Indigenous Data: Operationalizing Indigenous Data Governance in Repositories (recording)
<p>Riley Taitingfong, a Henry Luce Foundation postdoctoral scholar at the Native Nations Institute, delivered this keynote address on July 31st, 2024, at the Ethical Open Science for Past Global Change Data 2024 Symposium, in Keshena, Wisconsin, on the lands of the Menominee Nation.</p>
Online repository input data collection framework
Open the record for dataset details and reuse information.
Repositories for taxonomic data: Where we are and what is missing
<p>Natural history collections are leading successful large-scale projects of specimen digitization (images, metadata, DNA barcodes), transforming taxonomy into a big data science. Yet, little effort has been directed towards safeguarding and subsequently mobilizing the considerable amount of original data generated during the process of naming 15–20,000 species every year. From the perspective of alpha-taxonomists, we provide a review of the properties and diversity of taxonomic data, assess their volume and use, and establish criteria for optimizing data repositories. We surveyed 4113 alpha-taxonomic studies in representative journals for 2002, 2010, and 2018, and found an increasing yet comparatively limited use of molecular data in species diagnosis and description. In 2018, of the 2661 papers published in specialized taxonomic journals, molecular data were widely used in mycology (94%), regularly in vertebrates (53%), but rarely in botany (15%) and entomology (10%). Images play an important role in taxonomic research on all taxa, with photographs used in >80% and drawings in 58% of the surveyed papers. The use of omics (high-throughput) approaches or 3D documentation is still rare. Improved archiving strategies for metabarcoding consensus reads, genome and transcriptome assemblies, and chemical and metabolomic data could help to mobilize the wealth of high-throughput data for alpha-taxonomy. Because long term <span>—</span> ideally perpetual <span>—</span> data storage is of particular importance for taxonomy, energy footprint reduction via less storage-demanding formats is a priority if their information content suffices for the purpose of taxonomic studies. Whereas taxonomic assignments are quasi-facts for most biological disciplines, they remain hypotheses pertaining to evolutionary relatedness of individuals for alpha-taxonomy. For this reason, an improved re-use of taxonomic data, including machine-learning-based species identification and delimitation pipelines, <span>requires a cyberspecimen approach—linking data via unique specimen identifiers, and thereby making them </span>findable, accessible, interoperable, and reusable for taxonomic research<span>. This poses both qualitative challenges to adapt the </span>existing infrastructure of data cen<span>ters to a specimen-centered concept and quantitative challenges to host </span>and connect an estimated ≤2 million images produced per year by alpha-taxonomic studies, plus many millions of images from digitization campaigns. Of the 30–40,000 taxonomists globally, many are thought to be <span>non-professionals, and capturing the data for online storage and reuse therefore requires</span> low-complexity submission workflows and cost-free repository use. E<span>xpert taxonomists are the main stakeholders able to identify and formalize the needs of the discipline</span>; their expertise is needed to implement the envisioned<span> virtual collections of cyberspecimens.</span></p>
Data repository accompanying "Rapid microwave-only characterization and readout of quantum dots using multiplexed gigahertz-frequency resonators"
<p>Data repository including raw data and analysis scripts used for generating the figures appearing in the paper 'Rapid microwave-only characterization and readout of quantum dots using multiplexed gigahertz-frequency resonators'</p>
Training data for the GitHub repository "buildingsFromSentinel"
<p>Training and testing data for machine learning models predicting the building height and footprint from satellite data in urban areas.</p> <p>Sentinel-1 and -2 data are retrieved from https://scihub.copernicus.eu/ and the GHS built-up grid (here GHSBuilt10) from https://ghsl.jrc.ec.europa.eu/download.php?ds=buS2. GHSBuilt10 is derived from Sentinel-2 global image composite for the reference year 2018 using Convolutional Neural Networks (GHS-S2Net).</p> <p>The dataset contains the following folders:</p> <ul> <li>footprint: PNG images over urban areas with either three or four features: <ul> <li>XXX_labels.png: true-colour images (TCI) retrieved from Sentinel-2 data</li> <li>XXX_labels4.png: TCIs with the band 8 (i.e., near-infrared = NIR) as the fourth dimension in the image.</li> </ul> </li> <li>height: data for different cities <ul> <li>building_height.tif: real building height (only for the training data)</li> <li>sentinel_cropped: satellite images for the same area. Contains Sentinel-1 and -2 data as well as the GHS-Built data with a 10-m resolution.</li> <li>README.txt: information of the origin of the building height data</li> </ul> </li> </ul>
Cosegmentation for Plant Phenotyping (CosegPP) Data Repository Collected Via a High-Throughput Imaging System
<p>CosegPP is a data repository that contains four datasets for plant phenotyping. Each dataset contains: </p> <ol> <li>two species physically different for challenging segmentation. Buckwheat is a thin plant with a variety sizes of leaves and Sunflower is a bushy plant that contains flowering;</li> <li>the most commonly used induced environments in plant phenotyping such as a control and drought-induced; </li> <li>a temporal resolution that begins with the plants vegetative stage and ends with the plant fully matured;</li> <li>modalities (infrared, visible, near infrared) that are commonly used in plant phenotyping analysis; and </li> <li>multiple perspectives that are becoming widely acquired in plant phenotyping analysis due to its potential for three dimensional analysis.</li> </ol> <p>We thank Vincent Stoeger for acquiring the dataset using LemnaTec at the University of Nebraska-Lincoln.</p> <p>If you use this dataset, please cite this paper:</p> <p>Quiñones R, Munoz-Arriola F, Choudhury SD, Samal A (2021) Multi-feature data repository development and analytics for image cosegmentation in high-throughput plant phenotyping. PLoS ONE 16(9): e0257001. <a href="https://doi.org/10.1371/journal.pone.0257001">https://doi.org/10.1371/journal.pone.0257001</a></p>
Raw data repository for the article: "High spatial coherence and short pulse duration revealed by the Hanbury Brown and Twiss interferometry at the European XFEL"
<p>Raw data depository for the article in the Structural Dynamics journal: DOI: 10.1063/4.0000127. Details with the file information are given in the file "HBT_XFEL_Data_set_Info_final.pdf"</p>
Visualization of commits to papyri.info data repository (ca. 2011)
<p>This is a visualization of commits made to the papyri.info data repository: <a href="https://www.youtube.com/redirect?event=video_description&redir_token=QUFFLUhqbGRFdHJRZVFSXzNndHFaU19NUEVsMEhwT21UUXxBQ3Jtc0tsQV9jcnNzTnBoU29aZGtiODRTQjdwaWZSaU96N2laNDZfYzVmLXpBTUhVU0JXTXVETjdaSGxwbjVNWWhweUFKOFJJdUhSdUVFZ0tGSHFXMFRyZEM1UmRQOU9KODFLai1ubThMSEhPRml6MG1QS09TNA&q=https%3A%2F%2Fgithub.com%2Fpapyri%2Fidp.data&v=l7ujo41j_Ig">https://github.com/papyri/idp.data</a></p> <p>Made using gource: <a href="https://www.youtube.com/redirect?event=video_description&redir_token=QUFFLUhqa2dic2pFTzdjUW84MjNaWGUtZXJnNkdOLUpoUXxBQ3Jtc0trQWQwWi0zQU9ZNUgzdDBSc2U5cUNFemIyUEl0MFA4TGY0Qk9uSGZyU0djODRVejZhTFZmWVNoQndDRGYtRkZDOWJ4cnpJa2JNalhINXNiN3JadVFCUWlGMGwtRmxURkdRNFRZRkg4OHhRUkYxdU92cw&q=https%3A%2F%2Fgource.io%2F&v=l7ujo41j_Ig">https://gource.io/</a></p>
Copernicus Climate Change Service data for the pypsa-entsoe Github repository
<p>Files needed for the https://github.com/matteodefelice/pypsa-entsoe repository.</p>
DATA REPOSITORY FOR: All-Optical Nuclear Quantum Sensing Using Nitrogen-Vacancy Centers in Diamond
<p><strong>DATA REPOSITORY:<br> ALL-OPTICAL NUCLEAR QUANTUM SENSING USING NITROGEN-VACANCY CENTERS IN DIAMOND</strong></p> <p>This data repository contains the raw data as measured on the experimental setup, the files required to do the data processing we performed on the raw data, the scripts to run the simulations described in the journal article, and the scripts to reproduce the plots shown in the article's figures.<br> <br> Use MatLab R2019b or later to run these files.<br> See ReadMe.txt for more information.</p>
Dominance of contrasting fungal functional groups influence nutrient cycling across four Japanese cool-temperate forest soils - Data repository
<p>Soil data related to the publication <strong>Dominance of contrasting fungal functional groups influence nutrient cycling across four Japanese cool-temperate forest soils.</strong></p>
Data Repository for Chip-Chat: Challenges and Opportunities in Conversational Hardware Design
<p><strong># Data Repository for Chip-Chat: Challenges and Opportunities in Conversational Hardware Design</strong></p><p>This repository accompanies the manuscript accepted at MLCAD 2023, titled "Chip-Chat: Challenges and Opportunities in Conversational Hardware Design".</p><p>It contains the following:</p><p>- `free-chat-gpt4-tt03` - this contains the data used for the paper, which examines free-form process when exploring the potential applications for LLMs in hardware design. The task here was to generate the Verilog for a full (albeit small) processor design. Here, the chats are presented (and annotated) in the `/chats` subdirectory, which also includes a python script for extracting metadata (presented in table IV in the manuscript). Note that this directory also includes `/assembler` which provides a basic assembler (also written in Python) to make it easier to write demo programs (examples included) for the processor.</p><p>- `scripted-benchmarks` - this contains additional data not used in the paper, which examines a more rigid process when exploring the potential applications for LLMs in hardware design. Here, each model chats are separated by subdirectory.</p><p>- `scripted-benchmarks-gpt4-tt03` - this contains just the benchmarks not used in the paper, made by the first run of GPT-4, which were used for tapeout in Tiny Tapeout 3.</p><p>The two tt03 directories contain the GitHub action scripts required to invoke OpenLane and produce synthesis files, as well as used to perform simulation tests.</p>
Data repository Childprogramming + Debugging + Robotics
<p>Este repositorio de datos alberga información sobre las investigaciones realizadas en el marco de dos investigaciones relacionadas y centradas en el fomento del pensamiento computacional en la educación primaria mediante la metodología Childprogramming. El repositorio sirve de eje central para almacenar y organizar datos valiosos relativos a la aplicación de conceptos de depuración y robótica educativa. Contiene información relativa a los proyectos realizados en este ámbito. Incluye los resultados de los casos de estudio aplicados.</p> <p> </p>
Data repository for "Navigating trade-offs and sustainable development pathways in the Andean water-energy-food-ecosystem nexus"
<p>Data and code supporting the research article: Navigating trade-offs and sustainable development pathways in the Andean water-energy-food-ecosystem nexus.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.