Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

3,481

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

3,481 results for “data set”

Learn how ShareScore rates datasets ↗
zenodo28/100

De-identified article and author characteristics for a large data set of Web of Science

<p>This data set contains article and author characteristics for all records in the Web of Science, 2000-2020. Standard article identifiers have been removed and replaced with a document ID (`doc_id`), as linking to the original ID is not permitted.</p>

opencc-by-4.0Jan 2023View details →
zenodo28/100

Data set from Fischertechnik Smart Factory Model at University of St.Gallen (Standard Fischertechnik Configuration)

<p>This is about 90 mins worth of data collected via the MQTT interface of the Fischertechnik Industry 9.0V smart factory model available at the University of St.Gallen. Each entry in the file corresponds to one message (as JSON object) received on a specific topic via MQTT.</p> <p>The description of the MQTT interface can be found here: <a href="https://github.com/fischertechnik/txt_training_factory/blob/master/TxtSmartFactoryLib/doc/MqttInterface.md">https://github.com/fischertechnik/txt_training_factory/blob/master/TxtSmartFactoryLib/doc/MqttInterface.md</a></p> <p>Check the following publications to learn more about our research using the model factory:</p> <p>Malburg, L., Seiger, R., Bergmann, R., &amp; Weber, B. (2020). Using physical factory simulation models for business process management research. In&nbsp;<em>Business Process Management Workshops: BPM 2020 International Workshops, Seville, Spain, September 13&ndash;18, 2020, Revised Selected Papers 18</em>&nbsp;(pp. 95-107). Springer International Publishing.</p> <p>Seiger, R., Zerbato, F., Burattin, A., Garc&iacute;a-Ba&ntilde;uelos, L., &amp; Weber, B. (2020, October). Towards iot-driven process event log generation for conformance checking in smart factories. In&nbsp;<em>2020 IEEE 24th International Enterprise Distributed Object Computing Workshop (EDOCW)</em>&nbsp;(pp. 20-26). IEEE.</p> <p>Seiger, R., Malburg, L., Weber, B., &amp; Bergmann, R. (2022). Integrating process management and event processing in smart factories: A systems architecture and use cases.&nbsp;<em>Journal of Manufacturing Systems</em>,&nbsp;<em>63</em>, 575-592.</p> <p>&nbsp;</p>

opencc-by-4.0Feb 2023View details →
zenodo28/100

Raptor@GenomeBiology: RefSeq data set

<p>Size: 29 GiB</p> <p>SHA256:&nbsp;cbf98d727cb87df8dff249714c4a2e07e24ade532e7d56f8849e3e9d8c0f27db</p>

opencc-by-4.0Mar 2023View details →
zenodo28/100

Data Set "A simple and consistent quantum-chemical fragmentation scheme for proteins that includes two-body contributions"

<p>Data set accompanying the publication &quot;A simple and consistent quantum-chemical fragmentation scheme for proteins that includes two-body contributions&quot;</p> <p>This dataset contains:</p> <p>- PDB files of all structures used as test cases</p> <p>- PyADF scripts for running MFCC-MBE(2) calculations</p> <p>- Jupyter notebook for generating plots including raw numerical data</p> <p>- Total energies (in a.u.) for the different test cases</p>

opencc-by-4.0Aug 2022View details →
zenodo28/100

Data set and analytic codes supporting "The impacts of within-stream physical structure and riparian buffer strips on semi-aquatic bugs in Southeast Asian oil palm"

<p>This deposit contains data set and analytic codes&nbsp;(accompanied with a meta data) supporting&nbsp;&quot;The impacts of within-stream physical structure and riparian buffer strips on semi-aquatic bugs in Southeast Asian oil palm&quot;. We assessed the impacts of within-stream physical structure and riparian buffer strips&nbsp;on semi-aquatic bug (Gerromorpha, Hemiptera) communities in oil palm streams in Sabah, Malaysia. Collections of semi-aquatic bugs were conducted from&nbsp;oil palm with&nbsp;and without riparian buffer strips.</p> <p>Several environmental parameters were collected to represent within-stream physical structure. We investigated the impacts&nbsp;on the abundance, biomass, species richness, and community composition of semi-aquatic bugs. Additionally, we studied the effects on the proportion of juveniles as well as female&nbsp;<em>Ptilomera</em>&nbsp;sp. (a morphospecies with clear sexual dimorphism in this study).</p> <p>This research was funded by the Jardine Foundation, the Cambridge Trust, the Natural Environment Research Council (NERC) (studentship 1122589),&nbsp;Proforest, the Varley Gradwell Travelling Fellowship, the Tim Whitmore Fund, the Panton Trust, the Cambridge University Commonwealth Fund,&nbsp;the Hanne and Torkel Weis-Fogh Fund, and the S.T. Lee Fund.</p>

opencc-by-4.0Apr 2023View details →
zenodo28/100

Thermal Power Prediction Data set

<p>Haoning Jia &#39;s graduation thesis chapter three raw data.</p>

opencc-by-4.0Apr 2023View details →
zenodo28/100

Data Set Generated by the Fuzzy Model Constructed to Describe Execution Tracing Quality

<p>The uploaded data set was generated by the fuzzy model published in&nbsp;T. Galli, F. Chiclana, and F. Siewe. Genetic algorithm-based fuzzy inference system for describing execution tracing quality. Mathematics, 9(21), 2021. ISSN 2227-7390. doi: https://doi.org/10.3390/ma th9212822. URL https://www.mdpi.com/2571-5577/4/1/20.</p> <p>The goal of the data generation is to make the published model available in the form of data points in a 5D space, which facilitates the construction of simpler models to approximate the original model. The names of the columns in the .csv file constitute the quality properties of execution tracing: (1) accuracy, (2) legibility, (3) implementation, and (4) security, while column (5) contains&nbsp;execution tracing quality derived from the fuzzy model. The indices in brackets show the column indices in the .csv file.</p> <p>All variables lie in the continuous range [0, 100], where 100 means the best possible quality value and 0 the complete lack of&nbsp; quality or the lack of the given&nbsp;quality property. While generating the data, the inputs were&nbsp;increased by a step-size 5 and the model&#39;s output was collected, i.e. 4 inputs, from including 0 to 100 with 21 data points (21^4 = 194481).</p> <p>&nbsp;</p>

opencc-by-4.0Apr 2023View details →
zenodo28/100

Data set Quality of Life Autism

<p>Data set for Quality of life of Families of Children with Autism Spectrum Disorders&nbsp;in Jordan.</p>

opencc-by-4.0Apr 2023View details →
zenodo28/100

Data set

<p>Data for the article&nbsp;</p>

opencc-by-4.0Apr 2023View details →
zenodo28/100

Data set (Song et al., Nat. Water)

<p>Data set (Song et al., Nat. Water)</p>

opencc-by-4.0Jun 2023View details →
zenodo28/100

Data set for cardiovascular and mood responses to an acute bout of cold water immersion

<p>Data set for manuscript.&nbsp;</p>

opencc-by-4.0Jun 2023View details →
zenodo28/100

Data sets for 1489798

<p>Data sets for 1489798: Dissertation</p>

opencc-by-4.0Jun 2023View details →
dryad28/100

Raw motif mapping bedfile data and model training set class probabilities

<p>Leveraging prior viral genome sequencing data to make predictions on whether an unknown, emergent virus harbors a 'phenotype-of-concern' has been a long-sought goal of genomic epidemiology. A predictive phenotype model built from nucleotide-level information alone is challenging with respect to RNA viruses due to the ultra-high intra-sequence variance of their genomes, even within closely related clades. We developed a degenerate k-mer method to accommodate this high intra-sequence variation of RNA virus genomes for modeling frameworks. By leveraging a taxonomy-guided 'group-shuffle-split' cross validation paradigm on complete coronavirus assemblies from prior to October 2018, we trained multiple regularized logistic regression classifiers at the nucleotide k-mer level. We demonstrate the feasibility of this method by finding models accurately predicting withheld SARS-CoV-2 genome sequences as human pathogens and accurately predicting withheld Swine Acute Diarrhea Syndrome coronavirus (SADS-CoV) genome sequences as non-human pathogens. Feature selection using L1 regularization identified several degenerate nucleotide predictor motifs with high model coefficients for the human pathogen class that were present across widely disparate clades of coronaviruses. However, these motifs differed in which genes they were present in, what specific codons were used to encode them, and what the translated amino acid motif was. This emphasizes the importance of a phenetic view of emerging pathogenic RNA viruses, as opposed to the canonical phylogenetic interpretations most commonly used to track and manage viral zoonoses. Applying our model to more recent Orthocoronavirinae genomes deposited since October 2018 yields a novel contextual view of pathogen potential across bat-related, canine-related, porcine-related, and rodent-related coronaviruses and critical adaptations which may have contributed to the emergence of the pandemic SARS-CoV-2 virus. Finally, we discuss the next steps to achieve robust predictive ensembles and the utility of these models (and their associated predictor motifs) to novel biosurveillance protocols that substantially increase the 'pound-for-pound' information content of field-collected sequencing data and make a strong argument for the necessity of routine collection and sequencing of zoonotic viruses. </p>

opencc-zeroJun 2023View details →
zenodo28/100

Data set for Finding Equivalence of Layout Independent Water Distribution Network for Expansion/Reorganization Using Nonlinear Multi-Port Thevenin Theorem

<p>This is the dataset for our manuscript titled &quot;Finding Equivalence of Layout Independent Water Distribution Network for Expansion/Reorganization Using Nonlinear Multi-Port Thevenin Theorem&quot;.</p>

opencc-by-4.0Jul 2023View details →
zenodo28/100

M3-Seq imaging data set (Wang, et. al 2023)

<p>Raw image files for M3-Seq paper.</p>

opencc-by-4.0Sep 2023View details →
zenodo28/100

2023_Melo et al. BPMN x NI x Game - Data Set

<p>Full data from Caue&#39;s TCC</p>

opencc-by-4.0Jul 2023View details →
zenodo28/100

Code and experimental data for the ECAI 2023 paper "PARIS: Planning Algorithms for Reconfiguring Independent Sets"

<p>The archive chirsten-et-al-ecai2023-solvers contains the code necessary to generate the singularity images used in the experimental evaluation, except for CPLEX which is needed for some PARIS images. The solvers can be built using the build.sh script in the solver&#39;s sub-directory.</p> <p>The archive christen-et-al-ecai2023-data contains all the scripts necessary to run the experiments<br> presented in the paper, as well as all the raw and processed data.</p> <p>The archive christen-et-al-ecai2023-benchmarks contains all the benchmarks from the competition as well as our compiled PDDL and SAS+ versions, and the scripts used for the compilation.</p> <p>For a detailed explanation, see the README file within the archives. For licensing information about the solvers, consult the LICENSE file within the solvers archive.</p>

openother-atJul 2023View details →
zenodo28/100

EgC v3 - Emotion Data Set in Game Against Corruption (SBGames 2023)

<p>EgC v3 - Emotion Data Set in Game Against Corruption (SBGames 2023)</p>

opencc-by-4.0Aug 2023View details →
zenodo28/100

Data set of 1,275 images capturing interactions between flies and blooming flowers

<p>The dataset presented is a collection of 1,275 images capturing interactions between flies and blooming flowers. The images were sourced from internet repositories through searches conducted between August 2016 and August 2020, using the Google Chrome v. 33.x web browser. Internet searches focused on Google Images and three major social media platforms: Flickr, Instagram, and agefotostock. Photographs were taken by photographers worldwide and uploaded to these platforms, forming the basis of the dataset. The data encompasses various taxonomic and ecological information for both the flies and the flowers depicted in the images. For each image, detailed taxonomic information was recorded for the flies, including their suborders (Nematocera and Brachycera, grouped as Higher/Lesser), Family (wherever possible, distinguishing between Syrphidae and non-Syrphidae), and Genus and species (if available). Further characterization of the flies included recording their sex (identified based on morphology), feeding status (identified by visible proboscis extension into the flower), and the presence or absence of pollen particles on their bodies. Similarly, for each image, taxonomic information was collected for the flowers, including their Family and Genus and species (when identifiable). Flowers were categorized by petal color, which was grouped into four main categories based on the visible spectrum wavelength: purple to blue (ranging from 380-520 nm), green to yellow (ranging from 520-590 nm), orange to red (ranging from 590-740 nm), and white. Additionally, flowers were classified based on their shape, with four main categories: elongate cluster, round cluster, composite-shaped, or simple-shaped. To complement the taxonomic and morphological information, the dataset includes additional data for each image, such as web links to the original sources, geographic locations where the images were captured, and the date of image acquisition. This dataset offers a valuable resource for studying fly-flower interactions on a broad scale, using&nbsp;photographs contributed by photographers from around the world.&nbsp;</p>

opencc-by-4.0Aug 2023View details →
zenodo28/100

Interfacial Cherenkov radiation from ultralow-energy electrons - Data set

<p>Data set for &quot;Interfacial Cherenkov Radiation from Ultralow-Energy Electrons&quot;</p>

opencc-by-4.0Aug 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record