Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
3,481
datasets available to search
ShareScore release 0.9.0
Dataset results
3,481 results for “data set”
data sets for "Ultra-Conserved Elements and morphology reciprocally illuminate conflicting phylogenetic hypotheses in Chalcididae (Hymenoptera, Chalcidoidea)"
<p>data sets used in :</p> <p>Cruaud A, Delvare G, Nidelet S, Sauné L, Ratnasingham S, Chartois M, Blaimer BB, Gates M, Brady SG, Faure S, van Noort S, Rossi J-P, and Rasplus J-Y. in press. Ultra-Conserved Elements and morphology reciprocally illuminate conflicting phylogenetic hypotheses in Chalcididae (Hymenoptera, Chalcidoidea). Cladistics.</p> <p>1) morphological matrix (matrix_morphology_chalcididae_cladistics2020.nex) : Further details on characters and character states as well as illustrations can be found in the manuscript.</p> <p>2) concatenated UCE data set (merge_UCEs_chalcididae_cladistics2020.phy)</p> <p>3) all phylogenetic trees (morphology, UCEs, subset of UCEs; see paper for further details)</p> <p>4) a mesquite file with mapping of morphological characters on alternative UCE trees and the MJ consensus tree of the morphological analysis (mesquite_trees_and_morphological_transformations_chalcididae_cladistics2020.nex)</p>
Gullspång Pull-out Test Data Set
<p>This data set contains the outcome of a series of pull-out tests (and 3D scan analysis) in the investigation of the bond behavior of naturally-corroded, plain reinforcement sourced from a decommissioned structure. The contents of this data set includes 1) photos and test measurements for all tested rebars; 2) an SQL database containing all pullout (unprocessed) data; and 3) an SQL database containing processed data. Data is provided in SQL format to permit querying of entries across the otherwise large, and parametrically diverse, data set. A "Read Me" file is also provided for additional descriptions of the content within.</p>
The GitHub Sponsors Data Set
<p>This is the data set for our submission to MSR 2020. "Contributing by Paying: A Data Set and Data-Driven Study on the GitHub Sponsors Program". Please download and unzip the .zip archive to get the data set used in the paper (folder <em>2019-11-30</em>) and recent updates (folder <em>sponsorship_snapshots</em>) on the snapshots of sponsors.</p>
Data set of airborne and ground-based atmospheric measurements from Hyytiälä, Finland
<p>This data set includes airborne and ground-based measurements of aerosol particles and meteorological variables from Hyytiälä, Finland (61.85N, 24.28E). The airborne measurements were done between 2013-2015 and the ground-based measurements include selected data between 2006-2017.</p>
Data-sets for Indoor Photovoltaic Behavior in Low Lighting condition
<p>This file includes five data-sets from behavior of indoor photovoltaic modules under pure artificial lighting conditions with low light intensity.</p> <p>Three types of PV module are included by use of two types of light sources measured in a high accuracy controlled light testbed.</p> <p>One of data-sets includes data measured within a warehouse as a industrial environment mostly with pure artificial lighting with low light intensity.</p> <p>Each data-set includes multiple measurements at different light intensities and temperature condition.</p> <p>Each measurement includes both radiometry and photometry spectrum of the light, integrative light intensity and temperature in addition to the voltage-current relation of the PV module in that lighting condition.</p>
Ultimate-Guitar Data Set
<p>Data yang kami ambil adalah data pencarian chord lagu pada website ultimate guitar. Dataset yang kami gunakan berisi 5 kolom dan 150 baris daftar chord lagu yang ada pada website ultimate guitar. Kolom artist_name berisi nama artis yang menyanyikan lagu. penyanyi, kolom artist_songs berisi judul lagu, kolom song_rating berisi jumlah orang atau user yang melakukan rating ke chord lagu, kolom song_hits berisi jumlah chord lagu telah dilihat pada hari data di ambil, kolom type berisi bentuk dari chord lagu.</p>
Atmospheric aerosol, gases and meteorological parameters measured during the LAPSE-RATE campaign - Kansas State University data sets
<p>This publication summarizes the measurements and data sets generated by Kansas State University (KSU) during the LAPSE-RATE that took place in the San Luis Valley of Colorado during the summer of 2018. These data sets offer observations of atmospheric aerosols at the surface and in vertical column acquired by KSU rotary-wing Unmanned Aerial System. </p>
Repackaged Full ITIS Data Set (MS SQL Server)
<p>Retrieved 18 May 2020, from the Integrated Taxonomic Information System (ITIS) (<a href="http://www.itis.gov">http://www.itis.gov</a>). via https://www.itis.gov/downloads/itisMSSql.zip .</p> <p>The archive itisMSSql.zip was unzipped, and repackaged as individual gzipped files. The original zip file is included in this data publication.</p> <p>Files in this publication:</p> <p>1. itisMSSql.zip - file downloaded from https://www.itis.gov/downloads/itisMSSql.zip on 18 May 2020</p> <p>2. Files ending with .gz (e.g., taxonomic_units.gz, taxonomic_units.gz, synonym_links.gz) - repackaged, gzipped, content of itisMSSql.zip</p> <p> </p> <p> </p>
Data Sets for Utilizing Graceful Failure as An Opportunity for Flood Mitigation Downstream to Protect Communities and Infrastructure
<p>This spreadsheet provides the volume analysis calculations associated with the project. This project focused on exploring the potential feasibility to utilize other locations along the inland waterway system where “graceful failure” or planned breach of levees may be used as a means of flood protection for downstream communities and infrastructure. Spatial analysis techniques were used with development of specific criteria to screen national-level data sets to identify probable locations for such mitigative approaches. The criteria were primarily focused on identifying non-urbanized, non-developed land where intentional flooding for storage of flood waters would minimize impacts. Each location that was identified as a potential candidate was further evaluated for capacity for flood water detention. A consolidated set of areas were identified that could provide some storage capacity for flood mitigation. Additional engineering and localized analysis would be necessary to vet the areas for actual storage implementation. However, this study provides an example of an unconventional approach to flood mitigation on inland waterways which could reduce the need for disaster response and assist in transportation planning during extreme flood conditions.</p>
Financial data set used in INFORE project
<p>The data set contains exchange price data from different exchanges. The time span ranges from 01/01/2019 until 12/31/2019. Several Stock-, Future-, Index-, Commodity- and Currency-markets are covered.</p> <p>The data comes as tick by tick data (sub folder 'quotes') as well as condenced data (condenced to 1 -minute-format; in sub folder 'history').</p>
DATA SET - Sublethal exposure to deltamethrin impairs maternal egg care in the European earwig Forficula auricularia
<p>Data set of the manuscript entitled "Sublethal exposure to deltamethrin impairs maternal egg care in the European earwig Forficula auricularia" and published in the journal Chemosphere.</p>
San Solomon Springs Baseflow Period Hydrochemistry Data Set
<p>San Solomon Springs Dec 19' - April 20'</p>
Experimental Data Set for the study "Exploratory Landscape Analysis is Strongly Sensitive to the Sampling Strategy"
<p>This are the feature values used in the study "Exploratory Landscape Analysis is Strongly Sensitive to the Sampling Strategy".</p> <p>The dataset regroups feature values for every "cheap" features available in the R package <em>flacco </em>and are computed using 5 sampling strategies and in dimension <span class="math-tex">\($d=5$\)</span>:</p> <ol> <li>Random: the classical Mersenne-Twister algorithm;</li> <li>Randu: a random number generator that is notoriously bad;</li> <li>LHS: a centered Latin Hypercube Design;</li> <li>iLHS: an improved Latin Hypercube Design;</li> <li>Sobol: points extracted from a Sobol' low-discrepancy sequence.</li> </ol> <p>The csv file <em>features_summury_dim_5_ppsn.csv </em>regroups 100 values for every features whereas <em>features_summury_dim_5_ppsn_median.csv </em>regroups for every feature the median of the 100 values.</p> <p>In the folder <em>PPSN_feature_plots</em> are the histograms of feature values on the 24 COCO functions for 3 sampling strategies: Random, LHS and Sobol.</p> <p>The Python file <em>sampling_ppsn.py</em> is the code used to generate the sample points from which the feature values are computed.</p> <p>The file <em>stats50_knn_dt.csv</em> provide the raw data of median and IQR (inter quartile interval) for the heatmaps and boxplots available in the paper.</p> <p>Finally, the files <em>results_classif_knn100.csv</em> (resp. dt) provide the accuracy of 100 classifications for every settings.</p> <p> </p>
Data set for Neural-Network-Based Digital Predistortion for Active Antenna Arrays Under Load Modulation
<p>The dataset contains over-the-air measurements on a 64 active antenna array (Anokiwave AWMF-0129) operating at 28 GHz carrier frequency and transmitting a 200 MHz OFDM waveform with FFT size of 4096, 3168 active subcarriers, subcarrier spacing of 60 kHz and 5 times oversampling w.r.t the critical sampling rate. The dataset contains the I/Q samples of the TX waveform as well as the corresponding over-the-air received signals when the electrical beam is steered toward different beamforming directions. This dataset allows to observe and study the so-called beam-dependent load modulation, which is the phenomenon that causes the nonlinear characteristics of the antenna array to change with the beamforming direction. A Matlab script for data visualization is also provided.</p> <p>For further details please refer to the following papers:</p> <p>A. Brihuega <em>et al</em>., "Piecewise Digital Predistortion for mmWave Active Antenna Arrays: Algorithms and Measurements," in <em>IEEE Transactions on Microwave Theory and Techniques</em>, doi: 10.1109/TMTT.2020.2994311.</p> <p>A. Brihuega <em>et al</em>., “Neural-Network-Based Digital Predistortion for Active Antenna Arrays Under Load Modulation,” in <em>IEEE Microwave and Wireless Components Letters, </em>doi: 10.1109/LMWC.2020.3004003</p>
Data set for "Dynamic perceptual feature selectivity in primary somatosensory cortex upon reversal learning"
<p>This repository contains the data used to generate the figures and well as the main codes that were used for analyses.</p>
Sensor data set, electromechanical cylinder at ZeMA testbed (2kHz)
<p><strong>General information on the data set</strong></p> <p>The data set was generated at the ZeMA testbed. A working cycle lasts 2.8s and consists of a forward stroke, a waiting time and a return stroke. The data set does not consist of the entire working cycles. Only one second of the return stroke of each working cycle is used.</p> <p> </p> <p><strong>Structure of the data</strong></p> <ul> <li>data saved in HDF5 file as a 3D-matrix</li> <li>one row represents one second of the return stroke of one working cycle (6292 rows: 6292 cycles)</li> <li>one column represents one datapoint of the cycle, that is resampled to 2 kHz (2000 columns)</li> <li>one page represent one sensor (11 pages: 11 sensors)</li> </ul> <p> </p> <p><strong>Allocation of the pages to the sensors</strong></p> <p>page 1: microphone<br> page 2: acceleration plain bearing<br> page 3: acceleration piston rod<br> page 4: acceleration ball bearing<br> page 5: axial force<br> page 6: pressure<br> page 7: velocity<br> page 8: active current<br> page 9: motor current phase 1<br> page 10: motor current phase 2<br> page 11: motor current phase 3</p> <p> </p> <p><strong>Remark</strong></p> <p>The datasets are not in SI units. For conversion, you can use the PDF documentation.</p> <p> </p> <p><strong>Further information</strong></p> <p>For an introduction and tutorial to this data, a set of Jupyter notebooks is available <a href="https://github.com/harislulic/ZeMA-machine-learning-tutorials">here</a>. These notebooks contain Python code and a documentation of example machine learning tasks and analysis of this data set. In the near future, these will be extended to also include uncertainties in the input data.</p>
Sensor data set of 3 electromechanical cylinder at ZeMA testbed (2kHz)
<p><strong>General information on the data set</strong></p> <p>The data set was generated at the ZeMA testbed. A working cycle lasts 2.8s and consists of a forward stroke, a waiting time and a return stroke. The data set does not consist of the entire working cycles. Only one second of the return stroke of each working cycle is used.</p> <p> </p> <p><strong>Structure of the data</strong></p> <ul> <li>data saved in three HDF5 file as a 3D-matrix, one file is for one axis</li> <li>one row represents one second of the return stroke of one working cycle<br> axis 3: 6292 cycles<br> axis 5: 6083 cycles<br> axis 7: 5732 cycles</li> <li>one column represents one datapoint of the cycle, that is resampled to 2 kHz (2000 columns)</li> <li>one page represent one sensor (11 pages: 11 sensors)</li> </ul> <p> </p> <p><strong>Allocation of the pages to the sensors</strong></p> <p>page 1: microphone<br> page 2: acceleration plain bearing<br> page 3: acceleration piston rod<br> page 4: acceleration ball bearing<br> page 5: axial force<br> page 6: pressure<br> page 7: velocity<br> page 8: active current<br> page 9: motor current phase 1<br> page 10: motor current phase 2<br> page 11: motor current phase 3</p> <p> </p> <p><strong>Remark</strong></p> <p>The datasets are not in SI units. For conversion, you can use the PDF documentation.</p> <p> </p> <p><strong>Further information</strong></p> <p>For an introduction and tutorial to this data, a set of Jupyter notebooks is available <a href="https://github.com/harislulic/ZeMA-machine-learning-tutorials">here</a>. These notebooks contain Python code and a documentation of example machine learning tasks and analysis of this data set. In the near future, these will be extended to also include uncertainties in the input data.</p>
Data from: Hybrid zone barriers comparative data set
<p>Many recent studies have addressed the mechanisms operating during the early stages of speciation, but surprisingly few studies have tested theoretical predictions on the evolution of strong reproductive isolation (RI). To help address this gap, we first undertook a quantitative review of the hybrid zone literature for flowering plants in relation to reproductive barriers. Then, using Populus as an exemplary model group, we analysed genome-wide variation for phylogenetic tree topologies in both early- and late-stage speciation taxa to determine how these patterns may be related to the genomic architecture of RI. Our plant literature survey revealed variation in barrier complexity and an association between barrier number and introgressive gene flow. Focusing on Populus, our genome-wide analysis of tree topologies in speciating poplar taxa points to unusually complex genomic architectures of RI, consistent with earlier genome-wide association studies. These architectures appear to facilitate the 'escape' of introgressed genome segments from polygenic barriers even with strong RI, thus affecting their relationships with recombination rates. Placed within the context of the broader literature, our data illustrate how phylogenomic approaches hold great promise for addressing the evolution and temporary breakdown of RI during late stages of speciation. This article is part of the theme issue 'Towards the completion of speciation: the evolution of reproductive isolation beyond the first barriers'.</p>
Data sets v.1.1 for Perspective: SARS-CoV-2 may regulate cellular responses through depletion of specific host miRNAs
<p>The list of potential (bioinformatic predictions) interactions of human miRNA with 7 coronavirus genomes that include 3 pathogenic and 4 non-pathogenic coronaviruses. (<strong>Data Set 1</strong>). <em>The HCoVs' RNA genomes of pathogenic strains were SARS-CoV-2 (NC_045512.2), SARS-CoV (NC_004718.3), MERS-CoV (NC_019843.3). </em>The non-pathogenic strains were HCoV-OC43 (KU131570.1), HCoV-229E (NC_002645.1), HCoV-HKU1 (KF686346.1), and HCoV-NL63 (NC_005831.2). These coronaviruses were tested against the set of 896 confident mature human miRNA sequences that were obtained from the miRBbase v2.21 using the RNA22 v2 microRNA target discovery tool web-server. In order to reduce the false discovery rate of the MTS predictions, the most strict parameters were applied to the default computation workflow using a specificity of 92% versus a sensitivity of 22%.</p> <p><strong>Data set 2.</strong> The potential targets of miRNA that could be bound to either the pathogenic, the non-pathogenic or both groups of HCoVs. Predicted base on miRDIP database (with only top 1% of the most probable targets considered),</p> <p><strong>Data set 3. </strong> Pre-miRNA sequences in the <em>SARS-CoV-2</em> RNA sequence that could potentially enter the human RNAi pathway, base on miRNAFold webserver.</p>
Detrended SSH data set used for Rossby Wave Analysis, extraction at 39N from ORCA12.L46-MJM189 DRAKKAR simulation
<p>This data set corresponds to the detrended Sea Surface Heigh (SSH) simulated by the NEMO ocean circulation model, under the ORCA12.L46-MJM189 configuration, developped in the frame of the DRAKKAR project. This particular data set is an interpolation from the native numerical grid, covering the latitude 39N in the North Altantic ocean, for the period 1970 to 2015. The data are concatenated in a single file with 5-days average of SSH. This subset was used in Watelet et al. (2020) submitted paper, dealing with Rossby waves analysis.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.