Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
3,481
datasets available to search
ShareScore release 0.9.0
Dataset results
3,481 results for “data set”
Pultruded carbon fiber profiles - 3D x-ray tomography data-sets for two different pultruded profiles
<p>3D X-ray scan on Zeiss Xradia 520</p> <table> <tbody> <tr> <td> <table> <tbody> <tr> <td> </td> <td>A1</td> <td>A2</td> <td>A2S</td> <td>B1</td> <td>B1S</td> <td>B2</td> <td>B3</td> </tr> <tr> <td>Scanning Voxel size [μm]</td> <td>1.9767</td> <td>1.97</td> <td>1.9731</td> <td>1.9751</td> <td>1.9753</td> <td>1.9771</td> <td>1.9752</td> </tr> <tr> <td>FoV: (x; y; z) [mm]</td> <td>2.0x2.0</td> <td>2.0x2.0</td> <td>2.0x2.0</td> <td>2.0x2.0</td> <td>2.0x2.0</td> <td>2.0x2.0</td> <td>2.0x2.0</td> </tr> <tr> <td>Accelerating Voltage [kV]</td> <td>30</td> <td>30</td> <td>30</td> <td>30</td> <td>30</td> <td>30</td> <td>30</td> </tr> <tr> <td>Power [W]</td> <td>2</td> <td>2</td> <td>2</td> <td>2</td> <td>2</td> <td>2</td> <td>2</td> </tr> <tr> <td>Filter</td> <td>Air</td> <td>LE1</td> <td>LE1</td> <td>LE1</td> <td>LE1</td> <td>Air</td> <td>LE1</td> </tr> <tr> <td>Optical magnification</td> <td>4x</td> <td>4x</td> <td>4x</td> <td>4x</td> <td>4x</td> <td>4x</td> <td>4x</td> </tr> <tr> <td>Detector to sample distance [mm]</td> <td>25.5</td> <td>25.52</td> <td>26.82</td> <td>25.4</td> <td>26.75</td> <td>25.5</td> <td>25.8</td> </tr> <tr> <td>Source to sample distance [mm] </td> <td>10.6</td> <td>10.61</td> <td>11.11</td> <td>10.5</td> <td>11.1</td> <td>10.6</td> <td>10.7</td> </tr> <tr> <td>Exposure time [s]</td> <td>15</td> <td>18</td> <td>20</td> <td>18</td> <td>18</td> <td>15</td> <td>18</td> </tr> <tr> <td>No. of projections</td> <td>5201</td> <td>5201</td> <td>5201</td> <td>5201</td> <td>5201</td> <td>5201</td> <td>5201</td> </tr> <tr> <td>Rotation</td> <td>360</td> <td>360</td> <td>360</td> <td>360</td> <td>360</td> <td>360</td> <td>360</td> </tr> <tr> <td>Binning</td> <td>2</td> <td>2</td> <td>2</td> <td>2</td> <td>2</td> <td>2</td> <td>2</td> </tr> <tr> <td>Total scanning time [h]</td> <td>26</td> <td>33</td> <td>33</td> <td>32</td> <td>32</td> <td>26</td> <td>29</td> </tr> <tr> <td>File-size [GB]</td> <td>2</td> <td>2</td> <td>2</td> <td>2</td> <td>2</td> <td>2</td> <td>2</td> </tr> </tbody> </table> </td> </tr> </tbody> </table>
Data set for "Token-Level Multilingual Epidemic Dataset for Event Extraction"
<p>This is the data for the TPDL 2021 paper "<a href="https://zenodo.org/record/5780020">Token-Level Multilingual Epidemic Dataset for Event Extraction</a>". If you use this resource, please cite the paper:</p> <pre><code>@inproceedings{mutuvi2021dataset, title = "Token-level Multilingual Epidemic Dataset for Event Extraction", author = {Mutuvi, Stephen and Boros, Emanuela and Doucet, Antoine, and Lejeune, Gaël and Jatowt, Adam and Odeo, Moses}, booktitle = "Proceedings of the 25th International Conference on Theory and Practice of Digital Libraries, September 13–17, 2021, TPDL 2021", year = "2021", location = "Online" }</code></pre> <p> </p> <p>This work has been supported by the European Union Horizon 2020 research and innovation programme under grants 825153 (Embeddia) and 770299 (NewsEye).</p>
Data sets used to demonstrate the software MadHitter in the manuscript "The Landscape of Receptor-Mediated Precision Cancer Combination Therapy Via a Single-Cell Perspective"
<p>This is a zip archive of nine single-cell RNASeq data sets used in the manuscript entitled:</p> <p>"The Landscape of Receptor-Mediated Precision Cancer Combination Therapy Via A Single-Cell Perspective" by Saba Ahmadi, Pattara Sukprasert, Rahulsimham Vegesna, Sanju Sinha, Fiorella Schischlik, Natalie Artzi, Samir Khuller, Alejandro A. Schaffer, Eytan Ruppin,</p> <p>The README.txt describes the data sets in detail.</p> <p>The associated software can be found at https://github.com/ruppinlab/madhitter</p>
High quality rainfall data set for Australia's northern Murray-Darling Basin
<p>Monthly precipitation data were obtained from the Bureau of Meteorology’s homogeneous climate record through its climate change site network. (http://www.bom.gov.au/climate/change/?ref=ftr#tabs=Tracker&tracker=site-networks). These data have undergone complex quality control to address inconsistencies and errors. There are 8 precipitation stations comprising Augathella, Cunnamulla, Normandy, Miles, Surat, Bellata, Bingara and Curlewis. Additionally, monthly mean maximum temperature (TMax) and mean minimum temperature (TMin) data were obtained for Charleville, Miles, Thargomindah, St. George. Moree, Inverell, Walgett and Gunnedah from the Bureau of Meteorology’s climate change site network. </p>
Diffuse-optical data set measured with a smartphone-based sensor on Potato Hill, Oregon, USA
<p>This data set contains both raw data and derived data obtained on Potato Hill, Oregon, on December 17th 2021 using a diffuse-optical, smartphone-based sensor. The raw image files have been converted to an uncompressed Adobe-.dng file format, file names indicate whether the file contains data for the blue (405nm) or red (650nm) laser or spatial calibration data using a 9mm x 9mm calibration pattern. Spectral albedo measurements are contained in the subfilder ./Albedo, the raw images in ./Phone. The root directory contains the matlab code (Matlab R2021b) needed for analysis as well as the derived data.</p> <p>For analyzing the raw data set, use "CameraMatchPotatoHillFinal.m". It wraps around the function "CameraAnalysisFinal.m", which performs the image analysis and least-square fit to resorted and rescaled data, employing in turn the model function "theosurfGInf.m". It saves a derived data set (attenuation, absorption and scattering coefficients, albedos, absorption enhancement factor and snow density.</p> <p>The script "Albedo.m" analyzes the derived data set along with measured albedo and simulated albedo deposited in the file "snicar_120ppb.txt". The obtained albedo curves and black carbon mixing ratio are as shown in the below manuscript.</p> <p>If you wish to use this data set please contact Markus Allgaier at markusa@uoregon.edu with a description of the work and any questions so that we may offer guidance in regards to the best usage of our dataset. When using the data set within a publication, please cite:</p> <p>Markus Allgaier & Brian Smith, "A Smartphone-Based Sensor for Measuring the Optical Properties of Snow", in preparation, (2022)</p>
Data set of "Effect of Proximity, Burden, and Position on the Power Quality Accuracy Performance of Rogowski Coils"
<p>The uploaded data set contains all the measurements collected during the research that led to the publication of " "Effect of Proximity, Burden, and Position on the Power Quality Accuracy Performance of Rogowski Coils"</p>
Input features and benchmark data sets for protein complex prediction and E. coli proteome application by AF2Complex
<p>Benchmark data sets of AF2Complex, input features for application to E. coli proteome, and predicted structural models of E. coli Ccm I as described in</p> <p><strong>Predicting direct physical interactions in multimeric proteins with deep learning</strong></p> <p><em>Mu Gao, Davi Nakajima An, Jerry M. Parks, Jeffrey Skolnick</em></p> <ol> <li><a href="https://zenodo.org/api/files/087ae188-6f7a-4a71-a586-bbc7bdcc6843/af2complex_bench.tar.gz">af2complex_bench.tar.gz</a>: Benchmark data sets CP17, Dimer1193 and Oligomer562, including input features for AF2Complex/AF-Multimer, both paired and unpaired MSAs, as well as sequences, experimental structures, and results presented in the AF2Complex work (~90GB de-compressed size)</li> <li><a href="https://zenodo.org/api/files/087ae188-6f7a-4a71-a586-bbc7bdcc6843/ecoli_Ccm_I.tar.gz">ecoli_Ccm_I.tar.gz</a>: Computational models of the<em> E. coli</em> Ccm I system</li> <li><a href="https://zenodo.org/api/files/087ae188-6f7a-4a71-a586-bbc7bdcc6843/ecoli_set.tar.gz">ecoli_sets.tar.gz</a>: Lists of benchmark sets of positive and negative PPIs from <em>E. coli</em></li> <li><a href="https://zenodo.org/api/files/087ae188-6f7a-4a71-a586-bbc7bdcc6843/ecoli_af_fea.tar.gz">ecoli_af_fea.tar.gz: </a>Pre-generated input features of <em>E. coli</em> proteome for protein complex prediction and modeling by AF2Complex (4,429 proteins, ~800 GB de-compressed size). This data set can be used with AF2Complex to probe the interactions of any combinations among the 4,429 proteins of E. coli.</li> </ol> <p> </p>
Data sets for 'Mechanical Compliance of Individual Fractures in a Heterogeneous Rock Mass from Production-type Full-waveform Sonic Data', submitted to JGR: Solid Earth
<p>Synthetic and field FWS data sets are supplied for validation of proposed methods for compliance estimation of individual fractures. Each zip contains ‘ReadMe.txt ‘, which illustrates the files and corresponding parameters. More details about the setup in the submitted paper ‘Mechanical Compliance of Individual Fractures in a Heterogeneous Rock Mass from Production-type Full-waveform Sonic Data’.</p>
Data set for the article 'Temporal scaling in C. elegans larval development'
<p>This directory contains all analyzed data and data analysis scripts to create all figures for the article<br> Filina et al., Temporal scaling in C. elegans larval development, PNAS 2022 119:e2123110119</p>
The Good & Bad bird and face data sets
<p>The Good & Bad data set consists of successful as well as unsuccessful generated samples synthesized by a conditional text-to-image GAN model. It can be used to enhance the propability of synthesizing Good latent vectors.</p>
Data Set of Public Procurement in México (2013-2020)
<p>This repository contains all data sets, manuals, and codes to reproduce the main results shown in the paper "Practices of public procurement and the risk of corrupt behavior before and after the government transition in México" (on revision in EPJ Data Science - Springer.) (arxive: <a href="https://arxiv.org/abs/2108.02653">arXiv:2108.02653</a>)</p>
Atmospheric visibility inferred from continuous-wave Doppler wind lidar, data set
<p>Visibility data from Pershore, UK, between 2018 and 2020</p>
Compound activity data sets for 15 biological targets compiled from the ChEMBL and PubChem databases.
<p>Compound activity data sets for the 15 biological targets are deposited, along with structure-activity relationship matrices IDs. Active compounds were extracted from the ChEMBL database and inactive were from the PubChem database. Details of the data sets are described in the original publication. and the summary of the data sets is given in the readme.txt file. </p>
datacleanr manuscript data sets
<p>The data sets in the archive are used to generate figures for a research article introducing the datacleanr R application.</p>
TRINITY open access data repository data set 2 by Budapest University of Technology and Economics
<p>Horizon 2020 programme supports access to and reuse of research data generated by Horizon 2020 projects through the Open Research Data Pilot (ORDP). To support the validation of scientific results, the pilot focuses on providing access to data needed to validate the scientific results. There are several types of such data, e.g. machine learning data sets, models, measurements, statistical results of experiments, survey outcomes, etc.</p> <p>This deliverable summarizes the data that are expected to be collected in the course of the project and where and how they are stored. The aspect of providing open access to research data (as required by the European Commission’s Open Research Data Pilot, <a href="https://www.openaire.eu/what-is-the-open-research-data-pilot">https://www.openaire.eu/what-is-the-open-research-data-pilot</a>) is addressed in Section 3. Finally, in Section 4 we describe the data sets that were or are expected to be generated within the TRINITY projects and made freely available.</p>
Strateole-2 data set associated to the publication "A seismic network in the stratosphere"
<p>NetCDF files of the pressure and temperature data of TSEN sensors, and GPS coordinates, on board EUROS gondolas of Strateole-2 project (stratospheric balloons deployed during fall 2021).</p> <p>One file per gondola, associated to the 4 balloons detecting the Flores quake (2021/12/14 3:20:35.8 GMT) and to a single balloon detecting the Northern Peru quake (2021/11/28 10:52:25.8 GMT).</p> <p>These data cover one hour before and 2 hours after the quake. The rest of the Strateole-2 data will be released by the project. This SUbset is associated to the publication "A seismic network in the stratosphere.</p>
Data sets used in "Neural network processing of holographic images"
<p>Included are the training, validation, and testing data sets for synthetic holograms (netCDF), the HOLODEC data set containing the RF07 examples (netCDF), and the two splits of manually labeled HOLODEC image tiles (numpy arrays). The source code for using the data sets can be found at https://github.com/NCAR/holodec-ml </p>
Raw Data for the article: Portal vein puncture-related complications during transjugular intrahepatic portosystemic shunt creation: Colapinto needle set vs Rösch-Uchida needle set
<p>Transjugular portal vein puncture is considered the riskiest step in TIPS creation with possible incidence of portal vein puncture-related complications (PVPC). The Colapinto and the Rösch-Uchida needle sets are two different needle sets currently available. To date, there have been no randomized control trials or systematic reviews which compare the incidence of PVPC when using the two different needle sets. The aim of this literature review is to assess the rate of PVPC associated with the different needle sets used in the creation of TIPS. From the described search, 1500 articles were identified and 34 met the inclusion criteria. Outcome measured was the prevalence of PVPC using the different needle sets. Overall 212 (3.6%) PVPC were reported in 5865 patients; 142 (3.5%) reported in 4000 cases using the Rösch-Uchida set and 70 (3.7%) in 1865 patients using the Colapinto set (p = 0.69). PVPC in TIPS creation are not related to the choice of needle set used in the procedure. To our knowledge, this is the first review of its kind, the results of which support the theory that while the rate of PVPC is influenced by many factors, choice of needle set does not seem to be one of them.</p>
Data set for publication: Interfacial Deposition of Titanium Dioxide at the Polarized Liquid– Liquid Interface
<p>The attached files are the raw data used to write the manuscript (.txt; .csv; .jpg): Interfacial Deposition of Titanium Dioxide at the Polarized Liquid– Liquid Interface</p>
Raw data set for Negative Emissions in the Chemical Sector: Lifecycle CO2 Accounting for Biomass and CCS Integration into Ethanol, Ammonia, Urea, and Hydrogen Production.
<p>This repository contains the raw data and code used to generate the results in the paper:</p> <p>Tanzer S.E., Blok K., Ramirez Ramirez A. Negative Emissions in the Chemical Sector: Lifecycle CO2 Accounting for Biomass and CCS Integration into Ethanol, Ammonia, Urea, and Hydrogen Production. 15th International Conference on Greenhouse Gas Control Technologies, GHGT-15. 2021. doi: 10.2139/ssrn.3819778.</p> <p>also published as chapter 4 in the PhD dissertation ”Negative Emissions in the Industrial Sector”. The PhD was the department of Engineering Systems and Services, Faculty of Technology Policy, Management at the Delft University of Technology, between 2017-2022. </p> <p>This is intended to be a record of the exact data and code used to generate the results and graphics used in this publication. It is not necessarily designed for user-friendliness or tested to work on other machines and may contain extraneous data and files.</p> <p>To make use of the python black box modelling library for your own work, please check out the most recent public release, which can be found at https://zenodo.org/record/5800104#.YjUTnC8w30o</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.