Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
3,481
datasets available to search
ShareScore release 0.9.0
Dataset results
3,481 results for “data set”
Methylation-free E.coli nanopore sequencing (ONT R9.4.1) data set
<p>The data set consists of fast5 files divided into 5 zip files (fast5_[1-5].zip), a genome record (Ecoli_K12_MG1655.fasta), an Illumina assembly genome (illumina_contigs.fasta) and a fastq file from Guppy 5 (guppy_basecalled.fastq.gz). We sequenced the Ecoli non-methylated genomic DNA (D5016, Zymo Research) with an ONT MinION device. The sequencing libraries were prepared by fragmenting the genomic DNA using Covaris g-TUBE and a Ligation sequencing kit (SQK-LSK109, Oxford Nanopore) with Flow Cell chemistry R9.4.1. We also performed short-read Illumina sequencing on the same sample using the TruSeq PCR-free library preparation on the MiSeq sequencing platform (Illumina, USA), and constructed a draft assembly from the Illumina sequencing results using SPAdes v3.6.0. We also upload a reference genome directly obtained from the E.coli sample producer website. </p> <p>In addition, the data set contains two fastq files that produced by the Lokatt basecaller (lokatt_basecalled.fasta.gz) and local-trained Bonito basecaller (bonito_local_basecalled.fastq.gz ), respectively, which are used for benchmarking in the Lokatt basecaller paper.</p>
Data set used in "Effect of turbulence and viscosity models on wall shear stress derived biomarkers for aorta simulations"
<p>Data set used in "Effect of turbulence and viscosity models on wall shear stress derived biomarkers for aorta simulations"</p> <p>Includes the data for 20 heartbeats. Divided into external and internal walls regions. </p>
Pyrolysis Model Data Set Contribution for the MaCFP-3 Workshop October 2023
<p>Pyrolysis model data set contribution for the <a href="https://iafss.org/macfp/">MaCFP-3 workshop October 2023</a></p> <p>The contributions are also available via the materials database <a href="https://github.com/MaCFP/matl-db">MaCFP git repository.</a></p> <p>For the inverse modelling in general, please see the manuscript <a href="https://doi.org/10.48550/arXiv.2303.17446">"PMMA Pyrolysis Simulation -- from Micro- to Real-Scale"</a></p> <p> </p> <p>There are three different ZIP archives provided:</p> <ul> <li>"MaCFP_Gasification" contains the simulations reproducing the NIST Gasification Apparatus for validation of the parameter sets, as requested per MaCFP. There are three different runs conducted, two for the originally submitted parameter sets (BUW-FZJ Approach A/B). No changes to the parameter sets are conducted. A new parameter set is provided to the MaCFP repository (BUW-FZJ PMMA01) and its validation simulations are provided as well. Furthermore, Jupyter notebooks are provided to process the simulation data.</li> <li>"PyrolysisKinetics_PMMA01" contains the inverse modelling run to determine the pyrolysis reaction kinetics.</li> <li>"Thermophysical_PMMA01" contains the inverse modelling run to determine the thermophysical parameters.</li> </ul>
Evaluating a poroelastic model via pore pressure signals in seafloor sediments [data set]
<p>The csv files consist of pressure data collected off the coast of Camp Pendleton between 10 February 2021 and 25 February 2021, and include both pore pressure data from two instrumented surrogates and pressure data from a Nortek Signature. Timestamps are in posix time; pressure is in kPa. The time series for Surrogate A are prefixed "surrA"; those for Surrogate B are prefixed "surrB". The Nortek Signature time series is prefixed "Sig1000".</p>
Data set for "Membrane potential dynamics of excitatory and inhibitory neurons in mouse barrel cortex during active whisker sensing"
<p>Data set for: Kiritani T, Pala A, Gasselin C, Crochet S, Petersen CCH (2023) Membrane potential dynamics of excitatory and inhibitory neurons in mouse barrel cortex during active whisker sensing. PLOS ONE 18: e0287174. doi: 10.1371/journal.pone.0287174</p> <p>There are 2 files in this upload:</p> <p>1. The file named "2023_Kiritani_PLOSONE.pdf" is the Open Access pdf of the online publication in PLOS ONE.</p> <p>2. The file named "Kiritani_data_code.zip" (~5 GB) is a zipped version of a folder "Kiritani_data_code" (~5 GB), which contains the data analysed in the study along with the Matlab codes used to generate the published figures. To access the data and codes, first unzip the file. You need to install the Matlab 'Signal Processing' and 'Curve Fitting' Toolboxes. In Matlab, add the path of the folder 'Kiritani_data_code' and all subfolders. Directly from this folder, you should first run the codes in the folder 'Data_Analysis_Codes', sequentially executing 'Analysis_1.m' through to 'Analysis_9.m'. Note, execution of 'Analysis_9.m' can take a long time (~1 hour on a good desktop PC). You can then run the codes in the folder 'Figure_Plotting_Codes' to generate the figures published in the journal article. In the folder 'Data', you can also find a DataViewer to visualise the data sets, which you can run by executing 'DataViewer.m' directly from the subfolder ‘Data’.</p> <p> </p>
Rock magnetic data sets for Coupled detachment faulting and hydrothermal circulation at 49.7°E Southwest Indian Ridge revealed by seafloor magnetism
<p>The rock magnetic data sets in "Coupled detachment faulting and hydrothermal circulation at 49.7°E Southwest Indian Ridge revealed by seafloor magnetism" was studied, including the density, magnetic susceptibility, NRM, Q ratio and other parameters of rock samples , as well as the thermomagnetic curves, hysteresis loops, FORCs, AF and TD demagnetization.</p>
Data set from long-term wave, wind and response monitoring of the Bergsøysund Bridge
<p>Wind, wave, displacement and acceleration data have been collected in a measurement campaign on the Bergsøysund Bridge between the years 2014 and 2018. The data set is now available in this open-access research entry, for free access and download. The data is collected in two h5-files (hierachical data format), with sampling rates 2 Hz and 10 Hz, downsampled from the raw sampling rate of 200 Hz. Note that the data has undergone some minimal signal processing and adjustment, in line with that applied to the Hardanger Bridge data described in Fenerci et al. (2021). Tools and examples for import, data visualization and initial analysis are given in the opyndata Python package available on GitHub (Kvåle, 2022). Furthermore, a document briefly describing the hierarchy and structure of the data, is given. For more details on the measurement system and the bridge, it is referred to Kvåle and Øiseth (2017).</p> <p>The updated, copyrighted version of the appended preprint is published by ASCE with the following DOI: <a href="https://doi.org/10.1061/JSENDH.STENG-12095">10.1061/JSENDH.STENG-12095</a></p>
Simulated brainweb PET/MR data sets for denoising and deblurring
<p>The data set consists of 20 subjects based on the normal anatomical models from the brainweb phantom. See <a href="https://brainweb.bic.mni.mcgill.ca/">https://brainweb.bic.mni.mcgill.ca/</a>.</p> <p>For each subjectXX the data is organized as follows:</p> <ol> <li>image_0.nii.gz -> (1st simulated "random" PET contrast)</li> <li>image_1.nii.gz -> (2nd simulated "random" PET contrast)</li> <li>image_2.nii.gz -> (3rd simulated "random" PET contrast)</li> <li>attenuation_image.nii.gz -> (attenuation image)</li> <li>t1.nii.gz -> (high resolution T1 MR)</li> </ol> <p>All data sets have a shape of (220,220,184) and a voxel size of 1mm x 1mm x 1mm and are provided in nifti format.<br> The T1 MR scans and antomical models used to create the ground truth PET images are taken from the BrainWeb: Simulated Brain Database.</p> <p>See:</p> <p> http://www.bic.mni.mcgill.ca/brainweb/<br> C.A. Cocosco, V. Kollokian, R.K.-S. Kwan, A.C. Evans : "BrainWeb: Online Interface to a 3D MRI Simulated Brain Database"<br> NeuroImage, vol.5, no.4, part 2/4, S425, 1997 -- Proceedings of 3-rd International Conference on Functional Mapping of the Human Brain, Copenhagen, May 1997.<br> R.K.-S. Kwan, A.C. Evans, G.B. Pike : "MRI simulation-based evaluation of image-processing and classification methods"<br> IEEE Transactions on Medical Imaging. 18(11):1085-97, Nov 1999.<br> R.K.-S. Kwan, A.C. Evans, G.B. Pike : "An Extensible MRI Simulator for Post-Processing Evaluation"<br> Visualization in Biomedical Computing (VBC'96). Lecture Notes in Computer Science, vol. 1131. Springer-Verlag, 1996. 135-140.<br> D.L. Collins, A.P. Zijdenbos, V. Kollokian, J.G. Sled, N.J. Kabani, C.J. Holmes, A.C. Evans : "Design and Construction of a Realistic Digital Brain Phantom"<br> IEEE Transactions on Medical Imaging, vol.17, No.3, p.463--468, June 1998</p>
Data set for the journal article: Social life cycle assessment of green methanol and benchmarking against conventional fossil methanol
<p>Single File containing:</p> <ul> <li>Green Methanol Inventories: numerical data as displayed in Figure 4, Main social life cycle inventory data of the green methanol system. </li> <li>Conventional Methanol Inventories: numerical data as displayed in Figure 5, Main social life cycle inventory data of the conventional methanol system. </li> <li>Supplementary information: Diagrams and tables describing teh flowsheet of the simulations used in this work: <ul> <li> <p>Green methanol production process (flowsheet and stream table)</p> </li> <li> <p>Syngas production through Steam Methane Reforming (flowsheet and stream table)</p> </li> <li> <p>Conventional methanol production process (flowsheet and stream table)</p> </li> </ul> </li> </ul>
Multiparameter Water Quality Monitoring System for Continuous Monitoring of Fresh Waters Calibration and Measurement Data Set
<p>This data set contains calibration data for all sensors incorporated in the sensor node. It provides comparison measurements of TPL fluorescence taken by the node and reference spectrofluorimeter. Initial test measurements, as well as site measurements, are also provided. Finally, data from a heuristic method of TPL detection in the presence of algae and mud are also given.</p>
Data set - Measured in a context : making sense of open access book data
<p>For more than a decade, open access book platforms have been distributing titles in order to maximise their impact. Each platform offers some form of usage data, showcasing the success of their offering. However, the numbers alone are not sufficient to convey how well a book is actually performing.</p> <p>Our data set is consists of 18,014 books and chapters. The selected titles have been added to the OAPEN Library collection before 1 January 2022, and the usage data of twelve months (January to December 2022) has been captured. During that period, this collection of books and chapters has been downloaded more than 10 million times. Each title has been linked to one broad subject and the title’s language has been coded as either English, German or other languages.</p> <p>The titles are rated using the TOANI score.</p> <p>The acronym stands for Transparent Open Access Normalised Index. The transparency is based on the application of clear regulations, and by making all data used visible. The data is normalised, by using a common scale for the complete collection of an open access book platform. Additionally, there are only three possible values to score the titles: average, less than average and more than average. This index is set up to provide a clear and simple answer to the question whether an open access book has made an impact. It is not meant to give a sense of false accuracy; the complexities surrounding this issue cannot be measured in several decimal places.</p> <p>The TOANI score is based on the following principles:</p> <ul> <li>Select only titles that have been available for at least 12 months;</li> <li>Use the usage data of the same 12 months period for the whole collection;</li> <li>Each title is assigned one – high level – subject;</li> <li>Each title is assigned one language;</li> <li>All titles are grouped based on subject and language;</li> <li>The groups should consists of at least 100 titles;</li> <li>The following data must be made available for each title: <ul> <li>Platform</li> <li>Total number of titles in the group</li> <li>Subject</li> <li>Language</li> <li>Period used for the measurement</li> <li>Minimum value, maximum value, median, first and third quartile of the platform’s usage data</li> </ul> </li> <li>Based on the previous, titles are classified as: <ul> <li>“Less than average” – First quartile; 25 % of the titles</li> <li>“Average” – Second and third quartile; 50% of the titles</li> <li>“More than average” – Fourth quartile; 25 % of the titles</li> </ul> </li> </ul>
Data platform (genotyping data set) related to ERDF postdoctoral project No. 1.1.1.2/VIAA/4/20/718 "The role of vitamin D gene polymorphisms and its receptors in the modulation of intestinal inflammation in patients with relapsing and progressive forms of multiple sclerosis".
<p><strong>Data platform </strong><strong>(genotyping dataset)</strong> <strong>related to the ERDF postdoctoral project No. </strong><strong>1.1.1.2/VIAA/4/20/718</strong><strong> “</strong><strong>The role of vitamin D and its receptor gene polymorphisms in the modulation of intestinal inflammation in patients with relapsing and progressive forms of multiple sclerosis</strong><strong>”.</strong></p> <p><strong>About the project and gathered data:</strong></p> <p>The dataset contains genotyping data on 289 sex-balanced samples (approximately 60% women / 40% men)) were created at the the multiple sclerosis (MS) Clinic of the Latvian Maritime Medical Center (LMMC) in 2011 (disease duration of 1-51 years); the collection was updated within the framework of the ERDF MS project (2017-2020) and replenished during the ERDF postdoctoral project No. 1.1.1.2/VIAA/4/20/718 “The role of vitamin D and its receptor gene polymorphisms in the modulation of intestinal inflammation in patients with relapsing and progressive forms of multiple sclerosis” (2021-2023).</p> <p>For the <strong>Genotyping dataset </strong>relevant information for each patient from the MS disease cohort, referring to proteasomal gene genetic variations (microsatellites and SNPs): (HSMS006 <em>(PSMA6),</em> HSMS602 <em>(FAM177A1),</em> HSMS701 <em>(KIAA0391)</em>, HSMS702 <em>(KIAA0391)</em> HSMS801 <em>(KIAA0391)</em>, rs11543947<em>(PSMB5), </em>rs2277460 (mi110), rs1048990 (mi8)<em> (PSMA6),</em> rs1048990 (mi8)<em> (PSMA6),</em> rs2295826/rs2295827<em>(PSMC6),</em> rs2348071 <em>(PSMA3),</em> rs2071543, rs9357155 <em>(PSMB8),</em> rs17587<em>(PSMB9),</em> rs74421874 <em>(PSMD9); </em>rs9275596 from HLA region; vitamin D-related genes (VDR and GC) polymorphisms: rs2228570, rs1544410, rs7975232, rs731236 (<em>VDR</em>) and rs7041, rs4588 <em>(GC).</em></p>
Data set for the ensemble postprocessing of 2m surface temperature forecasts in Germany for 24 hours lead time
<p>Full data set for the ensemble postprocessing of 2m surface temperature forecasts at 462 observation stations in Germany for 24 hours lead time in the years 2015-2020. The data set is provided in .Rdata format supported by the statistical software <a href="https://www.r-project.org">R</a>. The ensemble forecasts are retrieved from <a href="https://www.ecmwf.int">ECMWF</a> and the observation data from the <a href="https://opendata.dwd.de/climate_environment/CDC/observations_germany/climate/hourly/air_temperature/historical/BESCHREIBUNG_obsgermany_climate_hourly_tu_historical_de.pdf">German Weather Service</a> (<a href="https://www.dwd.de/">DWD</a>). <br> <br> For more information about the data set see: <a href="https://github.com/jobstdavid/paper_gamvinereg">https://github.com/jobstdavid/paper_gamvinereg</a></p>
Data Set For Efficient Calculation of Dispersion Energy for Multireference Systems with Cholesky Decomposition. Application to Excited-state Interactions
<p>Data Set to Accompany:</p> <p>"Efficient Calculation of Dispersion Energy for Multireference Systems with Cholesky Decomposition. Application to Excited-state Interactions"</p>
Demontration Activities Data Set on the performances of the Photocatalytic pilot plant In Demosite 4 (Galeb, Omis, Croatia) Deliverable 5.6
<p>Raw data (TOC and emerging contaminants time profiles) concernig the validation experiment of the photocatalytic pilot plant for tertiary water treatment deployed in ProjectO demosite 4. The data are presented in D5.6 and pertain three months of demonstration activities</p>
Data set on firm characteristic and cash holding in Nigeria
<p>This data set is for research on firm characteristics and cash holding. The data comprise panel data for variables such as return on asset, asset tangibility, leverage, capital expenditure, dividend, growth, cash, firm size, firm age and foreign ownership, for 600 firm-year. </p>
WorldFloods extended data set
<p><strong>"Global Flood Extent Segmentation in Optical Satellite Images"</strong> article data.</p><p>This data set is the extended version of the <strong>WorldFloods </strong>dataset released by<a href="https://www.nature.com/articles/s41598-021-86650-z"> Mateo-Garcia et al. (2021)</a>. We filtered low-quality floodmaps, extended the period of coverage to include flood events up to 2023, and manually fixed the labels of several flood maps. The resulting dataset has 509 flood extent maps from 144 different flood events.</p><p>The flood extent masks were visually inspected, and manually corrected when necessary, in order to provide reliable data to train supervised ML algorithms for flood extent segmentation. Here we provide the <strong>floodmaps.zip </strong>with vectorized reference masks, containing polygons of flood water, permanent water, clouds, and area of interest for each flood map.</p><p>Additionally, the <strong>metadatas.zip </strong>contains all the necessary information to download corresponding Sentinel-2 images, as well as the location of each flood event and activation code (according to Copernicus EMS, UNOSAT, or GLOFMIR conventions).</p><p>Portalés-Julià, E., Mateo-García, G., Purcell, C., & Gómez-Chova, L. Global flood extent segmentation in optical satellite images. <i>Sci Rep</i> <strong>13</strong>, 20316 (2023). https://doi.org/10.1038/s41598-023-47595-7</p><p>This dataset is released under a Creative Commons non-commercial license (https://creativecommons.org/licenses/by-nc/4.0/legalcode.txt) </p><p>The development of this dataset has been supported by the Spanish Ministry of Science and Innovation project PID2019-109026RB-I00 (MINECO-ERDF MCIN/AEI/10.13039/501100011033).</p>
Data set for manuscript 'Quantifying geomorphically effective floods using satellite observations of river mobility'
<p>Data underlying the plots / used in the modelling work for the paper 'Quantifying geomorphically effective floods using satellite observations of river mobility', submitted to <em>Geophysical Review Letters.</em></p>
assessment of aldehydes to PTR-MS m/z 69 in indoor air measurements - data set
<ul> <li>contact: Lisa Ernle (lisa.ernle@mpic.de), Nijing Wang (nijing.wang@mpic.de), Jonathan Williams (jonathan.williams@mpic.de)</li> <li>instruments: fast GC-MS SOFIA (MPIC), PTR-ToF-MS 8000 (Ionicon)</li> <li>merged dataset</li> <li>calibrated with VOC standard gas mix (Apel-Riemer Environmental Inc., Colorado, USA)</li> <li>units (filename): <ul> <li>normalized counts per second [ncps] (20210426_p_ncps.txt, bar_mean.txt, bar_std.txt)</li> <li>parts per billion [ppb] (all_sub_20210426_ppb.txt)</li> </ul> </li> <li>for information concerning updated versions, please see ReadMe.txt</li> </ul>
TEAMx-PC22 (TEAMx pre-campaign 2022) – DWD Doppler wind lidar data set (SLXR172)
<p>This dataset contains data measured by DWD with a Doppler Wind Lidar SLXR172 during the TEAMx pre-campaign 2022. More details about TEAMx can be found at <a href="http://www.teamx-programme.org/">http://www.teamx-programme.org</a>.</p> <p><strong>DATA SET DESCRIPTION</strong></p> <p><strong>1. Measurement location and time period </strong></p> <p>Measurements with the SLXR172 were collected at the site of Brannenburg (47.741547 N / 12.122187 E / 456 m MSL) between 15.June – 19 October 2022.</p> <p><strong>2. Measurement setup</strong></p> <p>During the measurement period, two different scanning modes were applied:</p> <p>15. June - 26. July 2022 and 13. August – 19. October 2022 (VAD_CSM).</p> <ul> <li><strong>VAD (velocity-azimuth display) scans in </strong><strong>continuous scanning mode</strong><strong> : </strong>These scans were conducted at an elevation angle of 35°. Azimuth angle interval of the CSM data sampling was about 1.1°. </li> </ul> <p>27.July – 12. August 2022 (VAD_RHI)</p> <ul> <li><strong>VAD scans in step-stare mode: </strong>Step-stare scans were conducted at an elevation angle of 35° and with azimuth steps of 15°.</li> <li><strong>RHI (</strong><strong>range-height indicator) scans into the Inn Valley</strong>; The RHI scans were performed for 10 azimuth angles from 151° to 160° and covered elevation angles from 3° to 51°.</li> </ul> <p> </p> <p><strong><em>3. Data processing, corrections and filter</em></strong></p> <p>For <strong><em>VAD scans in continuous scanning mode</em></strong> the processed wind fields are provided. The data have not been corrected. The data can be filtered using the parameters R<sup>2 </sup>(coefficient of determination), CN (condition number) and NVRAD (number of radial velocities) as described in Päschke (2015):</p> <p>R<sup>2</sup>> 0.95 and CN<10 and NVRAD>12 </p> <p>Please note that in the postprocessing of the VAD CSM scans, the R<sup>2</sup> filter criterion was set to R<sup>2</sup>>0 in order to include all data and therefore might also include scans where the assumptions of homogeneity are not fulfilled. The parameter qwind is therefore not meaningful due to this configuration and should not be used to filter the data. We recommend the use of the above criterion from Päschke.</p> <p>For scans from the <strong><em>VAD scans in step stare mode</em></strong> as well as the <strong><em>RHI scans</em></strong> the raw data files are provided. They have not been corrected nor filtered.</p> <p><strong>4. Data file structure</strong></p> <p>The data are provided in NetCDF format. File names contain date and time information in UTC. The following wildcard characters are used in the file examples below: yyyy - year; mm - month, dd - day; HH - hour, MM - minute, `SS` - second. Files are sorted in monthly folders.</p> <p>The data are provided in two zip-files.</p> <ul> <li>VAD_CSM contains the processed wind fields from 15. June - 26. July 2022 and from 13. August – 19. October 2022)</li> <li>VAD+RHI the raw data files for 27.July – 12. August 2022.</li> </ul> <p>Raw data files of the VAD CSM scans can be provided upon request.</p> <p><strong>5. Contact</strong></p> <p>Contact Katrin.sedlmeier(at)dwd.de.at for any questions regarding the data set.</p> <p><strong>6. References</strong></p> <p>Päschke, E., Leinweber, R., and Lehmann, V.: An assessment of the performance of a 1.5 μm Doppler lidar for operational vertical wind profiling based on a 1-year trial, Atmos. Meas. Tech., 8, 2251–2266, https://doi.org/10.5194/amt-8-2251-2015, 2015.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.