Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
3,481
datasets available to search
ShareScore release 0.9.0
Dataset results
3,481 results for “data set”
Data set, combining epidemiological, genetics, and government stringency data of COVID-19 pandemic.
<p>This data set combines epidemiological, genetics, and government stringency data of COVID-19 pandemics, all from open data sources. The sources are: Our World in Data, Worldometer, GISAID-Nextstrain, and the Oxford COVID-19 Government Response Tracker (OxCGRT). The cut off date of the first version is at the end of June 2020. </p> <p>The simple data set is provided as an Excel workbook, where the first, "readme" worksheet describes the details of data of all the worksheets in the data set. </p> <p>This is a working data set, expected to be refreshed over time. </p> <p>Raw data are not cleaned - this collection is a tool to check various hypotheses regarding possible relations among the various data types. Simple visualisations of data relations are provided in a separate sheet. </p> <p> </p>
Southern Ocean Cloud and Aerosol data set: a compilation of measurements from the 2018 Southern Ocean Ross Sea Marine Ecosystems and Environment voyage
<p>Due to its remote location and extreme weather conditions, atmospheric in situ measurements are rare in the Southern Ocean. As a result, aerosol-cloud interactions in this region are poorly understood and remain a major source of uncertainty in climate models. This, in turn, contributes substantially to persistent biases in climate model simulations, numerical weather prediction models and reanalyses. It has been shown in previous studies that in situ and ground-based remote sensing measurements across the Southern Ocean are critical for complementing satellite data sets due to the importance of boundary layer and low-level cloud processes. These processes are poorly sampled by satellite-based measurements which are typically obscured by near-continuous overlying cloud cover observed in this region. Here we provide a comprehensive set of ship-based aerosol and meteorological observations collected on the TAN1802 voyage of R/V Tangaroa across the Southern Ocean, from Wellington, New Zealand, to the Ross Sea, Antarctica. The voyage was carried out from 8 February to 21 March, 2018. The compiled data set provides here includes measurements from a range of instruments, such as (i) meteorological conditions at the sea surface and profile measurements; (ii) the size and concentration of particles; (iii) trace gases dissolved in the ocean surface such as dimethyl sulfide and carbonyl sulfide; (iv) and remotely sensed observations of low clouds. We encourage the scientific community to use these measurements for further analysis and model evaluation studies, in particular, for studies of Southern Ocean clouds, aerosol and their interaction.</p>
Flick SMART multi-catch rodent station and bait station data sets: Council of the city of Sydney, October 2019 to July 2020
<p>Shortly after the enactment of preventative measures aimed at limiting the spread of COVID-19, local governments and public health authorities around the world reported an increased sighting of rats. We combined multi-catch rodent station data, rodent bait stations data, and rodent-related residents' complaints data to explore the effects that social distancing and lockdown measures might have had on the rodent population within the City of Sydney, Australia. We found that rodent captures, activity, and rodent related residents' complaints increased during the COVID-19 related lockdown period, followed by a steep decline post-lockdown. We found no changes in the geographical distribution of any of our indices of rodent abundance. We hypothesize that lockdown measures resulted in an increase in rodent activity driven by a reduction in human-derived food resources. This might have increased the mortality rate triggering a population crash. There is a high chance that the surviving individuals might be rodenticide resistant. It is possible that the onset of COVID-19 might have disrupted commensal rodent populations, with profound implications for the future management of these species. Here we make available multi-catch rodent station data and rodent bait stations data. We do not include rodent-related residents' complaints data due to potential identifier data that could be seen as a breach of private information sharing.</p>
Data from: Morphological characters can strongly influence early animal relationships inferred from phylogenomic data sets
<p>There are considerable phylogenetic incongruencies between morphological and phylogenomic data for the deep evolution of animals. This has contributed to a heated debate over the earliest-branching lineage of the animal kingdom: The sister to all other Metazoa (SOM). Here we use published phylogenomic datasets (∼45,000-400,000 characters in size with ∼15-100 taxa) that focus on early metazoan phylogeny to evaluate the impact of incorporating morphological datasets (∼15-275 characters). We additionally use small exemplar datasets to quantify how increased taxon sampling can help stabilize phylogenetic inferences. We apply a plethora of common methods, i.e. likelihood models and their "equivalent" under parsimony: character weighting schemes. Our results are at odds with the typical view of phylogenomics, i.e., that genomic-scale datasets will swamp out inferences from morphological data. Instead, weighting morphological data 2-10× in both likelihood and parsimony can in some cases "flip" which phylum is inferred to be the SOM. This typically results in the molecular hypothesis of Ctenophora as the SOM flipping to Porifera (or occasionally Placozoa). However, greater taxon sampling improves phylogenetic stability, with some of the larger molecular datasets (>200,000 characters and up to ∼100 taxa) showing node stability even with ≧100× up-weighting of morphological data. Accordingly, our analyses have three strong messages. A) The assumption that genomic data will automatically "swamp out" morphological data is not always true for the SOM question. Morphological data have a strong influence in our analyses of combined datasets, even when outnumbered thousands of times by morphological data. Morphology therefore should not be counted out a priori. B.) We here quantify for the first time how the stability of the SOM node improves for several genomic datasets when the taxon sampling is increased. C.) The patterns of "flipping points" (i.e., the weighting of morphological data it takes to change the inferred SOM) carry information about the phylogenetic stability of matrices. The weighting space is an innovative way to assess comparability of datasets that should be developed into a new sensitivity analysis tool.</p>
Data set with reference scenarios
<p>Data set with reference scenarios. As it is not possible to include the entire dataset in this report, we only include two Tables on final energy demand. Table 1 shows the Final Energy Demand projections per industrial subsector and EU28 country in the Reference Scenario and Table 2 the Final Energy Demand projections per industrial subsector and EU28 country in the Frozen Efficiency Scenario. The full dataset, including physical production (in ktonnes) and fuel and electricity demand (in TJ) per industrial sub-sector, per fuel type and per EU 28 country is available upon request to the project coordinator.</p>
Data set on Energy Efficiency potentials on top of reference scenarios
<p>As it is not possible to include the entire dataset in this report, we only include the Energy Efficiency potentials for Austria. The full dataset, showing the savings potentials for all EU27 (and UK) countries, is available upon request to the project coordinator.</p>
VOC analysis data set for polyfilament analysis of antibacterial PE, PA and PLA fibre samples with natural additive rosin and silver
<p>This data is a detail report related to the VOC measurements described and analyzed in the work<br> Weathering and safety of antibacterial polymer-rosin polyfilaments. The detailed description of the<br> method, samples preparations and main outcomes are given in the work. Here, Tables 1-7 indicate the<br> average and (range) concentration of emitted compounds from each polyfilament fibre sample analyzed<br> at different temperatures for their VOC emissions.</p>
Data Sets: An Assessment of Water Trusts: Drinking Water Quality and Provision in Six Low-income, Peri-urban Communities of Lusaka, Zambia
<p>Data set associated with the manuscript published in GeoHealth titled An Assessment of Water Trusts: Drinking Water Quality and Provision in Six Low-income, Peri-urban Communities of Lusaka, Zambia. This includes bacterial, nitrate, specific conductance, and Water Trust survey data collected from Lusaka, Zambia in 2013, 2014, 2016, an 2019.</p>
Token-based data sets for the analysis of the academic language of literary studies and linguistics
<p>These are the token-based data sets used for my PhD thesis ("Potentiale syntaktischer Annotationen für die datengeleitete Sprachbeschreibung am Beispiel der Wissenschaftssprachen der Germanistik", publication in progress).</p> <p>For python scripts and further data see https://github.com/melandresen/dissertation.</p> <p>Due to copyright law, the annotated texts of the corpus could only be published without the token layer. The files provided here include the token-based frequency data that have been derived from the origial texts and can be used as input to the analysis scripts in the GitHub-Repository.</p>
Data sets - Polish version of SARC-F
<p>These data sets correspond to the paper entitled Polish version of SARC-F to assess sarcopenia in older adults: an examination of reliability and validity</p>
VOC chromatogram data set for polyfilament analysis of antibacterial polyfilament fibre samples with natural additive rosin and silver
<p>This data is a detail report related to the VOC measurements described and analyzed in the paper Weathering of antibacterial melt-spun polyfilaments modified by pine rosin. The detailed description of the method, sample codes, samples preparations and main outcomes are given in the manuscript. Figures (1-7) are the direct graphs from the device software as ‘chromatograms’ (TIC, total ion chromatograms) for each polyfilament fibre sample analyzed at different temperatures. The graphs indicate the measured (computed) ion count (abundance, arbitrary units) as a function of time (minutes).</p> <p> </p>
Fiji/CellProfiler cell migration timelapse data set and code
<p>This zenodo upload consists of:</p> <p>Fiji IJMacro script create_LabelledMasks.ijm<br> CellProfiler pipeline TrackMate_CellProfiler.cppipe<br> Matlab script Plot_per_cell.m <br> <br> Saved manually curated TrackMate project crop_1_60_ManualCuration.xml<br> Saved Spots Results Table of manually curated TrackMate project: Spots in tracks statistics.csv<br> CellProfiler pipeline output cp_output.zip</p> <p>Example data set crop_1_60.tif - subset of image data previously described in:</p> <p><strong><a href="https://www.zotero.org/google-docs/?9YEDuq">Shafqat-Abbasi, H., Kowalewski, J. M., Kiss, A., Gong, X., Hernandez-Varas, P., Berge, U., Jafari-Mamaghani, M., Lock, J. G., and Strömblad, S. 2016. An analysis toolbox to explore mesenchymal migration heterogeneity reveals adaptive switching between distinct modes. eLife 5:e11384–e11384.</a></strong><br> Many thanks to Staffan Strömblad et al. for sharing the data.</p>
Data sets comparing Ca(2+) dynamics in whole-cell and β-escin-based perforated patch clamp recordings in adult mouse brain slices.
<p># Hess-et-al-beta-escin-2020<br> Data of the "Data in Brief" article concerning the added buffer approach with beta-escin perforated patch.</p> <p>A joint work by: Simon Hess (`simon.hess@uni-koeln.de`), Christophe Pouzat (`christophe.pouzat@math.unistra.fr`), and Peter Kloppenburg (`peter.kloppenburg@uni-koeln.de`).</p> <p>## Content</p> <p>This repository contains:</p> <p>- Directory `data_whole_cell` contains the experimental data in [HDF5](https://en.wikipedia.org/wiki/Hierarchical_Data_Format) format. The data contains recordings of _Substantia nigra_ dopaminergic neurons recorded in the whole-cell configuration.<br> - Directory `data_beta_escin` contains the experimental data in [HDF5](https://en.wikipedia.org/wiki/Hierarchical_Data_Format) format. The data contains recordings of _Substantia nigra_ dopaminergic neurons recorded in the β-escin perforated patch clamp configuration.</p>
Active Learning with RESSPECT: Data Set
<p>This folder contains pre-processed simulated data first made available by Rick Kessler for the <br> <a href="https://arxiv.org/abs/1008.1024">Supernova Photometric Classification Challenge (SNPCC)</a>.</p> <p>All data were feature extracted using the <a href="https://arxiv.org/pdf/0904.1066.pdf">Bazin parametric function</a>.</p> <p>This version of the data set was used to obtain the results reported in <a href="https://arxiv.org/pdf/2010.05941.pdf">Kennamer et al., 2020 - <em>Active learning with RESSPECT: resource allocation for extragalactic astronomical transients</em>.</a> Published during the <a href="http://www.ieeessci2020.org/symposiums/ciastro.html">2020 IEEE Symposium Series on Computational Intelligence</a>. The code used to obtain the results shown in the paper is available in the <a href="https://github.com/COINtoolbox/RESSPECT">COINtoolbox</a> (github). <br> <br> This work was developed under the <a href="https://cosmostatistics-initiative.org/resspect/">RESSPECT project</a>, an inter-collaboration agreement established between the <a href="https://lsstdesc.org/">LSST Dark Energy Science Collaboration (LSST-DESC)</a> and the <a href="https://cosmostatistics-initiative.org/">Cosmostatistics Initiative (COIN)</a> in order to develop an active learning pipeline to advise the allocation of telescope resources.</p>
Data set related to the manuscript "Simulations of ionic liquids confined in surface-functionalized nanoporous carbons: Implications for energy storage"
<p>Graphical files in the agr format for all the figures in the manuscript entitled "Simulations of ionic liquids confined in surface-functionalized nanoporous carbons: Implications for energy storage". Examples of input files for the three systems simulated.</p>
K2 Simulation Data Set
<p>Simulation results of rainfall-runoff events over the upper Arroyo Seco Basin using KINEROS2 described in the paper "The timing and magnitude of changes to Hortonian overland flow at the watershed scale during the post-fire recovery process" by <strong>Tao Liu</strong><strong>, Luke A. McGuire, Haiyan Wei, Francis K. Rengers, Hoshin Gupta, Lin Ji, David C. Goodrich </strong>submitted to Hydrological Processes.</p>
80 Initial Data-set Studies SMS Strategy CG
<p>Initial Data set of 80 primary studies SMS: "Empirical Strategies in Software Engineering Research: A Literature Survey", by Cathy Guevara-Vega<br> </p>
data set regarding to project –Polish version of Mini Sarcopenia Risk Assessment (MSRA) Questionnaire
<p><strong>This data set corresponds with the article titled Polish translation and validation of The Mini Sarcopenia Risk Assessment (MSRA) Questionnaire to assess nutritional and non-nutritional risk factors of sarcopenia in older adults.</strong></p>
Water quality data sets of the Gozyo catchment and the four U.S. watersheds
<p>This zipped file contains three excel files. Table GZ_Solute.xls and Table GZ_Tb.xls are high-frequency water quality data in the stream of the small forested catchment, Gozyo cattchment (12.14ha), Nara, Japan. The high-frequency on-site Water quality (WQ) measurements of K, Cl, and Na in the stream using flow-injection potentiometry method with 15-minute interval were interpolated to 10-minute interval data. The SS concentration in the stream was estimated by linear transformation from the observed 10-minute turbidity value. Table GZ_Solute.xls gives observed 10-minute solutes data with discharge, rainfall, and temperature from 2009 to 2011, and Table GZ_Tb.xls provides the observed 10-minute turbidity data from 2012 to 2014. SS value (mg/l) is calculated as 1.09 Turbidity + 6.19. The data value of -999 means missing observation.</p> <p>Tables US_daily_WQ_datasets.xls contains daily water quality (WQ) and discharge data from the four watersheds in the United States. The data value of -999 also means missing observation. The names (USGS station numbers) of four WQ monitoring sites are Muskingum River (03150000), Rock Creek (04197170), Skunk River (05474000), and Vermilion River (04195000). Daily WQ data were composed by following the procedure described in the Appendix S1 in Hirsch (2014). The WQ data for Muskingum, Rock, and Vermilion river were retrieved from Heidelberg University's National Center for Water Quality Research site (https://ncwqr.org/monitoring/data) and suspended sediment discharge data were downloaded from the USGS National Water Information System (<a href="http://waterdata.usgs.gov/nwis/">http://waterdata.usgs.gov/nwis/</a> or https://doi.org/: 10.5066/F7P55KJN). All the daily discharge data were acquired via the USGS National Water Information System.</p> <p>These data were used to evaluate the performance of the unbiased load estimates and confidence intervals of river loads based on the Rating curve method using importance sampling in the listed article below. To maintain the traceability of the proposed load estimation method and replicability the results in the article, the authors of the article upload the data used in this repository.</p> <p> </p> <p>References</p> <p>Hirsch, R. M. (2014). Large biases in regression-based constituent flux estimates: causes and diagnostic tools. <em>Journal of the American Water Resources Association</em>, 50(6), 1401-1424. https://doi.org/10.1111/jawr.12195</p> <p>Tada, A. and H. Tanakamaru. (Submitted) Unbiased estimates and confidence intervals for riverine loads, <em>Water Resources Research</em>.</p>
Data Sets and Code for Simple formulas for pseudoposition for electrical resistivity and IP in vertical boreholes based on mean positions of the sensistivity
<p>Data and Matlab code in figures.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.