Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
3,481
datasets available to search
ShareScore release 0.9.0
Dataset results
3,481 results for “data set”
Data set for: Research trends of Academic Performance and Social Networking Sites
<p>Data set for: Research trends of Academic Performance and Social Networking Sites.</p>
Agreement rate data set for GENETEX manuscript
<p><strong>Objectives:</strong> Clinico-Genomic Data (CGD) acquired through routine clinical practice has the potential to improve our understanding of clinical oncology. However, these data often reside in heterogeneous and semi-structured data, resulting in prolonged time-to-analyses.<br> <strong>Materials and Methods:</strong> We created GENETEX: an R package and Shiny application for text mining genomic reports from EHR and direct import into REDCap<sup>®</sup>.<br> <strong>Results:</strong> GENETEX facilitates the abstraction of CGD from EHR and streamlines capture of structured data into REDCap<sup>®</sup>. Its functions include natural language processing of key genomic information, transformation of semi-structured data into structured data and importation into REDCap. When evaluated with manual abstraction, GENETEX had >99% agreement and captured CGD in approximately one-fifth the time.<br> <strong>Conclusions:</strong> GENETEX is freely available under the Massachusetts Institute of Technology license and can be obtained from GitHub. GENETEX is executed in R and deployed as a Shiny application for non-R users. It produces high-fidelity abstraction of CGD in a fraction of the time.</p>
Data sets associated with "Deep structure of the lunar South Pole–Aitken basin"
<p>These grid files contain the data used in Figures 2 and 3 of the submitted manuscript, "Deep structure of the lunar South Pole–Aitken basin". Although the data are global, note that they are not meant to be used outside of the South Pole–Aitken basin.</p>
Data Set of Publication on Detecting Violations of Access Control and Information Flow Policies in Data Flow Diagrams
<p>This data set contains the evaluation data of our publication on "<a href="https://doi.org/10.1016/j.jss.2021.111138">Detecting Violations of Access Control and Information Flow Policies in Data Flow Diagrams</a>". The data set contains an overview on our implementation, description of the evaluation cases, and the evaluation results.</p> <p>The application to produce this data is available in another data set, which is linked to this data set.</p>
Adopt a Pixel 3 km: A Multiscale Data Set Linking Remotely Sensed Land Cover Imagery with Field Based Citizen Science Observation
<p>These datasets were used in an article submitted to the journal Frontiers in Climate in 2021: <a href="https://www.frontiersin.org/articles/10.3389/fclim.2021.658063/full">https://www.frontiersin.org/articles/10.3389/fclim.2021.658063/full</a></p> <p>Further supplemental links (including general information about GLOBE data) can be accessed at <a href="https://observer.globe.gov/get-data/mosquito-habitat-data">https://observer.globe.gov/get-data/mosquito-habitat-data</a>.</p>
Franck-Hertz experiment , 178 C processed data set
<p>For PHYS 3220 lab, <strong>Franck-Hertz experiment for Mercury, 178 C processed data set.</strong></p> <p>An experiment was conducted on 24th September 2021 at York University. </p>
Mimosa catherinensis SNP data sets
<p>To inform management strategies for conservation of <i>Mimosa catharinensis </i>– a narrow endemic, critically endangered plant species – we identified 1,497 unlinked SNP markers derived from a reduced representation sequencing method (i.e., ddRADseq). This set of molecular markers was employed to assess intrapopulation genetic parameters and the demographic history of one extremely small population of <i>M. catharinensis </i>located in the Brazilian Atlantic Forest. We observed a moderate level of genetic diversity for <i>M. catharinensis</i>. Interestingly, <i>M. catharinensis</i>, which is a lianescent shrub with no indication of seed production for at least two decades, presented high levels of outcrossing and no evidence of inbreeding. However, the reconstruction of demographic history of <i>M. catharinensis</i> indicate that the population should be suffered a recent bottleneck.</p>
Hydrogen Combustion Data Set (QCArchive View Formatted)
<p>Data curated by the QCArchive team. Original source is:</p> <p>A Benchmark Data Set for Hydrogen Combustion<br> https://doi.org/10.34974/0jr1-pb24</p> <p>Authors: Das, Akshaya and Stein, Christopher J. and Heidar-Zadeh, Farnaz and Bertels, Luke and Liu, Meili and Guan, Xingyi and Haghighatlari, Mojtaba and Li, Jie and Zhang, Oufan and Hao, Hongxia and Leven, Itai and Head-Gordon, Martin and Head-Gordon, Teresa</p>
Microsatellites data set: Correlated population genetic structure in a three-tiered host-parasite system: the potential for coevolution and adaptive divergence
<p><span><span><span><span><span><span><span><span><span><span><span>Three subspecies of Northern Bahamian Rock Iguanas, <i>Cyclura cychlura</i>, are currently recognized: <i>C. c. cychlura,</i>restricted to Andros Island, and <i>C. c. figginsi</i> and <i>C. c. inornata,</i> native to the Exuma Island chain. Populations on Andros are genetically distinct from Exuma Island populations, yet genetic divergence among populations in the Exumas is inconsistent with the two currently recognized subspecies from those islands. The potential consequences of this discrepancy might include the recognition of a single subspecies throughout the Exumas rather than two. That inference also ignores evidence that populations of <i>C. cychlura</i> are potentially adaptively divergent. We compared patterns of population relatedness in a three-tiered host-parasite system: <i>C. cychlura</i> iguanas, their ticks (genus <i>Amblyomma</i>, preferentially parasitizing these reptiles), and <i>Rickettsia </i>spp. endosymbionts (within tick ectoparasites). Our results indicate that while <i>C. c. cychlura</i> on Andros is consistently supported as a separate clade, patterns of relatedness among populations of <i>C. c. figginsi</i> and <i>C. c. inornata</i> within the Exuma Island chain are more complex. The distribution of the hosts, different tick species, and <i>Rickettsia</i> spp., supports the evolutionary independence of <i>C. c. inornata</i>. Further, these patterns are also consistent with two independent evolutionarily significant units within <i>C. c. figginsi</i>. Our findings suggest coevolutionary relationships between the reptile hosts, their ectoparasites, and rickettsial organisms, suggesting local adaptation. This work also speaks to the limitations of using neutral molecular markers from a single focal taxon as the sole currency for recognizing evolutionary novelty in populations of endangered species.</span></span></span></span></span></span></span></span></span></span></span></p>
Data Set Analisis Pengaruh COVID 19 Terhadap Perilaku Konsumtif Masyarakat dalam Penggunaan E-commerce
<p>Data set jawaban kuesioner Analisis Pengaruh COVID 19 Terhadap Perilaku Konsumtif Masyarakat dalam Penggunaan E-commerce</p>
Data set for "The effects of glycine to alanine mutations on the struc-ture of GPO collagen model proteins"
<p>Data set for "The effects of glycine to alanine mutations on the struc-ture of GPO collagen model proteins"</p> <p>The README files contain a description of all files and scripts in this repo.</p>
Data set for Predicting hospital occupancy for covid-19 patients: a simulation approach based on archetypes of empirical services' trajectories
<p>Data set for the paper: Predicting hospital occupancy for covid-19 patients: a simulation approach based on archetypes of empirical services’ trajectories</p> <p>Based on: Marin-Garcia, J. A., Ruiz, A., Julien, M., & Garcia-Sabater, J. P. (2021). A data generator for covid-19 patients’ care requirements inside hospitals. WPOM-Working Papers on Operations Management, 12(1), 76-115. https://doi.org/10.4995/wpom.15332</p> <p> </p>
Data set of 1) plant abundance in the herb layer and 2) Shrub and tree composition and structure following forest management along a chronosequence
<p>In each site (1200 m²), three circular plots (400 m²) were established (total of 198 plots). Data presented here for 1) plant of the herb layer and 2) shrub and tree are grouped by site (addition of three plots). </p> <p>In the herb layer, total plant species identity and abundance (percentage cover) were sampled in each plot using eight circular micro-plots of 4 m². Data were collected in 2016 and 2017, between June and August. To minimize seasonal variability and allow detection of early spring species, plants (identity and abundance) in the herb layer were measured twice (once in June to early-July, and once in late-July to August).</p> <p>For tree, in each plot, species identity, diameter at breast height (DBH, 1.3 m) and locations of each tree > 9.1 cm DBH were determined. In each plot, species identity and DBH of shrubs and small trees (DBH range: 1.1 to 9.1 cm) were measured in three circular micro-plots of 25 m². Data that were related to the shrub-canopy layer included all trees and shrubs with DBH > 1.1 cm. Here, in each site, forest composition and structure is represented by different combinations of DBH classes (1.1-4 cm, 4-9.1 cm, 9.1-20 cm, 20-35 cm, >35 cm) and species.</p> <p>Plant community composition and abundance were assessed in unmanaged forests (sites of old-growth forest > 100 years, with dominant and co-dominant trees older than 200 years, and no obvious sign of past harvesting), and in even-aged and uneven-aged managed forests along a chronosequence (< 5 years, 15 years, 30 years after forest harvesting).</p>
Data-Error Scaling Laws in Machine Learning on Combinatorial Mutation-prone Sets: Proteins and Small Molecules
<div> </div> <h3>Data</h3> <p> This folder contains the raw data used during this work. `out_seq_total.txt` contains information on the sequences used (mutations, number of mutations, etc.). `output_energies_total.txt` contains the response variables, which include: a) unrelaxed EvoEF energies (peptides); b) relaxed EvoEF energies (peptides, `*_repaired.txt`); and c) solvation energies (molecules). 3D structures are provided in `.xyz` format in the subfolder `XYZ` (molecules).</p> <div> <div><strong>GB1 dataset</strong></div> <br> <div>The GB1 dataset was <strong>not</strong> generated by us (https://doi.org/10.48550/arXiv.2405.05167). If you use the GB1 dataset, please cite the original paper:</div> <br> <div>Wu, N. C., Dai, L., Olson, C. A., Lloyd-Smith, J. O., & Sun, R. (2016). <em>Adaptation in protein fitness landscapes is facilitated by indirect paths. </em><strong>eLife</strong>, 5:e16965. doi:10.7554/eLife.16965</div> </div> <h3>Results</h3> <p>This folder contains the results (outputs) of the ML models trained using the provided scripts (see github repository). Such results are incuded in the form of `.npy` files. To load the files please include the option `allow_pickle=True`.</p> <p> Each `.npy` file contains the following keys:<br> * `initial_parameters`: script inputs.<br> * `d_encoder`: encoder used (not always included).<br> * `ns_train`: number of training points used for the LCs (rounded, integers).<br> * `ns_train_float`: number of training points used for the LCs (not rounded, float).<br> * `ns_train_norm`: number of training points used for the LCs (normalized, float).<br> * `res`: test MAEs.<br> * `res_tot`: (train,validation,test) MAEs.<br> * `res_tot_mut`: (train,validation,test) MAEs sorted by mutation number.<br> * `l_opt`: optimal kernel length used during the test.<br> * `ls`: kernel lengths used for grid search.<br> * `idx_seeds`: indices used to reshuffle the data. If one want to rebild the initial order use `np.argsort(idx_seeds)`.<br> * `alpha_opt`: optimal regression parameters used to calculate the test error. To sort the data use `alpha_opt[i][ii][np.argsort(idx_seeds[ii,0:arg_train_max].astype(int)[:ns_train[i]]`. Where `i` is the replicate number (0-99) and ii is the idex in the LC.<br> * `valid_errs`: validation error (MAE) calculated for each point in the hyperparameter (kernel scale) optimisation.<br> * `test_errs`: test error (MAE) calculated for each point in the hyperparameter (kernel scale) optimisation.</p>
HySN Data set
<p>HySN, a high resolution HYbrid SeNorge data set of daily near surface humidity and incident shortwave and longwave radiation, and surface pressure created by merging reanalysis data (Era-interim) with a national 1-by-1 km gridded temperature data set, SeNorgev2 (SeNorge v2.1 between 1979 and 2015, and SeNorge v2.0 from 2016). The assumptions and methods used to downscale humidity and longwave radiation are similar to those used in WATCH/PGMFD/NLDAS-1. The software used to compile the data is available at <a href="https://doi.org/10.5281/zenodo.1435555">GitHub</a>. The data is described, and compared to surface observations and other data sets here: <a href="https://doi.org/10.5194/essd-11-797-2019">Erlandsen, H. B., Tallaksen, L. M., and Kristiansen, J.: Merits of novel high-resolution estimates and existing long-term estimates of humidity and incident radiation in a complex domain, Earth Syst. Sci. Data, 11, 797-821, https://doi.org/10.5194/essd-11-797-2019, 2019</a>, and in its supplement: <a href="https://doi.org/10.5194/essd-11-797-2019-supplement">Supplement of Earth Syst. Sci. Data, 11, 797–821, 2019 https://doi.org/10.5194/essd-11-797-2019-supplement</a> .</p> <p>The data has the same projection as SeNorge, UTM33, and is stored for precipitation days, i.e. the datestamp is day-month-year 06 UTC and the variable represents average value of the preceding 24 hours. Monthly means of the data can be found here: <a href="https://doi.org/10.5281/zenodo.1993870">https://doi.org/10.5281/zenodo.1993870</a> .</p> <p>Note that HySN5, a high resolution HYbrid SeNorge data set of daily near surface humidity and surface incident shortwave and longwave radiation, and surface pressure created by merging reanalysis data (Era5) with a national 1-by-1 km gridded temperature data set, SeNorge2018, is available from 1979-2000 (<a href="http://10.5281/zenodo.3351430">10.5281/zenodo.3351430</a>) and 2001-2017 (<a href="http://10.5281/zenodo.3516560">10.5281/zenodo.3516560</a>). </p>
Monthly means of the HySN Data set
<p>Annual monthly means, and monthly means of the full time-series of HySN, a high resolution HYbrid SeNorge data set of daily near surface humidity and incident shortwave and longwave radiation, and surface pressure created by merging reanalysis data (Era-interim) with a national 1-by-1 km gridded temperature data set, SeNorge v2 (SeNorge v2.1 between 1979 and 2015, and SeNorge v2.0 from 2016). The assumptions and methods used to downscale humidity and longwave radiation are similar to those used in WATCH/PGMFD/NLDAS-1. The software used to compile the data is available at <a href="https://doi.org/10.5281/zenodo.1435555">GitHub</a>. The daily data is at <a href="https://doi.org/10.5281/zenodo.1970170">doi.org/10.5281/zenodo.1970170</a>.</p> <p>The data has the same projection as SeNorge, UTM33, and is stored for precipitation days, i.e. the datestamp is day-month-year 06 UTC and the variable represents average value of the preceding 24 hours.</p>
Turbulence in Natural Mangrove Pneumatophore Canopies - Data Set
<p>Raw and processed data files (format: zipped .mat files) for plotting Figures 3 - 15 in our paper entitled "Turbulence in Natural Mangrove Pneumatophore Canopies" published in the Journal of Geophysical Research: Oceans in 2018. </p>
Use of Augmented Reality in the preservation of architectural heritage: case of aqueduct Kuru Kopru _Data set
<p>Architectural preservation embeds all the activities dealing with the physical sustainability of the built heritage<strong>,</strong> its diffusion and comprehension by a wide scope public. Representation and diffusion of heritage take a core place in that process. Nowadays, Augmented Reality (AR) is one of the most used digital tools in the diffusion of architectural heritage. This dataset embeds data used for the modeling of reconstruction models of the Roman aqueduct Kuru Kopru from the Roman-Byzantıne period to the year 2017 and those used for the development of two AR applications for the diffusion of the aforementioned aqueduct.</p>
1-hour meteorological data set
<p>1-hour meteorological data set</p>
Improving soil aquifer treatment efficiency using air injection into the subsurface - Data set
<p>Data set for the research article "Improving soil aquifer treatment efficiency using air injection into the subsurface"</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.