Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
3,481
datasets available to search
ShareScore release 0.9.0
Dataset results
3,481 results for “data set”
Aminergic G protein-coupled receptor (GPCR) mutation data set
<p>Curated set of 6553 quantitative mutation data points covering 34 aminergic G protein-coupled receptors, annotated according to GPCRdb standards (http://gpcrdb.org/) building from the data set published by Kooistra et al. in BJP 2013, 170, 101.</p>
Chemokine G protein-coupled receptor (GPCR) mutation data set
<p>Curated set of 2004 quantitative mutation data points covering 10 chemokine G protein-coupled receptors, annotated according to GPCRdb standards (http://gpcrdb.org/) building from the data set published by Scholten et al. in BJP 2012, 165, 1617.</p>
Panasonic NCR18650B formation, RPT and half cell data set
<p>This dataset is a complement to a publication in "Batteries":</p> <p>Durability and Reliability of EV Batteries Under Electric Utility Grid Operations.</p> <p>Part1: Cell-to-cell variations and preliminary testing </p> <p>The data consists on matlab files for each tested cell.</p>
R data set: The Cancer Genome Atlas Gene Expression data
<p>This compound data set comprises the following information from the The Cancer Genome Atlas:</p> <ul> <li>RNA-Seq counts for 60483 genes across 11093 samples</li> <li>HuEx 1.0 ST gene expression data for 18632 genes across 1211 samples</li> <li>clinical indicators for 11160 patients</li> </ul> <p>All gene expression data is annotated across ENSEMBL, ENTREZ and symbols. Samples are annotated by TCGA barcodes.</p> <p>To read the data set into R (requires 6 GB of RAM) use:</p> <p>tcga <- readRDS("tcga.rds")</p>
IMP test data set
<p>This file contains the test data set used within the article:</p> <p><strong>IMP: a reproducible pipeline for reference-independent integrated metagenomic and metatranscriptomic analyses</strong></p> <p>Shaman Narayanasamy<sup>†</sup>, Yohan Jarosz<sup>†</sup>, Emilie E.L. Muller, Cédric C. Laczny, Malte Herold, Anne Kaysen, Anna Heintz-Buschart, Nicolás Pinel, Patrick May, and Paul Wilmes<sup>*</sup></p> <p>Preprint: http://biorxiv.org/content/early/2016/02/10/039263</p> <p>This test data set was used for benchmarking the run times of IMP. They are derived by selecting the first 5% of reads from a wastewater sludge microbial community dataset (see manuscript). Also included are the respective preprocessed FASTQ files such that IMP can be tested without running the preprocessing step. A README file inside the folder briefly describes the different FASTQ files contained in the folder.</p>
IMP simulated mock community data set
<p>This file contains the simulated mock (SM) metagenomic and metatranscriptomic dataset along with the original genomes used for simulation used within the article:</p> <p><strong>IMP: a reproducible pipeline for reference-independent integrated metagenomic and metatranscriptomic analyses</strong></p> <p>Shaman Narayanasamy<sup>†</sup>, Yohan Jarosz<sup>†</sup>, Emilie E.L. Muller, Cédric C. Laczny, Malte Herold, Anne Kaysen, Anna Heintz-Buschart, Nicolás Pinel, Patrick May, and Paul Wilmes<sup>*</sup></p> <p>Preprint: http://biorxiv.org/content/early/2016/02/10/039263</p> <p>The folder contains two subfolders (MG and MT), each containing the simulated metagenomic (MG) data and simulated metatranscriptomic (MT) data respectively. The files within these folders are in FASTQ format. The methods for generating these simulated data sets are described in the article.<br> </p> <p>The genomes and the resulting simulated metatranscriptomic data was generated and analysed within the article:</p> <p><strong>Comparison of assembly algorithms for improving rate of metatranscriptomic functional annotation</strong></p> <p>Albi Celaj, Janet Markle, Jayne Danska and John Parkinson; 2014; doi:10.1186/2049-2618-2-39</p> <p> </p> <p>It was provided upon request by the first author Albi Celaj, with permission to share the data. Please cite the aforementioned publication if this simulated metatranscriptomic data is used data is used.</p>
Example pRF data set for SamSrf
<p>This archive contains an example data set for a pRF mapping experiment. The functional data have already been motion corrected and coregistered to the structural scan. The structural has also been reconstructed in FreeSurfer. The archive contains the minimal files necessary to reproduce a pRF analysis from this data, including the cortical projection steps and the actual analysis.</p> <p>This assumes you will use the SamSrf toolbox for pRF mapping (although you could also analyze it with other packages):</p> <p>https://dx.doi.org/10.6084/m9.figshare.1344765</p>
Indexed Data Set From Molisan Regional Seismic Network Events
<p>Abstract:</p> <p><em>After the earthquake occurred in Molise (Central Italy) on 31st October 2002 (Ml 5.4, 29 people dead), the local Servizio Regionale per la Protezione Civile to ensure a better analysis of local seismic data, through a convention with the Istituto Nazionale di Geofisica e Vulcanologia (INGV), promoted the design of the Regional Seismic Network (RMSM) and funded its implementation. The 5 stations of RMSM worked since 2007 to 2013 collecting a large amount of seismic data and giving an important contribution to the study of seismic sources present in the region and the surrounding territory. This work reports about the dataset containing all triggers collected by RMSM since July 2007 to March 2009, including actual seismic events; among them, all earthquakes events recorded in coincidence to Rete Sismica Nazionale Centralizzata (RSNC) of INGV have been marked with S and P arrival timestamps. Every trigger has been associated to a spectrogram defined into a recorded time vs. frequency domain.<br> The dataset has been fully indexed in respect of the recorded spectra: list of all records, list of earthquakes, list of multiple earthquakes records.<br> The main aim of this structured dataset is to be used for further analysis with data mining and machine learning techniques on image patterns associated to the waveforms.</em></p>
Raw data sets for: A simple calculation algorithm to seperate high-resolution CH4 flux measurements into ebullition- and diffusion derived components (AMT)
<p>Raw data sets for the research article "A simple calculation algorithm to seperate high-resolution CH4 flux measurements into ebullition- and diffusion derived components", published in "Atmospheric Measurment Techniques" (AMT). Data sets include raw data sets for the field and laboratory study, as well as calculated CH4 fluxes (field).</p>
Chemokine G protein-coupled receptor (GPCR) mutation data set - revision
<p>Curated set of 2004 quantitative mutation data points covering 10 chemokine G protein-coupled receptors, annotated according to GPCRdb standards (http://gpcrdb.org/) building from the data set published by Scholten et al. in BJP 2012, 165, 1617.</p> <p>Errors in the previous version were corrected and new data points for CCR2 and CCR9 were added.</p>
Data set for ``Why is Differential Evolution Better than Grid Search for Tuning Defect Predictors?''
<p>One of the black arts of data mining is learning the magic parameters that control the learners. In software analytics, at least for defect prediction, several methods, like grid search and differential evolution(DE), have been proposed to learn those parameters. They’ve been proved to be able to improve learner performance.</p> <p>We want to evaluate which method can find better parameters in terms of performance score and runtime. This paper compares grid search to differential evolution, which is an evolutionary algorithm that makes extensive use of stochastic jumps around the search space. We find that the seemingly complete approach of grid search does no better, and sometimes worse, than the stochastic search. Yet, when repeated 20 times to check for conclusion validity, DE was over 210 times faster (6.2 hours for DE vs 54 days for grid search when both tuning Random Forest over 17 test data sets with F-measure as optimization objective).</p> <p>These results are puzzling: why does a quick partial search be just as effective as a much slower, and much more, extensive search? To answer that question, we turned to the theoretical optimization literature. Bergstra and Bengio conjecture that grid search is not more effective than more randomized searchers if the underlying search space is inherently low dimensional. This is significant since recent results show that defect prediction exhibits very low intrinsic dimensionality– an observation that explains why a fast method like DE may work as well as a seemingly more thorough grid search. This suggests, as a future research direction, that it might be possible to peek at data sets before doing any optimization in order to match the optimization algorithm to the problem at hand.</p>
Data set to article "Synthetic inversions for density using seismic and gravity data" by Blom, Boehm and Fichtner
<p><strong>Data set to “Synthetic inversions for density using seismic and gravity data” by Nienke Blom, Christian Boehm and Andreas Fichtner</strong></p> <p>This data set relates to our paper <em>“Synthetic inversions for density using seismic and gravity data”</em><em>, </em><em>in which we discuss the imaging of density variations inside the Earth as a separate, independent parameter using seismic waveform tomography and gravity measurements</em>. The research consists of synthetic experiments conducted using a home-written MATLAB wave propagation code. The data set contains the code itself, the input files and output files for each of the experiments described in the manuscript and its supplementary material, all the figures, some extra material (such as a video of Figure 1 in the manuscript) and some scripts.</p> <p>Below I’ll give a description of the contents of this data set and how they are structured, followed by an overview of the experiments conducted for the paper.</p> <p>In this data set, the following things can be found:</p> <ul> <li> <p>There is a directory with all the figures: FIGURES. This contains the figures in *.pdf, *.eps and *.png formats.</p> </li> <li> <p>There is a directory FD2D_ADJOINT_CODE with in it the MATLAB code fd2d-adjoint. If you plan on using our code, it would be awfully kind if you'd make a reference both to the code and to this paper. It was a lot of work to develop the code and the experiments. NOTE: the code supplied here is a snapshot of the code taken in February 2017. A more up-to-date version might be found on github (www.github.com/Phlos/fd2d-adjoint)</p> </li> <li> <p>For each (series of) experiment(s) described in the paper, there is a directory T1, T2, …, Tn. This also holds for the supplementary tests, the folders for which are designated with the suffix .SUPPLEMENTARY.</p> </li> <li> <p>For Figure 1 in the manuscript, there is a directory Fig1.snapshots. In this directory, everything pertaining to the snapshots figure and its corresponding video can be found.</p> </li> <li> <p>There is a separate directory SCRIPTS with a couple of useful scripts that might be used in addition to the ones in the fd2d-adjoint code.</p> </li> </ul> <p><br> In each of the test directories T1...Tn, there are subdirectories for each experiment conducted within that test framework. Each of the subdirectories has a name Systematic.test-[xxx]. Within those Systematic.. directories, the following can be found:</p> <ul> <li> <p>an input file Systematic….input_parameters.m that can be copied to [fd2d-adjoint]/input/input_parameters.m in order to re-run the experiment. As the code has been under development while the tests were run, it may be that some input parameters are missing from the earlier experiments.</p> </li> <li> <p>A mat-file obs.all-vars.mat. If this file is copied to [fd2d-adjoint]/output/Systematic.test… , this saves the recalculation of the ‘obs’ data when the code is run.</p> </li> <li> <p>A mat-file initial_misfits.mat. If this file is copied to [fd2d-adjoint]/output/Systematic.test… , this saves the recomputation of the initial misfits with respect to the obs data when the code is run.</p> </li> <li> <p>A file lbfgs_output_log.txt which monitors the misfit and gradient development across the iterations. If the inversion was restarted a couple of times, all of this remains in the logfile.</p> </li> <li> <p>For each iteration of the inversion iter[xxx], an iter[xxx].all-vars.mat file, which contains most of the matlab output files for this iteration.</p> </li> <li> <p>For each iteration of the inversion iter[xxx], some figures:</p> <ul> <li> <p>a model plot of the current model anomalies with respect to the background model iter[xxx].model-diff.rhovsvp.png.</p> </li> <li> <p>a gravity plot of the gravity vector difference between the current model and the background model iter[xxx].gravity_difference.png.</p> </li> <li> <p>a kernel plot of the total relative kernels (whether seis only or seis+grav) of the current model in rho-mu-lambda parametrisation: iter[xxx].rho-mu-lambda.png.</p> </li> </ul> </li> </ul> <p><br> </p> <p>Now follows a brief description of each of the (series of) tests conducted for the paper. The test numbers are mostly chronological, and so are the Systematic.test… subdirectories.</p> <ul> <li> <p><strong>Figure 1</strong>: shows snapshots of wave propagation past a density anomaly. The full data for this and the full video are given in the Fig1.snapshots. <em>Discussed in: Figure </em><em>1 of the manuscript.</em></p> </li> <li> <p><strong>T1: </strong><strong>reference.</strong> A reference test in which we assess to which density can be recovered as an independent parameter. <em>Discussed in: Figure </em><em>4</em></p> <ul> <li> <p>Reference experiment: Systematic.test-033</p> </li> </ul> </li> <li> <p><strong>T2: </strong><strong>ignored density.</strong> A test in which the effect is explored if density is ignored, i.e. if it is kept fixed to the starting model. <em>Discussed in: Figure </em><em>4</em></p> <ul> <li> <p>Fixing density: Systematic.test-040</p> </li> </ul> </li> <li> <p><strong>T3: </strong><strong>starting model</strong>. A series of test in which is explored to what extent the starting models of P and S seismic velocity influence the recovery of density. In the different sub-tests, different levels of information on P and S velocity are already present. <em>Discussed in: Figure </em><em>6</em></p> <ul> <li> <p>vs, vp 100% correct: Systematic.test-029</p> </li> <li> <p>vs, vp 75% correct: Systematic.test-037</p> </li> <li> <p>vs,vp 50% correct: Systematic.test-036</p> </li> </ul> </li> <li> <p><strong>T4: </strong><strong>fixed velocities</strong>. A series of tests in which is explored to what extent one can “get away with” only updating density, assuming that the models for P and S velocity are already sufficiently accurate. <em>Discussed in: Figure </em><em>7</em></p> <ul> <li> <p>vs,vp fixed at 50% correct: Systematic.test-038</p> </li> <li> <p>vs, vp fixed at 75% correct: Systematic.test-041</p> </li> <li> <p>vs, vp fixed at 100% correct: Systematic.test-039</p> </li> </ul> </li> <li> <p><strong>T5: </strong><strong>gravity</strong>. A set of tests in which the addition of gravity data to the (up until here purely) seismic inversion. Both the full gravity vector and its potential are used as gravity data. <em>Discussed in: Figure </em><em>8</em></p> <ul> <li> <p>seismic + full gravity vector (x,z) data: Systematic.test-045</p> </li> <li> <p>seismic + gravity potential data (‘geoid’): Systematic.test-046</p> </li> </ul> </li> <li> <p><strong>T6: noise</strong>. A series of tests in which the addition of noise to the seismic data is explored. Both correlated and uncorrelated noise are explored. Noise levels vary across frequencies. <em>Discussed in: Figure </em><em>9</em></p> <ul> <li> <p>correlated noise: Systematic.test-050</p> </li> <li> <p>uncorrelated noise: Systematic.test-052</p> </li> </ul> </li> <li> <p><strong>T7: impedance</strong>. A test in which the impedance contrast across anomaly boundaries are set to zero. It is explored to what extent the recovery of density relies on the presence of an impedance contrast. <em>Discussed in: Figure </em><em>10</em></p> <ul> <li> <p>no impedance contrast: Systematic.test-055</p> </li> </ul> </li> <li> <p><strong>T8: parametrisation (</strong><em><strong>supplementary</strong></em><strong>)</strong>. A test in which it is explored to what extent the inversion is affected if an inversion parametrisation using density and the elastic parameters mu and lambda is used, instead of the otherwise used parametrisation density-S velocity-P velocity. <em>Discussed in: </em><em>Supplementary </em><em>Figure </em><em>1,2 @ </em><em>Supplementary_material.pdf</em></p> <ul> <li> <p>inversion parametrisation rho-mu-lambda (reference target model): Systematic.test-032</p> </li> <li> <p>inversion parametrisation rho-mu-lambda with ‘scaling’ target model: Systematic.test-062a</p> </li> </ul> </li> <li> <p><strong>T9: scaling relations</strong>. A set of tests in which it is explored to what extent the recovery of density and seismic velocities is influenced if density is scaled to S velocity using a fixed scaling. <em>Discussed in: Figure </em><em>5</em></p> <ul> <li> <p>target model with density scaled to S velocity in different ways; all parameters free: Systematic.test-063</p> </li> <li> <p>same target model, but now density is scaled to S velocity with a fixed relationship: Systematic.test-067</p> </li> </ul> </li> <li> <p><strong>T10: anomaly strength (</strong><em><strong>supplementary</strong></em><strong>)</strong>. A set of tests in which the effect of the strength of the anomalies on the recovery of density and the other parameters is investigated. <em>Discussed in: </em><em>Supplementary </em><em>Figure </em><em>3-5 @ </em><em>Supplementary_material.pdf</em><em> </em></p> <ul> <li> <p>target model like reference case, but the anomalies 10% of PREM instead of 1%: Systematic.test-065</p> </li> <li> <p>target model like reference case, but the anomalies <em>in the upper mantle only</em> 10% of PREM instead of 1%: Systematic.test-064</p> </li> </ul> </li> </ul> <p><br> </p> <p>If you have any further questions, feel free to contact me.</p> <p>All the best,</p> <p>Nienke Blom, Utrecht University<br> n.a.blom@uu.nl<br> nienke.blom@posteo.net</p> <p> </p>
Data for the article "Performance of SCAN density functional method for a set of ionic liquids"
<p>The repository (https://github.com/vilab-tartu/SCAN) contains the database, geometries and an illustrative ipython notebook supporting the article "Performance of SCAN density functional method for a set of ionic liquids". </p>
The Kühtai data set: 25 years of lysimetric, and snow pillow and meteorological measurements
This dataset presents long-term observations from an experimental snow lysimeter plot in Kühtai (Austrian Alps). The data set includes 15 minutes data of snow water equivalent from a 10 m² snow pillow, snow melt outflow from a 10 m² snow lysimeter placed at the same location as the pillow, meteorological data (precipitation, incoming global radiation, reflected short wave radiation, air temperature, relative air humidity and wind speed), and other data (snow depths, snow temperatures at seven heights) from the period October, 1990 – May, 2015. All data have been quality checked, and gaps in the meteorological data have been filled in.
Data Set of the Thesis "Evaluation of Interaction Concepts in Virtual Reality Applications"
<p>This archive consists of videos files, log files and R tables on which the bachelor thesis 'Evaluation of Interaction Concepts in Virtual Reality Applications' by Hanna Holderied and subsequent publications are based. In addition to recorded data, it also contains the Unity VR application with which the case study of the thesis was performed. For details, please refer to the bachelor thesis published by the University of Goettingen, Institute of Computer Science.</p>
Word Embedding Data Sets Learned from Tweets and General Data
<p>This includes 10 word embedding data sets learned from about 400 million tweets and 7 billion words from general data. They can be used in tasks involving social media data, especially tweets, and other types of textual data. Users can choose different embedding sets based on their use cases; they can also easily try all of them to see which one provides the best performance for their application.</p> <p>More details about the training data collection, word embedding generation, preprocessing steps, and how to use them can be found from the following paper:</p> <p>Quanzhi Li, Sameena Shah, Xiaomo Liu, Armineh Nourbakhsh, Data Set: Word Embeddings Learned from Tweets and General Data, The 11th International AAAI Conference on Web and Social Media (ICWSM-17). Montreal, Canada. May 16-18, 2017</p>
data set of microscale strain for creep of Carrara marble
<p>Data set used for the construction of all microscale strain maps of creep of Carrara marble. The data is stored in Matlab structures (.mat files) and contains all markers coordinates as well as deformation tensors for a 9n point average (for more details see: Quintanilla-Terminel, A., and B. Evans (2016), Heterogeneity of inelastic strain during creep of Carrara marble: Microscale strain measurement technique, J. Geophys. Res. Solid Earth, 121, 5736–5760).</p>
Data-set of studies on mortality, sub-lethal and reproductive effect on amphibians and reptiles
<p>Amphibians and reptiles have not been considered in environmental risk assessments of chemicals, which has generated some debate about whether risk posed by some pollutants like pesticides on these animals are covered by surrogates in the groups of fish, mammals and birds. In order to develop a scientifically sound and robust risk assessment scheme it is necessary to have enough information available on the biological relevance of effects observed in laboratory studies in view of population level effects, to identify sensitive life stages and to compare sensitivity of our target study groups with that of their surrogates.</p> <p>With these objectives, a systematic review of toxicological literature on amphibians and reptiles a comprehensive search of relevant literature on toxicity data related to amphibians and reptiles was conducted, using the appropriate search strings and combinations, in six different source types: multidisciplinary databases of scientific literature (i.e. Web of Science and Scopus), literature included in general amphibian and reptile ecotoxicology compilations, literature compiled in technical reports previously prepared for EFSA, literature sources used for creating amphibian or reptile records in the ecotoxicological database created by De Zwart (see references for details), literature sources used for creating amphibian or reptile records in toxicological online databases (United States Environmental Protection Agency’s Ecotoxicology Knowledgebase–ECOTOX, and National Library of Medicine’s Hazardous Substances Data Bank–HSDB), and indexes of herpetological scientific journals of local scope not included in Web of Science or Scopus.</p> <p>Data extraction consisted of the retrieval of relevant information, including among other fields: species, age, sex, chemical substance, exposure route and duration, type of recorded endpoint, type of response, exposure concentration, mean effect value of the control and exposed groups and the reported variability measures of these mean effects, and statistical significance of the comparison.</p> <ul> <li>Data category 1: Reporting endpoints. Endpoints are defined here as benchmark values obtained from the integration of responses measured at different concentrations (e.g. LC<sub>50</sub>, EC<sub>50</sub>, NOEC, for which calculation it is necessary to make a regression with the percentage of effect at different exposure concentrations).</li> <li>Data category 2: Reporting responses for each tested level. These are studies in which different replicates of experimental units are exposed to different levels (doses, concentrations), including a control treatment, and a magnitude of effect is recorded at each level. For some of these studies it was possible to calculate an endpoint (e.g. LC<sub>50</sub>, EC<sub>50</sub>, NOEC) and for others it was not (e.g. if only one concentration was tested or if a statistically significant adjustment between exposure level and effect cannot be achieved).</li> </ul> <p> </p> <p>Three objectives corresponding to the review of the effects of chemicals on amphibians and reptiles, the main results of the study were:</p> <ol> <li>Identification of the most sensitive life stage.</li> <li>Extrapolation from laboratory data to mesocosm and field situations</li> <li>Comparison with surrogate taxa</li> </ol> <p>The review was conducted on 3642 full-text records for chemical exposure effects and 556 for life history traits, out of which 1332 and 204, respectively, were finally used for data extraction. The datasets comprised 23152 values corresponding to effects of chemicals and 1854 values corresponding to life history traits.</p>
Data set for manuscript 'A universal poroelastic mechanism for hydraulic signals in biomimetic and natural branches'
<p>The zip file contains all raw data associated with the manuscript 'A universal poroelastic mechanism for hydraulic signals in biomimetic and natural branches' with Python files used to process the data and plot the figures of the manuscript.</p>
Spectroscopic data set for an unusual white dwarf
<p>The tar file includes fits-formatted count (d*) and flux (ca*) spectra of an unusual white dwarf obtained at the MDM observatory. The spectra are wavelength calibrated. Users should exercise caution when quoting absolute flux. Heliocentric velocity corrections are not applied. The count spectra list the wavelength (angstrom) and the total count, and the flux spectra list the wavelength (angstrom) and the flux in units of erg/cm^2/s/angstrom.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.