Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
3,481
datasets available to search
ShareScore release 0.9.0
Dataset results
3,481 results for “data set”
Combining molecular data sets with strongly heterogeneous taxon coverage enlightens the peculiar biogeographic history of stoneflies (Insecta: Plecoptera)
<p class="Standard1">Extant members of the ancient insect order of stoneflies exhibit a disjunct, antitropical distribution, with one major lineage exclusively occurring in the Southern Hemisphere and the other, with few exceptions, on the Northern continents. Here, we address the biogeographic distribution and phylogenetic relationships of stoneflies using a phylogenetic workflow that combines both transcriptomic and Sanger sequence datasets with heterogeneous taxon coverage. We used a dataset comprising 2997 genes derived from the transcriptomes of 30 species and Sanger sequences of seven genes for 498 species. The backbone phylogeny was mainly inferred from the transcriptomic data, whereas the Sanger nucleotide sequence data provided high species density for divergence time estimation and diversification analyses. Our results show that the biogeographic pattern we observe today is primarily more likely shaped by long-distance over-land dispersal than by vicariance. We inferred that the ancestors of extant stoneflies originated in the Northern Hemisphere approximately 265 Ma and were presumably restricted to this area due to climatic and geographic boundaries. Our analyses suggest that with the break-up of Pangaea around 200 Ma and the associated climatic and geographical changes, two groups of stoneflies, the Anarctoperlaria and the Notonemouridae, dispersed to Gondwana and subsequently went extinct on the northern continents. Both groups likely dispersed across Gondwana before its break-up into the modern continents. At least one member of another group of 'northern' stoneflies, the Acroneuriinae, seems to have migrated from North America to South America around 67 Ma. We found four major net diversification rate shifts, indicating rapid radiation patterns that hampered a robust phylogenetic placement of these stonefly groups. Our study provides the first conclusive evolutionary explanation for the unique distribution pattern of stoneflies.</p>
Medical Equipment Image Data set
<p>This data set aims to realise the medical equipment recognition to aid visual search through three deep learning models. The data set contains ten medical equipment classes: commodes, wheelchairs, walking frames, blood pressure monitors, breast pumps, thermometers, rippled mattresses, oximeters, crutches, and therapeutic ultrasound machines. We collected from online resources around 220 images for each medical equipment class. Each image class in the test set has around 40 images.</p> <p>Refer to the paper <a href="http://doi.org/10.3233/JIFS-212786">here</a> or in research gate (preprint).</p>
Data sets used in "Neural network emulation of the formation of organic aerosols based on the explicit GECKO-A chemistry model"
<p>The training, validation, and testing data sets for toluene, dodecane, and alpha-pinene models described in the manuscript. A link to the manuscript will be added here when it becomes available. All trajectories in the data sets were generated using GECKO-A. The source code for using the data sets can be found at https://github.com/NCAR/gecko-ml </p>
Data set for the paper "Temperatures and cooling rates recorded by the New Caledonia ophiolite: implications for cooling mechanisms in young forearc sequences"
<p>Data set for the paper "Temperatures and cooling rates recorded by the New Caledonia ophiolite: implications for cooling mechanisms in young forearc sequences" by Secchiari et al.</p>
A method for identifying environmental stimuli and genes responsible for genotype-by-environment interactions from a large-scale multi-environment data set
<p>It has not been fully understood in real fields what environment stimuli cause the genotype-by-environment (G × E) interactions, when they occur, and what genes react to them. Large-scale multi-environment data sets are attractive data sources for these purposes because they potentially experienced various environmental conditions. In this study, we developed a data-driven approach termed <u>E</u>nvironmental <u>C</u>ovariate Search Affecting <u>G</u>enetic <u>C</u>orrelations (ECGC) to identify environmental stimuli and genes responsible for the G × E interactions from large-scale multi-environment data sets. ECGC was applied to a soybean (<i>Glycine max</i>) data set that consisted of 25,158 records collected at 52 environments. ECGC illustrated what meteorological factors shaped the G × E interactions in six traits including yield, flowering time, and protein content and when they were involved. For example, it illustrated the relevance of precipitation around sowing dates and hours of sunshine just before maturity to the interactions observed for yield. Moreover, genome-wide association mapping on the sensitivities to the identified stimuli discovered candidate and known genes responsible for the G × E interactions. Our results demonstrate the capability of data-driven approaches to bring novel insights on the G × E interactions observed in fields. This dataset provides the data used in this study and supplementary tables cited in the manuscript.</p>
VCF for neutral data set in Harpagifer bispinis along the Magellan Province
<p>Quaternary glacial cycles shaped the current distribution of polar and cold-temperate biotas. In the Magellan province of South America, ice covering during the last glacial maximum radically altered the landscape/seascape, speciation rate, and the distribution of species. Here we studied nototheniid fishes <i>Harpagifer </i>spp. along the Magellan province reported as two nominal species: <i>H. bispinis</i> in Patagonia and <i>H. palliolatus</i>, endemic to the Falkland/Malvinas Islands. Previous molecular analyses in <i>Harpagifer</i> showed that the genus may have recently colonized southern South America ~ 1 million years ago. The extensive use in systematics of molecular markers to determine evolutionary units has been improved due to the advances in NGS. Combining traditional DNA sequences and non-targeted GBS-SNPs we evaluated both, the presence of effective evolutionary units and contemporary patterns of genetic structure across the Magellan province. DNA sequences consistently showed an absence of phylogeographic structure, with shared dominant haplotypes between nominal species, pointing towards the presence of a single evolutionary unit. In contrast, SNPs identified three groups in Patagonia, two located north and south of the Strait of Magellan, and a third well-differentiated one in the Falkland/Malvinas Islands. Connectivity analyses using SNPs suggest limited and asymmetric gene flow from Patagonia to the Falkland/Malvinas. Contrasting rough- and fine-scale genetic evolutionary patterns recorded in <i>Harpagifer</i> enhance the relevance in the use of combined methodologies for species delimitation analyses. Depending on the question to be addressed, we could discriminate among phylogeographic structure discarding incipient speciation, and contemporary spatial differentiation processes linked to drift-migration equilibrium models.</p>
A Data Set for State and Parameter Estimation in Power Systems
<p>This data set consists of data from three power system models of different scales (IEEE 14, IEEE 118 and <a href="https://doi.org/10.5281/zenodo.2642175">PanTaGruEl</a>). For each of these systems, 5 different cases are provided, they are sorted from the least to the most "advanced" system operations.</p> <p>Data are stored in <a href="https://en.wikipedia.org/wiki/Hierarchical_Data_Format">HDF5</a> format (as H5T_NATIVE_FLOAT) which can be read by (mostly) any language (e.g. Python, Matlab or Julia).</p> <p><strong>Description of the different cases:</strong></p> <ul> <li><em>Case 1#</em> consists of 2000 samples. Each sample is obtained by: firstly, defining total active and reactive loads in the system which are then distributing to the buses and, secondly, dispatching generation (this is performed by running an OPF (Optimal Power Flow) with <a href="https://matpower.org/">Matpower</a>). The same distribution factors were used for every samples.</li> <li><em>Case 2#</em> is similar to <em>case 1#</em> with the addition of independent white noises to each bus load.</li> <li><em>Case 3#</em> differs from <em>case 1#</em> in that independent active and reactive bus loads are randomly drawn.</li> <li><em>Case 4# </em>is similar to <em>case 3#</em>, but some generators are randomly drawn to be in maintenance. This set of generators is independently generated for each sample.</li> <li><em>Case 5#</em> is similar to <em>case 4#</em>, plus the generation cost of each generator is randomly drawn from a predefined range. Costs are independently generated for each sample.</li> </ul> <p><strong>General Description:</strong></p> <p>Each data set case file contains the following elements:</p> <ul> <li>V (<span class="math-tex">\(N_{\rm bus} \times N_{\rm sample}\)</span> matrix): Voltage magnitudes,</li> <li>theta (<span class="math-tex">\(N_{\rm bus} \times N_{\rm sample}\)</span> matrix): Voltage phases,</li> <li>P (<span class="math-tex">\(N_{\rm bus} \times N_{\rm sample}\)</span> matrix): <a href="https://en.wikipedia.org/wiki/AC_power">Active</a> power injections (i.e. = generation - load),</li> <li>Q (<span class="math-tex">\(N_{\rm bus} \times N_{\rm sample}\)</span> matrix): <a href="https://en.wikipedia.org/wiki/AC_power">Reactive</a> power injections,</li> <li>idgen (<span class="math-tex">\(N_{\rm gen}\)</span> vector): index of generator buses,</li> <li>id_slack: index of the bus used as <a href="https://en.wikipedia.org/wiki/Slack_bus">slack bus</a>,</li> <li>epsilon (<span class="math-tex">\(N_{\rm line} \times 2\)</span> matrix): list of the lines in the system (Each row corresponds to a line. Entries are buses’ indices.),</li> <li>b (<span class="math-tex">\(N_{\rm line}\)</span> vector): line <a href="https://en.wikipedia.org/wiki/Admittance">susceptances</a>,</li> <li>g (<span class="math-tex">\(N_{\rm line}\)</span> vector): line <a href="https://en.wikipedia.org/wiki/Admittance">conductances</a>,</li> <li>bsh (<span class="math-tex">\(N_{\rm bus}\)</span> vector): shunt susceptances,</li> <li>gsh (<span class="math-tex">\(N_{\rm bus}\)</span> vector): shunt conductances.</li> </ul> <p><strong>Visualization:</strong></p> <p>The data set also includes bus coordinates.</p> <p><strong>Some theory:</strong></p> <p>The <a href="https://en.wikipedia.org/wiki/Incidence_matrix">incidence matrix</a> B is defined as</p> <p><span class="math-tex">\(B_{ij} = \left\{\begin{array}{l}-1,\; \text{if line $j$ starts at bus $i$,}\\1,\; \text{if line $j$ ends at bus $i$,}\\ 0,\; \text{otherwise.} \end{array}\right.\)</span></p> <p>(“Ends” and “starts” are purely conventional, but they have to be assigned to account for the direction power flows in the system. We use the first column of epsilon as "starts" and the second one as "ends".)</p> <p>The <a href="https://en.wikipedia.org/wiki/Nodal_admittance_matrix">admittance matrix</a> Y is obtained by</p> <p><span class="math-tex">\(y = g + ib,\\ y_{\rm sh} = g_{\rm sh} + ib_{\rm sh},\\ Y = B\,{\rm diag}(y)\,B^\top + {\rm diag}(y_{\rm sh}) .\)</span></p> <p>Defining the <a href="https://en.wikipedia.org/wiki/AC_power">complex</a> power injections and voltages, respectively, as</p> <p><span class="math-tex">\(S = P + iQ,\\ \underline{V} = V \cdot e^{i \theta}, \)</span></p> <p>where <span class="math-tex">\(\cdot\)</span> denotes the element-wise product. One has the following relation</p> <p><span class="math-tex">\(S = \underline{V} \cdot {\rm conj}(Y\, \underline{V}).\)</span></p> <p>This relation is equivalent to the <a href="https://en.wikipedia.org/wiki/Power-flow_study">power flow equations</a>.</p> <ul> </ul> <p> </p>
iMAC data set
<p>================================================================================<br> Title: Comparison between core-collapse supernova nucleosynthesis and meteoric <br> stardust grains: investigating magnesium, aluminium, and chromium <br> Authors: den Hartogh J., Peto M.K., Lawson T., Sieverding A., Brinkman H., <br> Pignatari M., Lugaro M. <br> ================================================================================<br> Description of contents: the data files in this data deposit are either stellar profiles or stellar yields as used in the paper to create figures. All files are machine readable .txt files. Any text editor is able to open these files. </p> <p> A complete list of the files in this repository are provided here:</p> <p> iMAC_profiles/m15exp_LAW_D_subm.txt<br> iMAC_profiles/m15exp_LAW_R_subm.txt<br> iMAC_profiles/m15exp_RIT_D_subm.txt<br> iMAC_profiles/m15exp_RIT_R_subm.txt<br> iMAC_profiles/m15exp_SIE_D_subm.txt<br> iMAC_profiles/m15exp_SIE_R_subm.txt<br> iMAC_profiles/m20exp_LAW_D_subm.txt<br> iMAC_profiles/m20exp_LAW_R_subm.txt<br> iMAC_profiles/m20exp_RIT_D_subm.txt<br> iMAC_profiles/m20exp_RIT_R_subm.txt<br> iMAC_profiles/m20exp_SIE_D_subm.txt<br> iMAC_profiles/m20exp_SIE_R_subm.txt<br> iMAC_profiles/m25exp_LAW_D_subm.txt<br> iMAC_profiles/m25exp_LAW_R_subm.txt<br> iMAC_profiles/m25exp_RIT_D_subm.txt<br> iMAC_profiles/m25exp_RIT_R_subm.txt<br> iMAC_profiles/m25exp_SIE_D_subm.txt<br> iMAC_profiles/m25exp_SIE_R_subm.txt</p> <p> iMAC_yields/m15_yields_D.txt<br> iMAC_yields/m20_yields_D.txt<br> iMAC_yields/m25_yields_D.txt</p> <p> For the profiles the *R_subm.txt files have the following columns:<br> mass,he4,c12,o16,ne20,si28,ni56,fe56,al26,al27,mg24,mg25,mg26,cr50,cr52,cr53,cr54<br> All isotopes are radioactive or non-decayed. </p> <p> For the profiles the *D_subm.txt files have the following columns:<br> mass,al27,mg24,mg25,mg26,cr50,cr52,cr53,cr54<br> All isotopes include all radioactive contributions and are thus fully decayed. </p> <p> The yield files have the following columns:<br> ref Mg24, Mg25, Mg26, Al26, Al27, Cr50, Cr52, Cr53, Cr54<br> All yields are fully decayed. The refs are equal to the labels in Figure 2. <br> <br> <br> <br> System requirements: -</p> <p>Additional comments: the Cr53 data for the middle row of Figure 3 is not included in these files.</p> <p>================================================================================</p>
Data set for this article tiled " the relationship between product innovation and customer satisfaction
<p>my data set for this article tiled " the relationship between product innovation and customer satisfaction</p>
Linkage of hospital records and death certificates by a search engine and machine learning: training and test set data
<p>INTRODUCTION: Vital status is of central importance to hospital clinical research. However, hospital information systems record only in-hospital death information. Recently, the French government released a publicly available dataset containing death-certificate data for over 25 million individuals. The objective of this study was to link French death certificates to the Bordeaux University Hospital records to complete the vital status information.</p> <p>MATERIALS AND METHODS: Our linkage strategy was composed of a search engine to reduce the number of comparisons and machine-learning algorithms. The overall pipeline was evaluated by assembling a file containing 3,565 in-hospital deaths and 15,000 alive persons.</p> <p>RESULTS: The recall and precision of our linkage strategy were 97.5% and 99.97% for the upper threshold and 99.4% and 98.9% for the lower threshold, respectively.</p> <p>CONCLUSION: In this article, we demonstrated the feasibility of accurately linking hospital records with death certificates using a search engine and machine learning.</p>
Data set for "Intertwined spin, charge, and pair correlations in the two-dimensional Hubbard model in the thermodynamic limit"
<p>This data set is for the paper "Intertwined spin, charge, and pair correlations in the two-dimensional Hubbard model in the thermodynamic limit". It contains the raw DCA HD5 and DQMC plain text output files, as well as the scripts and final processed data used to generate figures 1-5 of the main text and supplementary figures 1-12. Copies of the figures and latex files are also included for completeness. </p>
◂Fig. 2 Maximum likelihood trees. A Tree obtained when analysing 18S data set. B Tree obtained when analysing 28S data set. C Tree obtained when analysing COI data set. D Tree obtained when analysing 16S data set. Bootstrap support values below nodes. Syllis and Typosyllis species as they were originally described in Ramisyllis kingghidorahi n. sp., a new branching annelid from Japan
◂Fig. 2 Maximum likelihood trees. A Tree obtained when analysing 18S data set. B Tree obtained when analysing 28S data set. C Tree obtained when analysing COI data set. D Tree obtained when analysing 16S data set. Bootstrap support values below nodes. Syllis and Typosyllis species as they were originally described
Validation data Set: Development and validation of a quantitative method for 15 antiviral drugs in poultry muscle using liquid chromatography coupled to tandem mass spectrometry
<p>Validation dataset for paper published in the Journal of Chromatography A.</p> <p> </p> <p>Clément Douillet, Mary Moloney, Melissa Di Rocco, Christopher Elliott, Martin Danaher,<br> Development and validation of a quantitative method for 15 antiviral drugs in poultry muscle using liquid chromatography coupled to tandem mass spectrometry, Journal of Chromatography A, Volume 1665, 2022, 462793, ISSN 0021-9673,</p> <p><br> Abstract:</p> <p>The objective of this work was to develop a quantitative multi-residue method for analysing antiviral drug residues and their metabolites in poultry meat samples. Antiviral drugs are not licensed for the treatment of influenza in food producing animals. However, there have been some reports indicating their illegal use in poultry. In this study, a method was developed for the analysis of 15 antiviral drug residues in poultry muscle (chicken, duck, quail and turkey) using liquid chromatography coupled to tandem mass spectrometry. This included 13 drugs against influenza and associated metabolites, but also two drugs employed for the treatment of herpes (acyclovir and ganciclovir). The method required the development of a novel chromatographic separation using a hydrophilic interaction chromatographic (HILIC) BEH amide column, which was necessary to retain the highly polar compounds. The analytes were detected using a triple quadrupole mass spectrometer operating in positive electrospray ionization mode. A range of different sample preparation protocols suitable for polar compounds were evaluated. The most effective procedure was based on a simple acetonitrile-based protein precipitation step followed by a further dilution in a methanol/water solution. The confirmatory method was validated according to the EU 2021/808 guidelines on different species including chicken, duck, turkey and quail. The validation was performed using various calibration curves ranging from 0.1 µg kg−1to 200 µg kg−1, according to the analyte. Depending on the analyte sensitivity, decision limits achieved ranged from 0.12 µg kg−1 for arbidol to 34.7 µg kg−1 for ribavirin. Overall, the reproducibility precision values ranged from 2.8% to 22.7% and the recoveries from 84% to 127%. The method was applied to 120 commercial poultry samples from the Irish market, which were all found to be residue-free.<br> Keywords: Antiviral drug residues; Influenza; HILIC; LC-MS/MS; Poultry muscle</p>
Data sets for "Bridgmanite freezing in shocked meteorites due to amorphization-induced stress" by Nishi et al.
<p>This is the datasets for the article "Bridgmanite freezing in shocked meteorites due to amorphization-induced stress" by Nishi et al. Tables S1 and S2 contain Experimental conditions and results. </p>
Data-set of CO2, CH4, N2O dissolved concentrations and ancillary data in surface waters of 24 African lakes.
<p>Geo-referenced and timestamped data-set of water temperature, Specific conductivity (SpCond), oxygen saturation level (%O<sub>2</sub>), dissolved methane (CH<sub>4</sub>) concentration, dissolved nitrous oxide (N<sub>2</sub>O) concentration, partial pressure of carbon dioxide (pCO<sub>2</sub>), carbon stable isotope composition of dissolved inorganic carbon (δ<sup>13</sup>C-DIC), dissolved organic carbon (DOC) concentration, chlorophyll-a (Chl-a) concentration, cyanobacteria abundance (CHEMTAX), nitrate (NO<sub>3</sub><sup>-</sup>) and ammonia concentration (NH<sub>4</sub><sup>+</sup>), coloured dissolved organic matter slope ratio (CDOM SR) in surface waters of African 24 lakes (Victoria, Tanganyika, Albert, Kivu, Edward, Mai Ndombe, Tumba, George, Kamohonjo, Alaotra, Ndalaga, Nyamusingere, Kyamwinga, Mbita, Lukulu, Yandja, Mbalukira, Nkugute, Nyamunuka, Kitagata, Mrambi, Kyashanduka, Katinda, Lac Vert).</p>
Data sets and codes
<p>Data sets and codes used for MPRA working paper "Foreign Direct Investment, Growth, and Publication Bias in Latin America and the Caribbean.</p>
Data sets for "Water enhancement of Si self-diffusion in wadsleyite" by D. Druzhbin et al.
<p>This is FTIR and SIMS data for the article "Water enhancement of Si self-diffusion in wadsleyite" by D. Druzhbin et al.</p>
MINIMAL DATA SET FOR PHC CLIENT SATISFACTION SURVEY
<p>Minimal data set for '<strong>Client Satisfaction with Reproductive, Maternal, Newborn and Child Health (RMNCH) Services at Primary Health Care Centres in Edo State, Nigeria'</strong></p>
Data Set of Thesis on Architectural Data Flow Analysis for Detecting Violations of Confidentiality Requirements
<p>The data set contains the results of the validation, the docker image for conducting the validation and the source code of all developed projects to conduct the validation.</p>
Data set from study: Identification of research priorities of radiography science – A modified Delphi study in Europe
<p>Data set consists of two round Delphi study conducted in Europe. The aim of the study was to identify research priorities in radiography science. The objective was to chart the opinions of radiography experts from different fields of radiography and different countries in Europe. </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.