Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
3,481
datasets available to search
ShareScore release 0.9.0
Dataset results
3,481 results for “data set”
Teaching ESE Survey data set
<p>Teaching ESE Survey data set</p>
Data set for "Local and landscape correlates of coccinellid species richness, abundance, and assemblage change along a rural–urban gradient in Quintana Roo, Mexico."
<p>Data set for original article "Local and landscape correlates of coccinellid species richness, abundance, and assemblage change along a rural–urban gradient in Quintana Roo, Mexico"<strong> </strong>accepted at Biotropica.</p> <p>We studied changes in coccinellid assemblage composition, diversity, species richness and abundance in domestic gardens along a rural to urban gradient in southeastern Mexico, and identified local and lanscape variables that determine their species richness and abundance. </p>
Data set for: History-dependent domain and skyrmion formation in 2D van der Waals magnet Fe3GeTe2
<p>This data repository contains the experimental data utilised to produce the figures shown in History-dependent domain and skyrmion formation in 2D van der Waals magnet Fe3GeTe2, by M. T. Birch, L. Powalla, S. Wintz, O. Hovorka, K. Litzius, J. C. Loudon, L. A. Turnbull, V. Nehruji, K. Son, C. Bubeck, T. G. Rauch, M. Weigand, E. Goering, M. Burghard, G. Schütz.</p> <p>Experimental data analysis were performed using Python via Jupyter notebooks, utilising matplotlib, scipy, numpy and h5py libraries. Jupyter notebooks are provided to process some of the experimental data to form the figures shown in the paper.</p> <p><br> Data data is divided into folders:</p> <p>Characterisation_Data - contains EDX, TEM and AFM characterisation of the bulk FGT material and exfoliated flakes.</p> <p>Magnetometry_Data - Includes data files from the MPM3 magnetometry utilised to study the magnetisation of the bulk FGT crystal as a function of temperature and applied magnetic field. The two sub folders divide the files into data acquired with the field applied either parallel to the c axis (OOP - out of plane), or perpendicular to the c axis (IP - in plane).</p> <p>STXM_Data - Contains Scanning Transmission X-ray Microscopy data acquired at the MAXYMUS instrument at the BESSY II synchrotron, Berlin. The data is divided into sections acquired following different sample histories, as reported in the main paper: field sweep, field-cooled and zero-field cooled. Folders containing the extracted skyrmion sizes and stripe domain sizes are included, as well as Python notebooks for plotting Fig. 3, 4 and 5 of the main paper.</p> <p>Spectra_Data - Contains x-ray absorption spectra acquired on a bulk FGT crystal via total electron yield method (Bulk spectra), and spectra acquired on the exfoliated FGT flake by a transmission method (Flake Spectra). The accompanying jupyter notebook demonstrates the plotting and application of XMCD sum rules to the experimental data.</p> <p>LTEM_Data - Contains images acquired via Lorentz Transmission Electron Microscopy of an FGT flake, measured as a function of temperature and applied magnetic field at various tilt angles.</p> <p>Analysis_Scripts - Contains python notebook for plotting the magnetic phase diagrams, and plotting the thickness dependence estimation from the x-ray absorption spectra data.</p> <p>Meanfield_Simulations - Contains the raw output files of the simulations, the .h5 files, and Python notebook plotting scripts. Summary pdf files show the output of all simulation runs for the field sweep performed at each temperature. </p> <p><br> The latest version of this Data Set can be found at the Zenodo DOI:<br> 10.5281/zenodo.6346695</p> <p>Any further requests for materials are welcome, send any inquiries to Max Birch at birch@is.mpg.de</p>
Linearity Data Set
<p>Data behind Linearity Fig 1, Fig 2, and Fig 3, and the full data set for 175 SPARC galaxies</p>
OMEN data set: Tokheim pancancer data set with genome-wide interaction network
<p>OMEN is a Network-based Driver Gene Identification method that exploits Mutual Exclusivity.<br> This repository stores the data used in a pancancer experiment showcasing this method.</p> <p>It consists of</p> <p>- [tokheim_pancancer_somatic_CADD.pl] A file containing CADD probabilities for gene-patient pairs (derived from the pancancer data set in Tokheim, Collin J., et al. "Evaluating the evaluation of cancer driver genes." <em>Proceedings of the National Academy of Sciences</em> 113.50 (2016): 14330-14335.) formatted to be used by OMEN.<br> - [tokheim_pancancer_somatic_coverage_ranks.pl] A file containing gene coverage data (derived from the CADD data) formatted to be used by OMEN.<br> - [network.pl] A file containing a genome-wide interaction network consisting of high quality metabolic interactions from Recon X and literature curated interactions from Intact. Recon X and Intact data was acquired from Pathway Commons version 8.<br> The resulting network covers 7901 samples, and contains 15.694 nodes and 178.051 edges.</p>
In-vitro Major Arterial Cardiovascular Simulator: Benchmark Data Set for in-silico Model Validation
<p><strong>Background</strong><br> <br> The data described here supplements the paper "In-vitro Major Arterial Cardiovascular Simulator to generate Benchmark Data Sets for in-silico Model Validation" (to be submitted). It was created at Technische Hochschule Mittelhessen (THM) in Germany and uploaded to Zenodo. Please cite the paper M. Wisotzki, A. Mair, P. Schlett, B. Lindner, M. Oberhardt, S. Bernhard, In Vitro Major Arterial Cardiovascular Simulator to Generate Benchmark Data Sets for In Silico Model Validation (2022), Data 7(11), DOI: 10.3390/data7110145 and the Zenodo doi when using this dataset.</p> <p><strong>General description / Dataset Structure</strong></p> <p>Each mat-File describes a different stenosis degree at the popliteal artery of the in-vitro simulator MACSim (details can be found in the paper). There are 17 pressure signals for different positions, one flow sensor close to the stenosis location and one monitor signal of the proportional valve use to control the input curve. Total duration of each signal is 60s with a sampling rate of 1000 Hz. Each mat-file contains a header structure with metadata and struct array for signals of each sensor. Signals in each mat-File are aligned with respect to a common time axis, but this is not guaranteed between different measurements/files. The file format can either be loaded directly in Matlab or in Python with scipy's loadmat function.</p> <p>The different stenosis degrees for each degree are:<br> ScenarioI: 100 % Area fraction (no stenosis)<br> ScenarioII: 37,5 % Area fraction<br> ScenarioIII: 23,4 % Area fraction<br> ScenarioIV: 6,56 % Area fraction</p> <p><strong>Data fields for each file</strong></p> <table> <caption>headerStruct</caption> <thead> <tr> <th scope="col">field</th> <th scope="col">description</th> </tr> </thead> <tbody> <tr> <td>rate</td> <td>sampling rate in Hz</td> </tr> <tr> <td>description</td> <td>name of the scenario according to the paper, corresponds to filename</td> </tr> <tr> <td>configuration</td> <td>parameters of the trapezoidal input curve (offset and amplitude in mmHg, ascend times and descend times and smoothing window in a fraction the time period (1.2s))</td> </tr> </tbody> </table> <p> </p> <table> <caption>signalStruct</caption> <thead> <tr> <th scope="col">field</th> <th scope="col">description</th> </tr> </thead> <tbody> <tr> <td>nodeId</td> <td>corresponds to numbered nodes at which the sensor is placed, the corresponding location can be found in the paper (node numbering, not sensor numbers) or in the software SISCA (https://gitlab.com/agbernhard.lse.thm/sisca) in the example database.</td> </tr> <tr> <td>type</td> <td>'p' ... pressure or 'q' ... flow</td> </tr> <tr> <td>data</td> <td>double array, time series of each sensor, unit mmHg for type 'p' and ml/s for type 'q' </td> </tr> <tr> <td>anatomicalPosition</td> <td> <p>name of the corresponding anatomical position</p> </td> </tr> </tbody> </table>
Full data set for "Real-Space Observation of Fluctuating Antiferromagnetic Domains" article in Science Advances
<p>This is the full data set for the article, sufficient to reproduce any results shown. The data and data files description, and data retrieval instructions are in the Readme.txt file. </p>
A synthetic fraud detection data set.
<p>A synthetic fraud detection data set created using sklearn's make_blob for use in a blog.</p> <p>X, y = datasets.make_blobs(n_samples=[800000,200000], centers=None, cluster_std=[10.0, 2],random_state=42,n_features=4)</p>
Data sets for "Formation of an Al-rich niccolite-type silica in subducted oceanic crust: implications for water transport to the deep lower mantle"
<p>This is the XRD and IR datasets for the article "Formation of an Al-rich niccolite-type silica in subducted oceanic crust: implications for water transport to the deep lower mantle" by Liu and Yuan et al.</p>
X-ray computed tomography aided engineering approach for non-crimp fabric reinforced composites [Data set]
<p>The finite element models behind the publication</p> <p>Auenhammer, R.M., Jeppesen, N., Mikkelsen, L.P., Dahl, V.A., Blinzler, B.J., Asp, L.E. Robust numerical analysis of fibrous composites from X-ray computed tomography image data enabling low resolutions, <em>Composites Science and Technology, </em><strong>224</strong>, 109458, <a href="https://doi.org/10.1016/j.compscitech.2022.109458">https://doi.org/10.1016/j.compscitech.2022.109458</a>, 2022. </p> <p>The x-ray scan data which the model is based on can be found in the following publication:</p> <p>Jeppesen, N., V.A. Dahl, A.N. Christensen, A.B. Dahl, L.P. Mikkelsen, Characterization of the fiber orientations in non-crimp glass fiber reinforced composites using structure tensor. IOP Conf. Ser.: Mater. Sci. Eng. 942, 012037, <a href="https://doi.org/10.1088/1757-899X/942/1/012037">https://doi.org/10.1088/1757-899X/942/1/012037</a> 2020</p> <p>and data-set</p> <p>Jeppesen N, Dahl V A, Christensen A N, Dahl A B and Mikkelsen L P 2020 Characterization of the fiber orientations in non-crimp glass fiber reinforced composites using structure tensor [data set] Zenodo. <a href="http://dx.doi.org/10.5281/zenodo.3877522">http://dx.doi.org/10.5281/zenodo.3877522.</a></p> <p> </p>
Bulk components and diacids and related compounds in size-segregated aerosols at Tianjin, North China – Data set
<p>To better understand the sources of atmospheric aerosols in winter, we collected size-resolved (9-stage) aerosols at Tianjin, North China, and studied for inorganic ions, carbonaceous components, diacids and related compounds. Inorganic ions were measured using ion chromatography (ICS-5000 System, China, Dai An). Organic carbon (OC) and elemental carbon (EC) were measured using OC/EC analyzer (USA, Sunset Laboratory Inc.), based on thermal light reflectance following the IMPROVE protocol of the protective visual environment. Water-soluble organic carbon (WSOC) was measured using the total organic carbon (TOC) analyzer (model: OI, 1030W + 1088). Concentrations of diacids and related compounds were measured using a capillary gas chromatography (GC; Agilent 7890B). Peak identification was carried out with reference to authentic standards retention time measured by GC-FID and mass spectra performed using a GC-mass spectrometry (GC-MS) system. Based on the results obtained, we found that the atmospheric aerosols over the Tianjin region, North China were derived from mixed: both anthropogenic and biogenic, sources and the subsequent secondary formation/transformations were also important in winter.</p>
Data set for manuscript "Optimal utilization of PMTCT of HIV services among adolescents under group versus focused antenatal care"
<p>The derived variables in the data set are labeled "DER what variables they are derived from" except the variable "optimal utilization which is a composite outcome of other variables depending on HIV status as described in the manuscript.</p> <p>The variables "...cat" are categorized or recoded from an original numeric/ categorical variable</p> <p>The variables bin are used when a "cat" variable exists but further recoding is done to form a binary variable</p>
Data from: Implementation of a pediatric telemedicine and medication delivery service in a resource-limited setting: A pilot study for clinical safety and feasibility
<p>Objective: Determine the clinical safety and feasibility of implementing a telemedicine and medication delivery service (TMDS) to address gaps in nighttime healthcare access for children in low-resource settings.</p> <p>Results: A total of 391 cases were enrolled from September 9th, 2019 to January 19th, 2021; 89% (347) received a household visit. Most cases were triaged as mild or moderate (92%; 361). Among the severe cases, 83% (20) sought subsequent referred care. The most common complaint was a respiratory problem (63%; 246). At 10-days, 95% (329) of parents reported their child's condition as "improved" or "recovered". Ninety-nine percent (344) rated the TMDS as "good" or "great". The median phone consultation was 20 minutes, time to arrival at the household was 73 minutes and total workflow per case was 114 minutes.</p> <p>Conclusion: The TMDS was a feasible healthcare delivery model with high rates of improved clinical status at 10-days.</p>
Data set - Island area and historical geomorphological dynamics shape multifaceted diversity of barrier island floras
<div> <p><span>The influence of island dynamics and characteristics on taxonomic diversity, particularly species richness, are well studied. Yet, our knowledge on the influence of island dynamics and characteristics on other facets of diversity, namely functional and phylogenetic diversity, is limited, constraining our understanding of assembly processes on islands (e.g., biogeographic history, dispersal and environmental filtering, and species interactions). Using barrier islands, a highly dynamic and so far, understudied island type, we investigate how multiple facets of vascular plant diversity (functional, phylogenetic and taxonomic diversity) are shaped by island geomorphology, modern and historic area, and habitat heterogeneity. In line with our expectation, historical dynamics in island geomorphology affected phylogenetic and taxonomic diversity via habitat heterogeneity. However, island area was the best predictor across all facets of diversity. Specifically, larger islands had higher functional and phylogenetic diversity than expected by chance while most of the smaller islands had lower diversity. The influence of area on functional diversity acted via habitat heterogeneity, with habitat heterogeneity influencing negatively functional diversity. Our results suggest that larger islands accumulate functionally and phylogenetically unique species. Further, results for functional diversity pointed towards potential area-heterogeneity trade-offs, with these trade-offs likely resulting from increased interspecific competition favoring a specific set of trait values (of stronger competitors), particularly on smaller islands. Together, these results demonstrate that going beyond taxonomic diversity contributes to identifying underlying processes shaping diversity-area relationships. </span></p> </div>
Data Set used in "Full backward and forward dependencies through regional hypothetical extraction method"
<p>This set of data was obtained from EUREGIO database, developed by the Tinbergen Institute, which is a set of global IO tables with regional and sectoral disaggregation. The EUREGIO database collects the productive structure and commercial relations of the WIOD in the period 2000-2010. The table is broken down into 249 administrative regions at the NUTS2 level, from 24 EU countries, 16 non-EU countries, and a block that brings together countries from the rest of the world, making a total of 266 regions. The statistical information is organised in 11 IO tables, one for each year.</p> <p>The data base that we provid in this repository is used in our study with the aim to determine the key regions of the Spanish economy. In order to address this objective, IO tables of smaller dimensions are built, through an aggregation and disaggregation procedure. First, the 14 industries are grouped, then the 4 sectors of final demand and, lastly, the 4 components of value added. Below, the 266 EUREGIO regions are grouped into 21 regions. Of these, 19 regions correspond to Spain <a href="#_ftn1">[1]</a>, one region includes the rest of the NUTS2 in the EU and another region covers the rest of the world.</p> <p><a href="#_ftnref1">[1]</a> The 17 Spanish regions and the two autonomous cities of Ceuta and Melilla.</p>
Data set: The origin of sex differences in song in a tropical duetting wren
<p>The study of song development has focused on temperate zone birds in which typically only males sing. In the bay wren, Cantorchilus nigricapillus, both sexes sing, performing precisely timed, female-initiated duets in which birds alternate sex-specific song phrases. We investigated the origin of these sex differences by collecting bay wren eggs and nestlings and hand-raising them in individual acoustic isolation chambers. Each bird was tutored with either monophonic or stereophonic recordings of bay wren duets, or heard no song. As adults, each tutored bird individually sang complete duets, singing both male and female song phrases. On occasion, birds learned only the male or female part of a duet to which they were exposed. However, mono-tutored birds showed no sex-specificity in these solo songs, whereas stereo-tutored birds only sang solos consistent with their sex. In addition, stereo-tutored birds acquired songs over a longer period than did mono-tutored birds. In both groups, females showed more sex-specificity during the song learning process. Finally, we observed that tutored and acoustically isolated birds invented male-like songs, whereas only males invent songs in the wild. These results reveal the relative roles played by social versus innate influences in the development of sex-specific song in this species.</p>
Data sets compiled for review article on Gender Equity in Oceanography
<p>Data sets compiled for a review paper on Gender Equity in Oceanography, to appear in Annual Review of Marine Science, 2023. Includes: Annual Reviews of Marine Science invited author gender statistics; China and USA oceanography career progression gender statistics; Current employment fractions of 2010-2019 physical oceanography PhDs; Gordon Research Conference Oceanography gender statistics; JGR oceans author and reviewer gender statistics; Oceanography faculty gender statistics for China; Oceanography graduate degree gender statistics for China; Oceanography magazine author gender data. </p>
Data set for the article "Self-oscillation and Synchronisation Transitions in Elasto-Active Structures"
<p>This is the data set for the article "Self-oscillation and Synchronisation Transitions in Elasto-Active Structures", </p> <table summary="Additional metadata"> <tbody> <tr> <td><a href="https://doi.org/10.48550/arXiv.2106.05721">https://doi.org/10.48550/arXiv.2106.05721</a> <p> </p> </td> </tr> </tbody> </table> <p>Is contains raw images and processed data from images used to describe the self-oscillations under study in this article.</p> <p>All zip files corresponds to the single chain experiment except "Double_chain_Experiment.zip".</p> <p> </p>
data set regarding to project – Sarcopenia, obesity, sarcopenic obesity and risk of PNS in Polish older people
<p><strong>This data set corresponds with the article titled: Sarcopenia, obesity, sarcopenic obesity and risk of poor nutritional status in Polish community-dwelling older people aged 60 years and over.</strong></p>
Data set from "Absolute Treatment Effects for the Primary Outcome and All-cause Mortality in the Cardiovascular Outcome Trials of New Antidiabetic Drugs – A Meta-Analysis of Digitalized Individual Patient Data"
<p>This data set contains the complete digitalized individual patient data that are used in the manuscript "Absolute Treatment Effects for the Primary Outcome and All-cause Mortality in the Cardiovascular Outcome Trials of New Antidiabetic Drugs – A Meta-Analysis of Digitalized Individual Patient Data", which is accepted from "Acta Diabetologica".<br> The data file is in CSV format and contains the five variables "OutcomeType" (with values "AllcauseMortality" or "PrimaryOutcome"), "Study" (denoting the respective cardiovascular outcome trial), "Treatment" (denoting the respective study treatment or placebo), "Event" (denoting if the respective has been observed (Event=1) or not (Event=0)), and "SurvivalTimeMonths" (denoting the respective time to the event or censoring in months).</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.