Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,079
datasets available to search
ShareScore release 0.9.0
Dataset results
1,079 results for “source data”
Data and software associated with PHENOstruct: Prediction of human phenotype ontology terms using heterogeneous data sources
<p>Data and software associated with the paper:</p> <p>PHENOstruct: Prediction of human phenotype ontology terms using heterogeneous data sources</p>
Example Cytidine data set from I19-1 at Diamond Light Source
<p>Data set recorded in 6 scans, three omega and three phi, on the updated Diamond Light Source beamline I19-1, as part of routine commissioning. These data have been processed several times with different packages and are being made available to the community as an example set to allow other data processing software authors to verify the content of the data and headers.</p>
Experimental data for the publication: "Evaluating scintillator performance in time-resolved, hard X-ray studies at synchrotron light sources"
<p>In accordance with the expectations outlined in <em><strong>Clarifications of EPSRC expectations on research data management</strong></em> (09/10/14) this data has been made publicly available to complement the open access publication "Evaluating scintillator performance in time-resolved, hard X-ray studies at synchrotron light sources". </p> <p>There are six data sets, corresponding to the six experimental data sets presented in the article. In each data set, which may be identified by their file names and reference to the article, the 1st column is the RF trigger - to - ICCD exposure delay in [ns], and the second column in the intensity recorded on the ICCD in [counts]. This intensity accounts for any online and offline processing outlined in the article, such as on-CCD exposures, dark frame correction etc. </p>
Low dose, high multiplicity thermolysin X-ray diffraction data from Diamond Light Source beamline I03
<p>Low dose, high multiplicity X-ray diffraction data recorded from a thermolysin crystal prepared according to standard protocols as part of ongoing research. The data were recorded with low transmission to ensure minimal radiation damage, with the side-effect that the individual reflections are exceedingly weak even at low resolution, and the majority of background pixels have no counts.</p>
NiDPPE data recorded as part of beamline commissioning at Diamond Light Source I19-1
<p>This is an example data set recorded as part of beamline commissioning, consisting of a 360-degree phi scan with two-theta = 0 followed by three 170-degree omega scans with phi = 0, 120, 240 with two-theta offset to 30-degrees. Data recorded with a Pilatus 2M detector with 320um sensor, so the detector sensitivity needs to be considered carefully. Automatic processing using xia2 & DIALS gives:</p> <blockquote> <p>High resolution limit 0.65 3.56 0.65<br /> Low resolution limit 11.33 11.33 0.66<br /> Completeness 98.2 99.1 69.8<br /> Multiplicity 6.1 12.9 1.6<br /> I/sigma 12.0 71.2 0.9<br /> Rmerge 0.044 0.021 0.490<br /> Rmeas(I) 0.047 0.022 0.674<br /> Rmeas(I+/-) 0.047 0.022 0.674<br /> Rpim(I) 0.016 0.006 0.460<br /> Rpim(I+/-) 0.016 0.006 0.460<br /> CC half 1.000 1.000 0.732<br /> Total observations 57622 890 547<br /> Total unique 9400 69 332<br /> Assuming spacegroup: P 1 21/n 1<br /> Unit cell:<br /> 9.417 17.109 15.392<br /> 90.000 100.851 90.000</p> </blockquote>
L Cysteine Data collected 05/03/2016 at Diamond Light Source I19-1
<p>Data collected 05/03/2016 at Diamond Light Source I19-1 publicly available for users to test data reduction routines. Data are known to produce good merging statistics and final refinements. Three datasets collected on the same crystal with varying data collection times to explore beamline operational efficiency.</p> <p>l-cyst - collected with default set of runs that i19 used in the initial phase (i.e. 0.1 degree images, 0.1s per image)</p> <p>l-cyst_fast - collected with 0.1 degree images, 0.05s per image.</p> <p>l-cyst_very_fast - collected with 0.2 degree images, 0.05s per image.</p>
Example data from a grid scan of a thaumatin crystal recorded at Diamond Light Source beamline I03
<p>Example data recorded at Diamond Light Source beamline I03 as part of ongoing methods development. This is published to allow software authors to benchmark their analysis. Data were recorded with Pilatus3 6M detector as follows:</p> <p># 2016-04-27T10:02:44.731<br /> # Pixel_size 172e-6 m x 172e-6 m<br /> # Silicon sensor, thickness 0.001000 m<br /> # Exposure_time 0.0990000 s<br /> # Exposure_period 0.1000000 s<br /> # Tau = 0 s<br /> # Count_cutoff 1009797 counts<br /> # Threshold_setting: 6350 eV<br /> # Gain_setting: autog (vrf = 1.000)<br /> # N_excluded_pixels = 1161<br /> # Excluded_pixels: badpixel_mask.tif<br /> # Flat_field: FF_p60-0126_E12700_T6350_vrf_m0p100.tif<br /> # Trim_file: p60-0126_E12700_T6350.bin<br /> # Image_path: /ramdisk/2016/cm14451-2/20160427/gw/grid-test/0002/<br /> # Ratecorr_lut_directory: ContinuousStandard_v1.1<br /> # Retrigger_mode: 1<br /> # Wavelength 0.97625 A<br /> # Energy_range (0, 0) eV<br /> # Detector_distance 0.33810 m<br /> # Detector_Voffset 0.00000 m<br /> # Beam_xy (1246.42, 1208.60) pixels<br /> # Flux 0.000000<br /> # Filter_transmission 0.0100<br /> # Start_angle -0.0002 deg.<br /> # Angle_increment 0.0000 deg.<br /> # Detector_2theta 0.0000 deg.<br /> # Polarization 0.990<br /> # Alpha 0.0000 deg.<br /> # Kappa 0.0000 deg.<br /> # Phi 0.0000 deg.<br /> # Phi_increment 0.0000 deg.<br /> # Omega -0.0002 deg.<br /> # Omega_increment 0.0000 deg.<br /> # Chi 0.0000 deg.<br /> # Chi_increment 0.0000 deg.<br /> # Oscillation_axis X.CW<br /> # N_oscillations 1</p> <p> </p>
Figure 1. from Extending Marine Species Distribution Maps Using Non-Traditional Sources - Biodiversity Data Journal 3: e4900 (17 April 2015) https://doi.org/10.3897/BDJ.3.e4900
Figure 1. - The IUCN Red List Review Process (IUCN 2014a). Steps refer to the DOCUMENTATION STANDARDS AND CONSISTENCY CHECKS FOR IUCN RED LIST ASSESSMENTS AND SPECIES ACCOUNTS (IUCN 2013).
CALLISTO-SPK: A Stochastic Point Kinetics Code for Performing Low Source Nuclear Power Plant Start-up and Power Ascension Calculations Data Repository
<p>This dataset provides data to accompany the submission named "CALLISTO-SPK: A Stochastic Point Kinetics Code for Performing Low Source Nuclear Power Plant Start-up and Power Ascension Calculations" which has been submitted to Annals of Nuclear Energy. Details of the file included may be found in the readme file.</p>
Figures and data for the paper Perceptual Evaluation of Source Separation for Remixing Music
<p>Listening test results and figures for the paper:</p> <p>H. Wierstorf, D. Ward, R. Mason, E. M. Grais, C. Hummersone, M. D. Plumbley,<br> "Perceptual Evaluation of Source Separation for Remixing Music," in 143rd<br> Convention of the Audio Engineering Society, 2017.</p> <p>`fig02/data/` contains the results from single listeners and median results across<br> listeners.<br> `fig02/fig02.plt` is the code to regenerate `fig02/fig02.pdf`.</p> <p>`fig03/data/` contains average medians across all songs.<br> `fig03/fig03.plt` is the code to regenerate `fig03/fig03.pdf`.</p> <p>The listening test results were extracted from the raw data stored together with<br> the experimental procedure in https://doi.org/10.5281/zenodo.835191.<br> There the file `experiment/saves/ratings/analyze_results.py` can be run in order<br> to regenerate the result files provided with this publication.<br> </p>
Beta-Lactamase X-ray diffraction data recorded at Diamond Light Source I04 as part of commissioning & development
<p>X-ray diffraction data from crystals of beta-lactamase recorded during commissioning. The data were recorded in two omega scans of 3600 images @ 0.1 degrees / frame with different kappa and phi settings on a mini kappa device, with a Dectris PILATUS2 6M detector, at a wavelength of 1.239850A (10keV) for remote-SAD on the native Zn site.</p> <p> </p> <p>Data uploaded for education and tutorial purposes as a good quality example set using multi-axis geometry. </p> <p> </p> <p>For convenience the data are compressed with gzip and combined into one tar file for each sweep.</p>
Replication data for Anterior Cruciate Ligament Injury: identifying information sources and risk factor awareness among the general population
<p>Creating awareness about a disorder is important for prevention from the perspective of public health. However, for sports injuries, like the anterior cruciate ligament (ACL) injury, there is no study which has investigated the awareness of risk factors for the injury and prevention methods among the general population, to the best of our knowledge. The sources of information among the population are also unclear. The purpose of present study was to identify these aspects of public awareness about the ACL injury.</p> <p>A questionnaire was randomly distributed among the general population registered with a web based questionnaire supplier, to recruit 900 participants who were aware about the ACL injury. The questionnaire consisted of two parts: Question 1 asked them about their sources of information regarding the ACL injury; Question 2 asked them about the risk factors for ACL injury. Multivariate logistic regression was used to determine the information sources that provide a good understanding of the risk factors.</p> <p>The leading source of information for ACL injury was television (57.0%). However, the results of logistic regression analysis revealed that television was not an effective medium to create awareness about the risk factors, among the general population. Instead “Lecture by a coach”, “Classroom session on Health”, and “Newspaper” were significantly effective in creating a good awareness of the risk factors (p < 0.001).</p>
Scenario data, model source code and plotting routine for manuscript: Separating CO2 emission from removal targets comes with limited cost impacts
<p>This data archive contains REMIND model setup, results data and data analysis files for manuscript:<br><strong>Separating CO2 emission reduction from removal targets comes with limited cost impact.<br><br>plotting</strong>(directory) contains results data, manuscript specific data analysis and plotting routine scripts used to generate the figures of the manuscript.<br><strong>remind</strong>(directory) contains REMIND model source code and scenario set-up. Detailed scenario configurations are set in remind/config/scenario_config_SepMark.csv.<br><strong>remind2</strong>(directory) contains the slightly modified R-library package used for post-processing of REMIND output.<br><br>AMENDMENT<br><strong>Plots_SeparateMarkets_afterReviewProcess.Rmd</strong> After the review process, the new plotting script was added including the additional figures in the Supplementary Material. This file should replace the previous R-markdown file SepMark_essential/plotting/Plots_SeparateMarkets.Rmd.</p>
AnalyzAIRR: A user-friendly guided workflow for AIRR data analysis: example data and analysis source-code
<p>This repository contains:</p> <ul> <li>Annotated TCR-seq data files named <em>tripod-XX-XXXX</em></li> <li>The metadata corresponding to the annotated files</li> <li>The RepSeqExperiment object, which integrates the annotated files and the metadata and was used in the analysis pipeline</li> <li>The analysis script to generate the plots of the different figures</li> </ul>
Data Supporting: Identification of Nonlinearity Sources in a Flexible Wing
Open the record for dataset details and reuse information.
Source Data and Scripts - Reconstitution of Human Brain Cell Diversity in Organoids via Four Protocols
<p>Source data and scripts associated with the manuscript "<em>Reconstitution of Human Brain Cell Diversity in Organoids via Four Protocols</em>" (Naas et al. 2024, <em>bioRxiv</em>, <a href="https://doi.org/10.1101/2024.11.15.623576" rel="nofollow">DOI: 10.1101/2024.11.15.623576</a>). Corresponding scripts to reproduce all figures and tables presented in the manuscript are also available on <a href="https://github.com/jn-goe/brain_organoids_four_protocols">https://github.com/jn-goe/brain_organoids_four_protocols</a>.</p> <div> <div> <p>The therein introduced NEST-Score is available as R package on <a href="https://github.com/jn-goe/NESTScore">https://github.com/jn-goe/NESTScore</a>.</p> <p>The interactive Shiny App data explorer is available on <a href="https://vienna-brain-organoid-explorer.vbc.ac.at/">https://vienna-brain-organoid-explorer.vbc.ac.at</a>.</p> </div> </div>
The Impact of Ammonia on Particle Formation in the Asian Tropopause Aerosol Layer: data sources
<p>Model namelist, Configuration settings, Figure data for manuscript "The impact of ammonia on particle formation in the Asian Tropopause Aerosol Layer".</p>
Dataset: A continuous open source data collection platform for architectural technical debt assessment
<p>The dataset and replication package of the study "A continuous open source data collection platform for architectural technical debt assessment".</p> <p> </p> <p>Abstract</p> <p>Architectural decisions are the most important source of technical debt. In recent years, researchers spent an increasing amount of effort investigating this specific category of technical debt, with quantitative methods, and in particular static analysis, being the most common approach to investigate such a topic.</p> <p> </p> <p>However, quantitative studies are susceptible, to varying degrees, to external validity threats, which hinder the generalisation of their findings.</p> <p>In response to this concern, researchers strive to expand the scope of their study by incorporating a larger number of projects into their analyses. This practice is typically executed on a case-by-case basis, necessitating substantial data collection efforts that have to be repeated for each new study.</p> <p> </p> <p>To address this issue, this paper presents our initial attempt at tackling this problem and enabling researchers to study architectural smells at large scale, a well-known indicator of architectural technical debt. Specifically, we introduce a novel approach to data collection pipeline that leverages Apache Airflow to continuously generate up-to-date, large-scale datasets using Arcan, a tool for architectural smells detection (or any other tool).</p> <p>Finally, we present the publicly-available dataset resulting from the first three months of execution of the pipeline, that includes over 30,000 analysed commits and releases from over 10,000 open source GitHub projects written in 5 different programming languages and amounting to over a billion of lines of code analysed.</p>
Theory and implementation of inelastic Constitutive Artificial Neural Networks: Source code and data
<p>This dataset contains the source code of the inelastic Constitutive Artificial Neural Network (iCANN) as well as the data for the examples from the publication:</p> <p>Holthusen, H., Lamm, L., Brepols, T., Reese, S., & E. Kuhl.<em> Theory and implementation of inelastic Constitutive Artificial Neural Networks.</em></p> <p>arXiv: <a href="https://doi.org/10.48550/arXiv.2311.06380">https://doi.org/10.48550/arXiv.2311.06380</a></p> <p>Computer Methods in Applied Mechanics and Engineering: <a href="https://doi.org/10.1016/j.cma.2024.117063">https://doi.org/10.1016/j.cma.2024.117063</a></p> <p> </p> <p><strong>01_Example01: </strong> Artificially generated data</p> <p>This example investigates whether the iCANN is able to discover a model for the data generated by a continuum mechanical model.</p> <p> </p> <p><strong>02_Example02:</strong> Discovering a model for the polymer VHB 4910 subjected to cyclic loading</p> <p>Here, we investigate the ability of iCANN to discover and learn a model for the material response of VHB 4910 polymer subjected to cyclic loading at different stretch rates.</p> <p>The experimental data are taken from the literature:</p> <p>Hossain, M., Vu, D. K., & Steinmann, P. (2012). Experimental study and numerical modelling of VHB 4910 polymer. <em>Computational Materials Science</em>, <em>59</em>, 65-74.</p> <p><a href="https://doi.org/10.1016/j.commatsci.2012.02.027">https://doi.org/10.1016/j.commatsci.2012.02.027</a></p> <p> </p> <p><strong>03_Example03: </strong>Discovering a model for passive skeletal muscle subjected to relaxation</p> <p>In this example, we investigate whether the iCANN is able to discover a model for the material behavior of passive skeletal muscles. A total of five independent experiments are carried out in which the maximum applied compression stretch and the stretch rate are varied. In addition, the learning performance of the iCANN is investigated. Training is first carried out in each of the five experiments and then in each of four of the five experiments.</p> <p>The experimental data are taken from the literature:</p> <p>Van Loocke, M., Lyons, C. G., & Simms, C. K. (2008). Viscoelastic properties of passive skeletal muscle in compression: stress-relaxation behaviour and constitutive modelling. <em>Journal of biomechanics</em>, <em>41</em>(7), 1555-1566.</p> <p><a href="https://doi.org/10.1016/j.jbiomech.2008.02.007">https://doi.org/10.1016/j.jbiomech.2008.02.007</a></p> <p> </p> <p><strong>python_requirements.txt: </strong>File containing a list of installed Python modules used to implement the iCANN</p>
Distinct mobility patterns of BRCA2 molecules at DNA damage sites - source data
<p>This repository contains source data from the publication "Distinct mobility patterns of BRCA2 molecules at DNA damage sites"</p><p>- Source data of FRAP experiments. Text files contain time and intensity data of the individual FRAP curves including the normalized intensities as plotted in the figure.</p><p>- Source data for BRCA2 protein quantification </p><p>- dSTORM localization data</p><p>- Single-molecule tracks of BRCA2-Halo Mitomycin and untreated tracks. Original movies can be found here: http://doi.orig/10.5281/zenodo.10144073</p><p>For questions please contact Maarten Paul (m.w.paul [at] erasmusmc.nl</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.