Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,079
datasets available to search
ShareScore release 0.9.0
Dataset results
1,079 results for “source data”
Characterization of data sources for emerging risks identification - Data collection for the identification of emerging risks related to food and feed
<p>The data set presents the results of the quality assessment of data sources by the working group on data collection for the identification of emerging risks related to food and feed. For this assessment, the WG defined text descriptors and quality parameters (i.e. link with indicators, data type, geographic and period coverage, language, edition, timeliness, accessibility, clarity and comparability). These data sources were linked to eleven priority indicators (i.e. the ESCO indicators) and qualitatively assessed and profiled.</p>
Data presented in "A buffer gas beam source for short, intense and slow molecular pulses"
<p>Data presented in our paper "A buffer gas beam source for short, intense and slow molecular pulses". The files give the data shown in figures 4 and 5 of the paper.</p>
Open-source quality control routine and multi-year power generation data of 175 PV systems
<p><strong>Description</strong></p> <p>The repository contains an extensive dataset of PV power measurements and a python package (qcpv) for quality controlling PV power measurements. The dataset features four years (2014-2017) of power measurements of 175 rooftop mounted residential PV systems located in Utrecht, the Netherlands. The power measurements have a 1-min resolution.</p> <p><strong>PV power measurements</strong></p> <p>Three different versions of the power measurements are included in three data-subsets in the repository. Unfiltered power measurements are enclosed in <em>unfiltered_pv_power_measurements.csv</em>. Filtered power measurements are included as <em>filtered_pv_power_measurements_sc.csv </em>and<em> filtered_pv_power_measurements_ac.csv</em>. The former dataset contains the quality controlled power measurements after running single system filters only, the latter dataset considers the output after running both single and across system filters. The metadata of the PV systems is added in<em> metadata.csv</em>. This file holds for each PV system a unique ID, start and end time of registered power measurements, estimated DC and AC capacity, tilt and azimuth angle, annual yield and mapped grids of the system location (north, south, west and east boundary).</p> <p><strong>Quality control routine</strong></p> <p>An open-source quality control routine that can be applied to filter erroneous PV power measurements is added to the repository in the form of the Python package qcpv (<em>qcpv.py</em>). Sample code to call and run the functions in the qcpv package is available as <em>example.py.</em></p> <p><strong>Objective</strong></p> <p>By publishing the dataset we provide access to high quality PV power measurements that can be used for research experiments on several topics related to PV power and the integration of PV in the electricity grid.</p> <p>By publishing the qcpv package we strive to set a next step into developing a standardized routine for quality control of PV power measurements. We hope to stimulate others to adopt and improve the routine of quality control and work towards a widely adopted standardized routine. </p> <p><strong>Data usage</strong></p> <p>If you use the data and/or python package in a published work please cite: <em>Visser, L., Elsinga, B., AlSkaif, T., van Sark, W., 2022. Open-source quality control routine and multi-year power generation data of 175 PV systems. Journal of Renewable and Sustainable Energy.</em></p> <p><strong>Units</strong></p> <p>Timestamps are in UTC (YYYY-MM-DD HH:MM:SS+00:00).</p> <p>Power measurements are in Watt.</p> <p>Installed capacities (DC and AC) are in Watt-peak.</p> <p><em><strong>Additional information</strong></em></p> <p>A detailed discussion of the data and qcpv package is presented in: <em>Visser, L., Elsinga, B., AlSkaif, T., van Sark, W., 2022. Open-source quality control routine and multi-year power generation data of 175 PV systems. Journal of Renewable and Sustainable Energy. Corrections are discussed in: Visser, L., Elsinga, B., AlSkaif, T., van Sark, W., 2024. </em><em>Erratum: Open-source quality control routine and multiyear power generation data of 175 PV systems. Journal of Renewable and Sustainable Energy.</em></p> <p><strong>Acknowledgements </strong></p> <p>This work is part of the Energy Intranets (NEAT: ESI-BiDa 647.003.002) project, which is funded by the Dutch Research Council NWO in the framework of the Energy Systems Integration & Big Data programme. The authors would especially like to thank the PV owners who volunteered to take part in the measurement campaign. </p>
Source data for analysis of super-enhancer interactomes v1
<p><strong>Super-enhancer interactomes from single-cells link clustering and transcription</strong></p> <p>Derek Le, Antonina Hafner, Sadhana Gaddam, Kevin Wang, Alistair Boettiger</p> <p>This Zenodo repository contains data files associated with our analysis.</p> <p>Software for processing the data is available in the associated github repository: https://github.com/BoettigerLab/SEclustering-2024</p> <p>Additional information can be found in the associated manuscript, currently in preparation -- once it is posted on BioRxiv, it will be linked here. </p> <p>This deposition currently includes<br>1) SuperEnhancerLoci.xlsx - a master data table linking the genomic sequence barcode data from the corrected tables (described below) to the corresponding super-enhancer and their genomic coordinates in mm10.<br>2) Corrected_Data_Tables_by_FOV.zip -- contains drift corrected and chromatically corrected x,y,z coordinates and cellular barcode data to track cell type and coordinate barcode data to identify genomic sequences. <br>3) Processed_Seq_Data.zip -- re-processed sequencing based data used in this study.<br>4) FOF-CT_Spot_tables.zip -- draft versions of the 4DN FOF-CT formatted data-standard spot tables. See data format description here: https://fish-omics-format.readthedocs.io/en/latest/ <br>5) Probe_Sequences.zip -- fasta files, bed files, and codebook tables for the RNA and DNA probe sequences used in this study.</p>
Supporting material for: MoonIndex, an Open-Source Tool to Generate Spectral Indexes for the Moon from M3 Data
<p>Supplementary material for the paper called: MoonIndex, an Open-Source Tool to Generate Spectral Indexes for the Moon from M3 Data. The data without "indexes" in the name are map-projected M3 cubes, they can be used in the python library <i><strong>MoonIndex </strong></i>to obtain the spectral indexes stored in the files with "indexes" in the name.</p><p>This research was done on the framework of the EXPLORE project, that has received funding from the European Union's 2020 research and innovation program under grant agreement No 101004214. </p>
Source Data for main text and supplemental figures for manuscript "Community assessment of methods to deconvolve cellular composition from bulk gene expression"
<p>Source Data for main text and supplemental figures for manuscript "Community assessment of methods to deconvolve cellular composition from bulk gene expression"</p>
Dense vegetation hinders sediment transport towards saltmarsh interiors - Supporting data and source code (Part III: Extra runs)
<p>This is Part III of the supporting data and source code for the paper entitled "Dense vegetation hinders sediment transport towards saltmarsh interiors", submitted to <em>Limnology and Oceanography Letters.</em> It contains all input and output files for the extra simulations used in the paper (Figures S3, S8-S10).</p> <p>Each zip file corresponds to a model run. </p> <p>TIGER_XX.zip: Scenario XX, hydro-morphodynamics and vegetation dynamics, years 0-100.<br>TIGER_XX_100.zip: Scenario XX, hydro-morphodynamics and vegetation dynamics, years 100-200.<br>TIGER_XX_HYYY.zip: Scenario XX, hydro-morphodynamics only, year YYY.</p> <p>Main scenarios:<br>- 01: Spartina (Figures 1-5, S3-S10)<br>- 02: Salicornia (Figures 1-5, S3-S10)<br>- 83: No vegetation (Figures 1-5, S3, S8-S10)</p> <p>Additional scenarios:<br>- 146: Spartina, low bulk drag coefficient (Figure S3)<br>- 147: Spartina, very low bulk drag coefficient (Figure S3)<br>- 148: Salicornia, low bulk drag coefficient (Figure S3)<br>- 149: Salicornia, very low bulk drag coefficient (Figure S3)<br>- 122: Spartina, low settling velocity (Figure S8)<br>- 123: Spartina, high settling velocity (Figure S8)<br>- 124: Salicornia, low settling velocity (Figure S8)<br>- 125: Salicornia, high settling velocity (Figure S8)<br>- 126: No vegetation, low settling velocity (Figure S8)<br>- 127: No vegetation, high settling velocity (Figure S8)<br>- 128: Spartina, low critical bed erosion shear stress (Figure S8)<br>- 129: Spartina, high critical bed erosion shear stress (Figure S8)<br>- 130: Salicornia, low critical bed erosion shear stress (Figure S8)<br>- 131: Salicornia, high critical bed erosion shear stress (Figure S8)<br>- 132: No vegetation, low critical bed erosion shear stress (Figure S8)<br>- 133: No vegetation, high critical bed erosion shear stress (Figure S8)<br>- 134: Spartina, low Partheniades constant (Figure S8)<br>- 143: Spartina, high Partheniades constant (Figure S8)<br>- 136: Salicornia, low Partheniades constant (Figure S8)<br>- 144: Salicornia, high Partheniades constant (Figure S8)<br>- 138: No vegetation, low Partheniades constant (Figure S8)<br>- 145: No vegetation, high Partheniades constant (Figure S8)<br>- 150: Spartina, low sediment dry bulk density (Figure S8)<br>- 151: Spartina, high sediment dry bulk density (Figure S8)<br>- 152: Salicornia, low sediment dry bulk density (Figure S8)<br>- 153: Salicornia, high sediment dry bulk density (Figure S8)<br>- 154: No vegetation, low sediment dry bulk density (Figure S8)<br>- 155: No vegetation, high sediment dry bulk density (Figure S8)<br>- 76: Spartina, replicate #1 (Figures S9-S10)<br>- 77: Spartina, replicate #2 (Figures S9-S10)<br>- 78: Spartina, replicate #3 (Figures S9-S10)<br>- 88: Spartina, replicate #4 (Figures S9-S10)<br>- 80: Salicornia, replicate #1 (Figures S9-S10)<br>- 81: Salicornia, replicate #2 (Figures S9-S10)<br>- 82: Salicornia, replicate #3 (Figures S9-S10)<br>- 89: Salicornia, replicate #4 (Figures S9-S10)<br>- 85: No vegetation, replicate #1 (Figures S9-S10)<br>- 86: No vegetation, replicate #2 (Figures S9-S10)<br>- 87: No vegetation, replicate #3 (Figures S9-S10)<br>- 90: No vegetation, replicate #4 (Figures S9-S10)</p> <p> </p>
Dense vegetation hinders sediment transport towards saltmarsh interiors - Supporting data and source code (Part V: Figures)
<p>This is Part V of the supporting data and source code for the paper entitled "Dense vegetation hinders sediment transport towards saltmarsh interiors", submitted to <em>Limnology and Oceanography Letters.</em> It contains all input and output files to generate the figures of the paper.</p> <p>To be able to run the scripts as is, the path (at the beginning of each script) to the following folders must be updated:</p> <p>Runs (includes all model run folders from Part II and Part III)<br>Post (includes all post-processing folders from Part IV)</p>
Source data files for manuscript "Closed Magnetic Topology in the Venusian Magnetotail and Ion Escape at Venus"
<p>The zip file contains source data files for all figures in the manuscript "Closed Magnetic Topology in the Venusian Magnetotail and Ion Escape at Venus" published in Nature Communications. DOI: 10.1038/s41467-024-50480-0.</p>
FCH and FS Datasets for the paper "Integrating Multi-Source Remote Sensing Data for Mapping Boreal Forest Canopy Height and Species in interior Alaska in Support of Radar Modeling"
<p>This dataset provides forest canopy height and forest species in Delta Junction, interior Alaska in 2017. This dataset was produced based on the multi-source remote sensing datasets (AirMOSS, UAVSAR, Sentinel-1, Sentinel-2, topography), using a XGBoost approach.</p>
Data and source code for: Recent adaptation in a threatened salmonid revealed by museum genomics
<p>Steelhead/rainbow trout (Oncorhynchus mykiss) is an imperiled salmonid with two main life history strategies: migrate to the ocean or remain in freshwater. Domesticated hatchery forms of this species have been stocked into almost all California waterbodies, possibly resulting in introgression into natural populations and altered population structure. </p> <p>We compared whole-genome sequence data from contemporary populations against a set of museum population samples of steelhead from the same locations that were collected prior to most hatchery stocking. </p> <p>We observed minimal introgression and few steelhead-hatchery trout hybrids despite a century of extensive stocking. Our historical data show signals of introgression with a sister species and indications of an early hatchery facility. Finally, we found that migration-associated haplotypes have become less frequent over time, a likely adaptation to decreased opportunities for migration. Since contemporary migration-associated haplotype frequencies have been used to guide species management, we consider this to be a rare example of shifting baseline syndrome that has been validated with historical data. </p> <p>We suggest cautious optimism that a century of hatchery stocking has had minimal impact on California steelhead population genetic structure, but we note that continued shifts in life history may lead to further declines in the ocean-going form of the species. </p>
Source Data for Supplementary Information of "Expanding the substrate scope of PylRS enzymes to include non-⍺-amino acids in vitro and in vivo"
<p>The attached excel file contains the source data for LC-MS traces shown in the Supplementary Information of the paper "Expanding the substrate scope of PylRS enzymes to include non-⍺-amino acids in vitro and in vivo." Each graph is contained in a tab and labeled with the Supplementary Figure number and panel with which it is associated.</p>
Source data from Weiner et al 2024
<p>This repository provides the processed data necessay to produce the figures from the Weiner, <em>et al</em>., Nature Communications paper entitled "Inferring replication timing and proliferation dynamics from single-cell DNA sequencing data".</p> <p>The data is organized into directories based on figure numbers of the paper. For instance, the directory <span><code>source_data/fig2_S3/</code> contains all data pertaining to main figure 2 and supplementary figure 3. Within each directory, subdirectories are organized by figure panel, meaning that the file </span><span><code>source_data/fig2_S3/2fg_S3g/cell_metrics.tsv</code> only pertains to data which appears in Fig 2f, Fig 2g, and Fig S3g. </span></p> <p><span>Certain figure panels incorporate data from multiple samples. Sample ID subdirectories are used for these panels. For example, Fig S8a-c use data at the following paths </span><code><span>source_data/fig5_S8/S8abc/{sample_id}/s_phase_bafs.csv.gz</span></code></p> <p>All source code for upstream data preprocessing and downstream plotting can be found at the corresponding github repository for this manuscript: https://github.com/shahcompbio/scdna_replication_paper.<br><br>In addition to source data, we are also including two additional files which contain series of sample-specific heatmaps. Below are the descriptions of each file.</p> <p>Additional File 1: HMMcopy and SIGNALS heatmaps of high-quality G1/2-phase cells across all samples. Matrices of somatic copy number state called by HMMcopy (Ha, et al, 2012) (left) and allelic imbalance state called by SIGNALS (Funnell, et al, 2022) (right) for all cells in a sample. Only high-quality G1/2-phase cells are included in this analysis. The rows are sorted the same in both heatmaps to preserve mapping of cell IDs. Clone IDs for all cells are shown using the colorbar to the left of both heatmaps. Each page represents a unique sample in the metacohort of breast and ovarian cell lines and PDXs (Fig 4a). The sample ID and number of high quality G1/2-phase cells (i.e. SIGNALS cells) are shown at the top of each page. </p> <p>Additional File 2: PERT input and output matrices for gastric cancer cell lines at 500kb and 20kb resolution. Matrices of PERT input (left: reads per million and HMMcopy states) and PERT output (right: PERT somatic copy number and replication states) for the three gastric cancer cell lines sequenced with 10X Chromium single-cell DNA (Andor, et al, 2020). The top heatmaps on each page contains the cells predicted to be in S-phase by PERT and the bottom heatmaps contains the cells predicted to be in G1/2-phase by PERT. The rows are sorted the same in all four columns going from right to left to preserve mapping of cell IDs. Clone IDs for all cells are shown using the colorbar to the left of all four heatmaps. The first three pages represent PERT runs on each cell line at 500kb resolution. The final two pages represent PERT runs for two of the three cell lines at 20kb resolution. We did not run PERT at 20kb resolution for the SNU-668 cell line as there were too few S-phase cells to infer the RT profile at 20kb higher resolution.<br><br>For further information please reach out to Adam Weiner (weinera2@mskcc.org)</p>
Data and R scripts for the paper "New Evidence for a Directed Forgetting Effect in Source Memory and a Role of Source Feature Intrinsicality in the Item-Method"
<p>Data and R scripts from Experiment 1 and 2 of the paper "New Evidence for a Directed Forgetting Effect in Source Memory and a Role of Source Feature Intrinsicality in the Item-Method".</p>
Data archive: Niche overlap between a cold-water coral and an associated sponge for isotopically-enriched particulate food sources
<p>Data belonging to the paper: </p> <p>Dick van Oevelen, Christina E. Mueller, Tomas Lundälv, Fleur C. van Duyl, Jasper M. de Goeij, Jack J. Middelburg<span> </span>(In press) <strong>Niche overlap between a cold-water coral and an associated sponge for isotopically-enriched particulate food sources</strong>. PLOS ONE</p>
Research data, sources and documents for thesis on Exploring Complexity Metrics for Artifact-Centric Business Process Models
<p>Research data, sources and documents for thesis on Exploring Complexity Metrics for Artifact-Centric Business Process Models This repository contains the supplemental material for the <a href="https://pqdtopen.proquest.com/pubnum/10759956.html">thesis "Exploring Complexity Metrics for Artifact-Centric Business Process Models" by Marin, Mike A., Ph.D., University of South Africa (South Africa), 2017.</a></p>
Images and supporting data for high-resolution μCT of a mouse embryo using a compact laser-driven x-ray betatron source
<p>A high resolution x-ray CT scan of an embryonic mouse sample was performed with the betatron x-ray source produced by a laser wakefield accelerator. This data deposition includes all of the raw images of the mouse sample, information regarding their indexing, featured slices of the tomogram and some further raw data regarding the x-ray source characterisation.</p>
A dataset of regional operational programmes (ROP) and rural development program (PROW) expenditures and socio-economic features in 2007-2013, Poland (source: Bank of Local Data)
<p>Dataset prepared on the bases of Polish Central Statistical Office (Statistics Poland) Bank of Local Data system https://bdl.stat.gov.pl/BDL/dane/podgrup/tablica [access: 1.07.2018]. The data set the expenditure of funds for individual priority axes in the programmes of both policies in the 2007-2013 programming period and the change in socio-economic features at the local (<em>poviat</em>, NUTS4) level. The Pearson correlation coefficients are used to assess the relationship between the level of expenditure for RDP and ROP <em>per capita</em> and selected indicators describing the level of economic, social and demographic development of local government units. The results of the analysis (the article <strong>Regional approach to rural development? A case of regional and rural programs 2007-2015 in Poland) </strong>will consist of tables, texts and of numerical data. Article with data is available here: OI: 10.5604/01.3001.0012.2934 GICID: 01.3001.0012.2934 Available language versions: en. <strong>Issue: </strong>Annals PAAAE 2018; XX (4): 22-28, https://rnseria.com/resources/html/article/details?id=176837</p>
Single-crystal X-ray diffractometry data for a sample of NiCl₂-dppe collected on beamline I19-2 at Diamond Light Source
<p>Single-crystal X-ray diffractometry data for a sample of [1,2-Bis(diphenylphosphino)ethane]dichloronickel(II) (NiCl<sub>2</sub>-dppe, [(C<sub>6</sub>H<sub>5</sub>)<sub>2</sub>PCH<sub>2</sub>CH<sub>2</sub>P(C<sub>6</sub>H<sub>5</sub>)<sub>2</sub>]NiCl<sub>2</sub>).</p> <p>Data collected at Diamond Light Source I19-2 on 2015-05-18, publicly available for users to test data reduction routines. Data are known to produce good merging statistics and final refinements.</p> <p>The sample was prepared as follows:<br> Nickel chloride (II) hexahydrate (1 g, 2 mmol) was heated under vacuum to produce anhydrous nickel chloride (II) with a visible colour change from green to yellow. The resulting solid was taken up in ethanol (5 ml) and added to 1,2-bis(dimethylphosphine)ethane (dppe) (0.837 g, 2 mmol) in ethanol (10 ml). The solution was refluxed for 3 hour after which the solvent was evaporated. The small red crystals were purified by recrystallisation in acetone (70% yield).</p> <p>The sample was held at an approximate temperature of 150 K and the illuminating beam had a wavelength of 0.68890 Å (17.997 keV). The detector was held at 2θ = 25° throughout.</p> <p>Inventory of data:<br> <strong>010_Ni_dppe_Cl_2_150K01</strong> — 130° ω scan, 0.4° images, 0.4s per image, 325 images; κ = 45°, φ = 160°.<br> <strong>010_Ni_dppe_Cl_2_150K02</strong> — 130° ω scan, 0.4° images, 0.4s per image, 325 images; κ = 45°, φ = 40°.<br> <strong>010_Ni_dppe_Cl_2_150K03</strong> — 130° ω scan, 0.4° images, 0.4s per image, 325 images; κ = 45°, φ = -80°.<br> <strong>010_Ni_dppe_Cl_2_150K04</strong> — 198° ω scan, 0.4° images, 0.4s per image, 495 images; κ = 0°, φ = -80°.</p>
In-situ data recorded from thaumatin crystals on Diamond Light Source VMXi
<p>Data collected as part of routine beamline commissioning, from samples of thaumatin grown in 0.1 M Sodium citrate, 0.75M Sodium / Potassium tartrate. Data were collected in unattended mode with positions identified <em>via</em> SynchWeb from photographs taken with a Formulatrix imaging system (picture included.) Each data set consists of 200 images taken with an Eiger 2X 4M detector at a distance of 186mm, with exposure time of 2 ms, wavelength 0.97950 angstroms and 2% transmission with a DMM beam. Automated processing combined data using the xia2 multiplex tool to give a sufficiently complete data set for structure solution and refinement.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.