Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
410
datasets available to search
ShareScore release 0.9.0
Dataset results
410 results for “Data Repositories”
Data repository of the paper "The First Terrestrial Electron Beam Observed by The Atmosphere-Space Interactions Monitor" by D. Sarria et al.
<p>Data repository / Supporting information of the paper "The First Terrestrial Electron Beam Observed by The Atmosphere-Space Interactions Monitor" (2019) by D. Sarria et al.</p> <p>Access to the article: <a href="https://doi.org/10.1029/2019JA027071">https://doi.org/10.1029/2019JA027071</a></p>
Data repository - Spatial reconstruction of single enterocytes uncovers broad zonation along the intestinal villus axis
<p>Data associated with the manuscript entitled "Spatial reconstruction of single enterocytes uncovers broad zonation along the intestinal villus axis".</p> <p>Files:</p> <p>table_A_LCM_TPM_values.tsv: Gene expression levels of microdissected villus quintiles. First column is the ensemble gene id. Next 15 columns are the raw Kallisto TPM values for villus segments 1 (bottom) to 5 (top) for three different mice (a to c). Additional columns include the external gene name, description, and gene biotype.</p> <p>table_B_scRNAseq_UMI_counts.tsv: Raw UMI counts of cells that were utilized in this study. Analysis is based on raw data from the NCBI GEO datasets GSM2644349 and GSM2644350. Each of the columns represents a single cell, column headers are the corresponding cell barcodes and enable retrieval of tSNE coordinates from table_C_scRNAseq_tsne_coordinates_zones.tsv. Values represent raw UMI counts.</p> <p>table_C_scRNAseq_tsne_coordinates_zones.tsv: tSNE coordinates and reconstructed zones of cells that were utilized in this study. Tab separated text file. Analysis is based on raw data from the NCBI GEO datasets GSM2644349 and GSM2644350. Columns: cell_id: cell barcode, corresponds to column header of Table S2. Seurat tSNE coordinate 1 and tSNE coordinate 2. Last column is the inferred zone (Crypt, V1..V6).</p> <p>table_D_zonation_reconstruction.tsv: Zonation table of reconstructed scRNAseq data. Tab separated text file. Columns: Gene names: gene id, mean expression in each of the crypt zone and 6 villus zones, standard error of the means in the Crypt zone and 6 villus zones, p-value and q-value for the zonation profiles.</p> <p>raw_data.zip: The raw and intermediary data for runnning the scripts in <a href="https://github.com/aemoor/Code_spatial_reconstruction_enterocytes">https://github.com/aemoor/Code_spatial_reconstruction_enterocytes</a></p>
Data Usage Metrics at Repositories: A Survey
<p>Results of a survey undertaken by the Research Data Alliance Data (RDA) Usage Metrics Working Group during February and March 2019 and presented at the 13th RDA Plenary Meeting in Philadelphia on 3 April 2019.</p>
Data repository for manuscript "A new approach to Health Benefits Package design: an application of the Thanzi La Onse model in Malawi"
<p>Dataset to accompany the publication <em>“A new approach to Health Benefits Package design: an application of the Thanzi La Onse model in Malawi”</em> by Margherita Molaro, Sakshi Mohan, Bingling She, Martin Chalkley, Tim Colbourn, Joseph H. Collins, Emilia Connolly, Matthew M. Graham, Eva Janoušková, Ines Li Lin, Gerald Manthalu, Emmanuel Mnjowe, Dominic Nkhoma, Pakwanja D. Twea, Andrew N. Phillips, Paul Revill, Asif U. Tamuri, Joseph Mfutso-Bengo, Tara Mangal, and Timothy B. Hallett.</p> <p>The Thanzi La Onse (TLO) model used to produce this data is open source and available for review and usage at<a href="https://github.com/UCL/TLOmodel"> https://github.com/UCL/TLOmodel</a>. In particular, the outputs analysed in this study can be reproduced from model tag "Molaro_et_al_2024_HBP_design" (accessible at https://github.com/UCL/TLOmodel/tags) using the scenario file src/scripts/healthsystem/impact_of_policy/scenario_impact_of_policy.py. All analysis scripts used to generate the plots in the manuscript are located in the same directory and have filenames beginning with "analysis_impact_of_policy_".</p> <p>This repository contains post-processed simulation outputs, which were generated using the script src/scripts/healthsystem/impact_of_policy/analysis_extract_data.py (available from the same tag). The data included have the following structure:</p> <p>"Draw": Represents a specific prioritisation-policy, identified by the acronyms listed in Table 1 of the publication.</p> <p>"Run": Represents a single simulation instance of a draw. Each draw was simulated 10 times, each with independent random sampling, resulting in 10 "runs" per draw.</p> <p>The data files included in this repository are:</p> <p><strong>DALYS_by_cause_with_time.csv</strong>: DALYs (as defined in the publication) incurred on a given year due to each of the causes of DALYs considered.</p> <p><strong>HSIs_requested_by_type_and_facility_level_with_time.csv</strong>: total number of requested HSIs on a given year, broken down by HSI type and the facility level at which they were requested.</p> <p><strong>HSIs_delivered_by_type_and_facility_level_with_time.csv</strong>:total number of HSIs delivered on a given year broken down by HSI type and the facility level at which they were delivered.</p> <p><strong>Population_with_time.csv</strong>:total population size on a given year. </p> <p> </p> <p> </p>
Data and code repository for the research "Assessing the use of Airborne Electromagnetic Data for nitrate vulnerability assessment in the Central Valley, California"
<p>Data and code for "Assessing the use of Airborne Electromagnetic Data for nitrate vulnerability assessment in the Central Valley, California". This repository contains the analysis and post-processing code for generating results and figures used in the manuscript.</p>
Data and Analysis Files Repository: Repurposing Large-Format Microarrays for Scalable Spatial Transcriptomics
<p>Data and Analysis Files from "Repurposing Large-Format Microarrays for Scalable Spatial Transcriptomics"</p> <p>ArraySeq_Method.zip contains the following folder and contents:</p> <ul> <li>STARSolo: All code and count matrix output from fastq spatial barcode demultiplexing. </li> <li>Images: All resolution-downsampled H&E image scans from analyzed tissues</li> <li>Space_Ranger: All 10x Space Ranger output from Visium datasets generated in the paper. </li> <li>Analysis: All scripts for analyzing and plotting Array-seq and Visium datasets generated in this paper. Also contains output h5ad files. </li> </ul> <p>ArraySeq_Barcode_generation_n12.rmd: The script used to generate the Array-seq probes with 12-mer spatial barcodes. </p>
Parameter estimation data repository - "Spatial discordances between mRNAs and proteins in the intestinal epithelium"
<p>The repository contains data associated with the estimation of protein translation and decay rates in the manuscript "Spatial discordances between mRNAs and proteins in the intestinal epithelium". Specifically, it includes MCMC-chains approximating the posterior parameter distribution and figures showing the model's fit to the data for each gene as well as a summary table of all fit results for two different models, the constant translation-rate model ("constant_rate_model") and the declining translation-rate model ("declining_rate_model") as explained in the manuscript.</p> <p>Code associated with the parameter estimation is available at https://github.com/LiBuchauer/spatial_MP_discordances .</p>
Repository Analytics and Metrics Portal (RAMP) 2018 data
<p>The Repository Analytics and Metrics Portal (RAMP) is a web service that aggregates use and performance use data of institutional repositories. The data are a subset of data from RAMP, the Repository Analytics and Metrics Portal (<a href="http://ramp.montana.edu/">http://rampanalytics.org</a>), consisting of data from all participating repositories for the calendar year 2018. For a description of the data collection, processing, and output methods, please see the "methods" section below. Note that the RAMP data model changed in August, 2018 and two sets of documentation are provided to describe data collection and processing before and after the change.</p>
Repository Analytics and Metrics Portal (RAMP) 2017 data
<p>The Repository Analytics and Metrics Portal (RAMP) is a web service that aggregates use and performance use data of institutional repositories. The data are a subset of data from RAMP, the Repository Analytics and Metrics Portal (<a href="http://ramp.montana.edu/">http://rampanalytics.org</a>), consisting of data from all participating repositories for the calendar year 2017. For a description of the data collection, processing, and output methods, please see the "methods" section below.</p>
Panel: Dataverse Community and CoreTrustSeal: Certifying generalist data repositories
<p>CoreTrustSeal Certification is an important tool that helps researchers and practitioners evaluate the trustworthiness of a dataset and a data repository, yet its certification model can be challenging for generalist repositories to meet. This panel will discuss strategies for meeting certification standards for Trustworthy Data Repositories (TDR) across a variety of repositories with different methods of appraisal, curation, preservation, and organization. The audience will be encouraged to add to the discussion and invited to provide feedback on the concepts presented by the panelists.</p>
The EGFRvIII Transcriptome in glioblastoma - public data repository
<p>Compiled dataset from a large omics EGFRvIII study.</p>
Model data repository of "The role of sediment accretion and buoyancy on subduction dynamics and geometry"
<p>This dataset contains the code and data used in Brizzi et al. (2021): The role of sediment accretion and buoyancy on subduction dynamics and geometry</p>
ESPRESSO: Robust discovery and quantification of transcript isoforms from error-prone long-read RNA-seq data (repository for simulated ONT RNA-seq data)
<p>Simulated ONT direct RNA and 1D cDNA sequencing data of varying sequencing depths (0.5 million, 1 million, 3 million, and 5 million simulated reads) used for benchmark evaluations of transcript discovery and quantification in our paper "ESPRESSO: Robust discovery and quantification of transcript isoforms from error-prone long-read RNA-seq data". All details can be found in the <strong>Materials and Methods</strong> section of the paper. </p> <p><em>HEK293T_DirectRNA.transcriptome_quantification.tsv</em> and <em>HEK293T_DirectRNA.transcriptome_quantification.tsv </em>are tab-separated files containing estimated raw read counts and normalized abundance values (in TPM) of transcripts annotated in GENCODE v34lift37. Transcript quantification was done using NanoSim (version 3.1.0). </p> <p><em>HEK293T_DirectRNA.NanoSim_500k.fastq.gz</em>,<em> </em><em>HEK293T_DirectRNA.NanoSim_1M.fastq.gz</em>, <em>HEK293T_DirectRNA.NanoSim_3M.fastq.gz</em>, and<em> HEK293T_DirectRNA.NanoSim_5M.fastq.gz </em>are gzip compressed FASTQ files containing 0.5 million, 1 million, 3 million, and 5 million simulated ONT direct RNA sequencing reads respectively. </p> <p><em>HEK293T_1DcDNA.NanoSim_500k.fastq.gz</em>,<em> HEK293T_1DcDNA.NanoSim_1M.fastq.gz</em>, <em>HEK293T_1DcDNA.NanoSim_3M.fastq.gz</em>, and<em> HEK293T_1DcDNA.NanoSim_5M.fastq.gz </em>are gzip compressed FASTQ files containing 0.5 million, 1 million, 3 million, and 5 million simulated ONT 1D cDNA sequencing reads respectively. </p>
Data repository of FDTD GPR antenna optimization by Sam Stadler
<p>This is the data repository for the article by Sam Stadler and Jan Igel by the name "Developing realistic FDTD GPR antenna surrogates via full-waveform inversion by means of particle swarm optimization". In this repository, all the measurement data, simulated data, and high-resolution figures are stored for further use.</p>
Data repository for "Phonon-mediated room-temperature quantum Hall transport in graphene"
<p>This is the data presented in the manuscript "Phonon-mediated room-temperature quantum Hall transport in graphene", Nat Commun 14, 318 (2023). https://doi.org/10.1038/s41467-023-35986-3</p>
WorldCereal open global harmonized reference data repository (CC-BY-NC licensed data sets)
<p>Within the <strong>ESA funded</strong> WorldCereal project we have built an open harmonized reference data repository at global extent for model training or product validation in support of land cover and crop type mapping. Data from 2017 onwards were collected from many different sources and then harmonized, annotated and evaluated. These steps are explained in the harmonization protocol (10.5281/zenodo.7584463). This protocol also clarifies the naming convention of the shape files and the WorldCereal attributes (LC, CT, IRR, valtime and sampleID) that were added to the original data sets.</p> <p>This publication includes those harmonized data sets of which the original data set was published under the CC-BY-NC license or a license similar to CC-BY-NC. See document "_In-situ-data-World-Cereal - license - CC-BY-NC.pdf" for an overview of the original data sets. Currently this publication only includes a few small data sets for Tanzania originating from a disease monitoring program of the International Maize and Wheat Improvement Center (CIMMYT). CIMMYT made more data available for countries like Kenya, Ethiopia, Rwanda and, Malawi. However due project contraints these data sets were not yet harmonized.</p>
Raw data for the Github repository "Collection of scripts to download & process hydropower generation data in Argentina, Bolivia, Brazil, Uruguay"
<p>This is the raw data downloaded from the websites of the following system operators:</p> <p> - CNDC, Comité Nacional de Despacho de Carga (Bolivia): https://www.cndc.bo/home/index.php<br> - ONS, Operador Nacional do Sistema Elétrico (Brazil): https://www.ons.org.br/<br> - UTE, Usinas y Trasmisiones Eléctricas (Uruguay): https://www.ute.com.uy/<br> - CAMMESA, Compañía Administradora del Mercado Mayorista Eléctrico Sociedad Anónima (Argentina): https://cammesaweb.cammesa.com</p> <p>Hydropower generation data is extracted using the R scripts available here: https://github.com/matteodefelice/hydro-sam</p>
How to choose a research data repository software? Experience report. Table of requirements.
<p>In the age of digital transformation, scientific and social interest for data and data products is constantly on the rise. The quantity as well as the variety of digital research data is increasing significantly. This raises the question about the governance of this data. For example, how to store the data so that it is presented transparently, freely accessible and subsequently available for re-use in the context of good scientific practice. Research data repositories provide solutions to these issues.</p> <p>Considering the variety of repository software, it is sometimes difficult to identify a fitting solution for a specific use case. For this purpose a detailed analysis of existing software is needed. Presented table of requirements can serve as a starting point and decision-making guide for choosing the most suitable for your purposes repository software. This table is dealing as a supplementary material for the paper "How to choose a research data repository software? Experience report." (persistent identifier to the paper will be added as soon as paper is published).</p>
Glass-Like Random Catalogues for Two-Point Estimates on the Light Cone (Data Repository)
<p>This is the data repository for the article Glass-Like Random Catalogues for Two-Point Estimates on the Light Cone.</p> <p>arxiv link: https://arxiv.org/abs/2304.02040</p> <p>DOI:</p> <p> </p> <p>The file grlic_data.tar.gz contains three directories: /data, /correlations, and /randoms.</p> <p>Inside the /data folder, the data catalogues used in the article are stored: in "cat_part_1162568" the three columns correspond to the redshift, the cosine of the polar angle, and the azimuthal angle, respectively. The first three columns of the "cat_high_1164853" and "cat_low_1164853" catalogues correspond to the same properties for the high-mass and low-mass halos, respectively. In addition to that, these files also contain additional information about the halos: the number of particles (4th column), their M_200b mass (5th column) and their parent ID (6th colum), which is -1 if the halo is not a subhalo.</p> <p>Inside the /randoms folder, the random catalogues for each of the data catalogues in /data are stored. Files beginning with "part" refer to randoms based on the particle catalogue, in a similar fashion the files beginning with "high" correspond to randoms based on the high-mass halo catalogue and those beginning with "low" refer to the randoms based on the low-mass halo catalogue. Files with "..._glass<x>..." correspond to the glass-like random catalogues, and files with "..._rand<x>..." correspond to the Poisson-sampled randoms, where <x> is the value of <span class="math-tex">\(\alpha\)</span> used. <span class="math-tex">\(\alpha\)</span> is the factor by which the number of objects in the data catalogue is multiplied to get the number of objects in the random catalogues, <span class="math-tex">\(N_R = \alpha N_D\)</span>. For the glass-like randoms based on the high-mass halo catalogue, the additional suffix, "..._deltagrid<y>...", refers to the number of grid-cells used in the Zeldovich approximation, where <y> is the number of grid cells, and "..._Niter<z>..." refers to the number of Zeldovich iterations perfomed, which is given by <z>. Similarly to the data catalogues, the three columns represent the redshift, cosine of the polar angle, and the azimuthal angle of each object in the catalogue.</p> <p>Inside the /correlations folder, there are two subdirectories: /correlations/full and /correlations/multipoles. The /correlations/full directory contains the raw output from CUTE for each correlated pair of catalogues. For example, the subdirectory /correlations/full/low_low contains the outputs for the low-mass halo autocorrelation, or the subdirectory /correlations/full/high_part contains the outputs for the cross-correlation between the high-mass halo catalogue and the particle catalogue. The naming convention of these files is "<type><alpha>_<catalogues>_<i>", where <type> can either be "rand" or "glass", for either the Poisson-sampled or glass-like random catalogues, <alpha> is the factor <span class="math-tex">\(\alpha\)</span> already introduced above, <catalogues> again describes which two data catalogues have been cross- (or auto-) correlated, e.g. "highlow" refers to a cross-correlation between the high-mass halo catalogue and the low-mass halo catalogue, and finally <i> is a number between 0 and 19 for the 20 independent measurements of the correlations. For the high-mass halo catalogue autocorrelations involving glass-like randoms, the additional suffix, "..._deltagrid<y>...", refers to the number of grid-cells used in the Zeldovich approximation, where <y> is the number of grid cells, and "..._Niter<z>..." refers to the number of Zeldovich iterations perfomed, which is given by <z>. The format of the files is the standard CUTE format for the 3D correlation using binning in (mu,r), i.e. see the readme of https://github.com/damonge/CUTE/tree/master/CUTE, section 5. It reads:</p> <pre>For the 3-D correlation functions the output file has 7 columns with x1 x2 xi(x1,x2) D1D2(x1,x2) D1R2(x1,x2) R1D2(x1,x2) R1R2(x1,x2) where (x1,x2) is either (pi,sigma) or (mu,r). </pre> <p>The /correlations/multipoles subdirectory contains the estimated mean multipoles and their variance derived from the 20 individual correlation measurements for each data catalogue and type of random catalogue. Similarly to the /correlations/full subdirectory, it contains a separate directory for each pair of data catalogue for which the correlation was estimated. The files are named according to "<multipole>_<type><alpha>_<catalogues>_mean_var", where <multipole> is either l0 for the monopole, l1 for the dipole, or l2 for the quadrupole. The <type>, <alpha> and <catalogues> are identical to what was described above for the /correlations/full subdirectory. Again, for the high-mass halo catalogue autocorrelations involving glass-like randoms, the additional suffix, "..._deltagrid<y>...", refers to the number of grid-cells used in the Zeldovich approximation, where <y> is the number of grid cells, and "..._Niter<z>..." refers to the number of Zeldovich iterations perfomed, which is given by <z>. The columns in each of these files are the comoving separation d, the mean multipole <span class="math-tex">\(\xi_l\)</span>, and its variance <span class="math-tex">\(\sigma^2\)</span>of each bin.</p>
D-DUST Analysis Ready Data Repository
<p>Analysis-ready data repository (<em>D22_ARD_repository_v1_24022022.zip</em>) and Data Management Plan (<em>D-DUST_DMP_v1_22122022.pdf</em>) developed within the D-DUST Project (Data-driven moDelling of particUlate with Satellite Technology aid)</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.