Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,481
datasets available to search
ShareScore release 0.9.0
Dataset results
1,481 results for “data processing”
Data from: Cornerstones are the key stones: Using interpretable machine learning to probe the clogging process in 2D granular hoppers
Open the record for dataset details and reuse information.
LipidQuant 1.0: Automated data processing in lipid class separation - mass spectrometry quantitative workflows
Open the record for dataset details and reuse information.
Data from: DCDC2 READ1 regulatory element: how temporal processing differences may shape language
Open the record for dataset details and reuse information.
Data for: Image processing tools for petabyte-scale light sheet microscopy data (Part 2/2)
Open the record for dataset details and reuse information.
Data for: Image processing tools for petabyte-scale light sheet microscopy data (Part 1/2)
Open the record for dataset details and reuse information.
Data from: Inferring riverscape dispersal processes from fish biodiversity patterns
Open the record for dataset details and reuse information.
Data From: Estimation of genome-wide coupling in rattlesnake hybrids provides insight into the process of speciation and its progress
Open the record for dataset details and reuse information.
Data from: Fractal triads efficiently sample ecological diversity and processes across spatial scales
Open the record for dataset details and reuse information.
Scripts and data sets associated with: On testing homogeneity of the evolutionary process using alignments of homologous sequences
Open the record for dataset details and reuse information.
Processed single cell data from CODEX multiplexed imaging of the human intestine
Open the record for dataset details and reuse information.
Data for: Historical and contemporary processes drive global phylogenetic structure across geographical scales: Insights from bat communities
Open the record for dataset details and reuse information.
Wayqecha Amazon cloud curtain ecosystem experiment: Climate data and R processing code
Open the record for dataset details and reuse information.
Data for: Chronic exposure to odors at naturally occurring concentrations triggers limited plasticity in early stages of Drosophila olfactory processing
Open the record for dataset details and reuse information.
Data from: Fruit resources shape sexual selection processes in a lek mating system
Open the record for dataset details and reuse information.
Data for: Climate change is poised to alter mountain stream ecosystem processes via organismal phenological shifts
Open the record for dataset details and reuse information.
Northern elephant seal tracking and diving – processed data
Open the record for dataset details and reuse information.
Assembled file of one minute averages for high resolution surface meteorological (Met) and sea water intake (SWI) data from continuous underway measurements from CCE LTER process cruises in the CCE region, 2006 - 2019.
As the research vessel is underway for the duration of a CCE Process Cruise (since 2006, ongoing), 30 parameters are continuously measured regarding the oceanographic surface and atmospheric and navigational environment of the vessel, along the ship's trackline in the CCE region.
Plant species percent cover data: Biodiversity II: Effects of Plant Biodiversity on Population and Ecosystem Processes
Biodiversity II (E120) is designed to determine how the number of plant species affects the dynamics of ecological processes at the population, community, and ecosystem levels. By experimentally manipulating the number of species and the kinds of species, the amount of plant growth and the change from year to year, that result can be examined. Plots are large (9m x 9m actively maintained) and well-replicated, allowing responses of plant pathogens, insect herbivores, seed predators, soil parameters, invasive plant species and other variables to also be studied. Plots were seeded in May 1994 to have 1, 2, 4, 8, or 16 species, with roughly 30 replicates of each diversity level. The species composition of each plot was chosen by random draw from a pool of 18 grassland perennials that included four warm-season (C4) grasses, four cool-season (C3) grasses, four legumes, four non-legume forbs, and two woody species. All species occur in monoculture allowing comparison of responses of each species in monoculture to combinations of these same species. The experiment was established in 1994 by the lead investigators David Tilman, Peter Reich, Johannes Knops, and David Wedin. Experiment 120 is similar to Experiment 123, but it uses larger plots to provide a large capacity for long-term subexperiments.
LCI data for materials and processes comparing energy and water use of aqueous and gas-based metalworking fluids
<p>Datasets containing life cycle inventories for materials and processes, and results of the analysis in the article titled, "Comparing energy and water use of aqueous and gas-based metalworking fluids" published in the <em>Journal of Industrial Ecology.</em></p>
Pre-processed B cell receptor repertoire sequencing data from BioProject PRJNA527941
<p><strong>Data Processing</strong></p> <p> </p> <p>Samples were demultiplexed via their Illumina indices, and processed using the Immcantation toolkit(1,2). Raw fastq files were filtered based on a quality score threshold of 20. Paired reads were joined if they had a minimum length of 10 nt, maximum error rate of 0.3 and a significance threshold of 0.0001. Reads with identical UMI were collapsed to a consensus sequence. Reads with identical full-length sequence and identical constant primer but differing UMI were further collapsed. Sequences were then submitted to IgBlast (3) for VDJ assignment and sequence annotation. Constant region sequences were mapped to germline using Stampy(4). The number and type of V gene mutations was calculated using the shazam R package.(2)</p> <p> </p> <p><strong>software_versions</strong> pRESTO:0.5.3,Change-O:0.3.4,IgBlast 1.6.1, stampy1.0.21. shazam0.1.8</p> <p><strong>quality_thresholds</strong> FilterSeq.py pRESTO Q>20</p> <p><strong>paired_reads_assembly</strong> AssemblePairs.py pRESTO minlen 10 maxerror 0.3 alpha 0.0001</p> <p><strong>primer_match_cutoffs</strong> MaskPrimers.py pRESTO C primer & V primer maxerror 0.2</p> <p><strong>consensus_building</strong> BuildConsensus.py pRESTO maxerror 0.1 maxgap 0.5</p> <p><strong>collapsing_method</strong> CollapseSeq.py pRESTO</p> <p><strong>germline_database </strong>IMGT</p> <p> </p> <p><strong>Format</strong></p> <p> </p> <p>Processed sequences are provided in a tab delimited file format, including the following annotations:</p> <p> </p> <p><strong>C_CALL </strong>Isotype subclass</p> <p><strong>SEQUENCE_ID </strong>Sequence identifier</p> <p><strong>V_CALL </strong>V segment gene and allele</p> <p><strong>D_CALL </strong>D segment gene and allele</p> <p><strong>J_CALL </strong>J segment gene and allele</p> <p><strong>JUNCTION_LENGTH </strong>Junction length</p> <p><strong>CONSCOUNT </strong>Raw read count from which UMI consensus sequences were generated, summed over all UMIs for the given unique sequence.</p> <p><strong>DUPCOUNT </strong>UMI count for the given unique sequence</p> <p><strong>ISOTYPE </strong>Constant region primer (isotype)</p> <p><strong>MU_COUNT_CDR_R </strong>Number of replacement mutations in CDR region</p> <p><strong>MU_COUNT_CDR_S </strong>Number of silent mutations in CDR region</p> <p><strong>MU_COUNT_FWR_R </strong>Number of replacement mutations in FWR region</p> <p><strong>MU_COUNT_FWR_S </strong>Number of silent mutations in FWR region</p> <p><strong>MUT_TOTAL </strong>Total number of mutations in V gene </p> <p><strong>SEQUENCE_INPUT </strong>Full length sequence</p> <p><strong>SEQUENCE_IMGT </strong>Gapped IMGT sequence</p> <p><strong>V_GERM_START_VDJ </strong>position of the first nucleotide in ungapped V germline sequence alignment</p> <p><strong>JUNCTION </strong>Junction nucleotide sequence</p> <p><strong>GERMLINE_IMGT_D_MASK </strong>IMGT-gapped germline nucleotide sequence with ns masking the NP1-D-NP2 regions</p> <p><strong>Run </strong>ID of sequencing run</p> <p><strong>Sample_type </strong>The tissue sampled (e.g Peripheral Blood, bone marrow, ..)</p> <p><strong>Sex </strong>Sex of the Subject</p> <p><strong>Age </strong>Age of the subject</p> <p><strong>UNIQUE_ID </strong>Subject identifier </p> <p><strong>SAMPLE_ID </strong>Sample identifier, linking back to raw data</p> <p><strong>Subset </strong>Defined B cell subset </p> <p><strong>Repertoire </strong>Defined B cell repertoire (Naive, Memory IgM/IgD, IgA, IgG)</p> <p><strong>R_SCDR </strong>R/S ratio in CDR region</p> <p><strong>R_SFWR </strong>R/S ratio in FWR region</p> <p><strong>V_FAM </strong>V family gene</p> <p><strong>V_GENE </strong>V segment gene</p> <p><strong>D_GENE </strong>D segment gene</p> <p><strong>J_GENE </strong>J segment gene</p> <p><strong>Clust_Rank </strong>Cluster rank</p> <p><strong>Clust_REPRES </strong>Cluster representative</p> <p><strong>Clust_SIZE </strong>Cluster size</p> <p><strong>Clust_MAXFREQ </strong>Cluster maximum frequency</p> <p><strong>Clust_SHAREDNESS </strong>Cluster sharedness</p> <p><strong>CDR3_AA_GRAVY </strong>CDR3 hydrophobicity index</p> <p><strong>CDR3_AA_CHARGE </strong>CDR3 charge</p> <p><strong>CDRH3PDB </strong>CDRH3 PDB (Structure) code</p> <p><strong>H1Canon </strong>H1 Canonical class</p> <p><strong>H2Canon </strong>H2 Canonical class</p> <p><strong>H1_GERMLINE </strong>H1 Germline Canonical class</p> <p><strong>H2_GERMLINE </strong>H2 Germline Canonical class</p> <p> </p> <p><strong>References</strong></p> <p>1. Vander Heiden, J. A., G. Yaari, M. Uduman, J. N. H. Stern, K. C. O’Connor, D. A. Hafler, F. Vigneault, and S. H. Kleinstein. 2014. PRESTO: A toolkit for processing high-throughput sequencing raw reads of lymphocyte receptor repertoires. <em>Bioinformatics</em>30: 1930–1932.</p> <p>2. Gupta, N. T., J. A. Vander Heiden, M. Uduman, D. Gadala-Maria, G. Yaari, and S. H. Kleinstein. 2015. Change-O: A toolkit for analyzing large-scale B cell immunoglobulin repertoire sequencing data. <em>Bioinformatics</em>31: 3356–3358.</p> <p>3. Ye, J., N. Ma, T. L. Madden, and J. M. Ostell. 2013. IgBLAST: an immunoglobulin variable domain sequence analysis tool. <em>Nucleic Acids Res.</em>41.</p> <p>4. Lunter, G., and M. Goodson. 2011. Stampy: A statistical algorithm for sensitive and fast mapping of Illumina sequence reads. <em>Genome Res.</em>21: 936–939.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.