Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,481

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,481 results for “data processing”

Learn how ShareScore rates datasets ↗
dryad40/100

Data from: Cornerstones are the key stones: Using interpretable machine learning to probe the clogging process in 2D granular hoppers

Open the record for dataset details and reuse information.

publicJul 2025View details →
dryad40/100

LipidQuant 1.0: Automated data processing in lipid class separation - mass spectrometry quantitative workflows

Open the record for dataset details and reuse information.

publicJun 2021View details →
dryad40/100

Data from: DCDC2 READ1 regulatory element: how temporal processing differences may shape language

Open the record for dataset details and reuse information.

publicMay 2020View details →
dryad40/100

Data for: Image processing tools for petabyte-scale light sheet microscopy data (Part 2/2)

Open the record for dataset details and reuse information.

publicJul 2024View details →
dryad40/100

Data for: Image processing tools for petabyte-scale light sheet microscopy data (Part 1/2)

Open the record for dataset details and reuse information.

publicJul 2024View details →
dryad40/100

Data from: Inferring riverscape dispersal processes from fish biodiversity patterns

Open the record for dataset details and reuse information.

publicMar 2025View details →
dryad40/100

Data From: Estimation of genome-wide coupling in rattlesnake hybrids provides insight into the process of speciation and its progress

Open the record for dataset details and reuse information.

publicOct 2025View details →
dryad40/100

Data from: Fractal triads efficiently sample ecological diversity and processes across spatial scales

Open the record for dataset details and reuse information.

publicSep 2021View details →
dryad40/100

Scripts and data sets associated with: On testing homogeneity of the evolutionary process using alignments of homologous sequences

Open the record for dataset details and reuse information.

publicMay 2024View details →
dryad40/100

Processed single cell data from CODEX multiplexed imaging of the human intestine

Open the record for dataset details and reuse information.

publicSep 2023View details →
dryad40/100

Data for: Historical and contemporary processes drive global phylogenetic structure across geographical scales: Insights from bat communities

Open the record for dataset details and reuse information.

publicMar 2023View details →
dryad40/100

Wayqecha Amazon cloud curtain ecosystem experiment: Climate data and R processing code

Open the record for dataset details and reuse information.

publicDec 2024View details →
dryad40/100

Data for: Chronic exposure to odors at naturally occurring concentrations triggers limited plasticity in early stages of Drosophila olfactory processing

Open the record for dataset details and reuse information.

publicMay 2023View details →
dryad40/100

Data from: Fruit resources shape sexual selection processes in a lek mating system

Open the record for dataset details and reuse information.

publicOct 2024View details →
dryad40/100

Data for: Climate change is poised to alter mountain stream ecosystem processes via organismal phenological shifts

Open the record for dataset details and reuse information.

publicMar 2024View details →
dryad40/100

Northern elephant seal tracking and diving – processed data

Open the record for dataset details and reuse information.

publicMay 2025View details →
edi40/100

Assembled file of one minute averages for high resolution surface meteorological (Met) and sea water intake (SWI) data from continuous underway measurements from CCE LTER process cruises in the CCE region, 2006 - 2019.

As the research vessel is underway for the duration of a CCE Process Cruise (since 2006, ongoing), 30 parameters are continuously measured regarding the oceanographic surface and atmospheric and navigational environment of the vessel, along the ship's trackline in the CCE region.

openCC0Nov 2021View details →
edi40/100

Plant species percent cover data: Biodiversity II: Effects of Plant Biodiversity on Population and Ecosystem Processes

Biodiversity II (E120) is designed to determine how the number of plant species affects the dynamics of ecological processes at the population, community, and ecosystem levels. By experimentally manipulating the number of species and the kinds of species, the amount of plant growth and the change from year to year, that result can be examined. Plots are large (9m x 9m actively maintained) and well-replicated, allowing responses of plant pathogens, insect herbivores, seed predators, soil parameters, invasive plant species and other variables to also be studied. Plots were seeded in May 1994 to have 1, 2, 4, 8, or 16 species, with roughly 30 replicates of each diversity level. The species composition of each plot was chosen by random draw from a pool of 18 grassland perennials that included four warm-season (C4) grasses, four cool-season (C3) grasses, four legumes, four non-legume forbs, and two woody species. All species occur in monoculture allowing comparison of responses of each species in monoculture to combinations of these same species. The experiment was established in 1994 by the lead investigators David Tilman, Peter Reich, Johannes Knops, and David Wedin. Experiment 120 is similar to Experiment 123, but it uses larger plots to provide a large capacity for long-term subexperiments.

openCC0Dec 2020View details →
zenodo36/100

LCI data for materials and processes comparing energy and water use of aqueous and gas-based metalworking fluids

<p>Datasets containing life cycle inventories for materials and processes, and results of the analysis in the article titled, &quot;Comparing energy and water use of aqueous and gas-based metalworking fluids&quot; published in the <em>Journal of Industrial Ecology.</em></p>

opencc-by-4.0Dec 2019View details →
zenodo36/100

Pre-processed B cell receptor repertoire sequencing data from BioProject PRJNA527941

<p><strong>Data Processing</strong></p> <p>&nbsp;</p> <p>Samples were demultiplexed via their Illumina indices, and processed using the Immcantation toolkit(1,2).&nbsp;Raw fastq files were filtered based on a quality score threshold of 20. Paired reads were joined if they had a minimum length of 10 nt, maximum error rate of 0.3 and a significance threshold of 0.0001. Reads with identical UMI were collapsed to a consensus sequence. Reads with identical full-length sequence and identical constant primer but differing UMI were further collapsed. Sequences were then submitted to IgBlast (3) for VDJ assignment and sequence annotation. Constant region sequences were mapped to germline using Stampy(4). The number and type of V gene mutations was calculated using the shazam R package.(2)</p> <p>&nbsp;</p> <p><strong>software_versions</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;pRESTO:0.5.3,Change-O:0.3.4,IgBlast 1.6.1, stampy1.0.21. shazam0.1.8</p> <p><strong>quality_thresholds</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;FilterSeq.py pRESTO Q&gt;20</p> <p><strong>paired_reads_assembly</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;AssemblePairs.py pRESTO minlen 10 maxerror 0.3 alpha 0.0001</p> <p><strong>primer_match_cutoffs</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;MaskPrimers.py pRESTO C primer &amp; V primer maxerror 0.2</p> <p><strong>consensus_building</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;BuildConsensus.py pRESTO maxerror 0.1 maxgap 0.5</p> <p><strong>collapsing_method</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;CollapseSeq.py pRESTO</p> <p><strong>germline_database&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>IMGT</p> <p>&nbsp;</p> <p><strong>Format</strong></p> <p>&nbsp;</p> <p>Processed sequences are provided in a tab delimited file format, including the following annotations:</p> <p>&nbsp;</p> <p><strong>C_CALL&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>Isotype subclass</p> <p><strong>SEQUENCE_ID&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>Sequence identifier</p> <p><strong>V_CALL&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>V segment gene and allele</p> <p><strong>D_CALL&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>D segment gene and allele</p> <p><strong>J_CALL&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>J segment gene and allele</p> <p><strong>JUNCTION_LENGTH&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>Junction length</p> <p><strong>CONSCOUNT&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>Raw read count from which UMI consensus sequences were generated, summed over all UMIs for the given unique sequence.</p> <p><strong>DUPCOUNT&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>UMI count for the given unique sequence</p> <p><strong>ISOTYPE&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>Constant region primer (isotype)</p> <p><strong>MU_COUNT_CDR_R&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>Number of replacement mutations in CDR region</p> <p><strong>MU_COUNT_CDR_S&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>Number of silent mutations in CDR region</p> <p><strong>MU_COUNT_FWR_R&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>Number of replacement mutations in FWR region</p> <p><strong>MU_COUNT_FWR_S&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>Number of silent mutations in FWR region</p> <p><strong>MUT_TOTAL&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>Total number of mutations in V gene&nbsp;</p> <p><strong>SEQUENCE_INPUT&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>Full length sequence</p> <p><strong>SEQUENCE_IMGT&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>Gapped IMGT sequence</p> <p><strong>V_GERM_START_VDJ&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>position of the first nucleotide in ungapped V germline sequence alignment</p> <p><strong>JUNCTION&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>Junction nucleotide sequence</p> <p><strong>GERMLINE_IMGT_D_MASK&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>IMGT-gapped germline nucleotide sequence with ns masking the NP1-D-NP2 regions</p> <p><strong>Run&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>ID of sequencing run</p> <p><strong>Sample_type&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>The tissue sampled (e.g Peripheral Blood, bone marrow, ..)</p> <p><strong>Sex&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>Sex of the Subject</p> <p><strong>Age&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>Age of the subject</p> <p><strong>UNIQUE_ID&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>Subject identifier&nbsp;</p> <p><strong>SAMPLE_ID&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>Sample identifier, linking back to raw data</p> <p><strong>Subset&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>Defined B cell subset&nbsp;</p> <p><strong>Repertoire&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>Defined B cell repertoire (Naive, Memory IgM/IgD, IgA, IgG)</p> <p><strong>R_SCDR&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>R/S ratio in CDR region</p> <p><strong>R_SFWR&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>R/S ratio in FWR region</p> <p><strong>V_FAM&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>V family gene</p> <p><strong>V_GENE&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>V segment gene</p> <p><strong>D_GENE&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>D segment gene</p> <p><strong>J_GENE&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>J segment gene</p> <p><strong>Clust_Rank&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>Cluster rank</p> <p><strong>Clust_REPRES&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>Cluster representative</p> <p><strong>Clust_SIZE&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>Cluster size</p> <p><strong>Clust_MAXFREQ&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>Cluster maximum frequency</p> <p><strong>Clust_SHAREDNESS&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>Cluster sharedness</p> <p><strong>CDR3_AA_GRAVY&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>CDR3 hydrophobicity index</p> <p><strong>CDR3_AA_CHARGE&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>CDR3 charge</p> <p><strong>CDRH3PDB&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>CDRH3 PDB (Structure) code</p> <p><strong>H1Canon&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>H1 Canonical class</p> <p><strong>H2Canon&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>H2 Canonical class</p> <p><strong>H1_GERMLINE&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>H1 Germline Canonical class</p> <p><strong>H2_GERMLINE&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>H2 Germline Canonical class</p> <p>&nbsp;</p> <p><strong>References</strong></p> <p>1.&nbsp;Vander Heiden, J. A., G. Yaari, M. Uduman, J. N. H. Stern, K. C. O&rsquo;Connor, D. A. Hafler, F. Vigneault, and S. H. Kleinstein.&nbsp;2014. PRESTO: A toolkit for processing high-throughput sequencing raw reads of lymphocyte receptor repertoires.&nbsp;<em>Bioinformatics</em>30: 1930&ndash;1932.</p> <p>2. Gupta, N. T., J. A. Vander Heiden, M. Uduman, D. Gadala-Maria, G. Yaari, and S. H. Kleinstein.&nbsp;2015. Change-O: A toolkit for analyzing large-scale B cell immunoglobulin repertoire sequencing data.&nbsp;<em>Bioinformatics</em>31: 3356&ndash;3358.</p> <p>3. Ye, J., N. Ma, T. L. Madden, and J. M. Ostell. 2013. IgBLAST: an immunoglobulin variable domain sequence analysis tool.&nbsp;<em>Nucleic Acids Res.</em>41.</p> <p>4. Lunter, G., and M. Goodson. 2011. Stampy: A statistical algorithm for sensitive and fast mapping of Illumina sequence reads.&nbsp;<em>Genome Res.</em>21: 936&ndash;939.</p>

opencc-by-4.0Apr 2019View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record