Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

2,848

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

2,848 results for “sequence data”

Learn how ShareScore rates datasets ↗
zenodo32/100

Sample data for sequencing QC very short introduction

<p>Sample data for training</p>

opencc-by-4.0Mar 2024View details →
zenodo32/100

Single nucleus RNA-sequencing data related to Van Hijfte et al. 2024

Open the record for dataset details and reuse information.

opencc-by-4.0Dec 2023View details →
zenodo32/100

Datasets for "Physicochemical graph neural network for learning protein-ligand interaction fingerprints from sequence data"

<div> <p>Datasets used for implementing the <a href="https://github.com/huankoh/PSICHIC">PSICHIC</a> experiments shown in the <a href="https://doi.org/10.1101/2023.09.17.558145">manuscript</a>.</p> <p>&nbsp;</p> </div>

opencc-by-4.0Mar 2024View details →
dryad32/100

Data from: Estimating bloodstain age in the short term based on DNA fragment length using nanopore sequencer

<p>We used a nanopore sequencer to quantify DNA fragments &gt; 10,000 bp in size and then evaluated their relationship with short-term bloodstain age. Moreover, DNA degradation was investigated after bloodstains were wetted once with water. Bloodstain samples on cotton gauze were stored at room temperature and low humidity for up to 6 months. Bloodstains stored for 1 day were wetted with nuclease-free water, allowed to dry, and stored at room temperature and low humidity for up to 1 week. The proportion of fragments &gt; 20,000 bp in dry bloodstains tended to decrease over time, particularly for fragments &gt; 50,000 bp in size. This trend was modeled using a power approximation curve, with the highest R2 value (0.6475) noted for fragments &gt; 50,000 bp in size; lower values were recorded for shorter fragments. The proportion of longer fragments was significantly reduced in bloodstains that were dried after being wetted once, and there was significant difference in fragments &gt; 50,000 bp between dry conditions and once-wetted. This result suggests that even temporary exposure to water causes significant DNA fragmentation, but not extensive degradation. Thus, bloodstains that appear fresh but have a low proportion of long DNA fragments may have been wetted previously. Our results indicate that evaluating the proportion of long DNA fragments yields information on both bloodstain age and the environment in which they were stored.</p>

opencc-zeroApr 2024View details →
dryad32/100

Data from: Demographic inference from whole-genome and RAD sequencing data suggests alternating human impacts on goose populations since the last ice age

We investigated how population changes and fluctuations in the pink-footed goose might have been affected by climatic and anthropogenic factors. First, genomic data confirmed the existence of two separate populations: western (Iceland) and eastern (Svalbard/Denmark). Second, emographic inference suggests that the species survived the last glacial period as a single ancestral population with a low population size (100-1,000 individuals) that split into the current populations at the end of the Last Glacial Maximum with Iceland being the most plausible glacial refuge. While population changes during the last glaciation were clearly environmental, we hypothesize that more recent demographic changes are human-related: (1) the inferred population increase in the Neolithic is due to deforestation to establish new lands for agriculture, increasing available habitat for pink-footed geese (2) the decline inferred during the Middle Ages is due to human persecution and (3) improved protection explains the increasing demographic trends during the 20th century. Our results suggest both environmental (during glacial cycles) and anthropogenic effects (more recent) can be a threat to species survival.

opencc-zeroDec 2016View details →
zenodo32/100

Fig. 3 in Sequence capture data support the taxonomy of Pogonolepis (Asteraceae: Gnaphalieae) and show unexpected genetic structure

Fig. 3. Likelihood phylogeny of concatenated supermatrix of Pogonolepis and outgroups. Numbers above branches indicate UltraFast Bootstrap values, gene Concordance Factors, and site Concordance Factors. The dashed branch was shortened for the figure. Western Australian specimens of P. muelleriana are marked with (WA), all others are from eastern states. A and B indicate informally named clades inside P. stricta.

opennotspecifiedAug 2022View details →
zenodo32/100

Fig. 2 in Sequence capture data support the taxonomy of Pogonolepis (Asteraceae: Gnaphalieae) and show unexpected genetic structure

Fig. 2. Geographic spread of the specimens of Pogonolepis at CANB sampled for molecular analysis (large, pale circles) and ranges of species according to the Australasian Virtual Herbarium (dots; see https://doi.org/10.26197/ala.6556e234- 7160-49de-bf63-5123af624e94, accessed 31 May 2022). Red: P. muelleriana; blue: P. stricta. Note that specimens of P. muel-leriana geocoded in Canberra and Hobart were likely cultivated.

opennotspecifiedAug 2022View details →
zenodo32/100

Fig. 1 in Sequence capture data support the taxonomy of Pogonolepis (Asteraceae: Gnaphalieae) and show unexpected genetic structure

Fig. 1. (a) Pogonolepis muelleriana, New South Wales, SchmidtLebuhn 1633 (CANB); (b) Pogonolepis stricta, Western Australia, Schmidt-Lebuhn 1474 (CANB).

opennotspecifiedAug 2022View details →
zenodo32/100

SeQUeNCe Parallel Data

<p>This is the SeQUeNCe data used in preparation of our parallel simulation publication. The corresponding preprint may be found on <a href="https://doi.org/10.48550/arXiv.2111.03918">arXiv</a>.</p>

opencc-by-4.0Nov 2023View details →
zenodo32/100

RNA sequencing data from the prefrontal cortex and hippocampus of male (12 weeks old) hemizyguous CAG-HERV-W-env mice and wild-type controls

<p>RNA sequencing data from the prefrontal cortex (PFC) and hippocampus (HIPP) of male (12 weeks old) hemizyguous CAG<sup>HERV-Wenv </sup>mice ( C57BL6/J;129P2/Ola-Hprt mice; <em>n</em> = 3) relative to wild-type ( <em>n</em> = 3) littermates.&nbsp;Total RNA was extracted from prefrontal and hippocampal samples using the SPLIT RNA extraction kit (Lexogen, Austria) following the manufacturer&rsquo;s recommendations and was sent to the Functional Genomics Center in Zurich (FGCZ) for quality control and RNA sequencing.&nbsp;The quality of the isolated RNA was determined with a Fragment Analyzer (Agilent, Santa Clara, California, USA). Only those samples with a 260 nm/280 nm ratio between 1.8&ndash;2.1, a 28S/18S ratio within 1.5&ndash;2, and RIN (&gt;8) values qualified for a Poly-A enrichment strategy in order to generate the sequencing libraries applying the TruSeq mRNA Stranded Library Prep Kit (Illumina, Inc, California, USA). After Poly-A selection using Oligo-dT beads the mRNA was reverse-transcribed into cDNA. The cDNA was fragmented, end-repaired and poly-adenylated before ligation of TruSeq UD Indices (IDT, Coralville, Iowa, USA). The quality and quantity of the amplified sequencing libraries were validated using a Fragment Analyzer SS NGS Fragment Kit (1&ndash;6000 bp) (Agilent, Waldbronn, Germany). The equimolar pool of the samples was spiked into a NovaSeq6000 run targeting ~15M reads per sample on a S1 FlowCell (Novaseq S1 Reagent Kit, 100 cycles, Illumina, Inc, California, USA). Reads were quality-checked with FastQC. Sequencing adapters were removed with Trimmomatic and aligned to the reference genome and transcriptome of Mus Musculus (GENCODE, GRCm38,p5) with STAR v2.7.3. Distribution of the reads across genomic isoform expression was quantified using the R package GenomicRanges from Bioconductor Version 3.10. Minimum mapping quality, as well as minimum feature overlaps, was set to 10. Multi-overlaps were allowed. Differentially expressed genes (DEGs) were identified using the R package edgeR from Bioconductor Version 3.10, using a generalized linear model (glm) regression, a quasi-likelihood (QL) differential expression test and the trimmed means of M-values (TMM) normalization.</p>

restrictedcc-by-sa-4.0Nov 2024View details →
zenodo32/100

Whole animal single nuclei RNA sequencing data of catenulid Stenostomum brevipharyngium

<p>Provided is the count tables as well as the final h5ad.&nbsp;</p>

opencc-by-4.0Nov 2024View details →
zenodo32/100

Source data of the MultiSTAAR manuscript "A statistical framework for multi-trait rare variant analysis in large-scale whole-genome sequencing studies".

<p>This dataset serves as the source data for Figures 2-3 and Extended Data Figures 1-2 of the MultiSTAAR manuscript titled "A statistical framework for multi-trait rare variant analysis in large-scale whole-genome sequencing studies". MultiSTAAR is a statistical framework and computationally-scalable analytical pipeline for functionally-informed multi-trait rare variant analysis in large-scale WGS studies.<br><br><strong>Figure 2.</strong> Manhattan plots and Q-Q plots for unconditional gene-centric coding, noncoding and ncRNA multi-trait analysis of low-density lipoprotein cholesterol (LDL-C), high-density lipoprotein cholesterol (HDL-C) and triglycerids (TG) using TOPMed data (<em>n</em> = 61,838).<br><br><strong>Figure 3.</strong> TOPMed genetic region (2-kb sliding window) unconditional multi-trait analysis results of low-density lipoprotein cholesterol (LDL-C), high-density lipoprotein cholesterol (HDL-C) and triglycerides (TG) using TOPMed data (<em>n</em> = 61,838).<br><br><strong>Extended Data Figure 1.</strong> Manhattan plots and Q-Q plots for unconditional gene-centric coding, noncoding and genetic region (2-kb sliding window) multi-trait analysis of fasting glucose (FG) and fasting insulin (FI) using TOPMed data (<em>n</em> = 21,731).<br><br><strong>Extended Data Figure 2.</strong> Manhattan plots and Q-Q plots for unconditional gene-centric coding, noncoding and genetic region (2-kb sliding window) multi-trait analysis of C-reactive protein (CRP), interleukin 6 (IL-6), lipoprotein-associated phospholipase A2 (Lp-PLA2) activity, and lipoprotein-associated phospholipase A2 (Lp-PLA2) mass using TOPMed data (<em>n</em> = 9,380).</p>

opencc-by-4.0Nov 2024View details →
zenodo32/100

Whole genome re-sequencing workshop data: fastq files and reference genomes

<p>&nbsp;</p> <p>Whole genome re-sequencing data analysis workshop datasets. The files are necessary inputs for the workshop in https://github.com/PoODL-CES/Genomics_learning_workshop</p> <p>Tools and scripts listed in the https://github.com/PoODL-CES/Genomics_learning_workshop repository.</p> <p>&nbsp;</p> <p>These are subsampled fastq files from:<br><br>Khan, A., Patel, K., Shukla, H., Viswanathan, A., van der Valk, T., Borthakur, U., Nigam, P., Zachariah, A., Jhala, Y.V., Kardos, M. and Ramakrishnan, U., 2021. Genomic evidence for inbreeding depression and purging of deleterious genetic variation in Indian tigers. <em>Proceedings of the National Academy of Sciences</em>, <em>118</em>(49), p.e2023018118.</p> <p>The reference genome is from :</p> <p>Shukla, H., Suryamohan, K., Khan, A., Mohan, K., Perumal, R.C., Mathew, O.K., Menon, R., Dixon, M.D., Muraleedharan, M., Kuriakose, B. and Michael, S., 2023. Near-chromosomal de novo assembly of Bengal tiger genome reveals genetic hallmarks of apex predation. <em>GigaScience</em>, <em>12</em>, p.giac112.</p> <p>&nbsp;</p> <p>The reference has been indexed using:</p> <p>bwa index <a target="_blank" rel="noopener noreferrer">GCA_021130815.1_PanTigT.MC.v3_genomic.fna</a></p> <p>&nbsp;</p>

opencc-by-4.0Dec 2024View details →
zenodo32/100

FIGURE. Bayesian tree based on nuclear (ITS) sequence data showing phylogenetic position of Hedysarum sunhangii sp. nov. in Subsect. Crinifera. Bayesian posterior probability (PP) / maximum parsimony (MP) are given on each branch, respectively; maximum likelihood (ML) is below branches. in Hedysarum sunhangii (Fabaceae, Hedysareae), a new species from Pamir-Alay (Babatag Ridge - Uzbekistan)

FIGURE. Bayesian tree based on nuclear (ITS) sequence data showing phylogenetic position of Hedysarum sunhangii sp. nov. in Subsect. Crinifera. Bayesian posterior probability (PP) / maximum parsimony (MP) are given on each branch, respectively; maximum likelihood (ML) is below branches.

opennotspecifiedOct 2021View details →
zenodo32/100

FIGURE. Bayesian tree based on combined plastid (matK, trnL-trnF) sequence data showing phylogenetic position of Hedysarum sunhangii sp. nov. in Subsect. Crinifera. Bayesian posterior probability (PP) / maximum parsimony (MP) are given on each branch, respectively; maximum likelihood (ML) is below branches in Hedysarum sunhangii (Fabaceae, Hedysareae), a new species from Pamir-Alay (Babatag Ridge - Uzbekistan)

FIGURE. Bayesian tree based on combined plastid (matK, trnL-trnF) sequence data showing phylogenetic position of Hedysarum sunhangii sp. nov. in Subsect. Crinifera. Bayesian posterior probability (PP) / maximum parsimony (MP) are given on each branch, respectively; maximum likelihood (ML) is below branches

opennotspecifiedOct 2021View details →
dryad32/100

Targeted Next Generation Sequencing data for European Green Crab

<p class="western"><span>In the northeast Pacific Ocean there is high interest in developing eDNA-based survey methods to aid management of invasive populations of European green crab (<i>Carcinus maenas</i>). Expected benefits are improved sensitivity for early detection of secondary spread and to assess the outcome of eradication efforts. A new eDNA-based approach we term 'Targeted Next Generation Sequencing (tNGS)' is introduced here and shown to improve detection relative to qPCR at sites with lower green crab CPUE values measured by trapping. DNA standards (gBlock) with starting molecule copies that were 10- to 100- times lower than the qPCR limit of detection returned significant numbers of sequencing reads, which in our field assessments translated to a 7% - 10% increase in detection probability from tNGS relative to qPCR at sites with lower CPUE. We also found the number of sequencing reads from tNGS was significantly correlated with green crab CPUE whereas Ct values from qPCR were not. When sources of variation were partitioned for each assay, we found the difference between mean within-site and mean between-site variation was much larger and had non-overlapping confidence intervals for tNGS relative to qPCR, suggesting the former may offer more power for detecting spatial variation in eDNA availability. Results presented here indicate this approach is suitable for species of known low abundances where a positive detection has high economic or environmental consequences. Any species with an existing qPCR assay can be easily tested with a tNGS assay. We conclude with a discussion on the fit for purpose applications of tNGS vs. qPCR on how to best apply molecular surveys in management programs.</span></p>

opencc-zeroNov 2021View details →
zenodo32/100

16S rDNA sequence data of research--Different periods of high-fat diet during peri-pregnancy cause variant offspring gut microbiota

<p>&nbsp; These are the 16S rRNA gene sequencing data based on&nbsp;the Illumina HiSeq2500 platform. Group a, b, c, d stand for group CD-CD, CD-HFD, HFD-CD, HFD-HFD separately. Each mouse of one group&nbsp;had two samples such as a1_1 and a1_2.</p>

opencc-by-4.0Nov 2021View details →
dryad32/100

Sequencing data and taxonomic assignments from: Biodiversity and vector-borne diseases: host dilution and vector amplification occur simultaneously for Amazonian leishmaniases

<p>This is the sequencing data used in the paper:<strong> "</strong>Biodiversity and vector-borne diseases: host dilution and vector amplification occur simultaneously for Amazonian leishmaniases" by Kocher et al. The study aims at assessing the effects of biodiversity changes on Leishmania transmission using molecular analyses of sand fly pools and blood-fed dipterans. The data is split in three files corresponding to the PCR amplicons used in the study (Ins16S for insect identifications, 12SV5 for vertebrate identifications and leishmini for Leishmania identifications). Each file combines output from different Illumina Miseq and Hiseq runs, after read demultiplexing and adapter trimming, dereplication and removal of reads present in less than 10 copies (but before further read filtering). The data is presented in tabular format (similar to the output of the obitab command from the obitools package), together with information on the corresponging sample, sequencing run, and taxonomic assignments (performed with ecotag from the obitools).</p>

opencc-zeroNov 2021View details →
zenodo32/100

Magnetic bead epicPCR sequencing data, December 3 2021

<p>Samples:</p> <p>1. Chilomonas, 16S</p> <p>2. Rhodomonas, 16S</p> <p>3. Chilomonas+Rhodomonas, 16S</p> <p>4. Chilomonas+Rhodomonas+wastewater, 16S</p> <p>5. Wastewater, 16S</p> <p>6. Chilomonas, 18S</p> <p>7. Rhodomonas, 18S</p> <p>8. Chilomonas+Rhodomonas, 18S</p> <p>9. Chilomonas+Rhodomonas+wastewater, 18S</p> <p>10. Wastewater, 18S</p>

opencc-by-4.0Dec 2021View details →
dryad32/100

Hawksbill turtle ddRAD raw sequencing data

<p>Pleistocene environmental changes are generally assumed to have dramatically affected species' demography via changes in habitat availability, but this is challenging to investigate due to our limited knowledge of how Pleistocene ecosystems changed through time. Here, we tracked changes in shallow marine habitat availability resulting from Pleistocene sea level fluctuations throughout the last glacial cycle (120 – 14 thousand years ago; kya) and assessed correlations with past changes in genetic diversity inferred from genome-wide SNPs, obtained via ddRAD sequencing, in Caribbean hawksbill turtles, which feed in coral reefs commonly found in shallow tropical waters. We found sea level regression resulted in an average 75% reduction in shallow marine habitat availability during the last glacial cycle. Changes in shallow marine habitat availability correlated strongly with past changes in hawksbill turtle genetic diversity, which gradually declined to ~1/4th of present-day levels during the Last Glacial Maximum (LGM; 26 – 19 kya). Shallow marine habitat availability and genetic diversity rapidly increased after the LGM, signifying a population expansion in response to warming environmental conditions. Our results suggest a positive correlation between Pleistocene environmental changes, habitat availability and species' demography, and that demographic changes in hawksbill turtles were potentially driven by feeding habitat availability. However, we also identified challenges associated with disentangling the potential environmental drivers of past demographic changes, which highlight the need for integrative approaches. Our conclusions underline the role of habitat availability on species' demography and biodiversity, and that the consequences of ongoing habitat loss should not be underestimated.</p>

opencc-zeroDec 2020View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record