Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
2,848
datasets available to search
ShareScore release 0.9.0
Dataset results
2,848 results for “sequence data”
Data supporting publication: Ultra-deep Sequencing of Hadza Hunter-Gatherers Recovers Vanishing Gut Microbes
<p>Genomes, mapping databases, and data supporting the publication "Ultra-deep Sequencing of Hadza Hunter-Gatherers Recovers Vanishing Gut Microbes".</p> <p><strong>Please see the README on GitHub for descriptions of files and tutorials on how to use them: <a href="https://github.com/MrOlm/ZenodoREADME/blob/main/README.md">https://github.com/MrOlm/ZenodoREADME/blob/main/README.md</a></strong></p>
Spatial data sets of the paper Heinrich Stadial 1 continental sand dunes and Middle to Late Holocene paleosol sequences in SE Iberia: implications for human occupation and site formation processes
<p>Spatial data sets of the paper Heinrich Stadial 1 continental sand dunes and Middle to Late Holocene paleosol sequences in SE Iberia: implications for human occupation and site formation processes. This data set is composed by 3 shapefiles:</p> <ol> <li>Dune_field: Feature class polygon shapefile geometry representing the individual dunes identified in the Villena dune field.</li> <li>Sampled dunes: Shapefile of point geometry representing the location of the stratigraphic sequences of CC1, CC2 and CC3 sampled for texture, soil chemistry, OSL and radiocarbon dating. </li> <li>Sediment sourcing samples: Shapefile of point geometry representing the location of the reference samples of El Moron, El Arenal de la Virgen and Sierra del Castellar. </li> </ol> <p>The spatial reference system is EPSG 25830.</p>
Self-supervised learning of seismological data reveals undocumented eruptive sequences at the Mayotte submarine volcano - Supplementary Materials
<p>The following files are shared:<br> - The scripts used to train the model and generate the figures of the article<br> - The input images used to train the model as well as the final outputs (embedding matrix and the associated filenames matrix)<br> - The clusters organization with their associated images</p>
Training data for 'Exome sequencing data analysis' tutorial (Galaxy Training Material)
<p>The data used in this tutorial are a subset of the data published previously in <a href="https://zenodo.org/record/3243160">Training material for the course "Exome analysis with GALAXY"</a>. Credit for uploading the original data goes to Paolo Uva and Gianmauro Cuccuru!</p> <p>Specifically, you may need the following datasets for following the tutorial:</p> <p><strong>Raw sequencing reads</strong></p> <ul> <li><a href="https://zenodo.org/record/3243160/files/father_R1.fq.gz?download=1">https://zenodo.org/record/3243160/files/father_R1.fq.gz</a></li> <li><a href="https://zenodo.org/record/3243160/files/father_R2.fq.gz?download=1">https://zenodo.org/record/3243160/files/father_R2.fq.gz</a></li> <li><a href="https://zenodo.org/record/3243160/files/mother_R1.fq.gz?download=1">https://zenodo.org/record/3243160/files/mother_R1.fq.gz</a></li> <li><a href="https://zenodo.org/record/3243160/files/mother_R2.fq.gz?download=1">https://zenodo.org/record/3243160/files/mother_R2.fq.gz</a></li> <li><a href="https://zenodo.org/record/3243160/files/proband_R1.fq.gz?download=1">https://zenodo.org/record/3243160/files/proband_R1.fq.gz</a></li> <li><a href="https://zenodo.org/record/3243160/files/proband_R2.fq.gz?download=1">https://zenodo.org/record/3243160/files/proband_R2.fq.gz</a></li> </ul> <p><strong>Premapped sequencing reads</strong></p> <ul> <li><a href="https://zenodo.org/record/3243160/files/mapped_reads_father.bam?download=1">https://zenodo.org/record/3243160/files/mapped_reads_father.bam</a></li> <li><a href="https://zenodo.org/record/3243160/files/mapped_reads_mother.bam?download=1">https://zenodo.org/record/3243160/files/mapped_reads_mother.bam</a></li> <li><a href="https://zenodo.org/record/3243160/files/mapped_reads_proband.bam?download=1">https://zenodo.org/record/3243160/files/mapped_reads_proband.bam</a></li> </ul> <p><strong>Reference sequence (human chromosome 8)</strong></p> <ul> <li><a href="https://zenodo.org/record/3243160/files/hg19_chr8.fa.gz?download=1">https://zenodo.org/record/3243160/files/hg19_chr8.fa.gz</a></li> </ul> <p> </p> <p>If you would just like to play with GEMINI rather than work through the full tutorial, you'll find below a prebuilt GEMINI database (for GEMINI version 0.20.1) for the family trio. You can start exploring this database without having to run GEMINI load and, in fact, without having to install GEMINI's bundled annotation data.</p>
Catalog of NCBI sequence read archive (SRA) data for salamanders at the Hubbard Brook Experimental Forest 2012-2021
This project was designed to describe fine-scale population genetic differentiation of the stream salamander Gryinophilus porphyriticus among five study streams in the Hubbard Brook Experimental Forest. The data are paired with intensive capture-recapture data to assess direct fitness effects of individual genetic diversity, including effects of individual multilocus heterozygosity on stage-specific survival probabilities. This dataset publishes a manifest of the genomic sequence reads submitted to the National Center for Biotechnology Information (NCBI) Sequence Read Archive (SRA). These samples are published at NCBI under the BioProject ID 1090913 (https://www.ncbi.nlm.nih.gov/bioproject/1090913). The tables here include sample metadata and the NCBI URLs to each sample. These data were gathered as part of the Hubbard Brook Ecosystem Study (HBES). The HBES is a collaborative effort at the Hubbard Brook Experimental Forest, which is operated and maintained by the USDA Forest Service, Northern Research Station.
Sequence data for Claar et al. 2020 Scientific Reports
<p>Illumina Mi-Seq sequence data of the ITS2 marker (in .fasta format) associated with Claar et al. 2020 Scientific Reports "Chronic disturbance modulates symbiont (Symbiodiniaceae) beta diversity on a coral reef".</p>
Data for Bovine breed-specific augmented reference graphs facilitate accurate sequence read mapping and unbiased variant discovery
<p><strong>Description of the datasets</strong></p> <p>Data are organized as folders and compressed with tar.gz.</p> <p>There are two compressed data folder: <strong>data </strong>which used for cattle genome graphs experiment and <strong>data_human</strong> which we used for human genome graphs experiment. </p> <p><strong>Cattle genome graphs experiments</strong></p> <p>First you need to unzip the file using command <em>tar -xvzf data.tar.gz</em>. After unzipping, the data folder is organized as follows:</p> <ul> <li>Utilities: contain bovine ARS-UCD 1.2 fasta reference with the accompanying index.</li> <li>Bin: contain the softwares used in the paper (vg, liftover, vcf2diploid)</li> <li>Part1: data for analysis in variant prioritization section, further subdivided into: <ul> <li>vcf_sim: variant files from four animal in each breed used to simulate reads</li> <li>reads_sim: simulated short reads used for read mapping</li> <li>vcf_freq: variants augmented to graphs filtered based on allele frequency</li> </ul> </li> <li>Part2: data used for analysis in the section of graph mapping with breeds-filtered variants, further subdivided into: <ul> <li>vcf_breed: variant files used to graphs construction.</li> </ul> </li> <li>Part3: data used for analysis in the section of consensus genome, further subdivided into: <ul> <li>read_sims: simulated reads as in the part1, but the coordinates are liftovered to the new consensus genomes.</li> <li>reference: contain the original reference and consensus references.</li> <li>vcf_consensus: contain major allele variants to construct consensus genomes.</li> </ul> </li> <li>Part4: data analysis in the section of whole genome graph construction and variant genotyping. <ul> <li>vcf_construct: variants from chromosome 1-29 from 82 Brown Swiss used to construct BSW whole genome graph.</li> <li>BSW_graph: whole genome Brown Swiss graph with the three accompanying indexes (xg,gcsa, and gbwt).</li> </ul> </li> </ul> <p><strong>Human genome graphs experiments</strong></p> <p>First you need to unzip the <em>data_human</em> file using command <em>tar -xvzf data</em><em>_hum.tar.gz</em>. After unzipping, the data folder is organized as follows:</p> <ul> <li>reference: the g1k_v37 reference used as a graph backbone</li> <li>vcf_sim: variant files from four individuals in each population used to simulate reads</li> <li>reads_sim: simulated short reads used for read mapping</li> <li>vcf_freq: variants augmented to graphs filtered based on allele frequency</li> </ul>
Fig. 8 in A revision of the Thyropygus allevatus group. Part V: Nine new species of the extended opinatus subgroup, based on morphological and DNA sequence data (Diplopoda: Spirostreptida: Harpagophoridae)
Fig. 8. Thyropygus sutchariti sp. nov., from Kaeng Krachan, holotype (CUMZ-D00090), ♂, gonopods. A. Anterior view, left telopodite removed. B. Posterior view, left telopodite removed. C. Left telopodite, posterior-mesal view. D. Left telopodite, anterior-lateral view.
Fig. 11. A in A revision of the Thyropygus allevatus group. Part V: Nine new species of the extended opinatus subgroup, based on morphological and DNA sequence data (Diplopoda: Spirostreptida: Harpagophoridae)
Fig. 11. A. Thyropygus navychula sp. nov., specimen from Surin Islands, living ♂ (paratype, CUMZ-D00089-1). B. Thyropygus forceps sp. nov., specimen from Namwang Srithammasokrach, living ♂ (paratype, CUMZ-D00073-1).
Fig. 5 in A revision of the Thyropygus allevatus group. Part V: Nine new species of the extended opinatus subgroup, based on morphological and DNA sequence data (Diplopoda: Spirostreptida: Harpagophoridae)
Fig. 5. Thyropygus mesocristatus sp. nov., from Srikasorn, holotype (CUMZ-D00094), ♂, gonopods. A. Anterior view, left telopodite removed. B. Posterior view, left telopodite removed. C. Lateral view. D. Left telopodite, posterior-mesal view. E. Left telopodite, anterior-lateral view.
Fig. 2 in A revision of the Thyropygus allevatus group. Part V: Nine new species of the extended opinatus subgroup, based on morphological and DNA sequence data (Diplopoda: Spirostreptida: Harpagophoridae)
Fig. 2. Thyropygus cimi sp. nov., from Namwang Srithammasokrach, holotype (CUMZ-D00086), ♂, gonopods. A. Anterior view, left telopodite removed. B. Posterior view, left telopodite removed. C. Lateral view. D. Left telopodite, posterior-mesal view. E. Left telopodite, anterior-lateral view.
Fig. 1 in A revision of the Thyropygus allevatus group. Part V: Nine new species of the extended opinatus subgroup, based on morphological and DNA sequence data (Diplopoda: Spirostreptida: Harpagophoridae)
Fig. 1. Phylogenetic relationships of Thyropygus species based on maximum likelihood analysis (ML) and Bayesian Inference (BI) of 1147 bp of concatenated gene fragments of COI (660 bp) and 16S rRNA (487 bp). Numbers at nodes indicate branch support based on bootstrapping (ML) / posterior probability (BI). Scale bar = 0.06 substitutions/site. # indicates branches which received <50% ML bootstrap support, - indicates non-supported branches by posterior probability. Clade memberships and designations are shown as vertical bars; 1A1 = T. allevatus, 1A2 = cuisinieri subgroup, 1A3 = opinatus subgroup and 1A4 = induratus subgroup. The coloured area marks the T. opinatus subgroup. Abbreviations after species names refer to locality names as shown in Table 1.
Fig. 7 in A revision of the Thyropygus allevatus group. Part V: Nine new species of the extended opinatus subgroup, based on morphological and DNA sequence data (Diplopoda: Spirostreptida: Harpagophoridae)
Fig. 7. Thyropygus planispina sp. nov., from Tham Sua temple, holotype (CUMZ-D00088), ♂, gonopods. A. Anterior view, left telopodite removed. B. Posterior view, left telopodite removed. C. Lateral view. D. Left telopodite, posterior-mesal view. E. Left telopodite, anterior-lateral view.
Fig. 6 in A revision of the Thyropygus allevatus group. Part V: Nine new species of the extended opinatus subgroup, based on morphological and DNA sequence data (Diplopoda: Spirostreptida: Harpagophoridae)
Fig. 6. Thyropygus navychula sp. nov., from Surin Islands, holotype (CUMZ-D00095), ♂, gonopods. A. Anterior view, left telopodite removed. B. Posterior view, left telopodite removed. C. Left telopodite, posterior-mesal view. D. Left telopodite, anterior-lateral view.
Fig. 4 in A revision of the Thyropygus allevatus group. Part V: Nine new species of the extended opinatus subgroup, based on morphological and DNA sequence data (Diplopoda: Spirostreptida: Harpagophoridae)
Fig. 4. Thyropygus forceps sp. nov., gonopods. – A, C–E. Holotype (CUMZ-D00092), ♂, from Namwang Srithammasokrach. A. Anterior view, left telopodite removed. C. Posterior view, left telopodite removed. D. Left telopodite, posterior-mesal view. E. Left telopodite, anterior-lateral view. – B. Specimen from Tham Pha Deang temple (CUMZ-D00093), ♂. Anterior view, left telopodite removed.
Fig. 10 in A revision of the Thyropygus allevatus group. Part V: Nine new species of the extended opinatus subgroup, based on morphological and DNA sequence data (Diplopoda: Spirostreptida: Harpagophoridae)
Fig. 10. Thyropygus ursus sp. nov., from Lanta Islands, holotype (NMHW-Inv.7855), ♂, gonopods. A. Anterior view, left telopodite removed. B. Posterior view, left telopodite removed. C. Left telopodite, posterior-mesal view. D. Left telopodite, anterior-lateral view.
Fig. 9 in A revision of the Thyropygus allevatus group. Part V: Nine new species of the extended opinatus subgroup, based on morphological and DNA sequence data (Diplopoda: Spirostreptida: Harpagophoridae)
Fig. 9. Thyropygus undulatus sp. nov., from Khao Phanom Bencha, holotype (CUMZ-D00087), ♂, gonopods. A. Anterior view, left telopodite removed. B. Posterior view, left telopodite removed. C. Lateral view. D. Left telopodite, posterior-mesal view. E. Left telopodite, anterior-lateral view.
Fig. 3 in A revision of the Thyropygus allevatus group. Part V: Nine new species of the extended opinatus subgroup, based on morphological and DNA sequence data (Diplopoda: Spirostreptida: Harpagophoridae)
Fig. 3. Thyropygus culter sp. nov., from Rorn waterfall, holotype (CUMZ-D00091), ♂, gonopods. A. Anterior view, left telopodite removed. B. Posterior view, left telopodite removed. C. Left telopodite, posterior-mesal view. D. Left telopodite, anterior-lateral view.
Figure 1 in Evaluation of the taxonomy of Helix cincta (Muller, 1774) and Helix nucula (Mousson, 1854); insights using mitochondrial DNA sequence data
Figure 1. Map showing the localities of samples used in the present study representing the morphologically defined species and the distribution of Helix cincta (dash line, light grey) and Helix nucula (continuous line, dark grey).
Supporting data for Novel functional sequences uncovered through a bovine multi-assembly graph
<p><strong>Description of the datasets</strong></p> <p>Data are organized as a folder and compressed with tar.gz.</p> <p>You need to unzip the folder using the command <em>tar -xz</em><em>v</em><em>f</em> data.tar.gz. Unzipping will output a folder named <em>data_tidy</em>, which is organized as follow:</p> <ul> <li>graph.gfa : Graph in GFA format constructed from 6 cattle assemblies</li> <li>nonref.fa : Non-reference sequences extracted from the graph</li> <li>nonref.fa.masked: Hard masked repetitive regions version of nonref.fa</li> <li>nonref_woflanking.fa: Nonref.fa without flanking sequences</li> <li>nonref_woflanking.fa.masked: Masked version of nonref_woflanking.fa</li> <li>augustus_predict.gtf: Annotated gene models of Augustus from non-ref sequences</li> <li>augustus_prot.fa: Protein fasta of the predicted gene models from Augustus</li> <li>breeds_assembled.gtf: Annotation of the StringTie assembled across-breed transcriptome</li> <li>breeds_expressed.tsv: Expression data of breeds_assembled.gtf</li> <li>de_assembled.gtf: Annotation of the StringTie assembled differentially-expressed transcriptome on non-ref sequences</li> <li>de_expression.tsv: Differential expression results from de_assembled.gtf</li> <li>variant_nonref.tsv: Variants called from non-ref sequences (-1, 0, 1, 2 indicates no call, hom ref, het, and hom alt respectively)</li> </ul>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.