Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
2,848
datasets available to search
ShareScore release 0.9.0
Dataset results
2,848 results for “sequence data”
Collection and ddRadSeq sequencing data for Sitophilus zeamais from Oaxaca and Chiapas, Mexico
<p>The maize weevil, <em>Sitophilus zeamais</em>, is a ubiquitous pest of maize and other cereal crops worldwide and remains a threat to food security in subsistence communities. Few population genetic studies have been conducted on the maize weevil, but those that exist have shown that there is very little genetic differentiation between geographically dispersed populations and that it is likely the species has experienced a recent range expansion within the last few hundred years. While the previous studies found little genetic structure, they relied primarily on mitochondrial and nuclear microsatellite markers for their analyses. It is possible that more fine-scaled population genetic structure exists due to local adaptation, the biological limits of natural species dispersal, and the isolated nature of subsistence farming communities. In contrast to previous studies, here, we utilized genome-wide single nucleotide polymorphism data to evaluate the genetic population structure of the maize weevil from the southern and coastal Mexican states of Oaxaca and Chiapas. We employed strict SNP filtering to manage large next generation sequencing lane effects and this study is the first to find fine-scale genetic population structure in the maize weevil. Here, we show that although there continues to be gene flow between populations of maize weevil, that fine-scale genetic structure exists. It is possible that this structure is shaped by local adaptation of the insects, the movement and trade of maize by humans in the region, geographic barriers to gene flow, or a combination of these factors.</p>
DATA for Exploration of O-GlcNAc-transferase (OGT) glycosylation sites reveals a target sequence compositional bias
<p>Mass spectrometry data for identification of glycosylation sites in CBP ID3 and EWS LCRn</p> <p>Perl based implementation of glycosylation Site Predictor OGTcomPred</p>
Effect of different types of sequence data on palaeognath phylogeny
<div class="page"> <div class="layoutArea"> <div class="column"> <p>Palaeognathae consists of five groups of extant species: flighted tinamous (1) and four flightless groups: kiwi (2), cassowaries and emu (3), rheas (4), and ostriches (5). Molecular studies supported the groupings of extinct moas with tinamous and ele- phant birds with kiwi as well as ostriches as the group that diverged first among the five groups. However, phylogenetic re- lationships among the five groups are still controversial. Previous studies showed extensive heterogeneity in estimated gene tree topologies from conserved nonexonic elements, introns, and ultraconserved elements. Using the noncoding loci to- gether with protein-coding loci, this study investigated the factors that affected gene tree estimation error and the relation- ships among the five groups. Using closely related ostrich rather than distantly related chicken as the outgroup, concatenated and gene tree–based approaches supported rheas as the group that diverged first among groups (1)–(4). Whereas gene tree estimation error increased using loci with low sequence divergence and short length, topological bias in estimated trees oc- curred using loci with high sequence divergence and/or nucleotide composition bias and heterogeneity, which more occurred in trees estimated from coding loci than noncoding loci. Regarding the relationships of (1)–(4), the site patterns by parsimony criterion appeared less susceptible to the bias than tree construction assuming stationary time-homogeneous model and sug- gested the clustering of kiwi and cassowaries and emu the most likely with ∼40% support rather than the clustering of kiwi and rheas and that of kiwi and tinamous with 30% support each.</p> </div> </div> </div>
Fibertools: fast and accurate m6A calling using single-molecule long-read sequencing (ML data)
<p>Fibertools is a convolutional neural network that permits the fast and accurate identification of endogenous and exogenous N6-methyladenine (m6A)-marked bases using single-molecule long-read sequencing.<strong> </strong>This dataset (ML data) provides training and validation data for training fibertools supervised and semi-supervised CNN models for three long-read chemistries.</p>
Traing Data for "Assembly of metagenomic sequencing data" tutorial
<p>Metagenomics involves the extraction, sequencing and analysis of combined genomic DNA from <strong>entire microbiome</strong> samples. It includes then DNA from <strong>many different organisms</strong>, with different taxonomic background.</p> <p>Reconstructing the genomes of microorganisms in the sampled communities is critical step in analyzing metagenomic data. To do that, we can use <strong>assembly</strong> and assemblers, <em>i.e.</em> computational programs that stich together the small fragments of sequenced DNA produced by sequencing instruments.</p> <p>Assembling seems intuitively similar to putting together a jigsaw puzzle. Essentially, it looks for reads “that work together” or more precisely, reads that overlap. Tasks like this are <strong>not straightforward</strong>, but rather complex because of the complexity of the genomics (specially the repeats), the missing pieces and the errors introduced during sequencing.</p> <p>In this tutorial, we will learn how to run metagenomic assembly tool and evaluate the quality of the generated assemblies. To do that, we will use data from the study: <a href="https://www.ebi.ac.uk/metagenomics/studies/MGYS00005630#overview">Temporal shotgun metagenomic dissection of the coffee fermentation ecosystem</a>. For an in-depth analysis of the structure and functions of the coffee microbiome, a temporal shotgun metagenomic study (six time points) was performed. The six samples have been sequenced with Illumina MiSeq utilizing whole genome sequencing.</p> <p>Based on the 6 original dataset of the coffee fermentation system, we generated mock datasets for this tutorial.</p>
DNA sequence data for two Roscoea species, R. stenophylla and R. australis (Zingiberaceae)
<p><span>This dataset includes three genomic regions, nrITS (ITS1-5.8S–ITS2) and two chloroplast DNA (cpDNA) regions (psbA-trnH and trnL-F) for two <em>Roscoea</em> species <em>R. stenophylla</em> and <em>R. australis</em> (Zingiberaceae).</span></p>
75-kyr paired mean summer temperature (MST) and moisture balance (BIT-index) data for eastern equatorial Africa based on GDGT distributions in the DeepCHALLA sediment sequence
<p>The DeepCHALLA (DCH) GDGT-based climate proxy data series consist of paired measurements on 373 discrete 2-cm-thick sediment samples, represented by their DCH event-free depth in meter (mefd). Section depth center refers to the distance (in cm) of the sample’s center from the top of the individual core section as represented by its DCH core code. Sediment age is based on the age model as presented in Fig. 2 in Baxter et al. (2023). Mean Summer Temperature (MST) was calculated from the distribution of sedimentary brGDGTs according to Pearson et al. (2011) and rescaled using an ensemble reconstruction for eastern Africa covering the last 25 ka (Extended Data Fig. 4 in Baxter et al., 2023). The branched versus isoprenoid tetraether (BIT) index of sedimentary brGDGTs from Lake Chala, reflecting lake water-balance variation, was calculated according to Hopmans et al. (2004).</p>
Data from: A comparison of non-destructive visceral swab and tissue biopsy sampling methods for genotyping-by-sequencing in the freshwater mussel Fusconaia askewi
<p>Limiting harm to organisms via genetic sampling is an important consideration for rare species. Nondestructive sampling techniques have been developed to address this issue in freshwater mussels. Two methods, visceral swabbing and tissue biopsies, have proven to be effective for DNA sampling, though it is unclear as to which method is preferable for genotyping-by-sequencing (GBS). Tissue biopsies may cause undue stress and damage to organisms, while visceral swabbing potentially reduces the chance of such harm. Our study compared the efficacy of these two DNA sampling methods for generating GBS data for the Unionid freshwater mussel, Texas Pigtoe (<em>Fusconaia askewi</em>). Our results find both methods generate quality sequence data, though some considerations are in order. Tissue biopsies produced significantly higher DNA concentrations and larger numbers of reads when compared to swabs, though there was no significant association between starting DNA concentration and number of reads generated. Swabbing produced greater sequence depth (more reads per sequence) while tissue biopsies revealed greater coverage across the genome (at lower sequence depth). Patterns of genomic variation as characterized in principal component analyses were similar regardless of the sampling method, suggesting that the less invasive swabbing is a viable option for producing quality GBS data in these organisms.</p>
Additional raw data in `Cell-type-specific co-expression inference from single cell RNA-sequencing data'.
<p>This repository holds the additional raw data used to generate figures in the publication "<strong><em>Cell-type-specific co-expression inference from single cell RNA-sequencing data</em></strong>" (preprint version: <a href="http://source%20code%20repo%20for%20%60cell-type-specific%20co-expression%20inference%20from%20single%20cell%20rna-sequencing%20data%27./">https://www.biorxiv.org/content/10.1101/2022.12.13.520181v1</a>).</p> <p>Table of contents:</p> <ul> <li>Figure_1B.rds: <ul> <li>raw data of Figure 1B </li> <li>co-expression estimates of 500*499/2 gene pairs across 100 replicates for 7 methods under two settings of sequencing detph variations</li> </ul> </li> <li>Supplementary_Figure_1B.rds: <ul> <li>raw data of Supplementary Figure 1B </li> <li>co-expression estimates of 500*499/2 gene pairs across 100 replicates for 7 methods under two settings of sequencing detph variations</li> </ul> </li> <li>Supplementary_Figure_2.rds: <ul> <li>raw data of Supplementary Figure 2 </li> <li>empirical power evaluated for 4999 gene pairs and 6 methods</li> </ul> </li> <li>Figure_3B.rds: <ul> <li>raw data of Figure 3B</li> <li>co-expression estimates of a network of 500 genes for 9 methods across 100 replicates</li> </ul> </li> <li>Additional_Raw_Data.xlsx <ul> <li>raw data of Figure 3A: (geometric mean expression levels, co-expression estimates) for 4999 gene pairs and 11 methods</li> <li>raw data of Figure 3C: running times for 11 methods</li> <li>raw data of Supplementary Figure 3: (geometric mean expression levels, co-expression estimates) for 4999 gene pairs and 11 methods under two settings of sequencing detph variations</li> </ul> </li> </ul>
454-sequence data of Iron Age cattle from Althiburos – Tunisia
<p>The Maghreb is a key region for understanding the dynamics of cattle dispersal and admixture with local aurochs following their earliest domestication in the Fertile Crescent more than 10,000 years ago. Here, we present data on mitochondrial <em>D-loop</em> sequences obtained for 12 archaeological specimens of Iron Age (~2,800 cal BP–2,000 cal BP) domestic cattle from the Eastern Maghreb, i.e. Althiburos (El Kef, Tunisia). Maternal lineages were assigned to the elusive R and ubiquitous African-T1 haplogroups found in two and ten Althiburos specimens, respectively. Our results corroborate the introgression of aurochs females into the domestic stock of cattle from Althiburos. </p>
Dataset for: How eDNA data filtration, sequence coverage, and primer selection influence assessment of fish communities in northern temperate lakes
<p><span>For nearly 15 years now, environmental DNA</span><span> has demonstrated</span><span> its effectiveness in monitoring biodiversity. Methodological and technical improvements have significantly enhanced the field. However, the effect of factors such as sequence coverage, bioinformatic filtration and primer choice have been less explored or need to be optimized according </span><span>to </span><span>specific survey objectives and </span><span>study </span><span>site characteristics. We evaluated these factors </span><span>to </span><span>help optimize monitoring fish biodiversity in North American temperate lakes. We sampled water for fish community eDNA analysis in 12 lakes from southwestern Québec, Canada. The lakes were selected to encompass a wide range of surface areas and species richness. We sampled water from a total of </span><span>520</span><span> sites (25 to 50 per lake) and analyzed three mitochondrial DNA regions (12S rRNA; 16S rRNA; and cytb) using NovaSeq</span><span> sequencing. Our results, based on rarefied count matrices (from a sequencing depth of 100,000 to a minimum </span><span>depth </span><span>of 1,000 reads per sample), </span><span>showed</span><span> that </span><span>keeping only</span><span> species </span><span>in each sample if they</span><span> represented </span><span>at least one thousandth (species </span><span>minimum </span><span>read proportion threshold =</span><span> 0.001</span><span>)</span><span> of the </span><span>sample's</span><span> reads was adequate to remove false positives </span><span>and had a limited negative</span><span> impact on true positives</span><span> with low read counts. The</span><span> sequencing depth </span><span>was found to have</span><span> a negligible impact </span><span>on the accuracy</span><span> of fish </span><span>community assessment in a given lake. With the same sequencing depth and a complete local reference database for each primer set, </span><span>a single primer set </span><span>produced</span><span> similar species richness medians than the combination of two or three primer sets. Overall, 12S and 16S detected more species and provided more consistent community profiles than cytb. </span><span>Based on our observations, we suggest using the 12S MiFish-U primer set and applying a minimum proportion of 0.001 reads per species and site to monitor north-temperate lentic freshwater fish communities.</span></p>
Data files for 'Tan et al., (2022). Seismogenesis of the 2021 Mw 7.1 earthquake sequence near the northeastern Japan revealed by double-difference seismic tomography'
<p>catalog.dat : the selected earthquake phase data (originated from Hi-net, https://www.hinet.bosai.go.jp/?LANG=en)</p> <p>station.dat : the seismic station coordinates</p> <p>MOD : the initial velocity models (including the grid nodes, Vp & Vp/Vs)</p> <p>relocation.dat : the earthquake relocations by the DD tomography</p> <p>Vp_model.dat, Vs_model.dat, VpVs_model.dat : the inverted 3D velocity models by DD tomography (having exactly the same layout as MOD)</p>
Sequencing data for: Tracking climate-change induced biological invasions over 4 decades by metabarcoding archived natural eDNA samplers
<p><span>In a time of unprecedented global environmental change, understanding the response of biodiversity is paramount. However, our knowledge of anthropogenic impacts on ecosystems is limited by a lack of standardized retrospective biomonitoring data. Here, we use four-decade time series of archived blue mussels to trace spatiotemporal biodiversity change in coastal ecosystems. The filter-feeding mussels can serve as natural eDNA samplers, carrying an imprint of the surrounding aquatic community at the time of sampling. By sequencing the preserved DNA, we characterize highly diverse mussel-associated communities and reconstruct the invasion trajectory of an invasive species to the detriment of native taxa uncovering repeated population collapses and reinvasions after cold winters. Time series of natural eDNA samplers provide highly resolved temporal data on community assembly and global warming-driven invasion processes and overcome critical shortfalls in our understanding of biodiversity change in the Anthropocene.</span></p>
Data from: Insight into the structural and magnetotransport properties of epitaxial alpha-Fe2O3/Pt(111) heterostructures: The role of the reversed layer sequence
<p>We report on the chemical structure and spin Hall magnetoresistance (SMR) in epitaxial α-Fe<sub>2</sub>O<sub>3</sub>(hematite)(0001)/Pt(111) bilayers with hematite thicknesses of 6 nm and 15 nm grown by molecular beam epitaxy on a MgO(111) substrate. Unlike previous studies that involved Pt overlayers on hematite, the present hematite films were grown on a stable Pt buffer layer and displayed structural changes as a function of thickness. These structural differences (the presence of a ferrimagnetic phase in the thinner film) significantly affected the magnetotransport properties of the bilayers. We observed a sign change of the SMR from positive to negative when the thickness of hematite increased from 6 nm to 15 nm. For α-Fe<sub>2</sub>O<sub>3</sub>(15 nm)/Pt, we demonstrated room-temperature switching of the Néel order with rectangular, nondecaying switching characteristics. Such structures open the way to extending magnetotransport studies to more complex systems with double asymmetric metal/hematite/Pt interfaces.</p>
Serratia fonticola EBS19 Whole genome sequence data fasta file annotated
<p>Whole genome sequence data of <em>Serratia fonticola</em> <strong>EBS19 </strong>strain.</p>
Sample extraction and SNP sequencing data for: Identification of sex-linked SNP markers in wild populations of monomorphic birds
<p><span>Single-nucleotide polymorphism (SNP) analyses are a powerful tool for population genetics, pedigree reconstruction and phenotypic trait mapping. However, the untapped potential of SNP markers to discriminate the sex of individuals in species with reduced sexual dimorphism or of individuals during immature stages remains a largely unexplored avenue. Here, we develop a novel protocol for molecular sexing of birds based on the detection of unique Z- and W-linked SNP markers. Our method is based on the identification of two unique loci, one in each sexual chromosome. Individuals are considered males when they show no calls for the W-linked SNP and are heterozygotic or homozygotic for the Z-linked SNP, while females show both Z- and W-linked SNP calls. We validated the method in the Jackdaw (<em>Corvus</em> <em>monedula</em>). The reduced sexual dimorphism in this species makes it difficult to sex individuals in the wild. We assessed the reliability of the method using 36 individuals of known sex and found that their sex was correctly assigned in 100% of cases. The sex-linked markers also proved to be widely applicable to discriminate males and females from a sample of 927 genotyped individuals of different maturity stages with an accuracy of 99.5%. Given that SNP markers are increasingly used in quantitative genetic analyses of wild populations, the approach we propose has a great potential to be integrated into broader genetic research programmes without the need for additional sexing techniques.</span></p>
Study group charasteristics and sequencing data of patients with essential thrombocythemia and polycythemia vera
<p><span>Polycythemia vera (PV) and essential thrombocythemia (ET) are diseases driven by canonical mutations in <em>JAK2, CALR</em>, or <em>MPL </em>gene. Previous studies revealed that in addition to driver mutations, patients with PV and ET can harbor other mutations in various genes, with no established impact on disease phenotype. We hypothesized that the molecular profile of patients with PV and ET is dynamic throughout the disease. In this study we performed 37-gene targeted next-generation sequencing panel on the DNA samples collected from 49 study participants in two time points, separated by 78-141 months. We identified 78 variants across 37 analyzed genes in the study population. By analyzing the change in variant allele frequencies (VAFs) and revealing the acquisition of new mutations during the disease, we confirmed the dynamic nature of molecular profile of patients with PV and ET. We found connections of specific variants with the development of secondary myelofibrosis, thrombotic events, and response to treatment. We confronted our results with existing conventional and mutation-enhanced prognostic systems, showing the limited utility of available prognostic tools. Results of this study underline the significance of repeated molecular testing in patients with PV and ET and indicate the need for further research within this field to better understand the disease and improve available prognostic tools.</span></p>
Single Cell RNA sequencing data of ADT treated Prostate cancer patients
<p>The data was generated from a study that was conducted according to guidelines approved by the Review Board at the University of Texas Southwestern Medical Center. We procured patient biopsy samples from two distinct studies. The first is titled "Tissue Collection and Results Gathering for Radiotherapy Patients & Healthy Individuals" (STU 072010-098), and the second is a Phase I Clinical Study on Stereotactic Ablative Radiotherapy (SABR) for Pelvic and Prostate Areas in High-Risk Prostate Cancer Patients (STU062014-027). The single-cell RNA sequencing (scRNA-seq) took place in Dr. Douglas Strand's laboratory, adhering to the method outlined in Henry et al<sup>1</sup>. We used a 1-hour treatment with 5mg/ml of collagenase type I, 10mM of ROCK inhibitor, and 1mg of DNase. Barcode labeling for 3' GEX was done using a 10X machine, and the sequencing process utilized an Illumina NextSeq 500 device.</p> <p> </p> <p>1. Henry, G. H., Malewska, A., Joseph, D. B., Malladi, V. S., Lee, J., Torrealba, J., ... & Strand, D. W. (2018). A cellular anatomy of the normal adult human prostate and prostatic urethra. <em>Cell reports</em>, <em>25</em>(12), 3530-3542.</p>
Raw target enrichment of conserved element sequence data for 24 black coral species
<p><span>Deep-sea lineages are generally thought to arise from shallow-water ancestors, but this hypothesis is based on a relatively small number of taxonomic groups. Anthozoans, which include corals and sea anemones, are significant contributors to the faunal diversity of the deep sea, but the timing and mechanisms of their invasion into this biome remain elusive. Here, we reconstruct a fully resolved, time-calibrated phylogeny of 83 species in the order Antipatharia (black coral) to investigate their bathymetric evolutionary history. Our reconstruction indicates that extant black coral lineages first diversified in continental slope depths (~250–3,000 m) during the early Silurian (~437 Ma) and subsequently radiated into, and diversified within, both continental shelf (<250 m) and abyssal (>3,000 m) habitats. Ancestral state reconstruction analysis suggests that the appearance of morphological features that enhanced the ability of black corals to acquire nutrients coincided with their invasion of novel depths. Our findings have important conservation implications for anthozoan lineages, as the loss of "source" slope lineages could threaten millions of years of evolutionary history and confound future invasion events, thereby warranting protection. </span></p>
Models and Data associated with: Single-cell gene expression prediction from DNA sequence at large contexts
<p>This archive holds trained models and associated data for the <a href="https://www.biorxiv.org/content/10.1101/2023.07.26.550634v1">manuscript</a>:<br> "Single-cell gene expression prediction from DNA sequence at large contexts"</p> <p>Structure:</p> <ul> <li>configs - example configs for the workflows to produce publication data </li> <li>data_* - pre-processed single cell data used for publication</li> <li>models_* - model checkpoints, hyperparameters and training progress in tensorboard logs</li> <li>preprocessing - additional data required to reproduce the pre-processing workflow</li> </ul> <p> </p> <p>"Copyright 2023 GlaxoSmithKline Research & Development Limited. All rights reserved."</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.