Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
2,848
datasets available to search
ShareScore release 0.9.0
Dataset results
2,848 results for “sequence data”
Transcriptome sequencing data analysis of Bacillus subtilis NCIB 3610 wild-type strain, tmRNA mutant strain and MB revertant strain
GEO Series GSE199151. Bacillus subtilis subsp. subtilis NCIB 3610 = ATCC 6051 = DSM 10. 9 samples. Type: Expression profiling by high throughput sequencing.
High-throughput muscle fiber typing from RNA sequencing data
GEO Series GSE190489. Homo sapiens; Pan troglodytes. 2 samples. Type: Expression profiling by high throughput sequencing.
PING 2.0: An R/Bioconductor package for nucleosome positioning using next-generation sequencing data
GEO Series GSE47073. Saccharomyces cerevisiae. 2 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.
The data of complete chloroplast genome sequence of Tilia miqueliana (Malvaceae) in China
<p>This dataset includes the complete chloroplast genome of Tilia miqueliana (Malvaceae) in China.</p>
Data from: From environmental DNA sequences to ecological conclusions: how strong is the influence of methodological choices?
Aim: Environmental DNA (eDNA) is increasingly used for analysing and modelling all-inclusive biodiversity patterns. However, the reliability of eDNA-based diversity estimates is commonly compromised by arbitrary decisions for curating the data from molecular artefacts. Here, we test the sensitivity of common ecological analyses to these curation steps, and identify the crucial ones to draw sound ecological conclusions. Location : Valloire, French Alps. Taxon: Vascular plants and Fungi. Methods: Using soil eDNA metabarcoding data for plants and fungi from twenty plots sampled along a 1000-m elevation gradient, we tested how the conclusions from three types of ecological analyses: (i) the spatial partitioning of diversity, (ii) the diversity-environment relationship, and (iii) the distance-decay relationship, are robust to data curation steps. Since eDNA metabarcoding data also comprise erroneous sequences with low frequencies, diversity estimates were further calculated using abundance-based Hill numbers, which penalize rare sequences through a scaling parameter, namely the order of diversity q (Richness with q=0, Shannon diversity with q~1, Simpson diversity with q=2). Results: We showed that results from different ecological analyses had varying degrees of sensitivity to data curation strategies and that the use of Shannon and Simpson diversities led to more reliable results. We demonstrated that MOTU clustering, removal of PCR errors and of cross-sample contaminations had major impacts on ecological analyses. Main conclusions: In the Era of Big Data, eDNA metabarcoding is going to be one of the major tools to describe, model and predict biodiversity in space and time. However, ignoring crucial data curation steps will impede the robustness of several ecological conclusions. Here, we propose a roadmap of crucial curation steps for different types of ecological analyses.
Whole Genome Sequencing data
<p>The data corresponds to 3 whole genome sequencing experiments for determining the genotype of three mutant strains of the green microalga <em>C. reinhardtii</em> (mutants TSP1, TSP2 and TSP4). The respective description is to be published in the scientific journal "Fronteirs in Plant Science". </p>
Flesh ID: raw sequence data of the identification of meat source
<p>The data set represent the Fastq sequence files obtained from oxford Nanopore sequencing run of sample contain DNA of nine animal species extracted from meat samples (Horse, Chicken, Turkey, Cattle, Sheep, Duck, Rabbit, Goat and Donkey). </p>
Data from: 346 target gene sequences from Vitaceae for Hyb-Seq
<p>The Vitaceae (the grape family) consists of 16 genera and ca. 950 species. It is best known for the economically important fruit crop -- the grape <i>Vitis vinifera</i>. The deep phylogenetic relationships and character evolution of the grape family have attracted the attention of researchers in recent years. We herein reconstruct the phylogenomic relationships within Vitaceae using nuclear and plastid genes based on the Hyb-Seq approach and test the newly proposed classification system of the family. The five tribes of the grape family, including Ampelopsideae, Cayratieae, Cisseae, Parthenocisseae, and Viteae, are each robustly supported by both nuclear and chloroplast genomic data. The cupular floral disc (raised above and free from ovary at the upper part) is an ancestral state of Vitaceae, with the inconspicuous floral disc as derived in the tribe Parthenocisseae, and the state of adnate to the ovary as derived in the tribe Viteae. The 5-merous floral pattern was inferred to be the ancestral in Vitaceae, with the 4-merous flowers evolved at least two times in the family. The compound dichasial cyme (cymose with two secondary axes) is ancestral in Vitaceae and the thyrse inflorescence (a combination of racemose and cymose branching) in tribe Viteae is derived. The ribbon-like trichome only evolved once in Vitaceae, as a synapomorphy for the tribe Viteae.</p>
Cytosplore-Transcriptomics: a scalable inter-active framework for single-cell RNA sequenc-ing data analysis
<p>Allen institute 10X mouse data (<a href="https://portal.brain-map.org/atlases-and-data/rnaseq/mouse-whole-cortex-and-hippocampus-10x">https://portal.brain-map.org/atlases-and-data/rnaseq/mouse-whole-cortex-and-hippocampus-10x</a>) converted into h5 format including metadata, to be easily uploaded in Cytosplore-Transcriptomics.</p>
Data from: Annotation of pseudogenic gene segments by massively parallel sequencing of rearranged lymphocyte receptor loci
Background: The adaptive immune system generates a remarkable range of antigen-specific T-cell receptors (TCRs), allowing the recognition of a diverse set of antigens. Most of this diversity is encoded in the complementarity determining region 3 (CDR3) of the β chain of the αβ TCR, which is generated by somatic recombination of noncontiguous variable (V), diversity (D), and joining (J) gene segments. Deletion and non-templated insertion of nucleotides at the D-J and V-DJ junctions further increases diversity. Many of these gene segments are annotated as non-functional owing to defects in their primary sequence, the absence of motifs necessary for rearrangement, or chromosomal locations outside the TCR locus. Methods: We sought to utilize a novel method, based on high-throughput sequencing of rearranged TCR genes in a large cohort of individuals, to evaluate the use of functional and non-functional alleles. We amplified and sequenced genomic DNA from the peripheral blood of 587 healthy volunteers using a multiplexed polymerase chain reaction assay that targets the variable region of the rearranged TCRβ locus, and we determined the presence and the proportion of productive rearrangements for each TCRβ V gene segment in each individual. We then used this information to annotate the functional status of TCRβ V gene segments in this cohort. Results: For most TCRβ V gene segments, our method agrees with previously reported functional annotations. However, we identified novel non-functional alleles for several gene segments, some of which were used exclusively in our cohort to the detriment of reported functional alleles. We also saw that some gene segments reported to have both functional and non-functional alleles consistently behaved in our cohort as either functional or non-functional, suggesting that some reported alleles were not present in the population studied. Conclusions: In this proof-of-principle study, we used high-throughput sequencing of the TCRβ locus of a large cohort of healthy volunteers to evaluate the use of functional and non-functional alleles of individual TCRβ V gene segments. With some modifications, our method has the potential to be extended to gene segments in the α, γ, and δ TCR loci, as well as the genes encoding for B-cell receptor chains.
Data from: Small RNA sequencing reveals a novel tsRNA-26576 mediating tumorigenesis of breast cancer
Purpose: As a malignancy that develops from breast tissue, breast cancer has been widely regarded as the most common type of cancer threatening the health of women worldwide. Emerging evidence has demonstrated that tsRNAs might play a vital part in the tumorigenesis and progression of several types of cancers. However, the functions of tsRNAs in breast cancer remain largely unknown. Here, we investigated the functions of tsRNA-26576 in tumorigenesis of breast cancer. Patients and methods: In this study, the tsRNA deregulation states in breast cancer patients (four cancer tissues and four adjacent normal tissues) were evaluated using small RNA sequencing. And then, RT-PCR was used to detected the tsRNA-26576 expression level in breast cancer patients. Results: A total of 263 tsRNAs were identified as significantly differentially expressed, of which 75 were upregulated, and 188 were downregulated. The functional classification through KEGG pathway database illustrated that the most significant pathway enriched by the targets of differentially expressed tsRNAs was the pathway in cancer. Among these differently expressed tsRNAs, we found that tsRNA-26576 was remarkably upregulated in cancer tissue in comparison with adjacent normal tissue. Meanwhile, RT-PCR results verified that tsRNA-26576 expression level was highly upregulated in 10 paired samples from breast cancer patients. Besides, tsRNA-26576 was found to motivate cellular multiplication and migration while suppressing cellular apoptosis in MDA-MB-231 cells. Moreover, mRNA sequencing results showed that several tumor suppressor genes, including FAT4 and SPEN, were upregulated after delivering tsRNA-26576 inhibitor in MDA-MB-231 cells. Conclusion: We found tsRNA-26576 was upregulated in breast cancer tissue, and it could promote the cell growth while inhibite cell apoptosis. Therefore, tsRNA-26576 might serve as a potential clinical therapy target and a predictive marker for breast cancer.
Data from: Whole genome sequence accuracy is improved by replication in a population of mutagenized sorghum.
The accurate detection of induced mutations is critical for both forward and reverse genetics studies. Experimental chemical mutagenesis induces relatively few single base changes per individual. In a complex eukaryotic genome, false positive detection of mutations can occur at or above this mutagenesis rate. We demonstrate here, using a population of ethyl methanesulfonate (EMS) treated Sorghum bicolor BTx623 individuals, that using replication to detect false positive induced variants in next-generation sequencing data permits higher throughput variant detection with greater accuracy. We used a lower sequence coverage depth (average of 7X) from 586 independently mutagenized individuals and detected 5,399,493 homozygous SNPs. Of these, 76% originated from only 57,872 genomic positions prone to false positive variant calling. These positions are characterized by high copy number paralogs where the error-prone SNP positions are at copies containing a variant at the SNP position. The ability of short stretches of homology to generate these error prone positions suggests that incompletely assembled or poorly mapped repeated sequences are one driver of these error prone positions.. Removal of these false positives left 1,275,872 homozygous and 477,531 heterozygous EMS-induced SNPs which, congruent with the mutagenic mechanism of EMS, were greater than 98% G:C to A:T transitions. Through this analysis we generated a database of sequence indexed mutants of Sorghum. This collection contains 4,035 high impact homozygous mutations in 3,637 genes and 56,514 homozygous missense mutations in 23,227 genes. Each line contains, on average, 2,177 annotated homozygous SNPs per genome, including seven likely gene knockouts and 96 missense mutations. The number of mutations in a transcript was linearly correlated with the transcript length and also the G+C count, but not with the GC/AT ratio. Analysis of the detected mutagenized positions identified CG-rich patches, and flanking sequences strongly influenced EMS-induced mutation rates. Our method for detecting false-positive induced mutations is generally applicable to any organism, is independent of the choice of in silico variant-calling algorithm, and is most valuable when the true mutation rate is likely to be low, such as in laboratory induced mutations or somatic mutation detection in medicine.
Data from: Genetic diversity and population structure of Urochloa grass accessions from Tanzania using simple sequence repeat (SSR) markers
Urochloa (syn.—Brachiaria s.s.) is one of the most important tropical forages that transformed livestock industries in Australia and South America. Farmers in Africa are increasingly interested in growing Urochloa to support the burgeoning livestock business, but the lack of cultivars adapted to African environments has been a major challenge. Therefore, this study examines genetic diversity of Tanzanian Urochloa accessions to provide essential information for establishing a Urochloa breeding program in Africa. A total of 36 historical Urochloa accessions initially collected from Tanzania in 1985 were analyzed for genetic variation using 24 SSR markers along with six South American commercial cultivars. These markers detected 407 alleles in the 36 Tanzania accessions and 6 commercial cultivars. Markers were highly informative with an average polymorphic information content of 0.79. The analysis of molecular variance revealed high genetic variation within individual accessions in a species (92%), fixation index of 0.05 and gene flow estimate of 4.77 showed a low genetic differentiation and a high level of gene flow among populations. An unweighted neighbor-joining tree grouped the 36 accessions and six commercial cultivars into three main clusters. The clustering of test accessions did not follow geographical origin. Similarly, population structure analysis grouped the 42 tested genotypes into three major gene pools. The results showed the Urochloa brizantha (A. Rich.) Stapf population has the highest genetic diversity (I = 0.94) with high utility in the Urochloa breeding and conservation program. As the Urochloa accessions analyzed in this study represented only 3 of 31 regions of Tanzania, further collection and characterization of materials from wider geographical areas are necessary to comprehend the whole Urochloa diversity in Tanzania.
Data from: Modular tagging of amplicons using a single PCR for high-throughput sequencing
High-throughput sequencing (HTS) of PCR amplicons is becoming the method of choice to sequence one or several targeted loci for phylogenetic and DNA barcoding studies. Although the development of HTS has allowed rapid generation of massive amounts of DNA sequence data, preparing amplicons for HTS remains a rate-limiting step. For example, HTS platforms require platform-specific adapter sequences to be present at the 5′ and 3′ end of the DNA fragment to be sequenced. In addition, short multiplex identifier (MID) tags are typically added to allow multiple samples to be pooled in a single HTS run. Existing methods to incorporate HTS adapters and MID tags into PCR amplicons are either inefficient, requiring multiple enzymatic reactions and clean-up steps, or costly when applied to multiple samples or loci (fusion primers). We describe a method to amplify a target locus and add HTS adapters and MID tags via a linker sequence using a single PCR. We demonstrate our approach by generating reference sequence data for two mitochondrial loci (COI and 16S) for a diverse suite of insect taxa. Our approach provides a flexible, cost-effective and efficient method to prepare amplicons for HTS.
Data from: Scaling up DNA barcoding - primer sets for simple and cost efficient arthropod systematics by multiplex PCR and Illumina amplicon sequencing
1. The simplicity and cost efficiency of Illumina amplicon sequencing has greatly contributed to the advancement of DNA barcoding and metabarcoding applications. However, current amplicon sequencing based barcoding approaches are usually restricted to short, single-locus fragments, limiting their taxonomic and phylogenetic resolution. 2. Here, we establish a cost efficient and simple multiplex PCR protocol for arthropod systematics by Illumina amplicon sequencing. We introduce primer sets, including several new, generic primers, to reliably amplify nine loci across a wide range of arthropods. Using a diverse collection of arthropod species from 19 orders, we test loci for amplification efficiency and estimate the effect of cross-species amplification bias on taxon recovery from bulk community samples. We then explore the taxonomic and phylogenetic utility of the primer sets, focusing on a dataset of spiders that includes both deep and recent divergences. 3. The set of loci provides good phylogenetic support across a wide taxonomic spectrum, making it a useful addition to COI for resolving lineages within a comparative context. All loci recover sequences for the majority of arthropod taxa in separate PCRs. However, cross-species amplification bias in some primers prevents an exhaustive taxon recovery from bulk community samples. 4. Our protocol makes it possible to generate multilocus datasets for large numbers of arthropod taxa for a fraction of the price and workload of Sanger sequencing. This opens up the possibility for parallel phylogenetic and taxonomic analysis of large collections of arthropods, but also enables rapid exploratory analyses of target lineages. Primers for metabarcoding applications should be carefully evaluated for their performance in bulk community samples and chosen to minimize cross-species amplification bias.
Data from: Phylogenetically driven sequencing of extremely halophilic archaea reveals strategies for static and dynamic osmo-response
Organisms across the tree of life use a variety of mechanisms to respond to stress-inducing fluctuations in osmotic conditions. Cellular response mechanisms and phenotypes associated with osmoadaptation also play important roles in bacterial virulence, human health, agricultural production and many other biological systems. To improve understanding of osmoadaptive strategies, we have generated 59 high-quality draft genomes for the haloarchaea (a euryarchaeal clade whose members thrive in hypersaline environments and routinely experience drastic changes in environmental salinity) and analyzed these new genomes in combination with those from 21 previously sequenced haloarchaeal isolates. We propose a generalized model for haloarchaeal management of cytoplasmic osmolarity in response to osmotic shifts, where potassium accumulation and sodium expulsion during osmotic upshock are accomplished via secondary transport using the proton gradient as an energy source, and potassium loss during downshock is via a combination of secondary transport and non-specific ion loss through mechanosensitive channels. We also propose new mechanisms for magnesium and chloride accumulation. We describe the expansion and differentiation of haloarchaeal general transcription factor families, including two novel expansions of the TATA-binding protein family, and discuss their potential for enabling rapid adaptation to environmental fluxes. We challenge a recent high-profile proposal regarding the evolutionary origins of the haloarchaea by showing that inclusion of additional genomes significantly reduces support for a proposed large-scale horizontal gene transfer into the ancestral haloarchaeon from the bacterial domain. The combination of broad (17 genera) and deep (≥5 species in four genera) sampling of a phenotypically unified clade has enabled us to uncover both highly conserved and specialized features of osmoadaptation. Finally, we demonstrate the broad utility of such datasets, for metagenomics, improvements to automated gene annotation and investigations of evolutionary processes.
Data from: Whole genome amplification and reduced-representation genome sequencing of Schistosoma japonicum miracidia
Background: In areas where schistosomiasis control programs have been implemented, morbidity and prevalence have been greatly reduced. However, to sustain these reductions and move towards interruption of transmission, new tools for disease surveillance are needed. Genomic methods have the potential to help trace the sources of new infections, and allow us to monitor drug resistance. Large-scale genotyping efforts for schistosome species have been hindered by cost, limited numbers of established target loci, and the small amount of DNA obtained from miracidia, the life stage most readily acquired from humans. Here, we present a method using next generation sequencing to provide high-resolution genomic data from S. japonicum for population-based studies. Methodology/Principal Findings: We applied whole genome amplification followed by double digest restriction site associated DNA sequencing (ddRADseq) to individual S. japonicum miracidia preserved on Whatman FTA cards. We found that we could effectively and consistently survey hundreds of thousands of variants from 10,000 to 30,000 loci from archived miracidia as old as six years. An analysis of variation from eight miracidia obtained from three hosts in two villages in Sichuan showed clear population structuring by village and host even within this limited sample. Conclusions/Significance: This high-resolution sequencing approach yields three orders of magnitude more information than microsatellite genotyping methods that have been employed over the last decade, creating the potential to answer detailed questions about the sources of human infections and to monitor drug resistance. Costs per sample range from $50-$200, depending on the amount of sequence information desired, and we expect these costs can be reduced further given continued reductions in sequencing costs, improvement of protocols, and parallelization. This approach provides new promise for using modern genome-scale sampling to S. japonicum surveillance, and could be applied to other schistosome species and other parasitic helminthes
Data from: Estimations of evapotranspiration in an age sequence of Eucalyptus plantations in subtropical China
Eucalyptus species are widely planted for reforestation in subtropical China. However, the effects of Eucalyptus plantations on the regional water use remain poorly understood. In an age sequence of 2-, 4- and 6-year-old Eucalyptus plantations, the tree water use and soil evaporation were examined by linking model estimations and field observations. Results showed that annual evapotranspiration of each age sequence Eucalyptus plantations was 876.7, 944.1 and 1000.7 mm, respectively, accounting for 49.81%, 53.64% and 56.86% of the annual rainfall. In addition, annual soil evaporations of 2-, 4- and 6-year-old were 318.6, 336.1, and 248.7 mm of the respective Eucalyptus plantations. Our results demonstrated that Eucalyptus plantations would potentially reduce water availability due to high evapotranspiration in subtropical regions. Sustainable management strategies should be implemented to reduce water consumption in Eucalyptus plantations in the context of future climate change scenarios such as drought and warming.
Data from: Ontogenetic sequence reconstruction and sequence polymorphism in extinct taxa: an example using early tetrapods (Tetrapoda: Lepospondyli)
Ontogenetic sequence reconstruction is challenging particularly for extinct taxa because of when, where, and how fossils preserve. Different methods of reconstruction exist, but the effects of preservational bias, the applicability of size-independent methods, and the prevalence of sequence polymorphism (intraspecific variation) remain unexplored for paleontological data. Here I compare five different methods of ontogenetic sequence reconstruction and their effects on the detection of sequence polymorphism, using a large collection of the extinct vertebrates Microbrachis pelikani and Hyloplesion longicostatum. The postcranial ossification sequences presented here for those taxa are the first examples known for extinct lepospondyls. Sequences were reconstructed according to skull length, trunk length, increasing number of ontogenetic events, majority-rule consensus, and Ontogenetic Sequence Analysis (OSA). Results generally were in agreement, demonstrating that paleontological data may be used to robustly reconstruct developmental patterns. When reconstructing sequences based on fossils, size-based methods and OSA are more objective and less dependent on preservational bias than other techniques. Apart from the other methods, OSA also allows for statistical analysis of observed and predicted polymorphism. However, OSA requires a large sample size to yield meaningful results, and size-based methods are justified in paleontological studies when sample size is limited by poor preservation. Different methods of reconstruction detected different patterns of sequence polymorphism, although across all methods the magnitude of sequence variation for M. pelikani and H. longicostatum (1.3−3.4%) was within the lower range of values reported for extant vertebrates. Compared with other extinct and extant tetrapods, all sequence reconstruction methods consistently showed that M. pelikani and H. longicostatum exhibit advanced ossification of the pubis and delayed ossification of the scapula. However, the postcranial ossification sequences of these two taxa largely are congruent with those of other tetrapods, suggesting an underlying conservative ancestral pattern that evolved early in tetrapod history.
Data from: Angiosperm phylogeny based on 18S/26S rDNA sequence data: constructing a large dataset using next-generation sequence data
The utility of 18S and 26S in broad phylogenetic analyses has been much maligned due in large part to the low signal in both genes. However, few analyses have employed complete 26S rDNA sequences over a broad range of taxa, and most alignments of the two genes are done de novo, without taking into account the secondary structure of the two rRNA genes. Here we mine next-generation sequence data to compile large matrices (429 taxa) of complete 18S + 26S gene sequences, and we compare both de novo alignment methods with curated alignments done by eye that take into account secondary structure and hard-to-align regions (profile alignments). The combined 18S + 26S topology is overall very similar to recently published gene trees for the angiosperms based on three or more genes. Overall support for the backbone or framework of the combined tree is low (bootstrap support below 50%). Few major clades have bootstrap support above 50%. Most well-supported clades are tip clades (families and orders sensu APG III 2009). Importantly, the 18S + 26S rDNA topology is consistent with current estimates of relationships: the basalmost angiosperms are recovered (Amborellaceae, Nymphaeales, Austrobaileyales), as are most major clades, including Mesangiospermae, eudicots (Eudicotyledoneae sensu Cantino et al. 2007), core eudicots (Gunneridae sensu Cantino et al. 2007), rosids (Rosidae sensu Cantino et al. 2007), asterids (Asteridae sensu Cantino et al. 2007), and Caryophyllales. Most clades recognized at the ordinal level (sensu APG III 2009) are also recovered. However, there are also some unusual placements in the 18S + 26S topology, but none of these receives bootstrap support above 50%. The profile and de novo alignments gave very similar topologies. 18S + 26S trees remain useful sources of data in large combined analyses. This is the first time a large data set of complete 26S gene sequences has been employed at this scale; this gene in particular proved to be useful phylogenetically. Targeted sequencing of 18S/26S rDNA is not advocated here, but given that these regions provide useful phylogenetic information and are abundant in next-generation sequencing runs, we suggest that the data be used rather than discarded.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.