Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
2,848
datasets available to search
ShareScore release 0.9.0
Dataset results
2,848 results for “sequence data”
Data from Readsynth: short-read simulation for consideration of composition-biases in reduced metagenome sequencing approaches
<p><strong>Background</strong></p> <p>The application of reduced metagenomic sequencing approaches holds promise as a middle ground between targeted amplicon sequencing and whole metagenome sequencing approaches but has not been widely adopted as a technique. A major barrier to adoption is the lack of read simulation software built to handle characteristic features of these novel approaches. Reduced metagenomic sequencing (RMS) produces unique patterns of fragmentation per genome that are sensitive to restriction enzyme choice, and the non-uniform size selection of these fragments may introduce novel challenges to taxonomic assignment as well as relative abundance estimates.</p> <p><strong>Results</strong></p> <p>Through the development and application of simulation software, readsynth, we compare simulated metagenomic sequencing libraries with existing RMS data to assess the influence of multiple library preparation and sequencing steps on downstream analytical results. Based on read depth per position, readsynth achieved 0.79 Pearson's correlation and 0.94 Spearman's correlation to these benchmarks. Application of a novel estimation approach, fixed length taxonomic ratios, improved quantification accuracy of simulated human gut microbial communities when compared to estimates of mean or median coverage.</p> <p><strong>Conclusions</strong></p> <p>We investigate the possible strengths and weaknesses of applying the RMS technique to profiling microbial communities via simulations with readsynth. The choice of restriction enzymes and size selection steps in library prep are non-trivial decisions that bias downstream profiling and quantification. The simulations investigated in this study illustrate the possible limits of preparing metagenomic libraries with a reduced representation sequencing approach, but also allow for the development of strategies for producing and handling the sequence data produced by this promising application.</p>
Whole genome sequencing (WGS) data from invasive pine sawfly Diprion similis
<p>Biological introductions are unintended "natural experiments" that provide unique insights into evolutionary processes. Invasive phytophagous insects are of particular interest to evolutionary biologists studying adaptation, as introductions often require rapid adaptation to novel host plants. However, adaptive potential of invasive populations may be limited by reduced genetic diversity—a problem known as the "genetic paradox of invasions". One potential solution to this paradox is if there are multiple invasive waves that bolster genetic variation in invasive populations. Evaluating this hypothesis requires characterizing genetic variation and population structure in the invaded range. To this end, we assemble a reference genome and describe patterns of genetic variation in the introduced white pine sawfly, <em>Diprion</em> <em>similis</em>. This species was introduced to North America in 1914, where it has rapidly colonized the thin-needled eastern white pine (<em>Pinus</em> <em>strobus</em>), making it an ideal invasion system for studying adaptation to novel environments. To evaluate evidence of multiple introductions, we generated whole-genome resequencing data for 64 <em>D</em>. <em>similis</em> females sampled across the North American range. Both model-based and model-free clustering analyses supported a single population for North American <em>D</em>. <em>similis</em>. Within this population, we found evidence of isolation-by-distance and a pattern of declining heterozygosity with distance from the hypothesized introduction site. Together, these results support a single-introduction event. We consider implications of these findings for the genetic paradox of invasion and discuss priorities for future research in <em>D</em>. <em>similis</em>, a promising model system for invasion biology.</p>
Figure 2 in Diagnosability of mtDNA with Random Forests: Using sequence data to delimit subspecies
Figure 2. Illustration of steps in constructing a Random Forests ensemble of classification trees.
Data for "Shear Strain Evolution Spanning the 2020 Mw6.8 Elazığ and 2023 Mw7.8/Mw7.6 Kahramanmaraş Earthquake Sequence along the East Anatolian Fault Zone" manuscript
<p>Data necessary to support the analysis presented in "Shear Strain Evolution Spanning the 2020 Mw6.8 Elazığ and 2023 Mw7.8/Mw7.6 Kahramanmaraş Earthquake Sequence along the East Anatolian Fault Zone" manuscript</p>
Illumina Sequencing Data for "Phocaeicola vulgatus shapes the long-term growth dynamics and evolutionary adaptations of Clostridioides difficile"
<p>Illumina Sequencing Data for "<em>Phocaeicola vulgatus</em> shapes the long-term growth dynamics and evolutionary adaptations of <em>Clostridioides difficile</em>"</p>
Associated code and data for "A Practical Guideline for MicroRNA Sequencing Data Analysis in Chronic Lymphocytic Leukemia (doi: 10.1007/978-1-0716-4290-0_18)".
<p>This deposit contains the data, code, and analysis to recreate the results in the manuscript - Tuulikki Suomela, Liang Zhang, Julio Vera, Heiko Bruns, Xin Lai. A Practical Guideline for MicroRNA Sequencing Data Analysis in Chronic Lymphocytic Leukemia. Methods Mol. Biol., 2883, 403–426. <a href="https://www.researchgate.net/publication/387267721_A_Practical_Guideline_for_MicroRNA_Sequencing_Data_Analysis_in_Chronic_Lymphocytic_Leukemia">https://doi.org/10.1007/978-1-0716-4290-0_18</a>.</p> <p>The pipeline allows users to perform end-to-end analysis of bulk miRNA sequencing data, including quality control of FastQ files, mapping of read counts to miRNA genes using miRBase or Reference genome, quantification of miRNA read counts, differential gene expression analysis using DEseq2, gene set enrichment analysis using curated cancer hallmark gene sets, and identification of miRNA targets.</p> <p>If you have used the code for your research, please cite the original publication. Thank you very much.</p>
A HKU5-related Coronavirus identified and assembled from Short-Read sequencing data of Gossypium Barbadense
<p>This is the complete annotated genome of the Merbecovirus identified from Gossypium Barbadense sequencing data, <a href="https://trace.ncbi.nlm.nih.gov/Traces/sra/?run=SRR5885860">SRR5885860</a></p>
Raw data for "Condensates in RNA repeat sequences are heterogeneously organized and exhibit reptation dynamics"
<p>This is the raw data for the paper "Condensates in RNA repeat sequences are heterogeneously organized and exhibit reptation dynamics".</p> <p>There are 5 directories. Three correspond to the CAG repeats with different length (20, 31 and 47). The "scramble47" directory stores data for the scrambled sequence. "Electrostatics" has data for the electrostatics run (see details in Extended Data Fig. 9).</p> <p>Each directory of CAG contains multiple sub-directories corresponding to different concentrations.</p> <p> </p>
Raw data: multispecies amplicon sequencing (Loera, Studer, and Kölliker, 2021, Molecular Ecology Resources)
<p>Grasslands cover close to two fifths of Earth's land. They provide many ecosystem services related to the maintenance of soil integrity, and the regulation of water, carbon and nitrogen flows. Grasslands constitute the basis for sustainable roughage production for ruminant feeding. In Switzerland, grasslands cover more than 70% of the total agricultural land, which highlights their importance in the domestic food production chains.</p> <p>Plant genetic diversity (PGD), a component of biodiversity, influences ecosystem functioning in grasslands. High levels of grassland PGD are related to resistance against invasive plants and yield stabilization during environmental stress (e.g., drought or frost). The PGD of grasses and legumes —the two most economically relevant plant families found in grasslands, which naturally grow in a wide climate spectrum— harbors valuable genetic resources for forage breeding. Nevertheless, most PGD studies of natural or semi-natural grasslands (i.e., grasslands that are not sown) focus on a single or a few related species. Traditional PGD monitoring methods (e.g., simple sequence repeats, or SSRs) are ill-suited for large-scale, multispecies assessments. This limits our ability to study the ecological effects of grassland PGD, its spatiotemporal patterns, and its significance for grassland management.</p> <p>Looking to provide cost-effective tools for multispecies PGD monitoring in grasslands, we performed a sequence capture assay targeting 611 single-copy nuclear loci, followed by multispecies amplicon sequencing (i.e., amplicon sequencing using primer pairs that can be used in multiple species) on eleven selected loci.</p> <p>Our results indicate that multispecies amplicon sequencing is a cost-effective tool for genetic diversity assessment in grassland plant species. Furthermore, the sequence capture data provides the means to extend the number of multispecies amplicons for further research.</p>
single-nucleus RNA sequencing data from female Aedes aegypti maxillary palp
<p>Single-nucleus RNA sequencing data accompanying Herre*, Goldman* et al. (2022), "Non-Canonical Odor Coding in the Mosquito" (https://doi.org/10.1016/j.cell.2022.07.024)</p> <p>For further analysis see: https://github.com/VosshallLab/Younger_Herre_Vosshall2020/tree/main/snRNAseq_SupplementaryData</p> <p>For raw sequencing files see NCBI BioProject: PRJNA794050</p>
Data for Coral Growth and Sequences for Cyclin-E and G3P Primers for Orbicella faveolata
<p>Cyclin-E and glyceraldehyde 3-phosphate dehydrogenase primer sets were constructed using Primer3 from an annotated<br> transcriptome (Polato et al., 2011). For each sample and gene, reactions were performed in triplicate on a Step One Plus qPCR<br> machine (Applied Biosystems, Waltham, MA), using cDNA of <em>Orbicella faveolata</em> (Ellis & Solander, 1786) as a template. Singlepeak melt curve analysis was performed to test for nonspecific amplification products, but limited sample prevented primer<br> efficiency analysis. To verify primer specificity, the two primers were tested on cDNA from <em>Casseopia xamachana</em>, with no<br> amplification being observed.</p>
Data from: Diversity of land snail tribe Helicini (Gastropoda: Stylommatophora: Helicidae): where do we stand after 20 years of sequencing mitochondrial markers?
<p>Sequences of mitochondrial genes revolutionized the understanding of animal diversity and continue to be an important tool in biodiversity research. In the tribe Helicini, a prominent group of the western Palaearctic land snail fauna, mitochondrial data accumulating since the 2000s helped to newly delimit genera, inform species-level taxonomy, and reconstruct past range dynamics. We combined the published data with own unpublished sequences and provide a detailed overview of what they revealed about the diversity of the group. The delimitation of <i>Helix</i> is revised by placing <i>Helix godetiana</i> back in the genus and new synonymies are suggested within the genera <i>Codringtonia</i> and<i> Helix</i>. The spatial distribution of intraspecific mitochondrial lineages of several species is shown for the first time. Comparisons between species reveal considerable variation in distribution patterns of intraspecific lineages, from broad postglacial distributions to regions with a fine-scale pattern of allopatric lineage replacement. To provide a baseline for further research and information for anyone re-using the data, we thoroughly discuss the gaps in the current dataset, focusing on both taxonomic and geographic coverage. Thanks to the wealth of data already amassed and the relative ease with which they can be obtained, mitochondrial sequences remain an important source of information on intraspecific diversity over large areas and taxa.</p>
Whole genome sequence data of Lactiplantibacillus plantarum IMI507027
<p>The present data files are the annotation output from the whole genome sequencing of the strain Lactiplantibacillus plantarum IMI507027. </p>
Raw Data for the article: A retrospective molecular epidemiological scenario of carbapenemase-producing Klebsiella pneumoniae clinical isolates in a Sicilian transplantation hospital shows a swift polyclonal divergence among sequence types, resistome and virulome
<p>In this work, we assessed and characterized the epidemiological scenario of carbapenem-resistant Klebsiella pneumoniae strains (CR-Kp) at IRCCS-ISMETT, a transplantation hospital in Palermo, Italy, from 2008 to 2017. A total of 288 K. pneumoniae clinical isolates were selected based on their resistance to carbapenems. Molecular characterization was also done in terms of the presence of virulence and resistance genes. All patients were inpatients from our facility and clinical isolates were collected from several sources, either from infection or colonization cases. We observed that, in agreement with the Italian epidemiological scenario, initially only ST258 and ST512 clade II (but not from clade I) were identified from 2008 to 2011. From 2012 onwards, other STs have been observed, including the clinically relevant ST101 and ST307, but also others not previously observed in other Italian health settings, such as ST220 and ST753. The presence of genes involved in resistance and virulence was confirmed, and a heterogeneous genetic resistance profile throughout the years was observed. Our work highlights that resistance genes are rapidly disseminating between different and novel K. pneumoniae clones which, combined with resistance to multiple antibiotics, can derive into more aggressive and pathogenic multidrug-resistant strains of clinical importance. Our results stress the importance of continuous surveillance of CR Enterobacterales in health facilities so that novel STs carrying resistance and virulence genes that may become increasingly pathogenic can be identified and adequate therapies to adopted to avoid their dissemination and derived pathologies.</p>
Raw Data for the article: Liver Transplantation for Unresectable Intrahepatic Cholangiocarcinoma: The Role of Sequencing Genetic Profiling
<p>Intrahepatic cholangiocarcinoma (iCCA) is a rare and aggressive primary liver tumor, characterized by a range of different clinical manifestations and by increasing incidence and mortality rates even after curative treatment with radical resection. In recent years, growing attention has been devoted to this disease and some evidence supports liver transplantation (LT) as an appropriate treatment for intrahepatic cholangiocarcinoma; evolving work has also provided a framework for better understanding the genetic basis of this cancer. The aim of this study was to provide a clinical description of our series of patients complemented with Next-Generation Sequencing genomic profiling. From 1999 to 2021, 12 patients who underwent LT with either iCCA or a combined hepatocellular and cholangiocellular carcinoma (HCC-iCCA) were included in this study. Mutations were observed in gene activating signaling pathways known to be involved with iCCA tumorigenesis (KRAS/MAPK, P53, PI3K-Akt/mTOR, cAMP, WNT, epigenetic regulation and chromatin remodeling). Among several others, a strong association was observed between the Notch pathway and tumor size (point-biserial <em>rho<sub>pb</sub></em> = 0.93). Our results are suggestive of the benefit potentially derived from molecular analysis to improve our diagnostic capabilities and to devise new treatment protocols, and eventually ameliorate long-term survival of patients affected by iCCA or HCC-iCCA.</p>
Extended data for "The need to reassess single-cell RNA sequencing datasets: the importance of biological sample processing"
<p>Extended data for "The need to reassess single-cell RNA sequencing datasets: the importance of biological sample processing"</p>
Processed MinION genome sequencing data for strains ILHA G3AG5 and ILHA G3AA5
<p>Strain G3AA5 datasets include:</p> <p>Galaxy2750, Galaxy2754, Galaxy3522, Galaxy3523, Galaxy3524, Galaxy3541</p> <p> </p> <p>Strain G3AG4 datasets include:</p> <p>Galaxy2404, Galaxy2408, Galaxy3517, Galaxy3518, Galaxy3519, Galaxy3539</p>
Morphometric, gametic and parasitological data of clam Ameghynomia antiqua and DNAmt sequences of parasite
<p><span>The clam </span><em><span>Ameghynomia antiqua</span></em><span> is a highly important resource for fisheries due to its high catches volume. It is the bivalve mollusk with the highest fisheries landings from natural beds on the Pacific coast of southern South America; however, studies of the reproductive conditions of this species are scarce and date back many years. The object of the present work was to evaluate the reproductive characteristics of the species, analysing its gametogenic and gonadal cycle, and reproductive indices, in fishery locations that present the natural beds with the highest fisheries catches, as well as parasite loads in the species. The gonads of the individuals were sampled monthly over a year and classified into one of three states called: "in development", "ripe" and "spawned". Synchrony between the sexes was observed in the indicators of the Gonadosomatic Index and Condition Index in each of the locations, although no synchrony was observed between locations. In the gametogenic cycle, the "ripe" state was observed in females in spring-summer, followed by rapid recovery to new development of the gonads; in males the "ripe" state was observed throughout the year. It was observed that males entered the "spawned" state one month ahead of females. The presence of digenean parasites in the state of metacercariae was detected in the gonads and mantle. No significant differences were found in the prevalence or intensity of infection when analysed by sex and month. The metacercariae were identified, by sequencing of three DNA regions, as belonging to the clade shared by species of the genus <em>Parvatrema</em> and close to the <em>Gymnophalloides</em>; both these genera belong to the family Gymnophallidae of the superclass Digenea. Infection was observed to reduce the gonadal tissue, in some cases causing castration. This is the first record of the presence of these parasites of <em>A. antiqua</em>, with genetic identification at genus level. These results are relevant for act proper management of this resource, which is important for fishing.</span></p>
Data from: Benchmarking ultra-high molecular weight DNA preservation methods for long-read and long-range sequencing
<p>Studies in vertebrate genomics require sampling from a broad range of tissue types, taxa, and localities. Recent advancements in long-read and long-range genome sequencing have made it possible to produce high-quality chromosome-level genome assemblies for almost any organism. However, adequate tissue preservation for the requisite ultra-high molecular weight DNA (uHMW DNA) remains a major challenge. Here we present a comparative study of preservation methods for field and laboratory tissue sampling, across vertebrate classes and different tissue types. We find that no single method is best for all cases. Instead, the optimal storage and extraction methods vary by taxa, by tissue, and by down-stream application. Therefore, we provide sample preservation guidelines that ensure sufficient DNA integrity and amount required for use with long-read and long-range sequencing technologies across vertebrates. Our best practices generate the uHMW DNA needed for the high-quality reference genomes for Phase 1 of the Vertebrate Genomes Project (VGP), whose ultimate mission is to generate chromosome-level reference genome assemblies of all ~70,000 extant vertebrate species.</p>
Simulated exome-sequencing data for a family study of lymphoid cancer
<p>This repository contains all the data files for a simulated exome-sequencing study of 150 families ascertained to contain at least four members affected with lymphoid cancer.</p> <p>The simulated data can be found in the files section below. The files are:</p> <ol> <li>SLiM_output.txt - contains the SLiM-simulated, exome-wide, SNV data generated under an American-admixture demographic model, for the American-admixed sub-population only.</li> <li>SLiM_output_chr8&9.txt - contains the SLiM-simulated data above for all source populations as well as the American-admixed sub-population, but only for chromosomes 8 and 9.</li> <li>sample_info.txt - contains pedigree information of all the disease-affected individuals and individuals connecting them along a line of descent, for all 150 ascertained pedigrees.</li> <li>Genotypes.zip - a zipfile that contains 22 text files of genotypes for each chromosome. The genotypes are for simulated single-nucleotide variants on the exome and are in gene-dosage format. </li> <li>SNVmaps.zip - a zipfile that contains 22 text files giving the single-nucleotide variant information for each chromosome. </li> <li>familial_cRV.txt - contains the familial causal rare variants for all 150 ascertained pedigrees.</li> <li>study_peds.txt - contains the 150 pedigrees ascertained to contain four or more relatives affected with lymphoid cancer.</li> <li>PLINKfiles.zip - a zipfile that contains PLINK .fam, .bim and .bed files for all 22 of the chromosomes.</li> </ol> <p>All the scripts used to generate these data can be found in the GitHub repository archived at <a href="https://zenodo.org/record/6505385">https://zenodo.org/record/6505385</a> .</p> <p>We have also uploaded one intermediate .Rdata file, Chromwide.Rdata, to save the user substantial time when running the associated RMarkdown script for the simulation. We recommend loading Chromwide.Rdata into your R work-space rather than generating it from scratch.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.