Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
2,848
datasets available to search
ShareScore release 0.9.0
Dataset results
2,848 results for “sequence data”
Vector sequences in early WIV SRA sequencing data of SARS-CoV-2 inform on a potential large-scale security breach at the beginning of the COVID-19 pandemic
<p>DESCRIPTION</p> <p>Sequences identified as Influenza A virus, Spodoptera frugiperda rhabdovirus and Nipah henipavirus have been previously identified within the early HiSeq 1000 and HiSeq 3000 sequencing data of SARS-CoV-2, SRR11092059,SRR11092060,SRR11092061 and SRR11092062, and were being used to support the hypothesis that a "simultaneous outbreak of multiple zoonotic viruses" have happened in the Huanan Seafood market. https://doi.org/10.31219/osf.io/s4td6</p> <p>However, a closer examination of these sequences revealed that they were not sequences of actual wild viruses, but were in stead fragments left behind from PCR products and cloning vectors harboring both cDNA clones and infectious clones of such viruses, with evidence of viral sequences being joined directly to DNA sequences of vector and non-human origin within the same short reads.</p> <p>Here are the vector sequences and PCR product-like sequences recovered from the earliest WIV SRA sequencing data of Human SARS-CoV-2 from dataset SRR11092059,SRR11092060,SRR11092061,SRR11092062.</p> <p>Sequences associated with Vectors and PCR products from 3 distinct viral species have been obtained: The 3'-end of a Nipah Henipahvirus with fusion to a Hepatitis D virus Ribozyme, a T7 terminator and a Tetracycline resistance gene, The 5'-end of the same Nipah Henipahvirus with fusion to sequences found in diverse vectors, A complete vector genome encoding the HA gene of Influenza A virus subtype H7N9 under a CMV promoter and a bgH polyA terminator, and 221 Contiguous sequences corresponding to the Spodoptera frugiperda rhabdovirus reference genome fused to sequences that were homologous to multiple Plastid sequences and Notably Mitochondrial sequences of Rodents.</p> <p>As sequences corresponding to a rescued infectious clone of a BSL-4 organism (Nipah Henipahvirus) were found in sample sequences that supposedy represents patient samples that were obtained from Hospital ICU and sequenced in a pathogen diagnosis laboratory (which is separate from the Virology Research laboratory which is implied by the context of an Infectious Clone of such an organism, evident by the 3'-HDV ribozyme and T7 terminator fused directly to the 3'-terminus of the Nipah Henipahvirus reads), The discovery of artifact-containing sequences of at least 3 different pathogen species that are phylogenetically and methodologically distinct from each other in samples that were supposedly submitted by a laboratory that is Separate from the virological research laboratories that could have hosted such clone sequences imply extensive crosstalk and cross-contamination between the various laboratories within the Wuhan Institute of Virology, which includes at least one BSL-4 laboratory with evidence of containment breach of a BSL-4 organism and it's subsequent introduction into RNA-seq samples that were processed by a laboratory of distinct and separate purposes than the basic virological research evidenced by the Infectious Clone of the Hipah Henipahvirus.</p> <p>Such a discovery therefore likely imply a major security breach happening within the Wuhan institute of Virology at the time when the first sequences of SARS-CoV-2 was sampled and sequenced, which have important implications on the origins of the SARS-CoV-2 virus itself.</p> <p>METHODS</p> <p>The metagenomic sequencing datasets, SRR11092059,SRR11092060,SRR11092061 and SRR11092062 were first analyzed using the NCBI phylogenetic analysis tool, which identified viral sequences that is not related to SARS-CoV-2 itself. These include Influenza A virus (IAV, subtype H7N9), Spodoptera frugiperda rhabdovirus and Nipah Henipahvirus.</p> <p>The datasets were then subjected to BLAST search using MEGABLAST against the reference sequences of such viruses to verify the existence of the viral sequences and determine the exact sybtype of such viruses and the closest sequences on GenBank that corresponds to the reads. There seuqences are MH926031.1 for the Spodoptera frugiperda rhabdovirus, KY199425.1 for the Influenza A virus and AY988601.1 for the Nipah Henipahvirus.</p> <p>A second round BLAST analysis with these identified sequences were then performed, which unexpectedly revealed numerous reads corresponding to Cloning vectors and non-human Mitochondrial and Plastid sequences being fused directly to the sequences of the identified viral species. Reads were then downloaded and subjected to assembly using the CAP3 sequence assembly program and the EGASSEMBLER tool. Contig sequences were then queried against the NCBI nr/nt database which unanimously identified the original sample sequences as viral sequences inserted into cloning vectors.</p> <p>The complete sequence of the Influenza A virus Haemagluttinin (HA) gene clone was obtained from SRR11092061,SRR11092062 using multiple rounds of BLAST search and sequence assembly expansion on the existing vector-virus junction contigs, and a partial sequence corresponding the 3'-end of Nipah Henipahvirus AY988601.1 fused to a 3'-HDV ribozyme, T7 terminator and a Tet resistance gene was obtained from SRR11092059. In addition, 221 Contig sequences corresponding to the Rhabdovirus MH926031.1 fused to Chloroplast sequence MN524635.1 and Rodent Mitochondrial sequence MT241668.1 have been recovered from SRR11092061.</p> <p>We then performed a BLAST search using the identified vector sequences on SRR11092059,SRR11092060,SRR11092061 and SRR11092062, which confirms the existence of these two vetor sequences in all 4 datasets.</p>
Data for: Faster rates of molecular sequence evolution in reproduction-related genes and in species with hypodermic sperm morphologies
<p>This repository contains a record of analysis scripts and sequence alignments used for the analyses presented in the manuscript.</p> <p>Some of the R scripts depend on supplementary tables associated with the manuscript.</p>
Raw NGS data for the study 'Spouse-to-spouse Transmission and Evolution of Hypervariable Region 1 and 5’ Untraslated Region of Hepatitis C Virus Analyzed by Next-generation Sequencing'
<p>This file contains the original next-generation sequencing data (raw sequences in fastq format) which were analyzed in the study titled: "Spouse-to-spouse Transmission and Evolution of Hypervariable Region 1 and 5’ Untraslated Region of Hepatitis C Virus Analyzed by Next-generation Sequencing".</p> <p> </p> <p> </p>
Sequence data for the article "Targeted reduction of highly abundant transcripts with pseudo-random primers"
<p>Sequence data (Illumina MiSeq runs) for the article "Targeted reduction of highly abundant transcripts with pseudo-random primers". File names indicate unique run identifiers. In the manuscript, shorter names are used:</p> <ul> <li>NC12: 140602_M00528_0019_000000000-A88YD</li> <li>NC17: 140918_M00528_0047_000000000-A8GHU</li> <li>NC22b: 141105_M00528_0062_000000000-AAPC2</li> <li>NCki: 140207_M00528_0069_000000000-A5TY9</li> </ul> <p> </p>
FIGURES 34 37. Lamyctes hellyeri n in A new blind Lamyctes (Chilopoda: Lithobiomorpha) from Tasmania with an analysis of molecular sequence data for the Lamyctes Henicops Group
FIGURES 34 37. Lamyctes hellyeri n. sp. QVMAG 23: 23048, female, pretarsus of leg 14, scales 10 m. 34 36, anterior, posterior, and ventral views; 37, detail of lateral pore and ornament on scutes of main claw.
FIGURES 11 17. Lamyctes hellyeri n in A new blind Lamyctes (Chilopoda: Lithobiomorpha) from Tasmania with an analysis of molecular sequence data for the Lamyctes Henicops Group
FIGURES 11 17. Lamyctes hellyeri n. sp. 11, 14 17, QVMAG 23: 23046, female. 11, anterior part of head shield and basal part of antennae, scale 100 m; 14, sensilla on dorsal side of antenna, scale 10 m; 15 16, antennal articles, dorsal side, scales 50 m; 17, cephalic pleurite with Tömösváry organ, scale 50 m. 12 13, QVMAG 23: 23047, female. 12, ventral view of clypeus and labrum, scale 100 m; 13, labral midpiece and inner parts of sidepieces, scale 30 m.
FIGURES 1 4 in A new blind Lamyctes (Chilopoda: Lithobiomorpha) from Tasmania with an analysis of molecular sequence data for the Lamyctes Henicops Group
FIGURES 1 4. Lamyctes coeculus (Brölemann). 1, 3, AM KS 57961, female, Mellong Range, NSW, Australia. 2, 4, MCZ DNA 100472, female, Cerro San Javier, Tucumán, Argentina. 1 2, ventral view of head, scales 100 m; 3 4, dental margin of maxillipede coxosternite, scales 50 m.
FIGURES 18 25. Lamyctes hellyeri n in A new blind Lamyctes (Chilopoda: Lithobiomorpha) from Tasmania with an analysis of molecular sequence data for the Lamyctes Henicops Group
FIGURES 18 25. Lamyctes hellyeri n. sp. QVMAG 23: 23046, female. 18, ventral view of maxillipede, scale 100 m; 19 20, dental margin of maxillipede coxosternite, scales 50 m, 10 m; 21, tarsus and claw of second maxilla, scale 50 m; 22, distal part of tarsus and claw of second maxilla, scale 10 m; 23, coxal projections and telopods of first maxillae, scale 50 m; 24, first maxillae, scale 100 m; 25, plumose setae on inner margins of telopods of first maxillae, scale 10 m.
Graphing and tabulating next-generation sequencing and genotyping data
<p>Making figures and tables for publication. Each zip archive contains input data, shell script to initiate and log R script, one R script for generating several graphs and tables, and the output graphs and tables themselves.</p> <p>Data was generated by whole-genome resequencing of 22 individual D.melanogaster from Sussex-LHM population and 2 from the Sussex RG line, followed by read-mapping, then genotyping with Haplotype Caller and Genomestrip.</p> <p>Locations for raw data, code, logs, extended QC data:</p> <p>Sequence reads NCBI SRA268956</p> <p>NCBI dbSNP https://www.ncbi.nlm.nih.gov/projects/SNP/snp_viewBatch.cgi?sbid=1062461</p> <p>NCBI dbVar accession number pre-release nstd134</p> <p> </p> <p>The pre-print manuscript for this data is available on biorxiv: "Whole genome resequencing of a laboratory-adapted Drosophila melanogaster population sample" http://biorxiv.org/content/early/2016/10/17/081554 doi: http://dx.doi.org/10.1101/081554</p> <p> </p>
Read-mapping for next-generation sequencing data (Drosophila melanogaster)
<p>Code, logs and quality-control data for whole-genome resequencing of Sussex-LH<sub>M</sub> and RG <em>Drosophila melanogaster</em>.</p> <p>Mapping code is in the archive lhm_mapping_scripts.zip</p> <p>Mapping logs are in the in the archive lhm_mapping_logs.zip</p> <p>Other zip archives contain the quality control data.</p> <p>The pre-print manuscript for this data is available on biorxiv: "Whole genome resequencing of a laboratory-adapted Drosophila melanogaster population sample" http://biorxiv.org/content/early/2016/10/17/081554 doi: http://dx.doi.org/10.1101/081554</p>
Read-mapping for next-generation sequencing data (Wolbachia)
<p>Code, log files and QC data. NCBI SRA accession number SRP091004. Note that sequencing Wolbachia was not a central aim of the project, and was undertaken in order to maximise the amount of information that could be extracted from the raw genome sequence data targetted at the fruit-fly host.</p> <p>The pre-print manuscript for this data is available on biorxiv: "Whole genome resequencing of a laboratory-adapted Drosophila melanogaster population sample" http://biorxiv.org/content/early/2016/10/17/081554 doi: http://dx.doi.org/10.1101/081554</p>
Genotype reproducibility testing in next-generation sequencing data
<p>Code, log and results summary for testing the reproducibility of genotypes with three pairs of hemiclones in the Sussex LH<sub>M </sub><em>D.melanogaster </em>population sample. Discovery and genotyping of genomic sequence variants was done using GATK HaplotypeCaller, and Genomestrip. Numerical comparison of genotype calls within each pairs of hemiclone individuals was performed using GATK GenotypeConcordance.</p> <p> </p> <p>The pre-print manuscript for this data is available on biorxiv: "Whole genome resequencing of a laboratory-adapted Drosophila melanogaster population sample" http://biorxiv.org/content/early/2016/10/17/081554 doi: http://dx.doi.org/10.1101/081554</p>
16S rRNA sequences data set of Dicronocephalus species in this study.
Explanation note: This 16S rRNA data includes 46 individual sequences of the examined Dicronocephalus species in this study.
Data from: CIDER-Seq: unbiased virus enrichment and single-read, full length genome sequencing
<p>Raw and finished sequence data produced in the study: </p> <p>Mehta D, Hirsch-Hoffmann M, Patrignani A, Gruissem W, Vanderschuren H (2017) CIDER-Seq: unbiased virus enrichment and single-read, full length genome sequencing. <em><strong>bioRxiv</strong></em>. doi: https://doi.org/10.1101/168724</p>
Research data supporting "Sequence-Dependent Self-Assembly and Structural Diversity of Islet Amyloid Polypeptide-Derived β-Sheet Fibrils"
<p>Research data supporting the publication:</p> <p>Wang, S.-T. et al., 2017, Sequence-Dependent Self-Assembly and Structural Diversity of Islet Amyloid Polypeptide-Derived β-Sheet Fibrils, ACS Nano, http://dx.doi.org/10.1021/acsnano.7b02325</p>
Data for methylome sequencing: Enriching and Profiling Methylomes for Tumor Classification and Liquid Biopsies
<p>We benchmarked and demonstrated the versatility of FLEXseq (Fragment Ligation EXclusive methylation sequencing) across different sample types: genomic DNA from the K562 (leukemia) cell line, DNA mix-in titrations of four immune cell types (B cells, T cells, monocytes, and neutrophils), DNA titrations of three cancer cell lines (breast invasive carcinoma [BRCA], colon adenocarcinoma [COAD], and glioblastoma [GBM]) mixed with those four immune cell mixtures separately, input titrations of cell-free (cf) DNA from one plasma sample and DNA from formalin-fixed paraffin-embedded (FFPE) tissues, cfDNA from 106 cerebrospinal fluids (CSF) and 42 other body fluids, and DNA from 37 FFPE tissues.</p> <p>We sequenced all the samples mentioned above using FLEXseq. Paired-end reads were quality and length trimmed with cutadapt version 3.5, and all high-quality sequencing reads were then aligned to the hg38 reference genome using Bismark v0.23.0. We then filtered out reads with unmethylated cytosine in the non-CpG context with filter_non_conversion function. Next, we used the bismark_methylation_extractor function to extract the methylation calls (removing single-nucleotide polymorphisms [SNP]).</p> <p>We also used the bam2pat function from wgbs_tools, to convert bam files into .pat files for deconvolution, keeping reads covering at least three CpG sites. The .pat files preserve fragment-level data and were de-identified by removing SNPs using the mask_pat function. </p> <p>We used CNVkit (v0.9.10) to analyze and visualize genome-wide copy numbers. Our inputs into CNVkit were Bismark/Bowtie 2 aligned BAM files deduplicated by Bismark based on end positions and fragment lengths. We then generated log2copy ratio plots for all body fluid and FFPE samples based on the pooled reference and visualized them across all bins using the DNAcopy R package.</p> <p> </p>
Raw sequencing data for studying the colonization of soil communities after glacier retreat
<p>Glaciers show a pattern of retreat at the global scale. Deglaciated areas are exposed and colonized by multiple organisms, but lack of global studies hampers a complete understanding of the future of these ecosystems. Until now, the complete reconstruction of soil communities was hampered by the complex identification of organisms, thus analyses at broad geographical and taxonomic scale have been so far impossible. The dataset used for this study represents the assemblages of Bacteria, Mycota, Eukaryota, Collembola (springtails), Oligochaeta (Earth worms), Insecta, Arthropoda and Vascular Plants obtained using environmental DNA (eDNA) metabarcoding. eDNA was extracted from soil samples collected from multiple glacier forelands representative of some of the main mountain chains of Europe, Asia, the Americas and Oceania. We investigated chronosequences of glacier retreat (i.e., the chronological sequence of specific geomorphological features along deglaciated areas for which the date of glacier retreat is known) ranging from recent years to the Little Ice Age (~1850). We used this newly assembled global DNA metabarcoding dataset to obtain a complete reconstruction of community changes in novel ecosystems after glacier retreat. Information on assemblages can be then combined with analyses of soil, landscape and climate to identify the drivers of community changes.</p>
Data Set for the Journal Article "Automated Preparation of Nanoscopic Structures: Graph-Based Sequence Analysis, Mismatch Detection, and pH-Consistent Protonation with Uncertainty Estimates"
<p>This repository containes the data generated by ASAP and discussed in the journal article [Csizi, K.-S. and Reiher, M., 2023, arXiv:2307.16344], including Cartesian coordinates of training and test set molecules, and MD trajectories. </p>
Fig. 3 in Relationships Of The Heteronchocleidids (Heteronchocleidus, Eutrianchoratus And Trianchoratus) As Inferred From Ribosomal Dna Nucleotide Sequence Data
Fig. 3. Bayesian consensus tree for the anabantoids, channids and catfishes (silurids, bagrids, clariids) obtained using partial Cytochrome b sequences with cyprinids as outgroup. The heteronchocleidids genera present on the anabantoids and channids are shown with their geographical areas. Values shown at each node refer to Bayesian posterior probabilities. (*refer to Table 3 for names used in GenBank).
Fig. 2. Bayesian consensus tree generated from partial 28S in Relationships Of The Heteronchocleidids (Heteronchocleidus, Eutrianchoratus And Trianchoratus) As Inferred From Ribosomal Dna Nucleotide Sequence Data
Fig. 2. Bayesian consensus tree generated from partial 28S rDNA sequences (D1 domain) with Diplectanum spp. and Gyrodactylus spp. as outgroups. Values shown at each node refer to Bayesian (BI) posterior probabilities/maximum likelihood (ML) percentages of the bootstrap values with 100 replicates. Bootstrap values lower than 50 are given as dashes (-).
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.