Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
2,848
datasets available to search
ShareScore release 0.9.0
Dataset results
2,848 results for “sequence data”
Neutralization Data and Aligned ENV Sequences for Predicting Antibody Affinities using Artificial Neural Networks
<p>Sample file with neutralization data (IC<sub>50</sub>) for different antibodies and viral strains, adapted from J. Huang, G. Ofek, L. Laub, M. K. Louder, N. A. Doria-Rose, N. S. Longo, H. Imamichi, R. T. Bailer, B. Chakrabarti, S. K. Sharma, S. M. Alam, T. Wang, Y. Yang, B. Zhang, S. A. Migueles, R. Wyatt, B. F. Haynes, P. D. Kwong, J. R. Mascola, and M. Connors, “Broad and potent neutralization of HIV-1 by a gp41-specific human antibody.,” <em>Nature</em>, vol. 491, no. 7424, pp. 406–12, Nov. 2012.</p> <p> </p> <p>Aligned ENV sequences downloaded from the HIV Sequence Database (www.hiv.lanl.gov/content/sequence/HIV/mainpage.html). There are 4907 sequences and the alignment length is 1369.</p>
Sequence data for the article "Whole transcriptome analysis of thousands of FACS-sorted single cells with the single cell nanoCAGE protocol" - single cells dataset
<p>Sequence data (Illumina MiSeq runs) for the article "Whole transcriptome analysis of thousands of FACS-sorted single cells with the single cell nanoCAGE protocol". dataset of 2300 single cells. File names indicate unique sequencing runs. In the manuscripts, the informations about cell lines are found in the Supplemental Table 1. </p>
Sequence data for the article "Whole transcriptome analysis of thousands of FACS-sorted single cells with the single cell nanoCAGE protocol" - Protocol optimization
<p>Sequence data (Illumina MiSeq runs) for the article "Whole transcriptome analysis of thousands of FACS-sorted single cells with the single cell nanoCAGE protocol". Optimization of the protocol. Files names indicate unique run identifiers. In the manuscript, the link between unique run identifiers, cells and purpose of the experiment is found in the Supplemental Table 1. </p>
MRI raw data: RF-spoiled radial FLASH sequence without cardiac gating and during free breathing
<p>RF-spoiled radial FLASH sequence without cardiac gating and during free breathing<br> Sequence parameters:<br> TR 2.0ms<br> TE 1.3 ms<br> Flip angle 8°<br> Matrix Size 128x128 (two-fold oversampling, resulting in 256 sample points for each radial spoke)<br> Radial Spokes: 125<br> In Plane Resolution 2mmx2mm<br> Slice Thickness 8mm<br> 3T System<br> 32 channel body array coil, compressed to 12 virtual channels using svd based coil compression<br> ote: This data set was acquired with a radial acquisition. The corresponding k-Space trajectory is also included (256x125 matrix k). The Non-Uniform-FFT Toolbox is needed for reconstruction of this data set.<br> Florian Knoll (florian.knoll@tugraz.at) Date: 2.2.2011</p>
Massively parallel sequencing data of the HIV-1 pol region generated from the plasma of therapy-naïve chronically infected Brazilian blood donors
<p>The submitted massively parallel sequencing (MPS) data were partial data from the pol region of HIV-1 plasma viruses. Samples were obtained from 18 therapy-naive HIV-1 Brazilian blood donors with longstanding infection. Illumina ultra-deep sequencing technology (MiSeq platform) was used to generate the sequences. </p>
Cross-disease integration of single-cell RNA sequencing data from lung myeloid cells reveals TAM signature in in vitro model
<p>Single cells from a 3D human cell-based model comprising tumor cell line-derived spheroids, cancer-associated fibroblasts and primary monocytes were dissociated and analyzed using scRNAseq. 4 monocyte donors were used in the 3D model, and 3 monocyte donors were used for 2D differentiation of macrophages.</p>
Grain size and clay mineralogy data of the Esplugafreda sequence (Spain) with stable carbon, oxygen and clumped isotope data of soil carbonates
<p>This datasets contains 1) grain size distribution data of mudstone paleosols of the Esplugafreda sequence (Esplugafreda and Claret Formations) measured by laser diffraction, and 2) clay mineralogy measured by powder X-ray diffraction, as well as 3) stable carbon, oxygen and clumped (D47) isotope compositions of soil carbonates. The Esplugafreda sequence is found in the Tremp-Graus Basin in the southern forefront of the Pyrenees and consists of continental sediments formed in a coastal alluvial setting during the late Paleocene and early Eocene.</p> <p>In addition, a proxy dataset of late Paleocene and PETM continental temperatures of the northern hemisphere is also included.</p>
ParaMask, a new method to identify multicopy genomic regions, corrects major biases in whole-genome sequencing data. Additional Datasets.
<p>Data supporting the main figures of the "ParaMask, a new method to identify multicopy genomic regions, corrects major biases in whole-genome sequencing data" manuscript and a copy of the ParaMask software and scripts for analysis, and SV calls from longreads. README files are included.</p>
Supplementary data 'Mitochondrial genome sequence of the protist Ancyromonas sigmoides Kent, 1881 (Ancyromonadida) from the Sugluk Inlet, Hudson Strait, Nunavik, Québec'
<p>Fasta file with the transcripts obtained from RNAseq sequencing of Ancryomonas sigmoides (kmer 35) and file with datamining results.</p><p>Fasta file of the transcript matching the cox1 gene.</p><p>Scaffolds of the kmer 85 assembly of the genomic data of Ancryomonas sigmoides.</p><p>Databse used for datamining (fasta file)</p>
Next-generation Sequencing Data Associated with "Genome Editing Outcomes Reveal Mycobacterial NucS Participates in a Short-Patch Repair of DNA Mismatches"
Open the record for dataset details and reuse information.
Raw sequencing data and ngsfilters for snow track eDNA samples
<p>Continued advancements in environmental DNA (eDNA) research have made it possible to access intraspecific variation from eDNA samples, opening new opportunities to expand non-invasive genetic studies of wildlife populations. However, the use of eDNA samples for individual genotyping, as typically performed in non-invasive genetics, still remains elusive. We present the first successful individual genotyping of eDNA obtained from snow tracks of three large carnivores: brown bear (<em>Ursus</em> <em>arctos</em>), European lynx (<em>Lynx</em> <em>lynx</em>) and wolf (<em>Canis</em> <em>lupus</em>). DNA was extracted using a protocol for isolating water eDNA and genotyped using amplicon sequencing of short tandem repeats (STR) and, for brown bear, a sex marker, on a high-throughput sequencing platform. Individual genotypes were obtained for all species, but genotyping performance differed among samples and species. The proportion of samples genotyped to individuals was higher for brown bear samples (5/7) than for wolf (7/10) and lynx (4/9), but locus genotyping success was greater for brown bear (0.88). Results for three species show that reliable individual genotyping, including sex identification, is now possible from eDNA in snow tracks, underlining its vast potential to complement the non-invasive genetic methods used for wildlife. To fully leverage the application of snow track eDNA, improved understanding of the ideal species- and site-specific sampling conditions, as well as laboratory methods promoting genotyping success are needed. This will also inform efforts to retrieve and type nuclear DNA from other eDNA samples, thereby advancing eDNA–based individual and population-level studies.</p>
DNA sequence data generated using non-invasive feather and eggshell samples from the Grenada Dove for two gene regions: Cyt b and ND2
<p>As an island endemic with a decreasing population, the Critically Endangered Grenada Dove <em>Leptotila wellsi</em> is threatened by accelerated loss of genetic diversity resulting from ongoing habitat fragmentation. Small, threatened populations are difficult to sample directly but advances in molecular methods mean that non-invasive samples can be used. We performed the first assessment of genetic diversity of populations of Grenada Dove by a) assessing mtDNA genetic diversity in the only two areas of occupancy on Grenada, b) defining the number of haplotypes present at each site and c) evaluating evidence of isolation between sites. We used non-invasively collected samples from two locations: Mt Hartman (n=18) and Perseverance (n=12). DNA extraction and PCR were used to amplify 1,751 bps of mtDNA from two mitochondrial markers: NADH dehydrogenase 2 (<em>ND2</em>) and Cytochrome b (<em>Cyt b</em>). Haplotype diversity (<em>h</em>) of 0.4, a nucleotide diversity (π) of 0.00023 and two unique haplotypes were identified within the <em>ND2</em> sequences; a single haplotype was identified within the <em>Cyt b </em>sequences. Of the two haplotypes identified; the most common haplotype (haplotype A = 73.9%) was observed at both sites and the other (haplotype B = 26.1%) was unique to Perseverance. Our results show low mitochondrial genetic diversity and clear evidence for genetically isolated populations. The Grenada Dove needs urgent conservation action, including habitat protection and potential augmentation of gene flow by translocation in order to increase genetic resilience and diversity with the ultimate aim of securing the long-term survival of this Critically Endangered species. </p>
Data from: Evaluating genotyping-in-thousands by sequencing as a genetic monitoring tool for a climate sentinel mammal using non-invasive and archival samples
<p>Genetic tools for wildlife monitoring can provide valuable information on spatiotemporal population trends and connectivity, particularly in systems experiencing rapid environmental change. Though many DNA sequencing approaches still require high quality and quantity of DNA obtained from traditional sources (e.g. blood and tissue), rapid genotyping tools such as Genotyping-in-Thousands by sequencing (GT-seq) have improved our ability to make use of degraded and less concentrated DNA commonly obtained from non-invasive and archival samples. Here, we developed a multi-purpose GT-seq panel (307 single nucleotide polymorphisms) for a climate sentinel mammal (the American pika, <em>Ochotona princeps</em>) for use as a genetic tool for monitoring populations in the Canadian Rocky Mountains. We optimized the panel using contemporary tissue samples (n = 77) and subsequently applied it to archival tissue (n = 17) and contemporary fecal pellet samples (n = 129) to evaluate its effectiveness at identifying individuals and sex, estimating relatedness, and inferring population structure. The panel demonstrated high efficacy with contemporary and archival tissue samples (94.7% and 90.5% genotyping success, respectively) and negligible genotyping error (0.001% and 0.0%, respectively). Despite relatively high genotyping success for fecal pellet samples (79.7%), high genotyping error (28.4%) limited its power as a monitoring tool to assess genetic variation using non-invasive samples and highlighted the need for further optimization around sample and data collection.</p>
Data from: Genomic footprint of cladogenesis revealed through RADseq and Sanger sequencing demonstrates congruent patterns in the velvet worm Peripatopsis sedgwicki species complex (Onychophora: Peripatopsidae)
<p>In the present study, first generation DNA sequencing (mitochondrial cytochrome c oxidase subunit one, <em>COI</em>) and reduced-representative genomic RADseq data were used to understand the patterns and processes of diversification of the velvet worm, <em>Peripatopsis sedgwicki</em> species complex across its distribution range in South Africa. For the RADseq data, three datasets (two primary and one supplementary) were generated corresponding to 1259 - 11,468 SNPs, in order to assess the species diversity and phylogeographic of the species complex. Tree topologies for the two primary datasets were inferred using maximum likelihood and Bayesian inferences methods. Phylogenetic analyses using the <em>COI </em>datasets retrieved four distinct, statistically well-supported clades within the species complex. Five species delimitation methods applied to the <em>COI </em>data (ASAP, bPTP, bGMYC, STACEY, and iBPP) all showed support for the distinction of the Fort Fordyce Nature Reserve specimens. In the main <em>P. sedgwicki </em>species complex, the species delimitation methods revealed a variable number of operational taxonomic units and overestimated the number of putative taxa. Divergence time estimates coupled with the geographic exclusivity of species and phylogeographic results suggest recent cladogenesis during the Plio/Pleistocene. The RADseq were subjected to a principal components analysis and a discriminant analysis of principal components, under a maximum-likelihood framework. The latter results corroborate the four main clades observed using the <em>COI</em> data, however, applying additional filtering revealed additional diversity. The high overall congruence observed between the RADseq and <em>COI </em>data suggests that first generation sequence data remain a cheap and effective method for evolutionary studies, although RADseq does provide a far greater resolution of contemporary temporo-spatial patterns. </p>
(Extended Data) Amplicon deep sequencing of ama1 and mdr1 to track within-host P. falciparum diversity throughout treatment in a clinical drug trial
<p>These extended data accompany the manuscript: Targeted Amplicon deep sequencing of ama1 and mdr1 to track within-host <em>P. falciparum</em> diversity throughout treatment in a clinical drug trial</p> <p><strong>Table S1: Concentration ratios and resulting parasitemia in artificial dna mixtures of P. falciparum Lab Isolates 3D7 and Dd2.</strong> This table presents the parasitemia for the artificial mixtures of P. falciparum lab isolates 3D7 and Dd2. Each mixture was prepared at varying ratios of 3D7 to Dd2, starting from equal proportions to a complete presence of only 3D7. The original concentration of each isolate was approximately 50,000 parasites per microliter (pf/μl), and the table displays the proportion of each strain in the mixture and the resulting total parasitemia concentration.</p> <p><strong>Table S2. List of PCR and deep sequencing primers.</strong> This table shows the list of forward and reverse primers used for deep sequencing. In boldface are the MID tags, while in the regular face are the forward primers</p> <p><strong>Table S3. The relative frequencies of each ama1 variant and the number of samples with each variant.</strong> The relative frequencies (%) of the 33 AMA1 variants in pre-and post-treatment samples (n = 330) are shown as a 33 amino acid sequence. The frequencies were calculated by dividing the number of reads of each microhaplotype by the total number of reads obtained per sample (116,187,131).</p> <p><strong>Table S4. Distribution of microhaplotypes among samples.</strong> This table shows the occurrence of microhaplotypes across all participants, both with monoclonal and multiclonal ama1 infections. It presents the ama1 clonality – monoclonal or multiclonal (column 1) - participant IDs (column 2), microhaplotype IDs (column 3), and the relative frequencies of these microhaplotypes across timepoints from 0 to 1008 hours (day 42) (column 3). Dashes represent time points where microhaplotypes were missing or were not detected.</p> <p><strong>Table S5. Distribution of rare microhaplotypes among samples.</strong> This table shows the occurrence of rare microhaplotypes in various samples. It presents participant IDs (column 1), microhaplotype IDs (column 2), and the relative frequencies of these microhaplotypes across time points from 0 to 1008 hours (day 42) (column 3). Samples containing rare microhaplotypes - specifically from PID10, PID32, PID38, PID40, PID49, PID60, PID63, and PID65 - are shown in orange, along with the corresponding rare microhaplotypes and their time points of occurrence. Furthermore, participants are categorised by shared microhaplotypes to indicate instances of rarity and commonality. Except for one microhaplotype unique to PID30, rare microhaplotypes were detected in several samples, frequently exceeding a 5% relative frequency. Dashes represent time points where microhaplotypes were missing or were not detected.</p> <p><strong>Table S6. The parasitemia levels associated with each ama1 microhaplotype per timepoint.</strong> This table shows the parasitemia for each ama1 microhaplotype per timepoint and each participant. “Patient ID” represents the patient ID, “AMA1 COI at 0h” represents the complexity of infection (COI) for each participant at baseline, based on ama1 while subsequent columns represent the parasitemia for each ama1 microhaplotype from timepoint 0h to 1008h. Parasitemia was back-calculated using the COI and total parasitemia for each time point. For time points with a COI > 1, parasitemia for the respective ama1 microhaplotypes are separated by commas, cells in red indicate timepoints without sequencing data (ND = not determined). In contrast, cells in grey indicate time points where microhaplotypes were detected below 10 parasites/μl, hence at risk of falling below the sampling limit.</p> <p><strong>Figure S1. Performance of AmpSeq in the sequencing controls.</strong> Six aliquots were prepared for each control set to ensure sufficient control data in case of PCR or sequencing failure. The median read depth in the lab controls was 5,658 (range 4,310 – 12,603) and 704 (291 – 1,676). The x-axis represents the aliquot identifier across the five mixtures, starting from 1 to 6, while the y-axis represents the proportions of each variant across all aliquots. For ama1 (A), two variants (3D7 and Dd2) were detected, whereas in mdr1 (B), two variants were detected YY, FY and NY following amplification of Dd2 Copy I, Dd2 Copy II and 3D7, respectively. For ama1, sequencing failed for aliquot 6 of control set 1, while for mdr1, sequencing failed for aliquot 2 and 6 of control set 3, aliquots 1 and 6 of control set 4 and aliquots 1 and 5 of control set 5. Under the mdr1 control set 4, the Dd2 copy II (86F, 184Y) was not identified, possibly due to having very low concentrations that were not picked up in this aliquot. Based on our control mixtures, the minimum variant frequency we could detect was 5%.</p> <p><strong>Figure S2. Heatmaps of the successfully PCR amplified and sequenced samples for ama1 (A) and mdr1 (B).</strong> The rows represent the study participants, while the columns represent time in hours. Successfully sequenced samples are shown in blue, those that failed PCR are shown in red and those that failed sequencing are in black. The timepoint “ Rec” represents unscheduled visits where a recurrent sample was collected. The unshaded areas with "-" are time points where samples were not collected. For each time point, the number of samples successfully sequenced (n Successful) is indicated in the last row of each panel. The table in panel C shows the groupings of samples based on parasitemia, high (> 5,000), moderate (100-5,000) and low (< 100 parasites per microlitre). Many samples collected between 0h-12h had high parasitemia, samples collected between 18h–30h had moderate parasitemia, while samples collected after 30h were primarily of low parasitemia.</p> <p><strong>Figure S3. The mean complexity of infection (COI) by AMA1 throughout treatment.</strong> The mean COI (red diamonds) appeared to be stable (between 1.5 - 2) from baseline (0h) up to 72h and thereafter fluctuated due to the small sample sizes (<5) in the post-treatment samples. The black dots represent the COI per sample.</p> <p> </p>
Data required for "Low mutation rate of spontaneous mutants enables detection of causative genes by comparing whole genome sequences"
<p>In the early 1900s,mutation breeding to select varieties with desirable traits using spontaneous mutation was actively conducted around the world, including Japan. In rice, the number of fixed mutations per generation was estimated to be 1.38-2.25. Although this low mutation rate was a major problem for breeding in those days, in the modern era with the development of NGS technology, it was conversely considered to be an advantage for efficient gene identification. In this paper, we proposed an in silico approach using next-generation sequencing (NGS) to compare the whole genome sequence of a spontaneous mutant with that of a closely related strain with a nearly identical genome, to find polymorphisms that differ between them, and to identify the causal gene by predicting the functional variation of the gene caused by the polymorphism. Using this approach, we found four causal genes for the dwarf mutation, the round shape grain mutation and the awnless mutation. Three of these genes were the same as those previously reported, but one was a novel gene involved in awn formation. The novel gene was isolated from Bozu-Aikoku, a mutant of Aikoku with the awnless trait, in which nine polymorphisms were predicted to alter gene function by their whole-genome comparison. Based on the information on gene function and tissue-specific expression patterns of these candidate genes, Os03g0115700/LOC_Os03g02460, annotated as a shortchain dehydrogenase/reductase SDR family protein, is most likely to be involved in the awnless mutation. Indeed, complementation tests by transformation showed that it is involved in awn formation. Thus, this method is an effective way to accelerate genome breeding of various crop species by enabling the identification of useful genes that can be used for crop breeding with minimal effort for NGS analysis.</p>
Sequencing data for Flp-In T-Rex HeLa PINK1 CRISPR/CAS9 knock-out (CVCL_D5JI)
<p>PCR and sequencing data (.ab1 and .seq files) for CRISPR-Cas9 PINK1 knockout HeLa cells (clone 3C9) described in <a title="doi link" href="https://doi.org/10.1098/rsob.210264">doi.org/10.1098/rsob.210264</a>. Knock-out generated using the following guides (targeting exone 2), available from MRC-PPU Reagents and Services, University of Dundee: DU52528 and DU52530 (CRISPR Project #CR252). This cell line has been deposited to Cellosaurus: CVCL_D5JI.</p>
16S rDNA sequencing data for characterizing the endosymbionts of leaf curl plum aphid ( Brachycaudus helichrysi) clones
<p>Asexual lineages often exhibit broad distributions and can thrive in extreme habitats compared to their sexual counterparts. Several hypotheses can be proposed to explain this pattern. Asexual lineages could be versatile genotypes with wide environmental tolerance, enabling their dispersal and persistence across large geographic areas. Alternatively, asexual genotypes could be ecological specialists that thrive in specific environments and outcompete relatives colonizing distantly related areas with similar conditions in the process. Several aphid species feature widespread obligate asexual lineages, commonly known as "superclones". Yet it is often unknown whether these clones are widespead ecological generalists or successful specialists. To explore these hypotheses, we examined climatic niche differentiation among six globally distributed obligate asexual lineages of the cosmopolitan aphid pest, <em>Brachycaudus helichrysi</em>. To insure that we were investigating the aphid genotype niche and not a by-product of their association with endosymbionts mediating thermal tolerance, we first verified that clones hosted similar endosymbiont communities. Subsequently, we conducted multivariate analyses on clone occurrence data on a worldwide scale. Our results revealed that despite their global distribution, <em>B. helichrysi</em> superclones occupy different climatic niches. This study represents the first evidence that aphid superclones distribution can be mediated by distinctive ranges of climatic tolerance.</p>
MUFFIN : A suite of tools for the analysis of functional sequencing data - Example input data
<p>This repository contains the data required to run the example notebooks and to reproduce the figures from the paper : </p> <div> <div><strong>MUFFIN : A suite of tools for the analysis of functional sequencing data</strong></div> </div> <div><em>Pierre de Langen, Benoit Ballester</em></div> <div>bioRxiv 2023.12.11.570597; doi: <a href="https://doi.org/10.1101/2023.12.11.570597" target="_blank" rel="noopener">https://doi.org/10.1101/2023.12.11.570597</a></div> <div> </div> <div>Source code is located here :</div> <div><a href="https://github.com/pdelangen/Muffin" target="_blank" rel="noopener">https://github.com/pdelangen/Muffin</a></div> <div> </div> <ul> <li><strong>10k_pbmc_gene/ </strong>contains the data for 10k pbmc dataset in standard 10x sparse count table format.</li> <li><strong>genome_annot/</strong> contains gencode v38 and chromosomes for the human (used for gene set enrichment analyses)</li> <li><strong>GO_files/</strong> contains gene set information retrieved from the g:ProfileR website.</li> <li><strong>immune_chip/ </strong>contains the data required to re-run the ChIP-seq analyses, it will require to also launch the dl_data.smk to retrieve the data from ENCODE.</li> <li><strong>tcga_atac/</strong> contains the sample-genomic region ATAC tag count table, as well as the sample metadata and a gene set file of cancer hallmark genes.</li> <li><strong>scATAC/</strong> contains the cell barcode-genomic region ATAC tag count table, as well as the barcode metadata and the 10k pbmc dataset pre-analyzed in h5 AnnData format.</li> </ul> <div> </div>
Derived Protein Sequence Data From Kaggle
<p>It is the dataset containing protein sequence and its associated go-terms, used for priliminary training process</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.