Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
2,848
datasets available to search
ShareScore release 0.9.0
Dataset results
2,848 results for “sequence data”
Training Data for "Binning of metagenomic sequencing data" tutorial
<p><strong>Metagenomics is the study of genetic material recovered directly from environmental samples, such as soil, water, or gut contents, without the need for isolation or cultivation of individual organisms. Metagenomics binning is a process used to classify DNA sequences obtained from metagenomic sequencing into discrete groups, or bins, based on their similarity to each other</strong>. The goal of metagenomics binning is to assign the DNA sequences to the organisms or taxonomic groups that they originate from, allowing for a better understanding of the diversity and functions of the microbial communities present in the sample. This is typically achieved through computational methods that use sequence similarity, composition, and other features to group the sequences into bins.</p> <p>There are two main types of metagenomics binning: <strong>reference-based</strong> and <strong>de novo</strong>.</p> <ul> <li><strong>reference-based binning</strong> involves aligning the sequences to a database of known genomes or reference sequences</li> <li><strong>de novo binning</strong> involves clustering the sequences based on similarity without prior knowledge of the organisms or reference sequences present in the sample.</li> </ul> <p>Both methods have their strengths and limitations, and researchers often use a combination of approaches to improve the accuracy of their binning results. Metagenomics binning is an important tool for understanding the functional potential of microbial communities in various environments and has applications in fields such as biotechnology, environmental science, and human health.</p> <p>In this tutorial, we will learn how to run metagenomic binning tools and evaluate the quality of the results. In order to do that, we will use data from the study: <a href="https://www.ebi.ac.uk/metagenomics/studies/MGYS00005630#overview">Temporal shotgun metagenomic dissection of the coffee fermentation ecosystem</a> and MetaBAT2 algorithm. For an in-depth analysis of the structure and functions of the coffee microbiome, a temporal shotgun metagenomic study (six time points) was performed. The six samples have been sequenced with Illumina MiSeq utilizing whole genome sequencing.</p> <p>Based on the 6 original dataset of the coffee fermentation system, we generated mock datasets for this tutorial.</p>
DNA sequence and bioacoustic data of the Guibemantis liber complex from Madagascar (Amphibia, Mantellidae)
<p>Data from a taxonomic revision of the Guibemantis liber complex from Madagascar. The following data are included:</p> <p>- Advertisement call recordings of Guibemantis liber, G. razoky and G. razandry from different localities in wav format. See associated publication for metadata.</p> <p>- A table in Excel format with all DNA sequences used, metadata of the respective samples and specimens, and Genbank accession numbers.</p>
Data for: Human atlastin-3 is a constitutive ER membrane fusion catalyst (phylogenetic and sequence analysis)
<p>Homotypic membrane fusion catalyzed by the atlastin (ATL) GTPase sustains the branched endoplasmic reticulum (ER) network in metazoans. Our recent discovery that two of the three human ATL paralogs (ATL1/2) are C-terminally autoinhibited implied that relief of autoinhibition would be integral to the ATL fusion mechanism. An alternative hypothesis is that the third paralog ATL3 promotes constitutive ER fusion with relief of ATL1/2 autoinhibition used conditionally. However, published studies suggest ATL3 is a weak fusogen at best. Contrary to expectations, we demonstrate here that purified human ATL3 catalyzes efficient membrane fusion in vitro and is sufficient to sustain the ER network in triple knockout cells. Strikingly, ATL3 lacks any detectable C-terminal autoinhibition, like the invertebrate <em>Drosophila</em> ATL ortholog. Phylogenetic analysis of ATL C-termini indicates that C-terminal autoinhibition is a recent evolutionary innovation. We suggest that ATL3 is a constitutive ER fusion catalyst and that ATL1/2 autoinhibition likely evolved in vertebrates as a means of upregulating ER fusion activity on demand.</p>
Linkage maps and genotype data of strawberry produced with skim-sequencing data
<p>The following set of files contain the results and scripts to produce those results, described in Chapter 5 of the PhD thesis of Alejandro Thérèse Navarro, entitled "How to map a million markers: linkage mapping of skim-sequencing data in strawberry". In this study, a large dataset of markers produced by whole genome resequecning of a strawberry (<em>Fragaria </em>x <em>ananassa</em>) biparental population are used to generate linkage maps. To that end the software <a href="https://github.com/Alethere/SmoothDescent">Smooth Descent</a> is used, since it is oriented to obtaining linkage maps in usin low quality (error-prone) genotype data. With this methodology we were able to produce a linkage map of 27 out of 28 chromosomes of strawberry which containing 1.85M markers in ~2400 unique genetic mpositions. We also compare this map with a linkage map produced using SNP array data and with the genome sequence assembly "Camarosa".</p>
Methylation-free E.coli nanopore sequencing (ONT R9.4.1) data set
<p>The data set consists of fast5 files divided into 5 zip files (fast5_[1-5].zip), a genome record (Ecoli_K12_MG1655.fasta), an Illumina assembly genome (illumina_contigs.fasta) and a fastq file from Guppy 5 (guppy_basecalled.fastq.gz). We sequenced the Ecoli non-methylated genomic DNA (D5016, Zymo Research) with an ONT MinION device. The sequencing libraries were prepared by fragmenting the genomic DNA using Covaris g-TUBE and a Ligation sequencing kit (SQK-LSK109, Oxford Nanopore) with Flow Cell chemistry R9.4.1. We also performed short-read Illumina sequencing on the same sample using the TruSeq PCR-free library preparation on the MiSeq sequencing platform (Illumina, USA), and constructed a draft assembly from the Illumina sequencing results using SPAdes v3.6.0. We also upload a reference genome directly obtained from the E.coli sample producer website. </p> <p>In addition, the data set contains two fastq files that produced by the Lokatt basecaller (lokatt_basecalled.fasta.gz) and local-trained Bonito basecaller (bonito_local_basecalled.fastq.gz ), respectively, which are used for benchmarking in the Lokatt basecaller paper.</p>
Data from: Coding-sequence evolution does not explain divergence in petal anthocyanin pigmentation between Mimulus luteus var. luteus and M. l. variegatus
<p><span>Biologists have long been interested in understanding genetic constraints on the evolution of development. For example, noncoding changes in a gene might be favored relative to coding changes due to being less constrained by pleiotropic effects. Here we evaluate the importance of coding-sequence changes to the recent evolution of a novel anthocyanin pigmentation trait in the monkeyflower genus <em>Mimulus</em>. The magenta-flowered <em>Mimulus</em> <em>luteus</em> var. <em>variegatus</em> recently gained petal lobe anthocyanin pigmentation via a single-locus Mendelian difference from its sister taxon, the yellow-flowered <em>M. l. luteus</em>. Previous work showed that the differentially expressed transcription factor gene <em>MYB5a</em>/<em>NEGAN</em> is the single causal gene. However, it was not clear whether <em>MYB5a</em> coding-sequence evolution (in addition to the observed patterns of differential expression) might also have contributed to increased anthocyanin production in <em>M. l. variegatus</em>. Quantitative image analysis of tobacco leaves, transfected with <em>MYB5a</em> coding sequence from each taxon, revealed robust anthocyanin production driven by both alleles. Counter to expectations, significantly higher anthocyanin production was driven by the allele from the low-anthocyanin <em>M. l. luteus.</em> Together with previously-published expression studies, this supports the hypothesis that petal pigment in <em>M. l. variegatus</em> was not gained by protein-coding changes, but instead solely via non-coding cis-regulatory evolution. Finally, while constructing the transgenes needed for this experiment, we unexpectedly discovered two sites in <em>MYB5a</em> that appear to be post-transcriptionally edited – a phenomenon that has been rarely reported, and even less often explored, for nuclear-encoded plant mRNAs.</span></p>
Data for publication "Measuring the environment of a Cs qubit with dynamical decoupling sequences"
<p>Data sets as plotted in the preprint "Measuring the environment of a Cs qubit with dynamical decoupling sequences" are uploaded.<br> The zip file "data" contains a folder for each figure in the preprint (named after the figure). Each folder contains the data for all graphs in the respective figure.</p>
Supplementary data for: DNA sequences are as useful as protein sequences for inferring deep phylogenies
<p>Inference of deep phylogenies has almost exclusively used protein rather than DNA sequences, based on the perception that protein sequences are less prone to homoplasy and saturation or to issues of compositional heterogeneity than DNA sequences. Here we analyze a model of codon evolution under an idealized genetic code and demonstrate that those perceptions may be misconceptions. We conduct a simulation study to assess the utility of protein versus DNA sequences for inferring deep phylogenies, with protein-coding data generated under models of heterogeneous substitution processes across sites in the sequence and among lineages on the tree, and then analyzed using nucleotide, amino acid, and codon models. Analysis of DNA sequences under nucleotide-substitution models (possibly with the third codon positions excluded) recovered the correct tree at least as often as analysis of the corresponding protein sequences under modern amino acid models. We also applied the different data-analysis strategies to an empirical dataset to infer the metazoan phylogeny. Our results from both simulated and real data suggest that DNA sequences may be as useful as proteins for inferring deep phylogenies and should not be excluded from such analyses. Analysis of DNA data under nucleotide models has a major computational advantage over protein-data analysis, potentially making it feasible to use advanced models that account for among-site and among-lineage heterogeneity in the nucleotide-substitution process in inference of deep phylogenies.</p>
Coseismic Kinematics of the 2023 Kahramanmaras, Turkey Earthquake Sequence from InSAR and Optical Data
<p>We upload the supporting information and four high-resolution images Figures 1-4 in the manuscript accepted by the GRL journal (2023.07.31).</p>
Supplementary data for: Comparison of optical flow derivation techniques for retrieving tropospheric winds from satellite image sequences
Open the record for dataset details and reuse information.
Data from: Improved genome assembly of the whiteleg shrimp Penaeus (Litopenaeus) vannamei using long- and short-read sequences from public databases
Open the record for dataset details and reuse information.
Data for: High-throughput profiling of sequence recognition by tyrosine kinases and SH2 domains using bacterial peptide display
Open the record for dataset details and reuse information.
Green turtle ddRAD raw sequencing data
Open the record for dataset details and reuse information.
Data from: A cost-effective blood DNA methylation-based age estimation method in domestic cats, Tsushima leopard cats (Prionailurus bengalensis euptilurus), and Panthera species, using targeted bisulfite sequencing and machine learning models
Open the record for dataset details and reuse information.
Data for: Range and niche expansion through multiple interspecific hybridization - a genotyping by sequencing analysis of Cherleria (Caryophyllaceae)
Open the record for dataset details and reuse information.
Data from: Using transcriptome sequencing and pooled exome capture to study local adaptation in the giga-genome of Pinus cembra
Open the record for dataset details and reuse information.
Data from: Restriction site-associated DNA sequencing reveals local adaptation despite high levels of gene flow in Sardinella lemuru (Bleeker, 1853) along the northern coast of Mindanao, Philippines
Open the record for dataset details and reuse information.
Whole genome sequence data of Mycobacterium tuberculosis and Mycolicibacterium smegmatis mutants of the riboflavin biosynthetic pathway- Part 1
Open the record for dataset details and reuse information.
Data from: Latent generative landscapes as maps of functional diversity in protein sequence space
Open the record for dataset details and reuse information.
Data from: Microhaplotypes provide increased power from short-read DNA sequences for relationship inference
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.