Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
233
datasets available to search
ShareScore release 0.9.0
Dataset results
233 results for “long-read”
Metagenome-assembled genomes from Stordalen Mire, Sweden (2019) (MAGs from long-read, short-read, & hybrid assemblies)
<p>METHODS:</p> <p>Soil samples (6 total) were collected at the Stordalen Mire site in 2019 from two depths (1-5 & 20-24 cm below ground) across three habitats (Palsa, Bog, and Fen). DNA was extracted based on the protocol described by <a href="http://dx.doi.org/10.17504/protocols.io.yxmvm244bg3p/v1">Li et al. (2024)</a>. For short reads, libraries were prepared at the Joint Genome Institute (JGI) with the KAPA Hyperprep kit, and sequenced with Illumina NovaSeq 6000. For long reads, libraries were prepared with the SMRTbell Express Template Prep Kit 2.0 (PacBio), then sequenced using PacBio Sequel IIe at JGI. PacBio data was processed at JGI to form filtered CCS (Circular Consensus Sequencing) reads. </p> <p>Assemblies were generated with short-only, long-only, and hybrid read sources: <strong>Short-only</strong> was assembled with <a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC5411777/">metaSPAdes</a> (v3.15.4) using <a href="https://zenodo.org/records/10806928">Aviary</a> (v0.5.3) with default parameters. <strong>Long-only</strong> was assembled with <a href="https://www.nature.com/articles/s41592-020-00971-x">metaFlye</a> (v2.9-b1768) using <a href="https://zenodo.org/records/10806928">Aviary</a> (v0.5.3) with default parameters. <strong>Hybrid</strong> assembly was performed using <a href="https://zenodo.org/records/10806928">Aviary</a> v0.5.3 with default parameters. This involved a step-down procedure with long-read assembly through <a href="https://www.nature.com/articles/s41592-020-00971-x">metaFlye</a> (v2.9-b1768), followed by short-read polishing by <a href="https://genome.cshlp.org/content/27/5/737">Racon</a> (v1.4.3), <a href="https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0112963">Pilon</a> (v1.24) and then Racon again. Next, reads that didn't map to high-quality metaFlye contigs were hybrid assembled with <a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC5411777/">SPAdes (--meta option)</a> and binned out with <a href="https://peerj.com/articles/7359/">MetaBAT2</a> (v2.1.5). For each bin, the reads within the bin were hybrid assembled using <a href="https://journals.plos.org/ploscompbiol/article?id=10.1371/journal.pcbi.1005595">Unicycler</a> (v0.4.8). The high-coverage metaFlye contigs and Unicycler contigs were then combined to form the assembly fasta file. Genome recovery was performed using <a href="https://zenodo.org/records/10806928">Aviary</a> v0.5.3 with samples chosen for differential abundance binning by <a href="https://zenodo.org/records/10939393">Bin Chicken</a> (v0.4.2) using <a href="https://zenodo.org/records/7130825">SingleM metapackage S3.0.5</a>. This involved initial read mapping through <a href="https://zenodo.org/records/10531254">CoverM</a> (v0.6.1) using <a href="https://academic.oup.com/bioinformatics/article/34/18/3094/4994778">minimap2</a> (v2.18) and binning by <a href="https://peerj.com/articles/1165/">MetaBAT</a>, <a href="https://peerj.com/articles/7359/">MetaBAT2</a> (v2.1.5), <a href="https://www.nature.com/articles/s41587-020-00777-4">VAMB</a> (v3.0.2), <a href="http://doi.org/10.1038/s41467-022-29843-y">SemiBin</a> (v1.3.1), <a href="https://zenodo.org/records/10460259">Rosella</a> (v0.4.2), <a href="https://www.nature.com/articles/nmeth.3103">CONCOCT</a> (v1.1.0) and <a href="https://academic.oup.com/bioinformatics/article/32/4/605/1744462">MaxBin2</a> (v2.2.7). Genomes were analyzed using <a href="https://www.nature.com/articles/s41592-023-01940-w">CheckM2</a> (v1.0.2) and clustered at 95% ANI using <a href="https://zenodo.org/records/10526086">Galah</a> (v0.4.0).</p> <p> </p> <p>FILES:</p> <ul> <li><strong>EMERGE_MAGs_2019_long-short-hybrid.tar.gz</strong> - Archive containing the MAG files (.fna).</li> <li><strong>metadata_MAGs_2019_EMERGE.tsv</strong> - Table containing source sample names and accessions, GTDB classifications, CheckM2 quality information, NCBI GenomeBatch- and MIMAG(6.0)-formatted attributes, and other metadata for the MAGs.</li> </ul> <p> </p> <p>FUNDING:</p> <p>This research is a contribution of the EMERGE Biology Integration Institute (<a href="https://emerge-bii.github.io/">https://emerge-bii.github.io/</a>), funded by the National Science Foundation, Biology Integration Institutes Program, Award # 2022070.</p> <p>This study was also funded by the Genomic Science Program of the United States Department of Energy Office of Biological and Environmental Research, grant #s DE-SC0004632. DE-SC0010580. and DE-SC0016440.</p> <p>We thank the Swedish Polar Research Secretariat and SITES for the support of the work done at the Abisko Scientific Research Station. SITES is supported by the Swedish Research Council's grant 4.3-2021-00164.</p> <p>Data from the Joint Genome Institute (JGI) was collected under BER Support Science Proposal 503530 (DOI: <a href="https://doi.org/10.46936/10.25585/60001148">10.46936/10.25585/60001148</a>), conducted by the U.S. Department of Energy Joint Genome Institute (<a href="https://ror.org/04xm1d337">https://ror.org/04xm1d337</a>), a DOE Office of Science User Facility supported by the Office of Science of the U.S. Department of Energy under Contract No. DE-AC02-05CH11231.</p>
Long-read sequencing and structural variant characterization in 1,019 samples from the 1000 Genomes Project
SV analysis of the long-read sequencing data of 1,019 samples from the 1000 Genomes Project. The data is hosted at the International Genome Sample Resource (IGSR) in the <a href="https://ftp.1000genomes.ebi.ac.uk/vol1/ftp/data_collections/1KG_ONT_VIENNA/">1KG_ONT_VIENNA</a> directory. Please see the <a href="https://ftp.1000genomes.ebi.ac.uk/vol1/ftp/data_collections/1KG_ONT_VIENNA/README_1KG_ONT_VIENNA.md">README</a> and <a href="https://ftp.1000genomes.ebi.ac.uk/vol1/ftp/data_collections/1KG_ONT_VIENNA/README_1KG_ONT_VIENNA_datareuse_statement.md">data reuse statement</a> for further information about this dataset.
Machine learning classifiers for species classification of fungi using error-prone long-reads on extended metabarcodes
<p>Machine learning models used in the decision tree of linked machine learning models (<a href="https://github.com/teenjes/fungal_ML">https://github.com/teenjes/fungal_ML</a>)</p>
The SV callsets of the HG002 human sample produced by cuteSV with multi long-read sequencing platforms.
<p>The SV callsets of the HG002 human sample produced by cuteSV with multi long-read sequencing platforms.</p>
Metagenomics assemblies and high-quality MAGs for "Long-read metagenomics to retrieve high-quality metagenome-assembled genomes from canine feces"
<p>This dataset includes the different metagenomics assemblies analyzed and its summary (_info.txt file):</p> <p>- <a href="https://zenodo.org/api/files/3a502803-82f7-4b51-ac62-7c1dfcdcb680/100_assembly.fasta">100_assembly.fasta</a> is the Flye 2.7 metagenomics assembly merging HMW and non-HMW datasets</p> <p>- <a href="https://zenodo.org/api/files/3a502803-82f7-4b51-ac62-7c1dfcdcb680/75_assembly.fasta">75_assembly.fasta</a> is the Flye 2.7 metagenomics assembly including 75% of random data of the merged dataset.</p> <p>- <a href="https://zenodo.org/api/files/3a502803-82f7-4b51-ac62-7c1dfcdcb680/50_assembly.fasta">50_assembly.fasta</a> is the Flye 2.7 metagenomics assembly including 50% of random data of the merged dataset.</p> <p>- <a href="https://zenodo.org/api/files/3a502803-82f7-4b51-ac62-7c1dfcdcb680/HMW_assembly.fasta?versionId=749ff6fd-2642-4ad1-971a-7f3404baa595">HMW_assembly.fasta</a> is the Flye 2.7 metagenomics assembly for HMW dataset.</p> <p>Moreover, it also includes the eight frameshift-corrected high-quality MAGs analyzed in the manuscript. </p>
Tandem repeat catalog of the human genome generated from long-read assemblies
<p>Allele sequences of polymorphic loci (VCF) and README for all version 2 (2.0 + 2.1) files</p>
The PODER cross-ancestry long-read RNA-seq dataset
<p>Table legends:</p> <ul> <li>column_descriptions.docx: Extended description of columns in selected tables.</li> <li>00_sample_metadata: Sample and experimental metadata pertaining to each LR-RNA-seq sample.</li> <li>01_uma_gtf: Unfiltered merged annotation (UMA) GTF for all spliced intron chains discovered by each of the tools.</li> <li>02_uma_mt: Metadata table for SQANTI QC, Recount3 support, protein prediction, and other relevant characteristics for each transcript in the UMA annotation used to filter transcripts.</li> <li>03_poder_gtf: Final PODER GTF, including predicted CDSs for novel transcripts and annotated CDSs for known transcripts.</li> <li>04_poder_mt: Metadata table for SQANTI QC, Recount3 support, protein prediction, and other relevant characteristics for each transcript in PODER.</li> <li>06_poder_t_counts: PODER transcript counts in each sample as quantified by lr-kallisto.</li> <li>07_poder_g_counts: PODER gene counts in each sample as quantified by lr-kallisto.</li> <li>08_mage_t_counts_poder: Transcript counts for the MAGE RNA-seq dataset computed using the PODER annotation with kallisto.</li> <li>09_mage_t_counts_gencode: Transcript counts for the MAGE RNA-seq dataset computed using the GENCODE annotation with kallisto.</li> <li>10_mage_t_counts_enh: Transcript counts for the MAGE RNA-seq dataset computed using the Enhanced GENCODE annotation with kallisto.</li> <li>11_astu: Allele-specific transcript usage results.</li> <li>12_ase: Allele-specific expression results.</li> <li>13_mage_tau: Tau values for Enhanced GENCODE transcripts computed on the MAGE RNA-seq dataset.</li> <li>14_gwas_enrichments: GWAS enrichment results for ASTU genes overlapping GWAS genes.</li> <li>16_inter_catalog_overlap: Boolean detection of transcripts across PODER, annotations (GENCODE, RefSeq) and other RNA-seq / LR-RNA-seq-derived transcript catalogs (CHESS, GTEx, ENCODE4).</li> <li>17_personalized_hg38: Isoforms and their intron chains detected using personalized-GRCh38s.</li> <li>18_enh_gencode: Enhanced GENCODE GTF with novel transcripts from PODER added to all annotated transcripts from GENCODE v47</li> </ul>
Companion data deposit of manuscript: Evaluating and improving the representation of bacterial contents in long-read metagenome assemblies
<p>This upload contains the metagenome assemblies and their binning results generated & described in the manuscript "Evaluating and improving the representation of bacterial contents in long-read metagenome assemblies" (preprint version: arxiv2210.00098, "Towards complete representation of bacterial contents in metagenomic samples"). </p> <p>Mapping of sample names in the file names and the descriptors used as in the manuscript can be found in table S1, which is available along with the manuscript and also included in the supplementary_tables_and_figures tar archive here.</p>
Long-read sequencing reveals extensive gut phageome structural variations driven by genetic exchange with bacterial hosts
<p><span>Genetic variations are instrumental for unraveling phage evolution and deciphering their functional implications. Here we explore the underlying fine-scale genetic variations in the gut phageome, especially structural variations (SVs). By employing virome-enriched long-read metagenomics sequencing across 91 individuals, we identified a total of 14,438 non-redundant phage SVs, and revealed their prevalence within the human gut phageome. These SVs are mainly enriched in genes involved in recombination, DNA methylation, and antibiotic resistance. Strikingly, a substantial fraction of phage SV sequences share close homology with bacterial fragments, with most SVs enriched for horizontal gene transfer (HGT) mechanism. Further investigations showed that these SV sequences were genetic exchanged between specific phage-bacteria pairs, particularly between phages and their respective bacterial hosts. Temperate phages exhibits a higher frequency of genetic exchange with bacterial chromosomes then virulent phages. Collectively, our findings provide novel insights into the genetic landscape of the human gut phageome.</span></p>
Enhanced Protein Isoform Characterization Through Long-Read Proteogenomics - Workflow Results
<pre> </pre> <p>The detection of physiologically relevant protein isoforms encoded by the human genome is critical to biomedicine. Mass spectrometry (MS)-based proteomics is the preeminent method for protein detection, but isoform-resolved proteomic analysis relies on accurate reference databases that match the sample; neither a subset nor a superset database is ideal. Long-read RNA sequencing (e.g. PacBio, Oxford Nanopore) provides full-length transcript sequencing, which can be used to predict full-length proteins. Here, we describe a long-read proteogenomics approach for integrating matched long-read RNA-seq and MS-based proteomics data to enhance isoform characterization. We introduce a classification scheme for protein isoforms, discover novel protein isoforms, and present the first protein inference algorithm for the direct incorporation of long-read transcriptome data in protein inference to enable detection of protein isoforms that are intractable to MS detection. We have released an open-source Nextflow pipeline that integrates long-read sequencing in a proteomic workflow for isoform-resolved analysis.</p> <p>Companion Repositories:</p> <ol> <li><a href="https://doi.org/10.5281/zenodo.5920817">Long-Read-Proteogenomics Workflow GitHub Repository Release</a></li> <li><a href="https://doi.org/10.5281/zenodo.5920847">Long-Read-Proteogenomics Analysis GitHub Repository Release</a></li> </ol> <p>Companion Datasets</p> <ol> <li><a href="https://zenodo.org/deposit/5703754">Long-Read-Proteogenomics Workflow Sample and Reference Data</a></li> <li><a href="https://doi.org/10.5281/zenodo.5234651">TEST Data for Long-Read-Proteogenomics Workflow GitHub Actions</a></li> </ol> <p>This Repository contains the complete output from the execution of the <a href="https://doi.org/10.5281/zenodo.5920817">Long-Read-Proteogenomics Workflow</a>, using the input from <a href="https://zenodo.org/deposit/5703754">Jurkat Samples and Reference Data</a>. </p> <p>The file <em>jurkat.flnc.bam </em>was 6.5 GB had to be split into 13 separate files and for use should be rejoined -- here are the steps that were used to split the file up. </p> <p>1. Convert <em>jurkat.flnc.bam</em> (binary format) to sam file (text format) without header: <em>samtools view jurkat.flnc.bam > jurkat.flnc.sam</em></p> <p>2. Capture the header: <em>samtools view -H jurkat.flnc.bam > jurkat.flnc.header.sam</em></p> <p>3. Split <em>jurkat.flnc.sam</em> into smaller files (aim to get final size under 2GB): <em>split -l 400000 jurkat.flnc.sam jurkat.flnc.chunk.</em></p> <p>4. Convert each of these files back to bam for uploading: <em>samtools view -b jurkat.flnc.chunk.a* -o jurkat.flnc.chunk.a*.bam (*=a,b,c,d,e,f,g,h,i,j,k,l,m)</em></p> <p>After downloading, reverse this process including using the header file which is found in the LRPG-Manuscript-Results-results-results-jurkat-isoseq3-companion-files.tar.gz file></p> <p>1. Convert the bam files back to sam files: <em>samtools view jurkat.flnc.chunk.a*.bam > jurkat.flnc.chunk.a*.sam (*=a,b,c,d,e,f,g,h,i,j,k,l,m)</em></p> <p>2. Combine the header together with the sam files: <em>cat jurkat.flnc.chunk.a*sam > jurkcat.flnc.sam (</em>verified the same number of lines of the sam files is identical to the number of lines of the original without header: 4,956,761. Header file is 13 lines.</p> <p>3. Convert to bam files if desired: <em>samtools view -b jurkat.flnc.sam -o jurkat.flnc.bam</em></p> <p>4. Rehead with the header file: <em>samtools reheader -P -i jurkat.flnc.header.sam jurkat.flnc.bam</em></p>
Parent-of-origin detection and chromosome-scale haplotyping using long-read DNA methylation sequencing and Strand-seq
<p>Hundreds of loci in human genomes have alleles that are methylated differentially according to their parent of origin. These imprinted loci generally show little variation across tissues, individuals, and populations. We show that such loci can be used to distinguish the maternal and paternal homologs for all autosomes, without the need for the parental DNA. We integrate methylation-detecting nanopore sequencing with the long-range phase information in Strand-seq data to determine the parent of origin of chromosome-length haplotypes for both DNA sequence and DNA methylation in five trios with diverse genetic backgrounds.</p>
Data for "A replicable and modular benchmark for long-read transcript quantification methods"
<p>This archive contains the input necessary to run the inital (TranSigner-protocol and IsoQuant-protocol) benchmarks associated with the <a href="https://github.com/COMBINE-lab/lr_quant_benchmarks" target="_blank" rel="noopener"><code>lr_quant_benchmarks repository</code></a>. The archive can be decompressed with <code>tar</code> and <code>zstd</code> using the command <code>tar --use-compress-program=zstd -xf input.tar.zstd</code>.</p>
Polished Assemblies for "GoldPolish-Target: Targeted long-read genome assembly polishing"
<p>GoldPolish-Target is a targeted genome assembly polishing tool that uses long reads. We tested GoldPolish-Target on Oxford Nanopore Technologies datasets with a human cell line (NA24385) and Drosophila melanogaster (fruit fly). Here, we provide the data for the GoldRush baseline (unpolished) assembly and the GoldPolish-Target and medaka polished assemblies of these long-read datasets.</p>
TEST DATA for Enhanced protein isoform characterization through long-read proteogenomics
<p>Test data for The detection of physiologically relevant protein isoforms encoded by the human genome is critical to biomedicine. Mass spectrometry (MS)-based proteomics is the preeminent method for protein detection, but isoform-resolved proteomic analysis relies on accurate reference databases that match the sample; neither a subset nor a superset database is ideal. Long-read RNA sequencing (e.g. PacBio, Oxford Nanopore) provides full-length transcript sequencing, which can be used to predict full-length proteins. Here, we describe a long-read proteogenomics approach for integrating matched long-read RNA-seq and MS-based proteomics data to enhance isoform characterization. We introduce a classification scheme for protein isoforms, discover novel protein isoforms, and present the first protein inference algorithm for the direct incorporation of long-read transcriptome data in protein inference to enable detection of protein isoforms that are intractable to MS detection. We have released an open-source Nextflow pipeline that integrates long-read sequencing in a proteomic workflow for isoform-resolved analysis.</p> <p>Companion Repositories:</p> <ol> <li><a href="https://doi.org/10.5281/zenodo.5920817">Long-Read-Proteogenomics Workflow GitHub Repository Release</a></li> <li><a href="https://doi.org/10.5281/zenodo.5920847">Long-Read-Proteogenomics Analysis GitHub Repository Release</a></li> </ol> <p>Companion Datasets</p> <ol> <li><a href="https://zenodo.org/deposit/5703754">Jurkat Samples and Reference Data</a></li> <li><a href="http://10.5281/zenodo.5920920">Long-Read-Proteogenomics Workflow Results using Jurkat Sample data</a></li> </ol> <p>This Repository contains the test data, specifically:</p> <p><a href="https://doi.org/10.5281/zenodo.5234651">TEST Data for Long-Read-Proteogenomics Workflow GitHub Actions</a></p>
De novo assembly of a long-read Amblyomma americanum genome
<p>Genome assemblies of Amblyomma americanum generated from PacBio HiFi sequencing of 50 individual female ticks. This repository contains the unphased diploid assembly generated by the Flye assembler (Arcadia_Amblyomma_americanum_asm001.fasta). In addition, there are two associated fasta files containing sequences generated by submitting the unphased diploid assembly to separation by the Purge_Dups pipeline (purged pseudo-haploid assembly and haplotig assembly).</p> <p>Flye assembler: https://github.com/fenderglass/Flye</p> <p>Purge_Dups pipeline: https://github.com/dfguan/purge_dups</p> <p>NCBI Bioproject: PRJNA932813</p>
De novo assembly of a long-read Amblyomma americanum genome (NCBI/Genbank deposited genome)
<p>Genome assembly of Amblyomma americanum generated from PacBio HiFi sequencing of 50 individual female ticks. This repository contains the phased pseudo-haploid tick genome generated after assembly using Flye, phasing using Purge_Dups, and clean-up using custom python scripts generated in-house. </p> <p>NCBI Bioproject: PRJNA932813</p>
Long-read, chromosome-scale assembly of Vitis rotundifolia cv. Carlos and its unique resistance to Xylella fastidiosa subsp. fastidiosa.
<p>We assembled and annotated a new, long-read genome assembly for ‘Carlos’, a cultivar of muscadine that exhibits tolerance, to build upon the existing genetic resources available for muscadine. We are awaiting release of the genome through NCBI, so we have made the assembly and annotations available here.</p>
Reconstructing NOD-like receptor alleles with high internal conservation in Podospora anserina using long-read sequencing
Open the record for dataset details and reuse information.
A novel method to assess the integrity of frozen archival DNA samples: Alpha-diversity ratios of short and long-read 16S rRNA gene sequences
Open the record for dataset details and reuse information.
Raw Nanopore data for "Nanopore Long-Read Guided Complete Genome Assembly of Hydrogenophaga intermedia, and Genomic Insights into 4-Aminobenzenesulfonate, p-Aminobenzoic Acid and Hydrogen Metabolism in the Genus Hydrogenophaga"
<p>This is the raw Nanopore dataset (fast5) for Hydrogenophaga intermedia PBC. The gDNA was prepared using the now obsolete SQK-NSK007 kit and sequenced on a MINION R9 Flowcell. </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.