Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
525
datasets available to search
ShareScore release 0.9.0
Dataset results
525 results for “Long read”
Metagenome-assembled genomes from Stordalen Mire, Sweden (2019) (MAGs from long-read, short-read, & hybrid assemblies)
<p>METHODS:</p> <p>Soil samples (6 total) were collected at the Stordalen Mire site in 2019 from two depths (1-5 & 20-24 cm below ground) across three habitats (Palsa, Bog, and Fen). DNA was extracted based on the protocol described by <a href="http://dx.doi.org/10.17504/protocols.io.yxmvm244bg3p/v1">Li et al. (2024)</a>. For short reads, libraries were prepared at the Joint Genome Institute (JGI) with the KAPA Hyperprep kit, and sequenced with Illumina NovaSeq 6000. For long reads, libraries were prepared with the SMRTbell Express Template Prep Kit 2.0 (PacBio), then sequenced using PacBio Sequel IIe at JGI. PacBio data was processed at JGI to form filtered CCS (Circular Consensus Sequencing) reads. </p> <p>Assemblies were generated with short-only, long-only, and hybrid read sources: <strong>Short-only</strong> was assembled with <a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC5411777/">metaSPAdes</a> (v3.15.4) using <a href="https://zenodo.org/records/10806928">Aviary</a> (v0.5.3) with default parameters. <strong>Long-only</strong> was assembled with <a href="https://www.nature.com/articles/s41592-020-00971-x">metaFlye</a> (v2.9-b1768) using <a href="https://zenodo.org/records/10806928">Aviary</a> (v0.5.3) with default parameters. <strong>Hybrid</strong> assembly was performed using <a href="https://zenodo.org/records/10806928">Aviary</a> v0.5.3 with default parameters. This involved a step-down procedure with long-read assembly through <a href="https://www.nature.com/articles/s41592-020-00971-x">metaFlye</a> (v2.9-b1768), followed by short-read polishing by <a href="https://genome.cshlp.org/content/27/5/737">Racon</a> (v1.4.3), <a href="https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0112963">Pilon</a> (v1.24) and then Racon again. Next, reads that didn't map to high-quality metaFlye contigs were hybrid assembled with <a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC5411777/">SPAdes (--meta option)</a> and binned out with <a href="https://peerj.com/articles/7359/">MetaBAT2</a> (v2.1.5). For each bin, the reads within the bin were hybrid assembled using <a href="https://journals.plos.org/ploscompbiol/article?id=10.1371/journal.pcbi.1005595">Unicycler</a> (v0.4.8). The high-coverage metaFlye contigs and Unicycler contigs were then combined to form the assembly fasta file. Genome recovery was performed using <a href="https://zenodo.org/records/10806928">Aviary</a> v0.5.3 with samples chosen for differential abundance binning by <a href="https://zenodo.org/records/10939393">Bin Chicken</a> (v0.4.2) using <a href="https://zenodo.org/records/7130825">SingleM metapackage S3.0.5</a>. This involved initial read mapping through <a href="https://zenodo.org/records/10531254">CoverM</a> (v0.6.1) using <a href="https://academic.oup.com/bioinformatics/article/34/18/3094/4994778">minimap2</a> (v2.18) and binning by <a href="https://peerj.com/articles/1165/">MetaBAT</a>, <a href="https://peerj.com/articles/7359/">MetaBAT2</a> (v2.1.5), <a href="https://www.nature.com/articles/s41587-020-00777-4">VAMB</a> (v3.0.2), <a href="http://doi.org/10.1038/s41467-022-29843-y">SemiBin</a> (v1.3.1), <a href="https://zenodo.org/records/10460259">Rosella</a> (v0.4.2), <a href="https://www.nature.com/articles/nmeth.3103">CONCOCT</a> (v1.1.0) and <a href="https://academic.oup.com/bioinformatics/article/32/4/605/1744462">MaxBin2</a> (v2.2.7). Genomes were analyzed using <a href="https://www.nature.com/articles/s41592-023-01940-w">CheckM2</a> (v1.0.2) and clustered at 95% ANI using <a href="https://zenodo.org/records/10526086">Galah</a> (v0.4.0).</p> <p> </p> <p>FILES:</p> <ul> <li><strong>EMERGE_MAGs_2019_long-short-hybrid.tar.gz</strong> - Archive containing the MAG files (.fna).</li> <li><strong>metadata_MAGs_2019_EMERGE.tsv</strong> - Table containing source sample names and accessions, GTDB classifications, CheckM2 quality information, NCBI GenomeBatch- and MIMAG(6.0)-formatted attributes, and other metadata for the MAGs.</li> </ul> <p> </p> <p>FUNDING:</p> <p>This research is a contribution of the EMERGE Biology Integration Institute (<a href="https://emerge-bii.github.io/">https://emerge-bii.github.io/</a>), funded by the National Science Foundation, Biology Integration Institutes Program, Award # 2022070.</p> <p>This study was also funded by the Genomic Science Program of the United States Department of Energy Office of Biological and Environmental Research, grant #s DE-SC0004632. DE-SC0010580. and DE-SC0016440.</p> <p>We thank the Swedish Polar Research Secretariat and SITES for the support of the work done at the Abisko Scientific Research Station. SITES is supported by the Swedish Research Council's grant 4.3-2021-00164.</p> <p>Data from the Joint Genome Institute (JGI) was collected under BER Support Science Proposal 503530 (DOI: <a href="https://doi.org/10.46936/10.25585/60001148">10.46936/10.25585/60001148</a>), conducted by the U.S. Department of Energy Joint Genome Institute (<a href="https://ror.org/04xm1d337">https://ror.org/04xm1d337</a>), a DOE Office of Science User Facility supported by the Office of Science of the U.S. Department of Energy under Contract No. DE-AC02-05CH11231.</p>
Long-read sequencing and structural variant characterization in 1,019 samples from the 1000 Genomes Project
SV analysis of the long-read sequencing data of 1,019 samples from the 1000 Genomes Project. The data is hosted at the International Genome Sample Resource (IGSR) in the <a href="https://ftp.1000genomes.ebi.ac.uk/vol1/ftp/data_collections/1KG_ONT_VIENNA/">1KG_ONT_VIENNA</a> directory. Please see the <a href="https://ftp.1000genomes.ebi.ac.uk/vol1/ftp/data_collections/1KG_ONT_VIENNA/README_1KG_ONT_VIENNA.md">README</a> and <a href="https://ftp.1000genomes.ebi.ac.uk/vol1/ftp/data_collections/1KG_ONT_VIENNA/README_1KG_ONT_VIENNA_datareuse_statement.md">data reuse statement</a> for further information about this dataset.
Machine learning classifiers for species classification of fungi using error-prone long-reads on extended metabarcodes
<p>Machine learning models used in the decision tree of linked machine learning models (<a href="https://github.com/teenjes/fungal_ML">https://github.com/teenjes/fungal_ML</a>)</p>
The SV callsets of the HG002 human sample produced by cuteSV with multi long-read sequencing platforms.
<p>The SV callsets of the HG002 human sample produced by cuteSV with multi long-read sequencing platforms.</p>
Draft genome assembly of a Japanese Oikopleura dioica male individual (O3), using Nanopore long reads.
<p>This draft assembly was used to validate the chrY scaffolds of the OSKA2016 reference genome in the publication “A genome database for a Japanese population of the larvacean Oikopleura dioica”, Development Growth and Differentiation, Wang and coll., 2020 (in press). It is provided as supplemental data for the reproducibility of this work; please note that no further polishing has been done to correct sequencing errors.</p> <p>Genome sequence reads were produced on a MinION sequencer (Oxford Nanopore Technologies) using high-molecular weight DNA from a male individual of the Oikopleura dioica species of zooplankton. The individual was related to the laboratory strain established from a western Japanese population that was used to produce the OSKA2016 reference genome. The raw reads were basecalled with the Guppy software version 3.3.0 using its dna_r9.4.1_450bps algorithm, and deposited in the European Nucleotide Archive (Study ID: PRJEB38559). The draft assembly was made with the Flye software version 2.7 with the options --genome-size 65m and --min-overlap 3000.</p>
Metagenomics assemblies and high-quality MAGs for "Long-read metagenomics to retrieve high-quality metagenome-assembled genomes from canine feces"
<p>This dataset includes the different metagenomics assemblies analyzed and its summary (_info.txt file):</p> <p>- <a href="https://zenodo.org/api/files/3a502803-82f7-4b51-ac62-7c1dfcdcb680/100_assembly.fasta">100_assembly.fasta</a> is the Flye 2.7 metagenomics assembly merging HMW and non-HMW datasets</p> <p>- <a href="https://zenodo.org/api/files/3a502803-82f7-4b51-ac62-7c1dfcdcb680/75_assembly.fasta">75_assembly.fasta</a> is the Flye 2.7 metagenomics assembly including 75% of random data of the merged dataset.</p> <p>- <a href="https://zenodo.org/api/files/3a502803-82f7-4b51-ac62-7c1dfcdcb680/50_assembly.fasta">50_assembly.fasta</a> is the Flye 2.7 metagenomics assembly including 50% of random data of the merged dataset.</p> <p>- <a href="https://zenodo.org/api/files/3a502803-82f7-4b51-ac62-7c1dfcdcb680/HMW_assembly.fasta?versionId=749ff6fd-2642-4ad1-971a-7f3404baa595">HMW_assembly.fasta</a> is the Flye 2.7 metagenomics assembly for HMW dataset.</p> <p>Moreover, it also includes the eight frameshift-corrected high-quality MAGs analyzed in the manuscript. </p>
A survey of the sorghum transcriptome using single-molecule long reads
<p>Alternative splicing and alternative polyadenylation (APA) of pre-mRNAs greatly contribute to transcriptome diversity, coding capacity of a genome and gene regulatory mechanisms in eukaryotes. Second-generation sequencing technologies have been extensively used to analyze transcriptomes. However, a major limitation of short-read data is that it is difficult to accurately predict full-length splice isoforms. Here we sequenced the sorghum transcriptome using Pacific Biosciences single molecule real time long-read isoform sequencing and developed a pipeline called TAPIS (Transcriptome Analysis Pipeline for Isoform Sequencing) to identify full-length splice isoforms and APA sites. Our analysis reveals transcriptome-wide full-length isoforms at an unprecedented scale with over 11,000 novel splice isoforms. Additionally, we uncover APA of ~11,000 expressed genes and more than 2,100 novel genes. These results greatly enhance sorghum gene annotations and aid in studying gene regulation in this important bioenergy crop. The TAPIS pipeline will serve as a useful tool to analyze Iso-Seq data from any organism.</p>
Tandem repeat catalog of the human genome generated from long-read assemblies
<p>Allele sequences of polymorphic loci (VCF) and README for all version 2 (2.0 + 2.1) files</p>
The PODER cross-ancestry long-read RNA-seq dataset
<p>Table legends:</p> <ul> <li>column_descriptions.docx: Extended description of columns in selected tables.</li> <li>00_sample_metadata: Sample and experimental metadata pertaining to each LR-RNA-seq sample.</li> <li>01_uma_gtf: Unfiltered merged annotation (UMA) GTF for all spliced intron chains discovered by each of the tools.</li> <li>02_uma_mt: Metadata table for SQANTI QC, Recount3 support, protein prediction, and other relevant characteristics for each transcript in the UMA annotation used to filter transcripts.</li> <li>03_poder_gtf: Final PODER GTF, including predicted CDSs for novel transcripts and annotated CDSs for known transcripts.</li> <li>04_poder_mt: Metadata table for SQANTI QC, Recount3 support, protein prediction, and other relevant characteristics for each transcript in PODER.</li> <li>06_poder_t_counts: PODER transcript counts in each sample as quantified by lr-kallisto.</li> <li>07_poder_g_counts: PODER gene counts in each sample as quantified by lr-kallisto.</li> <li>08_mage_t_counts_poder: Transcript counts for the MAGE RNA-seq dataset computed using the PODER annotation with kallisto.</li> <li>09_mage_t_counts_gencode: Transcript counts for the MAGE RNA-seq dataset computed using the GENCODE annotation with kallisto.</li> <li>10_mage_t_counts_enh: Transcript counts for the MAGE RNA-seq dataset computed using the Enhanced GENCODE annotation with kallisto.</li> <li>11_astu: Allele-specific transcript usage results.</li> <li>12_ase: Allele-specific expression results.</li> <li>13_mage_tau: Tau values for Enhanced GENCODE transcripts computed on the MAGE RNA-seq dataset.</li> <li>14_gwas_enrichments: GWAS enrichment results for ASTU genes overlapping GWAS genes.</li> <li>16_inter_catalog_overlap: Boolean detection of transcripts across PODER, annotations (GENCODE, RefSeq) and other RNA-seq / LR-RNA-seq-derived transcript catalogs (CHESS, GTEx, ENCODE4).</li> <li>17_personalized_hg38: Isoforms and their intron chains detected using personalized-GRCh38s.</li> <li>18_enh_gencode: Enhanced GENCODE GTF with novel transcripts from PODER added to all annotated transcripts from GENCODE v47</li> </ul>
Companion data deposit of manuscript: Evaluating and improving the representation of bacterial contents in long-read metagenome assemblies
<p>This upload contains the metagenome assemblies and their binning results generated & described in the manuscript "Evaluating and improving the representation of bacterial contents in long-read metagenome assemblies" (preprint version: arxiv2210.00098, "Towards complete representation of bacterial contents in metagenomic samples"). </p> <p>Mapping of sample names in the file names and the descriptors used as in the manuscript can be found in table S1, which is available along with the manuscript and also included in the supplementary_tables_and_figures tar archive here.</p>
Long-read sequencing reveals extensive gut phageome structural variations driven by genetic exchange with bacterial hosts
<p><span>Genetic variations are instrumental for unraveling phage evolution and deciphering their functional implications. Here we explore the underlying fine-scale genetic variations in the gut phageome, especially structural variations (SVs). By employing virome-enriched long-read metagenomics sequencing across 91 individuals, we identified a total of 14,438 non-redundant phage SVs, and revealed their prevalence within the human gut phageome. These SVs are mainly enriched in genes involved in recombination, DNA methylation, and antibiotic resistance. Strikingly, a substantial fraction of phage SV sequences share close homology with bacterial fragments, with most SVs enriched for horizontal gene transfer (HGT) mechanism. Further investigations showed that these SV sequences were genetic exchanged between specific phage-bacteria pairs, particularly between phages and their respective bacterial hosts. Temperate phages exhibits a higher frequency of genetic exchange with bacterial chromosomes then virulent phages. Collectively, our findings provide novel insights into the genetic landscape of the human gut phageome.</span></p>
Supporting data for the manuscript "Severus: accurate detection and characterization of somatic structural variation in tumor genomes using long reads"
<p>Supporting data for the manuscript "Severus: accurate detection and characterization of somatic structural variation in tumor genomes using long reads".</p> <p>The archive contains files that are necessary to reproduce the cell line benchmarks from the paper, including:</p> <ul> <li>Scripts and command lines</li> <li>Original VCF outpurs of all tools used in benchmarking</li> <li>Minda evaluations and truthset VCF files</li> <li>Full Severus outputs + visualizations</li> <li>truvari calls</li> </ul>
Simulation data for benchmarking de novo long read transcriptome assembly software
<p>Method of simulation of differentially expressed biological replicates</p> <p>We first obtained a subset of transcripts that are widely expressed in the GTEx v9 dataset (92 samples) using Gencode comprehensive annotation (v44). We kept transcripts with more than 5 reads in at least 15 samples after Salmon quantification (18145 genes, 40509 transcripts), and stored their mean count per million (CPM) values as the control group’s baseline expression. We then generated a perturbed set of CPM values where transcript expression was changed by: (1) randomly selecting 1000 genes and changing all transcripts belonging to that gene concordantly (500 genes 2 fold up and 500 genes 2 fold down), (2) selected another 1000 genes randomly, and then select 2 random transcripts from the gene and swap their expression, (3) selected another 1000 genes randomly, and then select 1 random transcript to change its expression (500 transcripts 2 fold up and 500 transcripts 2 fold down). The updated CPM were stored as the perturbed group baseline expression. We then generated a count matrix and CPM matrix for 3 control replicates and 3 perturbed replicates with gamma distribution, followed by a Poisson distribution <a href="https://www.zotero.org/google-docs/?cUP4ui">(Baldoni et al., 2024)</a>. Both long-read and short-read FASTQ files were simulated using SQANTI-SIM with default settings and ONT R9.4 cDNA error profile (v 0.2.1) <a href="https://www.zotero.org/google-docs/?Qyopst">(Mestre-Tomás et al., 2023)</a>. The long read data contained 6 million reads in total, and an average read length of 1085 bp, and short read data was 100 bp paired-end. We then subsampled the short-read data to match the total number of base pairs in the long read data (6.5 billion bases). The simulated data was non-stranded, and contains 2000 DE genes, 2000 genes with DTU, 5927 transcripts with DTU and 6933 DE transcripts.</p> <p> </p>
Long reads training material for 'Quality Control' tutorial (Galaxy Training Material)
<p>The data provided here are part of a Galaxy Training Network tutorial for reads Quality Control.</p> <p>PacBio HiFi reads were provided by PacBio - GIAB sample HG002 (https://www.pacb.com/smrt-science/smrt-resources/datasets/) and was downsampled using seqtk (https://github.com/lh3/seqtk)</p> <p>Nanopore reads were provided by Tim Kahlke as part of "Long-Read, long reach Bioinformatics Tutorials" (https://timkahlke.github.io/LongRead_tutorials/) and was basecalled using Guppy v5.0.2 (dna_r9.4.1_450bps_sup.cfg).</p>
Enhanced Protein Isoform Characterization Through Long-Read Proteogenomics - Workflow Results
<pre> </pre> <p>The detection of physiologically relevant protein isoforms encoded by the human genome is critical to biomedicine. Mass spectrometry (MS)-based proteomics is the preeminent method for protein detection, but isoform-resolved proteomic analysis relies on accurate reference databases that match the sample; neither a subset nor a superset database is ideal. Long-read RNA sequencing (e.g. PacBio, Oxford Nanopore) provides full-length transcript sequencing, which can be used to predict full-length proteins. Here, we describe a long-read proteogenomics approach for integrating matched long-read RNA-seq and MS-based proteomics data to enhance isoform characterization. We introduce a classification scheme for protein isoforms, discover novel protein isoforms, and present the first protein inference algorithm for the direct incorporation of long-read transcriptome data in protein inference to enable detection of protein isoforms that are intractable to MS detection. We have released an open-source Nextflow pipeline that integrates long-read sequencing in a proteomic workflow for isoform-resolved analysis.</p> <p>Companion Repositories:</p> <ol> <li><a href="https://doi.org/10.5281/zenodo.5920817">Long-Read-Proteogenomics Workflow GitHub Repository Release</a></li> <li><a href="https://doi.org/10.5281/zenodo.5920847">Long-Read-Proteogenomics Analysis GitHub Repository Release</a></li> </ol> <p>Companion Datasets</p> <ol> <li><a href="https://zenodo.org/deposit/5703754">Long-Read-Proteogenomics Workflow Sample and Reference Data</a></li> <li><a href="https://doi.org/10.5281/zenodo.5234651">TEST Data for Long-Read-Proteogenomics Workflow GitHub Actions</a></li> </ol> <p>This Repository contains the complete output from the execution of the <a href="https://doi.org/10.5281/zenodo.5920817">Long-Read-Proteogenomics Workflow</a>, using the input from <a href="https://zenodo.org/deposit/5703754">Jurkat Samples and Reference Data</a>. </p> <p>The file <em>jurkat.flnc.bam </em>was 6.5 GB had to be split into 13 separate files and for use should be rejoined -- here are the steps that were used to split the file up. </p> <p>1. Convert <em>jurkat.flnc.bam</em> (binary format) to sam file (text format) without header: <em>samtools view jurkat.flnc.bam > jurkat.flnc.sam</em></p> <p>2. Capture the header: <em>samtools view -H jurkat.flnc.bam > jurkat.flnc.header.sam</em></p> <p>3. Split <em>jurkat.flnc.sam</em> into smaller files (aim to get final size under 2GB): <em>split -l 400000 jurkat.flnc.sam jurkat.flnc.chunk.</em></p> <p>4. Convert each of these files back to bam for uploading: <em>samtools view -b jurkat.flnc.chunk.a* -o jurkat.flnc.chunk.a*.bam (*=a,b,c,d,e,f,g,h,i,j,k,l,m)</em></p> <p>After downloading, reverse this process including using the header file which is found in the LRPG-Manuscript-Results-results-results-jurkat-isoseq3-companion-files.tar.gz file></p> <p>1. Convert the bam files back to sam files: <em>samtools view jurkat.flnc.chunk.a*.bam > jurkat.flnc.chunk.a*.sam (*=a,b,c,d,e,f,g,h,i,j,k,l,m)</em></p> <p>2. Combine the header together with the sam files: <em>cat jurkat.flnc.chunk.a*sam > jurkcat.flnc.sam (</em>verified the same number of lines of the sam files is identical to the number of lines of the original without header: 4,956,761. Header file is 13 lines.</p> <p>3. Convert to bam files if desired: <em>samtools view -b jurkat.flnc.sam -o jurkat.flnc.bam</em></p> <p>4. Rehead with the header file: <em>samtools reheader -P -i jurkat.flnc.header.sam jurkat.flnc.bam</em></p>
Long read proteogenomics to characterize protein isoform diversity in human umbilical vein endothelial cells (HUVECs)
<p>Endothelial cells (ECs) comprise the lumenal lining of all blood vessels and are critical for the functioning of the cardiovascular system and their phenotypes can be modulated by protein isoforms. To characterize the isoform landscape within EC, we applied a long read proteogenomics approach to analyze human umbilical vein endothelial cells (HUVECs). Transcripts delineated from PacBio sequencing serve as the basis for a sample-specific protein database used for downstream MS analysis to infer protein isoform expression. We detected 53,836 transcript isoforms from 10,426 genes, with 22,195 of those transcripts being novel. Furthermore, the predominant isoform in HUVECs does not correspond with the accepted “reference isoform” 25% of the time, with vascular pathway-related genes among this group. We found 2,597 protein isoforms supported through unique peptides, with an additional 2,280 isoforms nominated upon incorporation of long-read transcript evidence. We characterized a novel alternative acceptor for endothelial-related gene <em>CDH5</em>, suggesting potential changes in its associated signaling pathways. Finally, we identified novel protein isoforms arising from a diversity of splicing mechanisms supported by uniquely mapped novel peptides. Our results represent a high resolution atlas of known and novel isoforms of potential relevance to endothelial phenotypes and function.</p>
Parent-of-origin detection and chromosome-scale haplotyping using long-read DNA methylation sequencing and Strand-seq
<p>Hundreds of loci in human genomes have alleles that are methylated differentially according to their parent of origin. These imprinted loci generally show little variation across tissues, individuals, and populations. We show that such loci can be used to distinguish the maternal and paternal homologs for all autosomes, without the need for the parental DNA. We integrate methylation-detecting nanopore sequencing with the long-range phase information in Strand-seq data to determine the parent of origin of chromosome-length haplotypes for both DNA sequence and DNA methylation in five trios with diverse genetic backgrounds.</p>
Genome annotations of Drosophila melanogaster and Drosophila simulans wild-type strains from long read sequencing assemblies
<p>Genome assemblies were performed for eight wild-type strains of Drosophila melanogaster and Drosophila simulans from Oxford Nanopore long read sequencing (please refer to Mohamed et al. Cells 2020 (doi:10.3390/cells9081776)). Assemblies were deposited in the European Nucleotide Archive (ENA) at EMBL-EBI under accession number PRJEB50024 (<a href="https://www.ebi.ac.uk/ena/browser/view/PRJEBxxxx">https://www.ebi.ac.uk/ena/browser/view/</a>PRJEB50024).</p> <p>Transposable Element annotations: we used RepeatMasker 4.1.0 (<a href="http://repeatmasker.org/">http://repeatmasker.org/</a>) -species Drosophila, followed by OneCodeToFindThemAll (Bailly-Bechet et al. 2014) with default parameters.</p> <p>Gene annotations: We retrieved gtf files from FlyBase : <a>ftp.flybase.net/genomes/D</a><a>rosophila_melanogaster/dmel_r6,46_FB2022_03/gft/dmel-all-r6.46.gtf.gz</a> and <a>ftp.flybase.net/genomes/Drosophila_simulans/dsim_r2,02_FB2017_04/gtf/dsim-all-</a><a>r2,02.gtf.gz</a>. The corresponding fasta files were also downloaded from FlyBase: <a>ftp.flybase.net/genomes/Drosophila_melanogaster/dmel_r6,46_FB2022_03/</a><a>fasta</a><a>/dmel-all-</a><a>chromosome-</a><a>r6.46.</a><a>fasta</a><a>.gz</a> and <a>ftp.flybase.net/genomes/Drosophila_simulans/dsim_r2,02_FB2017_04/</a><a>fasta</a><a>/dsim-all-</a><a>chromosome-</a><a>r2,02.</a><a>fasta</a><a>.gz</a>. We used Liftoff (Shumate and Salzberg, 2020) to lift over gene annotations from the references to our genome assemblies. We used -flank 0.2 and only kept the “gene” and “exon” terms.</p>
Data for "A replicable and modular benchmark for long-read transcript quantification methods"
<p>This archive contains the input necessary to run the inital (TranSigner-protocol and IsoQuant-protocol) benchmarks associated with the <a href="https://github.com/COMBINE-lab/lr_quant_benchmarks" target="_blank" rel="noopener"><code>lr_quant_benchmarks repository</code></a>. The archive can be decompressed with <code>tar</code> and <code>zstd</code> using the command <code>tar --use-compress-program=zstd -xf input.tar.zstd</code>.</p>
Supplementary files for the manuscript "Assembly of Long Error-Prone Reads Using Repeat Graphs"
<p>Supplementary files for the manuscript "Assembly of Long Error-Prone Reads Using Repeat Graphs"</p> <p> </p> <p>Contents<br> ---------</p> <p>* `human_assemblues` - Flye assemblies of the human ONT sequencing data + QUAST benchmarking<br> of Flye, Canu and MaSuRCA assemblies. Scripts for assembly graph analysis are also included.</p> <p>* `nctc_assemblis` - Flye assemblies of the NCTC 21 bacterial dataset.</p> <p>* `yeast_assemblies` - working directories Flye, Canu, Falcon, Hinge and Miniasm assemblies of <br> yeast PB and ONT datasets + final assemblies + quast report. Some large files <br> (such as read alignments) were deleted.</p> <p>* `worm_assemblies` - working directories Flye, Canu, Falcon, Hinge and Miniasm assemblies of <br> the c. elegans dataset + final assemblies + quast report. Some large files <br> (such as read alignments) were deleted. `tandem_misassemblies` directory contain<br> the detailed analysis of nine tandem misassemblies. We recommend "gepard" dot-plotter for visualization.</p> <p>* `metagenome_assemblies` - Flye and Canu assemblies of a PacBio mock metagenome dataset.<br> In addition to metagenome assemblies, each bacteria was reassembled separately to<br> estimate the rate of divergence between the target genomes and the available references.</p> <p>* `simulated_data` - two assemblies of the simulated data illustrating Figure 1 (from Appendix I),<br> as well as simulated unbridged repeats benchmark.</p> <p><br> Software versions and parameters<br> --------------------------------</p> <p>* Flye - 2.3.5 (commit 20afeda)<br> * Canu - 1.7.1 (commit dfa60b8)<br> * Falcon - 0.3.0 (FALCON-Integrate commit 7498ef9)<br> * HINGE - 0.5.0 (commit 79fdf66)<br> * Miniasm - 0.2-r168-dirty (commit 40ec280) / Minimap2 2.8-r711-dirty (commit 8fc5f8d)<br> * Quast - 5.0.0 (commit de6973bb)</p> <p>Flye and Canu were run with the default parameters. The config files / scripts for<br> Falcon, HINGE and Miniasm could be found in the 'asm_config' archive folder.</p> <p>The HUMAN (but not the HUMAN+) assembly was generated with the earlier <br> Flye version 2.3.2 (released on Feb 20 2018) to provide a fair comparison <br> with the Canu and MaSuRCA assemblies (which were not updated since the release of Flye 2.3.2).<br> We note that the HUMAN assembly using the latest Flye version 2.3.5 has <br> NGA50 = 7.3 Mb and improves over the Flye 2.3.2 assembly (NGA50 = 6.3Mb). <br> HUMAN+ was assembled using the latest Flye and Canu versions (as of September 2018).</p> <p>The code for unbridged repeat resolution is currently available <br> in a separate 'flye-trestle' branch (commit 6100d32)</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.