Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
19,297
datasets available to search
ShareScore release 0.7.1
Dataset results
19,297 results for “Sequencing”
Database of fitted spectra for: Changing-Look AGNs - I. Tracking the transition on the main sequence of quasars
<h3>Results from the spectral fitting for a sample of changing-look active galactic nuclei (AGNs) with SDSS spectroscopy using PyQSOFit.</h3>
MMSEQS meets AntiRef: reference clusters of human antibody sequences
<p>This data set contains pre-computed mmseqs databases for the antiref fasta files created by <em>Briney et al.</em> </p> <p>Please cite the original work if you use any of the databases provided here.</p> <p>Sources:</p> <ul> <li><a href="https://github.com/brineylab/antiref">Antiref GitHub</a></li> <li><a href="../records/7474336">Antiref Zenodo</a></li> <li><a href="https://academic.oup.com/bioinformaticsadvances/article/3/1/vbad109/7247530?login=true">Antiref Paper</a></li> </ul> <p> </p> <p>The mmseqs databases were created as follows:</p> <p> </p> <p>```</p> <p>aria2x -x16 -s16 --input-file antiref_links.txt<br>snakemake -s antiref_mmseqs.smk --jobs 1 --cores 1 --local-cores 250</p> <p>```</p>
Supplementary dataset to publication: Complete Genome Sequence of Ovine Mycobacterium avium subsp. paratuberculosis Strain JIII-386 (MAP-S/type III) and Its Comparison to MAP-S/type I, MAP-C, and M. avium Complex Genomes.
<p>This is the modified supplemented material to the publication “Complete genome sequence of ovine Mycobacterium avium subsp. paratuberculosis strain JIII-386 (MAP-S/type III) and its comparison to MAP-S/type I, MAP-C, and M. avium complex genomes”.</p> <p>The complete circular genome of Mycobacterium avium subsp. paratuberculosis (MAP) strain JIII-386 from Germany, closed by Nanopore technology in this study, was presented and compared with the draft genome of JIII-386, previously published in [doi:10.1093/gbe/ew154], the closed genome of the MAP-S/type I strain Telford, the MAP-S/type III draft genome of strain S397, twelve closed MAP-C (type II) strains and eight closed Mycobacterium avium (M. a.) strains of subsp. hominissuis (MAH) and subsp. avium (MAA). Structural comparisons clearly revealed the mosaic nature of MAP genomes, the differences between MAP subtypes I, II and III, and the higher diversity of MAP-S compared to MAP-C genomes. </p> <p>The material provides a wealth of detailed results from these analyses and comparisons. These include a list of identified ncRNA and Riboswitches, as well as additional genes in finished JIII-386, the gene content of identified prophage regions, copy number of identified transposable elements and a list of selected virulence-associated genes in the different MAP-type (I - III) strains. The genomic islands identified and included genes along with their predicted functions were presented for six MAP genomes (belonging to MAP-S/type I and III, and MAP-C), one MAH genome and one MAA genome. One table shows the corresponding genomic islands in the genomes of JIII-386, Telford and three MAP-C genomes. Furthermore, homologous genes of known MAP-S specific Large Sequence Polymorphisms regions (LSP<sup>S</sup> = LSP-S) were recorded in different MAP-S type strains, one MAH and one MAA strain, as well as genes of deletions #1 (LSP<sup>A</sup>-20), #2, and s-delta-1, previously described as MAP-S-specific deletions, their presence or absence in 3 MAP-S, 12 MAP-C, 4 MAH, and 4 MAA strains were listed. Different presence or absence of genes, but also identified frameshifts or disruptions of various virulence-associated genes could lead to the different MAP-type specific phenotypic characteristics. Comprehensive core and pan genome analyses (results listed in six tables) revealed unique genes and genes likely to have been acquired by horizontal gene transfer in different MAP types and subtypes, but also emphasized the highly conserved and close relationship, and the complex evolution of M. a. strains.</p> <p> </p>
Peat characteristics, microbial PLFA, and fungal and actinobacterial sequences from Lakkasuo peatland drainage experiment, year 2004
<p>We analysed the response of microbial communities, characterized by phospholipid fatty acids (PLFAs), and fungal and actinobacterial communities, characterized by PCR-DGGE fingerprinting and direct sequencing, to changing hydrological conditions at three different sites in the boreal peatland complex Lakkasuo in southern Finland. Additionally, several peat characteristics were measured. The experimental design involved undrained controls as well as short-term (3 years) and long-term (43 years) water-level drawdown. The sites were, in their undrained state, a herb-rich sedge fen, a sedge fen, and a bog with hummock-lawn-hollow microtopography.</p> <p>Codes are explained in the Notes sheets of the Excel files. The contents of the csv files are identical to the corresponding Excel file data sheets.</p> <p>Please check the decimal separator! Comma is used in Finland, and that may have been carried over. All commas in data columns are decimal separators.</p>
Protein haplotype sequences obtained by ProHap from the Haplotype Reference Consortium Release 1.1 dataset
<p>Database of protein sequences obtained using ProHap (<a href="https://github.com/ProGenNo/ProHap">https://github.com/ProGenNo/ProHap</a>) on the data set of phased genotypes published by the Haplotype Reference Consortium, Release 1.1 (<a href="https://ega-archive.org/datasets/EGAD00001002729" target="_blank" rel="noopener">https://ega-archive.org/datasets/EGAD00001002729</a>). We used Ensembl v.110 for the mapping of coordinates between genes, exons, and transcripts.</p> <p>Release 1.1 of the HRC is provided aligned with the GRCh37 reference genome. We have performed a liftover to the GRCh38 reference using GeneBe (https://genebe.net/tools/liftover). Variants for which the reported alternative allele is considered as reference in GRCh38 were removed. A threshold of 1% minor allele frequency was applied to filter the remaining variants. After translation, a frequency threshold of 0.5% was applied to filter the resulting unique non-canonical sequences. The complete configuration file for the ProHap run is attached to this repository.</p> <p>This dataset contains one compressed directory, contains the following files:</p> <ul> <li>F1: The concatenated fasta file ready to be used with search engines, contains the following: <ul> <li>Protein haplotype sequences obtained by ProHap</li> <li>Reference proteome as per Ensembl v. 110</li> <li>Contaminant sequences from the cRAP project (<a href="https://www.thegpm.org/crap/">https://www.thegpm.org/crap/</a>)</li> <li>The file is provided in two formats - full and simplified. The simplified fasta contains only the artificial protein identifier and the matching gene name, and is optimised for compatibility with a wide range of tools. For annotation of peptides using the PeptideAnnotator, please provide the header (F1.2) in addition to the fasta file. </li> </ul> </li> <li>F2: Additional information about the haplotype sequences, to be used for mapping identified peptides to the original haplotypes</li> <li>F3: Translations of haplotype cDNA sequences, before merging with the reference proteome</li> </ul> <p>For further description of the files, please refer to <a href="https://github.com/ProGenNo/ProHap/wiki/Output-files">https://github.com/ProGenNo/ProHap/wiki/Output-files</a>.</p> <p>For the usage of these databases with search engines, and downstream anaylsis of identified peptides, please refer to the project's wiki page: <a href="https://github.com/ProGenNo/ProHap/wiki/Using-the-database-for-proteomic-searches">https://github.com/ProGenNo/ProHap/wiki/Using-the-database-for-proteomic-searches</a>.</p> <p>When using these databases in your publication, please cite: Vašíček, J., Kuznetsova, K.G., Skiadopoulou, D. <em>et al.</em> ProHap enables human proteomic database generation accounting for population diversity. <em>Nat Methods</em> (2024). <a href="https://doi.org/10.1038/s41592-024-02506-0">https://doi.org/10.1038/s41592-024-02506-0</a></p>
Motor sequence learning
Open the record for dataset details and reuse information.
Finger sequence planning
Open the record for dataset details and reuse information.
Darwin: an amino acid sequence collection of complete proteomes from eukaryotes with different phylogenetic affinities (v. 03_2020_137)
<p><strong>Background</strong></p> <p>Every time we find an interesting gene in an organism of interest, the first question is often “how widely is this gene distributed in the eukaryotic kingdom?”. Naturally, one could use NCBI BLAST search against the non-redundant sequence database provided by GenBank to answer this question. However, it can be cumbersome to parse the results and assign them to taxonomic units. It is also not straightforward to get an overview of which eukaryotic groups are represented in the results. Top BLAST hits can be crowded with sequences from closely-related organisms making it difficult gain an overview of the overall distribution across eukaryotes. To streamline this process, we developed an in-house database of complete eukaryotic proteomes. We tagged each sequence with a eukaryotic group handle (two-character symbol) and combined them into a single data set searchable by standalone BLAST on one’s own computer. We named this data set “Darwin” to reflect the diverse nature of the sequences it contains. </p> <p><strong>Methods</strong></p> <p>We downloaded predicted proteomes in FASTA format from different sources such as GenBank, Joint Genome Institute (Depart of Energy, USA), Broad Institute (Massachusetts Institute of Technology, USA), Phytozome and a number of other specialized websites catering for a specific organism such as the Arabidopsis Information Resource (TAIR), or the Saccharomyces Genome Database (SGD). All the organisms we included in Darwin are listed in Table 1. To reduce redundancy, we took care not to include the same species more than once unless subspecies were known to show wide diversity. Each sequence header was tagged with a eukaryotic group handle composed of two-character symbols (based on Keeling <em>et al</em>., 2005). These handles clearly appear in BLAST output and can be parsed easily. We combined sequences from all proteomes into a single data set and named it “Darwin”.</p> <p><strong>Results</strong></p> <p>The current version of Darwin (v. 03_2020_137) contains 2,601,132 amino acid sequences from 137 eukaryotes (Table 1, Data file 1). The sizes of the proteomes were diverse, ranging from ~4000 sequences in some alveolates to 60,000-76,000 in plants. Darwin represents most of the supergroups of eukaryotic kingdom described in Keeling <em>et al.,</em> (2005) except those in Rhizaria whose genomes were not available at the time of data set construction. The data set contains larger numbers of proteomes from fungi and plants reflecting areas of interest in our group. </p> <p><strong>Conclusions</strong></p> <p>Darwin is provided as a text fasta file that can be formatted for BLAST searches on standalone computers. The results from the BLAST searches can be parsed to determine how widely a gene of interest is distributed among different eukaryotes. Simple counting of the eukaryotic group handles would also yield an overview of the distribution across taxa. Darwin is also useful for rapidly finding out whether a gene is missing in particular taxa.</p> <p><strong>Reference</strong></p> <p>Keeling PJ, Burger G, Durnford DG, Lang BF, Lee RW, Pearlman RE, Roger AJ, Gray MW (2005) The tree of eukaryotes. <em>Trends Ecol. Evol.</em> <strong>20:</strong> 670-676</p>
Fine-scale structure of the 2016-2017 Central Italy Seismic Sequence from data recorded at the Italian National Network
<p><strong>Data Set </strong></p> <p>Catalog of 33,983 earthquakes located during the 2016-2017 Central Italy seismic sequence. The velocity model used is the 1D gradient P- and S-wave velocity models (after Carannante et al., 2013). We used the highest quality P- and S-wave arrival times manually picked by analysts of the National Institute of Geophysics and Volcanology (INGV) seismic monitoring room, having an uncertainty lower than 0.6 s. </p> <p>Events were located by means of a 2-step procedure: the INGV routine absolute locations computation for all events with ML ≥ 1.5 that occurred in the study area between August 2016 and January 2018, using the method described in Chiaraluce et al. (2017); the determination of relative locations by applying the HypoDD code (Waldhauser, 2001) to the catalog picks and phase delay times measured from waveform cross correlation.</p> <p>The time domain cross-correlation method (Schaff et al., 2004; Schaff and Waldhauser, 2005) was applied to seismograms of all pairs of events separated by 3 km or less and recorded at common stations. Seismograms were filtered in the 1-15 Hz frequency range using a 4 pole, zero phase band‐pass Butterworth filter. The correlations measurements were performed on 0.7 s long window for P-waves and 1 s windows for S-waves. Only measurements with correlation coefficients greater than 0.7 were kept, resulting in a total of ~4.4 million P and ~1.1 million S wave delay times. </p> <p>We sub-divided the entire dataset in 18 rectangular boxes, containing a maximum of 6000 earthquakes, orthogonal to and centered on the mean strike of the seismic sequence. The overlap between neighboring boxes is 50% with respect to the NW-SE extension. HypoDD is run separately on each box. Resulting relative locations from all boxes were combined into a single catalog, computing the weighted mean of double hypocenters in the overlapping regions (Waldhauser and Schaff, 2008).</p> <p>The final double-difference catalog includes 33,982 events occurring between 24<sup>th</sup> of August 2016 and 18<sup>th</sup> of January 2018.</p> <p>The catalog is in csv format, semicolon separator, ordered by origin time and the header content is the following:</p> <ul> <li>Id-ingv: ingv eventid, useful to link to the QuakeML phase file through the INGV fdsnws/event webservice (<a href="https://meet.google.com/linkredirect?authuser=0&dest=http%3A%2F%2Fwebservices.ingv.it%2Fswagger-ui%2Fdist%2F%3Furl%3Dhttps%3A%2F%2Fingv.github.io%2Fopenapi%2Ffdsnws%2Fevent%2F0.0.1%2Fevent.yaml">http://webservices.ingv.it/swagger-ui/dist/?url=https://ingv.github.io/openapi/fdsnws/event/0.0.1/event.yaml</a>) and to the reported magnitude;</li> <li>Latitude(°) expressed in decimal degrees;</li> <li>Longitude(°) expressed in decimal degrees;</li> <li>Depth(km) hypocentral depth expressed in kilometers;</li> <li>Year of origin time in the format yyyy;</li> <li>Month of origin time in the format mm;</li> <li>Day of origin time in the format dd; </li> <li>Hour of origin time in the format hh;</li> <li>Minute of origin time in the format min;</li> <li>Second of origin time in the format ??.?????? s;</li> <li>Magnitude: the value available at the phases downloading time (see Id-ingv fdsnws/event)</li> </ul> <p> </p> <p> </p> <p> </p> <p><br> </p>
Phylogeny of "Philoceanus complex" seabird lice (Phthiraptera: Ischnocera) inferred from mitochondrial DNA sequences
<p>Data from "Phylogeny of “<em>Philoceanus </em>complex” seabird lice (Phthiraptera: Ischnocera) inferred from mitochondrial DNA sequences". See the file index.html for details. Data includes NEXUS files for sequences, tree files output by MrBayes and PAUP, and host-parasite association files for TreeMap.</p>
Overcoming Limitation of AlphaFold2 by Deep-mutational Scanning and Stability-Selection of Protein Sequences
<p>This repository contains the processed datasets and corresponding code used in our study. While AlphaFold2 revolutionizes protein structure prediction, its accuracy critically depends on evolutionary information from natural homologs—limiting applications for proteins with sparse sequence families. Here, we bypass this bottleneck by employing deep mutational scanning and stability-guided selection to generate artificial homologs. Fed into AlphaFold2, these synthetic sequences match the accuracy achieved on well-predicted proteins with rich natural homology, while providing highly accurate predictions for difficult targets—including orphan proteins previously deemed "unpredictable." Our approach achieves high accuracy (<3 Å RMSD for 5/8 and <2 Å RMSD for 7/8 targets after excluding intrinsically flexible regions). Thus, integrating simple, scalable molecular biology (mutagenesis/selection) with high-throughput sequencing can deliver the accuracy similar to but at a fraction of the cost and time of traditional experimental structure-determination methods. This hybrid framework could democratize high-resolution structural biology, opening avenues to determine structures of protein complexes, modified proteins, and condition-dependent conformations. </p>
Public sequence accessions from INSDC, COG-UK and CNCB and EPI_SET from GISAID for SARS-CoV-2 genome sequences in 2023-08-01 UShER tree
<p>Genome sequences and metadata for the accessions in the .tsv.gz (gzip-compressed tab-separated text) files are freely available from their corresponding sources:</p><ul><li>insdc.accessionNameDate.tsv.gz: INSDC (GenBank, ENA, DDBJ) sequences and metadata may be downloaded using NCBI Datasets: https://www.ncbi.nlm.nih.gov/datasets/taxonomy/2697049/ (7,361,734 accessions used on 2023-08-01)</li><li>cog.accessionNameDate.tsv.gz: COG-UK sequences and metadata may be downloaded from https://cog-uk.s3.climb.ac.uk/phylogenetics/latest (as of publication); most COG-UK sequences have been submitted to ENA and are available from INSDC/NCBI Datasets as well. (724,978 accessions used on 2023-08-01)</li><li>cncb.accessionNameDate.tsv.gz: Sequences and metadata from several databases at the China National Center for Bioinformation (CNCB) may be downloaded from GenBase: https://ngdc.cncb.ac.cn/genbase/ (26,604 accessions used on 2023-08-01)</li></ul><p>GISAID data are subject to restrictions on sharing described in https://gisaid.org/terms-of-use/. Genome sequences and metadata are available to registered GISAID users as part of EPI_SET_231106ax at https://doi.org/10.55876/gis8.231106ax (7,718,061 accessions used on 2023-08-01).</p>
Strong sequence dependence in RNA/DNA hybrid strand displacement kinetics supplementary data and code
<p>Supplementary data and code needed to replicate figures and results for the paper: Strong sequence-dependence in RNA/DNA hybrid strand displacement kinetics - Francesca G. Smith, John P. Goertz, Molly M. Stevens and Thomas E. Ouldridge. README is included to explain each folder and file in the repository.</p>
Timema genome sequences and annotations. Version 8.
<p>Genome sequence (fasta) files and annotation (gff) files for ten <em>Timema </em>species: <em>T. bartmani, T. cristinae, T. poppensis, T. californicum, T. podura, T. tahoe, T. monikensis, T. douglasi, T. shepardi, and T. genevievae.</em><br> <br> Species are abbreviated as follows: Tbi = <em>T. bartmani</em>, Tce = <em>T. cristinae</em>, Tps = <em>T. poppensis</em>, Tcm = <em>T. californicum</em>, Tpa = <em>T. podura</em>, Tte = <em>T. tahoe</em>, Tms = <em>T. monikensis</em>, Tdi = <em>T. douglasi</em>, Tsi = <em>T. shepardi</em>, and Tge = <em>T. genevievae</em><br> </p> <p>For details of assembly and annotation see: <br> <br> Jaron, K. S*., Parker, D. J*., Anselmetti, Y., Tran Van, P. T., Bast, J., Dumas, Z., Figuet, E., François, C. M., Hayward, K., Rossier, V., Simion, P., Robinson-Rechavi, M., Galtier, N., Schwander, T. 2021. Convergent consequences of parthenogenesis on stick insect genomes. bioRxiv. doi: https://doi.org/10.1101/2020.11.20.391540</p> <p> </p> <p><strong>File list:</strong><br> <br> Tbi_b3v08.fasta = T. bartmani genome sequence file<br> Tbi_b3v08.max_arth_b2g_droso_b2g.gff = T. bartmani genome annotation file<br> Tce_b3v08.fasta = T. cristinae genome sequence file<br> Tce_b3v08.max_arth_b2g_droso_b2g.gff = T. cristinae genome annotation file<br> Tcm_b3v08.fasta = T. bartmani genome sequence file<br> Tcm_b3v08.max_arth_b2g_droso_b2g.gff = T. californicum genome annotation file<br> Tdi_b3v08.fasta = T. douglasi genome sequence file<br> Tdi_b3v08.max_arth_b2g_droso_b2g.gff = T. douglasi genome annotation file<br> Tge_b3v08.fasta = T. genevievae genome sequence file<br> Tge_b3v08.max_arth_b2g_droso_b2g.gff = T. genevievae genome annotation file<br> Tms_b3v08.fasta = T. monikensis genome sequence file<br> Tms_b3v08.max_arth_b2g_droso_b2g.gff = T. monikensis genome annotation file<br> Tpa_b3v08.fasta = T. podura genome sequence file<br> Tpa_b3v08.max_arth_b2g_droso_b2g.gff = T. podura genome annotation file<br> Tps_b3v08.fasta = T. poppensis genome sequence file<br> Tps_b3v08.max_arth_b2g_droso_b2g.gff = T. poppensis genome annotation file<br> Tsi_b3v08.fasta = T. shepardi genome sequence file<br> Tsi_b3v08.max_arth_b2g_droso_b2g.gff = T. shepardi genome annotation file<br> Tte_b3v08.fasta = T. tahoe genome sequence file<br> Tte_b3v08.max_arth_b2g_droso_b2g.gff = T. tahoe genome annotation file</p>
FASTA file containing to the MYB encoding gene Ant1 genomic sequences corresponding to wild and cultivated tomato accessions
<p>Fasta sequence correspond to the MYB encoding gene <em>An2-like</em>. The genomic sequences correspond to <em>Solanum galagpagnese</em> accession LA1141 (this study), <em>S. lycopersicum</em> variety OH8245 (this study), <em>S. lycopersicum</em> variety Heinz 1706 reference genome, and 84 tomato accessions published as part of The 100 Tomato Genome Sequencing Consortium (The 100 Tomato Genome Sequencing Consortium et al., 2014). Local sequences databases were made and retrieved using BLAST version/2018-08 for 84 accessions from The 100 Tomato Genome Sequencing Consortium (The 100 Tomato Genome Sequencing Consortium et al., 2014). Sequences corresponding to Heinz 1706 (Hosmani et al., 2018), were accessed using the Basic Local Alignment Search Tool (BLAST) tool available from the Sol Genomics Network (SGN) (available at <a href="https://solgenomics.net/tools/blast/">https://solgenomics.net/tools/blast/</a>).</p>
FASTA file containing the MYB encoding gene An2-like genomic sequences corresponding to wild and cultivated tomato accessions
<p>FASTA sequence corresponds to the MYB encoding gene <em>An2-like</em>. The genomic sequences correspond to <em>Solanum galagpagnese</em> accession LA1141 (this study), <em>S. lycopersicum</em> variety OH8245 (this study), <em>S. lycopersicum</em> variety Heinz 1706 reference genome (Hosmani et al., 2019), <em>S. lycopersicum </em>variety Indigo Rose (Yan et al., 2020), <em>S. lycopersicum</em> accession LA1996 [MN242011.1 (Colanero et al., 2020)], <em>S. chilense </em>accession LA1930 [MN242012.1 (Colanero et al., 2020)], and 84 tomato accessions published as part of The 100 Tomato Genome Sequencing Consortium (The 100 Tomato Genome Sequencing Consortium et al., 2014). Local sequences databases were made and retrieved using BLAST version/2018-08 for 84 accessions from The 100 Tomato Genome Sequencing Consortium (The 100 Tomato Genome Sequencing Consortium et al., 2014). Sequences corresponding to Heinz 1706 (Hosmani et al., 2018), Indigo Rose [MN433087 (Yan et al., 2020)], <em>S. lycopersicum </em>accession LA1996 [MN242011.1, EF433417.1 (Sapir et al., 2008; Colanero et al., 2020)], <em>S. chilense</em> accession LA1930 [MN242012.1 (Colanero et al., 2020)] were accessed using the Basic Local Alignment Search Tool (BLAST) tool available from the Sol Genomics Network (SGN) (available at <a href="https://solgenomics.net/tools/blast/">https://solgenomics.net/tools/blast/</a>) and the National Center for Biotechnology Information (NCBI)(available at NCBI: <a href="https://www.ncbi.nlm.nih.gov">https://www.ncbi.nlm.nih.gov</a>).</p>
FASTA file containing the MYB encoding genes at the Aft locus with genomic sequences corresponding to wild and cultivated tomato accessions
<p>FASTA sequences correspond to the MYB encoding genes <em>An2-like </em>and <em>Ant1</em>. The genomic sequences were combined correspond to <em>Solanum galagpagnese</em> accession LA1141 (this study), <em>S. lycopersicum</em> variety OH8245 (this study), <em>S. lycopersicum</em> variety Heinz 1706 reference genome (Hosmani et al., 2019), LA1996 [MN242011.1, EF433417.1(Sapir et al., 2008; Colanero et al., 2020)], and 84 tomato accessions published as part of The 100 Tomato Genome Sequencing Consortium (The 100 Tomato Genome Sequencing Consortium et al., 2014). Local sequences databases were made and retrieved using BLAST version/2018-08 for 84 accessions from The 100 Tomato Genome Sequencing Consortium (The 100 Tomato Genome Sequencing Consortium et al., 2014). Sequences corresponding to Heinz 1706 (Hosmani et al., 2018), <em>S. lycopersicum </em>accession LA1996 [MN242011.1, EF433417.1 (Sapir et al., 2008; Colanero et al., 2020)], <em>S. chilense</em> accession LA1930 [MN242012.1 (Colanero et al., 2020)] were accessed using the Basic Local Alignment Search Tool (BLAST) tool available from the Sol Genomics Network (SGN) (available at <a href="https://solgenomics.net/tools/blast/">https://solgenomics.net/tools/blast/</a>) and the National Center for Biotechnology Information (NCBI) (available at NCBI: <a href="https://www.ncbi.nlm.nih.gov/">https://www.ncbi.nlm.nih.gov</a>).</p>
xPore: Identification of differential RNA modifications from nanopore direct RNA sequencing
<p>xPore is a Python package for identification and quantification of differential RNA modifications from direct RNA sequencing.</p> <p>The detailed usage is documented at <a href="https://xpore.readthedocs.io/en/latest/">https://xpore.readthedocs.io/en/latest</a>, while all scripts and source code are available at <a href="https://github.com/GoekeLab/xpore">https://github.com/GoekeLab/xpore</a>.</p> <p>All the preprocessed datasets used in the paper are provided here. </p> <p>Please cite our paper below when using these data.<br> Ploy N. Pratanwanich et al. "Detection of differential RNA modifications from direct RNA sequencing of human cell lines." bioRxiv (2020).</p>
Clonal decomposition and DNA replication states defined by scaled single cell genome sequencing
<p><strong>OV2295 Tables</strong></p> <p>ov2295_breakpoint_counts.csv.gz: Table of breakpoint counts per cell</p> <ul> <li>prediction_id: identifier for the breakpoint</li> <li>cell_id: identifier for the cell</li> <li>read_count: number of reads</li> <li>library_id: identifier for the DNA library</li> <li>sample_id: identifier for the sequenced sample</li> <li>chromosome_1: chromosome of breakend 1</li> <li>strand_1: orientation of break end 1</li> <li>position_1: position of break end 1</li> <li>chromosome_2: chromosome of breakend 2</li> <li>strand_2: orientation of break end 2</li> <li>position_2: position of break end 2</li> </ul> <p>ov2295_cell_cn.csv.gz: Table of cell specific copy number</p> <ul> <li>cell_id: identifier for the cell</li> <li>sample_id: identifier for the sequenced sample</li> <li>library_id: identifier for the DNA library</li> <li>chr: chromosome of bin</li> <li>start: start of bin</li> <li>end: end of bin</li> <li>reads: number of reads</li> <li>copy: raw normalized copy number</li> <li>state: copy number state</li> <li>gc: percent gc of the bin</li> <li>map: average mappability of the bin</li> </ul> <p>ov2295_cell_metrics.csv.gz: Table of cell metrics</p> <ul> <li>cell_id: identifier of the cell</li> <li>unpaired_mapped_reads: number of unpaired mapped reads</li> <li>paired_mapped_reads: number of mapped reads that were properly paired</li> <li>unpaired_duplicate_reads: number of unpaired duplicated reads</li> <li>paired_duplicate_reads: number of paired reads that were also marked as duplicate</li> <li>unmapped_reads: number of unmapped reads</li> <li>percent_duplicate_reads: percentage of duplicate reads</li> <li>estimated_library_size: scaled total number of mapped reads</li> <li>total_reads: total number of reads, regardless of mapping status</li> <li>total_mapped_reads: total number of mapped reads</li> <li>total_duplicate_reads: number of duplicate reads</li> <li>total_properly_paired: number of properly paired reads</li> <li>coverage_breadth: percentage of genome covered by some read</li> <li>coverage_depth: average reads per nucleotide position in the genome</li> <li>median_insert_size: median insert size between paired reads</li> <li>mean_insert_size: mean insert size between paired reads</li> <li>standard_deviation_insert_size: standard deviation of the insert size between paired reads</li> <li>index_sequence: index sequence of the adaptor sequence</li> <li>column: column of the cell on the nanowell chip</li> <li>img_col: column of the cell from the perspective of the microscope</li> <li>index_i5: id of the i5 index adapter sequence</li> <li>sample_type: type of the sample</li> <li>primer_i7: id of the i5 index primer sequence</li> <li>experimental_condition: experimental treatment of the cell, includes controls</li> <li>index_i7: id of the i7 index adapter sequence</li> <li>cell_call: living/dead classification of the cell based on staining usually, C1 == living, C2 == dead</li> <li>sample_id: name of the sample</li> <li>primer_i5: id of the i5 index primer sequence</li> <li>row: row of the cell on the nanowell chip</li> <li>library_id: identifier for the DNA library</li> <li>index: ignored</li> <li>multiplier: during parameter searching, the set [1..6] that was chosen</li> <li>MSRSI_non_integerness: median of segment residuals from segment integer copy number states</li> <li>MBRSI_dispersion_non_integerness: median of bin residuals from segment integer copy number states</li> <li>MBRSM_dispersion: median of bin residuals from segment median copy number values</li> <li>autocorrelation_hmmcopy: hmmcopy copy autocorrelation</li> <li>cv_hmmcopy: ignored</li> <li>empty_bins_hmmcopy: number of empty bins in hmmcopy</li> <li>mad_hmmcopy: median absolute deviation of hmmcopy copy</li> <li>mean_hmmcopy_reads_per_bin: mean reads per hmmcopy bin</li> <li>median_hmmcopy_reads_per_bin: median reads per hmmcopy bin</li> <li>std_hmmcopy_reads_per_bin: standard deviation value of reads in hmmcopy bins</li> <li>total_halfiness: summed halfiness penality score of the cell</li> <li>total_mapped_reads_hmmcopy: total mapped reads in all hmmcopy bins</li> <li>scaled_halfiness: summed scaled halfiness penalty score of the cell</li> <li>mean_state_mads: mean value for all median absolute deviation scores for each state</li> <li>mean_state_vars: variance value for all median absolute deviation scores for each state</li> <li>mad_neutral_state: median absolute deviation score of the neutral 2 copy state</li> <li>breakpoints: number of breakpoints, as indicated by state changes not at the ends of chromosomes</li> <li>mean_copy: mean hmmcopy copy value</li> <li>state_mode: the most commonly occuring state</li> <li>log_likelihood: hmmcopy log likelihood for the cell</li> <li>true_multiplier: the exact decimal value used to scale the copy number for segmentation</li> <li>order: order of the cell in the hierarchical clustering tree</li> <li>quality: random forest classifier proability score that cell is good</li> </ul> <p>ov2295_clone_alleles.csv.gz: Table of clone specific allele data</p> <ul> <li>chr: chromosome of bin</li> <li>start: start of bin</li> <li>end: end of bin</li> <li>hap_label: haplotype block identifier</li> <li>clone_id: clone identifier</li> <li>allele_1_sum: number of reads for allele 1 of the haplotype block</li> <li>allele_2_sum: number of reads for allele 2 of the haplotype block</li> <li>total_counts_sum: total reads for the haplotype block</li> </ul> <p>ov2295_clone_breakpoints.csv.gz: Table of breakpoints per clone for OV2295 samples. Columns:</p> <ul> <li>prediction_id: identifier for the breakpoint</li> <li>chromosome_1: chromosome of breakend 1</li> <li>strand_1: orientation of break end 1</li> <li>position_1: position of break end 1</li> <li>chromosome_2: chromosome of breakend 2</li> <li>strand_2: orientation of break end 2</li> <li>position_2: position of break end 2</li> <li>clone_id: clone identifier</li> <li>read_count: number of reads</li> <li>is_present: presence=1, absent=0</li> </ul> <p>ov2295_clone_clusters.csv.gz: Table of cell clusters as putative clones</p> <ul> <li>cell_id: identifier for the cell</li> <li>clone_id: clone identifier</li> </ul> <p>ov2295_clone_cn.csv.gz: Table of allele specific copy number per clone for OV2295 samples. Columns:</p> <ul> <li>chr: chromosome of bin</li> <li>start: start of bin</li> <li>end: end of bin</li> <li>total_cn: HMMCopy predicted total copy number </li> <li>minor_cn: HMM predicted minor copy number </li> <li>major_cn: HMM predicted major copy number </li> <li>clone_id: clone identifier</li> </ul> <p>ov2295_clone_snvs.csv.gz: Table of SNVs per clone for OV2295 samples. Columns:</p> <ul> <li>chrom: chromosome</li> <li>coord: genome position</li> <li>ref: reference nucleotide</li> <li>alt: alternate nucleotide</li> <li>clone_id: clone identifier</li> <li>ref_counts: number of reads at this position matching the reference nucleotide</li> <li>alt_counts: number of reads at this position matching the alternate nucleotide</li> <li>total_counts: total number of reads at this position</li> <li>is_present: presence=0, absent=1</li> <li>is_het: is heterozygous</li> <li>is_hom: is homozygous for the alternate</li> </ul> <p>ov2295_nodes.csv.gz: Table of phylogenetic information for SNV evolution</p> <ul> <li>variant_id: identifier for the SNV as chrom:coord:ref:alt</li> <li>node: node in the phylogenetic tree</li> <li>loss: probability the SNV was lost at this node</li> <li>origin: probability the SNV originated at this node</li> <li>presence: probability the SNV is present at this node</li> <li>ml_origin: binary indicator the SNV originated at this node</li> <li>ml_presence: binary indicator the SNV is present at this node</li> <li>ml_loss: binary indicator the SNV was lost at this node</li> </ul> <p>ov2295_snv_counts.csv.gz: Table of SNV counts</p> <ul> <li>chrom: chromosome</li> <li>coord: genome position</li> <li>ref: reference nucleotide</li> <li>alt: alternate nucleotide</li> <li>ref_counts: number of reads at this position matching the reference nucleotide</li> <li>alt_counts: number of reads at this position matching the alternate nucleotide</li> <li>cell_id: identifier for the cell</li> <li>total_counts: total number of reads at this position</li> <li>sample_id: identifier for the sequenced sample</li> </ul> <p>ov2295_tree.pickle: Phylogenetic tree in python pickle format. Requires installation of the stochastic dollo code at: https://bitbucket.org/dranew/dollo, version 0.4.2.</p> <p>Note the following sample mapping: ‘SA922’: ‘OV2295(R2)’, ‘SA921’: ‘TOV2295(R)’, ‘SA1090’: ‘OV2295’,</p> <p><strong>Plots</strong></p> <p>ov_supp_clone_allele_cn.png: Clone allele ratios for each OV2295 sample.</p> <p>ov_supp_clone_total_cn.png: Clone copy number for each OV2295 sample.</p> <p>ov_supp_sample_total_cn.png: Bulk copy number for each OV2295 sample.</p> <p>ov_supp_sample_allele_cn.png: Bulk allele ratios for each OV2295 sample.</p>
DRM-labeled HIV-1 protease sequence dataset
<p>An HIV-1 protease dataset with labeled DRMs derived from the Stanford HIV Drug Resistance Database<sup>1,2</sup> is provided. It is in <em>fasta</em> format with major protease drug resistance mutations (as defined by Wensing et al.<sup>3</sup>) provided in the sequence name section (following the ">" symbol) as a comma-separated list. The dataset was used in the following paper: "Ahmed A., de Souza D. R., Link R. W., Nonnemacher M. R., Wigdahl B., Dampier W. Design of a SHERLOCK-based low resource screening assay for HIV-1 drug resistance, in preparation, 2021. "</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.