Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
2,848
datasets available to search
ShareScore release 0.9.0
Dataset results
2,848 results for “sequence data”
bulk-tumour-api: a programmatically accessible dataset of pre-processed bulk tumour sequencing data
<p><strong>This repository, including the API, are currently under development.</strong></p> <p><strong>bulk-tumour-api</strong>: A programmatically accessible dataset of pre-processed bulk tumour sequencing data. The python API can be found at https://github.com/tomouellette/bulk-tumour-api. All data stored in this repository have been collected from open access online sources. Original references and sources are provided in database.tsv (for empirical patient data) and synthetic.tsv (for simulated data).</p> <p><strong>A note on datasets: </strong></p> <ul> <li>All <em>empirical patient sequencing </em>samples have been processed into pseudo-VCF files which at minimum contain the following columns: sample identifier (sample), patient identifier (patient), chromosome (chr), position (pos), variant allele frequency (VAF), alternate read counts (t_alt_count), depth (DP), and total copy number (total_cn). However, if more data is required, unprocessed data including copy number segments or gene-level calls, clinical, and/or biopsy level information can be found in the /raw/. </li> <li>All <em>synthetic datasets </em>have also been processed in pseudo-VCF files. In some cases, all ground truth information (e.g. subclone frequency) is contained within the pseudo-VCF. In other cases, additional meta/ground-truth information are in separate files; any simulated sample with a column marked has_meta = True will have multiple files that will be downloaded together.</li> </ul>
Raw sequencing data PhD Mixoplankton spatio-temporal diversity and its environmental drivers in the North Sea
<p>Raw sequencing data PhD Mixoplankton spatio-temporal diversity and its environmental drivers in the North Sea</p>
MALDI-TOF-MS reference spectra and sequence data for domesticated equids (horse and donkey) collagen for Zooarchaeology by Mass Spectrometry (ZooMS)
<p>MALDI-TOF-MS spectra of extracted collagen from modern reference and archaeological bone samples to develop markers for Zooarchaeology by Mass Spectrometry (ZooMS) to distinguish between Equus species. For each sample digestions were done in both trypsin and chymotrypsin separately. Information about the species of the samples can be found in 'sample metadata.csv' file. Information on the extraction and digestion protocol can be found in the associated manuscript. The sequence data contains alignments of the proteins COL1A1 and COL1A2 for available Equus collagen protein sequences. More information on these files can be found in the corresponding manuscript to this dataset.<br> </p>
Control panel created from 30 Nanopore sequencing data from the Human Pangenome Reference Consortium
<p>This is control panel for <a href="https://github.com/friend1ws/nanomonsv">nanomonsv</a> software, which is expected to exclude many false positives as well as improve computational cost. This is made by aligning 30 Nanopore sequencing data from Human Pangenome Reference Consortium to the GRCh38 reference genome (obtained from <a href="https://console.cloud.google.com/storage/browser/genomics-public-data/resources/broad/hg38/v0;tab=objects">here</a>) with <a href="https://github.com/lh3/minimap2">minimap2</a> version 2.24. <strong>When you use these control panels and publish, do not forget to credit to <a href="https://humanpangenome.org/data-use-protocol/">HPRC</a>!</strong></p>
Data supporting "Transformer Model Generated Bacteriophage Genomes are Compositionally Distinct from Natural Sequences"
<p>Sequence and composition data supporting doi: <a href="https://doi.org/10.1101/2024.03.19.585716" target="_blank" rel="noopener">10.1101/2024.03.19.585716</a>. Uncompressed file size is ~5.8GB.</p> <p>Data in zip files is organized by sequence provenance (generRNA, natural, or transformer (megaDNA)). Common file types between folders include:</p> <ul> <li>Multi-record fasta file: Sequence data for all sequences of a given provenance. For generRNA sequences, these are found within the `seq` column of file "MFE_distribution_Fig4a.csv"</li> <li>Composition files: Individual sequence level compositional metrics for sliding 120 bp windows. Only structural metrics were used in this study.</li> <li>Genomad: Results from the genomad pipeline (https://portal.nersc.gov/genomad/)</li> <li>Stats: Aggregate statistics for all sequences of a given provenance.</li> </ul> <p>The natural folder also has a metadata file detailing the taxonomy for all natural sequences.<br><br>Figure datasets are the cleaned (sometimes aggregated) datasets that underly specific figures in the manuscript. The figure designations are based on the order in: https://www.biorxiv.org/content/10.1101/2024.03.19.585716v1.</p>
Supplementary data - Simultaneous polyclonal antibody sequencing and epitope mapping by cryo electron microscopy and mass spectrometry – a perspective
<p>Analysis files and scripts for <a href="https://doi.org/10.1101/2024.06.21.600107" target="_blank" rel="noopener">associated manuscript</a>. </p> <ul> <li>CR3022.zip: script (in Rust) and necessary data to run said script for CR3022 analysis with the results from running the script.</li> <li>MA-analysis-script.zip: script (in Rust) and necessary data to run said script for automated analysis of MA benchmark results.</li> <li>MA-analysis-data.zip: data from running the MA-analysis-script, containing all MA and Stitch output files.</li> <li>MA-analysis-data-EMPEM.zip: data from running MA and Stitch on the EMPEM benchmark.</li> </ul>
Data underpinning "Pulse sequence considerations for interleaved chemical exchange saturation transfer acquisition sequences."
<p>=================================================<br> Robert Casper Brand, PhD Candidate<br> Wellcome Centre for Integrative Neuroimaging, FMRIB Division, Nuffield Department of Clinical Neurosciences, University of Oxford, Oxford, UK.<br> =================================================</p> <p>This folder contains the images and datasets used to generate the figures of the paper named: "Pulse sequence considerations for interleaved chemical exchange saturation transfer acquisition sequences." </p> <p>Each figure of the paper, with its corresponding data, is contained in an opensource TikZ format file. The TikZ files include both information on the axis as well as the supporting data and can be opened with any generic text editor. For more information on TikZ, see:<br> https://www.sharelatex.com/learn/TikZ_package). </p> <p>Where datasets were too large to be run by standard TeX distributions, the data was attached in an alternative format, and a TikZ wrapper included.</p> <p>The figures can be generated through any of the opensource TeX distributions. For more information on LaTeX and TeX, please see: <br> https://www.latex-project.org/get/ and<br> https://www.sharelatex.com/learn/Pgfplots_package.</p> <p>A compilation example of all figures, which also lists any additional packages, is included in the "wrapper. Tex" file. The output of this process was added to this folder as well (wrapper.pdf).</p> <p>The included files were created using directly from Matlab using the matlab2tikz code:<br> https://www.mathworks.com/matlabcentral/fileexchange/22022-matlab2tikz-matlab2tikz</p>
Sample data for analysis of sequence variation in HIV
<p>These are downsampled interleaved paired fastq datasets from Jair et. 2019 (<a href="https://doi.org/10.1371/journal.pone.0214820">https://doi.org/10.1371/journal.pone.0214820</a>). The datasets were prepared by:</p> <ol> <li>Downloading original data from NCBI SRA (https://www.ncbi.nlm.nih.gov/bioproject/PRJNA517147)</li> <li>Trimming contaminating Nextera adapters using trim-galore</li> <li>Mapping reads against nxb2 reference of HIV genome (K03455.1) with BWA MEM</li> <li>Restricting mapped reads to <em>pol</em> gene vicinity (K03455.1:2000-5100)</li> <li>Downsampling mapped data to ~10% of the original with Picard's DownsampleSam</li> <li>Converting BAM to Interleaved Fastq with Picard's SamToFastq</li> <li>Gzipping resultant interleaved paired fastq files</li> </ol>
LevSeq epPCR data from ParPgb LQ from MinION sequencer
<p>This is the data published with the LevSeq preprint (https://doi.org/10.1101/2024.09.04.611255), these data provide a test case for users to confirm that they are able to run the pipelines and also the data to reproduce the results in the preprint. Each folder contains the raw fastq files from the MinION sequencer for a protoglobin variant for several plates, along with the required input file to run LevSeq.</p> <p>The data are basecalled reads, generated using the standard Oxford nanopore sequencing protocol. For experimental methods and details please see the paper.</p> <ol> <li><a href="../api/records/13694463/draft/files/20240421-YL-ParLQ-ep1.csv/content" target="_blank" rel="noopener noreferrer">20240421-YL-ParLQ-ep1.csv</a> is the reference file for one set of plates</li> <li><a href="../api/records/13694463/draft/files/20240502-YL-ParLQ-ep2.csv/content" target="_blank" rel="noopener noreferrer">20240502-YL-ParLQ-ep2.csv</a> is the reference file for the second set of plates</li> <li>20240421.zip is the raw data for 20240421-YL-ParLQ-ep1.csv </li> <li>20240502.zip 20240502-YL-ParLQ-ep2.csv</li> <li><a href="13694463" target="_blank" rel="noopener noreferrer">20240502-YL-ParLQ-ep2.fastq.zip</a> is the combined fastq files for the run from <a href="../api/records/13694463/draft/files/20240502-YL-ParLQ-ep2.csv/content" target="_blank" rel="noopener noreferrer">20240502-YL-ParLQ-ep2</a> for ease of use</li> <li><a href="../api/records/13694463/draft/files/20240422-YL-ParLQ-ep1.fastq.zip/content" target="_blank" rel="noopener noreferrer">20240422-YL-ParLQ-ep1.fastq.zip</a> is the combined fastq files for the run from <a href="../api/records/13694463/draft/files/20240421-YL-ParLQ-ep1.csv/content" target="_blank" rel="noopener noreferrer">20240421-YL-ParLQ-ep1</a> for ease of use</li> </ol>
Supplemental_Data_S1 for "Kmer Manifold Approximation and Projection for visualizing DNA sequences"
<p>This dataset includes the results generated by KMAP software applied to the htselexdata dataset. Each folder within the dataset contains outputs from multiple dimensionality reduction techniques, including KMAP, UMAP, t-SNE, and MDS. Additionally, motifs and logos have been derived using both KMAP and MEME methods. This data provides insights into motif patterns and structures, which can be beneficial for further bioinformatics and computational biology analyses.</p>
Verification of library complexity in the HEK-Cas9 sublibraries - sequence data of the generated sublibraries A and B
<p>Sequence data of the generated HEK-Cas9 sublibraries A and B, linked to the manuscript 10.1128/mbio.01925-24: The <em>Bordetella</em> effector protein BteA induces host cell death by disruption of calcium homeostasis by Martin Zmuda, Eliska Sedlackova, Barbora Pravdova, Monika Cizkova, Marketa Dalecka, Ondrej Cerny, Tania Romero Allsop, Tomas Grousl, Ivana Malcova, and Jana Kamanova</p>
Generation of transcriptional novelty by transposable element insertions in Arabidopsis, Genome Sequencing and eccDNA Data
<p><strong>Raw Illumina sequencing data from the Manuscript entitled "Generation of transcriptional novelty by transposable element insertions in Arabidopsis"</strong></p> <p><strong>A. Illumina genome sequencing reads of Arabidopsis control and hcLines that contain novel transposable element insertions.</strong></p> <p>To identify the genomic position of the new <em>ONSEN</em> insertions, the extracted DNA of the 11 selected lines (nine lines with new insertions and two control lines) was sent to BGI, Hong-Kong for Illumina paired-end 150 bp sequencing, aiming for a minimum of 20X sequencing coverage. Quality control of the raw reads was done using FastQC (Andrews S. (2010). FastQC: a quality control tool for high throughput sequence data. Available online at: <a href="http://www.bioinformatics.babraham.ac.uk/projects/fastqc">http://www.bioinformatics.babraham.ac.uk/projects/fastqc</a>) and trimming/clipping was done using Trimmomatic with parameters ILLUMINACLIP: TruSeq3:2:30:10 LEADING:20 TRAILING:20 SLIDINGWINDOW:4:20 and MINLEN:36. Quality of the reads was deemed excellent and no further actions were taken.</p> <p>Samples identifications: genome_hcLineX with "_1" indicating the forward and "_2" the reverse reads.</p> <p><strong>B. Illumina eccDNA sequencing of Arabidopsis control and hcLines following stress treatments</strong></p> <p>Extrachromosomal circular DNA was prepared and sequenced as follows: twenty plants from each petri dish were pooled separately and DNA was extracted using the CTAB method (<a href="https://dx.doi.org/10.17504/protocols.io.quidwue">dx.doi.org/10.17504/protocols.io.quidwue</a>). Following the mobilome-seq method described in (Lanciano et al., 2017), for all samples, we digested linear DNA from 2 µg of total DNA for 17 hours at 37<sup>o</sup>C using 10 U of PlasmidSafe (<em>LubioScience cat# E3101K</em>), followed by enzyme denaturation (30 mins at 70<sup>o</sup>C). Digested DNA was precipitated with isopropanol supplemented with 1 µg of GlycoBlue coprecipitant (<em>Fisher Scientific cat# 10391565</em>). Circular DNA was then amplified through rolling circle amplification (RCA) with the Illustra TempliPhi kit (<em>GE Healthcare cat# 25-6400-10</em>), following the manufacturer recommendation and leaving the reaction for 16h at 30<sup>o</sup>C. DNA was once again precipitated with isopropanol and sent for Illumina paired end 150 bp sequencing at BGI, Hong Kong. </p> <p>Samples identification: </p> <p>eccDNA_A.thaliana_ctrl: control reads</p> <p>eccDNA_A.thaliana_HS: heat stressed plants reads</p> <p>eccDNA_A.thaliana_AZ_HS: reads of alpha-amanitin, zebularine and heat-stressed plants</p> <p>"R1" indicates forward and "R2" reverse reads.</p> <p> </p>
xPore: Identification of differential RNA modifications from nanopore direct RNA sequencing - SGNEx data
<p>xPore is a Python package for identification and quantification of differential RNA modifications from direct RNA sequencing.</p> <p>The detailed usage is documented at <a href="https://xpore.readthedocs.io/en/latest/">https://xpore.readthedocs.io/en/latest</a>, while all scripts and source code are available at <a href="https://github.com/GoekeLab/xpore">https://github.com/GoekeLab/xpore</a>.</p> <p>All the preprocessed datasets used in the paper are provided here. </p> <p>Please cite our paper below when using these data.<br> Ploy N. Pratanwanich et al. "Detection of differential RNA modifications from direct RNA sequencing of human cell lines." bioRxiv (2020).</p>
Field data for: Enterovirus sequence data obtained from primate samples in Central Africa suggest a high prevalence of enteroviruses with possible zoonotic potential
<p>Enteroviruses infect humans and animals, can cause disease, and some may be transmitted across species barriers. We collected different types of samples from various species of Central African wildlife, including data on sampling location and tested the samples for the presence of Enterovirus RNA using a family level PCR. Specimen collection was approved by an Institutional Animal Care and Use Committee (IACUC) of the University of California Davis, and the Governments of Cameroon and the Democratic Republic of the Congo. Enterovirus RNA was detected in samples from 17 primates and 2 rodents. Some sequences were very similar while others were dissimilar to known species, highlighting the unexplored enterovirus diversity in wildlife.</p> <p>The samples and filed data were collected by field ecologists as part of the USAID funded PREDICT project (https://ohi.vetmed.ucdavis.edu/programs-projects/predict-project) and screened for enterovirus RNA using consensus PCR. Maps were generated using basic maps from Paintmaps (http://www.paintmaps.com), a free tool for educational and academic use. The dataset contains the metadata on enterovirus screening among wildlife in Cameroon and the Democratic Republic of the Congo from 2003-2014 as part of the USAID funded PREDICT project. Please refer to the article for more information on methods and references.</p>
Data from the article "Plastome sequencing of South American Podocarpus species reveals low rearrangement rates despite ancient Gondwanan disjunctions"
<p>Input data, intermediate and final analysis output files associated to the manuscript "Plastome sequencing of South American <em>Podocarpus </em>species reveals low rearrangement rates despite ancient Gondwanan disjunctions"</p> <p>We sequenced the plastomes of four South American species of <em>Podocarpus</em> from Patagonia, southern Yungas, and Brazilian subtropical forests: <em>P. nubigenus, P. parlatorei, P. salignus </em>and <em>P. selowii</em>. We compared their plastomes to those published from Brazil, Africa, New Zealand, and Southeast Asia, along with representatives from other genera within Podocarpaceae as outgroups. The four newly sequenced plastomes ranged in size between 133,791 bp and 133,991 bp. Gene content and order among chloroplasts from South American, African and Asian <em>Podocarpus</em> were conserved and different from the plastome of <em>P. totara</em>, from New Zealand. Most genes showed substitution patterns consistent with a conservative selective regime. Phylogenies inferred from either complete sequences or protein coding regions were mostly congruent with previous studies, but showed earlier branching of <em>P. salignus</em>, <em>P. totara</em> and <em>P. sellowii</em>.</p>
16S rRNA sequencing gene datasets for CRC data
<p>Used datasets: </p> <table> <thead> <tr> <th scope="col"> <table> <thead> <tr> <th>Dataset</th> <th>16S rRNA Region</th> <th>Control (n)</th> <th>Adenoma (n)</th> <th>CRC (n)</th> <th>Available metadata</th> </tr> </thead> <tbody> <tr> <td><a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4823848/">Baxter</a></td> <td>V4</td> <td>171</td> <td>198</td> <td>120</td> <td>Gender, age, weight, height, BMI, country, race</td> </tr> <tr> <td><a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4221363/">Zackular</a></td> <td>V4</td> <td>30</td> <td>30</td> <td>30</td> <td>Gender, age, weight, height, BMI, country, race, FOBT, medication</td> </tr> <tr> <td><a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4299606/">Zeller</a></td> <td>V4</td> <td>50</td> <td>38</td> <td>41</td> <td>Gender, age, BMI, country, FOBT</td> </tr> <tr> <td><strong>TOTAL</strong></td> <td>V4</td> <td>251</td> <td>266</td> <td>191</td> <td><em>All of the above</em></td> </tr> </tbody> </table> </th> </tr> </thead> <tbody> <tr> <td> </td> </tr> </tbody> </table> <p>Data processing & sharing</p> <p>All datasets were processed using <a href="https://docs.qiime2.org/2021.11/">qiime2</a> pipeline with <a href="https://benjjneb.github.io/dada2/">DADA2</a> for Sequence quality control and feature table construction and <a href="https://www.arb-silva.de/">SILVA</a> database for taxonomic assignment, and then a <em>phyloseq </em>object was constructed.</p> <ul> <li>Abundance table at genus level is in file <em>genus.csv</em> (Sample counts with NO filtering).</li> <li>Clean metadata is in <em>metadata.csv</em> file (Countries: CA - Canada. USA - United States of America. FRA - France.)</li> <li>Phyloseq object is in file <em>physeq.RDS</em> (Saved as an RDS object in R)</li> </ul> <p>More information is <a href="https://hackmd.io/nbsLqCLlSNSRFc5RBX9c5Q?view">here</a>. </p> <p> </p>
Sequencing data and code - Prince et al 2023
<p>Raw sequencing data, code and intermediate analysis files from "Antiviral activity of molnupiravir precursor NHC against SARS-CoV-2 Variants of Concern (VOCs) and implications for the therapeutic window and resistance" (Prince et al, 2023).</p> <p>Please see sample_metadata.xlsx for all metadata relating to the files contained in this repository.</p> <p>Code for data visualisation can be found in: mut_sub_all_pts_serial-pass.R. Specific paths to data will have to be changes to refer to where you have downloaded the data in this repository. Metadata for use with the R script is serial-pass-nimagen-metadata-forR.csv.</p>
Data For: Identifying rare variants inconsistent with identity-by-descent in population-scale whole-genome sequencing data
<p>Simulation output and Genome-wide scan for nIBD variants in UK10K data as reported in:</p> <p>Identifying rare variants inconsistent with identity-by-descent in population-scale whole-genome sequencing data</p> <p>Johnson KE, Adams CJ, Voight BF. Methods Ecol Evol 2022 Nov;13(11): 2429–2442.</p> <p>Code available at: https://github.com/kelsj/EVICORD</p>
Data for: Regularized sequence-context mutational trees capture variation in mutation rates across the human genome
<p>Additional data on output models from Bayer as reported in:</p> <p>Regularized sequence-context mutational trees capture variation in mutation rates across the human genome</p> <p>Adams CJ, Conery M, Auerbach BJ, Jensen ST, Mathieson I, Voight BF. BioRxiv https://doi.org/10.1101/2022.10.14.512160</p> <p>Accepted, PLoS Genetics. </p> <p>Code Available at: https://github.com/bvoightlab/Baymer</p>
Timed Sequence Task - Data
<p>Timed Sequence Task - data</p> <p>Raw and generated data from Timed Sequence Task. Data are generated by set of codes for standard operant boxes equipped with two fixed levers and a feeder (Med Associates, MED-307A-B1) and operated by Med-PC V software. </p> <p>Names of the tasks indicate the required sequences of lever presses (eg left - left - right - right). Individual inter-press intervals are specifically defined and are logged in great detail (ie. different types of incorrect presses are distinguished and recorded separately).</p> <p> </p> <p>Operant boxes code is available here: <a href="https://github.com/mjtecka/Timed-Sequence-Task-OB">Timed Sequence Task - OB</a>.</p> <p>Python scripts for the automated analysis of the resulting task logs are available on GitHub: <a href="https://github.com/mjtecka/Timed-Sequence-Task-Analysis">Timed Sequence Task Analysis</a></p> <p> </p> <p>Dataset contains following files</p> <p>raw_data/data_vpa_final.csv</p> <ul> <li>original, complete log of all tasks as exported from Med-PC V Software</li> <li>columns are described in separate file <a href="https://zenodo.org/api/files/3084f324-a303-47a4-8e58-e3506c990d37/Data%20VPA%20final%20-%20columns%20structure.txt">Data VPA final - columns structure.txt</a></li> </ul> <p>raw_data/mastersheet_complete_one_file.csv :</p> <ul> <li>further context for the raw log allowing to split the original log in multiple task logs; format [ mouse - day - task ] ; this can be done using scripts provided here: https://doi.org/10.5281/zenodo.7881058 </li> </ul> <p> </p> <p>generated data/data_by_task/[task-log].csv </p> <ul> <li>data from complete log filtered according to mastersheet into specific [task-log].csv files</li> </ul> <p>generated_data/subanalysis_results/[subanalysis-name]/[task-name].csv</p> <ul> <li>results of specific subanalyses divided by task names into separate files. <ul> <li>e.g. count_all_presses_test/LLPP-count_all_presses_test.csv = subanalysis count all presses for LLPP task</li> </ul> </li> <li>note that csv files in this directory have different structure, according to selected subanalysis</li> </ul> <p> </p> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.