Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

2,848

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

2,848 results for “sequence data”

Learn how ShareScore rates datasets ↗
zenodo36/100

Raw sequencing data of NGS of stomach contents of sympatric species of weakly electric fish (genus: Campylomormyrus)

<p>This dataset contains the raw sequencing data of stomach contents of sympatric species of weakly electric fish (genus: <em>Campylomormyrus</em>) using next generation sequencing.</p> <p>Stomach content samples&nbsp;were collected from five <em>Campylomormyrus</em> species (<em>C. alces</em>, ; <em>C. compressirostris</em>, ; <em>C. curvirostris</em>, ; <em>C. numenius</em>, ; <em>C. tshokwe</em>) and samples of <em>Gnathonemus petersii</em> (<em>G. petersii, </em>), a sister genus of <em>Campylomormyrus.</em></p> <p>The fish specimens, from which these stomach content samples are extracted, were collected during an expedition to the Republic of the Congo in fall 2012.</p> <p>The dataset files are in FASTA format.</p>

opencc-by-4.0May 2022View details →
zenodo36/100

Data for "Predicting aggregate morphology of sequence-defined macromolecules with Recurrent Neural Networks"

<p>These are the data associated with the paper, &quot;Predicting aggregate morphology of sequence-defined macromolecules with Recurrent Neural Networks&quot; (DOI 10.1039/D2SM00452F). Three of the directories contains subdirectories with `GSD` files dumped from HOOMD. The other contains pretrained RNN models as TorchScript binaries exported from PyTorch.</p>

opencc-by-4.0May 2022View details →
dryad36/100

Data from: A genotyping-in-thousands by sequencing panel to inform invasive deer management using non-invasive fecal and hair samples

<p>Studies in ecology, evolution, and conservation often rely on non-invasive samples, making it challenging to generate large amounts of high-quality genetic data for many elusive and at-risk species. We developed and optimized a Genotyping-in-Thousands by sequencing (GT-seq) panel using non-invasive samples to inform the management of invasive Sitka black-tailed deer (<em>Odocoileus hemionus sitkensis</em>) in Haida Gwaii (Canada). We validated our panel using paired high-quality tissue and non-invasive fecal and hair samples to simultaneously distinguish individuals, identify sex and reconstruct kinship among deer sampled across the archipelago, then provided a proof-of-concept application using field-collected feces on SGang Gwaay, an island of high ecological and cultural value. Genotyping success across 244 loci was high (90.3%) and comparable to that of high-quality tissue samples genotyped using restriction-site associated DNA sequencing (92.4%), while genotyping discordance between paired high-quality tissue and non-invasive samples was low (0.50%). The panel will be used to inform future invasive species operations (culls or eradications) in Haida Gwaii by providing individual and population information to inform management. More broadly, our GT-seq workflow that includes quality control analyses for targeted SNP selection and a modified protocol may be of wider utility for other studies and systems where non-invasive genetic sampling is employed.</p>

opencc-zeroDec 2021View details →
dryad36/100

Data for a preliminary molecular phylogeny of the family Hydroptilidae (Trichoptera): exploring the combination of targeted enrichment data and legacy Sanger sequence data

<p><span>The purpose of this study is to provide a proof-of-concept that the use of molecular data, particularly targeted enrichment data, and statistically supported methods of analysis can result in the construction of a stable phylogenetic framework for the microcaddisflies (Trichoptera: Hydroptilidae). Here, we use a combination of targeted enrichment data for ca. 300 nuclear protein-coding genes and legacy (Sanger-based) sequence data for the mitochondrial COI gene and partial sequence from the 28S rRNA gene.</span></p>

opencc-zeroJun 2022View details →
dryad36/100

DNA metabarcoding sequence data for diet analysis of caribou

<p>Woodland caribou (<em>Rangifer tarandus caribou</em>) are threatened in Canada due to the drastic decline in population size caused primarily by human-induced landscape changes that decrease habitat and increase predation risk. Conservation efforts have largely focused on reducing predators and protecting critical habitat, whereas research on dietary niches and the role of potential food constraints in lichen-poor environments is limited. To improve our understanding of dietary niche variability, we used a next-generation sequencing approach with metabarcoding of DNA extracted from faecal pellets of woodland caribou located on Lake Superior in lichen-rich (mainland) and lichen-poor (island) environments. Amplicon sequencing of fungal ITS2 region revealed lichen-associated fungi as predominant in samples from both populations, but amplification at the chloroplast <em>trnL </em>region, which was only successful on island samples, revealed primary consumption of yew based on relative read abundance (<em>Taxus spp.</em>; 83.68%) with dogwood (<em>Cornus spp</em>.; 9.67%) and maple (<em>Acer spp.</em>; 4.10%) also prevalent. These results suggest that conservation efforts for caribou need to consider the availability of food resources beyond lichen to ensure successful outcomes.  More broadly, we provide a reliable methodology for assessing ungulate diet from archived faecal pellets that could reveal important dietary shifts over time in response to climate change.</p>

opencc-zeroJun 2022View details →
zenodo36/100

Data in support of Using target sequence capture to improve the phylogenetic resolution of a rapid radiation in New Zealand Veronica

<p>Includes alignments and trees for the analysis found in Thomas et al. 2021, Using target sequence capture to improve the phylogenetic resolution of a rapid radiation in New Zealand Veronica;&nbsp;American Journal of Botany, Special Issue: Exploring Angiosperms353: a Universal Toolkit for Flowering Plant Phylogenomics. Alignments comprise subsets of Angiosperms353 genes given each filtering scheme (full, intersection, sortadate_BP, sortadate_TL) and gene type/subset (exons, introns, supercontigs), and for markers downloaded from GenBank, as explained in the Methods section of Thomas et al. 2021. Trees were included for each of these alignments from IQtree and Astral; SVDquartets tree was only estimated for the full set of supercontigs. Gene trees were generated with IQtree. Tree files are named differently than the final manuscript; refer to the number of genes specified in Fig 1 of Thomas et al, 2021 and specified in each filename to identify filtering scheme.&nbsp;Raw sequence reads are available on the Sequence Read Archive at <a href="http://www.ncbi.nlm.nih.gov/bioproject/715342">http://www.ncbi.nlm.nih.gov/bioproject/715342</a>.</p>

opencc-by-4.0Jun 2022View details →
zenodo36/100

Data and processing scripts for PRISM barcode sequencing data used in "Massively parallel pooled screening reveals genomic determinants of nanoparticle-cell interactions"

<p>Sequencing data for the PRISM barcodes generated after nano-particle treatment is presented in this repository alongside the code to process the sequencing counts to generate the binning probabilities and weighted scores.&nbsp;<br> <br> For the details please see the original publication or the bioarxiv preprint:&nbsp;&nbsp;https://doi.org/10.1101/2021.04.05.438521<br> <br> The raw data is provided in PILOT_DATA_COUNTS.csv and EXPERIMENT_DATA_COUNTS.csv files, for the pilot and the actual experiment.&nbsp;<br> <br> For each of these files an R script is provided to process them, along with the output of the scripts (PILOT_DATA_PROBABILITIES.csv and EXPERIMENT_DATA_PROBABILITIES.csv)</p>

opencc-by-4.0Jun 2022View details →
zenodo36/100

Discovering molecular regulators of ageing using mixture models with RNA-sequencing data

<p>Identifying the molecular regulators that control ageing is challenging because the ageing process is influenced by a combination of genetic and environmental factors which makes it difficult to source the contribution of a single gene. Multiple studies have demonstrated that as humans age, increased gene expression heterogeneity results in the dysregulation of key regulators and pathways. Given the dynamic nature of gene expression, it is vital that this data be modelled by statistical approaches that can appropriately account for changes in variability to understand the contribution of heterogeneity during the aging process and properly identify its regulators. This study demonstrates the utility of using mixture models to model biological variability of gene expression occurring during ageing and how novel potential regulators of ageing can be identified.</p> <p>Our mixture modelling approach was applied to gene expression data from the Genotype-Tissue Expression (GTEx) cohort. For every gene, the expression profile was modelled using a mixture model across the cohort where the subset of donors corresponding to each mode was tested for a significant change in age group. The multi-tissue aspect of GTEx was leveraged to find ageing regulators based on this mixture model approach genes that were common across multiple tissues, suggesting that the regulation of ageing may also be controlled through a set of genes that have non-tissue-specific activity.</p> <p>Our approach identified well-documented ageing regulators <em>mTOR </em>and <em>RICTOR</em> and other potential ageing regulators such as <em>IL4</em> and <em>GPR4</em> which were detected only by our approach. Genes identified by edgeR, DESeq2 and the mixture model-based approach were enriched for similar biological pathways. This suggests that while the specific ageing regulators identified from our approach may be distinct, they generally belong in the same pathways as the genes identified by standard approaches. Overall, these results indicate that modelling gene expression variability using mixture models in conjunction with standard differential gene expression can help uncover new regulators that have a potential role for understanding human ageing.</p> <p>I</p>

opencc-by-4.0Apr 2022View details →
zenodo36/100

Supplementary Data for "The shaky foundations of simulating single-cell RNA sequencing data"

<p>Supplementary Data for &quot;The shaky foundations of simulating single-cell RNA sequencing data&quot;</p> <p>See description.txt, Supplementary Text, Methods and&nbsp;https://github.com/HelenaLC/simulation-comparison for further description of the files available here.</p>

opencc-by-4.0Aug 2022View details →
zenodo36/100

Data and scripts for the manuscript of svaRetro and svaNUMT: modular packages for annotating retrotransposed transcripts and nuclear integration of mitochondrial DNA in genome sequencing data

<p>This upload include data and scripts supporting&nbsp;the results described in the manuscript of&nbsp;<em>svaRetro and svaNUMT: modular packages for annotating retrotransposed transcripts and nuclear integration of mitochondrial DNA in genome sequencing data</em><em>.&nbsp;</em>Detailed description of the contents can be found in README.txt.</p>

opencc-by-4.0Feb 2022View details →
dryad36/100

Ultra-short response-guided Hepatitis C treatment with sofosbuvir and daclatasvir: the SEARCH study HCV sequence data

<p><strong>Background</strong></p> <p>WHO has called for research into predictive factors for selecting persons who could be successfully treated with shorter durations of antiviral therapy for Hepatitis C. We evaluated early virological response as a means of shortening treatment and explored host, viral and pharmacokinetic contributors to treatment outcome.</p> <p class="MsoNormal"><strong><span>Methods</span></strong></p> <p class="MsoNormal"><span>Duration of sofosbuvir and daclatasvir (SOF/DCV)</span> was determined according to <span>day 2 (D2) virologic response</span> for <span>HCV genotype (gt) 1- or 6-infected adults in Vietnam with mild liver disease. Participants received 4 or 8 weeks of treatment according to whether D2 HCV RNA was above or below 500 IU/ml (standard duration is 12 weeks). Primary endpoint was sustained virological response (SVR12). Those failing therapy were retreated with 12 weeks SOF/DCV. Host IFNL4 genotype and viral sequencing was performed at baseline, with repeat viral sequencing if virological rebound was observed. Levels of SOF, its inactive metabolite</span> <span>GS-331007 and DCV were measured on day 0 and 28. </span></p> <p class="MsoNormal"><strong>Findings</strong></p> <p class="MsoNormal"><span>Of 52 adults enrolled, 34 received 4 weeks SOF/DCV, 17 got 8 weeks and one withdrew. SVR12 was achieved in 21/34 (62%) treated for 4 weeks, and 17/17 (100%) treated for 8 weeks. Overall, 38/51 (75%) were cured with first-line treatment </span><span>(mean duration of 37 days). Despite a high prevalence of putative </span>NS5A-inhibitor <span>resistance-associated substitutions (RAS), all first-line treatment failures were cured after retreatment (13/13). We found no evidence treatment failure was associated with host </span><span>IFNL4 genotype, viral</span><span> subtype, baseline RAS or DCV levels. SOF metabolite levels were higher in those failing 4-week therapy.</span></p> <p class="MsoNormal"><strong>Interpretation</strong></p> <p class="MsoNormal">Shortened SOF/DCV therapy with retreatment if needed, reduces DAA use while maintaining high cure rates. D2 virologic response alone does not adequately predict SVR12 with 4 weeks of treatment.</p>

opencc-zeroAug 2022View details →
dryad36/100

Simulated data from: Reference-free assembly of long-read transcriptome sequencing data with RNA-Bloom2

<p><span>Long-read sequencing technologies have improved significantly since their emergence. Their read lengths, potentially spanning entire transcripts, is advantageous for reconstructing transcriptomes. Existing long-read transcriptome assembly methods are primarily reference-based and to date, there is little focus on reference-free transcriptome assembly. We introduce RNA-Bloom2, a reference-free assembly method for long-read transcriptome sequencing data. </span>RNA-Bloom2 is available on GitHub at: <a href="https://github.com/bcgsc/RNA-Bloom">https://github.com/bcgsc/RNA-Bloom</a>.</p> <p><span>We benchmarked the assembly quality and the computational performance of RNA-Bloom2 on simulated data. We prepared two mouse simulated datasets with Trans-NanoSim</span><span> for the cDNA and dRNA sequencing protocols model on experimental ONT data</span><span>. The datasets were simulated </span><span>based on the mouse ENSEMBL annotation for GRCm39.</span><span> To investigate the effect of sequencing depth, we subsampled each dataset to 2, 10, and 18 million reads, resulting in a total of six sets of reads for our benchmarking experiments. Using the simulated data, w</span><span>e showed that the transcriptome assembly quality of RNA-Bloom2 is competitive to those of reference-based methods.</span></p>

opencc-zeroSep 2022View details →
zenodo36/100

Mielke & Carvalho 2022 Chimpanzee play sequences are structured hierarchically as games - Data

<p>Data and scripts for the 2022 manuscript &#39;Chimpanzee play sequences are structured hierarchically as games&#39; - preprint here:&nbsp;</p> <p>https://doi.org/10.1101/2022.06.14.496075</p> <p>Dataset and scripts generated on 20/09/2022. For potential changes and all information see:</p> <p>https://github.com/AlexMielke1988/Mielke-Carvalho_Chimpanzee-Play</p>

opencc-by-4.0Sep 2022View details →
zenodo36/100

Code and Data for "Real-time dynamic single-molecule protein sequencing on an integrated semiconductor device"

<p><strong>Code and Data for &quot;Real-time dynamic single-molecule protein sequencing on an integrated semiconductor device&quot;.</strong></p> <pre>Code to analyze data produced by the Quantum-Si benchtop device and semiconductor chip is provided in a Python library <strong>qsi_algo</strong> under several submodules: - <strong>rs_caller.py</strong>: Algorithm for calling RS segments (also called ROI segments throughout code). - <strong>rs_caller_controller.py</strong>: Code framework for executing RS calling and property computation in a distributed manner - <strong>rs_properties</strong>: Code for computing properties of identified RS - <strong>rs_classifier</strong>: Algorithms for identifying peptide states (i.e. residue calls) associated with an RS - <strong>utils.py</strong>: shared helper code - <strong>pulse_reader</strong>: reader for binary pulse file - <strong>filters</strong>: ROI and pulse filtering utilities - <strong>plotting</strong>: functions for visualization of data relevant to the analyses presented Jupyter notebooks (<strong>.ipynb</strong>) files are named according to the manuscript figure they are associated with. Analysis code inside uses provided RS (recognition segment) data to demonstrate filtering and residue-calling techniques required to replicate analyses shown in manuscript figures. Please note: several methods rely on randomization for model initialization and/or data sampling which can cause small deviations from equivalent analyses in published figures. The raw data produced from the Quantum-Si benchtop device and semiconductor chip for the assays presented in the accompanying study is presented in a pulse-called binary file format. Pulses can be used as input for RS identification and peptide state identification. Pre-segmented (RS-identified) files are included for convenience. The data contained in the files include: <strong>{run_id}.bin</strong>: Binary format for storing pulse info. The reader provided in <strong>qsi_algo.pulse_reader</strong> produces the following columns: - <strong>aperture_index</strong>: unique aperture index on chip - <strong>start_f</strong>: index of first frame in pulse, counted from the beginning of the run - <strong>end_f</strong>: index of last frame in pulse, counted from the beginning of the run - <strong>dur_f</strong>: duration of pulse in frames - <strong>dur_s</strong>: duration of pulse in seconds - <strong>ipd_f</strong>: interpulse duration in frames (number of frames since end of preceding pulse) - <strong>ipd_s</strong>: interpulse duration in seconds (time in seconds elapsed since end of preceding pulse) - <strong>snr</strong>: signal-to-noise ratio (bin1_intensity / bin1_bg_std) - <strong>intensity</strong>: intensity of pulse (counts above baseline in bin1) - <strong>bin0_intensity</strong>: counts above baseline in bin0 - <strong>intensity_display</strong>: bin1_intensity + bin1_bg_mean - <strong>binratio</strong>: bin0_intensity / bin1_intensity - <strong>bg_mean</strong>: bin1 background mean in region of pulse - <strong>bg_std</strong>: bin1 background standard deviation in region pulse - <strong>bin0_bg_mean</strong>: bin0 background mean in region of pulse - <strong>bin0_bg_std</strong>: bin0 background standard deviation in region pulse <strong>{run_id}.csv.gz</strong>: Compressed comma-separated value file containing RS/ROI properties computed from raw pulses.bin file by included RS caller (example in <strong>rs_caller.py</strong>). - <strong>ap</strong>: unique aperture index on chip - <strong>ROI</strong>: ordinal ROI number in the aperture, 0-indexed - <strong>start_p</strong>: index (.loc) of first pulse in the ROI (inclusive) in pulse dataframe - <strong>end_p</strong>: index (.loc) of last pulse in the ROI (inclusive) in pulse dataframe - <strong>start_f</strong>: first frame of the first pulse in the ROI (inclusive) - <strong>end_f</strong>: Last frame of the last pulse in the ROI (exclusive) - <strong>start_s</strong>: Time (in seconds elapsed from beginning of run) of the start of the ROI - <strong>end_s</strong>: Time (in seconds elapsed from beginning of run) of the end of the ROI - <strong>dur_f</strong>: Duration in frames of the ROI - <strong>dur_s</strong>: Duration in seconds of the ROI - <strong>num_pulses</strong>: Number of pulses in the ROI (that also passed filtering during ROI-calling) - <strong>pw_mean</strong>: Mean pulse duration (in seconds) of pulses in the ROI - <strong>ipd_mean</strong>: Mean inter-pulse duration (in seconds) of pulses in the ROI - <strong>snr_mean</strong>: Mean signal-to-noise ratio of pulses in the ROI - <strong>intensity_mean</strong>: Mean intensity above baseline of pulses in the ROI - <strong>binratio_norm</strong>: Estimated pulse bin ratio of pulses in the ROI, according to the following equation: sum(bin0_intensity*dur_f) / np.sum(bin1_intensity*dur_f) - <strong>ROI_score</strong>: ROI quality score (0-1 from least to most likely to contain recognizer-peptide recognition pulsing) - <strong>binratio_skew</strong>: bin ratio correction factor accounting for binning signal timing differences across the chip. This factor has already been applied to the binratio_norm column</pre>

opencc-by-4.0Aug 2022View details →
dryad36/100

Nanopore sequencing data analysis using Microsoft Azure cloud computing service

<p>Genetic information provides insights into the exome, genome, epigenetics and structural organisation of the organism. Given the enormous amount of genetic information, scientists are able to perform mammoth tasks to improve the standard of health care such as determining genetic influences on outcome of allogeneic transplantation. Cloud-based computing has increasingly become a key choice for many scientists, engineers and institutions as it offers on-demand network access and users can conveniently rent rather than buy all required computing resources. With the positive advancements of cloud computing and nanopore sequencing data output, we were motivated to develop an automated and scalable analysis pipeline utilizing cloud infrastructure in Microsoft Azure to accelerate HLA genotyping service and improve the efficiency of the workflow at lower cost. In this study, we describe (i) the selection process for suitable virtual machine sizes for computing resources to balance between the best performance versus cost-effectiveness; (ii) the building of Docker containers to include all tools in the cloud computational environment; (iii) the comparison of HLA genotype concordance between the in-house manual method and the automated cloud-based pipeline to assess data accuracy. In conclusion, the Microsoft Azure cloud-based data analysis pipeline was shown to meet all the key imperatives for performance, cost, usability, simplicity and accuracy. Importantly, the pipeline allows for the ongoing maintenance and testing of version changes before implementation. This pipeline is suitable for data analysis from MinION sequencing platforms and could be adopted for other data analysis application processes.</p>

opencc-zeroOct 2022View details →
zenodo36/100

VADER TyrT Illumina sequencing data

<p>FASTQ files used in data analysis for &quot;Virus-assisted directed evolution of enhanced suppressor tRNAs in mammalian cells.&quot;</p>

opencc-by-4.0Oct 2022View details →
zenodo36/100

VADER PylT-PyOtR Illumina sequencing data

<p>FASTQ files used in data analysis for &quot;PyOtR: a novel Pyrrolysyl tRNA evolved for enhanced unnatural amino acid incorporation.&quot;</p>

opencc-by-4.0Oct 2022View details →
zenodo36/100

VADER TyrT-MARIO Illumina sequencing data

<p>FASTQ files used in data analysis for &quot;Evolution of MARIO, an improved Tyrosyl tRNA for more efficient genetic code expansion.&quot;</p>

opencc-by-4.0Oct 2022View details →
zenodo36/100

Raw DNA sequence data of an individual known as "whitequark" (part 1)

<p>Whole genome sequenced on NovaSeq 6000, paired-end 2x150bp with&nbsp;350bp insert.&nbsp;30-40&times;&nbsp;coverage.</p> <p>This dataset can be used by anyone, with attribution.</p>

opencc-by-nc-4.0Sep 2017View details →
zenodo36/100

Raw DNA sequence data of an indivdual known as "whitequark" (part 2)

<p>Whole genome sequenced on NovaSeq 6000, paired-end 2x150bp with 350bp insert. 30-40× coverage.</p> <p>This dataset can be used by anyone, with attribution.</p>

opencc-by-nc-4.0Sep 2017View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record