Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

2,848

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

2,848 results for “sequence data”

Learn how ShareScore rates datasets ↗
zenodo36/100

Training data for 'Mapping-by-sequencing' tutorial (Galaxy Training Material)

<p>The data provided here are part of a Galaxy Training Network tutorial that demonstrates mapping-by-sequencing analysis and represent a subsample of the data used in Sun &amp; Schneeberger, 2015 (DOI:10.1007/978-1-4939-2444-8_19).</p>

opencc-by-4.0Dec 2017View details →
zenodo36/100

Data for paper: Transfer learning for cross-context prediction of protein expression from 5'UTR sequence

<p>This depsit contains data for the paper entitled: "<strong>Transfer learning for cross-context prediction of protein expression from 5'UTR sequence</strong>".</p> <p>The <strong>rebeca.zip</strong> file contains a snapshot of the rebeca package which can be used to train, fine tune and test the CONV-LSTM model used in this study.</p> <p>The <strong>datasets.zip</strong> file contains the compiled sequence to expression datasets from across all Flow-seq expressions considered in this study.&nbsp;</p> <p>The <strong>analysis.zip</strong> file contains all data files and jupyter notebooks necessary to reproduce our analysis. Each Flow-seq study has a dedicated folder (e.g., `fepB') with two sub-folders: 1. The `data\_split' folder, which contains the steps necessary to split the Flow-seq data for our ML experiments (a `readme.txt' file describes the input and output files and a jupyter notebook is available to reproduce the data split); 2. The `data\_analysis' folder, which contains a jupyter notebook and the necessary input files to reproduce the analysis of our experiments.</p>

opencc-by-4.0Apr 2024View details →
zenodo36/100

Experimental data for "An End-to-End Coding Scheme for DNA-Based Data Storage With Nanopore Sequenced Reads"

<p>The experimental dataset used in "An End-to-End Coding Scheme for DNA-Based Data Storage With Nanopore Sequenced Reads."</p> <p>A set of 91,766 150-nt oligos were synthesised with GenScript (oligos.fasta). Each oligo consists of a pseudo-random 110-nt payload flanked by 20-nt primers at each end. The strands are split in three roughly equal groups (two groups of 30,589 and one group of 30,588). Each group has a dedicated primer pair for targeted PCR amplification (the primer pairs used for amplification are provided in primers_synthesis.fasta). The pseudo-random payload was designed to avoid primer-payload collisions.</p> <p>For each file, a sample from the synthesised pool was PCR amplified using the corresponding primer pair and sequenced using Oxford Nanopore Technologies MinION sequencing device following the standard library preparation protocol for amplicon DNA. The raw reads were basecalled using guppy, either in fast- ("acc-false") or high-accuracy ("acc-true") regime. The basecaller generated two groups of reads&mdash;"passQ-true" for the reads that passed the quality-score threshold of 8 and "passQ-false" for those that did not. For each group of reads, a BLAST-based fuzzy search for primer sequences was performed and, based on the resulting alignments, the segments containing the correct primer pairs and located at a distance of 150+-15nt were extracted (separately for forward and reverse-complemented reads). The segments are then assigned to the closest synthesized strand based on Levenshtein distance. The resulting clusters are used to estimate the parameters of the end-to-end DNA storage channel model and to test the proposed error-correction scheme.</p> <p>The archive clustered_read_segments.tar.gz contains 12 sub-archives, for each file (0,1,2), accuracy ("acc-true" or "acc-false"), and Q-score ("passQ-true" or "passQ-false"). Within each sub-archive, there are two folders (one for forward read segments and one for backward read segments), and each folder contains two files: one for the reference synthesised (or "transmitted") sequences that correspond to the file in question ("TX__" &mdash; e.g., "TX__file=0_accBaCa=true_passQ=true_filter=true_forward_.txt") and another file for the sequenced (or "received") segment clusters ("RX__" &mdash; e.g., "RX__file=0_accBaCa=true_passQ=true_filter=true_forward_.txt"). The received clusters in the "RX__" file are ordered in correspondence with the synthesised sequences in the "TX__" file, and a line "===============================" is used as a separator.</p>

opencc-by-4.0May 2024View details →
zenodo36/100

Data supporting publication "Metagenomic Immunoglobulin Sequencing (MIG-Seq) Exposes Patterns of IgA Antibody Binding in the Healthy Human Gut Microbiome"

<p>Data supporting publication "Metagenomic Immunoglobulin Sequencing (MIG-Seq) Exposes Patterns of IgA Antibody Binding in the Healthy Human Gut Microbiome"</p>

opencc-by-4.0Nov 2023View details →
dryad36/100

Sequencing data for seabird eDNA in long-nosed fur seal diets from southeastern Australia

<p>Wildlife conflicts require robust quantitative data on incidence and impacts, particularly among species of conservation and cultural concern. We apply a multi-assay framework to quantify predation in a southeastern Australian scenario where complex management implications and calls for predator culling have grown despite a paucity of data on seabird predation by recovering populations of long-nosed fur seals (<em>Arctocephalus forsteri</em>). We apply two ecological surveillance techniques to analyse this predator's diet – traditional morphometric (prey hard-part) and environmental DNA metabarcoding (genetic) analyses using an avian specific primer for the 12S ribosomal RNA (rRNA) gene – to provide managers with estimated predation incidence, number of seabird species impacted and inter-prey species relative importance to the predator. DNA metabarcoding identified additional seabird taxa and provided relative quantitative information where multiple prey species occur within a sample; while parallel use of both genetic and hard-part analyses revealed a greater diversity of taxa than either method alone. Using data from both assays, the estimated frequency of occurrence of predation on seabirds by long-nosed fur seals ranged from 9.1–29.3% of samples and included up to 6 detected prey species. The most common seabird prey was the culturally valued little penguin (Eudyptula minor) that occurred in 6.1–25.3% of samples, higher than previously reported from traditional morphological assays alone. We then explored DNA haplotype diversity for little penguin genetic data, as a species of conservation concern, to provide a preliminary estimate of the number of individuals consumed. Polymorphism analysis of consumed little penguin DNA identified five distinct mitochondrial haplotypes – representing a minimum of 16 individual penguins consumed across 10 fur seal scat samples from 99 sampled across southeastern Australia. We recommend rapid uptake and development of cost-effective genetic techniques and broader spatiotemporal sampling of fur seal diets to further quantify predation and hotspots of concern for wildlife conflict management.</p>

opencc-zeroMay 2024View details →
dryad36/100

Data from: Targeted genotyping-by-sequencing of potato and data analysis with R/polyBreedR

<p>"Mid-density" targeted genotyping-by-sequencing (GBS) combines trait-specific markers with thousands of genomic markers at an attractive price for linkage mapping and genomic selection. A 2.5K targeted GBS assay for potato was developed using the DArTag<sup>TM</sup> technology and later expanded to 4K targets. Genomic markers were selected from the potato Infinium<sup>TM</sup> SNP array to maximize genome coverage and polymorphism rates. The DArTag and SNP array platforms produced equivalent dendrograms in a test set of 298 tetraploid samples, and 83% of the common markers showed good quantitative agreement, with RMSE (root-mean-squared-error) less than 0.5. DArTag is suited for genomic selection candidates in the clonal evaluation trial, coupled with imputation to a higher-density platform for the training population. Using the software polyBreedR, an R package for the manipulation and analysis of polyploid marker data, the RMSE for imputation by linkage analysis was 0.15 in a small half-diallel population (N=85), which was significantly lower than the RMSE of 0.42 with the Random Forest method. Regarding high-value traits, the DArTag markers for resistance to potato virus Y, golden cyst nematode, and potato wart appeared to track their targets successfully, as did multi-allelic markers for maturity and tuber shape. In summary, the potato DArTag assay is a transformative and publicly available technology for potato breeding and genetics.</p>

opencc-zeroFeb 2025View details →
zenodo36/100

Control panel created from 30-40 Nanopore or PacBio HiFi sequencing data from the Human Pangenome Reference Consortium

<p>This is control panel for <a href="https://github.com/friend1ws/nanomonsv">nanomonsv</a> software, which is expected to exclude many false positives as well as improve computational cost. This is made by aligning 30-40 Nanopore or PacBio HiFi sequencing data from Human Pangenome Reference Consortium (HPRC) to the GRCh38 or CHM13 reference genomes with <a href="https://github.com/lh3/minimap2">minimap2</a> version 2.24.</p> <p><strong>When you use these control panels and publish, do not forget to credit to <a href="https://humanpangenome.org/data-use-protocol/">HPRC</a>!</strong></p> <div> <div> <div> <p>Reference genomes:</p> <ul> <li>GRCh38: <a href="https://ftp.ncbi.nlm.nih.gov/genomes/all/GCA/000/001/405/GCA_000001405.15_GRCh38/seqs_for_alignment_pipelines.ucsc_ids/GCA_000001405.15_GRCh38_no_alt_analysis_set.fna.gz">Download GRCh38</a></li> <li>CHM13: <a href="https://s3-us-west-2.amazonaws.com/human-pangenomics/T2T/CHM13/assemblies/analysis_set/chm13v2.0_maskedY_rCRS.fa.gz">Download CHM13</a> <div> <div> <div> <div>&nbsp;</div> </div> </div> </div> </li> </ul> </div> </div> </div>

opencc-by-4.0Jun 2024View details →
zenodo36/100

Datasets for evaluating SCEMENT: Scalable and Memory Efficient Integration of Large-scale Single Cell RNA-sequencing Data

<p>This resource contains pre-processed A. thaliana root , the H. sapiens aortic valve datasets, PBMC Covid atlas and public 10x datasetse used in the paper, SCEMENT: Scalable and Memory Efficient Integration of Large-scale Single Cell RNA-sequencing Data. The raw datasets provided in the links below are pre-processed for quality control with respect to both cells and genes.&nbsp;</p> <p>A. thaliana datasets are sourced from the following locations at <a href="https://www.ebi.ac.uk/gxa/sc/home">Single-cell Gene expression Atlas </a>and <a href="https://www.ncbi.nlm.nih.gov/geo/">Gene Expression Omnibus (GEO)</a>:</p> <ol> <li>E-GEOD-121619 : <a title="E-GEOD-121619" href="https://www.ebi.ac.uk/gxa/sc/experiments/E-GEOD-121619/results">https://www.ebi.ac.uk/gxa/sc/experiments/E-GEOD-121619/results</a></li> <li>E-GEOD-152766 : <a href="https://www.ebi.ac.uk/gxa/sc/experiments/E-GEOD-152766/results">https://www.ebi.ac.uk/gxa/sc/experiments/E-GEOD-152766/results</a></li> <li>E-GEOD-158761 : &nbsp; <a href="https://www.ebi.ac.uk/gxa/sc/experiments/E-GEOD-158761/results">https://www.ebi.ac.uk/gxa/sc/experiments/E-GEOD-158761/results</a></li> <li>E-GEOD-123013 : <a href="https://www.ebi.ac.uk/gxa/sc/experiments/E-GEOD-123013/results">https://www.ebi.ac.uk/gxa/sc/experiments/E-GEOD-123013/results</a></li> </ol> <p>H. sapiens datasets are obtained from the NCBI database : <a href="https://www.ncbi.nlm.nih.gov/bioproject/PRJNA562645/">https://www.ncbi.nlm.nih.gov/bioproject/PRJNA562645/&nbsp;</a></p> <ol> <li>GSE152766: <a href="https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE152766">https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE152766</a></li> <li>GSE158761: <a href="https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE158761">https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE158761</a></li> </ol> <p>All COVID atlas datasets are from: <a href="http://covid19.cancer-pku.cn">http://covid19.cancer-pku.cn</a> . covid_atlas_data1.zip contains the h5ad files and covid_atlas_data2.zip contains the Seurat rds files.</p> <p>PBMC datasets are from the following public sources:</p> <table> <tbody> <tr> <td>Dataset Name</td> <td>Chemistry Version</td> <td>Web Link</td> </tr> <tr> <td>10k Human PBMCs, 3' v3.1, Chromium X</td> <td>v3.1</td> <td><a href="https://www.10xgenomics.com/datasets/10k-human-pbmcs-3-ht-v3-1-chromium-x-3-1-high">https://www.10xgenomics.com/datasets/10k-human-pbmcs-3-ht-v3-1-chromium-x-3-1-high</a></td> </tr> <tr> <td>20k Human PBMCs, 3' HT v3.1, Chromium X</td> <td>v3.1</td> <td><a href="https://www.10xgenomics.com/datasets/20-k-human-pbm-cs-3-ht-v-3-1-chromium-x-3-1-high-6-1-0">https://www.10xgenomics.com/datasets/20-k-human-pbm-cs-3-ht-v-3-1-chromium-x-3-1-high-6-1-0</a></td> </tr> <tr> <td>10k Human PBMCs, 3' v3.1, Chromium Controller</td> <td>v3.1</td> <td><a href="https://www.10xgenomics.com/datasets/10k-human-pbmcs-3-v3-1-chromium-controller-3-1-high">https://www.10xgenomics.com/datasets/10k-human-pbmcs-3-v3-1-chromium-controller-3-1-high</a></td> </tr> <tr> <td>Healthy PBMC Chromium Connect (channel 1)</td> <td>v3.1</td> <td><a href="https://www.10xgenomics.com/datasets/peripheral-blood-mononuclear-cells-pbm-cs-from-a-healthy-donor-chromium-connect-channel-1-3-1-standard-3-1-0">https://www.10xgenomics.com/datasets/peripheral-blood-mononuclear-cells-pbm-cs-from-a-healthy-donor-chromium-connect-channel-1-3-1-standard-3-1-0</a></td> </tr> <tr> <td>Healthy PBMC Chromium Connect (channel 5)</td> <td>v3.1</td> <td><a href="https://www.10xgenomics.com/datasets/peripheral-blood-mononuclear-cells-pbm-cs-from-a-healthy-donor-chromium-connect-channel-5-3-1-standard-3-1-0">https://www.10xgenomics.com/datasets/peripheral-blood-mononuclear-cells-pbm-cs-from-a-healthy-donor-chromium-connect-channel-5-3-1-standard-3-1-0</a></td> </tr> <tr> <td>10k PBMCs from a Healthy Donor (v3 chemistry)</td> <td>v3.0</td> <td><a href="https://www.10xgenomics.com/datasets/10-k-pbm-cs-from-a-healthy-donor-v-3-chemistry-3-standard-3-0-0">https://www.10xgenomics.com/datasets/10-k-pbm-cs-from-a-healthy-donor-v-3-chemistry-3-standard-3-0-0</a></td> </tr> <tr> <td>1k PBMCs from a Healthy Donor (v2 chemistry)</td> <td>v2.0</td> <td><a href="https://www.10xgenomics.com/datasets/1-k-pbm-cs-from-a-healthy-donor-v-2-chemistry-3-standard-3-0-0">https://www.10xgenomics.com/datasets/1-k-pbm-cs-from-a-healthy-donor-v-2-chemistry-3-standard-3-0-0</a></td> </tr> <tr> <td>1k PBMCs from a Healthy Donor (v3 chemistry)</td> <td>v3.0</td> <td><a href="https://www.10xgenomics.com/datasets/1-k-pbm-cs-from-a-healthy-donor-v-3-chemistry-3-standard-3-0-0">https://www.10xgenomics.com/datasets/1-k-pbm-cs-from-a-healthy-donor-v-3-chemistry-3-standard-3-0-0</a></td> </tr> <tr> <td>Fresh 68k PBMCs (Donor A)</td> <td>v1.0</td> <td><a href="https://www.10xgenomics.com/datasets/fresh-68-k-pbm-cs-donor-a-1-standard-1-1-0">https://www.10xgenomics.com/datasets/fresh-68-k-pbm-cs-donor-a-1-standard-1-1-0</a></td> </tr> <tr> <td>Frozen PBMCs (Donor A)</td> <td>v1.0</td> <td><a href="https://www.10xgenomics.com/datasets/frozen-pbm-cs-donor-a-1-standard-1-1-0">https://www.10xgenomics.com/datasets/frozen-pbm-cs-donor-a-1-standard-1-1-0</a></td> </tr> <tr> <td>Frozen PBMCs (Donor B)</td> <td>v1.0</td> <td><a href="https://www.10xgenomics.com/datasets/frozen-pbm-cs-donor-b-1-standard-1-1-0">https://www.10xgenomics.com/datasets/frozen-pbm-cs-donor-b-1-standard-1-1-0</a></td> </tr> <tr> <td>Frozen PBMCs (Donor C)</td> <td>v1.0</td> <td><a href="https://www.10xgenomics.com/datasets/frozen-pbm-cs-donor-c-1-standard-1-1-0">https://www.10xgenomics.com/datasets/frozen-pbm-cs-donor-c-1-standard-1-1-0</a></td> </tr> <tr> <td>PBMCs from a Healthy Donor: Whole Transcriptome Analysis</td> <td>v3.1</td> <td><a href="https://www.10xgenomics.com/datasets/pbm-cs-from-a-healthy-donor-whole-transcriptome-analysis-3-1-standard-4-0-0">https://www.10xgenomics.com/datasets/pbm-cs-from-a-healthy-donor-whole-transcriptome-analysis-3-1-standard-4-0-0</a></td> </tr> <tr> <td>PBMC 600K</td> <td>v1</td> <td><a href="https://www.ebi.ac.uk/gxa/sc/experiments/E-HCAD-4/downloads">https://www.ebi.ac.uk/gxa/sc/experiments/E-HCAD-4/downloads</a></td> </tr> <tr> <td>GSM4560071</td> <td>v2.0</td> <td><a href="https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSM4560071">https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSM4560071</a></td> </tr> <tr> <td>GSM4560074</td> <td>v2.0</td> <td><a href="https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSM4560074">https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSM4560074</a></td> </tr> <tr> <td>GSM4560070</td> <td>v2.0</td> <td><a href="https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSM4560070">https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSM4560070</a></td> </tr> </tbody> </table> <p>References for the Datasets :</p> <ol> <li>H. sapiens dataset: Kang Xu, Shangbo Xie,Yuming Huang,Tingwen Zhou, Ming Liu, Peng Zhu, Chunli Wang, Jiawei Shi, Fei Li,Frank W. Sellke and Nianguo Dong (2020) Cell-Type Transcriptome Atlas of Human Aortic Valves Reveal Cell Heterogeneity and Endothelial to Mesenchymal Transition Involved in Calcific Aortic Valve Disease.</li> <li>E-GEOD-152766: Shahan R, Hsu C, Nolan TM, Cole BJ, Taylor IW et al. (2020) A single cell Arabidopsisroot atlas reveals developmental trajectories in wild type and cell identity mutants.</li> <li>E-GEOD-121619: Jean-Baptiste K, McFaline-Figueroa JL, Alexandre CM, Dorrity MW, Saunders L et al. (2019) Dynamics of Gene Expression in Single Root Cells of Arabidopsis thaliana.</li> <li>E-GEOD-123013: Ryu KH, Huang L, Kang HM, Schiefelbein J. (2019) Single-Cell RNA Sequencing Resolves Molecular Relationships Among Individual Plant Cells.</li> <li>E-GEOD-158761: Gala HP, Lanctot A, Jean-Baptiste K, Guiziou S, Chu JC et al. (2020) A single cell view of the transcriptome during lateral root initiation in Arabidopsis thaliana.</li> <li>COVID Atlas Reference: Xianwen Ren, Wen Wen, Xiaoying Fan et.al. (2021) COVID-19 immune features revealed by a large-scale single-cell transcriptome atlas</li> <li>PBMC data are downloaded from respective links</li> </ol>

opencc-by-4.0Jun 2024View details →
dryad36/100

Data from: Affordable de novo generation of fish mitogenomes using amplification-free enrichment of mitochondrial DNA and deep sequencing of long fragments

<p>Biomonitoring surveys from environmental DNA make use of metabarcoding tools to describe the community composition. These studies match their sequencing results against public genomic databases to identify the species. However, mitochondrial genomic reference data are yet incomplete, only a few genes may be available, or the suitability of existing sequence data is suboptimal for species-level resolution. Here we present a dedicated and cost-effective workflow with no DNA amplification for generating complete fish mitogenomes for the purpose of strengthening fish mitochondrial databases. Two different long-fragment sequencing approaches using Oxford Nanopore sequencing coupled with mitochondrial DNA enrichment were used. One where the enrichment is achieved by preferential isolation of mitochondria followed by DNA extraction and nuclear DNA depletion ('mitoenrichment').  A second enrichment approach takes advantage of the CRISPR-Cas9 targeted scission on previously dephosphorylated DNA ('targeted mitosequencing'). The sequencing results varied between tissue, species, and integrity of the DNA. The mitoenrichment method yielded 0.17-12.33 % of sequences on target and a mean coverage ranging from 74.9 to 805-fold. The targeted mitosequencing experiment from native genomic DNA yielded 1.83-55 % of sequences on target and a 38 to 2123-fold mean coverage. This produced complete the mitogenome of species with homopolymeric regions, tandem repeats, and gene rearrangements. We demonstrate that deep sequencing of long fragments of native fish DNA is possible and can be achieved with low computational resources in a cost-effective manner, opening the discovery of mitogenomes of non-model or understudied fish taxa to a broad range of laboratories worldwide.</p>

opencc-zeroJun 2024View details →
zenodo36/100

Data for Isotype-aware Inference of B cell Clonal Lineage Trees from Single-cell Sequencing Data

<p>This is the accompanying data to the manuscript titled<em> Isotype-aware Inference of B cell Clonal Lineage Trees from Single-cell Sequencing Data</em>. To reproduce the TRIBAL output please use this <a href="https://doi.org/10.5281/zenodo.12741290">code repository</a> as the arguments and codebase may have changed since release.&nbsp;</p>

opencc-by-4.0Jul 2024View details →
zenodo36/100

Illumina RNA-Sequencing fastq data from insecticide resistant Anopheles gambiae s.l

<p>This is a dataset of Illumina RNA sequencing reads, for a project investigating resistance to Pirimiphos-methyl in the major malaria vectors, Anopheles gambiae and Anopheles coluzzii. There are four biological replicates for the following conditions:</p> <p>&nbsp;</p> <p>Ngousso (susceptible)</p> <p>Kisumu (susceptible)</p> <p>Bouake gambiae unexposed</p> <p>Bouake gambiae PM survivors</p> <p>Bouake coluzzii unexposed&nbsp;</p> <p>Bouake coluzzii PM&nbsp; survivors&nbsp;</p> <p>&nbsp;</p> <p>SRA submission: SUB14596876</p> <p>&nbsp;</p>

opencc-by-4.0Jul 2024View details →
zenodo36/100

single-nucleus RNA sequencing data of Minimal Change Disease and Focal Segmental Glomerulosclerosis patients

<p>single-nucleus RNA sequencing data of Minimal Change Disease and Focal Segmental Glomerulosclerosis patients</p>

opencc-by-4.0Jul 2024View details →
zenodo36/100

Fig. 13 Erythraeus regalis, larva. a Gnathosoma, ventral view. b in Towards resolving the double classification in Erythraeus (Actinotrichida: Erythraeidae): matching larvae with adults using 28S sequence data and experimental rearing

Fig. 13 Erythraeus regalis, larva. a Gnathosoma, ventral view. b Palp tarsus

opencc-by-4.0May 2016View details →
zenodo36/100

Data from: Individual Movement - Sequence Analysis Method (IM-SAM): characterising spatio-temporal patterns of animal trajectories across scales and landscapes

<p>Dataset included in Zenodo supports the analyses performed in &quot;<em>Individual Movement - Sequence Analysis Methods (IM-SAM) characterising spatio-temporal patterns of animal trajectories across scales and landscapes.</em>&quot;</p> <p>The dataset includes one RDS file, that can be easily loaded into R using the readRDS function. The RDS file consists out of a list including two objects per animal:</p> <ul> <li>Object 1 contains a data frame with the real and simulated sequences for an animal. e.g., ls[[1]][[1]]&nbsp;</li> <li>Object 2 contains the home range in raster format of an animal. e.g., ls[[1]][[2]]</li> </ul> <p>The data frames in object 1 contain real habitat use sequences and corresponding simulated habitat use sequences generated in the home range of the specific individual (900 simulated sequences: 6 habitat selection rules x 3 selection coefficients x 50 repetitions). Open and closed habitats are respectively encoded by 0 and 1. The first 96 columns of each row in a data frame represent a 16-day habitat use sequence, with a fixed 4-hour relocation interval (0, 4, 8, 12, 16 and 20h). Column names are named as follows: Day_1_0h, Day_1_4h,..., Day_16_20h. In the next columns we provide the selection coefficients (columns 97-99), the habitat selection rules (or pattern, columns 100-102) and the number of missing values (mvs, columns, 103-104) for each of the real and simulated sequences. Note that simulated sequences have no missing values (i.e. values are always 0.00) and for real sequences there is no selection coefficient or habitat selection rule (i.e. values are always xxx).</p> <p>Rownames of simulated sequences are composed out of the habitat selection rule (c, o, a24, a33, a42 and u), the selection coefficient (5, 10, 50) and the replicate (1 to 50), separated by dashes. For example, the first simulated sequence in the first data frame (ls[[1]][[1]][1,]) is described as a24_10_1. The rownames of real sequences instead are composed out of the individuals&#39; identifier, the biweekly period (1 to 23) and the year. For example, the first real sequence in the first data frame (ls[[1]][[1]][901,]) is described as 1_5_2006.</p> <p><br> &nbsp;</p>

opencc-by-4.0May 2018View details →
zenodo36/100

Species abundance information improves sequence taxonomy classification accuracy - Qiita data

<p><a href="https://qiime2.org">QIIME 2</a> Artifacts containing microbiome samples, organised by EMPO 3 classification.</p> <p>Used to test <a href="https://library.qiime2.org/plugins/q2-clawback/">q2-clawback</a>.</p> <p>Downloaded from <a href="https://qiita.ucsd.edu">Qiita</a>.&nbsp;Draws from the following Qiita study ids:</p> <p>11113[1], 11444, 1716, 10369[2], 990[3], 2080, 1713, 894, 1289, 1883, 1673, 1288, 10353, 2192[4], 10323, 678, 1773, 662, 1799, 864, 1481, 1024[5], 1064, 2182, 10934, 1674, 1795[6], 10273, 10283[7], 10422[8], 804, 10308, 1056[9], 2382[5], 1240, 889, 1041, 1717, 1222, 11149, 11669, 807[10], 10245, 1711, 1721, 910, 1001, 895, 550[11], 1747[12], 713[13], 755, 861, 958[14], 11161[15], 11154[16], 945, 723, 1715, 1714, 10798.</p> <p>References<br> 1. Schulfer, A. F. et al. Nat Microbiol 3, 234&ndash;242 (2017).<br> 2. Ruhe, J. et al. Front Plant Sci 7 (2016).<br> 3. O&#39;Brien, S. L. et al. Environ Microbiol 18, 2039&ndash;2051 (2016).<br> 4. Lax, S. et al. Science 345, 1048&ndash;1052 (2014).<br> 5. Zarraonaindia, I. et al. mBio 6 (2015).<br> 6. Navas-Molina, J. A. et al. in Methods Enzymol 371&ndash;444 (2013).<br> 7. Fang, X. et al. Front Microbiol 9 (2018).<br> 8. Tripathi, A. et al. mSystems 3 (2018).<br> 9. Delsuc, F. et al. Mol Ecol 23, 1301&ndash;1317 (2013).<br> 10. Gibbons, S. M. et al. PLoS ONE 9, e97435 (2014).<br> 11. Caporaso, J. G. et al. Genome Biol 12, R50 (2011).<br> 12. Hyde, E. R. et al. mSystems 1 (2016).<br> 13. Brazelton, W. J., Nelson, B. &amp; Schrenk, M. O. Front Microbiol 2 (2012).<br> 14. Vitaglione, P. et al. Am J Clin Nutr 101, 251&ndash;261 (2014).<br> 15. Spirito, C. M., Marzilli, A. M. &amp; Angenent, L. T. Environ Sci Technol 52, 13438&ndash;13447 (2018).<br> 16. Pham, V. T. H. et al. Sci Rep 7 (2017).</p>

opencc-by-4.0Jan 2019View details →
zenodo36/100

Species abundance information improves sequence taxonomy classification accuracy - HMP and NCBI data

<p>Data used to test&nbsp;<a href="https://library.qiime2.org/plugins/q2-clawback/">q2-clawback</a>&nbsp;using taxonomic weights derived from shotgun sequencing experiments.</p> <p>Includes reference sequences and taxonomies derived from the NCBI RefSeq database[1] and paired amplicon and shotgun sequencing results downloaded from the Human Microbiome Project[2].</p> <p>References<br> 1.&nbsp;O&rsquo;Leary, N. A. et al. en. Nucleic Acids Res. 44, D733&ndash;45 (2016).<br> 2. Huttenhower, C. et al. Nature 486, 207 (2012).</p>

opencc-by-4.0Jan 2019View details →
zenodo36/100

Data for 'Comparative Analysis of Single-Cell RNA Sequencing Methods'

<p>Raw sequencing data to &quot;Comparative Analysis of Single-Cell RNA Sequencing Methods&quot;.&nbsp;</p> <p>https://www.ncbi.nlm.nih.gov/pubmed/28212749</p> <p>&nbsp;</p> <p>In addition to the GEO submission&nbsp;https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE75790, you can find here raw bam files for UMI-methods tagged with cell barcode and UMI sequences.</p> <p>MD5 checksum:&nbsp;f10825509952fffd9c4dc0c1dcb9eb8e</p>

opencc-by-nc-sa-4.0Feb 2017View details →
zenodo36/100

Sequencing data: Fungi, fire and insects: Protea infructescences as reservoirs for fungal biodiversity in fire-prone environments

<p>This is the dataset for a submitted manuscript entitled &quot;Fungi, fire and insects: <em>Protea</em> infructescences as reservoirs for fungal biodiversity in fire-prone environments&quot;.</p> <p>The upload contains all forward and reverse paired end .fastq files. The files are already demultiplexed. The metadata.xlsx file contains sample information, gps coordinates, etc.</p> <p>&nbsp;</p>

opencc-by-4.0Dec 2018View details →
zenodo36/100

Data for: Beyond signal quality: The value of unmaintained pH, dissolved oxygen, and oxidation-reduction potential sensors for remote performance monitoring of on-site sequencing batch reactors

<p>Sensor maintenance is time-consuming and is a bottleneck for monitoring on-site wastewater treatment systems. Hence, we compare maintained and unmaintained sensors to monitor the biological performance of a small-scale sequencing batch reactor (SBR). The sensor types are ion-selective pH, optical dissolved oxygen (DO), and oxidation-reduction potential (ORP) with platinum electrode. We created soft sensors using engineered features: ammonium valley for pH, oxidation ramp for DO, and nitrite ramp for the ORP. Four soft sensors based on unmaintained pH sensors correctly identified the completion of the ammonium oxidation (89 to 91 out of 107 cycles), about as many times as soft sensors based on a maintained pH sensor (91 out of 107 cycles). In contrast, the DO soft sensor using data from a maintained sensor showed slightly better (89 out of 96 cycles) detection performance than that using data from two unmaintained sensors (77, respectively 82 out of 96 correct). Furthermore, the DO soft sensor using maintained data is much less sensitive to the optimisation of cut-off frequency and slope tolerance than the soft sensor using unmaintained data. The nitrite ramp provided no useful information on the state of nitrite oxidation, so no comparison of maintained and unmaintained ORP sensors was possible in this case. We identified two hurdles when designing soft sensors for unmaintained sensors: i) Sensors&#39; type- and design-specific deterioration affects performance. ii) Feature engineering for soft sensors is sensor type specific, and the outcome is strongly influenced by operational parameters such as the aeration rate. In summary, the results with the provided soft sensors show that frequent sensor maintenance is not necessarily needed to monitor the performance of SBRs. Without sensor maintenance monitoring smalls-scale SBRs becomes practicable, which could improve the reliability of unstaffed on-site treatment systems substantially.</p>

opencc-zeroDec 2018View details →
zenodo36/100

Data file with manuscript titled 'A Structurally Validated Sequence Alignment of 497 Human Protein Kinase Domains'

<p>The files used in different analysis reported in the manuscript titled - &#39;A Structurally-Validated Multiple Sequence Alignment of 497 Human Protein Kinase Domains&#39; are shared at two locations. Following is a brief description of these files.</p> <p>Location -&nbsp; https://github.com/DunbrackLab/Kinases<br> 1. HMM profile files - HMM files for each of the nine groups computed separately labeled as Groupname.hmm, like AGC.hmm<br> 2. HMM profile file - HMM file computed from the full alignment including all the sequences - Human-PK.hmm<br> 3. Score files - HMM scores of each kinase sequence against all the groupwise HMMs both for iteration1 (HMM-iter1-scores-tables.txt) and iteration2 (HMM-iter1-scores-tables.txt)<br> 4. Jalview session file - Kinase alignment with sequences colored by secondary structure information from PDB file if the structure is known; or predicted secondary structure if the experimental structure is not known. The file could be opened in Jalview - kinases-PDB-SSPred.jvp</p> <p>Location - https://zenodo.org/record/3445533<br> 1. The file contains list of residue pairs aligned in pairwise structural alignments of 272 human protein kinases which were used as a benchmark in the study. The alignments were created by FATCAT and optimized by SE program.</p>

opencc-by-4.0Sep 2019View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record