Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
2,848
datasets available to search
ShareScore release 0.9.0
Dataset results
2,848 results for “sequence data”
RNA sequencing data of intestinal tissue from wild-type C57BL/6 mice and ITGB6-overexpressing transgenic mice
GEO Series GSE188361. Mus musculus. 10 samples. Type: Expression profiling by high throughput sequencing.
RNA sequencing data of differentially treated cancer associated fibroblast, myeloid-derived supression cells, and ICC tumor cells in ICC patient
GEO Series GSE158755. Homo sapiens. 18 samples. Type: Expression profiling by high throughput sequencing.
Digital transformation of herbal medicine: Conversion to biological entity data using tonifying herbal medicine-induced transcriptome sequencing_HepG2_batchC
GEO Series GSE244704. Homo sapiens. 60 samples. Type: Expression profiling by high throughput sequencing.
Bulk RNA sequencing data of intact and digested rice root for Single-cell RNA-seq protoplasting induced gene filtering
GEO Series GSE283509. Oryza sativa. 8 samples. Type: Expression profiling by high throughput sequencing.
RNA sequencing data from triple-negative breast cancer patient-derived xenografts (PDX)
GEO Series GSE110626. Homo sapiens. 21 samples. Type: Expression profiling by high throughput sequencing.
Mutating Zta(N182) to S, Q, T, I, and V changes sequence specific DNA binding to four types of DNA (65k data set)
GEO Series GSE126588. synthetic construct; Mus musculus. 36 samples. Type: Other.
Childhood Cancer Data Initiative (CCDI): Comprehensive Genomic Sequencing of Pediatric Cancer Cases (CMRI/KUCC)
This study provides paired tumor normal genomic sequencing data from approximately 200 children with cancer, including both solid tumors and leukemias, done by the Children's Mercy Research Institute (CMRI) and University of Kansas Cancer Center (KUCC). These data include whole genome sequencing, whole exome sequencing, bulk RNA sequencing, and single-cell RNA and ATAC sequencing. Additional phenotypic, pathologic, and genetic data, gathered clinically for these samples, are also provided.
RNA-sequencing data from Pseudomonas putida KT2440 wild-type and Hfq-tagged strain form co-immunoprecipitation experiment
GEO Series GSE85581. Pseudomonas putida. 12 samples. Type: Expression profiling by high throughput sequencing.
RNA sequencing data of microglia isolated from brains of WT, Hexb-tdTomato, Hexb-CreERT2 and Hexb-KO mice. {Reporter]
GEO Series GSE148412. Mus musculus. 16 samples. Type: Expression profiling by high throughput sequencing.
Discovery and verification of liver cancer marker genes and variable scission based on second-generation sequencing data analysis
GEO Series GSE136846. Homo sapiens. 12 samples. Type: Expression profiling by high throughput sequencing.
Transcriptome profiling of derived-hepatocyte progenitors from human iPSCs with nanoCAGE - part2 - sequencing data (FASTQ files)
<p>This repository contains raw sequencing data (FASTQ files) produced from Illumina MiSeq run IDs "181114_M00528_0390_000000000-C7P58" (aka "NC_LIMMS3") and "190218_M00528_0406_000000000-CB4HR" (aka "NC_LIMMS4") . Sequencing libraries were prepared following the latest version of the nanoCAGE protocol (Poulain et al., Methods Mol Biol. 2017;1543:57-109. doi: 10.1007/978-1-4939-6716-2_4). They respectively contain a mix of 13 ("NC_LIMMS3") and 24 ("NC_LIMMS4") samples tagged by specific barcode sequences at the 5'-ends (see tables below). The tagmentation step included in the protocol was performed using an equimolar mix of 12 Nextera XT N-series index primers (N701 to N712), therefore "NNNNNNNN" was indicated as index sequence on the Illumina Sample Sheet for the demultiplexing (see tables below). Libraries were sequenced paired-end on Illumina MiSeq system with the MiSeq Reagent Kit v3 (150 cycles: 58 cycles used for READ1, 8 cycles used for the Index, and 84 cycles used for READ2). Genomic alignments (BED files) of paired-end reads on human genome assemblies hg19 and hg38 using the MOIRAI pipeline (Hasegawa et al. BMC Bioinformatics 2014 May 16;15:144. doi: 10.1186/1471-2105-15-144) were deposited at Zenodo under the following Digital Object Identifier: 10.5281/zenodo.2572394.</p> <p><em><strong>"181114_M00528_0390_000000000-C7P58" ("NC_LIMMS3"):</strong></em></p> <p><strong>sample_name group barcode_sequence index_sequence</strong><br> LIMMS43_04_PETRI_S4D7_rep1 iPSC_CLONE_TODAI ACAGAT NNNNNNNN<br> LIMMS44_24_PETRI_S4D7_rep2 iPSC_CLONE_TODAI ATCGTG NNNNNNNN<br> LIMMS45_31_PETRI_S4D7_rep3 iPSC_CLONE_TODAI CACGAT NNNNNNNN<br> LIMMS46_36_PETRI_S4D14_rep1 iPSC_CLONE_TODAI CACTGA NNNNNNNN<br> LIMMS47_46_PETRI_S4D14_rep2 iPSC_CLONE_TODAI CTGACG NNNNNNNN<br> LIMMS48_63_PETRI_S4D14_rep3 iPSC_CLONE_TODAI GAGTGA NNNNNNNN<br> LIMMS49_79_PETRI_CELLARTIS_rep1 iPSC_CLONE_TODAI GTATAC NNNNNNNN<br> LIMMS50_92_PETRI_CELLARTIS_rep2 iPSC_CLONE_TODAI TCGAGC NNNNNNNN<br> LIMMS51_09_PETRI_CELLARTIS_rep3 iPSC_CLONE_TODAI ACATGA NNNNNNNN<br> LIMMS52_21_PETRI_TODAI_rep1 iPSC_CLONE_CELLARTIS ATCATA NNNNNNNN<br> LIMMS53_33_PETRI_TODAI_rep2 iPSC_CLONE_CELLARTIS CACGTG NNNNNNNN<br> LIMMS54_45_PETRI_TODAI_rep3 iPSC_CLONE_CELLARTIS CGATGA NNNNNNNN<br> LIMMS55_57_iPSC_rep1 CONTROL_iPSC GAGATA NNNNNNNN</p> <p><em><strong>"190218_M00528_0406_000000000-CB4HR" ("NC_LIMMS4"):</strong></em></p> <p><strong>sample_name group barcode_sequence index_sequence</strong><br> LIMMS56_04_iPSC_rep4 CONTROL_iPSC ACAGAT NNNNNNNN<br> LIMMS57_24_LSECS_1_11 LSECS_PETRI_MONO ATCGTG NNNNNNNN<br> LIMMS58_31_LSECS_2_11 LSECS_PETRI_MONO CACGAT NNNNNNNN<br> LIMMS59_36_LSECS_3_11 LSECS_PETRI_MONO CACTGA NNNNNNNN<br> LIMMS60_46_LSECS_1-06 LSECS_PETRI_MONO CTGACG NNNNNNNN<br> LIMMS61_63_B3_MONO_11_D3 BC_MONO_D3 GAGTGA NNNNNNNN<br> LIMMS62_79_B9_CO_10_D14 BC_CO_D14 GTATAC NNNNNNNN<br> LIMMS63_92_B13_CO_11_D3 BC_CO_D3 TCGAGC NNNNNNNN<br> LIMMS64_09_P2_10_D14 PETRI_MONO ACATGA NNNNNNNN<br> LIMMS65_21_P3_10_D14 PETRI_MONO ATCATA NNNNNNNN<br> LIMMS66_33_P3_11_D14 PETRI_MONO CACGTG NNNNNNNN<br> LIMMS67_45_B1_MONO_10_D14 BC_MONO_D14 CGATGA NNNNNNNN<br> LIMMS68_57_B2_MONO_10_D14 BC_MONO_D14 GAGATA NNNNNNNN<br> LIMMS69_69_B1_MONO_11_D14 BC_MONO_D14 GCTCTC NNNNNNNN<br> LIMMS70_81_B2_MONO_11_D14 BC_MONO_D14 GTATGA NNNNNNNN<br> LIMMS71_93_B6_CO_10_D14 BC_CO_D14 TCGATA NNNNNNNN<br> LIMMS72_11_B7_CO_10_D14 BC_CO_D14 AGTAGC NNNNNNNN<br> LIMMS73_23_B8_CO_10_D14 BC_CO_D14 ATCGCA NNNNNNNN<br> LIMMS74_35_B9_CO_11_D3 BC_CO_D3 CACTCT NNNNNNNN<br> LIMMS75_47_B11_CO_11_D14 BC_CO_D14 CTGAGC NNNNNNNN<br> LIMMS76_59_B12_CO_11_D14 BC_CO_D14 GAGCGT NNNNNNNN<br> LIMMS77_71_B14_CO_11_D14 BC_CO_D14 GCTGCA NNNNNNNN<br> LIMMS78_83_B15_CO_11_D14 BC_CO_D14 TATAGC NNNNNNNN<br> LIMMS79_95_iPSC_rep1_4 CONTROL_iPSC TCGCGT NNNNNNNN</p>
Necroptotic response in ALL_CRISPRscreen sequencing data
<p>FASTQ.</p> <p>sgRNA CRISPR screen sequencing data for PID0117 LC.sg.Library.RFP657</p> <table> <tbody> <tr> <td>ID</td> <td>treatment</td> <td>replicate</td> </tr> <tr> <td>6618</td> <td>vehicle</td> <td>1st</td> </tr> <tr> <td>6619</td> <td>SM 15 mg/kg</td> <td>2nd</td> </tr> <tr> <td>6620</td> <td>SM 5 mg/kg</td> <td>3rd</td> </tr> <tr> <td>6621</td> <td>vehicle</td> <td>1st</td> </tr> <tr> <td>6622</td> <td>SM 15 mg/kg</td> <td>2nd</td> </tr> <tr> <td>6623</td> <td>SM 5 mg/kg</td> <td>3rd</td> </tr> <tr> <td>6624</td> <td>vehicle</td> <td>1st</td> </tr> <tr> <td>6625</td> <td>SM 15 mg/kg</td> <td>2nd</td> </tr> <tr> <td>6626</td> <td>SM 5 mg/kg</td> <td>3rd</td> </tr> </tbody> </table> <p> </p>
Nanopore sequencing data
<p>Data generated with nanopore sequencing (MinION, Oxford Nanopore Technologies). Data generated with MinKNOW (raw read reports). Every report represents a different sequencing run. The data are linked to PhD thesis by Michaela Eleni Christodoulaki.</p>
16S Amplicon sequence variants (ASVs) data of NEREA Augmented Observatory
<p>Cleaned raw 16S paired-end sequences were imported into the QIIME2 pipeline v. 2022.2.0. Leftover primers and adapters’ sequences were removed through cutadapt. The amplicon sequence variants (ASV) table, which represent true biological sequences within each sample, was generated using the denoised-paired method including truncation, denoising, dereplication, and chimera filtering of the DADA2 (Divisive Amplicon Denoising Algorithm 2) plugin inside QIIME2. Default parameters were used with the exception of the forward and reverse sequence length (--p-trunc-len-f and --p-trunc-len-r), that were set to 220 and 180, respectively. For taxonomy classification, the V4-V5 region were extracted from the pre-formatted reference sequences and taxonomy file build on the SILVA 138 99% OTUS database and the vsearch v. 2.6.2 global alignment implemented in QIIME2 was used.</p>
18S Amplicon sequence variants (ASVs) data of NEREA Augmented Observatory
<p>The quality of the Illumina paired-end V9-18S raw reads (FASTQ format, 2 X 150 PE) was checked using vsearch (vsearch --fastq_stats), then pre-processed with cutadapt and vsearch to remove primer sequences, trim low quality bases and unify mixed orientation reads produced in the ligation-based library preparation. Processed reads were then used to generate amplicon sequence variants (ASVs) using the DADA2 R library; the pipeline was adapted from the one described on the program website (https://benjjneb.github.io/dada2/tutorial.html); no further quality filtering was implemented at this stage, except for discarding all reads with ambiguities (parameter maxN=0 of function filterAndTrim). Filtered F and R reads were used to train the error model and then denoised by applying the trained error model to generate ASVs. Finally, F and R reads were merged and checked for chimeras; up to 9 mismatches were allowed for read merging (parameter maxMismatch=9 of function mergePairs). ASVs were then classified with BLAST against the PR2 v5.01 reference database, integrated with 1293 sequences from Gulf of Naples protist strains and fungi environmental sequences. Highest bit score matches with the best taxonomic resolution were then selected among the returned results.</p>
Data for "Mental Programming of Spatial Sequences in Working Memory in Macaque Frontal Cortex"
<p>36 recording sessions of monkey O.</p> <p>Each file contains three variables. TrialInfo: trial information about the targets, responses, event timing etc. Spk_channel: recording channel id. Spk: spike time for each channel.</p> <p>Access will be fulfilled by sending email of reasonable request to lead corresponding author (Liping Wang).</p> <p> </p>
Comparative analysis of 43 distinct RNA modifications by nanopore tRNA sequencing - RNA004 data
<p>This project focuses on nanopore direct tRNA sequencing method refinement and benchmarking of RNA modification detection between Oxford Nanopore's previous direct RNA sequencing chemistry, RNA002, and the RNA004 chemistry released in November 2023. The majority of the data was collected during a December 2023 tRNA sequencing workshop sponsored by the Hesselberth lab at the University of Colorado, in which participants prepared matched sequencing libraries using both chemistries from tRNA isolated from six different species.</p> <p>Fastq data has been submitted to the SRA (GSE272876); here we are hosting the raw POD5 data from all libraries prepared with RNA004 chemistry as well as select libraries prepared with the deprecated RNA002 chemistry for signal reanalysis.</p>
Data repository for "Flexible Control of Sequence Working Memory in Macaque Frontal Cortex"
<p>Access will be fulfilled by sending email of reasonable request to lead corresponding author (Liping Wang).</p>
FIGURE 1 in Description of nymphs and female subimago of Sparsorythus multilabeculatus Sroka & Soldán, 2008 (Ephemeroptera: Tricorythidae) associated with male imago based on DNA sequence data
FIGURE 1. Map showing Nakhon Ratchasima Province, Thailand; black spot = sampling locality.
FIGURE 13 in Description of nymphs and female subimago of Sparsorythus multilabeculatus Sroka & Soldán, 2008 (Ephemeroptera: Tricorythidae) associated with male imago based on DNA sequence data
FIGURE 13. Sparsorythus multilabeculatus, female subimago forewing. Scale bar: 1 mm.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.