Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
25,372
datasets available to search
ShareScore release 0.7.1
Dataset results
25,372 results for “Transcriptomics”
the transcriptome GTFs, FASTA and SQANTI reports for short-read assembled isoforms, long-read assembled isoforms and our assembled isoforms
Open the record for dataset details and reuse information.
Comparative transcriptomics and phylostratigraphy of Argentine ant odorant receptors
<p>Nestmate recognition in ants is regulated through the detection of cuticular hydrocarbons by odorant receptors (ORs) in the antennae. These ORs are crucial for maintaining colony cohesion that allows invasive ant species to dominate colonized environments. In the invasive Argentine ant, <em>Linepithema humile</em>, ORs regulating nestmate recognition are thought to be present in a clade of nine exon odorant receptors, but the identity of the specific genes remains unknown. We sought to narrow down the list of candidate genes using transcriptomics and phylostratigraphy. Comparative transcriptomic analyses were conducted on the antennae, head, thorax, and legs of Argentine ant workers. We have identified a set of twenty-one nine-exon odorant receptors enriched in the antennae compared to the other tissues, allowing for downstream verification of whether they can detect Argentine ant cuticular hydrocarbons. Further investigation of these ORs could allow us to further understand the mechanisms underlying nestmate recognition and colony cohesion in ants.</p>
Transcriptome data of the analysis of two isolates of the tomato pathogen Cladosporium fulvum during host interaction
<p>This dataset contains sequences of assembled transcripts from isolates Race 5 and Race 4 of the tomato pathogen Cladosporium fulvum during interaction with its host.</p> <p><strong>transcripts:</strong> Assembled transcripts in FASTA and GTF fomats. The GTF files have coordinates of the transcripts in the reference genomes of isolates Race 5 (GCA_020509005.2) and Race 4 (GCA_035196885.1). The other FASTA files include the predicted open reading frames (ORFs) in the transcripts. The nucleotide coding sequence and translated amino acid sequences of the ORFs are in separated FASTA files. In the file names isolate Race 5 is indicated with '*R5*', and isolate Race 4 is indicated with '*R4*'. Description of the files is shown below:</p> <ul> <li><code>representatives_R4_diff_introns.fasta</code>: full-length transcript sequences from isolate Race 4.</li> <li><code>representatives_R5_diff_introns.fasta</code>: full-length transcript sequences from isolate Race 5.</li> <li><code>representatives_R4_diff_introns.gtf</code>: coordinates of the transcripts from isolate Race 4 in the genome of Race 4.</li> <li><code>representatives_R5_diff_introns.gtf</code>: coordinates of the transcripts from isolate Race 5 in the genome of Race 5.</li> <li><code>representatives_orfs_aa_R4_diff_introns.fasta:</code> predicted protein sequences encoded in the transcripts from isolate Race 4.</li> <li><code>representatives_orfs_aa_R5_diff_introns.fasta</code>: predicted protein sequences encoded in the transcripts from isolate Race 5.</li> <li><code>representatives_orfs_cds_R4_diff_introns.fasta</code>: predicted coding sequences in the transcripts from isolate Race 4.</li> <li><code>representatives_orfs_cds_R5_diff_introns.fasta</code>: predicted coding sequences in the transcripts from isolate Race 5.</li> </ul> <p><strong>expression:</strong> Contains tab-separated files with the expression values (transcripts per million - TPM) of the transcripts from isolates Race 5 and Race 4 at specific time points (2, 4, 6, 8, 10, 12, and 14 dpi) during interaction with tomato. TPM values were estimated with the alignment-free method Salmon.</p>
Technology comparison (image-based spatial transcriptomics)- annotated datasets
<p>This repository contains all the AnnData datasets, regionally annotated, used in the comparison of image-based spatial transcriptomics technologies (Marco Salas et al. 2024)</p>
Datasets for the manuscript: "Metabolic disruption of zebrafish (Danio rerio) embryos by bisphenol A. An integrated metabolomic and transcriptomic approach"
<h1>Metabolomics datasets for the manuscript: Metabolic disruption of zebrafish (Danio rerio) embryos by bisphenol A. An integrated metabolomic and transcriptomic approach</h1> <h2><em>Instrumental conditions</em></h2> <p>LC-MS analyses were carried out using an Agilent Infinity 1200 series LC system coupled with an orthogonal G1385-44300 interface (Agilent Technologies, Waldbronn, Germany) to a 6220 oa-TOF LC/MS mass spectrometer (Agilent Technologies). LC control and separation data acquisition were performed using ChemStation software (Agilent Technologies) that was running in combination with the MassHunter workstation software (Agilent Technologies) for control and data acquisition of the TOF mass spectrometer. For the chromatographic separations, an HILIC TSK Gel Amide-80 column (250 mm length, 2.1 mm inner diameter and 5 μm particle size, Tosoh Bioscience, Tokyo, Japan) was used at 25 °C with gradient elution at a flow rate of 0.15 mL·min<sup>−1</sup>. Elution gradient was performed using solvent A (acetonitrile) and solvent B (5 mM of ammonium acetate adjusted to pH 5.5 with acetic acid) as follows: 0–8 min, linear gradient from 25 to 30% B; 8–12 min, from 30 to 60% B; 12–17 min, 60% B; 17–20 min, back linearly from 60% to 25% B; and from 20 to 27 min, 25% B. Solvents were degassed for 15 min by sonication before use. Sample injection was performed with an autosampler at 4 °C, and the injection volume was 5 μL. All samples (six replicates per treatment: control, 4.4 μM BPA, 8.8 μM BPA and 17.5 μM BPA) were randomly injected. Several blank samples and calibration standards were also randomly injected to further assess the stability of the instrument among runs.</p> <p>The TOF mass spectrometer operated both in positive and negative mode using the following parameters: capillary voltage 4000 V, drying gas temperature 350 °C, drying gas flow rate 8 L·min<sup>−1</sup>, nebulizer gas 32 psi, fragmentor voltage 150 V, skimmer voltage 65 V and OCT 1 RF Vpp voltage 300 V. Data were collected in profile mode at 1 spectrum/s (approximately 10 000 transients/spectrum) with an <em>m/z</em> range of 85–1000 working in the extended dynamic range mode (2 GHz) with the mass range set to standard.</p> <h2><em>List of files</em></h2> <h3>Negative ionization</h3> <ul> <li>Control x 12 samples - 6 x 2 replicates</li> <li>BPA 1 ppm x 12 samples - 6 x 2 replicates</li> <li>BPA 2 ppm x 12 samples - 6 x 2 replicates</li> <li>BPA 4 ppm x 12 samples - 6 x 2 replicates</li> </ul> <h3>Positive ionization</h3> <ul> <li>Control x 12 samples - 6 x 2 replicates</li> <li>BPA 1 ppm x 12 samples - 6 x 2 replicates</li> <li>BPA 2 ppm x 12 samples - 6 x 2 replicates</li> <li>BPA 4 ppm x 12 samples - 6 x 2 replicates</li> </ul>
Transcriptome analysis and functional study of phospholipase A2 in Galleria mellonella larvae lipid metabolism in response to envenomation by an ectoparasitoid, Iseropus kuwanae
<p>The file is the raw data of "Transcriptome analysis and functional study of phospholipase A2 in <em>Galleria mellonella</em> larvae lipid metabolism in response to envenomation by an ectoparasitoid, <em>Iseropus kuwanae</em>".</p>
An evolutionary framework of Acanthaceae based on transcriptomes and genome skims
<p>Acanthaceae is a family of tropical flowering plants with approximately 4000 species. Despite remarkable variation in morphological traits, research on patterns of character evolution has been limited by uncertain relationships among some of the major lineages. We sampled from these major lineages to estimate a phylogenomic framework using a combination of newly sequenced shotgun genome skims plus new and publicly available transcriptomes. We used OrthoFinder2 to infer a species tree with strong branch support. Except for the placement of Crabbea, our results corroborate the most recent chloroplast and nrITS sequence-based topology. Of 587 single copy loci, 10 were recovered for all 16 species; a RAxML tree estimated from these 10 loci resulted in the same topology as other datasets assembled in this study, with the exception of relationships among three sampled species of Barleria; however, branch support was lower compared to the tree reconstructed using more data. ABBA-BABA tests were conducted to investigate patterns of introgression involving Crabbea; few nucleotides supported alternative topologies. SplitsTree networks of the 587 loci and 6,136 trees revealed conflict among the branches leading to Andrographideae, Whitfieldieae, and Neuracanthus. A principal components analysis in treespace found no distinct clusters of trees. Our results strongly corroborate the previously published chloroplast and nr-ITS-based phylogeny of Acanthaceae with increased resolution among Barlerieae, Andrographideae, Whitfieldieae, and Neuracanthus. We propose that the tree presented here is the best estimate to date of Acanthaceae phylogeny. This advance in our knowledge of relationships will allow us to investigate character evolution and other phenomena within this diverse group of plants.</p>
Transcriptome profiling associated with CARD11 overexpres-sion in Colorectal Cancer implicates a potential role for Tumour Immune Microenvironment and Cancer pathways modulation via NF-κB
<p>tables for CARD11</p>
Deep Clustering Representation for Spatially Resolved Transcriptomics Data via Multi-view Variational Graph Auto-Encoders with Consensus Clustering
Open the record for dataset details and reuse information.
SQUID: Transcriptomic Structural Variation Detection from RNA-seq -- simulation data part 3
<p>Simulation data part 3 for SQUID software.</p>
SQUID: Transcriptomic Structural Variation Detection from RNA-seq -- simulation data part 2
<p>Simulation data part 2 for SQUID software.</p>
SQUID: Transcriptomic Structural Variation Detection from RNA-seq -- simulation data part 1
<p>Simulation data part 1 for SQUID software.</p>
Solving for X: evidence for sex-specific autism biomarkers across multiple transcriptomic studies: Figure Data
<p>Figure Data for DOI:10.1101/309518.</p>
Transcriptome profiling of derived-hepatocyte progenitors from human iPSCs with nanoCAGE - part 1 - sequencing data (FASTQ files)
<p>This repository contains raw sequencing data (FASTQ files) produced from Illumina MiSeq run IDs "170630_M00528_0292_000000000-B9JY8" (aka "NC_LIMMS") and "180221_M00528_0334_000000000-B6PJM" (aka "NC_LIMMS2") . Sequencing libraries were prepared following the latest version of the nanoCAGE protocol (Poulain et al., Methods Mol Biol. 2017;1543:57-109. doi: 10.1007/978-1-4939-6716-2_4). They respectively contain a mix of 24 ("NC_LIMMS") and 18 ("NC_LIMMS2") samples tagged by specific barcode sequences at the 5'-ends (see tables below). The tagmentation step included in the protocol was performed using an equimolar mix of 12 Nextera XT N-series index primers (N701 to N712), therefore "NNNNNNNN" was indicated as index sequence on the Illumina Sample Sheet for the demultiplexing (see tables below). Libraries were sequenced paired-end on Illumina MiSeq system with the MiSeq Reagent Kit v3 (150 cycles: 58 cycles used for READ1, 8 cycles used for the Index, and 84 cycles used for READ2). Genomic alignments (BED files) of paired-end reads on human genome assemblies hg19 and hg38 using the MOIRAI pipeline (Hasegawa et al. BMC Bioinformatics 2014 May 16;15:144. doi: 10.1186/1471-2105-15-144) were deposited at Zenodo under the following Digital Object Identifier: 10.5281/zenodo.1017276.</p> <p> </p> <p><em><strong>"170630_M00528_0292_000000000-B9JY8" ("NC_LIMMS") :</strong></em></p> <p><strong>ID Sample_name Barcode_number Barcode_sequence Index_sequence</strong></p> <p>1 iPSC_control_rep1 4 ACAGAT NNNNNNNN</p> <p>2 iPSC_control_rep2 24 ATCGTG NNNNNNNN</p> <p>3 iPSC_control_rep3 31 CACGAT NNNNNNNN</p> <p>4 S3P1_OK_rep1 36 CACTGA NNNNNNNN</p> <p>5 S3P1_OK_rep2 46 CTGACG NNNNNNNN</p> <p>6 S3P1_OK_rep3 63 GAGTGA NNNNNNNN</p> <p>7 S4P1_OK_rep1 79 GTATAC NNNNNNNN</p> <p>8 S4P1_OK_rep2 92 TCGAGC NNNNNNNN</p> <p>9 S4P1_OK_rep3 9 ACATGA NNNNNNNN</p> <p>10 S4P2_OK_rep1 21 ATCATA NNNNNNNN</p> <p>11 S4P2_OK_rep2 33 CACGTG NNNNNNNN</p> <p>12 S4P2_OK_rep3 45 CGATGA NNNNNNNN</p> <p>13 S1P1_rep1 57 GAGATA NNNNNNNN</p> <p>14 S1P1_rep2 69 GCTCTC NNNNNNNN</p> <p>15 S1P1_rep3 81 GTATGA NNNNNNNN</p> <p>16 S3P1_FAILED_rep1 93 TCGATA NNNNNNNN</p> <p>17 S3P1_FAILED_rep2 11 AGTAGC NNNNNNNN</p> <p>18 S3P1_FAILED_rep3 23 ATCGCA NNNNNNNN</p> <p>19 S4P1_FAILED_rep1 35 CACTCT NNNNNNNN</p> <p>20 S4P1_FAILED_rep2 47 CTGAGC NNNNNNNN</p> <p>21 S4P1_FAILED_rep3 59 GAGCGT NNNNNNNN</p> <p>22 S4P2_FAILED_rep1 71 GCTGCA NNNNNNNN</p> <p>23 S4P2_FAILED_rep2 83 TATAGC NNNNNNNN</p> <p>24 S4P2_FAILED_rep3 95 TCGCGT NNNNNNNN</p> <p> </p> <p><em><strong>"180221_M00528_0334_000000000-B6PJM" ("NC_LIMMS2"):</strong></em></p> <p><strong>ID Sample_name Barcode_number Barcode_sequence Index_sequence</strong></p> <p>25 PETRI_rep1 04 ACAGAT NNNNNNNN</p> <p>26 PETRI_rep2 24 ATCGTG NNNNNNNN</p> <p>27 PETRI_rep3 31 CACGAT NNNNNNNN</p> <p>28 BIOCHIP_E_rep1 6 CACTGA NNNNNNNN</p> <p>29 BIOCHIP_M_rep1 46 CTGACG NNNNNNNN</p> <p>30 BIOCHIP_S_rep1 63 GAGTGA NNNNNNNN</p> <p>31 BIOCHIP_E_rep2 79 GTATAC NNNNNNNN</p> <p>32 BIOCHIP_M_rep2 92 TCGAGC NNNNNNNN</p> <p>33 BIOCHIP_S_rep2 09 ACATGA NNNNNNNN</p> <p>34 BIOCHIP_E_rep3 21 ATCATA NNNNNNNN</p> <p>35 BIOCHIP_M_rep3 33 CACGTG NNNNNNNN</p> <p>36 BIOCHIP_S_rep3 45 CGATGA NNNNNNNN</p> <p>37 HEPATOCYTES_rep1 57 GAGATA NNNNNNNN</p> <p>38 HEPATOCYTES_rep2 69 GCTCTC NNNNNNNN</p> <p>39 iPSC_control_rep1-2 81 GTATGA NNNNNNNN</p> <p>40 BIOCHIP_E_rep2-2 93 TCGATA NNNNNNNN</p> <p>41 BIOCHIP_M_rep1-2 11 AGTAGC NNNNNNNN</p> <p>42 BIOCHIP_S_rep2-2 23 ATCGCA NNNNNNNN</p>
Contractile activity-specific transcriptome response to acute endurance exercise and training in human skeletal muscle
<p>Supplemental Figure for a study <strong>Contractile activity-specific transcriptome response to acute endurance exercise and training in human skeletal muscleи </strong> <a href="https://www.ncbi.nlm.nih.gov/pubmed/30779632#">Am J Physiol Endocrinol Metab.</a> 2019 Feb 19. <strong> </strong> https://www.ncbi.nlm.nih.gov/pubmed/30779632</p>
Contractile activity-specific transcriptome response to acute endurance exercise and training in human skeletal muscle
<p>Supplemental Tables for a study <strong>Contractile activity-specific transcriptome response to acute endurance exercise and training in human skeletal muscle </strong> <a href="https://www.ncbi.nlm.nih.gov/pubmed/30779632#">Am J Physiol Endocrinol Metab.</a> 2019 Feb 19. <strong> </strong>https://www.ncbi.nlm.nih.gov/pubmed/30779632</p>
Supporting data and code for: Dissecting the transcriptomic basis of phenotypic evolution in the aquatic keystone grazer Daphnia: Part I
<p>Main codes for Dissecting the transcriptomic basis of phenotypic evolution in the aquatic keystone grazer Daphnia.</p>
An evaluation of pool-sequencing transcriptome-based exon capture for population genomics of non-model species.
<p>This archive is associated with the article “An evaluation of pool-sequencing transcriptome-based exon capture for population genomics of non-model species.”. Authors: Emeline Deleury, Thomas Guillemaud, Aurelie Blin & Eric Lombaert.</p> <p>The archive contains :<br> - The sequences of the 5,717 Harmonia axyridis randomly selected CDS (5717-targeted-CDS-sequences.gff3, sequence in FASTA format at the end of the file)<br> - For the subset of 3,161 targeted CDS that have a genomic match over their entire length, the positions of exons on transcripts (3161-targeted-CDS-EXON-POSITIONS.csv)</p>
additional file 1 and 2 for De novo transcriptome sequencing of Serangium japonicum (Coleoptera: Coccinellidae)
<p>file1:some comman statastic results of transcriptome sequences of 6 S.japonicum samples. file2:GO classed different expression genes.</p>
Transcriptome association studies of neuropsychiatric traits in African Americans implicates PRMT7 in schizophrenia
<p>Genome-wide association study of Schizophrenia and Bipolar Disorder in African Americans 2019</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.