Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

102

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

102 results for “core genes”

Learn how ShareScore rates datasets ↗
zenodo40/100

CORE: Gene Expression-Cancer Knowledge Base

<p>This repository contains the Gene Expression-Cancer&nbsp;Knowledge Base&nbsp;generated by the CORE system.<br> The &#39;schema.owl&#39; file contains the KB schema, whereas the &#39;data.ttl&#39; file contains the actual data.</p>

opencc-by-4.0Jan 2023View details →
zenodo36/100

TagSeq for gene expression in non-model plants: a pilot study at the Santa Rita Experimental Range NEON core site

<p>TagSeq analysis scripts and assembled transcriptomes for four vascular plant species from the Santa Rita Experimental Range, AZ. Transcriptomes for each species were sequenced and assembled as described below. Additional details available in the associated manuscript: MS LINK. Raw reads for each available at NCBI BioProject #PRJNA599443.</p> <p>&nbsp;</p> <p><strong>Taxon selection and sampling&nbsp;</strong></p> <p>This study focused on four commonly-occurring species at the Santa Rita Experimental Range Long Term Research and Core NEON site (SRER). These include the native species <em>Tidestromia</em> <em>lanuginosa</em> (Nutt.) Standl. (Amaranthaceae; &lsquo;woolly tidestromia&rsquo;), <em>Parkinsonia</em> <em>florida</em> (Benth. ex A. Gray) S. Watson. (Fabaceae; &lsquo;blue palo verde&rsquo;), and <em>Bouteloua</em> <em>aristidoides</em> (Kunth) Griseb. (Poaceae; &lsquo;needle grama&rsquo;), as well as the introduced species <em>Eragrostis</em> <em>lehmanniana</em> Nees (Poaceae; &lsquo;Lehmann lovegrass&rsquo;; native to southern Africa). All species were identified using a combination of the historical flora of the Santa Rita Experimental Range (Medina, 2003), the Arizona Flora (Kearney et al., 1960), and the Flora of North America (Flora of North America Editorial Committee, eds. 1993). Vouchers were deposited in the University of Arizona herbarium (ARIZ). Tissue from mature plants was collected from an apparently healthy individual representing each target species during the 2017 growing season. An entire stem was sampled for <em>B. aristidoides</em> (with flowers and fruits) and <em>E. lehmanniana</em> (without flowers or fruits). Leaves and leaflets only were sampled for <em>P. florida</em> and <em>T. lanuginosa</em>.</p> <p>&nbsp;</p> <p><strong>RNA extraction and RNA-seq</strong></p> <p>Total RNA was extracted from tissue using the Spectrum Plant Total RNA Kit (Sigma-Aldrich Co., St. Louis, MO, USA) following Protocol A. RNA was used to prepare cDNA using Nugen&rsquo;s Ovation RNA-Seq System via single primer isothermal amplification (Catalogue # 7102-A01) and automated on the Apollo 324 liquid handler (Wafergen). cDNA was quantified on the Nanodrop (Thermo Fisher Scientific) and was sheared to approximately 300 bp fragments using the Covaris M220 ultrasonicator. Libraries were generated using Kapa Biosystem&rsquo;s library preparation kit (KK8201). Fragments were end repaired and A-tailed, and individual indexes and adapters (Bioo, catalogue #520999) were ligated on each separate sample. The adapter ligated molecules were cleaned using AMPure beads (Agencourt Bioscience/Beckman Coulter, A63883), and amplified with Kapa&rsquo;s HIFI enzyme (KK2502). Each library was then analyzed for fragment size on an Agilent&rsquo;s Tapestation, and quantified by qPCR (KAPA Library Quantification Kit, KK4835) on Thermo Fisher Scientific&rsquo;s Quantstudio 5 before multiplex pooling (13-16 samples per lane) and paired-end sequencing at 2x150 bp on the Illumina NextSeq500 platform at Arizona State University&rsquo;s CLAS Genomics Core facility. Raw read quality was assessed using fastQC (Andrews, 2010).</p> <p>&nbsp;</p> <p><strong><em>De novo</em> transcriptome assembly</strong></p> <p>Raw sequence reads were processed using the SnoWhite pipeline (Barker et al., 2010a; Dlugosch et al., 2013), which included trimming adapter sequences and bases with a quality score below 20 from the 3&#39; ends of all reads, removing reads that are entirely primer and/or adapter fragments using TagDust (Lassmann et al., 2009), and removing polyA/T tails with SeqClean (https://sourceforge.net/projects/seqclean/). All transcriptomes were assembled with SOAPdenovo-Trans v1.03 (Xie et al., 2014) using a k-mer of 57. Assembled sequences for each species are in the files ending &quot;.scafSeq&quot;.</p> <p>&nbsp;</p> <p><strong>Protein Translations</strong></p> <p>We used TransPipe (Barker et al., 2010) to identify plant proteins within the assembled transcripts for each reference transcriptome and provide protein and in-frame nucleic acid sequences for each species. The reading frame and protein translation for each sequence was identified by comparison to protein sequences from 25 sequenced and annotated plant genomes from Phytozome (Goodstein et al., 2012). Using BLASTX (Wheeler et al., 2008), best hit proteins were paired with each gene at a minimum cutoff of 30% sequence similarity over at least 150 sites. Genes that did not have a best hit protein at this level were removed. To determine the reading frame and generate estimated amino acid sequences, each gene was aligned against its best hit protein by Genewise 2.2.2 (Birney et al., 2004). Based on the highest scoring Genewise DNA-protein alignments, stop and &#39;N&#39; containing codons were removed to produce estimated amino acid sequences for each gene. Output included paired DNA and protein sequences with the DNA sequence reading frame corresponding to each protein sequence. Nucleic acid sequence files end in &ldquo;.fna&rdquo;, whereas amino acid sequence files end in &ldquo;.faa&rdquo;. Numbers of sequences in each of these files correspond to the position of the sequence in the associated assembly file.</p> <p>&nbsp;</p> <p><strong>Custom scripts</strong></p> <p>&ldquo;removePCRdups57.pl&rdquo; is a Perl script that takes an input FASTQ file and removes exact duplicates identified over a supplied length at the beginning (3&rsquo; end) of the read.&nbsp;</p> <p>Run: perl removePCRdups57.pl &lt;inputFASTQ&gt; &lt;length&gt;</p> <p>&nbsp;</p> <p>&ldquo;create_GTF.pl&rdquo; is a Perl script that takes an input FASTA file and creates a GTF file suitable for input into HtSeq-count v.0.5.4 (Anders et al., 2015).</p> <p>Run: perl create_GTF.pl &lt;inputFASTA&gt;</p> <p>&nbsp;</p> <p>&ldquo;combine_HtSeq.pl&rdquo; is a Perl script that takes a set of htseq output files and makes a tab delim table of counts with header of sample names and first col of row names. The input file list file should be a text file with lists of Htseq files to combine on each line, where lines are tab delimited of the following form:</p> <p>&nbsp;&nbsp;&nbsp;&lt;NameForOutputFile&gt; &lt;firstHtseqFile&gt; &lt;NextHtseqFile&gt; &lt;...etc...&gt;</p> <p>Run: perl combine_HtSeq.pl &lt;inputFileList&gt;</p> <p>&nbsp;</p>

opencc-by-4.0Apr 2020View details →
zenodo36/100

De novo assembly of 20 chicken genomes reveals the undetectable phenomenon for thousands of core genes on micro-chromosomes and sub-telomeric regions

<p>The gene numbers and evolutionary rates of birds were assumed to be much lower than those&nbsp;of mammals, which is&nbsp;in sharp contrast to the huge species number and morphological diversity of birds. It is therefore&nbsp;necessary to construct a complete avian genome and analyze its evolution. We constructed a chicken pan-genome from 20 <em>de novo</em>&nbsp;assembled&nbsp;genomes&nbsp;with high sequencing depth, and&nbsp;identified 1,335 protein-coding genes and 3,011 long noncoding RNAs not found in GRCg6a. The majority of these novel genes were detected across most individuals of the examined transcriptomes but were seldomly&nbsp;measured in each of the DNA sequencing data regardless of Illumina or PacBio technology. Furthermore, different from previous pan-genome models, most of these novel genes were overrepresented on chromosomal sub-telomeric regions&nbsp;and micro-chromosomes, surrounded by&nbsp;extremely high proportions of tandem repeats, which&nbsp;strongly blocks&nbsp;DNA sequencing. These hidden genes were proved to be shared by all chicken genomes, included many housekeeping genes, and enriched in immune pathways. Comparative genomics revealed the novel genes had three-fold elevated substitution rates than known ones, updating the knowledge about&nbsp;evolutionary rates in&nbsp;birds. Our study provides a framework for constructing a better chicken genome, which will contribute towards the understanding of avian evolution and improvement of poultry breeding.</p>

opencc-by-4.0Jan 2022View details →
zenodo36/100

Core-shell structured chitosan-polyethylenimine nanoparticles for gene delivery: Improved stability, cellular uptake, and transfection efficiency

<p>Gene therapy has emerged as a promising treatment option for various acquired and inherited diseases. The delivery of nucleic acids relies on so-called vectors that condense and encapsulate their cargo, generating stable nano-sized particles. Especially non-viral gene delivery systems are of increasing interest. However, accomplishing therapeutic levels of transgene expression and limited tolerability of these systems remain a challenge. Therefore, we investigate in the present study the improvement of nucleic acid delivery using depolymerized chitosan &ndash; polyethylenimine DNA core complexes (dCS-PEI/DNA). These core complexes are further entrapped into a variety of dCS-based shells, functionalized with poly(ethylene glycol) (PEG) spacers conjugated to ionic moieties (amino or carboxylate groups) and cell penetrating peptides. This modular approach allowed to evaluate the effect of the shell functional components on the physico-chemical particle characteristics and biological effects <em>in vitro</em>. The optimized ternary complex combines a core-dCS-LPEI/DNA complex with a shell consisting of dCS-PEG-COOH, which resulted in improved encapsulation of nucleic acid, accelerated cellular uptake, enhanced transfection efficiency, and superior transfection potency in human hepatoma HuH-7 cells and mouse primary hepatocytes. Effects on transgene expression are confirmed <em>in vivo</em> in wild-type mice following retrograde intrabiliary infusion. After administration to mice of only 100 ng complexed nanovector DNA, ternary complexes induce a high reporter gene signal for three days. We conclude that ternary core-shell structured particles comprising functionalized chitosan are a promising gene delivery technology for both <em>in vitro</em> as well as <em>in vivo </em>applications. The modular design will facilitate the development of chemically modified derivatives.</p>

opencc-by-4.0Jul 2022View details →
zenodo36/100

Speos: An ensemble graph representation framework to predict core genes for complex diseases (Datasets)

<p>the &quot;data.tar.gz&quot; tarball contains the unprocessed or minimally processed data used by Speos. If you intend to use the framework or want to inspect the data, download this part of the dataset.</p> <p>The &quot;final_datasets.tar.gz&quot; tarball contains tsv-formatted, processed data matrices which are directly used as input features for the ensemble models.&nbsp;</p> <p>There are two tsv-files&nbsp;per disease, one labeled &quot;normal&quot;, which contains the p input features alongside the gene identifiers and a column which indicates if the gene is labeled as Mendelian or not, and another file labeled &quot;with_n2v_vectors&quot;, which also contains the 100-dimensional vectors produced by Node2Vec so the N2V+MLP method can be reproduced with the exact same parameters.&nbsp;</p> <p>All files contain a header row which describes the column and no index column.</p> <p>The &quot;model_parameters.tar.gz&quot; tarball contains the model parameters for all ensemble models used to produce the candidate genes.</p>

opencc-by-4.0Jan 2023View details →
dryad32/100

Genome reduction is associated with bacterial pathogenicity across different scales of temporal and ecological divergence - between species core gene alignments

<p><span>Emerging bacterial pathogens threaten global health and food security, and so it is important to ask whether these transitions to pathogenicity have any common features. We present a systematic study of the claim that pathogenicity is associated with genome reduction and gene loss. We compare broad-scale patterns across all bacteria, with detailed analyses of <i>Streptococcus suis</i>, an emerging zoonotic pathogen of pigs, which has undergone multiple transitions between disease and carriage forms. We find that pathogenicity is consistently associated with reduced genome size across three scales of divergence (between species within genera, and between and within genetic clusters of <i>S. suis</i>). While genome reduction is also found in mutualist and commensal bacterial endosymbionts, genome reduction in pathogens cannot be solely attributed to the features of their ecology that they share with these species, i.e. host restriction or intracellularity. Moreover, other typical correlates of genome reduction in endosymbionts (reduced metabolic capacity, reduced GC content, and the transient expansion of non-functional elements) are not consistently observed in pathogens. Together, our results indicate that genome reduction is a predictive marker of pathogenicity in bacteria.</span></p>

opencc-zeroNov 2020View details →
dryad32/100

Data from: Core genes evolve rapidly in the long-term evolution experiment with Escherichia coli

Bacteria can evolve rapidly under positive selection owing to their vast numbers, allowing their genes to diversify by adapting to different environments. We asked whether the same genes that evolve rapidly in the long-term evolution experiment with Escherichia coli (LTEE) have also diversified extensively in nature. To make this comparison, we identified ~2000 core genes shared among 60 E. coli strains. During the LTEE, core genes accumulated significantly more nonsynonymous mutations than flexible (i.e., noncore) genes. Furthermore, core genes under positive selection in the LTEE are more conserved in nature than the average core gene. In some cases, adaptive mutations appear to modify protein functions, rather than merely knocking them out. The LTEE conditions are novel for E. coli, at least in relation to its evolutionary history in nature. The constancy and simplicity of the environment likely favor the complete loss of some unused functions and the fine-tuning of others.

opencc-zeroDec 2016View details →
dryad32/100

Data from: Core genes evolve rapidly in the long-term evolution experiment with Escherichia coli

Open the record for dataset details and reuse information.

publicApr 2017View details →
dryad32/100

Genome reduction is associated with bacterial pathogenicity across different scales of temporal and ecological divergence - between species core gene alignments

Open the record for dataset details and reuse information.

publicNov 2020View details →
dryad28/100

Data from: The core planar cell polarity gene, Vangl2, directs adult corneal epithelial cell alignment and migration

This study shows that the core planar cell polarity (PCP) genes direct the aligned cell migration in the adult corneal epithelium, a stratified squamous epithelium on the outer surface of the vertebrate eye. Expression of multiple core PCP genes was demonstrated in the adult corneal epithelium. PCP components were manipulated genetically and pharmacologically in human and mouse corneal epithelial cells in vivo and in vitro. Knockdown of VANGL2 reduced the directional component of migration of human corneal epithelial (HCE) cells without affecting speed. It was shown that signalling through PCP mediators, dishevelled, dishevelled-associated activator of morphogenesis and Rho-associated protein kinase directs the alignment of HCE cells by affecting cytoskeletal reorganization. Cells in which VANGL2 was disrupted tended to misalign on grooved surfaces and migrate across, rather than parallel to the grooves. Adult corneal epithelial cells in which Vangl2 had been conditionally deleted showed a reduced rate of wound-healing migration. Conditional deletion of Vangl2 in the mouse corneal epithelium ablated the normal highly stereotyped patterns of centripetal cell migration in vivo from the periphery (limbus) to the centre of the cornea. Corneal opacity owing to chronic wounding is a major cause of degenerative blindness across the world, and this study shows that Vangl2 activity is required for directional corneal epithelial migration.

opencc-zeroDec 2015View details →
dryad28/100

Data from: Plasticity of promoter-core sequences allows bacteria to compensate for the loss of a key global regulatory gene

Transcription regulatory networks (TRNs) are of central importance for both short-term phenotypic adaptation in response to environmental fluctuations and long-term evolutionary adaptation, with global regulatory genes often being targets of natural selection in laboratory experiments. Here, we combined evolution experiments, whole-genome resequencing, and molecular genetics to investigate the driving forces, genetic constraints, and molecular mechanisms that dictate how bacteria can cope with a drastic perturbation of their TRNs. The crp gene, encoding a major global regulator in Escherichia coli, was deleted in four different genetic backgrounds, all derived from the Long-Term Evolution Experiment (LTEE) but with different TRN architectures. We confirmed that crp deletion had a more deleterious effect on growth rate in the LTEE-adapted genotypes; and we showed that the ptsG gene, which encodes the major glucose-PTS transporter, gained CRP dependence over time in the LTEE. We then further evolved the four crp-deleted genotypes in glucose minimal medium, and we found that they all quickly recovered from their growth defects by increasing glucose uptake. We showed that this recovery was specific to the selective environment and consistently relied on mutations in the cis regulatory region of ptsG, regardless of the initial genotype. These mutations affected the interplay of transcription factors acting at the promoters, changed the intrinsic properties of the existing promoters, or produced new transcription initiation sites. Therefore, the plasticity of even a single promoter region can compensate by three different mechanisms for the loss of a key regulatory hub in the E. coli TRN.

opencc-zeroDec 2018View details →
ClinicalTrials.gov28/100

Gene Expression in Tumor Tissue From Women Undergoing Surgery for Breast Cancer or Core Biopsy of the Breast

ClinicalTrials.gov study NCT00898937. IPD Sharing: Not stated. Countries: 1. Publications: 0.

restrictedIPD-UNDECIDEDFeb 2026View details →
dryad28/100

Data from: The core planar cell polarity gene, Vangl2, directs adult corneal epithelial cell alignment and migration

Open the record for dataset details and reuse information.

publicOct 2016View details →
dryad28/100

Data from: Sequencing-based gene network analysis provides a core set of gene resource for understanding thermal adaptation in Zhikong scallop Chlamys farreri

Open the record for dataset details and reuse information.

publicSep 2013View details →
dryad28/100

Data from: Plasticity of promoter-core sequences allows bacteria to compensate for the loss of a key global regulatory gene

Open the record for dataset details and reuse information.

publicMar 2019View details →
dryad28/100

Data from: Interacting networks of resistance, virulence and core machinery genes identified by genome-wide epistasis analysis

Open the record for dataset details and reuse information.

publicNov 2017View details →
geo24/100

Transcriptome profiles reveals potential core genes in the black skin of Xichuan Black-bone chickens(Gallus gallus)

GEO Series GSE128028. Gallus gallus. 2 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenJun 2019View details →
geo24/100

Integrated transcriptomic analysis of Trichosporon asahii uncovers the core genes and pathways of fluconazole resistance

GEO Series GSE106454. Trichosporon asahii var. asahii CBS 2479; Trichosporon asahii var. asahii CBS 8904. 10 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenFeb 2018View details →
geo24/100

Mixed-species RNAseq analysis of human lymphoma cell adhesion to mouse stromal cells identifies a core gene set that is also differentially expressed in the lymph node microenvironment of MCL and CLL

GEO Series GSE99501. Mus musculus; Homo sapiens. 16 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenJun 2018View details →
geo24/100

CDK12 catalytic activity is rate-limiting for RNAPII processivity on core DNA replication genes and G1/S progression (nuclear RNA)

GEO Series GSE120071. Homo sapiens. 6 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenJul 2019View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record