Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

665

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

665 results for “exon”

Learn how ShareScore rates datasets ↗
dryad32/100

Data from: Genome-wide exon-capture approach identifies genetic variants of Norway spruce genes associated with susceptibility to Heterobasidion parviporum infection

Root and butt rot caused by members of the Heterobasidion annosum species complex is the most economically important disease of conifer trees in boreal forests. Wood decay in the infected trees dramatically decreases their value and causes considerable losses to forest owners. Trees vary in their susceptibility to Heterobasidion infection, but the genetic determinants underlying the variation in the susceptibility are not well-understood. We performed the identification of Norway spruce genes associated with the resistance to Heterobasidion parviporum infection using genome-wide exon-capture approach. Sixty-four clonal Norway spruce lines were phenotyped, and their responses to H. parviporum inoculation were determined by lesion length measurements. Afterwards, the spruce lines were genotyped by targeted resequencing and identification of genetic variants (SNPs). Genome-wide association analysis identified 10 SNPs located within 8 genes as significantly associated with the larger necrotic lesions in response to H. parviporum inoculation. The genetic variants identified in our analysis are potential marker candidates for future screening programs aiming at the differentiation of disease-susceptible and resistant trees.

opencc-zeroDec 2017View details →
dryad32/100

Data from: Major histocompatibility complex class I evolution in songbirds: universal primers, rapid evolution and base compositional shifts in exon 3

Genes of the Major Histocompatibility Complex (MHC) have become an important marker for the investigation of adaptive genetic variation in vertebrates because of their critical role in pathogen resistance. However, despite significant advances in the last few years the characterization of MHC variation in non-model species still remains a challenging task due to the redundancy and high variation of this gene complex. Here we report the utility of a single pair of primers for the cross-amplification of the third exon of MHC class I genes, which encodes the more polymorphic half of the peptide-binding region (PBR), in oscine passerines (songbirds; Aves: Passeriformes), a group especially challenging for MHC characterization due to the presence of large and complex MHC multigene families. In our survey, although the primers failed to amplify exon 3 from two suboscine passerine birds, they amplified exon 3 of multiple MHC class I genes in all 16 species of oscine songbirds tested, yielding a total of 120 sequences. The 16 songbird species belong to 14 different families, primarily within the Passerida, but also in the Corvida. Using a conservative approach based on the analysis of cloned amplicons (n = 16) from each species, we found between 3 and 10 MHC sequences per individual. Each allele repertoire was highly divergent, with the overall number of polymorphic sites per species ranging from 33 to 108 (out of 264 sites) and the average number of nucleotide differences between alleles ranging from 14.67 to 43.67. Our survey in songbirds allowed us to compare macroevolutionary dynamics of exon 3 between songbirds and non-passerine birds. We found compelling evidence of positive selection acting specifically upon peptide-binding codons across birds, and we estimate the strength of diversifying selection in songbirds to be about twice that in non-passerines. Analysis using comparative methods suggest weaker evidence for a higher GC content in the 3rd codon position of exon 3 in non-passerine birds, a pattern that contrasts with among-clade GC patterns found in other avian studies and may suggests different mutational mechanisms. Our primers represent a useful toolfor the characterization of functional and evolutionarily relevant MHC variation across the hyperdiverse songbirds.

opencc-zeroDec 2012View details →
zenodo32/100

Files From The Exon Analysis Pipeline of Pediatric Solid Tumors

<p>Long-read iso-seq validation of exon regions:</p> <p>Methods of the Long-read Sequencing:</p> <p>Libraries for each tumor were prepared and sequencing on a Pacific Biosciences (PacBio) RSII instrument at JAX. The raw data files were processed according to the PacBio Isoseq3 pipeline which utilizes a number of command line tools provided in PacBio SMRT Tools v10.2 (<a title="https://www.pacb.com/support/software-downloads/" href="https://www.pacb.com/support/software-downloads/">https://www.pacb.com/support/software-downloads/</a>).&nbsp;</p> <p>The pipeline generates non-redundant full-length (FL) transcripts in the following steps for each tumor sample: (i) compute consensus sequences and read quality, (ii) remove primers and adapters, (iii) remove&nbsp;polyA&nbsp;tail and artificial concatemers, (iii) de novo isoform-level clustering, (iv) minimap2 aligns FL transcripts to human reference (GENCODEv40), (v) transcripts were collapsed based on genomic mapping, long-read abundance was estimated, and GTF annotation file generated, (vi) SQANTI3 performed transcript classification and created both a reference corrected transcriptome&nbsp;fasta&nbsp;file and a corrected GTF file for each tumor.&nbsp;</p> <p>&nbsp;</p> <p>Gene target exon PSI value calculation:</p> <p>The &ldquo;generateEvents&rdquo; operation from SUPPA2v2.3 (<a href="https://github.com/comprna/SUPPA">https://github.com/comprna/SUPPA</a>) was used to extract splice sites from each tumor GTF, next &ldquo;psiPerEvent&rdquo; operation was applied to calculate Percent Spliced In (PSI) statistic for each alternative splice site using Kallisto estimated transcript abundance values. Gene target exons of interest were selected from output tables using an absolute splice site genomic coordinate match.</p> <p>Please contact Dr Ching Lau for additional inquiries (<a title="mailto:ching.lau@jax.org" href="mailto:ching.lau@jax.org">ching.lau@jax.org</a>).&nbsp;</p>

opencc-by-4.0Jan 2024View details →
dryad32/100

Comparing ultraconserved elements and exons for phylogenomic analyses of Middle American cichlids: When data agree to disagree

<p>Choosing among types of genomic markers to be used in a phylogenomic study can have a major influence on the cost, design, and results of a study. Yet few attempts have been made to compare categories of next-generation sequence markers limiting our ability to compare the suitability of these different genomic fragment types. Here we explore properties of different genomic markers to find if they vary in the accuracy of component phylogenetic trees and to clarify the causes of conflict obtained from different datasets or inference methods. As a test case, we explore the causes of discordance between phylogenetic hypotheses obtained using a novel dataset of ultraconserved elements (UCEs) and a recently published exon dataset of the cichlid tribe Heroini. Resolving relationships among heroine cichlids has historically been difficult, and the processes of diversification and colonization of Middle America and the Greater Antilles are not yet well understood. Despite differences in informativeness and levels of gene tree discordance between UCEs and exons, the resulting phylogenomic hypotheses generally agree on most relationships. The independent datasets disagreed in areas with low phylogenetic signal that were overwhelmed by noise and non-phylogenetic signals. For UCEs, high levels of incomplete lineage sorting (ILS) seem to be a major cause of noise, whereas, for exons, non-phylogenetic signal may be caused by a reduced number of highly informative loci. This paucity of informative loci in exons might be due to heterogeneous substitution rates that are problematic to model (i.e., computationally restrictive) resulting in systematic errors that UCEs (being less informative individually but more uniform) are less prone to. These results generally demonstrate the robustness of phylogenomic methods to accommodate genomic markers with different biological and phylogenetic properties. However, we identify common and unique pitfalls of different categories of genomic fragments when inferring enigmatic phylogenetic relationships.</p>

opencc-zeroOct 2021View details →
dryad32/100

Flatfish exon-capture

<p>This dataset contains alignments used to study phylogenetic relationships of flatfishes based on exon-capture data. The samples in this dataset represent 89 species: 86 flatfishes and 3 outgroup Carangidae. 57 samples were extracted and assembled from tissues collected for this project, while 39 were sourced from previously assembled data that were prepared as part of the FishLife project. For this study we targeted 4,434 markers developed by Jiang et al. (2019) for capture efficiency in ray-finned fishes (Actinopterygii). Library preparation followed the protocol from Li et al. (2013) involving a double capture method used to increase concentrations of hybridized DNA. Raw reads were assembled into loci using the Assexon bioinformatics pipeline (Yuan et al. 2020). Assembled exons were aligned on amino acids using MAFFT (Katoh et al., 2002), then translated back to codon-based alignment using a custom perl script mafft_aln.pl, and poorly aligned markers were removed. These aligned data represent the unfiltered dataset in the associated study, which was further processed into subsets based on missing data, clocklikeness, nucleotide composition, and evolutionary rate.</p>

opencc-zeroNov 2021View details →
dryad32/100

Exon-based phylogenomics and the relationships of African cichlids: Tackling the challenges of reconstructing phylogenies with repeated rapid radiations

<p>African cichlids (subfamily: Pseudocrenilabrinae) are among the most diverse vertebrates, and their propensity for repeated rapid radiation has made them a celebrated model system in evolutionary research. Nonetheless, despite numerous studies, phylogenetic uncertainty persists, and riverine lineages remain comparatively underrepresented in higher-level phylogenetic studies. Heterogeneous gene histories resulting from incomplete lineage sorting (ILS) and hybridization are likely sources of uncertainty, especially during episodes of rapid speciation. We investigate relationships of Pseudocrenilabrinae and its close relatives while accounting for multiple sources of genetic discordance using species tree and hybrid network analyses with hundreds of single-copy exons. We improve sequence recovery for distant relatives, thereby extending the taxonomic reach of our probes, with a hybrid reference guided/<em>de novo</em> assembly approach. Our analyses provide robust hypotheses for most higher-level relationships and reveal widespread gene heterogeneity, including in riverine taxa. ILS and past hybridization are identified as sources of genetic discordance in different lineages. Sampling of various Blenniiformes (formerly Ovalentaria) adds strong phylogenomic support for convict blennies (Pholidichthyidae) as sister to Cichlidae, and points to other potentially useful protein-coding markers across the order. A reliable phylogeny with representatives from diverse environments will support ongoing taxonomic and comparative evolutionary research in the cichlid model system.</p>

opencc-zeroAug 2022View details →
zenodo32/100

FIGURE 5. Maximum likelihood phylogenetic tree constructed with UCEs and exon loci dataset for the novel species P in A new species of Plumarella (Octocorallia: Calcaxonia: Primnoidae) from the Northeast Pacific, and the redescription of Plumarella longispina Kinoshita, 1908

FIGURE 5. Maximum likelihood phylogenetic tree constructed with UCEs and exon loci dataset for the novel species P. williamsi (in bold), the redescribed species P. longispina (in red), the related taxa and rooted to outgroup genera. ML bootstrap support values&gt;70% are shown above branches.

opennotspecifiedJul 2024View details →
zenodo32/100

An evaluation of pool-sequencing transcriptome-based exon capture for population genomics of non-model species.

<p>This archive is associated with the article &ldquo;An evaluation of pool-sequencing transcriptome-based exon capture for population genomics of non-model species.&rdquo;. Authors: Emeline Deleury, Thomas Guillemaud, Aurelie Blin &amp;&nbsp; Eric Lombaert.</p> <p>The archive contains :<br> - The sequences of the 5,717 Harmonia axyridis randomly selected CDS (5717-targeted-CDS-sequences.gff3, sequence in FASTA format at the end of the file)<br> - For the subset of 3,161 targeted CDS that have a genomic match over their entire length, the positions of exons on transcripts (3161-targeted-CDS-EXON-POSITIONS.csv)</p>

opencc-by-4.0Mar 2019View details →
zenodo32/100

Widespread naturally variable human exons aid genetic interpretation (Preprint)

<p>Supplemental files for Widespread naturally variable human exons aid genetic interpretation (<a href="https://www.biorxiv.org/content/10.1101/2024.09.09.612029v1">Preprint link</a>)</p> <p>&nbsp;</p> <p>For newest version, click here: <a href="https://zenodo.org/records/15790343">https://zenodo.org/records/15790343</a></p> <p>&nbsp;</p>

opencc-by-4.0Aug 2024View details →
dryad32/100

Data from: An exon-capture system for the entire class Ophiuroidea

Exon-capture studies have typically been restricted to relatively shallow phylogenetic scales due primarily to hybridisation constraints. Here, we present an exon-capture system for an entire class of marine invertebrates, the Ophiuroidea, built upon a phylogenetically diverse transcriptome foundation. The system captures ~90% of the 1552 exon target, across all major lineages of the quarter-billion year old extant crown group. Key features of our system are: 1) basing the target on an alignment of orthologous genes determined from 52 transcriptomes spanning the phylogenetic diversity and trimmed to remove anything difficult to capture, map or align, 2) use of multiple artificial representatives based on ancestral state reconstructions rather than exemplars to improve capture and mapping of the target, 3) mapping reads to a multi-reference alignment, and 4) using patterns of site polymorphism to distinguish among paralogy, polyploidy, allelic differences and sample contamination. The resulting data gives a well-resolved tree (currently standing at 417 samples, 275,352 sites, 91% data-complete) that will transform our understanding of ophiuroid evolution and biogeography.

opencc-zeroDec 2014View details →
dryad32/100

Malinae481 exonic probe set

<p><span><span><span><span><span><span><span><span><span><span><span>The subtribe Malinae (Rosaceae) comprises close to 1,000 species and up to 30 genera. It is defined, among other characters, by a derived base chromosome number of x = 17. </span></span></span></span></span></span></span></span></span></span></span><span><span><span><span><span><span><span><span><span><span><span>A parsimonious pattern of chromosome breakage and fusion explains the derivation of the x = 17 karyotype from a polyploidisation event of two x = 9 genomes. </span></span></span></span></span></span></span></span></span></span></span><span><span><span><span><span><span><span><span><span><span><span>High collinearity between the genetic and physical maps of <em>Pyrus</em> and <em>Malus</em> as well as their identical karyotypes suggest that the reorganization of the Malinae genome occurred before the divergence of the two genera, and presumably before speciation of the whole subtribe. The preservation of extensive duplicated chromosomal segments in the genome of members of the Malinae, dating back to the genome-wide duplication, necessitates an assessment of the homology of markers used in phylogenetic studies. First, we identified single-copy orthologs and relatively recently diverged paralogs in the genome of <i>Malus</i> <i>domestica</i>. Average sequence divergence among orthologs was 3.65% in a genome-wide comparison of <i>Malus</i> and <i>Pyrus</i>, which is about one third of the average sequence divergence of 9.36% between Maleae and Amygdaleae orthologs and around half of the divergence based on fourfold degenerate site transversion of 8% between paralogs from the most recent WGD within each of these genomes. Based on this information, we developed a strategy to identify in the <i>Malus</i> genome (1) single-copy loci (presumably located within non-duplicated regions, or hidden paralogs due to differential gene loss) and (2) loci only duplicated once (one paralog only present in the genome; avoidance of multi-gene families). We further constrained our locus selection by imposing the minimum divergence between them to 6% as minimum sequence divergence between orthologs and paralogs. Specifically, we blasted 28,695 mRNAs (queries thereafter) in the range of 800–3,000 bp from the <i>Malus domestica</i> 'Golden Delicious' (GDDH13) genome v.1.1, downloaded from the Genome Database for Rosaceae, against the <i>Malus domestica</i> GDDH13 genome v.1.1 (subject thereafter) using the nucleotide BLAST search. Default settings were used except for the e-value, which was lowered to 0.00001. In an initial filtering step, we only kept the hits exceeding 70 bp and 10% of query length with ≥80% sequence similarity between query and subject. We then assigned the hits (usually corresponding to exons) to loci based on the criterion that the length of introns separating the hits does not exceed 10,000 bp. Queries that showed hits with &gt;6 loci were not taken into account. In a second, refined filtering step we only kept loci that fulfilled the following criteria: length cover and sequence similarity of the sum of all hits for a particular locus ≥90%, length of single hits ≥100 bp (in accordance with the bait length of 100 bp), intron length ≤1,200 bp, number of loci per query ≤2, sequence divergence among loci ≥6%. We then blasted the 1,280 <i>Malus</i> queries that fulfilled our selection criteria against the <i>Pyrus communis</i> Bartlett DH genome v.2.0 and applied the same selection criteria. 616 of these queries fulfilled them. Subject sequences from the <i>Malus</i> (799) and <i>Pyrus</i> (764) genomes, matching chosen mRNAs and representing full loci, were then extracted and exon-intron boundaries inferred based on the alignments together with the blasted queries. After filtering for identical numbers of loci in both genomes, sequence divergence between exons of single-copy loci in <i>Malus</i> and <i>Pyrus</i> ≤15% and exon length ≥80 bp, we ended up with 713 loci (481 loci if pairs of paralogous loci are treated as single loci), which corresponded to the 546 mRNAs. Extracted sequences were collapsed at ≥95% similarity and used for bait design. The final exonic probe set covers 2,008,479 bp in total. This way, we expect to get baits with good correspondence to the Malinae genomes, i.e. both orthologs and paralogs are targeted.</span></span></span></span></span></span></span></span></span></span></span></p>

opencc-zeroApr 2022View details →
zenodo32/100

Figure 1 in Major histocompatibility complex Class II DRB exon-2 diversity of the Eurasian lynx (Lynx lynx) in China

Figure 1. Neighbour-joining tree of the 13 Eurasian lynx DRB nucleotide sequences. Distance was based on Kimura's two-parameter distance. The bootstrap values (percentage of 1000 replicates) supporting particular clusters are shown next to the node. Scale bar represents substitutions per site.

opennotspecifiedJan 2009View details →
ClinicalTrials.gov32/100

Regorafenib in GIST With Secondary C-KIT Exon 17 Mutation

ClinicalTrials.gov study NCT02606097. IPD Sharing: Not stated. Countries: 1. Publications: 1.

restrictedIPD-UNDECIDEDFeb 2026View details →
ClinicalTrials.gov32/100

Study to Evaluate Ultevursen in Subjects With Retinitis Pigmentosa (RP) Due to Mutations in Exon 13 of the USH2A Gene

ClinicalTrials.gov study NCT06627179. IPD Sharing: NO. Countries: 10. Publications: 0.

closedIPD-NOFeb 2026View details →
ClinicalTrials.gov32/100

Phase III Trial to Evaluate the Elortinib vs Gefitinib in Advanced NSCLC With EGFR Exon 19 or 21 Mutations

ClinicalTrials.gov study NCT01024413. IPD Sharing: UNDECIDED. Countries: 1. Publications: 14.

restrictedIPD-UNDECIDEDFeb 2026View details →
ClinicalTrials.gov32/100

Erlotinib Versus Gemcitabine/Carboplatin in Chemo-naive Stage IIIB/IV Non-Small Cell Lung Cancer Patients With Epidermal Growth Factor Receptor (EGFR) Exon 19 or 21 Mutation

ClinicalTrials.gov study NCT00874419. IPD Sharing: Not stated. Countries: 1. Publications: 2.

restrictedIPD-UNDECIDEDFeb 2026View details →
ClinicalTrials.gov32/100

Study to Evaluate Safety and Tolerability of QR-421a in Subjects With RP Due to Mutations in Exon 13 of the USH2A Gene

ClinicalTrials.gov study NCT03780257. IPD Sharing: NO. Countries: 3. Publications: 2.

closedIPD-NOFeb 2026View details →
ClinicalTrials.gov32/100

Tepotinib Plus Paclitaxel in MET Amplified or MET Exon 14 Alterated Gastric and GEJ Carcinoma

ClinicalTrials.gov study NCT05439993. IPD Sharing: NO. Countries: 1. Publications: 1.

closedIPD-NOFeb 2026View details →
ClinicalTrials.gov32/100

Study of Tarloxotinib in Pts With NSCLC (EGFR Exon 20 Insertion, HER2-activating Mutations) & Other Solid Tumors With NRG1/ERBB Gene Fusions

ClinicalTrials.gov study NCT03805841. IPD Sharing: NO. Countries: 3. Publications: 1.

closedIPD-NOFeb 2026View details →
ClinicalTrials.gov32/100

A Study of Ripretinib vs Sunitinib in Patients With Advanced GIST With Specific KIT Exon Mutations Who Were Previously Treated With Imatinib

ClinicalTrials.gov study NCT05734105. IPD Sharing: Not stated. Countries: 15. Publications: 1.

restrictedIPD-UNDECIDEDFeb 2026View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record