Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

416

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

416 results for “Introns”

Learn how ShareScore rates datasets ↗
zenodo48/100

U2- and U12-type intron classifications for Physarum polycephalum in BED format

<p>A BED file containing intron information for <em>Physarum polycephalum</em>&nbsp;introns classified as U2- or U12-type by <a href="https://github.com/glarue/intronIC">intronIC</a>. This data is associated with the following manuscript:&nbsp;https://doi.org/10.1101/2020.10.12.336362; the genome and annotation file&nbsp;used to identify the introns are available here:&nbsp;https://doi.org/10.5281/zenodo.4086119.</p> <p>&nbsp;</p> <p>The file columns are:</p> <p>1. Genome FASTA record name (scaffold)</p> <p>2. Intron start coordinate (0-indexed)</p> <p>3. Intron end coordinate (1-indexed)</p> <p>4. Intron label from intronIC</p> <p>5. U12-type probability score (0-100); introns with scores &gt; 95 were considered U12-type in the manuscript</p> <p>6. Strand</p>

opencc-by-4.0Oct 2020View details →
zenodo44/100

Splicing accuracy varies across human introns, tissues, age and disease

<p>Alternative splicing impacts most multi-exonic human genes. Inaccuracies during this process may have an important role in ageing and disease. Here, we investigated splicing accuracy using RNA-sequencing data from &gt;14K control samples and 40 human body sites, focusing on split reads partially mapping to known transcripts in annotation. We show that splicing inaccuracies occur at different rates across introns and tissues and are primarily affected by the abundance of core components of the spliceosome assembly and its regulators. Using publicly available data from RNA-knockdowns and CLIP-seq binding sites of numerous spliceosomal components and related regulators, we demonstrated the importance of RNA-binding proteins in splicing accuracy. We found that age is positively correlated with a global decline in splicing fidelity, mostly affecting genes implicated in neurodegenerative diseases. We found further support for the latter by observing a genome-wide increase in splicing inaccuracies in samples affected with Alzheimer's disease as compared to neurologically normal individuals. This in-depth characterisation of splicing has important implications for our understanding of the role of inaccuracies in ageing and human disease, particularly in neurodegenerative disorders.</p>

opencc-by-4.0Mar 2023View details →
zenodo40/100

Supplementary material for "Genes divided according to the relative position of the longest intron show increased representation in different KEGG pathways"

<p>This is a 4th version of the Supplementary Information related to the article "Genes divided according to the relative position of the longest intron show increased representation in different KEGG pathways". It contains upgraded primary and control data used for gene set enrichment analyzes, results of these analyzes as well as our code for calculation of intron lenghts. Detailed information can be found in the above-mentioned article, which deals with the connection between the position of the longest introns in genes and biological functions of genes.</p>

opencc-by-4.0Jun 2024View details →
zenodo40/100

Fig. 5 in Loss and Gain of Group I Introns in the Mitochondrial Gene of the Scleractinia (Cnidaria; Anthozoa).

Fig. 5. Bayesian estimates of divergence times in scleractinians. The basal axis is a geologic time scale in units of million years ago (mya). Different time intervals are labeled with abbreviations (Cam, Cambrian; Ord, Ordovician; Sil, Silurian; Dev, Devonian; Car, Carboniferous; Per, Permian; Tri, Triassic; Jur, Jurassic; Cre, Cretaceous; Pal, Paleogene; Neo, Neogene). The chart below the phylogenetic tree gives the extinction rate (solid line) and origination rate (dashed line) in different geological periods, which were modified from Kiessling (2004) with major extinction events labeled with abbreviations (Rhae, Rhaetian; Plie, Pliensbachian; Kimm, Kimmeridgian; Ceno, Cenomanian; Maa, Maastrichtian, KT-extinction). Species with different types of intron are labeled with symbols: ●, Intron-729 (I729); ▲, Intron-893 (I893); ■, Intron-876 (I876).

opencc-by-4.0May 2017View details →
zenodo40/100

Fig. 4 in Loss and Gain of Group I Introns in the Mitochondrial Gene of the Scleractinia (Cnidaria; Anthozoa).

Fig. 4. Comparison of phylogenetic trees between the cox1 exon (left side) and intron (right side) in complex corals and corallimorpharians (A) and in sponges and robust corals (B). Tree topologies presenting the phylogenetic relationships of exons and introns were consensus trees between the maximum-likelihood analysis and Bayesian algorism. Numbers on branches are Shimedaira- Hasegawa-like/posterior probabilities. Dashed lines are potential changes in phylogenetic positions between the exon and intron trees.

opencc-by-4.0May 2017View details →
zenodo40/100

Fig. 3 in Loss and Gain of Group I Introns in the Mitochondrial Gene of the Scleractinia (Cnidaria; Anthozoa).

Fig. 3. Phylogeny and characteristics of cox1 intron traits in hexacorals. The tree topology was constructed with Mrbayes. Numbers labeled on branches are Shimodaira-Hasegawa-like support/posterior probabilities. Species with different types of introns are labeled with symbols: ●, Intron-729 (I729); ▲, Intron-893 (I893); ■, Intron-876 (I876).

opencc-by-4.0May 2017View details →
zenodo40/100

Fig. 1 in Loss and Gain of Group I Introns in the Mitochondrial Gene of the Scleractinia (Cnidaria; Anthozoa).

Fig. 1. Secondary structures of representative cox1 introns in anthozoans. A: Corallimorpharian (Rhodactis howesii); B: basal and complex corals (Gardeneris hawaiinesis); C: robust corals (Diploastrea heliopora); D: actiniarian (Metridinium senile); E: poriferian (Plakortis angulospiculatus); F: zoantharian (Savalia savaglia). Features of the secondary structure indicate the characteristics of group I introns: 10 helical elements P1~P10; consensus primary structures P, Q, R, and S in hollow letters; internal guide sequence, IGS. Initial and terminal sites of the predicted open reading frame are labeled "ORF start" and "ORF stop", respectively.

opencc-by-4.0May 2017View details →
zenodo40/100

dataset relate to article: "Functional Characterization of Two Variants at the Intron 6-Exon 7 Boundary of the KCNQ2 Potassium Channel Gene Causing Distinct Epileptic Phenotypes"

<p><strong>Sequencing Analysis performed at Fondazione Besta and carried out as part of the study mentioned at title</strong></p>

opencc-by-4.0Feb 2023View details →
zenodo40/100

Dataset related to article: "Congenital insensitivity to pain a novel mutation affecting a U12-type intron causes multiple aberrant splicing of SCN9A"

<p>raw data related to article reported at title</p>

opencc-by-4.0Feb 2023View details →
zenodo36/100

Modified version of intronIC for classifying U12-type introns in Physarum polycephalum

<p>A modified version of <a href="https://github.com/glarue/intronIC">intronIC</a>&nbsp;with position-specific BPS score weighting to reflect the unique U12-type BPS motif found in&nbsp;<em>Physarum</em>. See the README inside the included directory for instructions on recreating the intron classifications present in the manuscript.</p> <p>&nbsp;</p> <p>IMPORTANT UPDATE (08/2021): There is a typo in the README file within the archive&mdash;in order to run the program correctly, the argument flags &quot;-r12&quot; and &quot;-r2&quot; need to be inverted from how they are listed in the example command. In other words, the &quot;-r12&quot; flag needs to be placed where the &quot;-r2&quot; flag is, and vice-versa. Apologies for the error!</p>

opencc-by-4.0Nov 2020View details →
zenodo36/100

PRKN gene Introns in families from Order Primates show common Open Reading Frames

<p>Supplementary material:</p> <p><strong>01 Supplementary material. </strong>Contains Table 1, named: &ldquo;<u>PRKN</u><em> gene determined in Primates in Gene database NCBI. Supplementary material</em>&rdquo;. Data are shown for the primate species identified for this study. In addition, a Folder with title &ldquo;<em>PRKN</em> database&rdquo;, that contains the subfolders of the 34 primate species for which the <em>PRKN</em> gene was reported until April 2023. Species used for this study are indicated with the title &ldquo;0 Selected-species&rdquo;, and those not used only with the species name. In turn, all have four subfolders:</p> <ul> <li><strong>01 <em>PRKN</em> gene and mRNA variants-Fasta format</strong>: contains the sequences of the mRNA variants in FASTA format and of the <em>PRKN</em> gene.</li> <li><strong>02 Alignment&nbsp;<em>PRKN</em> gene and mRNA variants</strong>: contains the <em>PRKN</em> gene and mRNA sequences in FASTA format and the result of the alignment between the gene and the variants. Additionally, a Word document with the manual verification of the previous result. This document is not included in the primate species folders that were not used in this study.</li> <li><strong>03&nbsp;<em>PRKN</em> aminoacid variants-Fasta format: </strong>contains the amino acid sequences of the <em>PRKN </em>gene variants in FASTA format.</li> <li><strong>04 PRKN Protein variants domains</strong>: contains documents with the graphical representation of the PRKN isoform domains, in Scalable Vector Graphic (SVG) format.</li> </ul> <p><strong>02 Supplementary material.</strong>&nbsp;Table 2 in Excel document containing the positions of the introns and exons of <em>PRKN </em>gene the 16 primate species used in this study. The first sheet contains the size of the introns/exons and their location by species. In the following sheets, the species name, intron/exon position and size are indicated.</p> <p><strong>03 Supplementary material. </strong>Contains two files:</p> <ul> <li><strong>Figure 1.</strong> Document with the alignments of the Primate PRKN proteins encoded by each mRNA variant. Also, are indicated the amino acids whose codons contain the intron with their respective phase.</li> <li><strong>Table 3, named &ldquo;<em>Intron Phase for </em>PRKN<em> gene</em>&rdquo;</strong>, in Excel document containing the phases, codons, their respective amino acids and the location of the introns of the primate <em>PRKN</em> gene. The first sheet contains the consensus data for each species, grouped according to families. In the following sheets, the species name, the Phases, the codon and the respective amino acid for each intron are indicated. Phases 0, 1, and 2 are indicated in purple, red, and green letters, respectively. To Phases 1 and 2, the position of the intron in the codon is indicated with a vertical bar, whereas, for Phase 0, each codon is in separate boxes. The mRNA variants that showed loss or fusion of introns are indicated as &ldquo;Non determine&rdquo;. Variants with alternatively spliced mRNA and phase change are indicated in yellow boxes.</li> </ul> <p><strong>04 Supplementary material. </strong>Contains two Excel documents and a subfolder with the putative <em>PRKN</em> introns Open Reading Frame (ORF):</p> <ul> <li><strong>Excel file: Table 4, named &ldquo;<em>Positive ORFs</em>&rdquo;</strong>. The first sheet contains the general data of 5&rsquo;-3&rsquo; and 3&rsquo;-5&rsquo; ORFs, characterized for each species by intron and organized by families. The second sheet corresponds to the total number of exons determined for each ORF, with the previously mentioned organization. The additional sheets are divided by intron and the species are organized by families along with the different characterized ORFs.</li> <li><strong>Excel file: Table 5 named &ldquo;<em>ORFs Introns</em>&rdquo;</strong>. In the first sheet, the number of ORFs of the introns is indicated, according to the grouping performed in this study. The other Sheets correspond to both unique and common ORFs determined for each species per intron, organized by family.</li> <li><strong>Subfolder named &ldquo;<em>ORFs alignments</em>&rdquo;</strong>: Contains the alignments between the ORFs of each intron by species. Exclusive and common ORFs are also indicated.</li> </ul> <p><strong>05 Supplementary material.</strong> Contains two Excel documents with putative functional sites of PRKN introns ORFs protein:</p> <ul> <li><strong>Excel file: Table 6 named &ldquo;<em>Family Common Functional Sites</em>&rdquo;</strong>. The first sheet contains the general information of the functional sites characterized in the ORFs, determined in the introns of each species ordered by family. Also, the total count of ORFs per family. The other sheets correspond to the categories of cellular processes. Within these are the functional site according to the species that presented it, ordered by family and the total count of each site.</li> <li><strong>Excel file: Table 7 named &ldquo;<em>Protein Domain ORFs</em>&rdquo;</strong>. Functional sites determined only for ORFs common to two or more species are found. The position of each functional site is indicated according to the amino acid sequence and in relation to each species.</li> </ul> <p><strong>06 Supplementary material. </strong>Contains two subfolders that comprise BLAST and Protein Data Bank search from mRNA and PRKN ORF intron proteins:</p> <ul> <li><strong>BLAST subfolder comprises &nbsp;a Word document and 7 plain text documents. </strong>The Word document corresponds to the sequences reported in the Transcriptome Shotgun Assembly protein (tsa_nr) and Expressed sequence tags (est) databases, for ORFs that are common between Homo sapiens and the other species, according to the intron. The plaintext-like documents correspond to BLAST analyses that yielded more than six sequences similar to the reference ORF. The documents are named according to the sequences analyzed: the ORF name and database (Protein or Nucleotide).</li> <li><strong>Protein Data Bank subfolder. </strong>Contains 6 files in .PDF format, that correspond to protein structures that showed similarities with human ORFs 32 and 11, from Introns 6 and 9, respectively. Each document contains the protein structure information indicated in the Protein Data Bank (PDB) and the amino acid alignment between the Intron ORF sequence and the PDB protein sequence. The documents are named according to the protein ID obtained in the PDB and the respective ORF.</li> </ul>

openApr 2024View details →
dryad36/100

Data from: An intronic transposon insertion associates with a trans-species color polymorphism in Midas cichlid fishes

<p><span><span><span><span><span><span><span><span><span><span><span>Polymorphisms have fascinated biologists for a long time, but their genetic underpinnings often remained elusive. Here, we aimed to uncover the genetic basis of the gold/dark polymorphism that is eponymous of Midas cichlid fish (<i>Amphilophus </i>spp.) adaptive radiations in Nicaraguan crater lakes. While most Midas cichlids are of the melanic "dark morph", about 10% of individuals lose their melanic pigmentation during their ontogeny and transition into a conspicuous "gold morph". Using a new haplotype-resolved long-read assembly we discovered an 8.2kb, transposon-derived inverted repeat in an intron of an undescribed gene, which we term <i>goldentouch</i> in reference to the Greek myth of King Midas. The gene <i>goldentouch</i> is differentially expressed between morphs, likely due to structural implications of inverted repeats in both DNA and RNA (cruciform and hairpin formation). The near-perfect association with the phenotype across several independent populations suggests that this insertion likely underlies this trans-specific, stable polymorphism.</span></span></span></span></span></span></span></span></span></span></span></p>

opencc-zeroNov 2021View details →
zenodo36/100

Intronization signatures in coding exons reveal the evolutionary fluidity of eukaryotic gene architecture

<p>This dataset accompanies the manuscript &quot;Intronization signatures in coding exons reveal the evolutionary fluidity of eukaryotic gene architecture&quot;. It contains the relevant tables, scripts and figures used for and created during data analysis.</p>

opencc-by-4.0Jun 2022View details →
zenodo36/100

List of intron retention associated variants from Sequence Read Archive

<p>List of intron retention associated variants via <a href="https://github.com/friend1ws/iravnet">IRAVNet</a>&nbsp;applied to 219,615 transcriptome sequence data from Sequence Read Archive</p>

opencc-by-4.0Sep 2022View details →
zenodo36/100

intron-innovation/AfriMed-QA: AfriMed-QA-v-1.2

<h2>What's Changed</h2> <ul> <li>Evalauting Pipeline for AfriMed-QA by @owos in https://github.com/intron-innovation/AfriMed-QA/pull/1</li> <li>PR to support few shot evalution and instruction tuning by @owos in https://github.com/intron-innovation/AfriMed-QA/pull/2</li> <li>Gpt4o by @tobiolatunji in https://github.com/intron-innovation/AfriMed-QA/pull/3</li> <li>PR to fix evaluation error. by @owos in https://github.com/intron-innovation/AfriMed-QA/pull/4</li> <li>Ab eval by @tobiolatunji in https://github.com/intron-innovation/AfriMed-QA/pull/7</li> <li>edit prompt structure by @owos in https://github.com/intron-innovation/AfriMed-QA/pull/8</li> <li>Gpt4o by @tobiolatunji in https://github.com/intron-innovation/AfriMed-QA/pull/10</li> <li>Gpt3.5turbo results by @SanniM3 in https://github.com/intron-innovation/AfriMed-QA/pull/9</li> <li>Gpt4o by @tobiolatunji in https://github.com/intron-innovation/AfriMed-QA/pull/11</li> <li>Gpt4o by @tobiolatunji in https://github.com/intron-innovation/AfriMed-QA/pull/14</li> <li>all results from my end by @owos in https://github.com/intron-innovation/AfriMed-QA/pull/15</li> <li>results for gpt4turbo and claude sonnet by @SanniM3 in https://github.com/intron-innovation/AfriMed-QA/pull/16</li> <li>updated readme by @princenimo in https://github.com/intron-innovation/AfriMed-QA/pull/18</li> <li>Inference results for meditron, pmc llama, mistral, and openbio by @princenimo in https://github.com/intron-innovation/AfriMed-QA/pull/17</li> <li>claude opus results by @SanniM3 in https://github.com/intron-innovation/AfriMed-QA/pull/19</li> <li>Push results from claude_3_haiku run by @timfaniran in https://github.com/intron-innovation/AfriMed-QA/pull/20</li> <li>Results for MedLM by @owos in https://github.com/intron-innovation/AfriMed-QA/pull/21</li> <li>Results for 7 models by @princenimo in https://github.com/intron-innovation/AfriMed-QA/pull/22</li> <li>Support to Eval MedQA by @owos in https://github.com/intron-innovation/AfriMed-QA/pull/23</li> <li>push new data version by @tobiolatunji in https://github.com/intron-innovation/AfriMed-QA/pull/24</li> <li>Update README.md by @princenimo in https://github.com/intron-innovation/AfriMed-QA/pull/25</li> <li>PR for results from google models by @owos in https://github.com/intron-innovation/AfriMed-QA/pull/26</li> <li>instruction tuning results by @SanniM3 in https://github.com/intron-innovation/AfriMed-QA/pull/27</li> <li>updated results for instruct and few shot - gpt4o and claude by @SanniM3 in https://github.com/intron-innovation/AfriMed-QA/pull/29</li> <li>Add files via upload by @princenimo in https://github.com/intron-innovation/AfriMed-QA/pull/30</li> <li>MedQA Google Models by @owos in https://github.com/intron-innovation/AfriMed-QA/pull/28</li> <li>More Google Results on MedQA by @owos in https://github.com/intron-innovation/AfriMed-QA/pull/31</li> <li>Google Models: Fews shots + instruct Prompt by @owos in https://github.com/intron-innovation/AfriMed-QA/pull/32</li> <li>Add files via upload by @tejuafonja in https://github.com/intron-innovation/AfriMed-QA/pull/33</li> <li>Medqa afrimed qa 3 shot by @princenimo in https://github.com/intron-innovation/AfriMed-QA/pull/34</li> <li>Google models, Instruct Only by @owos in https://github.com/intron-innovation/AfriMed-QA/pull/35</li> <li>Update README.md by @princenimo in https://github.com/intron-innovation/AfriMed-QA/pull/36</li> </ul> <h2>New Contributors</h2> <ul> <li>@owos made their first contribution in https://github.com/intron-innovation/AfriMed-QA/pull/1</li> <li>@tobiolatunji made their first contribution in https://github.com/intron-innovation/AfriMed-QA/pull/3</li> <li>@SanniM3 made their first contribution in https://github.com/intron-innovation/AfriMed-QA/pull/9</li> <li>@princenimo made their first contribution in https://github.com/intron-innovation/AfriMed-QA/pull/18</li> <li>@timfaniran made their first contribution in https://github.com/intron-innovation/AfriMed-QA/pull/20</li> <li>@tejuafonja made their first contribution in https://github.com/intron-innovation/AfriMed-QA/pull/33</li> </ul> <p><strong>Full Changelog</strong>: https://github.com/intron-innovation/AfriMed-QA/commits/version-1</p>

opencc-by-nc-4.0Jun 2024View details →
zenodo36/100

Evidence for the association between the intronic haplotypes of ionotropic glutamate receptors and schizophrenia

<p>VCF and BED files for the publication &quot;Evidence for the association between the intronic haplotypes of ionotropic glutamate receptors and schizophrenia&quot;.</p>

opencc-by-4.0Aug 2021View details →
zenodo36/100

An intron SNP rs2069837 in IL-6 is associated with osteonecrosis of the femoral head development

<p>Two variants (rs2069837, and rs13306435) in the IL-6 gene were identified and genotyped from 566 patients with ONFH and 566 healthy controls. The associations between IL-6 polymorphisms and ONFH susceptibility were assessed using odds ratio (OR) and 95% confidence interval (95% CI) via logistic regression. The results of the overall analysis revealed that IL-6 rs2069837 is correlated with decreased risk of ONFH among the Chinese Han population (<em>p</em> &lt; 0.05). In stratified analysis, rs2069837 also reduced the susceptibility to ONFH in older people (&gt; 51 years), males, nonsmokers, and nondrinkers (<em>p</em> &lt; 0.05).</p>

opencc-by-4.0Sep 2021View details →
ClinicalTrials.gov36/100

A Study of PEGASYS (Peginterferon Alfa-2a (40KD)) in Combination With Ribavirin in Patients With Chronic Hepatitis C (CHC) Previously Treated With PEG-Intron + Ribavirin

ClinicalTrials.gov study NCT00087568. IPD Sharing: Not stated. Countries: 1. Publications: 2.

restrictedIPD-UNDECIDEDFeb 2026View details →
ClinicalTrials.gov36/100

Study of PEG-Intron for Plexiform Neurofibromas

ClinicalTrials.gov study NCT00396019. IPD Sharing: Not stated. Countries: 1. Publications: 1.

restrictedIPD-UNDECIDEDFeb 2026View details →
ClinicalTrials.gov36/100

A Phase II Study of Pegylated Interferon Alfa 2b (PEG-Intron(Trademark)) in Children With Diffuse Pontine Gliomas

ClinicalTrials.gov study NCT00036569. IPD Sharing: Not stated. Countries: 1. Publications: 3.

restrictedIPD-UNDECIDEDFeb 2026View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record