Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

665

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

665 results for “exon”

Learn how ShareScore rates datasets ↗
zenodo44/100

Predicting Exon Criticality from Protein Sequence

<p>Exon ByPASS (predicting Exon-skipping Based in Protein amino acid SequenceS), predictions on test exons from Human and Mouse transcripts. The exons in the test set from the two genomes are those that are not predicted to be skippable in hg38 and mm10 annotation and are also exons that in-frame when skipped. The preprocessed data includes the ensemble transcript id and exon rank as well the amino acid sequence for the upstream, downstream, and exon of interest. Additionally, the table contains the output probability from the model in the last column. The input data is&nbsp;transformed data of the amino acid sequence that is the Exon ByPASS model can use as an input.</p>

opencc-by-4.0Jan 2022View details →
zenodo44/100

GTEx exon-level expression summaries for Snaptron

<p>GTEx exon-level expression summary from the Snaptron collection. &nbsp;Format is a tab-separated text file compressed and indexed using BGZip, along with supplementary files containing a Tabix index for the data (ending in tbi) as well as two files describing the samples represented in the columns of the data file (ending in tsv). &nbsp;Uses GENCODE v25 annotation for quantification. &nbsp;Source data for the quantification are the bigWig files produced as part of recount2. &nbsp;More information at http://snaptron.cs.jhu.edu.</p>

opencc-by-4.0Feb 2019View details →
zenodo44/100

GTEx exon-exon-junction-level expression summary for Snaptron

<p>GTEx exon-level expression summary from the Snaptron collection. &nbsp;Format is a tab-separated text file compressed and indexed using BGZip, along with supplementary files containing a Tabix index for the data (ending in tbi) as well as two files describing the samples represented in the columns of the data file (ending in tsv). &nbsp;Uses Rail-RNA for spliced alignment. &nbsp;More information at http://snaptron.cs.jhu.edu.</p>

opencc-by-4.0Feb 2019View details →
zenodo40/100

Mapping of protein domains to exons

<p>Database table maps transcripts and exons to their isoforms and Pfam domains.&nbsp;</p> <p>Originally generated from Ensembl Biomart&nbsp;</p> <p>If you use the database, please cite:</p> <p>Zakaria Louadi, Kevin Yuan, Alexander Gress, Olga Tsoy, Olga V Kalinina, Jan Baumbach, Tim Kacprowski*, Markus List*.&nbsp;<a href="https://doi.org/10.1093/nar/gkaa768">DIGGER: exploring the functional role of alternative splicing in protein interactions</a>, Nucleic Acids Research.<br> * Joint last authors.</p> <p>Source code of DIGGER:&nbsp;<a href="https://github.com/louadi/DIGGER">https://github.com/louadi/DIGGER</a></p>

opencc-by-4.0Jun 2020View details →
zenodo40/100

Data and Analysis Scripts for "Diviner uncovers hundreds of novel human (and other) exons through comparative analysis of proteins"

<p>This compressed directory contains the complete set of data and analysis scripts represented in "<em>Diviner</em> uncovers novel coding regions on eukaryotic genomes using targeted homology search."</p> <p>&nbsp;</p>

opencc-by-4.0May 2024View details →
zenodo40/100

Exon capture museomics deciphers the nine-banded armadillo species complex and identifies a new species endemic to the Guiana Shield

<h2><strong>Exon capture museomics deciphers the nine-banded armadillo species complex and identifies a new species endemic to the Guiana Shield</strong></h2> <p>Mathilde Barthe*, Lo&iuml;s Rancilhac, Maria C. Arteaga, Anderson Feij&oacute;, Marie-Ka Tilak, Fabienne Justy, W. J. Loughry, Colleen M. McDonough, Benoit de Thoisy, Fran&ccedil;ois Catzeflis, Guillaume Billet, Lionel Hautier, Benoit Nabholz, and Fr&eacute;d&eacute;ric Delsuc*</p> <p>*Corresponding authors: mathilde.barthe.pro@gmail.com; frederic.delsuc@umontpellier.fr</p> <p>&nbsp;</p> <h2><strong>Description of available files.&nbsp;</strong></h2> <p><strong>01_Figures_&amp;_tables_of_the_main_text.zip&nbsp;</strong><br>- &nbsp; &nbsp;Figure 1: Phylogenetic relationships reconstructed by maximum likelihood and maps representing the distribution of individuals according to their lineage.<br>- &nbsp; &nbsp;Figure 2: Assignment of individuals to lineages according to phylogenetic analyses, admixture analysis and phylogenetic delimitation.<br>- &nbsp; &nbsp;Figure 3: Principal Component Analysis of genetic variance.<br>- &nbsp; &nbsp;Figure 4: Distribution map and genetic composition of individuals of the four recognized species.</p> <p><strong>02_Supplementary_tables_&amp;_figures.zip&nbsp;</strong><br>- &nbsp; &nbsp;Figure S1: Distribution of targeted nuclear loci along a chromosome scale assembly.<br>- &nbsp; &nbsp;Figure S2: Mitochondrial genome depth of coverage.<br>- &nbsp; &nbsp;Figure S3: Calculation of mitochondrial lineage support for detecting contamination.<br>- &nbsp; &nbsp;Figure S4: Mitochondrial lineage support for each individual.&nbsp;<br>- &nbsp; &nbsp;Figure S5: a) Inbreeding coefficient and b) heterozygosity estimate for individuals according to cleaning steps.<br>- &nbsp; &nbsp;Figure S6: Percentage of missing data per captured locus.&nbsp;<br>- &nbsp; &nbsp;Figure S7: Summary information of the 837 cleaned nuclear loci (number of sequences, the proportion of variable sites and the percentage of missing data).&nbsp;<br>- &nbsp; &nbsp;Figure S8: &nbsp;Phylogenetic relationships of the 62 Dasypus individuals obtained using Astral on the 832 ML gene trees from the captured nuclear loci reconstructed with IQ-Tree and ModelFinder.<br>- &nbsp; &nbsp;Figure S9: Results of analyses to detect introgression.<br>- &nbsp; &nbsp;Figure S10: Cross validation errors according to the number of clusters (K) investigated.<br>- &nbsp; &nbsp;Figure S11: Detailed analysis of the substructure within the newly recognized D. novemcinctus (Southern lineage).<br>- &nbsp; &nbsp;Figure S12: Species delimitation estimated by bPTP-h.&nbsp;<br>- &nbsp; &nbsp;Figure S13: Comparison of the three best models from the model selection estimated with PHRAPL.<br>- &nbsp; &nbsp;Figure S14: Species delimitation estimated using GMYC.<br>- &nbsp; &nbsp;Figure S15: Heatmaps of pairwise genetic indexes between lineages.&nbsp;<br>- &nbsp; &nbsp;Figure S16: Effect of filters on admixture results.&nbsp;<br>- &nbsp; &nbsp;Figure S17: Updated map from Arteaga et al. (2020).&nbsp;<br>- &nbsp; &nbsp;Figure S18: Maximum likelihood phylogenetic tree of 212 pb of the 16s ribosomal RNA of five individuals analyzed in Abba et al., (2018) and three from this study.<br>- &nbsp; &nbsp;Table S1: List of biological samples with detailed information.<br>- &nbsp; &nbsp;Table S2a: Quality statistics by locus after filtering steps.<br>- &nbsp; &nbsp;Table S2b: Quality statistics by individuals after filtering steps.<br>- &nbsp; &nbsp;Table S3: Species delimitation estimated using PHRAPL for the four combinations.<br>- &nbsp; &nbsp;Table S4: Comparison of the lineage of the nineteen individuals in common with Arteaga et al. (2020) and our study.&nbsp;<br>- &nbsp; &nbsp;Table S5: Adult cranial measurements (in millimeters) of the four Dasypus species recognized in this study following Feij&oacute; &amp; Cordeiro-Estrela (2016).&nbsp;<br>- &nbsp; &nbsp;Table S6: Adult external measurements (in millimeters) of the four Dasypus species recognized in this study.&nbsp;</p> <p><br><strong>03_Mitogenomes.zip</strong><br>- &nbsp; &nbsp;Mitogenome_reference_Dasypus_novemcinctus.fasta: Mitogenome reference used to mapped reads and extract mitochondrial DNA.<br>- &nbsp; &nbsp;Concatenated_mitochondrial_genes.fasta: Concatenated nucleotide sequences of 15 mitochondrial genes (13 protein-coding + 2 rRNAs). Sites with more than 50% missing data were excluded resulting in a total of 13,924 sites. &nbsp;<br>- &nbsp; &nbsp;Concatenated_mitochondrial_genes_partition.txt: Partition file of concatenated sequences of the 15 mitochondrial genes (13 protein-coding + 2 rRNAs).&nbsp;<br>- &nbsp; &nbsp;Concatenated_mitochondrial_genes_TESTNEW.treefile : Maximum likelihood phylogenetic tree inferred from the concatenated sequences of the 15 mitochondrial genes using IQ-TREE under a partitioned model applying ModelFinder on each partition.<br>- &nbsp; &nbsp;Depth_coverage_mitogenomes.csv: Table of mean depth of coverage and proportion of missing data (Ns) of the 72 reconstructed mitochondrial genomes sequenced for this study.&nbsp;</p> <p><br><strong>04_Reanalyses.zip</strong><br>- &nbsp; &nbsp;Dloop_alignment.fasta: Alignment of the D-loop sequences obtained in this study with those from Arteaga et al. (2020).&nbsp;<br>- &nbsp; &nbsp;Dloop_alignment.treefile: Maximum likelihood phylogenetic tree inferred from the D-loop alignment using IQ-TREE (GTR+G model).<br>- &nbsp; &nbsp;Abba_shotgun.fasta: Alignment of the 16S rRNA of individuals from this study and those from Abba et al. (2018).<br>- &nbsp; &nbsp;Abba_shotgun.fasta.treefile: Maximum likelihood phylogenetic tree inferred from the 16S rRNA alignment using IQ-TREE (GTR+G model).</p> <p><br><strong>05_Contamination_exploration.zip</strong><br>- &nbsp; &nbsp;Mitochondrial_diagnostic_positions.csv: Table of the 350 diagnostic mitochondrial positions used to estimate proportion of reads supporting each lineage. Position number refers to the Complete_mitogenome_alignment.fasta file.<br>- &nbsp; &nbsp;Read_support_to_diagnostic_positions.csv: For each individual, this table reports the Diagnostic Rate (proportion of diagnostic positions per lineage supported by at least 3 reads), the Read Proportion (mean read proportion supporting diagnostic positions per lineage), Index (proportion of synapomorphies per lineage normalized by average frequency of reads supporting these synapomorphies) and the type of tissue (museum or fresh tissue).<br>- &nbsp; &nbsp;Contamination_exploration.R: R script used to plot read support to lineages and the effect of tissue type (fresh or museum).</p> <p>&nbsp;</p> <p><strong>06_Nuclear_dataset.zip</strong><br>- &nbsp; &nbsp;TATU_1000exons4baits.fasta: Reference sequences of 1,000 exons and flanking regions used to define the probes for exon capture extracted from the Dasypus novemcinctus genome.<br>- &nbsp; &nbsp;Dasypus_capture_Final_Baits_Set.fas: Sequences of the 16,146 probes used to capture the 997 nuclear loci (exons and flanking regions).<br>- &nbsp; &nbsp;Diploid_837_nuclear_loci.fasta: Diploid sequences of the 837 nuclear loci for the 62 individuals in PopPhyl format (Locus|lineage|individual|Allele).<br>- &nbsp; &nbsp;Mean_coverage_by_individuals.csv: Table of mean depth of coverage, horizontal coverage, and number of loci per individual after filtering.&nbsp;<br>- &nbsp; &nbsp;Mean_coverage_by_loci.csv: Table of mean depth of coverage, horizontal coverage and number of loci per loci after filtering.<br>- &nbsp; &nbsp;Location_loci_targeted.bed : list of the loci targeted by exon capture, with their genomic locations on the chromosome scale assembly of <em>Dasypus novemcinctus</em> (mDasNov1.hap2)</p> <p>&nbsp;</p> <p><strong>07_Disentangling_genotyping_errors.zip</strong><br>- &nbsp; &nbsp;Table_of_heterozygosity_and_inbreeging_coefficient.csv: Table of heterozygosity (He) and inbreeding coefficient (F) estimated for each cleaning steps: initial data, after correction of heterozygous positions (must be supported by a proportion of reads between 0.3 and 0.7), and after exclusion of 159 potentially paralogous loci.<br>- &nbsp; &nbsp;Plot_effect_of_cleaning_on_He&amp;F.R: R script used to plot the effect of cleaning steps on heterozygosity (He) and inbreeding coefficient (F).</p> <p>&nbsp;</p> <p><strong>08_Distribution_maps.zip&nbsp;</strong><br>- &nbsp; &nbsp;Coordinates_according_mito_nuclear_lineages.csv: Table of GPS coordinates of individuals according to their mitochondrial and nuclear lineages.<br>- &nbsp; &nbsp;Plot_mito_nuclear_distribution.R: R script used to plot individuals on the Neotropical map according to their mitochondrial and nuclear lineages in Figure 1.<br>- &nbsp; &nbsp;Mitochondrial_distribution.pdf: Geographical distribution of the 75 individuals according to their mitochondrial lineage.<br>- &nbsp; &nbsp;Nuclear_distribution.pdf: Geographical distribution of the 58 individuals according to their nuclear lineage.</p> <p>&nbsp;</p> <p><strong>09_Phylogenetic_inference.zip</strong><br>● &nbsp; &nbsp;Phylogram_Tree&nbsp;<br>- &nbsp; &nbsp;Concatenated_nuclear_loci.fasta: Concatenated sequences of the 837 nuclear loci &nbsp;representing a total of 506,355 sites.&nbsp;<br>- &nbsp; &nbsp;Concatenated_nuclear_loci_partition.txt: Partition file for the 837 nuclear loci concatenation.<br>- &nbsp; &nbsp;Concatenated_nuclear_loci_TESTNEW.treefile: Maximum likelihood phylogenetic tree inferred from the 837 nuclear loci concatenation using IQ-TREE under a partitioned model applying ModelFinder on each partition.</p> <p>● &nbsp; &nbsp;Ultrametric_Tree<br>- &nbsp; &nbsp;Ultrametric_tree_concatenated_nuclear_loci.treefile: Ultrametric tree inferred from the 837 nuclear loci concatenation (Concatenated_nuclear_loci.fasta in Phylogram_Tree folder) using a partitioned model applying ModelFinder on each partition (Concatenated_nuclear_loci_partition.txt in Phylogram_Tree folder). The ML phylogram (Concatenated_nuclear_loci_TESTNEW.treefile in Phylogram_Tree folder) was used as a guide tree. The root was dated at 6 Mya.</p> <p>● &nbsp; &nbsp;Gene_Tree&nbsp;<br>- &nbsp; &nbsp;Concatenate_gene_tree.treefile: File containing all gene trees reconstructed using IQ-TREE applying ModelFinder to each gene.<br>- &nbsp; &nbsp;Astral_consensus_tree.txt: Summary species tree reconstructed with Astral using Concatenate_gene_tree_TESTNEW.treefile</p> <p>● &nbsp; &nbsp;Introgression analyses:<br>- &nbsp; &nbsp;Concordance_factors_Dasypus.csv &nbsp;<br>- &nbsp; &nbsp;Topology_Weighting_Dasypus_plots.R<br>- &nbsp; &nbsp;SnaQ_results_hmax0.out &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp;<br>- &nbsp; &nbsp;SnaQ_results_hmax1.out &nbsp;<br>- &nbsp; &nbsp;SnaQ_results_hmax2.out &nbsp; &nbsp; &nbsp;&nbsp;<br>- &nbsp; &nbsp;SnaQ_results_hmax3.out &nbsp; &nbsp;&nbsp;<br>- &nbsp; &nbsp;twisst_guianensis_spmap.txt &nbsp; &nbsp; &nbsp; &nbsp;<br>- &nbsp; &nbsp;twisst_guianensis_Weights<br>- &nbsp; &nbsp;twisst_mexico_spmap.txt<br>- &nbsp; &nbsp;twisst_mexico_Weights</p> <p>&nbsp;</p> <p><strong>10_Species_delimitation.zip</strong><br>● &nbsp; &nbsp;BPP&nbsp;<br>- &nbsp; &nbsp;input_for_bpp.phy: Sequence alignments of the 837 nuclear loci in phylip format.<br>- &nbsp; &nbsp;lineage_for_BPP: Correspondence file between individuals and lineages.<br>- &nbsp; &nbsp;r1 and r2: folders containing config files (bpp.ctl) and outputs of the BPP analysis.&nbsp;</p> <p>● &nbsp; &nbsp;bPTP&nbsp;<br>- &nbsp; &nbsp;PTPh_Support_Partition.txt: Details of the most supported species partition.&nbsp;<br>- &nbsp; &nbsp;PTPh_tree_partition.png: Tree illustrating the most supported species partition.&nbsp;</p> <p>● &nbsp; &nbsp;GMYC&nbsp;<br>- &nbsp; &nbsp;Script_GMYC.R: R script used to run the GMYC delimitation method on the ultrametric tree (11_Phylogenetic_inference/Ultrametric_Tree/Ultrametric_tree_concatenated_nuclear_loci.treefile).<br>- &nbsp; &nbsp;Figure_GMYC.png: Figure illustrating the results of the GMYC species delimitation analysis.</p> <p>● &nbsp; &nbsp;PHRAPL<br>- &nbsp; &nbsp;Script_PHRAPL.R: R script used to run the PHRAPL delimitation method on the 09_Phylogenetic_inference.zip/Gene_Tree /Concatenate_gene_tree.treefile</p> <p>&nbsp;</p> <p><strong>11_Population_genetic_analyses.zip</strong><br>● &nbsp; &nbsp;PCA<br>- &nbsp; &nbsp;Input_for_PCA.fasta: Diploid sequences of the 57 individuals (DNO-MC21 and DPI-L29 excluded) in PopPhyl format (Locus|species|individual|allele).<br>- &nbsp; &nbsp;PCA_Output: Output of the PopPhyl2PCA analysis using the Input_for_PCA.fasta file.<br>- &nbsp; &nbsp;Script_to_plot_PCA.R: R script used to plot PCA according to the mitochondrial lineage and nuclear composition (Admixture results).</p> <p>● &nbsp; &nbsp;ADMIXTURE<br>- &nbsp; &nbsp;lineage_for_Admixture.list: Correspondence between individuals and lineages file.<br>- &nbsp; &nbsp;Input_Admixture.*: 19,872 SNPs from nuclear data across the Dasypus complex.<br>- &nbsp; &nbsp;Output_Admixture.k.*: Output from the Admixture analysis according to K values (from 1 to 7).<br>- &nbsp; &nbsp;Output_Admixture.cv.error: Summary of the error value according to K.&nbsp;<br>- &nbsp; &nbsp;Plot_Admixture.R: R script used to plot Admixture results reordered by phylogeny.&nbsp;<br>- &nbsp; &nbsp;Plot_map_distribution_admixture.R: R script used to plot Admixture results on the Neotropical map.</p> <p>● &nbsp; &nbsp;Stats_Da_Dxy_GDI<br>- &nbsp; &nbsp;Pairwise_genetic_statistics.csv: Summary statistics computed using ABCstat_global.txt from the DILSmcsnp program for all pairwise combinations of individuals from the different lineages.&nbsp;<br>- &nbsp; &nbsp;Pairwise_GDI.csv: Genetic Differentiation Index estimates for all pairwise combinations of individuals from the different lineages.&nbsp;<br>- &nbsp; &nbsp;Plot_genetic_statistics.R: R script used to plot mean genetic statistics between lineages.</p> <p>● &nbsp; &nbsp;Sublineage_structure&nbsp;<br>○ &nbsp; &nbsp;ADMIXTURE<br>- &nbsp; &nbsp;Plot_map_distribution_sublineage_admixture.R: R script used to plot Admixture results on the Neotropical map.<br>○ &nbsp; &nbsp;PCA<br>- &nbsp; &nbsp;Input_for_PCA_southern_lineage.fasta: &nbsp;Diploid sequences of the 24 individuals of the Southern lineage in PopPhyl format (Locus|species|individual|allele).<br>- &nbsp; &nbsp;PCA_Output_southern_lineage: Output of the PopPhyl2PCA analysis using the Input_for_PCA_southern_lineage.fasta file.<br>- &nbsp; &nbsp;Script_to_plot_PCA_sublineage.R: R script to plot PCA according to the mitochondrial lineage and nuclear composition (Admixture results) focussing on individuals from the Southern lineage.</p> <p><br><strong>12_Morpho_molecular_distribution.zip</strong><br>- &nbsp; &nbsp;Coordinates_according_morphogroup_lineages.csv: GPS coordinates of individuals used in Hautier et al. (2017) according to their morphogroup.<br>- &nbsp; &nbsp;Plot_map_distribution_morpho_admixture.R: R script used to plot Admixture results and the individuals from Hautier et al. (2017) on the Neotropical map in Figure 5.<br>- &nbsp; &nbsp;Skull_lateral_*.png: Illustration of the lateral view of the skull of four individuals representing each species.<br>- &nbsp; &nbsp;Skull_sinuses_*.png: Illustration of the skull and paranasal sinuses of four individuals representing each species.</p> <p>&nbsp;</p>

opencc-by-4.0Jan 2024View details →
dryad40/100

Data from: Exon capture museomics deciphers the nine-banded armadillo species complex and identifies a new species endemic to the Guiana Shield

<p>The nine-banded armadillo (<em>Dasypus novemcinctus</em>) is the most widespread xenarthran species across the Americas. Recent studies have suggested it is composed of four morphologically and genetically distinct lineages of uncertain taxonomic status. To address this issue, we used a museomic approach to sequence 80 complete mitogenomes and capture 997 nuclear loci for 71 <em>Dasypus</em> individuals sampled across the entire distribution. We carefully cleaned up potential genotyping errors and cross contaminations that could blur species boundaries by mimicking gene flow. Our results unambiguously support four distinct lineages within the <em>D. novemcinctus</em> complex. We found cases of mito-nuclear phylogenetic discordance but only limited contemporary gene flow confined to the margins of the lineage distributions. All available evidence including the restricted gene flow, phylogenetic reconstructions based on both mitogenomes and nuclear loci, and phylogenetic delimitation methods consistently supported the four lineages within <em>D. novemcinctus</em> as four distinct species. Comparable genetic differentiation values to other recognized <em>Dasypus</em> species further reinforced their status as valid species. Considering congruent morphological results from previous studies, we provide an integrative taxonomic view to recognise four species within the <em>D. novemcinctus </em>complex: <em>D. novemcinctus</em>, <em>D. fenestratus</em>, <em>D. mexicanus</em>, and <em>D. guianensis </em>sp. nov.<em>, </em>a new species endemic of the Guiana Shield that we describe here. The two available individuals of <em>D. mazzai</em> and <em>D. sabanicola</em> were consistently nested within <em>D. novemcinctus </em>lineage and their status remains to be assessed. The present work offers a case study illustrating the power of museomics to reveal cryptic species diversity within a widely distributed and emblematic species of mammals.</p>

opencc-zeroJun 2024View details →
zenodo40/100

Scripts and analysis files for categorization of PKZILLA matching proteomic peptides into protein-unique, protein-multimatch & exon-unique, exon-multimatch categories.

<p>A .zip file containing the source data files &amp; Jupyter notebook for analysis of the <em>Prymnesium parvum</em> 12B1 PKZILLA-detecting proteomic results (<a href="https://doi.org/10.5281/zenodo.10023441">https://doi.org/10.5281/zenodo.10023441</a>), and the resulting files from the workflow. See "Analysis of proteomic results" section of the manuscript Materials and Methods for further detail.&nbsp;</p> <p><strong>Key files:</strong></p> <ul> <li>'PKZILLA-1_classify_peptides.txt' - A plaintext report of the # of classified peptides for PKZILLA-1</li> <li>'PKZILLA-2_classify_peptides.txt' - A plaintext report of the # of classified peptides for PKZILLA-2</li> <li>'./hierarchical_classified_xlsx/' - Excel spreadsheets with the classified peptides for PKZILLA-1 and PKZILLA-2</li> <li>'./Process_into_polypeptide_coordinates/' - Workflow, results, and plots for back-alignment of peptides back to PKZILLA-1 and PKZILLA-2 genomic loci</li> </ul>

opencc-by-4.0Oct 2023View details →
zenodo40/100

GTEx exon-exon junction-level splicing summary

<p>GTEx exon-exon junction-level splicing summary from the Snaptron collection. &nbsp;Format is a tab-separated text file compressed and indexed using BGZip. &nbsp;The accompanying Tabix index is also available. &nbsp;More information at http://snaptron.cs.jhu.edu.</p>

opencc-by-4.0Feb 2019View details →
zenodo40/100

GTEx exon-level expression summary

<p>GTEx exon-level expression summary from the Snaptron collection. &nbsp;Format is a tab-separated text file compressed and indexed using BGZip. &nbsp;The accompanying Tabix index is also available. &nbsp;Uses GENCODE v25 annotation for quantification. &nbsp;Source data for the quantification are the bigWig files produced as part of recount2. &nbsp;More information at http://snaptron.cs.jhu.edu.</p>

opencc-by-4.0Feb 2019View details →
zenodo40/100

Sanger sequencing traces of specific exons of the Sm.TRPM_PZQ gene from schistosome field samples

<p>Praziquantel (PZQ) is the only drug available to treat schistosomiasis, which is caused by schistosome blood flukes. In <em>Schistosoma mansononi</em>, the transient receptor potential (TRP) channel Sm.TRPM<sub>PZQ</sub> is strongly suspected to be the target of PZQ. Our <a href="https://doi.org/10.1101/2021.06.09.447779">genetic analysis of <em>S. mansoni</em> response to PZQ</a> revealed a QTL on its chromosome 3 which contains the <em>Sm.TRPM<sub>PZQ</sub></em> gene, strongly suggesting that <em>Sm.TRPM<sub>PZQ</sub></em> could be responsible for PZQ resistance. Therefore, understanding the natural variation in this gene and identifying potential resistance alleles will be a valuable tool for monitoring mass treatment programs aimed at schistosomiasis elimination.</p> <p>We investigated our schistosome collection to examine mutations present in <em>Sm.TRPM<sub>PZQ</sub></em> in natural schistosome populations. We analyzed exome sequencing data from 259 miracidia, cercariae or adult parasites from 3 African countries (Senegal, Niger, Tanzania), the Middle East (Oman) and South America (Brazil). We were able to sequence 36/41 exons of <em>Sm.TRPM<sub>PZQ</sub></em> from 122/259 parasites on average (s.e. = 18.65). We identified several mutations in critical areas of the channel. However, these mutations were supported by a limited number of reads only and required confirmation by Sanger sequencing.</p> <p>The present dataset corresponds to the sequencing effort done on specific exons which carried the mutations to be confirmed. We generated PCR products which were sequenced on on ABI sequencer using Eurofins Genomics services. The SCF files were then analyzed using PolyPhred (see manuscript for details about PCR conditions and data analysis). The trace files are available in the traces folder. Each filename carries a barcode which corresponds to a combination of sample, exon, and primer. All the combinations and corresponding barcodes are listed in the barcode_list.tsv file.</p> <p>Table header details of the barcode list:</p> <ul> <li><em>Sample</em>: the name of sample. The sample coding is as follows: species.country_patientID. BR: Brazil, SN: Senegal, NE: Niger, TZ: Tanzania, OM: Oman.</li> <li><em>Exon</em>: the exon targeted. The exon number corresponds to the exon number of isoform 5 and not the exon number of the gene.</li> <li><em>Barcode</em>: the barcode provided by Eurofins Genomics.</li> <li><em>Primer</em>: the primer used for sequencing. F: forward, R: reverse.</li> </ul>

opencc-by-4.0Aug 2021View details →
zenodo40/100

Huntingtin structure is orchestrated by HAP40 and shows a polyglutamine expansion-specific interaction with exon 1

<p>Supplementary Data files to accompany manuscript by Harding et al &quot;Huntingtin structure is orchestrated by HAP40 and shows a polyglutamine expansion-specific interaction with exon 1&quot;</p> <ul> <li>Supplementary Data 1 - Multiple sequence alignment for HTT used for Consurf analysis</li> <li>Supplementary Data 2 - Multiple sequence alignment for HAP40 used for Consurf analysis</li> <li>Supplementary Data 3 - Apo HTT cryo-EM map&nbsp;</li> <li>Supplementary Data 4 - HTT-HAP40 Q23 regularised SAXS profile</li> <li>Supplementary Data 5 - HTT-HAP40 Q54 regularised SAXS profile</li> <li>Supplementary Data&nbsp;6 - HTT-HAP40 &Delta;exon 1 regularised SAXS&nbsp;profile</li> <li>Supplementary Data 7 - XL-MS data</li> <li>Supplementary Data 8 - HTT-HAP40 ensemble weightings</li> <li>Supplementary Data 9 - HTT-HAP40 Q23 ensemble models</li> <li>Supplementary Data 10 - HTT-HAP40 Q54 ensemble models</li> <li>Supplementary Data 11 - HTT-HAP40&nbsp;&Delta;exon 1&nbsp;ensemble models</li> </ul>

opencc-by-4.0Sep 2021View details →
zenodo40/100

Datset related to article "DUPLICATION OF EXONS 15 AND 16 IN MATRIN-3: A PHENOTYPE BRIDGING AMYOTROPHIC LATERAL SCLEROSIS AND IMMUNE-MEDIATED DISORDERS"

<p><strong>NGS analysis performed at Fondazione Besta carried out as part of the study reported at title&nbsp;</strong></p>

opencc-by-4.0Feb 2023View details →
zenodo40/100

dataset relate to article: "Functional Characterization of Two Variants at the Intron 6-Exon 7 Boundary of the KCNQ2 Potassium Channel Gene Causing Distinct Epileptic Phenotypes"

<p><strong>Sequencing Analysis performed at Fondazione Besta and carried out as part of the study mentioned at title</strong></p>

opencc-by-4.0Feb 2023View details →
ClinicalTrials.gov40/100

Basket Study of Neratinib in Participants With Solid Tumors Harboring Somatic HER2 or EGFR Exon 18 Mutations

ClinicalTrials.gov study NCT01953926. IPD Sharing: YES. Countries: 13. Publications: 10.

controlledIPD-YESFeb 2026View details →
dryad40/100

Data from: Exon capture museomics deciphers the nine-banded armadillo species complex and identifies a new species endemic to the Guiana Shield

Open the record for dataset details and reuse information.

publicJul 2024View details →
zenodo36/100

A 63-bp insertion in exon 2 of the porcine KIF21A gene is associated with arthrogryposis multiplex congenita

<p>This dataset includes genotype and phenotype data of 11 affected (case) and 23 unaffected (control) pigs for AMC case-control haplotype association testing and a VCF file of 809 candidate variants compatible with recessive inheritance of AMC.</p> <ol> <li>Genotype data (AMC.bed, AMC.bim and AMC.fam): &nbsp;34 pigs genotyped with&nbsp;Illumina PorcineSNP60 Genotyping Bead chips.</li> <li>Phenotype data (AMC.pheno): the first, second and third column is Family ID, Animal ID and phenotype (1=case;2=control), respectively.</li> <li>AMC candidate variants compatible with recessive inheritance (AMC_candidate_muations.vcf.gz and&nbsp;AMC_candidate_muations.vcf.gz.tbi). Sequencing data of five animals in AMC pedigree have also been deposited at the Sequence Read Archive of the NCBI at the BioProject PRJNA622908&nbsp;under sample accessions SAMN14532191, SAMN14532478, SAMN14532769, SAMN14532790 and SAMN14532792.</li> </ol>

opencc-by-4.0Mar 2020View details →
dryad36/100

Novel DMD mouse model carrying a multi-exonic Dmd deletion exhibit progressive muscular dystrophy and early-onset cardiomyopathy

Duchenne muscular dystrophy (DMD) is a life-threatening neuromuscular disease caused by the lack of dystrophin, resulting in progressive muscle wasting and locomotor dysfunctions. By adulthood, almost all patients also develop cardiomyopathy, which is the primary cause of death in DMD. While there has been extensive effort in creating animal models to study treatment strategies for DMD, most fail to recapitulate the complete skeletal and cardiac disease manifestations that are presented in affected patients. Here, we generated a mouse model mirroring a patient deletion mutation of exons 52-54 (<i>Dmd &amp;[Delta]52-54</i>). The <i>Dmd &amp;[Delta]52-54</i> mutation led to the absence of dystrophin, resulting in progressive muscle deterioration with weakened muscle strength. Moreover, <i>Dmd &amp;[Delta]52-54</i> present with early-onset cardiomyopathy which is absent in current pre-clinical dystrophin deficient mouse models. Therefore, <i>Dmd &amp;[Delta]52-54</i> presents itself as an excellent pre-clinical model to evaluate the impact on skeletal and cardiac muscles for both mutation dependent and independent approaches.

opencc-zeroAug 2020View details →
zenodo36/100

EBAII - Linux training toy dataset - bed file of hg38 exons

<p>This is a toy dataset for the Linux training session at EBAII (Initiation au traitement des donn&eacute;es de g&eacute;nomique obtenues par s&eacute;quen&ccedil;age &agrave; haut d&eacute;bit)</p>

opencc-by-4.0Nov 2021View details →
zenodo36/100

Computational models from: Allele-specific activation, enzyme kinetics, and inhibitor sensitivities of EGFR exon 19 deletion mutations in lung cancer

<p>Computational models,&nbsp;compressed molecular dynamics (MD) simulation trajectories, and sample input files&nbsp;for &quot;Allele-specific activation, enzyme kinetics, and inhibitor sensitivities of EGFR exon 19 deletion mutations in lung cancer&quot;. An early version of this manuscript is available as a preprint here:&nbsp;https://www.biorxiv.org/content/10.1101/2022.03.16.484661v1</p>

opencc-by-4.0May 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record