Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
19
datasets available to search
ShareScore release 0.7.1
Dataset results
19 results for “exon capture”
Exon capture museomics deciphers the nine-banded armadillo species complex and identifies a new species endemic to the Guiana Shield
<h2><strong>Exon capture museomics deciphers the nine-banded armadillo species complex and identifies a new species endemic to the Guiana Shield</strong></h2> <p>Mathilde Barthe*, Loïs Rancilhac, Maria C. Arteaga, Anderson Feijó, Marie-Ka Tilak, Fabienne Justy, W. J. Loughry, Colleen M. McDonough, Benoit de Thoisy, François Catzeflis, Guillaume Billet, Lionel Hautier, Benoit Nabholz, and Frédéric Delsuc*</p> <p>*Corresponding authors: mathilde.barthe.pro@gmail.com; frederic.delsuc@umontpellier.fr</p> <p> </p> <h2><strong>Description of available files. </strong></h2> <p><strong>01_Figures_&_tables_of_the_main_text.zip </strong><br>- Figure 1: Phylogenetic relationships reconstructed by maximum likelihood and maps representing the distribution of individuals according to their lineage.<br>- Figure 2: Assignment of individuals to lineages according to phylogenetic analyses, admixture analysis and phylogenetic delimitation.<br>- Figure 3: Principal Component Analysis of genetic variance.<br>- Figure 4: Distribution map and genetic composition of individuals of the four recognized species.</p> <p><strong>02_Supplementary_tables_&_figures.zip </strong><br>- Figure S1: Distribution of targeted nuclear loci along a chromosome scale assembly.<br>- Figure S2: Mitochondrial genome depth of coverage.<br>- Figure S3: Calculation of mitochondrial lineage support for detecting contamination.<br>- Figure S4: Mitochondrial lineage support for each individual. <br>- Figure S5: a) Inbreeding coefficient and b) heterozygosity estimate for individuals according to cleaning steps.<br>- Figure S6: Percentage of missing data per captured locus. <br>- Figure S7: Summary information of the 837 cleaned nuclear loci (number of sequences, the proportion of variable sites and the percentage of missing data). <br>- Figure S8: Phylogenetic relationships of the 62 Dasypus individuals obtained using Astral on the 832 ML gene trees from the captured nuclear loci reconstructed with IQ-Tree and ModelFinder.<br>- Figure S9: Results of analyses to detect introgression.<br>- Figure S10: Cross validation errors according to the number of clusters (K) investigated.<br>- Figure S11: Detailed analysis of the substructure within the newly recognized D. novemcinctus (Southern lineage).<br>- Figure S12: Species delimitation estimated by bPTP-h. <br>- Figure S13: Comparison of the three best models from the model selection estimated with PHRAPL.<br>- Figure S14: Species delimitation estimated using GMYC.<br>- Figure S15: Heatmaps of pairwise genetic indexes between lineages. <br>- Figure S16: Effect of filters on admixture results. <br>- Figure S17: Updated map from Arteaga et al. (2020). <br>- Figure S18: Maximum likelihood phylogenetic tree of 212 pb of the 16s ribosomal RNA of five individuals analyzed in Abba et al., (2018) and three from this study.<br>- Table S1: List of biological samples with detailed information.<br>- Table S2a: Quality statistics by locus after filtering steps.<br>- Table S2b: Quality statistics by individuals after filtering steps.<br>- Table S3: Species delimitation estimated using PHRAPL for the four combinations.<br>- Table S4: Comparison of the lineage of the nineteen individuals in common with Arteaga et al. (2020) and our study. <br>- Table S5: Adult cranial measurements (in millimeters) of the four Dasypus species recognized in this study following Feijó & Cordeiro-Estrela (2016). <br>- Table S6: Adult external measurements (in millimeters) of the four Dasypus species recognized in this study. </p> <p><br><strong>03_Mitogenomes.zip</strong><br>- Mitogenome_reference_Dasypus_novemcinctus.fasta: Mitogenome reference used to mapped reads and extract mitochondrial DNA.<br>- Concatenated_mitochondrial_genes.fasta: Concatenated nucleotide sequences of 15 mitochondrial genes (13 protein-coding + 2 rRNAs). Sites with more than 50% missing data were excluded resulting in a total of 13,924 sites. <br>- Concatenated_mitochondrial_genes_partition.txt: Partition file of concatenated sequences of the 15 mitochondrial genes (13 protein-coding + 2 rRNAs). <br>- Concatenated_mitochondrial_genes_TESTNEW.treefile : Maximum likelihood phylogenetic tree inferred from the concatenated sequences of the 15 mitochondrial genes using IQ-TREE under a partitioned model applying ModelFinder on each partition.<br>- Depth_coverage_mitogenomes.csv: Table of mean depth of coverage and proportion of missing data (Ns) of the 72 reconstructed mitochondrial genomes sequenced for this study. </p> <p><br><strong>04_Reanalyses.zip</strong><br>- Dloop_alignment.fasta: Alignment of the D-loop sequences obtained in this study with those from Arteaga et al. (2020). <br>- Dloop_alignment.treefile: Maximum likelihood phylogenetic tree inferred from the D-loop alignment using IQ-TREE (GTR+G model).<br>- Abba_shotgun.fasta: Alignment of the 16S rRNA of individuals from this study and those from Abba et al. (2018).<br>- Abba_shotgun.fasta.treefile: Maximum likelihood phylogenetic tree inferred from the 16S rRNA alignment using IQ-TREE (GTR+G model).</p> <p><br><strong>05_Contamination_exploration.zip</strong><br>- Mitochondrial_diagnostic_positions.csv: Table of the 350 diagnostic mitochondrial positions used to estimate proportion of reads supporting each lineage. Position number refers to the Complete_mitogenome_alignment.fasta file.<br>- Read_support_to_diagnostic_positions.csv: For each individual, this table reports the Diagnostic Rate (proportion of diagnostic positions per lineage supported by at least 3 reads), the Read Proportion (mean read proportion supporting diagnostic positions per lineage), Index (proportion of synapomorphies per lineage normalized by average frequency of reads supporting these synapomorphies) and the type of tissue (museum or fresh tissue).<br>- Contamination_exploration.R: R script used to plot read support to lineages and the effect of tissue type (fresh or museum).</p> <p> </p> <p><strong>06_Nuclear_dataset.zip</strong><br>- TATU_1000exons4baits.fasta: Reference sequences of 1,000 exons and flanking regions used to define the probes for exon capture extracted from the Dasypus novemcinctus genome.<br>- Dasypus_capture_Final_Baits_Set.fas: Sequences of the 16,146 probes used to capture the 997 nuclear loci (exons and flanking regions).<br>- Diploid_837_nuclear_loci.fasta: Diploid sequences of the 837 nuclear loci for the 62 individuals in PopPhyl format (Locus|lineage|individual|Allele).<br>- Mean_coverage_by_individuals.csv: Table of mean depth of coverage, horizontal coverage, and number of loci per individual after filtering. <br>- Mean_coverage_by_loci.csv: Table of mean depth of coverage, horizontal coverage and number of loci per loci after filtering.<br>- Location_loci_targeted.bed : list of the loci targeted by exon capture, with their genomic locations on the chromosome scale assembly of <em>Dasypus novemcinctus</em> (mDasNov1.hap2)</p> <p> </p> <p><strong>07_Disentangling_genotyping_errors.zip</strong><br>- Table_of_heterozygosity_and_inbreeging_coefficient.csv: Table of heterozygosity (He) and inbreeding coefficient (F) estimated for each cleaning steps: initial data, after correction of heterozygous positions (must be supported by a proportion of reads between 0.3 and 0.7), and after exclusion of 159 potentially paralogous loci.<br>- Plot_effect_of_cleaning_on_He&F.R: R script used to plot the effect of cleaning steps on heterozygosity (He) and inbreeding coefficient (F).</p> <p> </p> <p><strong>08_Distribution_maps.zip </strong><br>- Coordinates_according_mito_nuclear_lineages.csv: Table of GPS coordinates of individuals according to their mitochondrial and nuclear lineages.<br>- Plot_mito_nuclear_distribution.R: R script used to plot individuals on the Neotropical map according to their mitochondrial and nuclear lineages in Figure 1.<br>- Mitochondrial_distribution.pdf: Geographical distribution of the 75 individuals according to their mitochondrial lineage.<br>- Nuclear_distribution.pdf: Geographical distribution of the 58 individuals according to their nuclear lineage.</p> <p> </p> <p><strong>09_Phylogenetic_inference.zip</strong><br>● Phylogram_Tree <br>- Concatenated_nuclear_loci.fasta: Concatenated sequences of the 837 nuclear loci representing a total of 506,355 sites. <br>- Concatenated_nuclear_loci_partition.txt: Partition file for the 837 nuclear loci concatenation.<br>- Concatenated_nuclear_loci_TESTNEW.treefile: Maximum likelihood phylogenetic tree inferred from the 837 nuclear loci concatenation using IQ-TREE under a partitioned model applying ModelFinder on each partition.</p> <p>● Ultrametric_Tree<br>- Ultrametric_tree_concatenated_nuclear_loci.treefile: Ultrametric tree inferred from the 837 nuclear loci concatenation (Concatenated_nuclear_loci.fasta in Phylogram_Tree folder) using a partitioned model applying ModelFinder on each partition (Concatenated_nuclear_loci_partition.txt in Phylogram_Tree folder). The ML phylogram (Concatenated_nuclear_loci_TESTNEW.treefile in Phylogram_Tree folder) was used as a guide tree. The root was dated at 6 Mya.</p> <p>● Gene_Tree <br>- Concatenate_gene_tree.treefile: File containing all gene trees reconstructed using IQ-TREE applying ModelFinder to each gene.<br>- Astral_consensus_tree.txt: Summary species tree reconstructed with Astral using Concatenate_gene_tree_TESTNEW.treefile</p> <p>● Introgression analyses:<br>- Concordance_factors_Dasypus.csv <br>- Topology_Weighting_Dasypus_plots.R<br>- SnaQ_results_hmax0.out <br>- SnaQ_results_hmax1.out <br>- SnaQ_results_hmax2.out <br>- SnaQ_results_hmax3.out <br>- twisst_guianensis_spmap.txt <br>- twisst_guianensis_Weights<br>- twisst_mexico_spmap.txt<br>- twisst_mexico_Weights</p> <p> </p> <p><strong>10_Species_delimitation.zip</strong><br>● BPP <br>- input_for_bpp.phy: Sequence alignments of the 837 nuclear loci in phylip format.<br>- lineage_for_BPP: Correspondence file between individuals and lineages.<br>- r1 and r2: folders containing config files (bpp.ctl) and outputs of the BPP analysis. </p> <p>● bPTP <br>- PTPh_Support_Partition.txt: Details of the most supported species partition. <br>- PTPh_tree_partition.png: Tree illustrating the most supported species partition. </p> <p>● GMYC <br>- Script_GMYC.R: R script used to run the GMYC delimitation method on the ultrametric tree (11_Phylogenetic_inference/Ultrametric_Tree/Ultrametric_tree_concatenated_nuclear_loci.treefile).<br>- Figure_GMYC.png: Figure illustrating the results of the GMYC species delimitation analysis.</p> <p>● PHRAPL<br>- Script_PHRAPL.R: R script used to run the PHRAPL delimitation method on the 09_Phylogenetic_inference.zip/Gene_Tree /Concatenate_gene_tree.treefile</p> <p> </p> <p><strong>11_Population_genetic_analyses.zip</strong><br>● PCA<br>- Input_for_PCA.fasta: Diploid sequences of the 57 individuals (DNO-MC21 and DPI-L29 excluded) in PopPhyl format (Locus|species|individual|allele).<br>- PCA_Output: Output of the PopPhyl2PCA analysis using the Input_for_PCA.fasta file.<br>- Script_to_plot_PCA.R: R script used to plot PCA according to the mitochondrial lineage and nuclear composition (Admixture results).</p> <p>● ADMIXTURE<br>- lineage_for_Admixture.list: Correspondence between individuals and lineages file.<br>- Input_Admixture.*: 19,872 SNPs from nuclear data across the Dasypus complex.<br>- Output_Admixture.k.*: Output from the Admixture analysis according to K values (from 1 to 7).<br>- Output_Admixture.cv.error: Summary of the error value according to K. <br>- Plot_Admixture.R: R script used to plot Admixture results reordered by phylogeny. <br>- Plot_map_distribution_admixture.R: R script used to plot Admixture results on the Neotropical map.</p> <p>● Stats_Da_Dxy_GDI<br>- Pairwise_genetic_statistics.csv: Summary statistics computed using ABCstat_global.txt from the DILSmcsnp program for all pairwise combinations of individuals from the different lineages. <br>- Pairwise_GDI.csv: Genetic Differentiation Index estimates for all pairwise combinations of individuals from the different lineages. <br>- Plot_genetic_statistics.R: R script used to plot mean genetic statistics between lineages.</p> <p>● Sublineage_structure <br>○ ADMIXTURE<br>- Plot_map_distribution_sublineage_admixture.R: R script used to plot Admixture results on the Neotropical map.<br>○ PCA<br>- Input_for_PCA_southern_lineage.fasta: Diploid sequences of the 24 individuals of the Southern lineage in PopPhyl format (Locus|species|individual|allele).<br>- PCA_Output_southern_lineage: Output of the PopPhyl2PCA analysis using the Input_for_PCA_southern_lineage.fasta file.<br>- Script_to_plot_PCA_sublineage.R: R script to plot PCA according to the mitochondrial lineage and nuclear composition (Admixture results) focussing on individuals from the Southern lineage.</p> <p><br><strong>12_Morpho_molecular_distribution.zip</strong><br>- Coordinates_according_morphogroup_lineages.csv: GPS coordinates of individuals used in Hautier et al. (2017) according to their morphogroup.<br>- Plot_map_distribution_morpho_admixture.R: R script used to plot Admixture results and the individuals from Hautier et al. (2017) on the Neotropical map in Figure 5.<br>- Skull_lateral_*.png: Illustration of the lateral view of the skull of four individuals representing each species.<br>- Skull_sinuses_*.png: Illustration of the skull and paranasal sinuses of four individuals representing each species.</p> <p> </p>
Data from: Exon capture museomics deciphers the nine-banded armadillo species complex and identifies a new species endemic to the Guiana Shield
<p>The nine-banded armadillo (<em>Dasypus novemcinctus</em>) is the most widespread xenarthran species across the Americas. Recent studies have suggested it is composed of four morphologically and genetically distinct lineages of uncertain taxonomic status. To address this issue, we used a museomic approach to sequence 80 complete mitogenomes and capture 997 nuclear loci for 71 <em>Dasypus</em> individuals sampled across the entire distribution. We carefully cleaned up potential genotyping errors and cross contaminations that could blur species boundaries by mimicking gene flow. Our results unambiguously support four distinct lineages within the <em>D. novemcinctus</em> complex. We found cases of mito-nuclear phylogenetic discordance but only limited contemporary gene flow confined to the margins of the lineage distributions. All available evidence including the restricted gene flow, phylogenetic reconstructions based on both mitogenomes and nuclear loci, and phylogenetic delimitation methods consistently supported the four lineages within <em>D. novemcinctus</em> as four distinct species. Comparable genetic differentiation values to other recognized <em>Dasypus</em> species further reinforced their status as valid species. Considering congruent morphological results from previous studies, we provide an integrative taxonomic view to recognise four species within the <em>D. novemcinctus </em>complex: <em>D. novemcinctus</em>, <em>D. fenestratus</em>, <em>D. mexicanus</em>, and <em>D. guianensis </em>sp. nov.<em>, </em>a new species endemic of the Guiana Shield that we describe here. The two available individuals of <em>D. mazzai</em> and <em>D. sabanicola</em> were consistently nested within <em>D. novemcinctus </em>lineage and their status remains to be assessed. The present work offers a case study illustrating the power of museomics to reveal cryptic species diversity within a widely distributed and emblematic species of mammals.</p>
Data from: Exon capture museomics deciphers the nine-banded armadillo species complex and identifies a new species endemic to the Guiana Shield
Open the record for dataset details and reuse information.
Data from: Genome-wide exon-capture approach identifies genetic variants of Norway spruce genes associated with susceptibility to Heterobasidion parviporum infection
Root and butt rot caused by members of the Heterobasidion annosum species complex is the most economically important disease of conifer trees in boreal forests. Wood decay in the infected trees dramatically decreases their value and causes considerable losses to forest owners. Trees vary in their susceptibility to Heterobasidion infection, but the genetic determinants underlying the variation in the susceptibility are not well-understood. We performed the identification of Norway spruce genes associated with the resistance to Heterobasidion parviporum infection using genome-wide exon-capture approach. Sixty-four clonal Norway spruce lines were phenotyped, and their responses to H. parviporum inoculation were determined by lesion length measurements. Afterwards, the spruce lines were genotyped by targeted resequencing and identification of genetic variants (SNPs). Genome-wide association analysis identified 10 SNPs located within 8 genes as significantly associated with the larger necrotic lesions in response to H. parviporum inoculation. The genetic variants identified in our analysis are potential marker candidates for future screening programs aiming at the differentiation of disease-susceptible and resistant trees.
Flatfish exon-capture
<p>This dataset contains alignments used to study phylogenetic relationships of flatfishes based on exon-capture data. The samples in this dataset represent 89 species: 86 flatfishes and 3 outgroup Carangidae. 57 samples were extracted and assembled from tissues collected for this project, while 39 were sourced from previously assembled data that were prepared as part of the FishLife project. For this study we targeted 4,434 markers developed by Jiang et al. (2019) for capture efficiency in ray-finned fishes (Actinopterygii). Library preparation followed the protocol from Li et al. (2013) involving a double capture method used to increase concentrations of hybridized DNA. Raw reads were assembled into loci using the Assexon bioinformatics pipeline (Yuan et al. 2020). Assembled exons were aligned on amino acids using MAFFT (Katoh et al., 2002), then translated back to codon-based alignment using a custom perl script mafft_aln.pl, and poorly aligned markers were removed. These aligned data represent the unfiltered dataset in the associated study, which was further processed into subsets based on missing data, clocklikeness, nucleotide composition, and evolutionary rate.</p>
An evaluation of pool-sequencing transcriptome-based exon capture for population genomics of non-model species.
<p>This archive is associated with the article “An evaluation of pool-sequencing transcriptome-based exon capture for population genomics of non-model species.”. Authors: Emeline Deleury, Thomas Guillemaud, Aurelie Blin & Eric Lombaert.</p> <p>The archive contains :<br> - The sequences of the 5,717 Harmonia axyridis randomly selected CDS (5717-targeted-CDS-sequences.gff3, sequence in FASTA format at the end of the file)<br> - For the subset of 3,161 targeted CDS that have a genomic match over their entire length, the positions of exons on transcripts (3161-targeted-CDS-EXON-POSITIONS.csv)</p>
Data from: An exon-capture system for the entire class Ophiuroidea
Exon-capture studies have typically been restricted to relatively shallow phylogenetic scales due primarily to hybridisation constraints. Here, we present an exon-capture system for an entire class of marine invertebrates, the Ophiuroidea, built upon a phylogenetically diverse transcriptome foundation. The system captures ~90% of the 1552 exon target, across all major lineages of the quarter-billion year old extant crown group. Key features of our system are: 1) basing the target on an alignment of orthologous genes determined from 52 transcriptomes spanning the phylogenetic diversity and trimmed to remove anything difficult to capture, map or align, 2) use of multiple artificial representatives based on ancestral state reconstructions rather than exemplars to improve capture and mapping of the target, 3) mapping reads to a multi-reference alignment, and 4) using patterns of site polymorphism to distinguish among paralogy, polyploidy, allelic differences and sample contamination. The resulting data gives a well-resolved tree (currently standing at 417 samples, 275,352 sites, 91% data-complete) that will transform our understanding of ophiuroid evolution and biogeography.
Data from: Exon capture phylogenomics: efficacy across scales of divergence
Open the record for dataset details and reuse information.
Flatfish exon-capture
Open the record for dataset details and reuse information.
Data from: An exon-capture system for the entire class Ophiuroidea
Open the record for dataset details and reuse information.
Data from: Genome-wide exon-capture approach identifies genetic variants of Norway spruce genes associated with susceptibility to Heterobasidion parviporum infection
Open the record for dataset details and reuse information.
Mimicry and mitonuclear discordance in nudibranchs: new insights from exon capture phylogenomics
<p>Phylogenetic inference and species delimitation can be challenging in taxonomic groups that have recently radiated and where introgression produces conflicting gene trees, especially when species delimitation has traditionally relied on mitochondrial data and colour pattern. <i>Chromodoris</i>, a genus of colourful and toxic nudibranch in the Indo-Pacific, has been shown to have extraordinary cryptic diversity and mimicry, and has recently radiated, ultimately complicating species delimitation. In these cases, additional genome-wide data can help improve phylogenetic resolution and provide important insights about evolutionary history. Here, we employ a transcriptome-based exon capture approach to resolve <i>Chromodoris</i> phylogeny with data from 2,925 exons and 1,630 genes, derived from 15 nudibranch transcriptomes. We show that some previously identified mimics instead show mitonuclear discordance, likely deriving from introgression or mitochondrial capture, but we confirm one 'pure' mimic in Western Australia. Sister-species relationships and species-level entities were recovered with high support in both concatenated Maximum Likelihood (ML) and summary coalescent phylogenies, but the ML topologies were highly variable while the coalescent topologies were consistent across datasets. Our work also demonstrates the broad phylogenetic utility of 149 genes that were previously identified from eupulmonate gastropods. This study is one of the first to i) demonstrate the efficacy of exon capture for recovering relationships among recently radiated invertebrate taxa, ii) employ genome-wide nuclear markers to test mimicry hypotheses in nudibranchs and iii) provide evidence for introgression and mitochondrial capture in nudibranchs.</p>
Data from: An evaluation of transcriptome-based exon capture for frog phylogenomics across multiple scales of divergence (Class: Amphibia, Order: Anura)
Custom sequence capture experiments are becoming an efficient approach for gathering large sets of orthologous markers in nonmodel organisms. Transcriptome-based exon capture utilizes transcript sequences to design capture probes, typically using a reference genome to identify intron–exon boundaries to exclude shorter exons (<200 bp). Here, we test directly using transcript sequences for probe design, which are often composed of multiple exons of varying lengths. Using 1260 orthologous transcripts, we conducted sequence captures across multiple phylogenetic scales for frogs, including outgroups ~100 Myr divergent from the ingroup. We recovered a large phylogenomic data set consisting of sequence alignments for 1047 of the 1260 transcriptome-based loci (~561 000 bp) and a large quantity of highly variable regions flanking the exons in transcripts (~70 000 bp), the latter improving substantially by only including ingroup species (~797 000 bp). We recovered both shorter (<100 bp) and longer exons (>200 bp), with no major reduction in coverage towards the ends of exons. We observed significant differences in the performance of blocking oligos for target enrichment and nontarget depletion during captures, and differences in PCR duplication rates resulting from the number of individuals pooled for capture reactions. We explicitly tested the effects of phylogenetic distance on capture sensitivity, specificity, and missing data, and provide a baseline estimate of expectations for these metrics based on a priori knowledge of nuclear pairwise differences among samples. We provide recommendations for transcriptome-based exon capture design based on our results, cost estimates and offer multiple pipelines for data assembly and analysis.
Data from: Identification and qualification of 500 nuclear, single-copy, orthologous genes for the Eupulmonata (Gastropoda) using transcriptome sequencing and exon capture
The qualification of orthology is a significant challenge when developing large, multiloci phylogenetic data sets from assembled transcripts. Transcriptome assemblies have various attributes, such as fragmentation, frameshifts and mis-indexing, which pose problems to automated methods of orthology assessment. Here, we identify a set of orthologous single-copy genes from transcriptome assemblies for the land snails and slugs (Eupulmonata) using a thorough approach to orthology determination involving manual alignment curation, gene tree assessment and sequencing from genomic DNA. We qualified the orthology of 500 nuclear, protein-coding genes from the transcriptome assemblies of 21 eupulmonate species to produce the most complete phylogenetic data matrix for a major molluscan lineage to date, both in terms of taxon and character completeness. Exon capture targeting 490 of the 500 genes (those with at least one exon >120 bp) from 22 species of Australian Camaenidae successfully captured sequences of 2825 exons (representing all targeted genes), with only a 3.7% reduction in the data matrix due to the presence of putative paralogs or pseudogenes. The automated pipeline Agalma retrieved the majority of the manually qualified 500 single-copy gene set and identified a further 375 putative single-copy genes, although it failed to account for fragmented transcripts resulting in lower data matrix completeness when considering the original 500 genes. This could potentially explain the minor inconsistencies we observed in the supported topologies for the 21 eupulmonate species between the manually curated and 'Agalma-equivalent' data set (sharing 458 genes). Overall, our study confirms the utility of the 500 gene set to resolve phylogenetic relationships at a range of evolutionary depths and highlights the importance of addressing fragmentation at the homolog alignment stage for probe design.
Data from: Assexon: assembling exon using gene capture data
Open the record for dataset details and reuse information.
Data from: Identification and qualification of 500 nuclear, single-copy, orthologous genes for the Eupulmonata (Gastropoda) using transcriptome sequencing and exon capture
Open the record for dataset details and reuse information.
Data from: SNP discovery in candidate adaptive genes using exon capture in a free-ranging alpine ungulate
Open the record for dataset details and reuse information.
Data from: An evaluation of transcriptome-based exon capture for frog phylogenomics across multiple scales of divergence (Class: Amphibia, Order: Anura)
Open the record for dataset details and reuse information.
Mimicry and mitonuclear discordance in nudibranchs: new insights from exon capture phylogenomics
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.