Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,293

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,293 results for “gene sequencing”

Learn how ShareScore rates datasets ↗
dryad40/100

Complex models of sequence evolution improve fit, but not gene tree discordance, for tetrapod mitogenomes

<p>Variation in gene tree estimates is widely observed in empirical phylogenomic data and is often assumed to be the result of biological processes. However, a recent study using tetrapod mitochondrial genomes to control for biological sources of variation due to their haploid, uniparentally inherited, and non-recombining nature found that levels of discordance among mitochondrial gene trees were comparable to those found in studies that assume only biological sources of variation. Additionally, they found that several of the models of sequence evolution chosen to infer gene trees were doing an inadequate job of fitting the sequence data. These results indicated that significant amounts of gene tree discordance in empirical data may be due to poor fit of sequence evolution models and that more complex and biologically realistic models may be needed. To test how the fit of sequence evolution models relates to gene tree discordance, we analyzed the same mitochondrial datasets as the previous study using two additional, more complex models of sequence evolution that each model a different biologically realistic aspect of the evolutionary process: a covarion model to incorporate heterotachy, and a model partitioned model to incorporate variable evolutionary patterns by codon position. Our results show that both additional models fit the data better than the models used in the previous study, with the covarion being consistently and strongly preferred as tree size increases. However, even these more preferred models still inferred highly discordant mitochondrial gene trees, thus deepening the mystery around what we label the "Mito-Phylo Paradox" and leading us to ask whether the observed variation could be biological after all.</p>

opencc-zeroMar 2024View details →
zenodo40/100

FIGURE 4 Deduced amino-acid sequences for the Gnrh2 in Differential expression of HPG-axis genes in autotetraploids derived from red crucian carp Carassius auratus red var., × blunt snout bream Megalobrama amblycephala,

FIGURE 4 Deduced amino-acid sequences for the Gnrh2 () and Lhr () genes in Carassius auratus red var. (RCC) and autotetraploid C. auratus red var. ♀ × Megalobrama amblycephala ♂ (4nRR)

opencc-by-4.0Dec 2018View details →
zenodo40/100

FI GU R E 3 Maximum likelihood phylogenetic tree of the Hyalospheniformes with a focus on Apodera, Alocodera, and Padaungiella based on COI gene sequences. Bootstrap values (bs) and Bayesian posterior probabilities (p.p.) are indicated respectively between branches. COI sequences from genera other than Apodera were retrieved from GenBank in Superficially described and ignored for 92 years, rediscovered and emended: Apodera angatakere (Amoebozoa: Arcellinida: Hyalospheniformes) is a new flagship testate amoeba taxon from Aotearoa (New Zealand)

FI GU R E 3 Maximum likelihood phylogenetic tree of the Hyalospheniformes with a focus on Apodera, Alocodera, and Padaungiella based on COI gene sequences. Bootstrap values (bs) and Bayesian posterior probabilities (p.p.) are indicated respectively between branches. COI sequences from genera other than Apodera were retrieved from GenBank

opencc-by-4.0Aug 2021View details →
dryad40/100

Data from: Pitfalls and pointers: an accessible guide to marker gene amplicon sequencing in ecological applications

<p>Next Generation Sequencing (NGS) is a powerful tool that has been rapidly adopted by many ecologists studying microbial communities. Despite the exciting demonstration of NGS technology as a tool for ecological research, cryptic pitfalls inherent to its use can obscure correct interpretation of NGS data. Here, we provide an accessible overview of a NGS process that uses marker gene amplicon sequences (MGAS) that will allow scientists, particularly community ecologists, to make appropriate methodological choices and understand limits on inference about community composition and diversity that can be drawn from MGAS data.</p> <p>We describe the MGAS pipeline, focusing specifically on cryptic sources of variation that have received less emphasis in the ecological literature, but which may substantially impact inference about microbial community diversity and composition. By simulating communities from published microbiome data, we demonstrate how these sources of variation can generate inaccurate or misleading patterns.</p> <p>We specifically highlight sample dilution without researcher awareness and lane-to-lane variability, two cryptic sources of variation arising during the MGAS pipeline. These sources of variation affect estimates of species presence and relative abundance, particularly for species with moderate to low abundances. Each of these sources of bias can lead to errors in the estimation of both absolute and relative abundance within, and turnover among, microbial communities.</p> <p>Awareness and understanding of what happens and, specifically, why it happens during MGAS generation is key to generating a strong data set and building a robust community matrix. Requesting sample dilution information from the sequencing center, including technical replicates across sequencing lanes, and understanding how sampling intensity and community taxa distribution patterns shape the measurement of community richness, evenness, and diversity are critical for drawing correct ecological inferences using MGAS data.</p>

opencc-zeroNov 2021View details →
zenodo40/100

Fig. 17. Phylogenetic hypothesis using nuclear gene sequences TMO-4C4 and 18S in A New Genus and Species of Pygmy Pipehorse from Taitokerau Northland, Aotearoa New Zealand, with a Redescription of Acentronura Kaup, 1853 and Idiotropiscis Whitley, 1947 (Teleostei, Syngnathidae)

Fig. 17. Phylogenetic hypothesis using nuclear gene sequences TMO-4C4 and 18S retrieved with Maximum Likelihood (ML), Maximum Parsimony (MP), and Bayesian Inference (MrBayes), representing 17 species from clade 6 from the analysis of Hamilton et al. (2017) and the new taxon. Tree rooted with the southern Australian trunk-brooder pipefish Heraldia nocturna. Nodal support at the generic level is shown in ML/MP/MrBayes order. See Data Accessibility for tree file.

opencc-by-4.0Sep 2021View details →
dryad40/100

Data from: Restriction site-associated DNA sequencing reveals local adaptation despite high levels of gene flow in Sardinella lemuru (Bleeker, 1853) along the northern coast of Mindanao, Philippines

<p>Stock identification and delineation are important in the management and conservation of marine resources. These were highlighted as priority research areas for Bali sardinella (<em>Sardinella lemuru</em>) which is among the most commercially important fishery resources in the Philippines. Previous studies have already assessed the stocks of <em>S. lemuru</em> between Northern Mindanao Region (NMR) and Northern Zamboanga Peninsula (NZP), yielding conflicting results. Phenotypic variation suggests distinct stocks between the two regions, while mitochondrial DNA did not detect evidence of genetic differentiation for this high gene flow species. This paper tested the hypothesis of regional structuring using genome-wide single nucleotide polymorphisms (SNPs) acquired through restriction-site associated DNA sequencing (RADseq). We examined patterns of population genomic structure using a full panel of 3,573 loci, which was then partitioned into a neutral panel of 3,348 loci and an outlier panel of 31 loci. Similar inferences were obtained from the full and neutral panels, which were contrary to the inferences from the outlier panel. While the full and neutral panels suggested a panmictic population (global F<sub>ST</sub> ~ 0, p &gt; 0.05), the outlier panel revealed genetic differentiation between the two regions (global F<sub>ST</sub> = 0.161, p = 0.001; F<sub>CT</sub> = 0.263, p &lt; 0.05). This indicated that while gene flow is apparent, selective forces due to environmental heterogeneity between the two regions play a role in maintaining adaptive variation. Annotation of the outlier loci returned five genes that were mostly involved in organismal development. Meanwhile, three unannotated loci had allele frequencies that correlated with sea surface temperature. Overall, our results provided support for local adaptation despite high levels of gene flow in <em>S. lemuru</em>. Management therefore should not only focus on demographic parameters (e.g., stock size, catch volume), but also consider the preservation of adaptive variation.</p>

opencc-zeroFeb 2022View details →
zenodo40/100

Raw gene sequences

<p>Raw sequences used for the phylogenetic analysis&nbsp;</p>

opencc-by-4.0Feb 2022View details →
zenodo40/100

Aligned gene sequences

<p>Aligned gene sequences used for the phylogenetic analysis</p>

opencc-by-4.0Feb 2022View details →
zenodo40/100

Trimmed gene sequences

<p>Trimmed gene sequences used for phylogenetic analysis</p>

opencc-by-4.0Feb 2022View details →
zenodo40/100

Illumina sequencing data of Agro-mediated gene-edited apple lines

<p>FastaQ pair-end Illumina sequencing data of the Dipm1/4, Hipm1, and Mlo19 genes&nbsp;from the different apple Agro-edited lines obtained in the project. GA = Gala; GD = Golden Delicious</p> <p>Lines included are:</p> <p>GA1, GA3, GA4, GA5, and GA WT</p> <p>GD1, 2, 6, 10, 11, 12, 15, 17, 18, 19, 20, 21, 23, 24, 26, 27, 29, 31, 33, 36, 37, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, and WT</p>

opencc-by-4.0Apr 2022View details →
zenodo40/100

MGBC-26640: nucleotide sequences for gene annotations

<p>Nucleotide sequences of annotated genes from the 26,640 high-quality, non-redundant genomes of the MGBC.</p>

opencc-by-4.0Apr 2022View details →
zenodo40/100

Fig. 3 in Application Of Dna Barcoding In Taxonomy And Phylogeny: An Individual Case Of Coi Partial Gene Sequencing From Seven Animal Species

Fig. 3. Phylogenetic position of Macrobiotus sp., Bayesian inference phylogenetic tree. Sequences obtained by us are written in bold.

opencc-by-4.0Sep 2019View details →
zenodo40/100

Fig. 1 in Application Of Dna Barcoding In Taxonomy And Phylogeny: An Individual Case Of Coi Partial Gene Sequencing From Seven Animal Species

Fig. 1. Phylogenetic position of D. lindholmi and L. a. exigua, Bayesian inference phylogenetic tree. Sequences obtained by us are written in bold.

opencc-by-4.0Sep 2019View details →
zenodo40/100

Insertion sequences and other mobile elements associated with antibiotic resistance genes in Enterococcus isolates from an inpatient with prolonged bacteremia.

<p>Insertion sequences (ISs) and other transposable elements are associated with the mobilization of antibiotic resistance determinants and the modulation of pathogenic characteristics. In this work, we aimed to investigate the association between ISs and antibiotic resistance genes, and their role in dissemination and modification of the antibiotic resistant phenotype. To that end, we leveraged fully resolved <em>Enterococcus faecium</em> and <em>Enterococcus faecalis</em> genomes of isolates collected over five&nbsp;days from an inpatient with prolonged bacteremia. Isolates from both species harbored similar IS family content but showed significant species-dependent differences in copy number and arrangements of ISs throughout their replicons. Here, we describe two inter-specific IS-mediated recombination events and IS-mediated excision events in plasmids of <em>E. faecium</em> isolates. We also characterize a novel arrangement of the ISs in a Tn1546-like transposon in <em>E. faecalis</em> isolates likely implicated in a vancomycin genotype-phenotype discrepancy. Furthermore, an extended analysis revealed a novel association between daptomycin resistance mutations in <em>liaSR</em> genes and a putative composite transposon in<em> E. faecium</em>, offering a new paradigm for the study of daptomycin resistance and novel insights into the dissemination of daptomycin resistance. In conclusion, our study highlights the role ISs and other transposable elements play in the rapid adaptation and response to clinically relevant stresses such as aggressive antibiotic treatment in enterococci.</p>

opencc-by-4.0Mar 2022View details →
dryad40/100

Joint analysis of microsatellites and flanking sequences enlightens complex demographic history of interspecific gene flow and vicariance in rear-edge oak populations

<p><span>Inference of recent population divergence requires fast evolving markers and necessitates to differentiate shared genetic variation caused by ancestral polymorphism and gene flow. Theoretical research shows that the use of compound marker systems integrating linked polymorphisms with different mutational dynamics, such as a microsatellite and its flanking sequences, can improve estimation of population structure and inference of demographic history, especially in the case of complex population dynamics. However, empirical application in natural populations has so far been limited by lack of suitable methods for data collection. A solution comes from the development of sequence-based microsatellite genotyping which we used to study molecular variation at 36 sequenced nuclear microsatellites in seven <em>Quercus canariensis</em> and four <em>Q. faginea</em> rear-edge populations across Algeria. We aim to decipher their taxonomic relationship, past evolutionary history and recent demographic trajectory. First, we compare the estimation of population genetics parameters and simulation-based inference of demographic history from microsatellite sequence alone, flanking sequence alone or the combination of linked microsatellite and flanking sequence variation. Second, we apply random forest approximate Bayesian computation to identify which of these sequence types is most informative. Whereas analysing microsatellite variation alone indicates recent interspecific gene flow, additional information gained by integrating nucleotide variation in flanking sequences, by reducing homoplasy, suggests ancient interspecific gene flow followed by drift in isolation instead. The weight of each polymorphism in the inference also demonstrates the value of linked variations with contrasted mutation dynamic to improve estimation of both demographic and mutational parameters.</span></p>

opencc-zeroJun 2022View details →
zenodo40/100

Targeted Gene Panel Sequencing Data of RELN paper - Meyer Children's Hospital IRCCS

<h3>Dataset description</h3> <p>&nbsp;</p> <p>The dataset has been prepared according the Minimal Information about a high throughput SEQuencing Experiment (MINSEQE) as reported in: <a href="https://doi.org/10.5281/zenodo.5706412">https://doi.org/10.5281/zenodo.5706412</a></p> <p>This dataset includes:</p> <ul> <li>The Targeted Gene Panel Sequencing Raw Data (FASTQ files) from two individuals harbouring RELN variants</li> <li>The &lsquo;final&rsquo; processed data,&nbsp;submitted both as VCF and TXT files, and obtained from the ANNOVAR annotations of the two patients</li> </ul> <p>The gene panel list used in the targeted capture and the essential experimental and data processing protocols has been reported in the RELN paper.</p> <h3>Identifiers</h3> <p>The 444D indentifier correspond to&nbsp;<strong>DN1 patient</strong> in the RELN paper.</p> <p>Tissue: peripheral blood sample</p> <p>Sex: female</p> <p>Age at sequencing: 21 years</p> <p>The 528T indentifier correspond to <strong>DN2 patient </strong>in the RELN paper.</p> <p>Tissue: peripheral blood sample</p> <p>Sex: female</p> <p>Age at sequencing: 1.5 years</p>

opencc-by-4.0May 2024View details →
zenodo40/100

Fig. 9 in Constraints on Phylogenetic Interrelationships among Four Free-living Litostomatean Lineages Inferred from 18S rRNA gene-ITS Region sequences and Secondary Structure of the ITS2 molecule

Fig. 9. Evolutionary hypothesis of interrelationships among the four free-living litostomatean lineages studied. This scenario was suggested on the basis of morphology and the consensus secondary structure of the ITS2 molecules. CK – circumoral kinety, DB – dorsal brush, OB – oral bulge, OO – oral bulge opening, P – proboscis, PE – perioral kinety, PR – preoral kineties, SK – somatic kineties.

opencc-by-4.0Dec 2017View details →
zenodo40/100

Fig. 5 in Constraints on Phylogenetic Interrelationships among Four Free-living Litostomatean Lineages Inferred from 18S rRNA gene-ITS Region sequences and Secondary Structure of the ITS2 molecule

Fig. 5. Quartet likelihood-mapping showing distribution of phylogenetic signal in the 18S-A and the CON-1 alignment for three possible relationships among the four main free-living litostomatean lineages studied. The corners of the triangles show the percentage of fully resolved trees, i.e., phylogenetically informative signal. The rectangular areas show the percentage of trees that are in conflict. The central triangle shows the percentage of unresolved star-like trees, i.e., phylogenetically uninformative signal. Coding of free-living litostomatean lineages: H – Haptorida, P – Pleurostomatida, R – Rhynchostomatia, S – Spathidiida.

opencc-by-4.0Dec 2017View details →
zenodo40/100

Fig. 4 in Constraints on Phylogenetic Interrelationships among Four Free-living Litostomatean Lineages Inferred from 18S rRNA gene-ITS Region sequences and Secondary Structure of the ITS2 molecule

Fig. 4. Super-network of 66 free-living litostomatean taxa constructed from 80 randomly selected post-burn-in trees from the Bayesian inference of the 18S-A–D, ITSR-C and ITSR-D as well as the CON-1 and CON-2 alignments. The super-network was constructed in the program SplitsTree, using the Z-closure option, tree size weighted mean, ten runs, and the refined heuristic technique. For details on taxa and characteristics of the alignments analyzed, see Supplementary Table S1 and S2.

opencc-by-4.0Dec 2017View details →
zenodo40/100

Fig. 3 in Constraints on Phylogenetic Interrelationships among Four Free-living Litostomatean Lineages Inferred from 18S rRNA gene-ITS Region sequences and Secondary Structure of the ITS2 molecule

Fig. 3. Phylogeny based on the 18S rRNA gene and the ITS1-5.8S-ITS2 region of 56 free-living litostomatean taxa (alignment CON-1). Posterior probabilities for the Bayesian inference and bootstrap values for maximum likelihood were mapped onto the 50% majority rule ML tree. Dashes indicate posterior probabilities below 0.50 and ML bootstrap values below 50%. The scale bar indicates five substitutions per ten nucleotide positions. For details on taxa, evolutionary model used, and characteristics of the CON-1 alignment, see Supplementary Table S1 and S2.

opencc-by-4.0Dec 2017View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record