Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,293

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,293 results for “gene sequencing”

Learn how ShareScore rates datasets ↗
dryad36/100

DNA sequence data generated using non-invasive feather and eggshell samples from the Grenada Dove for two gene regions: Cyt b and ND2

<p>As an island endemic with a decreasing population, the Critically Endangered Grenada Dove <em>Leptotila wellsi</em> is threatened by accelerated loss of genetic diversity resulting from ongoing habitat fragmentation. Small, threatened populations are difficult to sample directly but advances in molecular methods mean that non-invasive samples can be used. We performed the first assessment of genetic diversity of populations of Grenada Dove by a) assessing mtDNA genetic diversity in the only two areas of occupancy on Grenada, b) defining the number of haplotypes present at each site and c) evaluating evidence of isolation between sites. We used non-invasively collected samples from two locations: Mt Hartman (n=18) and Perseverance (n=12). DNA extraction and PCR were used to amplify 1,751 bps of mtDNA from two mitochondrial markers: NADH dehydrogenase 2 (<em>ND2</em>) and Cytochrome b (<em>Cyt b</em>). Haplotype diversity (<em>h</em>) of 0.4, a nucleotide diversity (π) of 0.00023 and two unique haplotypes were identified within the <em>ND2</em> sequences; a single haplotype was identified within the <em>Cyt b </em>sequences. Of the two haplotypes identified; the most common haplotype (haplotype A = 73.9%) was observed at both sites and the other (haplotype B = 26.1%) was unique to Perseverance. Our results show low mitochondrial genetic diversity and clear evidence for genetically isolated populations. The Grenada Dove needs urgent conservation action, including habitat protection and potential augmentation of gene flow by translocation in order to increase genetic resilience and diversity with the ultimate aim of securing the long-term survival of this Critically Endangered species. </p>

opencc-zeroNov 2023View details →
zenodo36/100

Data required for "Low mutation rate of spontaneous mutants enables detection of causative genes by comparing whole genome sequences"

<p>In the early 1900s,mutation breeding to select varieties with desirable traits using &nbsp;spontaneous mutation was actively conducted around the world, including Japan. In rice, the number of fixed mutations per generation was estimated to be 1.38-2.25. Although this low mutation rate was a major problem for breeding in those days, in the modern era with the development of NGS technology, it was conversely considered to be an advantage for efficient gene identification. In this paper, we proposed an in silico approach using next-generation sequencing (NGS) to compare the whole genome sequence of a spontaneous mutant with that of a closely related strain with a nearly identical genome, to find polymorphisms that differ between them, and to identify the causal gene by predicting the functional variation of the gene caused by the polymorphism. Using this approach, we found four causal genes for the dwarf mutation, the round shape grain mutation and the awnless mutation. Three of these genes were the same as those previously reported, but one was a novel gene involved in awn formation. The novel gene was isolated from Bozu-Aikoku, a mutant of Aikoku with the awnless trait, in which nine polymorphisms were predicted to alter gene function by their whole-genome comparison. Based on the information on gene function and tissue-specific expression patterns of these candidate genes, Os03g0115700/LOC_Os03g02460, annotated as a shortchain dehydrogenase/reductase SDR family protein, is most likely to be involved in the awnless mutation. Indeed, complementation tests by transformation showed that it is involved in awn formation. Thus, this method is an effective way to accelerate genome breeding of various crop species by enabling the identification of useful genes that can be used for crop breeding with minimal effort for NGS analysis.</p>

opencc-by-4.0Mar 2024View details →
dryad36/100

Sequence-dependent model of genes with dual σ factor preference

<p class="MsoNormal"><em>Escherichia coli</em> uses <span>s</span> factors to quickly control large gene cohorts during stress conditions. While most of its genes respond to a single <span>s</span> factor, approximately 5% of them have dual <span>s</span> factor preference. The most common are those responsive to both <span>s</span><sup>70</sup>, which controls housekeeping genes, and <span>s</span><sup>38</sup>, which activates genes during stationary growth and stresses. Using RNA-seq and flow-cytometry measurements, we show that 'σ<sup>70+38</sup> genes' are nearly as upregulated in stationary growth as 'σ<sup>38</sup> genes'. Moreover, we find a clear quantitative relationship between their promoter sequence and their response strength to changes in σ<sup>38</sup> levels. We then propose and validate a sequence dependent model of σ<sup>70+38</sup> genes, with dual sensitivity to <span>s</span><sup>38 </sup>and <span>s</span><sup>70</sup>, that is applicable in the exponential and stationary growth phases, as well in the transient period in between. We further propose a general model, applicable to other stresses and σ factor combinations. Given this, promoters controlling σ<sup>70+38</sup> genes (and variants) could become important building blocks of synthetic circuits with predictable, sequence-dependent sensitivity to transitions between the exponential and stationary growth phases.</p>

opencc-zeroApr 2022View details →
zenodo36/100

Supplementary material (Sanger sequencing photo) of the case of THRB mosaicism presented in Clinically Symptomatic Resistance to Thyroid Hormone β Syndrome Because of THRB Gene Mosaicism of Donnars et al. (https://doi.org/10.1210/clinem/dgac347)

<p>Supplementary material (Sanger photo) of the case of THRB mosaicism presented in &quot;Clinically Symptomatic Resistance to Thyroid Hormone &beta; Syndrome Because of THRB Gene Mosaicism&quot;, by Donnars et al. (https://doi.org/10.1210/clinem/dgac347)</p>

opencc-by-4.0Jun 2022View details →
dryad36/100

Long-read-based draft genome sequence of Indian black gram IPU-94-1 'Uttara': Insights into disease resistance and seed storage protein genes

<p>Black gram [Vigna mungo (L.) Hepper var. <em>mung<a>o</a></em>] [LAV1] is a warm-season legume highly prized for its protein content along with significant folate and iron proportions. To expedite the genetic enhancement of black gram, a high-quality draft genome from the center of origin of the crop is indispensable. Here, we established a draft genome sequence of an Indian black gram cultivar, 'Uttara' (IPU 94-1), known for its high resistance to mungbean yellow mosaic virus. Pacific Biosciences of California, Inc. (PacBio) single-molecule real-time (SMRT) and Illumina sequencing assembled a draft reference-guided assembly with a cumulative size of ~454.4 Mb, of which, 444.4 Mb was anchored on 11 pseudomolecules corresponding to 11 chromosomes. Uttara assembly denotes features of a high-quality draft genome illustrated through high N50 value (42.88 Mb), gene completeness (benchmarking universal single-copy ortholog [BUSCO] score 94.17%), and low levels of ambiguous nucleotides (N) percent (0.0005%). Gene discovery using transcript evidence predicted 28,881 protein-coding genes, from which, ~95% were functionally annotated. A global survey of genes associated with disease resistance revealed 119 nucleotide binding site–leucine rich repeat (NBS-LRR) proteins, while 23 genes encoding seed storage proteins (SSPs) were discovered in black gram. A large set of microsatellite loci were discovered for marker development in the crop. Our draft genome of an Indian black gram provides the foundational genomic resources for the improvement of important agronomic traits and ultimately will help in accelerating black gram breeding programs.</p>

opencc-zeroJul 2023View details →
zenodo36/100

Structural variants in the barley gene pool: precision and sensitivity to detect them using short-read sequencing and their association with gene expression and phenotypic variation

<p>SNV of 23 parental barley inbreds of the double round robin population (DRR) (<a href="https://doi.org/10.1111/pbi.13746">https://doi.org/10.1111/pbi.13746</a>) used in the publication &quot;Structural variants in the barley gene pool: precision and sensitivity to detect them using short-read sequencing and their association with gene expression and phenotypic variation&quot;. SV, INDELs, and additional data are available via figshare (https://doi.org/10.6084/m9.figshare.16802473).</p>

opencc-by-4.0Apr 2022View details →
dryad36/100

Cytb gene sequences of Fejervarya species from Lesser Sunda, Indonesia and other Asian countries

<p>Cyt b gene sequences of Fejervarya species were done and later submitted to DDBJ. Then, we received accession number which are used in our manuscript "Postmating isolation and evolutionary relationships among Fejervarya species from Lesser Sunda, Indonesia and other Asian countries revealed by crossing experiments and mtDNA Cytb sequences analyses". These are genuine data. There is no conflict of interest of these data.</p>

opencc-zeroOct 2022View details →
zenodo36/100

Supplementary material 1 from: Scacchetti P, Pansonato-Alves J, Utsunomia R, Oliveira C, Foresti F (2011) Karyotypic diversity in four species of the genus Gymnotus Linnaeus, 1758 (Teleostei, Gymnotiformes, Gymnotidae): physical mapping of ribosomal genes and telomeric sequences. Comparative Cytogenetics 5(3): 223-235. https://doi.org/10.3897/compcytogen.v5i3.1375

Nexus file of aligned COI and COII nucleotide sequences.

opencc-by-4.0Aug 2011View details →
zenodo36/100

Gene and protein sequence features augment HLA class I ligand predictions

<p>Dataset and analyses supporting the manuscript "Gene and protein sequence features augment HLA class I ligand predictions".</p> <p>The "peptides" files contain the mass-spec detected peptides obtained from HLA ligandomics performed on the indicated tumor lines.&nbsp;</p> <p>The "protein data" files contain the RNAseq data (TPM) and Ribosome profiling data (ribosome occupancy) per protein, for each tumor line.&nbsp;</p> <p>The "source data" zip archive contains the source data underlying the figures of the manuscript.</p> <p>The "HLA ligandome analyses" zip archive contains the R scripts used for all data analysis in the manuscript, including all data and output files. These analyses can also be found at https://github.com/kasbress/HLA_Ligandome_Analyses/</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0May 2024View details →
dryad36/100

Selection pressure analysis of dengue virus complete genome and E gene nucleotide sequences from Pakistan

<p>This dataset comprises 43 E gene and 44 complete genome nucleotide sequences of the dengue virus from serotypes DENV-1 to DENV-4, representing all documented sequences in Pakistan to date, sourced from the Virus Pathogen Resource (ViPR) database and NCBI. The E gene is critical as it is involved in serotype changes of the dengue virus, making it a pivotal target for understanding shifts in viral pathogenicity and immune escape mechanisms. The aim of compiling this dataset is to facilitate comprehensive genetic analysis and enhance understanding of the evolutionary dynamics of the dengue virus within the region. To assess the evolutionary pressures acting on these sequences, we conducted a selection pressure analysis utilizing computational methods. These methods include the Single Likelihood Ancestor Counting (SLAC), Fixed Effects Likelihood (FEL), adaptive Branch Site Random Effects Likelihood (aBSREL), Mixed Effects Model of Evolution (MEME), and the Genetic Algorithm for Recombination Detection (GARD), all implemented in the HyPhy software package. Our analysis focused on identifying genomic sites under both positive and negative selection pressures, providing insights into the adaptive evolutionary processes affecting the E gene of the dengue virus in Pakistan. Understanding the molecular evolution of this gene is crucial for predicting serotype evolution, potentially aiding in the development of effective vaccines and therapeutic strategies.</p>

opencc-zeroMay 2024View details →
zenodo36/100

Nucleotide sequence database of Copper-containing membrane monooxygenases genes for analysing primer pairs targeting the ammonia monooxygenase subunit A gene of complete ammonia oxidising Nitrospira

<p>Nucleotide sequences of 487 Cu-mmo genes, including amoA comammox clade A and clade B, amoA ammonia oxidizing bacteria as well as other Cu-mmo genes.</p>

opencc-by-4.0Dec 2023View details →
zenodo36/100

Fig. 2 in Molecular phylogeny of Indonesian Lymantria Tussock Moths (Lepidoptera: Erebidae) based on CO I gene sequences

Fig. 2. Pairwise sequences divergence based on K2P model versus Transition/Transversion (Ts/Tv).

opencc-by-4.0Feb 2014View details →
zenodo36/100

Fig. 1 in Mitochondrial DNA diversity in the acanthocephalan Prosthenorchis elegans in Colombia based on cytochrome c oxidase I (COI) gene sequence

Fig. 1. Photo showing the characteristic external morphology of Prosthenorchis elegans.

opencc-by-4.0Dec 2015View details →
zenodo36/100

Fig. 2 in Molecular phylogeny of Indonesian Zeuzera (Lepidoptera: Cossidae) wood borer moths based on CO I gene sequence

Fig. 2. Scatter plots of K2P model distance for Transition (Ts) versus Transversion (Tv).

opencc-by-4.0Feb 2015View details →
zenodo36/100

Genus level DNA sequence data for three genes (matK, rbcL, trnH-psbA) for the paper: A comprehensive, genus-level time-calibrated phylogeny of the tree flora of Mediterranean Europe and an assessment of its vulnerability

<p>This data file contains the consensus DNA sequences in fasta format, of 64 tree genera found in Mediterranean Europe, following the checklist of M&eacute;dail et al. (2019).&nbsp;</p> <p>The data are used in a manuscript submitted for publication to Botany Letters and currently under revision. The manuscript is entitled: &quot;<em>A comprehensive, genus-level time-calibrated phylogeny of the tree flora of Mediterranean Europe and an assessment of its vulnerability</em>&quot;. Its authors are: Marwan Cheikh Albassatneh, Marcial Escudero, Loic Ponge<sup>*</sup>, Anne-Christine Monnet, Juan Arroyo, Toni Nikolic, Gianluigi Bacchetta, Francesca Bagnoli, Panayotis Dimopoulos, Agathe Leriche, Fr&eacute;d&eacute;ric M&eacute;dail, Anne Roig, Ilaria Spanu, Giovanni Giuseppe Vendramin, Arndt Hampe, Bruno Fady.</p>

opencc-by-4.0Sep 2019View details →
zenodo36/100

Treponema pallidum Cytoplasmic filament protein gene partial sequence

<p>A partial sequence of the &nbsp;Cytoplasmic filament protein gene (cfpA)&nbsp; of treponema pallidum from non human primates</p>

opencc-by-4.0Nov 2019View details →
zenodo36/100

Saved model and preprocessed data for "CRMnet:a deep learning model for predicting gene expression from large regulatory sequence datasets"

<p>Saved TUNet model and preprocessed training data&nbsp;for &quot;CRMnet: a deep learning model for predicting&nbsp;gene expression from large regulatory&nbsp;sequence datasets&quot;</p> <p>To load the trained model:</p> <pre><code class="language-python">import tensorflow as tf tf.keras.models.load_model("path to the model folder")</code></pre> <p>for more information please find our repository:&nbsp;https://github.com/jiayuwen/CRMnet</p>

opencc-by-4.0Nov 2022View details →
dryad36/100

Data from: Higher evolutionary dynamics of gene copy number for Drosophila glue genes located near short repeat sequences

<p><strong>Background</strong></p> <p>During evolution, genes can experience duplications, losses, inversions and gene conversions. Why certain genes are more dynamic than others is poorly understood. Here we examine how several <em>Sgs</em> genes encoding glue proteins, which make up a bioadhesive that sticks the animal during metamorphosis, have evolved in <em>Drosophila</em> species.</p> <p><strong>Results</strong></p> <p>We examined high-quality genome assemblies of 24 <em>Drosophila</em> species to study the evolutionary dynamics of four glue genes that are present in <em>D. melanogaster</em> and are part of the same gene family <em>–</em> <em>Sgs1, Sgs3, Sgs7 and Sgs8 –</em> across approximately 30 millions of years. We annotated a total of 102 <em>Sgs</em> genes and grouped them into 4 subfamilies. We present here a new nomenclature for these <em>Sgs</em> genes based on protein sequence conservation, genomic location and presence/absence of internal repeats. Two types of glue genes were uncovered. The first category (<em>Sgs1, Sgs3x, Sgs3e</em>) showed a few gene losses but no duplication, no local inversion and no gene conversion. The second group (<em>Sgs3b, Sgs7, Sgs8</em>) exhibited multiple events of gene losses, gene duplications, local inversions and gene conversions. Our data suggest that the presence of short "new glue" genes near the genes of the latter group may have accelerated their dynamics.</p> <p><strong>Conclusions</strong></p> <p>Our comparative analysis suggests that the evolutionary dynamics of glue genes is influenced by genomic context. Our molecular, phylogenetic and comparative analysis of the four glue genes <em>Sgs1, Sgs3, Sgs7</em> and <em>Sgs8 </em>provides the foundation for investigating the role of the various glue genes during <em>Drosophila</em> life.</p>

opencc-zeroJan 2023View details →
zenodo36/100

Sequence diversity in MAX effectors and other genes in 120 isolates of the rice blast fungus Magnaporthe oryzae

<p>- list_of_accessions_and_assembly_statistics.xlsx: list of 120 isolate and genome assembly statistics</p> <p>- assemblies.zip: genome assemblies with repeats were not masked</p> <p>- orthogroups.txt: list of orthogroups in orthogroups.zip</p> <p>-orthogroups.zip: sequences of orthogroups, as identified using Orthofinder. Sequences were aligned using translatorX (https://doi.org/10.1093/nar/gkq291)</p> <p>- single_copy_orthologs.zip: folder which contains aligned sequences of single-copy orthologs (alignment with translatorX&nbsp;https://doi.org/10.1093/nar/gkq291); three types of genes were distinguished: MAX effectors, other secreted proteins, and other genes;&nbsp;note that to produce this dataset, the 11 orthogroups that included paralogous copies of MAX effectors were split into sets of orthologous sequences using genealogies inferred using RAXML v8, yielding a total of 94 single-copy MAX orthologs; for each split orthogroup, sets of orthologous sequences were assigned a number that was added to the orthogroup&rsquo;s identifier as a suffix (for instance paralogous sequences of orthogroup OG0000244 were split into orthogroups OG0000244_1 and OG0000244_2)</p>

opencc-by-4.0Feb 2023View details →
dryad36/100

Dataset from: The origin and fate of fungal mitochondrial horizontal gene transferred sequences in orchids (Orchidaceae)

<p>The transfer of DNA among distantly related organisms is relatively common in bacteria but less prevalent in eukaryotes. Among fungi and plants, few of these events have been reported. Two segments of fungal mitochondrial DNA have been discovered in the mitogenome of orchids. Here, we build on their work to understand the timing of those transfer events, which orchids retain the fungal DNA, and the fate of the foreign DNA during orchid evolution. We update the content of the large DNA fragment and establish that it was transferred to the most recent common ancestor of a highly diverse clade of epidendroid orchids that lived ~28–43 Mya. Also, we present hypotheses of the origin of the small transferred fragment. Our findings deepen the knowledge of these interesting DNA transfers among organelles and we formulate a probable mechanism for these horizontal gene transfer events.</p>

opencc-zeroJun 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record