Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

232

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

232 results for “centromeres”

Learn how ShareScore rates datasets ↗
zenodo40/100

Chromosome images used for "Centromere detection of human metaphase chromosome images using a candidate based method"

<p>Chromosome image data used in the paper &quot;Centromere detection of human metaphase chromosome images using a candidate based method&quot;. Images are in tiff format</p>

opengpl-2.0Jun 2016View details →
zenodo40/100

Chromosome images used for "Centromere detection of human metaphase chromosome images using a candidate based method"

<p>Chromosome image data used in the paper &quot;Centromere detection of human metaphase chromosome images using a candidate based method&quot;. Images are in tiff format.</p>

opengpl-2.0Jun 2016View details →
zenodo40/100

Regional Centromere Configuration in the Fungal Pathogens of Pneumocystis Genus

<p>Supplementary material&nbsp;</p>

opencc-by-4.0Jan 2024View details →
zenodo40/100

The structure of simple satellite variation in the human genome and its correlation with centromere ancestry (Supplemental Data)

<p>Accompanying <a href="https://github.com/is-the-biologist/1KGP_SATS" target="_blank" rel="noopener">Github</a></p> <p><strong>Supplemental File 1.</strong> BLAST results of k-mer concatemers against T2T-CHM13-v2.0.</p> <p><strong>Supplemental File 2.</strong> Annotations of centromeres, and telomeres of T2T-CHM13-v2.20. Table of abundance of k-mers in annotated regions as numpy file from BLAST hits. Abundance of k-mers across genome in 100kb bins from BLAST hits as .npz files accessible by example:</p> <p>&nbsp; &nbsp; import numpy as np<br>&nbsp; &nbsp; dense = np.load("filename.npz")<br>&nbsp; &nbsp; dense["chr1"]<br>&nbsp;&nbsp;<br><strong>Supplemental File 3</strong>. Table of pairwise R2 between simple satellites and table of pairwise interspersion OR between simple satellites. Folder containing QQ plots of negative binomial fit of satellite copy number distribution used to qualitatively asses model fit.</p> <p><strong>Supplemental File 4. </strong>Materials and results of cenGRM analysis. Boundaries used for centromeric regions of each cenGRM, cenGRMs in GCTA format, and tables with the results of cenGRM GCTA runs. Also provide pdfs of the dendrograms/heatmaps produced from UPGMA clustering of each cenGRM.&nbsp;</p> <p><strong>Supplemental File 5</strong>&nbsp;Non-human significant BLAST hits from BLAST-ing k-mer concatamers to non-human sequences.</p> <p><strong>Supplemental Table 1.</strong> Copy number normalized to 1x depth given GC bias of 126 most abundant satellites analyzed in paper in each individual. Additional columns represent metadata of the individual:</p> <ul> <li>instrument: sequencer instrument name used to sequence library.</li> <li>run: sequencer run of the library.</li> <li>flow: flowcell ID of the ibrary.</li> <li>pop: 1,000 Genomes Project population ID.</li> <li>superpop: 1,000 Genomes Project superpopulation ID.</li> <li>reads: average autosomal read depth of the library.</li> </ul> <p><strong>Supplemental Table 2. </strong>Copy number normalized to 1x depth given GC bias of the top 126 most abundant satellites analyzed in paper in each individual of the 1KGP, plus estimates of the same satellites in CHM13 short-read libraries subsampled from 18x-0.5x, 18x depth simulated library of the T2T-CHM13v2.0 assembly analyzed using k-Seek, and Tandem Repeat Finder results of the T2T-CHM13v2.0 asembly <a href="https://doi.org/10.1126/science.abk3112" target="_blank" rel="noopener">Hoyt 2022</a>.</p> <p><strong>Supplemental Table 3.</strong> Copy number normalized to 1x depth given GC bias of all tandem repeats with k-mer &lt;= 20 (6,309) found collectively in the CHM13 short-read libraries subsampled from 18x-0.5x, 18x depth simulated library of the T2T-CHM13v2.0 assembly analyzed using k-Seek, and Tandem Repeat Finder results of the T2T-CHM13v2.0 asembly <a href="https://doi.org/10.1126/science.abk3112" target="_blank" rel="noopener">Hoyt 2022</a>.</p>

opencc-by-4.0Mar 2024View details →
dryad40/100

Turnover of retroelements and satellite DNA drives centromere reorganization over short evolutionary timescales in Drosophila

Open the record for dataset details and reuse information.

publicNov 2024View details →
dryad36/100

Data from: Centromere innovations within a mouse species

<p class="MsoNormal">Mammalian centromeres direct faithful genetic inheritance and are typically characterized by regions of highly repetitive and rapidly evolving DNA. We focused on a mouse species, <em>Mus pahari, </em>that we found has evolved to house centromere-specifying CENP-A nucleosomes at the nexus of a satellite repeat that we identified and term p-satellite (p-sat), a small number of recruitment sites for CENP-B, and short stretches of perfect telomere repeats. One <em>M. pahari</em> chromosome, however, houses a radically divergent centromere harboring ~6 Mbp of a homogenized p-sat-related repeat, p-sat<sup>B</sup>, that contains &gt;20,000 functional CENP-B boxes. There, CENP-B abundance drives accumulation of microtubule-binding components of the kinetochore, as well as a microtubule-destabilizing kinesin of the inner centromere. The balance of pro- and anti-microtubule-binding by the new centromere permits it to segregate during cell division with high fidelity alongside the older ones whose sequence creates a markedly different molecular composition.</p>

opencc-zeroNov 2023View details →
dryad36/100

Data from: Repositioning of centromere-associated repeats during karyotype evolution in Oryzias fishes

<p>The karyotype, which is the number and shape of chromosomes, is a fundamental characteristic of all eukaryotes. Karyotypic changes play an important role in many aspects of evolutionary processes, including speciation. In organisms with monocentric chromosomes, it was previously thought that chromosome number changes were mainly caused by centric fusions and fissions, whereas chromosome shape changes, that is changes in arm numbers, were mainly due to pericentric inversions. However, recent genomic and cytogenetic studies have revealed examples of alternative cases, such as tandem fusions and centromere repositioning, found in the karyotypic changes within and between species. Here, we employed comparative genomic approaches to investigate whether centromere repositioning occurred during karyotype evolution in medaka fishes. In the medaka family (Adrianichthyidae), the three phylogenetic groups differed substantially in their karyotypes. The <em>Oryzias latipes</em> species group has larger numbers of chromosome arms than the other groups, with most chromosomes being metacentric. The <em>O. javanicus</em> species group has similar numbers of chromosomes to the <em>O. latipes</em> species group, but smaller arm numbers, with most chromosomes being acrocentric. The <em>O. celebensis</em> species group has fewer chromosomes than the other two groups and several large metacentric chromosomes that were likely formed by chromosomal fusions. By comparing the genome assemblies of <em>O. latipes</em>, <em>O. javanicus</em>, and <em>O. celebensis</em>, we found that repositioning of centromere-associated repeats might be more common than simple pericentric inversion. Our results demonstrated that centromere repositioning may play a more important role in karyotype evolution than previously appreciated.</p>

opencc-zeroNov 2023View details →
zenodo36/100

Comprehensive characterisation of the genomic insertion site of a transgene in highly repetitive, centromeric region of Anopheles mosquitoes

<p>The availability of the genomic sequence of the malaria mosquito <em>Anopheles</em> <em>gambiae</em> has sparked in recent years the development of transgenic technologies with the potential to be used as novel tools for vector control. These technologies rely on genome editing that confers features able to affect vector capacity. This can be achieved by either reducing the mosquito population or by making mosquitoes refractory to the parasite infection. Although sophisticated molecular techniques such as those based on AttB/AttP site-specific recombination and CRISPR/Cas9 systems can lead to the integration of transgenes in specific sites of the genome, methods that allow semi-random integration are still in use due to their high efficiency; PiggyBac transposon-mediated integrations fall in this category. Characterization of the insertion site of transgenes in transgenic strains generated via PiggyBac integration can be hampered when the transgene is inserted in regions of the genome rich in repetitive sequences. Here we describe a number of techniques that were used to identify the genomic location of the transgene in a repetitive region in the <em>Anopheles gambiae</em> strain Ag(PMB)1 which was initially reported on Chromosome 3R 36D. Whilst Inverse PCR used in previous analysis was unable to distinguish between multiple genomic locations as potential insertion sites of the transgene, here we demonstrate that the use of FISH identifies clearly the integration of the transgene in a poorly annotated centromeric region of Chromosome 2R 19D. This study emphasises the need for accuracy in sequencing data for the genome of organisms of medical importance such as <em>Anopheles </em>mosquitoes. An effort to further improve reference genomes is of paramount importance to support and facilitate vector control interventions based on genome editing.</p>

opencc-by-4.0Dec 2021View details →
dryad36/100

Fixed allele differences associated with the centromere reveal chromosome morphology and rearrangements in a reptile (Varanus acanthurus Boulenger)

<p>Chromosome rearrangements are often implicated with genomic divergence and are proposed to be associated with species evolution. Rearrangements alter the genomic structure and interfere with homologous recombination by isolating a portion of the genome. Integration of multi-platform next generation DNA sequencing technologies has enabled putative identification of chromosome rearrangements in many taxa, however, integrating these data sets with cytogenetics is still uncommon beyond model genetic organisms. Therefore, to achieve the ultimate goal for the genomic classification of eukaryotic organisms, physical chromosome mapping remains critical. The ridge-tailed goannas (<em>Varanus</em> <em>acanthurus</em> BOULENGER) are a group of dwarf monitor lizards comprised of several species found throughout Northern Australia. These lizards exhibit extreme divergence at both the genic and chromosomal levels. The chromosome polymorphisms are widespread extending across much of their distribution, raising the question if these polymorphisms are homologous within the <em>V. acanthurus</em> complex. We used a combined genomic and cytogenetic approach to test for homology across divergent populations with morphologically similar chromosome rearrangements. We showed that more than one chromosome pair was involved with the widespread rearrangements. This finding provides evidence to support <em>de novo</em> chromosome rearrangements have occurred within populations. These chromosome rearrangements are characterised by fixed allele differences originating in the vicinity of the centromeric region. We then compared this region with several other assembled genomes of reptiles, chicken and the platypus. We demonstrated that the synteny of genes in chordates remains conserved despite centromere repositioning across these taxa.</p>

opencc-zeroMay 2023View details →
dryad36/100

Data from: Repositioning of centromere-associated repeats during karyotype evolution in Oryzias fishes

Open the record for dataset details and reuse information.

publicNov 2023View details →
dryad36/100

Fixed allele differences associated with the centromere reveal chromosome morphology and rearrangements in a reptile (Varanus acanthurus Boulenger)

Open the record for dataset details and reuse information.

publicMay 2023View details →
dryad36/100

Data from: Evidence of centromeric histone 3 chaperone involved in DNA damage repair pathway in budding yeast

Open the record for dataset details and reuse information.

publicOct 2025View details →
dryad36/100

Selfish chromosomal drive shapes recent centromeric histone evolution in monkeyflowers

Open the record for dataset details and reuse information.

publicMar 2021View details →
dryad36/100

Data from: Functional monocentricity with holocentric characteristics and chromosome-specific centromeres in a stick insect

Open the record for dataset details and reuse information.

publicDec 2024View details →
dryad36/100

Data from: Centromere innovations within a mouse species

Open the record for dataset details and reuse information.

publicNov 2023View details →
dryad32/100

Data from: Comparative analysis of DNA repeats and identification of novel Fesreba centromeric element in fescues and ryagrasses

<p>Background<br> Cultivated grasses are an important source of food for domestic animals worldwide. Better knowledge of their genomes can speed up the development of new cultivars with better quality and resistance to biotic and abiotic stresses. The most widely grown grasses are tetraploid ryegrass species ( Lolium spp.) and diploid and hexaploid fescue species (Festuca spp.). In this work, we characterized repetitive DNA sequences and their contribution to genome size in five fescue and two ryegrass species, as well as one fescue and two ryegrass cultivars.</p> <p>Results<br> Partial genome sequences produced by Illumina technology were used for genome-wide comparative analyses using RepeatExplorer pipeline. Retrotransposons were found to be the most abundant repeat types in all seven grass species. Athila element of Ty3/gypsy family showed the most striking differences in copy number between fescues and ryegrasses. The sequence data enabled the assembly of an LTR element Fesreba, which is highly enriched in centromeric and (peri)centromeric regions in all species. A combination of FISH with a probe specific to Fesreba element and immunostaining with CENH3 antibody showed their colocalization and indicated a possible role of Fesreba in centromere function.</p> <p>Conclusions<br> Comparative repeatome analysis in a set of fescues and ryegrasses provided new insights into their genome organization and divergence, including the assembly of LTR element Fesreba. A new LTR element Fesreba was identified and found abundant in centromeric regions of the fescues and ryegrasses. It may have a role in the function of their centromeres.</p>

opencc-zeroJul 2020View details →
dryad32/100

Data from: Islands of retroelements are major components of Drosophila centromeres

Centromeres are essential chromosomal regions that mediate kinetochore assembly and spindle attachments during cell division. Despite their functional conservation, centromeres are amongst the most rapidly evolving genomic regions and can shape karyotype evolution and speciation across taxa. Although significant progress has been made in identifying centromere-associated proteins, the highly repetitive centromeres of metazoans have been refractory to DNA sequencing and assembly, leaving large gaps in our understanding of their functional organization and evolution. Here, we identify the sequence composition and organization of the centromeres of Drosophila melanogaster by combining long-read sequencing, chromatin immunoprecipitation for the centromeric histone CENP-A, and high-resolution chromatin fiber imaging. Contrary to previous models that heralded satellite repeats as the major functional components, we demonstrate that functional centromeres form on islands of complex DNA sequences enriched in retroelements that are flanked by large arrays of satellite repeats. Each centromere displays distinct size and arrangement of its DNA elements but is similar in composition overall. We discover that a specific retroelement, G2/Jockey-3, is the most highly enriched sequence in CENP-A chromatin and is the only element shared among all centromeres. G2/Jockey-3 is also associated with CENP-A in the sister species Drosophila simulans, revealing an unexpected conservation despite the reported turnover of centromeric satellite DNA. Our work reveals the DNA sequence identity of the active centromeres of a premier model organism and implicates retroelements as conserved features of centromeric DNA.

opencc-zeroDec 2018View details →
dryad32/100

Data from: Evolution of long centromeres in fire ants

Background: Centromeres are essential for accurate chromosome segregation, yet sequence conservation is low even among closely related species. Centromere drive predicts rapid turnover because some centromeric sequences may compete better than others during female meiosis. In addition to sequence composition, longer centromeres may have a transmission advantage. Results: We report the first observations of extremely long centromeres, covering on average 34 % of the chromosomes, in the red imported fire ant Solenopsis invicta. By comparison, cytological examination of Solenopsis geminata revealed typical small centromeric constrictions. Bioinformatics and molecular analyses identified CenSol, the major centromeric satellite DNA repeat. We found that CenSol sequences are very similar between the two species but the CenSol copy number in S. invicta is much greater than that in S. geminata. In addition, centromere expansion in S. invicta is not correlated with the duplication of CenH3. Comparative analyses revealed that several closely related fire ant species also possess long centromeres. Conclusions: Our results are consistent with a model of simple runaway centromere expansion due to centromere drive. We suggest expanded centromeres may be more prevalent in hymenopteran insects, which use haplodiploid sex determination, than previously considered.

opencc-zeroDec 2015View details →
zenodo32/100

De novo assemblies for the manuscrip "Candida albicans isolates contain frequent heterozygous structural variants and transposable elements within genes and centromeres"

Open the record for dataset details and reuse information.

opencc-by-4.0Jun 2024View details →
zenodo32/100

Supporting Data for "Pushing the limits of HiFi assemblies reveals centromere diversity between two Arabidopsis thaliana genomes"

<p>This dataset contains supporting files referenced by the following publication:</p> <p>&bull; Rabanal FA, Gr&auml;ff M, Lanz C, Fritschi K, Llaca V, Lang M, Carbonell-Bejerano P, Henderson I, Weigel D. <strong>Pushing the limits of HiFi assemblies reveals centromere diversity between two&nbsp;<em>Arabidopsis thaliana</em>&nbsp;genomes</strong>.&nbsp;<em>Nucleic Acids Research</em>. doi: 10.1093/nar/gkac1115</p> <p>&nbsp;</p> <p>Directory structure:</p> <ul> <li><strong>Bionano_optical_maps_based_assemblies</strong>: this directory contains results from the Bionano optical map based scaffolding&nbsp;for the main long-read assemblers analysed in the study for <em>Arabidopsis thaliana</em> accession&nbsp;Ey15-2 (9994): <ul> <li><strong>9994.CLR_Canu</strong></li> <li><strong>9994.HiFi_FALCON</strong></li> <li><strong>9994.HiFi_HiCanu</strong></li> <li><strong>9994.HiFi_Hifiasm</strong></li> <li><strong>9994.HiFi_IPA</strong></li> <li><strong>9994.HiFi_Peregrine</strong></li> </ul> </li> </ul> <p>&nbsp;</p> <ul> <li><strong>Col-0_HiFi-Hifiasm_assembly</strong>: this directory contains the Pacbio HiFi based chromosome level assembly (fasta file) and repeat annotation (gff file) of <em>Arabidopsis thaliana</em> accession&nbsp;Col-0 (6909).&nbsp;</li> </ul> <p>&nbsp;</p> <ul> <li><strong>Ey15-2_HiFi-Hifiasm_plus_CLR-Canu_assembly</strong>:&nbsp;this directory contains the Pacbio HiFi+CLR based chromosome level assembly (fasta file) and repeat annotation (gff file) of <em>Arabidopsis thaliana</em> accession&nbsp;Ey15-2 (9994).&nbsp;</li> </ul> <p>&nbsp;</p> <ul> <li><strong>Naish2021_Wang2021_repeat_annotation</strong>: this directory contains the repeat annotation (gff files) for the&nbsp;<em>Arabidopsis thaliana</em>&nbsp;Col-0 (6909) assemblies performed by Naish <em>et al.</em> (doi: 10.1126/science.abi7489) and Wang <em>et al.</em> (doi: 10.1016/j.gpb.2021.08.003).&nbsp;</li> </ul> <p>&nbsp;</p> <ul> <li><strong>TAIR10_masked</strong>: this directory contains the repeat-hard-masked version of the TAIR10&nbsp;<em>Arabidopsis thaliana</em>&nbsp;Col-0 (6909) reference genome that was used for in silico scaffolding of contigs with RagTag (<a href="http://github.com/malonge/RagTag">https://github.com/malonge/RagTag</a>).</li> </ul> <p>&nbsp;</p>

opencc-by-4.0Nov 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record