Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

409

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

409 results for “genome organization”

Learn how ShareScore rates datasets ↗
zenodo44/100

Nuclear Genome Organization in Fungi: From Gene folding to Rabl Chromosomes

<p>We discuss the current knowledge on the fungal genome organization, from the association of chromosomes within the nucleus to topological structures at individual genes and the genetic factors required for the hierarchical organization. Chromosome conformation capture followed by high-throughput sequencing (Hi-C) has elucidated how fungal genomes are globally organized in Rabl configuration where centromere or telomere bundles are associated with opposite faces of the nuclear envelope. Here, we explore the presence, in fungal taxa, of the typical proteins associated with genome organization in eukaryotes.</p>

opencc-by-4.0Dec 2022View details →
zenodo40/100

Implications of the three-dimensional chromatin organization for genome evolution in a fungal plant pathogen

<p><span>The spatial organization of eukaryotic genomes is linked to their biological functions, although it is not clear how this impacts the overall evolution of a genome. Here, we uncover the three-dimensional (3D) genome organization of the phytopathogen <em>Verticillium dahliae</em>,<em> </em>known to possess distinct genomic regions, designated adaptive genomic regions (AGRs), enriched in transposable elements and genes that mediate host infection. Short-range DNA interactions form clear topologically associating domains (TADs) with gene-rich boundaries that show reduced levels of gene expression and reduced genomic variation. Intriguingly, TADs are less clearly insulated in AGRs than in the core genome. At a global scale, the genome contains bipartite long-range interactions, particularly enriched for AGRs and more generally containing segmental duplications. Notably, the patterns observed for <em>V. dahliae </em>are also present in other <em>Verticillium</em> species. Thus, our analysis links 3D genome organization to evolutionary features conserved throughout the <em>Verticillium</em> genus.</span></p>

opencc-by-4.0Dec 2023View details →
zenodo40/100

Figure 3 in De novo mutations in the genome organizer CTCF cause intellectual disability

Figure 3. - Principal component analysis (PCA) of morphometric data. The ellipse highlights the group formed by the albino and normally pigmented specimens of the same size class. The albino is represented by the white square. Squares (N3 = 20-30 cm size-class); inverse triangles (N4 = 30-40 cm), circles (N5 = 40-50 cm), lozenges (N6 = 50-60 cm), and triangles (N7 = 60-70 cm).

opencc-by-4.0Dec 2013View details →
zenodo40/100

Figure 2 in De novo mutations in the genome organizer CTCF cause intellectual disability

Figure 2. - Regressions of log of disc width vs. log of weight (A) and log of total length vs. log of weight (B) using the albino specimen and data from 28 female individuals of G. micrura. The albino specimen is represented by the white dot.

opencc-by-4.0Dec 2013View details →
zenodo40/100

Loss of multi-level 3D genome organization during breast cancer progression - Processed LAD files

<p>This entry contains the processed LAD files produced as part of the following study:<br><strong>Loss of multi-level 3D genome organization during breast cancer progression</strong></p>

opencc-by-4.0Aug 2024View details →
zenodo40/100

Loss of multi-level 3D genome organization during breast cancer progression - Third-party datasets

<p>This entry contains the following datasets:</p> <p>Datasets used by <a href="https://github.com/dixonlab/hic_breakfinder" target="_blank" rel="noopener">hic_breakfinder</a>:</p> <ul> <li>inter_expect_1Mb.hg38.txt</li> <li>intra_expect_100kb.hg38.txt</li> </ul> <p>FIles were originally downloaded from <a href="https://salkinstitute.box.com/s/m8oyv2ypf8o3kcdsybzcmrpg032xnrgx" target="_blank" rel="noopener">this</a> URL.</p> <p>Datasets used by <a href="https://github.com/parklab/HiNT" target="_blank" rel="noopener">HiNT</a>:</p> <ul> <li>backgroundMatrices_hg38.zip</li> <li>refData_hg38.zip</li> </ul> <p>Files were originally downloaded from the following URLs: <a href="http://compbio.med.harvard.edu/hint/refData/" target="_blank" rel="noopener">link1</a>, <a href="http://compbio.med.harvard.edu/hint/backgroundMatrices/" target="_blank" rel="noopener">link2</a>.</p> <p>The above datasets are used by the data analysis workflows hosted at <a href="https://github.com/paulsengroup/2022-mcf10a-cancer-progression" target="_blank" rel="noopener">paulsengroup/2022-mcf10a-cancer-progression.</a><br>The results produced by running the workflows from the above repository were used as part of the following study:<br><strong>Loss of multi-level 3D genome organization during breast cancer progression</strong></p>

opencc-by-4.0Aug 2024View details →
zenodo40/100

Code and data associated with Christiansen et al. 2021 "Facilitating population genomics of non-model organisms through optimized experimental design for reduced representation sequencing"

<p>All code and data input and output files (except reference genome and raw sequencing data) needed to reproduce the results of Christiansen et al. 2021&nbsp;as released on&nbsp;<a href="https://github.com/notothen/radpilot">https://github.com/notothen/radpilot</a> alongside journal publication. See published paper:</p> <p>Christiansen, H., Heindler, F.M., Hellemans, B.&nbsp;<em>et al.</em>&nbsp;Facilitating population genomics of non-model organisms through optimized experimental design for reduced representation sequencing.&nbsp;<em>BMC Genomics</em>&nbsp;<strong>22,&nbsp;</strong>625 (2021). <a href="https://doi.org/10.1186/s12864-021-07917-3">https://doi.org/10.1186/s12864-021-07917-3</a></p>

openother-openJun 2021View details →
zenodo36/100

Genome-Scale Imaging of the 3D Organization and Transcriptional Activity of Chromatin

<p>We prepared these&nbsp;datasets&nbsp;associated with the paper &ldquo;Genome-scale imaging of the 3D organization and transcriptional activity of chromatin&rdquo; published in Cell: <a href="https://doi.org/10.1016/j.cell.2020.07.032">https://doi.org/10.1016/j.cell.2020.07.032</a>.</p> <p>Please find detailed descriptions of individual data files in the README_August 2020.txt.</p> <p>We provide example codes to load and analyze these datasets in:&nbsp;<a href="https://github.com/ZhuangLab/Chromatin_Analysis_2020_cell">https://github.com/ZhuangLab/Chromatin_Analysis_2020_cell</a>.</p> <p>If you use these datasets, please cite our Cell paper.</p>

opencc-by-4.0Aug 2020View details →
dryad36/100

The Rhododendron genome and chromosomal organization provide insight into shared whole-genome duplications across the heath family (Ericaceae)

<p>The genus <em>Rhododendron</em> (Ericaceae), which includes horticulturally important plants such as azaleas, is a highly diverse and widely distributed genus of &gt;1,000 species. Here, we report the chromosome-scale de novo assembly and genome annotation of <em>Rhododendron williamsianum</em> as a basis for continued study of this large genus. We created multiple short fragment genomic libraries, which were assembled using ALLPATHS-LG. This was followed by contiguity preserving transposase sequencing (CPT-seq) and fragScaff scaffolding of a large fragment library, which improved the assembly by decreasing the number of scaffolds and increasing scaffold length. Chromosome-scale scaffolding was performed by proximity-guided assembly (LACHESIS) using chromatin conformation capture (Hi-C) data. Chromosome-scale scaffolding was further refined and linkage groups defined by restriction-site associated DNA (RAD) sequencing of the parents and progeny of a genetic cross. The resulting linkage map confirmed the LACHESIS clustering and ordering of scaffolds onto chromosomes and rectified large-scale inversions. Assessments of the <em>R. williamsianum</em> genome assembly and gene annotation estimate them to be 89% and 79% complete, respectively. Predicted coding sequences from genome annotation were used in syntenic analyses and for generating age distributions of synonymous substitutions/site between paralgous gene pairs, which identified whole-genome duplications (WGDs) in <em>R. williamsianum</em>. We then analyzed other publicly available Ericaceae genomes for shared WGDs. Based on our spatial and temporal analyses of paralogous gene pairs, we find evidence for two shared, ancient WGDs in <em>Rhododendron</em> and <em>Vaccinium</em> (cranberry/blueberry) members that predate the Ericaceae family and, in one case, the Ericales order.</p>

opencc-zeroOct 2020View details →
zenodo36/100

Loss of multi-level 3D genome organization during breast cancer progression - FISH dataset

<p>This entry contains the raw and processed FISH images produced by the following study:<br><strong>Loss of multi-level 3D genome organization during breast cancer progression</strong></p> <p>The raw images contained in file 2022-mcf10a-cancer-progression-fish-db.tar.gz were processed using fish data analysis workflow (<a href="https://github.com/paulsengroup/2022-mcf10a-cancer-progression/blob/main/run_fish.sh" target="_blank" rel="noopener">link</a>) hosted at <a href="https://github.com/paulsengroup/2022-mcf10a-cancer-progression" target="_blank" rel="noopener">paulsengroup/2022-mcf10a-cancer-progression</a>.<br>The resulting files have been archived in file 2022-mcf10a-cancer-progression-fish-processed-data.tar.gz.</p>

opencc-by-4.0Aug 2024View details →
zenodo36/100

Comparative whole genome phylogeny of animal, environmental and human strains confirms the genogroups organization and the diversity of Stenotrophomonas maltophilia

<p>Reannotation of Smc genomes from Refseq (Prokka&nbsp;v1.13) and&nbsp;gene presence and absence spreadsheet from Roary.</p>

opencc-by-4.0Jun 2019View details →
zenodo36/100

CodonTransformer - Genomic and CodonTransformer-Generated Sequences for Fine-tuned Organisms

<p>This dataset is used in creating Fig. 2a and Supplementary Figs. 2-16 of the paper, mainly including the predictions of base (pretrained) and finetuend CodonTransformer model along with various metrics.&nbsp;</p>

opencc-by-4.0Aug 2024View details →
dryad36/100

The Rhododendron genome and chromosomal organization provide insight into shared whole-genome duplications across the heath family (Ericaceae)

Open the record for dataset details and reuse information.

publicOct 2020View details →
dryad32/100

Data from: Diversification in wild populations of the model organism Anolis carolinensis: a genome-wide phylogeographic investigation

The green anole (Anolis carolinensis) is a lizard widespread throughout the southeastern United States and is a model organism for the study of reproductive behavior, physiology, neural biology, and genomics. Previous phylogeographic studies of A. carolinensis using mitochondrial DNA and small numbers of nuclear loci identified conflicting and poorly supported relationships among geographically structured clades; these inconsistencies preclude confident use of A. carolinensis evolutionary history in association with morphological, physiological, or reproductive biology studies among sampling localities and necessitate increased effort to resolve evolutionary relationships among natural populations. Here, we used anchored hybrid enrichment of hundreds of genetic markers across the genome of A. carolinensis and identified five strongly supported phylogeographic groups. Using multiple analyses, we produced a fully resolved species tree, investigated relative support for each lineage across all gene trees, and identified mito-nuclear discordance when comparing our results to previous studies. We found fixed differences in only one clade—southern Florida restricted to the Everglades region—while most polymorphisms were shared between lineages. The southern Florida group likely diverged from other populations during the Pliocene, with all other diversification during the Pleistocene. Multiple lines of support, including phylogenetic relationships, a latitudinal gradient in genetic diversity, and relatively more stable long-term population sizes in southern phylogeographic groups, indicate that diversification in A. carolinensis occurred northward from southern Florida.

opencc-zeroDec 2015View details →
dryad32/100

Data from: "De novo assembled transcriptome of organs involved in reproduction in an endangered endemic Iberian cyprinid fish (Squalius pyrenaicus)" in Genomic Resources Notes Accepted 1 June 2015 to 31 July 2015

Sex determination systems are diverse, especially among fish, and include genetic and/or environmental components. Unexpectedly for such a basic aspect of development, sex determination systems change rapidly during evolution and gonadal fate is not ultimate, being actively maintained lifelong. Here, sequences of expressed genes involved in maintenance of gonad identity and reproduction processes were obtained through transcriptome assembly of the brain-gonadal axis tissues of a freshwater fish inhabiting highly variable environments, the gonochoristic Iberian fish Squalius pyrenaicus. Through Illumina total RNA-sequencing, male and female transcriptomes of brain and gonad tissues were assembled with Trans-ABySS software and merged to produce a more comprehensive S. pyrenaicus transcriptome. Coding sequences (CDS) predicted by TransDecoder were annotated using blastx. By means of read mapping against the reference transcriptome and CDS datasets, using Bowtie2, the accuracy of read mapping was assessed. This first endemic Iberian cyprinid transcriptome of organs involved in reproduction processes may serve as a valuable genomic resource for studying sexual mechanisms and other aspects of evolution, such as speciation and responses to environmental changes, and may be a useful tool for conservation studies since S. pyrenaicus is an endangered species.

opencc-zeroDec 2014View details →
dryad32/100

Data from: Whole organism lineage tracing by combinatorial and cumulative genome editing

Multicellular systems develop from single cells through distinct lineages. However, current lineage-tracing approaches scale poorly to whole, complex organisms. Here, we use genome editing to progressively introduce and accumulate diverse mutations in a DNA barcode over multiple rounds of cell division. The barcode, an array of clustered regularly interspaced short palindromic repeats (CRISPR)/Cas9 target sites, marks cells and enables the elucidation of lineage relationships via the patterns of mutations shared between cells. In cell culture and zebrafish, we show that rates and patterns of editing are tunable and that thousands of lineage-informative barcode alleles can be generated. By sampling hundreds of thousands of cells from individual zebrafish, we find that most cells in adult organs derive from relatively few embryonic progenitors. In future analyses, genome editing of synthetic target arrays for lineage tracing (GESTALT) can be used to generate large-scale maps of cell lineage in multicellular systems for normal development and disease.

opencc-zeroDec 2015View details →
dryad32/100

Data from: Genome-wide single nucleotide polymorphism (SNP) identification and characterization in a non-model organism, the African buffalo (Syncerus caffer), using next generation sequencing

This study aimed to develop a set of SNP markers with high resolution and accuracy within the African buffalo. Such a set can be used, among others, to depict subtle population genetic structure for a better understanding of buffalo population dynamics. In total, 18.5 million DNA sequences of 76 bp were generated by next generation sequencing on an Illumina Genome Analyzer II from a reduced representation library using DNA from a panel of 13 African buffalo representative of the four subspecies. We identified 2534 SNPs with high confidence within the panel by aligning the short sequences to the cattle genome (Bos taurus). The average sequencing depth of the complete aligned set of reads was estimated at 5x, and at 13x when only considering the final set of putative SNPs that passed the filtering criterion. Our set of SNPs was validated by PCR amplification and Sanger sequencing of 15 SNPs. Of these 15 SNPs, 14 amplified successfully and 13 were shown to be polymorphic (success rate: 87%). The fidelity of the identified set of SNPs and potential future applications are finally discussed.

opencc-zeroDec 2015View details →
zenodo32/100

Metagenome-assembled genomes for "Impacts of beaver ponds on biogeochemical cycling of organic nitrogen within a fire-impacted watershed"

<p>This dataset includes all of the metagenome-assembled genomes (MAGs) used in Roth et al.:&nbsp;&quot;Impacts of beaver ponds on biogeochemical cycling of organic nitrogen within a fire-impacted watershed&quot; (in prep.). The metagenomic sequencing was completed on a suite of sediment samples collected from the sediment-water interface of beaver ponds within wildfire burn scars.</p>

opencc-by-4.0Dec 2021View details →
zenodo32/100

Draft genome sequences of Arabidopsis thaliana-associated micro-organisms from Reijerscamp soil, the Netherlands

<p><strong>Methodological summary and relevant references</strong></p> <p>Compressed tar archive containing 447 draft bacterial genomes and their annotations used in several studies including Fourie&nbsp;<em>et al</em>. (2024; in review) and Selten et al. (2024; in prep). Genome sequences are obtained by Illumina-only sequencing of microbial cultures. Illumina reads were demultiplexed and cleaned with cutadapt (version 2.8) (Martin, 2011) and assembled into genomes using A5 (A5-miseq version 20160825) (Coil et al., 2014). Genome contamination and heterogeneity was checked with CheckM (version 1.1.3) (Parks et al., 2015) and any genomes with multiple single copy gene occurrences were subjected to MaxBin (version 2.2.7) (Wu et al., 2014) to separate the genomes from contaminated bacterial cultures. Any non-bacterial contigs in the genome assemblies were removed using MMSeqs2 (version 13.45111) (Steineigger &amp; Sch&ouml;ding, 2017). Open reading frames were found and annotated by PROKKA (version 1.14.6) (Seemann, 2014) and EggNOG (version 2.1.4-2) (Cantalapiedra et al., 2021) respectively.&nbsp;Microbial cultures were derived from&nbsp;<em>Arabidopsis thaliana</em> roots grown in Reijerscamp soil, described in Stringlis <em>et al</em>., 2018 https://doi.org/10.1073/pnas.1722335115.</p> <p><strong>The uploaded files are</strong></p> <ol> <li>Genome assemblies</li> <li>Prokka gene predictions in GFF3 format</li> <li>Predicted transcripts from genes in (2)</li> <li>Predicted proteins from genes in (2), and</li> <li>EggNOG annotations for the proteins in (4)</li> </ol> <p><strong>Genomes and annotations pending upload om NCBI GenBank (April 2024)</strong></p>

opencc-by-nc-nd-4.0Apr 2024View details →
zenodo32/100

FIGURE 6 in Complete nucleotide sequence and organization of the mitochondrial genome of Sirthenea flavipes (Hemiptera: Reduviidae: Peiratinae) and comparison with other assassin bugs

FIGURE 6. Phylogenetic tree of four sequenced assassin bugs. Bayesian inference and Maximum likelihood analysis inferred from all genes recovered the same topological structure. Bootstrap values and Bayesian posterior probabilities are indicated at each node.

opennotspecifiedJun 2013View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record