Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

29

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

29 results for “Fragmented genomes”

Learn how ShareScore rates datasets ↗
dryad40/100

Virus classification for viral genomic fragments using PhaGCN2

<p>Viruses are the most ubiquitous and diverse entities in the biome. Due to the rapid growth of newly identified viruses, there is an urgent need for accurate and comprehensive virus classification, particularly for novel viruses. Here, we present PhaGCN2, which can rapidly classify the taxonomy of viral sequences at family level and supports the visualization of the associations of all families. We evaluate the performance of PhaGCN2 and compare it with the state-of-the-art virus classification tools, such as vConTACT2, CAT, and VPF-Class, using the widely accepted metrics. The results show that PhaGCN2 largely improves the precision and recall of virus classification, increases the number of classifiable virus sequences in the Global Ocean Virome dataset (v2.0) by 4 times, and classifies more than 90% of the Gut Phage Database. PhaGCN2 makes it possible to conduct high-throughput and automatic expansion of the database of the International Committee on Taxonomy of Viruses.</p>

opencc-zeroApr 2022View details →
zenodo40/100

Рис. 1. ФиΛогенетические Αеревья хантавируса AMRV и его прироΑного носитеΛя восточноазиатской мыши Apodemus peninsulae Thomas, 1906. А. ФиΛогенетическое Αерево восточноазиатской мыши Apodemus peninsulae, построенное метоΑом «максимаΛьного правΑопоΑобия» (ML) и поΛученное на основе анаΛиза участка гена цитохрома b мтΔНК (744 п.н.). В узΛах ветвΛения указаны бутстреп-поΑΑержки, рассчитанные ΑΛя 1000 повторов. Цветными Λиниями обозначены фиΛогенетические Λинии: Αве Китайские (зеΛеный), Корейская «Korea» (синий), Амурская «Amur» (красный). ПоΛужирным шрифтом выΑеΛены собственные образцы. Названия образцов из GenBank/NCBI быΛи сокращены; B. ФиΛогенетическое Αерево из работы Α. Н. Яшиной с ΑопоΛнениями, построенное метоΑом «бΛижайшего сосеΑа» (NJ) на основе посΛеΑоватеΛьностей фрагмента М-сегмента (2737–2980 н.п.) генома хантавирусов. В узΛах ветвΛения указаны бутстреппоΑΑержки, рассчитанные ΑΛя 1000 повторов. Жирным выΑеΛены иссΛеΑованные РНК изоΛяты (Яшина 2012; Яшина и Αр. 2019) Fig. 1. Phylogenetic trees of AMRV and its natural reservoir host — the Korean field mouse Apodemus peninsulae Thomas, 1906. A. Phylogenetic tree of the Korean field mouse Apodemus peninsulae constructed by the "maximum likelihood" method (ML). The data are obtained from the analysis of the cytochrome b mtDNA gene fragments (744 bp). Bootstrap supports calculated for 1,000 repeats are indicated in the branching nodes. Colored lines indicate phylogenetic lines: two Chinese (green), Korea (blue), and Amur (red). Own samples are highlighted in bold. The names of the samples from GenBank/NCBI have been shortened; B. Phylogenetic tree from L. N. Yashina's work with additions constructed by the neighbour joining method (NJ). It is based on the sequences of an M-segment fragment (2737–2980 bp) of the hantavirus genome. Bootstrap supports calculated for 1,000 repeats are indicated in the branching nodes. The researched RNA isolates are highlighted in bold (Yashina 2012; Yashina et al. 2019) in Variability of the gene cyt b in the Korean field mouse Apodemus peninsulae Thomas, 1906 - a reservoir host of AMRV in the Khasansky District of Primorsky Krai

Рис. 1. ФиΛогенетические Αеревья хантавируса AMRV и его прироΑного носитеΛя восточноазиатской мыши Apodemus peninsulae Thomas, 1906. А. ФиΛогенетическое Αерево восточноазиатской мыши Apodemus peninsulae, построенное метоΑом «максимаΛьного правΑопоΑобия» (ML) и поΛученное на основе анаΛиза участка гена цитохрома b мтΔНК (744 п.н.). В узΛах ветвΛения указаны бутстреп-поΑΑержки, рассчитанные ΑΛя 1000 повторов. Цветными Λиниями обозначены фиΛогенетические Λинии: Αве Китайские (зеΛеный), Корейская «Korea» (синий), Амурская «Amur» (красный). ПоΛужирным шрифтом выΑеΛены собственные образцы. Названия образцов из GenBank/NCBI быΛи сокращены; B. ФиΛогенетическое Αерево из работы Α. Н. Яшиной с ΑопоΛнениями, построенное метоΑом «бΛижайшего сосеΑа» (NJ) на основе посΛеΑоватеΛьностей фрагмента М-сегмента (2737–2980 н.п.) генома хантавирусов. В узΛах ветвΛения указаны бутстреппоΑΑержки, рассчитанные ΑΛя 1000 повторов. Жирным выΑеΛены иссΛеΑованные РНК изоΛяты (Яшина 2012; Яшина и Αр. 2019) Fig. 1. Phylogenetic trees of AMRV and its natural reservoir host — the Korean field mouse Apodemus peninsulae Thomas, 1906. A. Phylogenetic tree of the Korean field mouse Apodemus peninsulae constructed by the "maximum likelihood" method (ML). The data are obtained from the analysis of the cytochrome b mtDNA gene fragments (744 bp). Bootstrap supports calculated for 1,000 repeats are indicated in the branching nodes. Colored lines indicate phylogenetic lines: two Chinese (green), Korea (blue), and Amur (red). Own samples are highlighted in bold. The names of the samples from GenBank/NCBI have been shortened; B. Phylogenetic tree from L. N. Yashina's work with additions constructed by the neighbour joining method (NJ). It is based on the sequences of an M-segment fragment (2737–2980 bp) of the hantavirus genome. Bootstrap supports calculated for 1,000 repeats are indicated in the branching nodes. The researched RNA isolates are highlighted in bold (Yashina 2012; Yashina et al. 2019)

opencc-by-4.0Jul 2024View details →
zenodo40/100

Data accompanying "Genomic data recover previously undetectable fragmentation effects in an endangered amphibian"

<p>Target sequences and&nbsp;SNP and genotype calls from the manuscript&nbsp;&quot;Genomic data recover previously undetectable fragmentation effects in an endangered amphibian&quot;.</p>

opencc-by-4.0Oct 2018View details →
zenodo40/100

An ultra-dense haploid genetic map for evaluating the highly fragmented genome assembly of Norway spruce (Picea abies)

<p>Data files for construction of the haploid genetic map for Norway spruce (<em>Picea abies</em>). &nbsp;Available at&nbsp;&nbsp;<a href="https://doi.org/10.1101/292151">https://doi.org/10.1101/292151</a></p>

opencc-by-4.0Mar 2018View details →
zenodo40/100

Fig. 2 in Comparative analyses of the fragmented mitochondrial genomes of wild pig louse Haematopinus apri from China and Japan

Fig. 2. The complete mitochondrial genome of wild pig louse Haematopinus apri form China. Each minichromosome has a coding region and a non-coding region (NCR, in black). The names and transcript orientation of genes are indicated in the coding region and the minichromosomes are placed in alphabetical order of protein-coding genes and rRNA genes. Abbreviations: atp6 and atp8, ATP synthase F0 subunits 6 and 8; cytb, cytochrome b; cox1-3, cytochrome c oxidase subunits 1–3; nad1-6 and nad4L, NADH dehydrogenase subunits 1–6 and 4L; rrnS and rrnL, small and large subunits of ribosomal RNA. tRNA genes are indicated with their single-letter abbreviations of the corresponding amino acids.

opencc-by-4.0Aug 2022View details →
dryad40/100

Virus classification for viral genomic fragments using PhaGCN2

Open the record for dataset details and reuse information.

publicApr 2022View details →
zenodo36/100

The alignments of chloroplast genome sequences and nuclear ribosomal DNA fragments of six oak species sampled in the hot-dry valley of the Jinsha River, southwestern China

<p>Both chloroplast (cp) genome sequences and nuclear ribosomal (nr) DNA were assembled using GetOrganelle v.1.7.6.1 for 18 oak trees sampled in the Panzhihua Cycad National Nature Reserve, Sichuan Province, China. These trees belong to six oak species, including Quercus cocciferoides, Q. dolicholepis, Q. franchetii, Q. griffithii, Q. longispica, and Q. variabilis. We used PhyloSuite v.1.1.152 to extract coding sequences (CDSs), tRNA genes, rRNA genes, introns, and intergenic spacers (IGSs) of the 18 oak cp genomes. These sequences were aligned separately using MAFFT v.7.3.13 and manually adjusted with BioEdit v.7.2.5. Length variations in mononucleotide repeats were excluded and inversions were replaced with their reverse complements because of their tendency for homoplasy. Other indels were coded as binary characters according to the simple gap coding method using GapCoder. Separate assignments were concatenated according to their respective positions in the cp genome to obtain the alignments of LSC, SSC, IRb, and the whole cp genome.</p>

opencc-by-4.0Jan 2024View details →
dryad36/100

Population genomics of flat-tailed horned lizards (Phrynosoma mcallii) informs conservation and management across a fragmented Colorado Desert landscape

<p><em>Phrynosoma mcallii</em> (flat-tailed horned lizards) is a species of conservation concern in the Colorado Desert of the United States and Mexico. We analyzed ddRADseq data from 45 lizards to estimate population structure, infer phylogeny, identify migration barriers, map genetic diversity hotspots, and model demography. We identified the Colorado River as the main geographic feature contributing to population structure, with the populations west of this barrier further subdivided by the Salton Sea. Phylogenetic analysis confirms that northwestern populations are nested within southeastern populations. The best-fit demographic model indicates Pleistocene divergence across the Colorado River, with significant bidirectional gene flow, and a severe Holocene population bottleneck. These patterns suggest that management strategies should focus on maintaining genetic diversity on both sides of the Colorado River and Salton Sea. We recommend additional lands in the U.S. and Mexico that should be considered for similar conservation goals as those in the Rangewide Management Strategy (RMS). We also recommend periodic rangewide genomic sampling to monitor ongoing attrition of diversity, hybridization, and changing structure due to habitat fragmentation, climate change and other long-term impacts.</p>

opencc-zeroApr 2024View details →
zenodo36/100

Nanopore MinION Run Metrics and genomic DNA fragment size analysis data from automated phenol-chloroform extractions (RBI LabDroid Maholo)

<p>Nanopore MinION run MinKNOW statistical metrics output, Agilent Femto Pulse and Tape Station gDNA fragment size analysis reports of genomic DNA isolated from automated&nbsp;RBI LabDroid&nbsp;Maholo organic extractions.</p>

opencc-by-4.0Jan 2022View details →
dryad36/100

Population genomics of flat-tailed horned lizards (Phrynosoma mcallii) informs conservation and management across a fragmented Colorado Desert landscape

Open the record for dataset details and reuse information.

publicApr 2024View details →
dryad36/100

A spatial genomic approach identifies time lags and historic barriers to gene flow in a rapidly fragmenting Appalachian landscape

Open the record for dataset details and reuse information.

publicJun 2020View details →
dryad32/100

Data from: Sun skink landscape genomics: assessing how microevolutionary processes shape genetic and phenotypic diversity across a heterogeneous and fragmented landscape

Incorporating genomic data sets into landscape genetic analyses allows for powerful insights into population genetics, explicitly geographical correlates of selection, and morphological diversification of organisms across the geographical template. Here, we utilize an integrative approach to examine gene flow and detect selection, and we relate these processes to genetic and phenotypic population differentiation across South-East Asia in the common sun skink, Eutropis multifasciata. We quantify the relative effects of geographic and ecological isolation in this system and find elevated genetic differentiation between populations from island archipelagos compared to those on the adjacent South-East Asian continent, which is consistent with expectations concerning landscape fragmentation in island archipelagos. We also identify a pattern of isolation by distance, but find no substantial effect of ecological/environmental variables on genetic differentiation. To assess whether morphological conservatism in skinks may result from stabilizing selection on morphological traits, we perform FST–PST comparisons, but observe that results are highly dependent on the method of comparison. Taken together, this work provides novel insights into the manner by which micro-evolutionary processes may impact macro-evolutionary scale biodiversity patterns across diverse landscapes, and provide genomewide confirmation of classic predictions from biogeographical and landscape ecological theory.

opencc-zeroDec 2014View details →
dryad32/100

Data from: Genome-wide assessment of population structure and genetic diversity and development of a core germplasm set for sweet potato based on specific length amplified fragment (SLAF) sequencing

Sweet potato, Ipomoea batatas (L.) Lam., is an important food crop that is cultivated worldwide. However, no genome-wide assessment of the genetic diversity of sweet potato has been reported to date. In the present study, the population structure and genetic diversity of 197 sweet potato accessions most of which were from China were assessed using 62,363 SNPs. A model-based structure analysis divided the accessions into three groups: group 1, group 2 and group 3. The genetic relationships among the accessions were evaluated using a phylogenetic tree, which clustered all the accessions into three major groups. A principal component analysis (PCA) showed that the accessions were distributed according to their population structure. The mean genetic distance among accessions ranged from 0.290 for group 1 to 0.311 for group 3, and the mean polymorphic information content (PIC) ranged from 0.232 for group 1 to 0.251 for group 3. The mean minor allele frequency (MAF) ranged from 0.207 for group 1 to 0.222 for group 3. Analysis of molecular variance (AMOVA) showed that the maximum diversity was within accessions (89.569%). Using CoreHunter software, a core set of 39 accessions was obtained, which accounted for approximately 19.8% of the total collection. The core germplasm set of sweet potato developed will be a valuable resource for future sweet potato improvement strategies.

opencc-zeroDec 2016View details →
zenodo32/100

CRAG: De novo characterization of cell-free DNA fragmentation hotspots in plasma whole-genome sequencing

<p><strong>Supplementary dataset for the manuscript:</strong></p> <p><strong>CRAG: De novo characterization of cell-free DNA fragmentation hotspots in plasma whole-genome sequencing</strong></p> <p>&nbsp;Xionghui Zhou1,*, Haizi Zheng1,*, Hailu Fu1,*, Kelsey L. Dillehay McKillip2-3, Susan M. Pinney2,4, Yaping Liu1-2,5-7 #</p> <p>Affiliations:</p> <p>1 Division of Human Genetics, Cincinnati Children&rsquo;s Hospital Medical Center, Cincinnati, OH 45229</p> <p>2 University of Cincinnati Cancer Center, Cincinnati, OH 45229</p> <p>3 Department of Pathology &amp; Laboratory Medicine, University of Cincinnati College of Medicine, Cincinnati, OH 45229</p> <p>4 Department of Environmental and Public Health Sciences, University of Cincinnati College of Medicine, Cincinnati, OH 45229</p> <p>5 Division of Biomedical Informatics, Cincinnati Children&rsquo;s Hospital Medical Center, Cincinnati, OH 45229</p> <p>6 Department of Pediatrics, University of Cincinnati College of Medicine, Cincinnati, OH 45229</p> <p>7 Department of Electrical Engineering and Computing Sciences, University of Cincinnati College of Engineering and Applied Science, Cincinnati, OH 45229</p> <p>* These authors contributed equally</p> <p># Email: lyping1986@gmail.com</p>

opencc-by-4.0Jun 2022View details →
dryad32/100

Genome assembly of the Australian black tiger shrimp (Penaeus monodon) reveals a novel fragmented IHHNV EVE sequence

<p>Abstract Shrimp are a valuable aquaculture species globally; however, disease remains a major hindrance to shrimp aquaculture sustainability and growth. Mechanisms mediated by endogenous viral elements have been proposed as a means by which shrimp that encounter a new virus start to accommodate rather than succumb to infection over time. However, evidence on the nature of such endogenous viral elements and how they mediate viral accommodation is limited. More extensive genomic data on Penaeid shrimp from different geographical locations should assist in exposing the diversity of endogenous viral elements. In this context, reported here is a PacBio Sequel-based draft genome assembly of an Australian black tiger shrimp (Penaeus monodon) inbred for 1 generation. The 1.89 Gbp draft genome is comprised of 31,922 scaffolds (N50: 496,398 bp) covering 85.9% of the projected genome size. The genome repeat content (61.8% with 30% representing simple sequence repeats) is almost the highest identified for any species. The functional annotation identified 35,517 gene models, of which 25,809 were protein-coding and 17,158 were annotated using interproscan. Scaffold scanning for specific endogenous viral elements identified an element comprised of a 9,045-bp stretch of repeated, inverted, and jumbled genome fragments of infectious hypodermal and hematopoietic necrosis virus bounded by a repeated 591/590 bp host sequence. As only near complete linear ∼4 kb infectious hypodermal and hematopoietic necrosis virus genomes have been found integrated in the genome of P. monodon previously, its discovery has implications regarding the validity of PCR tests designed to specifically detect such linear endogenous viral element types. The existence of joined inverted infectious hypodermal and hematopoietic necrosis virus genome fragments also provides a means by which hairpin double-stranded RNA could be expressed and processed by the shrimp RNA interference machinery.</p>

opencc-zeroDec 2022View details →
dryad32/100

Genome assembly of the Australian black tiger shrimp (Penaeus monodon) reveals a novel fragmented IHHNV EVE sequence

Open the record for dataset details and reuse information.

publicDec 2022View details →
dryad32/100

Data from: Genome-wide assessment of population structure and genetic diversity and development of a core germplasm set for sweet potato based on specific length amplified fragment (SLAF) sequencing

Open the record for dataset details and reuse information.

publicFeb 2018View details →
dryad32/100

Data from: Sun skink landscape genomics: assessing how microevolutionary processes shape genetic and phenotypic diversity across a heterogeneous and fragmented landscape

Open the record for dataset details and reuse information.

publicMar 2015View details →
dryad28/100

Data from: Mitochondrial genome fragmentation unites the parasitic lice of eutherian mammals

Organelle genome fragmentation has been found in a wide range of eukaryotic lineages; however, its use in phylogenetic reconstruction has not been demonstrated. We explored the use of mitochondrial (mt) genome fragmentation in resolving the controversial suborder-level phylogeny of parasitic lice (order Phthiraptera). There are ~5,000 species of parasitic lice in four suborders (Amblycera, Ischnocera, Rhyncophthirina and Anoplura), which infest mammals and birds. The phylogenetic relationships among these suborders are unresolved despite decades of studies. We sequenced the mt genomes of eight species of parasitic lice and compared them with 17 other species of parasitic lice sequenced previously. We found that the typical single-chromosome mt genome is retained in the lice of birds but fragmented into many minichromosomes in the lice of eutherian mammals. The shared derived feature of mt genome fragmentation unites the eutherian mammal lice of Ischnocera (family Trichodectidae) with Anoplura and Rhyncophthirina to the exclusion of the bird lice of Ischnocera (family Philopteridae). This novel clade is also supported by phylogenetic analysis of mt genome and cox1 gene sequences. Our results demonstrate, for the first time, that organelle genome fragmentation is informative for resolving controversial high-level phylogenies.

opencc-zeroDec 2017View details →
dryad28/100

Data from: Mitochondrial genome fragmentation unites the parasitic lice of eutherian mammals

Open the record for dataset details and reuse information.

publicSep 2018View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record