Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

244

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

244 results for “genomic variants”

Learn how ShareScore rates datasets ↗
zenodo36/100

dataset relared to article "Childhood‑onset dystonia‑causing KMT2B variants result in a distinctive genomic hypermethylation profile"

<p>Sanger sequences (.abi files) of all patients of Fondazione Besta cohort.</p>

opencc-by-4.0Sep 2021View details →
zenodo36/100

From Forensics to Clinical Research: Expanding the Variant Calling Pipeline for the Precision ID mtDNA Whole Genome Panel

<p>In this dataset we provide the 1000 Genomes Project&#39;s samples processed in our manuscript:</p> <p>Cortes-Figueiredo, F.; Carvalho, F.S.; Fonseca, A.C.; Paul, F.; Ferro, J.M.; Sch&ouml;nherr, S.; Weissensteiner, H.; Morais, V.A. From Forensics to Clinical Research: Expanding the Variant Calling Pipeline for the Precision ID mtDNA Whole Genome Panel. <em>Int. J. Mol. Sci</em>. <strong>2021</strong>, <em>22</em>, 12031. <a href="https://doi.org/10.3390/ijms222112031">https://doi.org/10.3390/ijms222112031</a><em>.</em></p> <p><strong>Abstract</strong></p> <p>Despite a multitude of methods for the sample preparation, sequencing, and data analysis of mitochondrial DNA (mtDNA), the demand for innovation remains, particularly in comparison with nuclear DNA (nDNA) research. The Applied Biosystems&trade; Precision ID mtDNA Whole Genome Panel (Thermo Fisher Scientific, USA) is an innovative library preparation kit suitable for degraded samples and low DNA input. However, its bioinformatic processing occurs in the enterprise Ion Torrent Suite&trade; Software (TSS), yielding BAM files aligned to an unorthodox version of the revised Cambridge Reference Sequence (rCRS), with a heteroplasmy threshold level of 10%. Here, we present an alternative customizable pipeline, the PrecisionCallerPipeline (PCP), for processing samples with the correct rCRS output after Ion Torrent sequencing with the Precision ID library kit. Using 18 samples (3 original samples and 15 mixtures) derived from the 1000 Genomes Project, we achieved overall improved performance metrics in comparison with the proprietary TSS, with optimal performance at a 2.5% heteroplasmy threshold. We further validated our findings with 50 samples from an ongoing independent cohort of stroke patients, with PCP finding 98.31% of TSS&rsquo;s variants (TSS found 57.92% of PCP&rsquo;s variants), with a significant correlation between the variant levels of variants found with both pipelines.</p> <p><br> Please refer to our the github page <a href="https://github.com/filcfig/PCP.git">filcfig/PCP</a>, for more details on running the PrecisionCalllerPipeline.</p>

opencc-by-4.0Oct 2021View details →
zenodo36/100

Detection of SARS-CoV-2 variants by genomic analysis of wastewater ampliconic samples (Galaxy Training Material)

<p>The tutorial aims to train how to run workflows to analyze lineages abundances in SAR-CoV-2 wastewater ampliconic samples. (https://training.galaxyproject.org/training-material/)</p>

opencc-by-4.0Dec 2022View details →
zenodo36/100

Detection of SARS-CoV-2 variants by genomic analysis of wastewater metatranscriptomic samples (Galaxy Training Material)

<p>The tutorial aims to train how to run workflows to analyze lineages abundances in SAR-CoV-2 wastewater metatranscriptomic samples. (https://training.galaxyproject.org/training-material/)</p>

opencc-by-4.0Dec 2022View details →
zenodo36/100

Multimodal learning of noncoding variant effects using genome sequence and chromatin structure

<p>ncVarPred-1D3D:</p> <p>The data used for testing the inconsistency among genome sequence, epigenetic profile, and later, to show its relation to 3D chromatin structure can be found in sanity_check_data.tar.gz.</p> <p>Some trained model for noncoding mutation effect prediction (mapping genome sequence to&nbsp;epigenetic profile) can be found in CNN_MLP, CNN_GCN, CNN_RNN_MLP, CNN_RNN_GCN.tar.gz.</p> <p>The trained model for pathogenic variants prediction can be found in fewshot_pathogenic_model.tar.gz.&nbsp;</p> <p>The training data can be found in training_data.tar.gz.</p> <p>Some noncoding variant&nbsp;effects prediction results, e.g. eQTL and pathogenic variants, can be replicated using the data shared in ncVar_data.tar.gz.</p>

opencc-by-4.0Apr 2023View details →
zenodo36/100

TP53 synthetic genomics data for benchmarking variant callers

<p>This is a&nbsp;synthetic genomics dataset generated with&nbsp;<a href="https://github.com/ncsa/NEAT">NEAT </a>&nbsp;for the gene TP53 for the use case of&nbsp;benchmarking somatic variant callers. The reports for all bam files where created using&nbsp;<a href="https://github.com/genome/bam-readcount">bam-readcount</a>.</p> <p>To find out more about our pipeline please visit&nbsp;<a href="https://github.com/BiodataAnalysisGroup/synth4bench">the Biodata Analysis Group GitHub</a>&nbsp;and also our&nbsp;<a href="https://biodataanalysisgroup.github.io/">GitHub page</a>&nbsp;:)</p>

opencc-by-4.0Jun 2023View details →
zenodo36/100

Multimodal learning of noncoding variant effects using genome sequence and chromatin structure

<p>ncVarPred-1D3D: pretrained models of Sei (PMID: 35817977) + our 3D structure embedding models are shared. The models are trained and validated&nbsp;using&nbsp;DeepSEA (PMID: 26301843) selected 200 bp regions (we extended to 4K bp neighboring) to predict the epigenetic profile containing 21907 epigenetic events Sei processed.</p> <p>The pretrained DeepSEA (PMID: 26301843) and reproduced DanQ (PMID: 27084946) can be found in SOTA.tar.gz.</p>

opencc-by-4.0Apr 2023View details →
zenodo36/100

Dataset related to "CFH and CFHR structural variants in atypical Hemolytic Uremic Syndrome: Prevalence, genomic characterization and impact on outcome "

<p>The upload consists of&nbsp;2 excel files, including&nbsp;genetic and clinical&nbsp;data, and 6&nbsp;power point files with western blot images.</p> <p>SMRT sequencing&nbsp;data are deposited in the EBI European Nucletide Archive; Accession number: PRJEB44176.&nbsp;</p>

opencc-by-4.0Aug 2023View details →
zenodo36/100

The Genomic Reference Resource for African Cattle: genome sequences and high-density array variants.

<p><em>The diversity in genome resources is fundamental to designing genomic strategies for local breed improvement and utilisation. These resources also support gene discovery and enhance our understanding of the mechanisms of resilience with applications beyond local breeds. We report here the genome sequences of 573 samples (198 new genomes) and high-density (HD) array genotyping of 1,082 samples (537 new samples) from indigenous African cattle populations. The new sequences have an average genome coverage of ~30X, three times higher than the average (~10X) of the over 300 sequences already in the public domain. Following variant quality checks, we identified approximately 32.4 million sequence variants and 661,943 HD autosomal variants mapped to the Bos taurus reference genome (ARS-UCD1.2). &nbsp;The new datasets were generated as part of the Centre for Tropical Livestock Genetic and Health (CTLGH) Genomic Reference Resource for African Cattle (GRRFAC) initiative, which aspires to facilitate the generation of this livestock resource. We hope this resource will be utilised by the global scientific community and breeders for sustainable global livestock improvement.</em></p>

opencc-by-4.0Sep 2023View details →
dryad36/100

A whole-genome reference panel of 14,393 individuals for East Asian populations accelerates discovery of rare functional variants

<p>Underrepresentation of non-European populations hinders growth of global precision medicine. Resources such as imputation reference panels that match the study population are necessary to find low-frequency variants with substantial effects. We created a reference panel consisting of 14,393 whole-genome sequences including more than 11,000 Asian individuals. Genome-wide association studies were conducted using the reference panel and a population-specific genotype array of 72K subjects for eight phenotypes. This panel yields improved imputation accuracy of rare and low-frequency variants within East Asian populations compared with the largest reference panel. Thirty-nine previously unidentified associations were found, and more than half of the variants were East-Asian-specific. We discovered genes with rare protein-altering variants, including LTBP1 for height and GPR75 for body mass index, as well as putative regulatory mechanisms for rare noncoding variants with cell-type-specific effects. We suggest this data set will add to the potential value of Asian precision medicine.</p>

opencc-zeroOct 2023View details →
dryad36/100

Variants and QTL associated with variation in genome-wide crossover number

Open the record for dataset details and reuse information.

publicDec 2024View details →
dryad36/100

The role of structural variants in pest adaptation and genome evolution of the Colorado potato beetle, Leptinotarsa decemlineata (Say)

Open the record for dataset details and reuse information.

publicJun 2024View details →
dryad36/100

A whole-genome reference panel of 14,393 individuals for East Asian populations accelerates discovery of rare functional variants

Open the record for dataset details and reuse information.

publicOct 2023View details →
dryad36/100

Data from: Genome divergence between European anchovy ecotypes fuelled by structural variants originating from trans-equatorial admixture

Open the record for dataset details and reuse information.

publicSep 2025View details →
dryad36/100

Large scale across-breed genome-wide association study reveals a variant in HMGA2 associated with inguinal cryptorchidism risk in dogs

Open the record for dataset details and reuse information.

publicMay 2022View details →
dryad36/100

Data for: Raw count data, transcribed variant count data, and reference genomic annotation files for Boocock et al. 2024

Open the record for dataset details and reuse information.

publicMay 2024View details →
dryad36/100

Data for: Growth cone advance requires EB1 as revealed by genomic replacement with a light-sensitive variant

Open the record for dataset details and reuse information.

publicMar 2023View details →
dryad36/100

Splice altering variant predictions in four archaic hominin genomes

Open the record for dataset details and reuse information.

publicDec 2022View details →
zenodo32/100

Variant analysis of SARS-CoV-2 genomes

<p>These are supplemental files accompanying a publication.</p>

opencc-by-4.0May 2020View details →
dryad32/100

Data from: Long read reference genome-free reconstruction of a full-length transcriptome from Astragalus membranaceus reveals transcript variants involved in bioactive compound biosynthesis

Astragalus membranaceus, also known as Huangqi in China, is one of the most widely used medicinal herbs in Traditional Chinese Medicine. Traditional Chinese Medicine formulations from Astragalus membranaceus have been used to treat a wide range of illnesses, such as cardiovascular disease, type 2 diabetes, nephritis and cancers. Pharmacological studies have shown that immunomodulating, anti-hyperglycemic, anti-inflammatory, antioxidant and antiviral activities exist in the extract of Astragalus membranaceus. Therefore, characterising the biosynthesis of bioactive compounds in Astragalus membranaceus, such as Astragalosides, Calycosin and Calycosin-7-O-β-D-glucoside, is of particular importance for further genetic studies of Astragalus membranaceus. In this study, we reconstructed the Astragalus membranaceus full-length transcriptomes from leaf and root tissues using PacBio Iso-Seq long reads. We identified 27 975 and 22 343 full-length unique transcript models in each tissue respectively. Compared with previous studies that used short read sequencing, our reconstructed transcripts are longer, and are more likely to be full-length and include numerous transcript variants. Moreover, we also re-characterised and identified potential transcript variants of genes involved in Astragalosides, Calycosin and Calycosin-7-O-β-D-glucoside biosynthesis. In conclusion, our study provides a practical pipeline to characterise the full-length transcriptome for species without a reference genome and a useful genomic resource for exploring the biosynthesis of active compounds in Astragalus membranaceus.

opencc-zeroDec 2016View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record