Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

492

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

492 results for “sequence modeling”

Learn how ShareScore rates datasets ↗
zenodo36/100

Discovering molecular regulators of ageing using mixture models with RNA-sequencing data

<p>Identifying the molecular regulators that control ageing is challenging because the ageing process is influenced by a combination of genetic and environmental factors which makes it difficult to source the contribution of a single gene. Multiple studies have demonstrated that as humans age, increased gene expression heterogeneity results in the dysregulation of key regulators and pathways. Given the dynamic nature of gene expression, it is vital that this data be modelled by statistical approaches that can appropriately account for changes in variability to understand the contribution of heterogeneity during the aging process and properly identify its regulators. This study demonstrates the utility of using mixture models to model biological variability of gene expression occurring during ageing and how novel potential regulators of ageing can be identified.</p> <p>Our mixture modelling approach was applied to gene expression data from the Genotype-Tissue Expression (GTEx) cohort. For every gene, the expression profile was modelled using a mixture model across the cohort where the subset of donors corresponding to each mode was tested for a significant change in age group. The multi-tissue aspect of GTEx was leveraged to find ageing regulators based on this mixture model approach genes that were common across multiple tissues, suggesting that the regulation of ageing may also be controlled through a set of genes that have non-tissue-specific activity.</p> <p>Our approach identified well-documented ageing regulators <em>mTOR </em>and <em>RICTOR</em> and other potential ageing regulators such as <em>IL4</em> and <em>GPR4</em> which were detected only by our approach. Genes identified by edgeR, DESeq2 and the mixture model-based approach were enriched for similar biological pathways. This suggests that while the specific ageing regulators identified from our approach may be distinct, they generally belong in the same pathways as the genes identified by standard approaches. Overall, these results indicate that modelling gene expression variability using mixture models in conjunction with standard differential gene expression can help uncover new regulators that have a potential role for understanding human ageing.</p> <p>I</p>

opencc-by-4.0Apr 2022View details →
zenodo36/100

PredictION: A predictive model to establish the performance of Oxford sequencing reads of SARS-CoV-2

<p>Dataset (1) that included 1461 samples and a dataset (2) with 471 samples that was a subset of dataset 1 that included Number of sequenced reads per genome, CT (Cycle threshold) value [N2 target gene], mean coverage depth, coverage genome (percentage), and quantification cDNA (ng/&micro;l).</p>

opencc-by-4.0Sep 2022View details →
zenodo36/100

Ranger models for predicting isoform abundance from UTR sequence features

<p>Each RDS file contains a ranger object, trained on transcripts after removing those associated with the held-out genes in one of the five cross-validation folds. The day and replicate number in the file name corresponds to the neuronal differentiation sample on which the model was trained. The file gene_folds.txt indicates the fold from which each gene was excluded during model training. The file transcript_gene_associations.txt contains transcript-gene associations. The file predictors.RDS contains the matrix of predictor variables.</p>

opencc-by-4.0Jul 2024View details →
zenodo36/100

Evolution models of helium white dwarf-main-sequence star merger remnants: the mass distribution of single low-mass white dwarfs

<p>Inlists and data for &quot;<a href="https://ui.adsabs.harvard.edu/#abs/2018MNRAS.474..427Z/abstract">Evolution models of helium white dwarf-main-sequence star merger remnants: the mass distribution of single low-mass white dwarfs</a>&quot;</p>

opencc-by-4.0Mar 2019View details →
zenodo36/100

Spectral models for binary products: Unifying subdwarfs and Wolf-Rayet stars as a sequence of stripped-envelope stars

<p>MESA inlists associated with <a href="https://ui.adsabs.harvard.edu/#abs/2018A&amp;A...615A..78G/abstract">G&ouml;tberg et al. (2018)</a>. MESA version 8118.</p> <p>Publication DOI: <a href="https://doi.org/10.1051/0004-6361/201732274">10.1051/0004-6361/201732274</a></p>

opencc-by-4.0Mar 2019View details →
zenodo36/100

Geodetic model of the March 2021 Thessaly seismic sequence inferred from seismological and InSAR data

<p>A selection of Sentinel-1 (S1) wrapped and unwrapped measurements used in this study (from &quot;a&quot; to &quot;u&quot; files in tiff format as indicated in the word file attached). S1 data were processed by using our own internally developed InSAR&nbsp;processing chain.<br> <br> Earthquakes data locations.</p> <p><br> &nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Jun 2021View details →
zenodo36/100

Training data for the Sei framework sequence model

<p>Training data for the Sei framework deep learning model. The data contains chromatin profiles from the Cistrome Project: <strong>please agree to the terms of usage at the Cistrome Project (http://cistrome.org/db/#/bdown) before downloading.&nbsp;</strong></p>

opencc-by-4.0Jun 2021View details →
zenodo36/100

Long reads and Hi-C sequencing illuminate the two-compartment genome of the model arbuscular mycorrhizal symbiont Rhizophagus irregularis

<p>This repository contains annotations for the strains of <em>R. irregularis</em> chromosome assemblies.</p>

opencc-by-4.0Aug 2021View details →
zenodo36/100

Data for "Unsupervised learning of sequence-specific aggregation behavior for a model copolymer"

<p>These are the data associated with the paper, &quot;Unsupervised learning of sequence-specific aggregation behavior for a model copolymer&quot; (DOI 10.1039/D1SM01012C). Each of the directories contains subdirectories with `GSD` files dumped from HOOMD. Each subdirectory roughly corresponds to one or two of the figures in the paper.</p>

opencc-by-4.0Aug 2021View details →
zenodo36/100

Saved model and preprocessed data for "CRMnet:a deep learning model for predicting gene expression from large regulatory sequence datasets"

<p>Saved TUNet model and preprocessed training data&nbsp;for &quot;CRMnet: a deep learning model for predicting&nbsp;gene expression from large regulatory&nbsp;sequence datasets&quot;</p> <p>To load the trained model:</p> <pre><code class="language-python">import tensorflow as tf tf.keras.models.load_model("path to the model folder")</code></pre> <p>for more information please find our repository:&nbsp;https://github.com/jiayuwen/CRMnet</p>

opencc-by-4.0Nov 2022View details →
zenodo36/100

Modeling the Sequence Dependence of Differential Antibody Binding in the Immune Response to Infectious Disease

<p>Raw peptide microarray data of fluorescence intensities representing relative binding of antibodies in&nbsp;sera samples collected from different cohorts of patients diagnosed with a number of viral infections. Healthy controls are also included.</p> <p>Columns:</p> <p>Sequence - peptide sequence</p> <p>HCV - Hepatitis Virus C</p> <p>Dengue - Dengue virus</p> <p>WNV - West Nile Virus</p> <p>HBV - Hepatitis Virus B</p> <p>Chagas - Chagas disease</p> <p>ND - negative/healthy donor</p> <p>LowCV - low coefficient of variation</p> <p>HighCV - high coefficient of variation</p>

opencc-by-4.0Nov 2022View details →
zenodo36/100

Models and Data associated with: Single-cell gene expression prediction from DNA sequence at large contexts

<p>This archive holds trained models and associated data&nbsp;for the <a href="https://www.biorxiv.org/content/10.1101/2023.07.26.550634v1">manuscript</a>:<br> &quot;Single-cell gene expression prediction from DNA sequence at large contexts&quot;</p> <p>Structure:</p> <ul> <li>configs&nbsp;- example configs for the workflows to produce publication data&nbsp;</li> <li>data_* - pre-processed single cell data used for publication</li> <li>models_* - model checkpoints, hyperparameters and training progress in tensorboard logs</li> <li>preprocessing - additional data required to reproduce the pre-processing workflow</li> </ul> <p>&nbsp;</p> <p>&quot;Copyright 2023 GlaxoSmithKline Research &amp; Development Limited. All rights reserved.&quot;</p>

opencc-by-nc-nd-4.0Sep 2023View details →
dryad36/100

Colonization history of the Canary Islands endemic Lavatera acerifolia, (Malvaceae) unveiled with Genotyping-by-Sequencing data and niche modeling

Open the record for dataset details and reuse information.

publicFeb 2020View details →
dryad36/100

Generation of synthetic whole-slide image tiles of tumours from RNA-sequencing data via cascaded diffusion models

Open the record for dataset details and reuse information.

publicApr 2024View details →
dryad36/100

Visitation sequence data from: Alternative flowers affect model and mimic flower discrimination performance of bumble bees

Open the record for dataset details and reuse information.

publicApr 2021View details →
dryad36/100

Data from: Interaction of sequence data and paleogeographic priors in biogeographic dating: How could biological data inform time-constrained geological models?

Open the record for dataset details and reuse information.

publicDec 2025View details →
dryad36/100

Supplementary table, figures and DNA sequences of sorghum gene models SbiRTx430.01G455400 and SbiRTx.02G006600 that feature primers, gRNAs and indels created

Open the record for dataset details and reuse information.

publicMar 2025View details →
dryad36/100

Sequence-dependent model of genes with dual σ factor preference

Open the record for dataset details and reuse information.

publicApr 2022View details →
dryad36/100

Single-Cell RNA-sequencing of neural precursor cells from an Alzheimer's mouse model, wild-type mice, and Alzheimer's mice rescued with Usp16 haploinsufficiency

Open the record for dataset details and reuse information.

publicApr 2022View details →
dryad36/100

Data for: Accurate sequence-to-affinity models for SH2 domains from multi-round peptide binding assays coupled with free-energy regression

Open the record for dataset details and reuse information.

publicAug 2025View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record