Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

492

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

492 results for “sequence modeling”

Learn how ShareScore rates datasets ↗
zenodo32/100

Geodetic Model of 2017 Kerman earthquake sequences estimated from joint inversion of InSAR and offset tracking techniques

<p>We upload here the&nbsp;inSAR and Pixel offset tracking results obtained based on Radar, Sentinel-1, PlanetScope, and Sentinel-2 satellite data.&nbsp;&nbsp;&nbsp;</p>

opencc-by-4.0Oct 2022View details →
zenodo32/100

An evaluation of pool-sequencing transcriptome-based exon capture for population genomics of non-model species.

<p>This archive is associated with the article &ldquo;An evaluation of pool-sequencing transcriptome-based exon capture for population genomics of non-model species.&rdquo;. Authors: Emeline Deleury, Thomas Guillemaud, Aurelie Blin &amp;&nbsp; Eric Lombaert.</p> <p>The archive contains :<br> - The sequences of the 5,717 Harmonia axyridis randomly selected CDS (5717-targeted-CDS-sequences.gff3, sequence in FASTA format at the end of the file)<br> - For the subset of 3,161 targeted CDS that have a genomic match over their entire length, the positions of exons on transcripts (3161-targeted-CDS-EXON-POSITIONS.csv)</p>

opencc-by-4.0Mar 2019View details →
zenodo32/100

Deep Learning Based Models for Preimplantation Mouse and Human Embryos Based on Single Cell RNA Sequencing

Open the record for dataset details and reuse information.

opencc-by-4.0Sep 2024View details →
zenodo32/100

Massive pre-main-sequence stars in M17 - Modelling hydrogen and dust in mYSO disks

<p>This is a basic reproduction package for the paper &quot;Massive pre-main-sequence stars in M17 - Modelling hydrogen and dust in mYSO disks&quot; by F. Backs, et al. 2023. It provides the data products required to reproduce the results.</p>

opencc-by-4.0Aug 2022View details →
zenodo32/100

Case studies from doubleHelix: nucleic acid sequence identification, assignment and validation tool for cryo-EM and crystal structure models

<p>Case studies from &quot;doubleHelix: nucleic acid sequence identification, assignment and validation tool for cryo-EM and crystal structure models&quot;</p>

opencc-by-4.0Feb 2023View details →
zenodo32/100

Case studies from: Sequence assignment validation in protein crystal structure models with checkMySequence

<p>Case studies from&nbsp;&quot;Sequence assignment validation in protein crystal structure models with checkMySequence&quot;</p>

opencc-by-4.0Feb 2023View details →
zenodo32/100

Comprehensive benchmark and architectural analysis of deep learning models for Nanopore sequencing basecalling

<p>Placeholder data for the Lambda phage data used in: Comprehensive benchmark and architectural analysis of deep learning models for Nanopore sequencing basecalling.</p> <p>For the complete dataset see the Sequence Read Archive under the PRJNA926802 bioproject ID.</p>

openunlicenseFeb 2023View details →
zenodo32/100

Complementary sequence and model data for fungal E3BP

<p>Multiple Sequance Alignments (MSA) and atomic models of E3BP predicted using AlphaFold (AF), as supplementary datasets to the publication &quot;The structure and evolutionary diversity of the fungal E3-binding protein&quot; (2023)</p>

opencc-by-4.0Mar 2023View details →
zenodo32/100

Enhancing Protein Sequence Annotation in Viral Genomics Using Large Language Models and Soft Alignments.

<p>List of 200 most abundant VOG descriptions.</p>

opencc-by-4.0Jul 2023View details →
zenodo32/100

Protein sequences matching TIGRFAM models

<p>A set of 411 large and 3,001 small TIGRFAM&nbsp;protein families originally used for benchmarking sequence clustering programs. Large families contain at least 20,000 sequences with a genus label, whereas small contain fewer than 20,000. Sequences were&nbsp;downloaded from <a href="https://www.ncbi.nlm.nih.gov/genome/annotation_prok/tigrfams/">NCBI</a>&nbsp;and renamed by their original name (accession number)&nbsp;followed by their semi-colon separated NCBI taxonomy of the originating organism. For example, the first sequence in TIGRFAM_large/TIGRFAM00005.fas.gz is named:</p><blockquote><p>WP_000005837.1 RluA family pseudouridine synthase, partial [Bacillus anthracis];TIGR00005(group);cellular organisms(no rank);Bacteria(superkingdom);Terrabacteria group(clade);Firmicutes(phylum);Bacilli(class);Bacillales(order);Bacillaceae(family);Bacillus(genus);Bacillus cereus group(species group);Bacillus anthracis(species)</p></blockquote>

opencc-by-4.0Oct 2023View details →
ClinicalTrials.gov32/100

Mechanisms of Inherited Retinal Dystrophies Using Whole Genome Sequencing and in Vitro and in Vivo Models

ClinicalTrials.gov study NCT05793515. IPD Sharing: UNDECIDED. Countries: 1. Publications: 1.

restrictedIPD-UNDECIDEDFeb 2026View details →
ClinicalTrials.gov32/100

Single-cell Sequencing and Establishment of Models in Neuroendocrine Neoplasm

ClinicalTrials.gov study NCT04927611. IPD Sharing: NO. Countries: 1. Publications: 21.

closedIPD-NOFeb 2026View details →
ClinicalTrials.gov32/100

A Multi-omics Sequencing-based Model for Predicting Efficacy and Dynamic Monitoring of Treatment in Small Cell Lung Cancer

ClinicalTrials.gov study NCT07026669. IPD Sharing: NO. Countries: 1. Publications: 12.

closedIPD-NOFeb 2026View details →
dryad32/100

Data from: Ultraconserved elements sequencing as a low-cost source of complete mitochondrial genomes and microsatellite markers in non-model amniotes

Open the record for dataset details and reuse information.

publicSep 2016View details →
dryad32/100

Data from: Targeted re-sequencing of coding DNA sequences for SNP discovery in non-model species

Open the record for dataset details and reuse information.

publicJun 2018View details →
dryad32/100

Data from: "Genome-wide microsatellite marker development from next-generation sequencing of two non-model bat species impacted by wind turbine mortality: Lasiurus borealis and L. cinereus (Vespertilionidae)" in Genomic Resources Notes accepted 1 October 2013 to 30 November 2013

Open the record for dataset details and reuse information.

publicJan 2014View details →
dryad32/100

Data from: Utility of pooled sequencing for association mapping in non-model organisms

Open the record for dataset details and reuse information.

publicMar 2018View details →
dryad32/100

Data from: Using a continuum model to decipher the mechanics of embryonic tissue spreading from time-lapse image sequences: an approximate Bayesian computation approach

Open the record for dataset details and reuse information.

publicOct 2019View details →
dryad32/100

Data from: Mixture models of nucleotide sequence evolution that account for heterogeneity in the substitution process across sites and across lineages

Open the record for dataset details and reuse information.

publicJun 2014View details →
dryad32/100

The genome sequence of Samia ricini, a new model species of lepidopteran insect

Open the record for dataset details and reuse information.

publicSep 2020View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record