Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
492
datasets available to search
ShareScore release 0.9.0
Dataset results
492 results for “sequence modeling”
Geodetic Model of 2017 Kerman earthquake sequences estimated from joint inversion of InSAR and offset tracking techniques
<p>We upload here the inSAR and Pixel offset tracking results obtained based on Radar, Sentinel-1, PlanetScope, and Sentinel-2 satellite data. </p>
An evaluation of pool-sequencing transcriptome-based exon capture for population genomics of non-model species.
<p>This archive is associated with the article “An evaluation of pool-sequencing transcriptome-based exon capture for population genomics of non-model species.”. Authors: Emeline Deleury, Thomas Guillemaud, Aurelie Blin & Eric Lombaert.</p> <p>The archive contains :<br> - The sequences of the 5,717 Harmonia axyridis randomly selected CDS (5717-targeted-CDS-sequences.gff3, sequence in FASTA format at the end of the file)<br> - For the subset of 3,161 targeted CDS that have a genomic match over their entire length, the positions of exons on transcripts (3161-targeted-CDS-EXON-POSITIONS.csv)</p>
Deep Learning Based Models for Preimplantation Mouse and Human Embryos Based on Single Cell RNA Sequencing
Open the record for dataset details and reuse information.
Massive pre-main-sequence stars in M17 - Modelling hydrogen and dust in mYSO disks
<p>This is a basic reproduction package for the paper "Massive pre-main-sequence stars in M17 - Modelling hydrogen and dust in mYSO disks" by F. Backs, et al. 2023. It provides the data products required to reproduce the results.</p>
Case studies from doubleHelix: nucleic acid sequence identification, assignment and validation tool for cryo-EM and crystal structure models
<p>Case studies from "doubleHelix: nucleic acid sequence identification, assignment and validation tool for cryo-EM and crystal structure models"</p>
Case studies from: Sequence assignment validation in protein crystal structure models with checkMySequence
<p>Case studies from "Sequence assignment validation in protein crystal structure models with checkMySequence"</p>
Comprehensive benchmark and architectural analysis of deep learning models for Nanopore sequencing basecalling
<p>Placeholder data for the Lambda phage data used in: Comprehensive benchmark and architectural analysis of deep learning models for Nanopore sequencing basecalling.</p> <p>For the complete dataset see the Sequence Read Archive under the PRJNA926802 bioproject ID.</p>
Complementary sequence and model data for fungal E3BP
<p>Multiple Sequance Alignments (MSA) and atomic models of E3BP predicted using AlphaFold (AF), as supplementary datasets to the publication "The structure and evolutionary diversity of the fungal E3-binding protein" (2023)</p>
Enhancing Protein Sequence Annotation in Viral Genomics Using Large Language Models and Soft Alignments.
<p>List of 200 most abundant VOG descriptions.</p>
Protein sequences matching TIGRFAM models
<p>A set of 411 large and 3,001 small TIGRFAM protein families originally used for benchmarking sequence clustering programs. Large families contain at least 20,000 sequences with a genus label, whereas small contain fewer than 20,000. Sequences were downloaded from <a href="https://www.ncbi.nlm.nih.gov/genome/annotation_prok/tigrfams/">NCBI</a> and renamed by their original name (accession number) followed by their semi-colon separated NCBI taxonomy of the originating organism. For example, the first sequence in TIGRFAM_large/TIGRFAM00005.fas.gz is named:</p><blockquote><p>WP_000005837.1 RluA family pseudouridine synthase, partial [Bacillus anthracis];TIGR00005(group);cellular organisms(no rank);Bacteria(superkingdom);Terrabacteria group(clade);Firmicutes(phylum);Bacilli(class);Bacillales(order);Bacillaceae(family);Bacillus(genus);Bacillus cereus group(species group);Bacillus anthracis(species)</p></blockquote>
Mechanisms of Inherited Retinal Dystrophies Using Whole Genome Sequencing and in Vitro and in Vivo Models
ClinicalTrials.gov study NCT05793515. IPD Sharing: UNDECIDED. Countries: 1. Publications: 1.
Single-cell Sequencing and Establishment of Models in Neuroendocrine Neoplasm
ClinicalTrials.gov study NCT04927611. IPD Sharing: NO. Countries: 1. Publications: 21.
A Multi-omics Sequencing-based Model for Predicting Efficacy and Dynamic Monitoring of Treatment in Small Cell Lung Cancer
ClinicalTrials.gov study NCT07026669. IPD Sharing: NO. Countries: 1. Publications: 12.
Data from: Ultraconserved elements sequencing as a low-cost source of complete mitochondrial genomes and microsatellite markers in non-model amniotes
Open the record for dataset details and reuse information.
Data from: Targeted re-sequencing of coding DNA sequences for SNP discovery in non-model species
Open the record for dataset details and reuse information.
Data from: "Genome-wide microsatellite marker development from next-generation sequencing of two non-model bat species impacted by wind turbine mortality: Lasiurus borealis and L. cinereus (Vespertilionidae)" in Genomic Resources Notes accepted 1 October 2013 to 30 November 2013
Open the record for dataset details and reuse information.
Data from: Utility of pooled sequencing for association mapping in non-model organisms
Open the record for dataset details and reuse information.
Data from: Using a continuum model to decipher the mechanics of embryonic tissue spreading from time-lapse image sequences: an approximate Bayesian computation approach
Open the record for dataset details and reuse information.
Data from: Mixture models of nucleotide sequence evolution that account for heterogeneity in the substitution process across sites and across lineages
Open the record for dataset details and reuse information.
The genome sequence of Samia ricini, a new model species of lepidopteran insect
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.