Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

302

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

302 results for “protein sequence”

Learn how ShareScore rates datasets ↗
geo24/100

MP3-seq: Massively parallel measurement of protein-protein interactions by sequencing

GEO Series GSE271790. Saccharomyces cerevisiae. 36 samples. Type: Other.

openGEO-OpenJul 2024View details →
zenodo24/100

SARS-CoV-2 vs. Homo sapiens BLASTP protein sequence analysis results

<p>SARS-CoV-2 vs. Homo sapiens BLASTP protein sequence analysis results</p>

opencc-by-4.0Apr 2020View details →
dryad24/100

The whole protein sequences of the nuclear genome of Chrysosplenium sinicum

<p>We used PacBio and Illumina sequencing to de novo assemble the nuclear genome of Chrysosplenium sinicum, and further annotated the whole protein  sequences of the nuclear genome of Chrysosplenium sinicum.</p>

opencc-zeroAug 2020View details →
zenodo24/100

Structure prediction of protein-ligand complexes from sequence information with Umol

<p>posebusters_benchmark_set.tar.zst - files for the prediction (features to Umol) and scoring of the pose busters benchmark&nbsp;</p><p>posebusters_pred_native.tar.zst - pdb and sdf files of proteins and ligands. Includes native structures, predicted structures and relaxed predicted structures with plDDT in the B factor column.</p><p>posebusters_scores.csv - contains ligand RMSD and other metrics for the unrelaxed structures predicted with Umol.</p><p>PDBBind_processed.tar.zst &nbsp;- files for the training (features to Umol) using PDBbind version 2020</p><p>&nbsp;</p><p>&nbsp;</p>

opencc-by-4.0Oct 2023View details →
zenodo24/100

Predicting biophysical characteristics of proteins from their amino acid sequence

<p>As part of the Galaxy&nbsp;Training Material, this is the tutorial workflow&#39;s input file in FASTA format. It contains the amino acid sequence&nbsp;of protein P04156 (&quot;PRIO_HUMAN OS=Homo sapiens OX=9606 GN=PRNP PE=1 SV=1&quot;).</p>

opencc-by-4.0Sep 2022View details →
zenodo24/100

Representations and associated fitness values of protein sequences

<p>We studied the ability to predict protein fitness from sequence using our method&nbsp;<a href="https://github.com/amillig/MERGE" target="_blank" rel="noopener">MERGE</a> and other methods (i.e., ECNet, eUniRep, EVmutation, One-Hot, and UniRep). The following files are included in this repository:</p> <ul> <li>The&nbsp;folder&nbsp;<em>ECNet</em>&nbsp;contains the braw, csv, and fasta files required to run <a href="https://github.com/luoyunan/ECNet" target="_blank" rel="noopener">ECNet</a>.</li> <li>The folders&nbsp;<em>eUniRep, </em><em>MERGE, One-Hot</em>, and <em>UniRep</em> contain csv files including&nbsp;the names and fitness values of protein variants as well as a numerical representation of their sequence. The folder <em>MERGE </em>also contains representations of the wild type sequences as npy files in the&nbsp;<em>wts</em> folder.</li> <li>The folder <em>Params</em> contains params files generated with <a href="https://github.com/debbiemarkslab/plmc" target="_blank" rel="noopener">PLMC</a>.</li> <li>The script <em>get_performances.py</em> enables to generate models using different methods (i.e., eUniRep, EVmutation, MERGE, One-Hot, Pure_ML, and UniRep) and to determine their performance for predicting the fitness of protein variants from sequence.</li> </ul> <p><strong>References</strong></p> <table> <tbody> <tr> <td><a href="https://doi.org/10.1038/s41467-021-25976-8" target="_blank" rel="noopener">ECNet</a></td> <td> <p>Luo, Y., Jiang, G., Yu, T. et al. ECNet is an evolutionary context-integrated deep learning framework for protein engineering. Nat Commun 12, 5743 (2021).</p> </td> </tr> <tr> <td><a href="https://doi.org/10.1038/s41592-021-01100-y" target="_blank" rel="noopener">eUniRep</a></td> <td>Biswas, S., Khimulya, G., Alley, E.C. et al. Low-N protein engineering with data-efficient deep learning. Nat Methods 18, 389&ndash;396 (2021)</td> </tr> <tr> <td><a href="https://doi.org/10.1038/nbt.3769" target="_blank" rel="noopener">EVmutation</a></td> <td>Hopf, T., Ingraham, J., Poelwijk, F. et al. Mutation effects predicted from sequence co-variation. Nat Biotechnol 35, 128&ndash;135 (2017).</td> </tr> <tr> <td><a href="https://doi.org/10.1038/s41592-019-0598-1" target="_blank" rel="noopener">UniRep</a></td> <td>Alley, E.C., Khimulya, G., Biswas, S. et al. Unified rational protein engineering with sequence-based deep representation learning. Nat Methods 16, 1315&ndash;1322 (2019).</td> </tr> </tbody> </table> <p>&nbsp;</p>

openApr 2024View details →
zenodo24/100

A joint embedding of protein sequence and structure enables robust variant effect predictions

<p>Data related to the GitHub repository KULL-Centre/_2023_Blaabjerg_SSEmb, which is also stored on Zenodo here: <span><span><a href="../doi/10.5281/zenodo.13765792" target="_blank" rel="noopener noreferrer">https://zenodo.org/doi/10.5281/zenodo.13765792</a>.</span></span></p>

opencc-by-4.0Dec 2023View details →
ClinicalTrials.gov24/100

Tissue Collection for Correlation Between ATM Alterations by Next-Generation Sequencing and ATM Loss-of-Protein

ClinicalTrials.gov study NCT04976803. IPD Sharing: NO. Countries: 2. Publications: 0.

closedIPD-NOFeb 2026View details →
geo24/100

Transcriptome-wide identification of 5-methylcytosine by deaminase and reader protein-assisted sequencing

GEO Series GSE254194. Homo sapiens. 20 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenMar 2025View details →
geo24/100

FMR1 targets distinct mRNA sequence elements to regulate protein expression [PAR-CLIP]

GEO Series GSE39682. Homo sapiens. 6 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenDec 2012View details →
geo24/100

Using combined single-cell gene expression, TCR sequencing and cell surface protein barcoding to characterize and track CD4 T cell clones from murine tissues

GEO Series GSE240041. Mus musculus. 3 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenAug 2023View details →
geo24/100

Structural annotation of equine protein-coding genes determined by mRNA sequencing

GEO Series GSE21925. Equus caballus. 8 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenNov 2010View details →
geo24/100

PROPER-seq (PROtein Protein intERaction sequencing)

GEO Series GSE150818. Homo sapiens. 6 samples. Type: Other.

openGEO-OpenAug 2021View details →
geo24/100

Computational frameworks for predicting protein interactions via single-cell proximity sequencing

GEO Series GSE196130. Homo sapiens. 192 samples. Type: Other.

openGEO-OpenFeb 2022View details →
dryad24/100

Data from: Manipulation of the N-terminal sequence of the Borna disease virus X protein improves its mitochondrial targeting and neuroprotective potential

Open the record for dataset details and reuse information.

publicDec 2015View details →
dryad24/100

The whole protein sequences of the nuclear genome of Chrysosplenium sinicum

Open the record for dataset details and reuse information.

publicAug 2020View details →
dryad24/100

Data from: SNPdryad: predicting deleterious non-synonymous human SNPs using only orthologous protein sequences

Open the record for dataset details and reuse information.

publicJan 2014View details →
geo24/100

Genome-wide binding of tomato AP1/FUL-like proteins determined by DNA-affinity purification sequencing (DAP-seq)

GEO Series GSE271397. Solanum lycopersicum. 21 samples. Type: Other.

openGEO-OpenJul 2025View details →
geo24/100

RNA-sequencing to investigate transcriptional changes caused by ZNF384 fusion proteins in murine pre-B cells

GEO Series GSE112558. Mus musculus. 8 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenOct 2018View details →
geo24/100

Centromere location in Arabidopsis is unaltered by extreme divergence in CENH3 protein sequence

GEO Series GSE88907. Arabidopsis thaliana. 11 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.

openGEO-OpenOct 2016View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record