Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
302
datasets available to search
ShareScore release 0.9.0
Dataset results
302 results for “protein sequence”
MP3-seq: Massively parallel measurement of protein-protein interactions by sequencing
GEO Series GSE271790. Saccharomyces cerevisiae. 36 samples. Type: Other.
SARS-CoV-2 vs. Homo sapiens BLASTP protein sequence analysis results
<p>SARS-CoV-2 vs. Homo sapiens BLASTP protein sequence analysis results</p>
The whole protein sequences of the nuclear genome of Chrysosplenium sinicum
<p>We used PacBio and Illumina sequencing to de novo assemble the nuclear genome of Chrysosplenium sinicum, and further annotated the whole protein sequences of the nuclear genome of Chrysosplenium sinicum.</p>
Structure prediction of protein-ligand complexes from sequence information with Umol
<p>posebusters_benchmark_set.tar.zst - files for the prediction (features to Umol) and scoring of the pose busters benchmark </p><p>posebusters_pred_native.tar.zst - pdb and sdf files of proteins and ligands. Includes native structures, predicted structures and relaxed predicted structures with plDDT in the B factor column.</p><p>posebusters_scores.csv - contains ligand RMSD and other metrics for the unrelaxed structures predicted with Umol.</p><p>PDBBind_processed.tar.zst - files for the training (features to Umol) using PDBbind version 2020</p><p> </p><p> </p>
Predicting biophysical characteristics of proteins from their amino acid sequence
<p>As part of the Galaxy Training Material, this is the tutorial workflow's input file in FASTA format. It contains the amino acid sequence of protein P04156 ("PRIO_HUMAN OS=Homo sapiens OX=9606 GN=PRNP PE=1 SV=1").</p>
Representations and associated fitness values of protein sequences
<p>We studied the ability to predict protein fitness from sequence using our method <a href="https://github.com/amillig/MERGE" target="_blank" rel="noopener">MERGE</a> and other methods (i.e., ECNet, eUniRep, EVmutation, One-Hot, and UniRep). The following files are included in this repository:</p> <ul> <li>The folder <em>ECNet</em> contains the braw, csv, and fasta files required to run <a href="https://github.com/luoyunan/ECNet" target="_blank" rel="noopener">ECNet</a>.</li> <li>The folders <em>eUniRep, </em><em>MERGE, One-Hot</em>, and <em>UniRep</em> contain csv files including the names and fitness values of protein variants as well as a numerical representation of their sequence. The folder <em>MERGE </em>also contains representations of the wild type sequences as npy files in the <em>wts</em> folder.</li> <li>The folder <em>Params</em> contains params files generated with <a href="https://github.com/debbiemarkslab/plmc" target="_blank" rel="noopener">PLMC</a>.</li> <li>The script <em>get_performances.py</em> enables to generate models using different methods (i.e., eUniRep, EVmutation, MERGE, One-Hot, Pure_ML, and UniRep) and to determine their performance for predicting the fitness of protein variants from sequence.</li> </ul> <p><strong>References</strong></p> <table> <tbody> <tr> <td><a href="https://doi.org/10.1038/s41467-021-25976-8" target="_blank" rel="noopener">ECNet</a></td> <td> <p>Luo, Y., Jiang, G., Yu, T. et al. ECNet is an evolutionary context-integrated deep learning framework for protein engineering. Nat Commun 12, 5743 (2021).</p> </td> </tr> <tr> <td><a href="https://doi.org/10.1038/s41592-021-01100-y" target="_blank" rel="noopener">eUniRep</a></td> <td>Biswas, S., Khimulya, G., Alley, E.C. et al. Low-N protein engineering with data-efficient deep learning. Nat Methods 18, 389–396 (2021)</td> </tr> <tr> <td><a href="https://doi.org/10.1038/nbt.3769" target="_blank" rel="noopener">EVmutation</a></td> <td>Hopf, T., Ingraham, J., Poelwijk, F. et al. Mutation effects predicted from sequence co-variation. Nat Biotechnol 35, 128–135 (2017).</td> </tr> <tr> <td><a href="https://doi.org/10.1038/s41592-019-0598-1" target="_blank" rel="noopener">UniRep</a></td> <td>Alley, E.C., Khimulya, G., Biswas, S. et al. Unified rational protein engineering with sequence-based deep representation learning. Nat Methods 16, 1315–1322 (2019).</td> </tr> </tbody> </table> <p> </p>
A joint embedding of protein sequence and structure enables robust variant effect predictions
<p>Data related to the GitHub repository KULL-Centre/_2023_Blaabjerg_SSEmb, which is also stored on Zenodo here: <span><span><a href="../doi/10.5281/zenodo.13765792" target="_blank" rel="noopener noreferrer">https://zenodo.org/doi/10.5281/zenodo.13765792</a>.</span></span></p>
Tissue Collection for Correlation Between ATM Alterations by Next-Generation Sequencing and ATM Loss-of-Protein
ClinicalTrials.gov study NCT04976803. IPD Sharing: NO. Countries: 2. Publications: 0.
Transcriptome-wide identification of 5-methylcytosine by deaminase and reader protein-assisted sequencing
GEO Series GSE254194. Homo sapiens. 20 samples. Type: Expression profiling by high throughput sequencing.
FMR1 targets distinct mRNA sequence elements to regulate protein expression [PAR-CLIP]
GEO Series GSE39682. Homo sapiens. 6 samples. Type: Expression profiling by high throughput sequencing.
Using combined single-cell gene expression, TCR sequencing and cell surface protein barcoding to characterize and track CD4 T cell clones from murine tissues
GEO Series GSE240041. Mus musculus. 3 samples. Type: Expression profiling by high throughput sequencing.
Structural annotation of equine protein-coding genes determined by mRNA sequencing
GEO Series GSE21925. Equus caballus. 8 samples. Type: Expression profiling by high throughput sequencing.
PROPER-seq (PROtein Protein intERaction sequencing)
GEO Series GSE150818. Homo sapiens. 6 samples. Type: Other.
Computational frameworks for predicting protein interactions via single-cell proximity sequencing
GEO Series GSE196130. Homo sapiens. 192 samples. Type: Other.
Data from: Manipulation of the N-terminal sequence of the Borna disease virus X protein improves its mitochondrial targeting and neuroprotective potential
Open the record for dataset details and reuse information.
The whole protein sequences of the nuclear genome of Chrysosplenium sinicum
Open the record for dataset details and reuse information.
Data from: SNPdryad: predicting deleterious non-synonymous human SNPs using only orthologous protein sequences
Open the record for dataset details and reuse information.
Genome-wide binding of tomato AP1/FUL-like proteins determined by DNA-affinity purification sequencing (DAP-seq)
GEO Series GSE271397. Solanum lycopersicum. 21 samples. Type: Other.
RNA-sequencing to investigate transcriptional changes caused by ZNF384 fusion proteins in murine pre-B cells
GEO Series GSE112558. Mus musculus. 8 samples. Type: Expression profiling by high throughput sequencing.
Centromere location in Arabidopsis is unaltered by extreme divergence in CENH3 protein sequence
GEO Series GSE88907. Arabidopsis thaliana. 11 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.