Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
11,855
datasets available to search
ShareScore release 0.7.1
Dataset results
11,855 results for “Proteins”
Prediction of Conformational Variability for RRM proteins in inter3m data base
<p>Predictions for protein Conformational Variability for the entries in InteR3M (<a href="https://inter3mdb.loria.fr/">https://inter3mdb.loria.fr/</a>), performed with the software ConforMine (in preparation).</p>
Multi-Dimensional Data Viewer (MDV) user manual for data exploration: "Systematic analysis of YFP traps reveals common discordance between mRNA and protein across the nervous system"
<table> <tbody> <tr> <td> <p> Please also see the latest version of the repository:<br> <a href="https://doi.org/10.5281/zenodo.6374011">https://doi.org/10.5281/zenodo.6374011</a> and<br> our website: <a href="https://ilandavis.com/jcb2023-yfp">https://ilandavis.com/jcb2023-yfp</a></p> </td> </tr> </tbody> </table> <p> </p> <p>The explosion in the volume of biological imaging data challenges the available technologies for data interrogation and its intersection with related published bioinformatics data sets. Moreover, intersection of highly rich and complex datasets from different sources provided as flat csv files requires advanced informatics skills, which is time consuming and not accessible to all. Here, we provide a “user manual” to our new paradigm for systematically filtering and analysing a dataset with more than 1300 microscopy data figures using Multi-Dimensional Viewer (MDV) -<a href="https://mdv.molbiol.ox.ac.uk/projects/mdv_project/7012?view=RNA+%2F+Protein+Distribution">link</a>, a solution for interactive multimodal data visualisation and exploration. The primary data we use are derived from our published systematic analysis of 200 YFP traps reveals common discordance between mRNA and protein across the nervous system (<a href="https://doi.org/10.1083/jcb.202205129">eprint link</a>). This manual provides the raw image data together with the expert annotations of the mRNA and protein distribution as well as associated bioinformatics data. We provide an explanation, with specific examples, of how to use MDV to make the multiple data types interoperable and explore them together. We also provide the open-source python code <a href="https://github.com/ilandavislab/Annotate.OMERO.Fig">(github link)</a> used to annotate the figures, which could be adapted to any other kind of data annotation task.</p>
Data set for the journal article: Site-Specific Protein Ubiquitylation Using an Engineered, Chimeric E1 Activating Enzyme and E2 SUMO Conjugating Enzyme Ubc9
<p>Mutations observed in evolved chimeric E1 variants. Top row (1.X to 4.X) describes rounds of evolutions with respective variants in the round. </p> <p>Residues that appear to be enriched are highlighted with gray fill. Star (★) marks residues subjected to saturation mutagenesis in the round 4.</p>
PFAM Protein Families Dataset for Machine Learning
<p>A cleaned dataset of protein sequences and protein families for classification. The dataset is exported from PFAM as of June 2023 and curated to achieve the following characteristics:</p> <ul> <li>only protein families included with >=100 sequences</li> <li>families with >2000 sequences are truncated and only represented by 2000 sequences (chosen randomly)</li> <li>only proteins with sequence lengths between 100 and 1000</li> <li>amino acid sequences are form PDB; chains are concatenated only if not similar</li> </ul> <p>The dataset is not balanced, numbers of sequences per family in PFAM and in in dataset are:</p> <pre><code>families: 62, sequences: 46872 total (in PFAM) -> included (in dataset) Number in family ALLERGEN: 122 -> 122 Number in family APOPTOSIS: 381 -> 381 Number in family BIOSYNTHETIC PROTEIN: 346 -> 346 Number in family BIOTIN BINDING PROTEIN: 165 -> 165 Number in family BLOOD CLOTTING: 138 -> 138 Number in family CALCIUM BINDING PROTEIN: 135 -> 135 Number in family CELL ADHESION: 1116 -> 1116 Number in family CELL CYCLE: 511 -> 511 Number in family CHAPERONE: 964 -> 964 Number in family CONTRACTILE PROTEIN: 158 -> 158 Number in family CYTOKINE: 191 -> 191 Number in family DE NOVO PROTEIN: 253 -> 253 Number in family DNA BINDING PROTEIN: 1008 -> 1008 Number in family ELECTRON TRANSPORT: 841 -> 841 Number in family FLUORESCENT PROTEIN: 348 -> 348 Number in family GENE REGULATION: 607 -> 607 Number in family HORMONE: 272 -> 272 Number in family HORMONE GROWTH FACTOR: 159 -> 159 Number in family HORMONE RECEPTOR: 121 -> 121 Number in family HYDROLASE: 19551 -> 2000 Number in family HYDROLASE ANTIBIOTIC: 120 -> 120 Number in family HYDROLASE HYDROLASE INHIBITOR: 2890 -> 2000 Number in family HYDROLASE INHIBITOR: 315 -> 315 Number in family IMMUNE SYSTEM: 3333 -> 2000 Number in family IMMUNOGLOBULIN: 155 -> 155 Number in family ISOMERASE: 2457 -> 2000 Number in family ISOMERASE ISOMERASE INHIBITOR: 139 -> 139 Number in family LECTIN: 139 -> 139 Number in family LIGASE: 1780 -> 1780 Number in family LIGASE LIGASE INHIBITOR: 163 -> 163 Number in family LIPID BINDING PROTEIN: 421 -> 421 Number in family LIPID TRANSPORT: 115 -> 115 Number in family LUMINESCENT PROTEIN: 221 -> 221 Number in family LYASE: 4150 -> 2000 Number in family LYASE LYASE INHIBITOR: 298 -> 298 Number in family MEMBRANE PROTEIN: 1338 -> 1338 Number in family METAL BINDING PROTEIN: 951 -> 951 Number in family METAL TRANSPORT: 409 -> 409 Number in family MOTOR PROTEIN: 195 -> 195 Number in family OXIDOREDUCTASE: 11531 -> 2000 Number in family OXIDOREDUCTASE OXIDOREDUCTASE INHIBITOR: 766 -> 766 Number in family OXYGEN STORAGE: 127 -> 127 Number in family OXYGEN STORAGE TRANSPORT: 260 -> 260 Number in family OXYGEN TRANSPORT: 414 -> 414 Number in family PHOTOSYNTHESIS: 173 -> 173 Number in family PLANT PROTEIN: 255 -> 255 Number in family PROTEIN BINDING: 1613 -> 1613 Number in family PROTEIN TRANSPORT: 693 -> 693 Number in family RECEPTOR: 108 -> 108 Number in family REPLICATION: 161 -> 161 Number in family RNA BINDING PROTEIN: 546 -> 546 Number in family SIGNALING PROTEIN: 2312 -> 2000 Number in family STRUCTURAL PROTEIN: 869 -> 869 Number in family SUGAR BINDING PROTEIN: 1250 -> 1250 Number in family TOXIN: 546 -> 546 Number in family TRANSCRIPTION REGULATION: 3283 -> 2000 Number in family TRANSFERASE: 14724 -> 2000 Number in family TRANSFERASE INHIBITOR: 126 -> 126 Number in family TRANSFERASE TRANSFERASE INHIBITOR: 2465 -> 2000 Number in family TRANSLATION: 370 -> 370 Number in family TRANSPORT PROTEIN: 2782 -> 2000 Number in family VIRAL PROTEIN: 2150 -> 2000</code></pre> <p>Files:</p> <ul> <li>families.csv: list of protein families with frequencies</li> <li>pfam_46872x62.csv: full dataset with amino acid sequences as string (one-letter code)</li> <li>pfam-trn-xy.csv: training dataset with amino acid sequences as tokens (1..25) and padded to a common length of 1000 with padding token 0:</li> </ul> <pre><code> Amino acid | Token | Description -------------------------------- C | 1 | Cysteine S | 2 | Serine T | 3 | Threonine A | 4 | Alanine G | 5 | Glycine P | 6 | Proline D | 7 | Aspartic acid E | 8 | Glutamic acid Q | 9 | Glutamine N | 10 | Asparagine H | 11 | Histidine R | 12 | Arginine K | 13 | Lysine M | 14 | Methionine I | 15 | Isoleucine L | 16 | Leucine V | 17 | Valine W | 18 | Tryptophan Y | 19 | Tyrosine F | 20 | Phenylalanine B | 21 | Aspartic acid or Asparagine Z | 22 | Glutamic acid or Glutamine J | 23 | Leucine or Isoleucine U | 24 | Selenocysteine X | 25 | Unknown amino acid . | 0 | padding token</code></pre> <p> </p> <ul> <li>pfam-trn-labels.csv: plain-text labels for training data</li> <li>pfam-tst-xy.csv</li> <li>pfam-tst-labels.csv: test data</li> <li>pfam-balanced-trn-xy.csv</li> <li>pfam-balanced-trn-labels.csv:</li> <li>pfam-balanced-tst-xy.csv</li> <li>pfam-balanced-tst-labels.csv: balanced datasets, created by oversampling.</li> </ul>
Project files provided as supporting information to the manuscript "Ligand-protein interactions in lysozyme investigated through a dual-resolution model"
<p><strong>README file for the project files provided as supporting information to the manuscript "Ligand-protein interactions in lysozyme investigated through a dual-resolution model"</strong></p> <p>February 12, 2020</p> <p>Authors: Raffaele Fiorentini, Kurt Kremer and Raffaello Potestio</p> <p>================================</p> <p>Overview</p> <p>The dataset is organised in three (compressed) subfolders (see the tree diagrams in each section):</p> <p>- annihilation<br> - decoupling<br> - density</p> <p>The figure deltaG_binding_ann_dec_comparison.png shows the results of binding free energy calculations comparing the values obtained both for annihilation and decoupling.</p> <p>The figure deltaG_binding_annih_gromacs_espp.png displays the results for Binding FE, comparing the values obtained in GROMACS and ESPResSo++.</p> <p>The README.pdf file contains detailed information about these folders and their content.</p> <p>================================</p> <p>The "annihilation" folder contains all results concerning the calculation of binding free energy in case of annihilation and it is divided in two parts: </p> <p>- complex<br> - ligand</p> <p>In "complex" are reported the results of Ligand-Protein FE both in ESPResSo++ and GROMACS. All simulations are fully-atomistic. </p> <p>In "ligand" are reported the results of ligand solvation free energy both in ESPResSo++ and GROMACS. All simulations are fully-atomistic. </p> <p>====</p> <p>The "decoupling" folder contains all results concerning the calculation of binding free energy in case of decoupling and it is divided in three parts: </p> <p>- complex-DualRes<br> - complex-FullyAT<br> - ligand</p> <p>In "complex-DualRes" are reported the results of Ligand-Protein FE only in ESPResSo++ (GROMACS cannot do decoupling). The system is simulated in Dual-Resolution. It is possible to find the trajectory files in the sub-directories "lambdaindex-0" and "lambdaindex-30".</p> <p>In "complex-fullyAT" are reported the results of Ligand-Protein FE only in ESPResSo++. The system simulated is fully-atomistic. It is possible to find the trajectory file in the sub-directories "lambdaindex-0" and "lambdaindex-30".</p> <p>In "ligand" are reported the results of ligand solvation free energy only in ESPResSo++. All simulations are fully-atomistic. It is possible to find the trajectory file in the sub-directories "lambdaindex-0" and "lambdaindex-20".</p> <p>====</p> <p>The "density" folder contains the data for the tuning of the c parameter of the steric repulsion among residues. This parameter is tuned so that the water density attains the value computed in all-atom simulations.</p>
Protein Graphs Dataset from PDB
<p>This dataset contains the protein graphs constructed from PDB, the Protein Data Bank (www.rcsb.org/pdb), used in the paper:</p> <p>Nilothpal Talukder and <strong>Mohammed J. Zaki</strong>. <strong>A distributed approach for graph mining in massive networks.</strong> <em>Data Mining and Knowledge Discovery: Special Issue on ECML/PKDD 2016 Journal Track Papers</em>, 30(5):1024–1052, 2016. URL: <a href="http://link.springer.com/article/10.1007/s10618-016-0466-x">http://link.springer.com/article/10.1007/s10618-016-0466-x</a>.</p> <p>The format of graphs is as follows:</p> <p>t # GID</p> <p>v VID VLABEL</p> <p>e VID1 VID2 ELABEL</p> <p>where</p> <p>GID is a graph identifier (integer)</p> <p>VID is a vertex identifier (integer) with VLABEL its vertex label (integer)</p> <p>VID1 VID2 denotes an edge between the two vertices, with ELABEL the edge label (integer)</p>
SIRAH-CoV2 initiative: NSP9 RNA binding protein (PDBid:6W4B)
<p>This dataset contains the trajectory of a 10 microseconds-long coarse-grained molecular dynamics simulation of SARS-CoV2 NSP9 RNA binding protein (PDB id: 6W4B, Bioassembly 1). Simulations have been performed using the SIRAH force field running with the Amber18 package at the Uruguayan National Center for Supercomputing (ClusterUY) under the conditions reported in <a href="https://pubs.acs.org/doi/10.1021/acs.jctc.9b00006">Machado et al. JCTC 2019</a>, adding 150 mM NaCl according to <a href="https://pubs.acs.org/doi/10.1021/acs.jctc.9b00953">Machado & Pantano JCTC 2020</a>. </p> <p>The file 6W4B_SIRAHcg_rawdata.tar contains all the raw information required to visualize (on VMD), analyze, backmap, and eventually continue the simulations using Amber18 or higher. Step-By-Step tutorials for running, visualizing, and analyzing CG trajectories using <a href="https://academic.oup.com/bioinformatics/article/32/10/1568/1743152">SirahTools</a> can be found at www.sirahff.com.</p> <p>Additionally, the file 6W4B_SIRAHcg_10us_prot.tar contains only the protein coordinates, while 6W4B_SIRAHcg_10us_prot_skip10ns.tar contains one frame every 10ns.</p> <p>To take a quick look at the trajectory:</p> <p>1- Untar the file 6W4B_SIRAHcg_10us_prot_skip10ns.tar</p> <p>2- Open the trajectory on VMD using the command line:</p> <p>vmd 6W4B_SIRAHcg_prot.prmtop 6W4B_SIRAHcg_prot.ncrst 6W4B_SIRAHcg_prot_10us_skip10ns.nc -e sirah_vmdtk.tcl</p> <p>Note that you can use normal VMD drawing methods as vdw, licorice, etc., and coloring by restype, element, name, etc. </p> <p>This dataset is part of the SIRAH-CoV2 initiative.</p> <p>For further details, please contact Sergio Pantano (spantano@pasteur.edu.uy).</p>
SIRAH-CoV2 initiative: Nucleocapsid protein N-terminal RNA binding domain (PDB id:6M3M)
<p>This dataset contains the trajectory of a 10 microseconds-long coarse-grained molecular dynamics simulation of SARS-CoV2 Nucleocapsid protein N-terminal RNA binding domain (PDB id:6M3M). Simulations have been performed using the SIRAH force field running with the Amber18 package at the Uruguayan National Center for Supercomputing (ClusterUY) under the conditions reported in <a href="https://pubs.acs.org/doi/10.1021/acs.jctc.9b00006">Machado et al. JCTC 2019</a>, adding 150 mM NaCl according to <a href="https://pubs.acs.org/doi/10.1021/acs.jctc.9b00953">Machado & Pantano JCTC 2020</a>. </p> <p>The files 6M3M_SIRAHcg_rawdata.tar contains all the raw information required to visualize (on VMD), analyze, backmap, and eventually continue the simulations using Amber18 or higher. Step-By-Step tutorials for running, visualizing, and analyzing CG trajectories using <a href="https://academic.oup.com/bioinformatics/article/32/10/1568/1743152">SirahTools</a> can be found at www.sirahff.com.</p> <p>Additionally, the file 6M3M_SIRAHcg_10us_prot.tar contains only the protein coordinates, while 6M3M_SIRAHcg_10us_prot_skip10ns.tar contains one frame every 10ns.</p> <p>To take a quick look at the trajectory:</p> <p>1- Untar the file 6M3M_SIRAHcg_10us_prot_skip10ns.tar</p> <p>2- Open the trajectory on VMD using the command line:</p> <p>vmd 6W4B_SIRAHcg_prot.prmtop 6W4B_SIRAHcg_prot.ncrst 6W4B_SIRAHcg_prot_10us_skip10ns.nc -e sirah_vmdtk.tcl</p> <p>Note that you can use normal VMD drawing methods as vdw, licorice, etc., and coloring by restype, element, name, etc. </p> <p>This dataset is part of the SIRAH-CoV2 initiative.</p> <p>For further details, please contact Florencia Klein (fklein@pasteur.edu.uy) or Sergio Pantano (spantano@pasteur.edu.uy).</p>
Exploring Data Evaluation Strategies for Enhanced Identification of Host Cell Proteins in Drug Products of Therapeutic Antibodies and Fc-Fusion Proteins
<p>This data repository contains all previously unpublished raw data files for the manuscript “Exploring Data Evaluation Strategies for Enhanced Identification of Host Cell Proteins in Drug Products of Therapeutic Antibodies and Fc-Fusion Proteins” by Wolfgang Esser-Skala, Marius Segl, Therese Wohlschlager, Veronika Reisinger, Johann Holzmann, and Christian G. Huber. See <em>readme.md</em> for further information.</p>
SIRAH-CoV2 initiative: ORF7A enconded accessory protein (PDB id:6W37)
<p>This dataset contains the trajectory of a 10 microseconds-long coarse-grained molecular dynamics simulation of SARS-CoV2 ORF7a encoded accessory protein (PDB id: 6W37). Simulations have been performed using the SIRAH force field running with the Amber18 package at the Uruguayan National Center for Supercomputing (ClusterUY) under the conditions reported in <a href="https://pubs.acs.org/doi/10.1021/acs.jctc.9b00006">Machado et al. JCTC 2019</a>, adding 150 mM NaCl according to <a href="https://pubs.acs.org/doi/10.1021/acs.jctc.9b00953">Machado & Pantano JCTC 2020</a>. </p> <p>The file 6W37_SIRAHcg_rawdata.tar contains all the raw information required to visualize (on VMD), analyze, backmap, and eventually, continue the simulations using Amber18 or higher. Step-By-Step tutorials for running, visualizing, and analyzing CG trajectories using <a href="https://academic.oup.com/bioinformatics/article/32/10/1568/1743152">SirahTools</a> can be found at www.sirahff.com.</p> <p>Additionally, the file 6W37_SIRAHcg_10us_prot.tar contains only the protein coordinates, while 6W37_SIRAHcg_10us_prot_skip10ns.tar contains one frame every 10ns.</p> <p>To take a quick look at the trajectory:</p> <p>1- Untar the file 6W37_SIRAHcg_10us_prot_skip10ns.tar</p> <p>2- Open the trajectory on VMD using the command line:</p> <p>vmd 6w37_SIRAHcg_prot.prmtop 6w37_SIRAHcg_prot.ncrst 6w37_SIRAHcg_prot_10us_skip10ns.nc -e sirah_vmdtk.tcl</p> <p>Note that you can use normal VMD drawing methods as vdw, licorice, etc., and coloring by restype, element, name, etc. </p> <p>This dataset is part of the SIRAH-CoV2 initiative.</p> <p>For further details, please contact Martín Soñora (msonora@pasteur.edu.uy) or Sergio Pantano (spantano@pasteur.edu.uy).</p>
Dynamics of SARS-CoV-2 spike protein in open and closed states and identification of key structural perturbations upon mutations
<p>The SARS-Cov-2 spike protein resides on the exterior surface of the coronavirus, and therefore, acts as the first point of contact that mediates cell attachment and fusion. During this process, it undergoes dramatic conformational changes upon host receptor binding. We are leveraging high-performance computing to identify these structural perturbations in wildtype and mutant spike protein models. The files contain structures from molecular dynamics simulations of closed SARS-Cov-2 spike protein embedded in POPC membrane.</p>
Scored protein-protein interactions accompanying "A pan-plant protein complex map reveals deep conservation and novel assemblies"
<p><a href="http://plants.proteincomplexes.org/static/data/panplant_cfms_scores_annot.txt.gz">All scored pairwise protein-protein interactions with CF-MS scores (3,076,999 unique pairwise interactions)</a></p> <ul> <li>Description: Scores between Orthogroups with the corresponding CF-MS score and eggNOG generated orthogroup descriptions.</li> <li>Note: Only the highest scoring pairs are considered significant. A CF-MS score >= 0.509 corresponds to 10% FDR, >= 0.207 corresponds to 50% FDR</li> <li>Format: OrthogroupID1 [tab] OrthogroupID2 [tab] Score [tab] Annotation1 [tab] Annotation2</li> </ul>
UnityMol COVID19 Spike protein 360 video
<p>A 360 degree video with a camera path through the COVID19 spike protein-ACE2 complex (example 1 of our paper on biorxiv), illustrating the use of specific cameras to export enriched media.</p> <p>NB: not all video players allow you to experience the 360 degree navigability. The youtube link should work fine.</p>
Human Pleckstrin Homology domain Interacting Protein (PHIP); A Target Enabling Package
<p>SGC Oxford has expressed, purified and crystallized the second bromodomain of PHIP as part of the probe programme. Fragment screening and X-ray crystallography identified binders, some of which optimised to uM affinity. However, molecules with probe properties were not obtained. Consequently it has been decided to put the information generated into the public domain.</p>
Associated Data: RASPD+: Fast protein-ligand binding free energy prediction using simplified physicochemical features
<p>Additional digital data to "RASPD+: Fast protein-ligand binding free energy prediction using simplified physicochemical features" (ChemRxiv preprint:<a href="https://doi.org/10.26434/chemrxiv.12636704.v1">https://doi.org/10.26434/chemrxiv.12636704</a>).</p> <p>Associated code can be found at: <a href="https://github.com/HITS-MCM/RASPDplus">https://github.com/HITS-MCM/RASPDplus</a></p> <p>Files:</p> <ul> <li>weights.tar.gz: contains the model weights of one random dataset split and its associated crossvalidation folds. Used for standard RASPD+ evaluation.</li> <li>additional_model_replicates.tar.gz: contains the remaining models trained on the full set of descriptors.</li> <li>external_test_sets.tar.gz: contains the descriptor tables for all external test sets used</li> <li>dude.tar.gz: contains the descriptor tables for and several identifier lists for evaluation on the Directory of Useful Decoys - Enhanced (DUD-E)</li> <li>run_outputs.tar.gz: Performance metric data and predicted values created during the model training and evaluation runs. Basis for the figures and metrics in the manuscript.</li> </ul> <p> </p>
Elongation of very long chain fatty acids protein 7 (ELOVL7); A Target Enabling Package
<p>The long-chain fatty acid elongases (ELOVL) catalyse the first rate-limiting step in the two carbon elongation of the acyl chains of fatty acids (FAs) greater than 12 carbons in length. Defects in these ELOVL elongases cause severe genetic diseases, such as Stargardt disease-3 and several ataxias, and knockout studies suggest roles in insulin resistance and hepatic steatosis. This TEP provides the first structural information for this family of enzymes which, coupled with mutagenesis and biophysical studies, demonstrates how substrates and products bind within the active site.</p>
Dehydrogenase E1 and transketolase domain-containing protein 1 (DHTKD1); A Target Enabling Package
<p>Inherited mutations of the <em>GCDH </em>gene for glutaryl-CoA dehydrogenase, catalysing the sixth enzymatic step in lysine catabolism, lead to the rare neurometabolic disorder Glutaric Aciduria type 1 (GA1). There is a rationale that inhibition of the fifth lysine catabolising step, catalysed by the DHTKD1 enzyme, could provide therapeutic benefit for GA1 by means of substrate reduction. This TEP provides early tools to develop DHTKD1 inhibitors, including recombinant protein, structure, biophysical (activity and stability) assays and fragment hits of human DHTKD1. This work also reports the interaction of DHTKD1 with its functional partner DLST as a binary complex, and an EM reconstruction of the DLST catalytic core.</p>
List of human genes and their probabilities of being intolerant to heterozygous protein truncating variants.
<p>Recalculation of the Supplementary Table 2 (doi:10.1371/journal.pcbi.1004647.s002) of the journal article "The Characteristics of Heterozygous Protein Truncating Variants in the Human Genome" by Bartha and Rausell published in PLoS Computational Biology (http://dx.doi.org/10.1371/journal.pcbi.1004647). Probabilities in this dataset were computed using human variation data from the Exome Aggregation Consortium (http://exac.broadinstitute.org/).</p> <p>Methods described in that article is relevant for this dataset. All author and affiliation information in that article is relevant for this dataset.</p> <p>Credit for the original human variation data is for the Exome Aggregation Consortium (http://exac.broadinstitute.org/, doi:10.1038/nature19057).</p>
SeMRA Protein Complex Mappings Database
<p>Analyze the landscape of protein complex nomenclature resources, species-agnostic. See instructions for reproduction and usage in the attached README.md.</p>
Protein and DNA alignments for ILS and Entropy calculations
<p>Datasets for reproducing the Entropy and ILS calculations of the 3rd Chapter of my PhD.</p><p>Data include Protein, Exon and Intron alignments for 12 genes.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.