Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
2,848
datasets available to search
ShareScore release 0.9.0
Dataset results
2,848 results for “sequence data”
scooby: Modeling multi-modal genomic profiles from DNA sequence at single-cell resolution - Supplementary data and code
<p>Data and code to reproduce the analyses from the study: "scooby: Modeling multi-modal genomic profiles from DNA sequence at single-cell resolution". </p>
Sequence-independent, site-specific incorporation of chemical modifications to generate light-activated plasmids (Source Data)
<p>Source data for Chemical Science paper "Sequence-independent, site-specific incorporation of chemical modifications to generate light-activated plasmids" DOI: 10.1039/D3SC02761A</p>
Spatial transcriptome mapping of the desmoplastic growth pattern of colorectal liver metastases by in situ sequencing - Image and In SItu Sequencing data
<p>Image and ISS data for Spatial transcriptome mapping of the desmoplastic growth pattern of colorectal liver metastases by in situ sequencing reveals a biologically relevant zonation of the desmoplastic rim.</p><p>Image data consists of:</p><ul><li>Nucler stain (DAPI)</li><li>Masks for liver, rim and tumor regions</li><li>H&E images of parallel tissue sections</li></ul><p>Gene and cluster marker data is collected in the <i>markers.h5ad</i> file which can be read using AnnData (<a href="https://anndata.readthedocs.io/en/latest/">https://anndata.readthedocs.io/en/latest/)</a>. </p><p> </p>
FIGURE 2. Phylogenetic results. A, Maximum likelihood tree from COI dataset rooted with Ophelia limacina. B, Maximum likelihood tree from ITS1 in Validation of three sympatric Thoracophelia species (Annelida: Opheliidae) from Dillon Beach, California using mitochondrial and nuclear DNA sequence data
FIGURE 2. Phylogenetic results. A, Maximum likelihood tree from COI dataset rooted with Ophelia limacina. B, Maximum likelihood tree from ITS1 dataset rooted according to the result for the COI dataset. Support values are shown as jackknife from parsimony analysis and bootstrap from maximum likelihood respectively separated by /. * indicates 100% values for each support measure.
FIGURE 1. The three sympatric Thoracophelia spp. from Dillon Beach. A, Thoracophelia dillonensis. B in Validation of three sympatric Thoracophelia species (Annelida: Opheliidae) from Dillon Beach, California using mitochondrial and nuclear DNA sequence data
FIGURE 1. The three sympatric Thoracophelia spp. from Dillon Beach. A, Thoracophelia dillonensis. B, Pectinate branchiae of T. dillonensis. C, Thoracophelia williamsi. D, Bifurcated branchiae with pinnules of T. williamsi. E, Thoracophelia mucronata. F, Bifurcated branchiae of T. mucronata. Scale bars all 1 mm.
Integrated single-cell RNA-sequencing data of unwounded and wounded mouse skin and fibroblasts.
<p>This repository contains the .h5ad files that store the integrated scRNA-seq data we generated for the work, Almet et al. (2023), "Fibroblasts evolve in single-cell state to drive extracellular matrix and signaling changes across wound healing", to be published in the Journal of Investigative Dermatology.</p><p>The integrated* files contain both raw counts, normalized counts, as well as unspliced and spliced count estimates that were obtained using kallisto|bustools and velocyto. We integrated the data from the following published datasets:</p><ol><li><a href=" https://doi.org/10.7554/eLife.60066">Phan et al. (2021)</a>: Unwounded P21 mice and small wound P21 + 7 mice</li><li><a href="https://doi.org/10.1016/j.celrep.2020.02.091">Haensel et al. (2020)</a>: Unwounded P49 mice and small wound P49 + 4 mice</li><li><a href="https://doi.org/10.1038/s41467-018-08247-x)">Guerrero-Juarez et al. (2019)</a>: Large wound day 12 mice</li><li><a href="https://doi.org/10.1016/j.stem.2020.07.008">Abbasi et al. (2020)</a>: Large wound day 14 mice</li><li><a href="https://doi.org/10.1126/sciadv.aay3704">Gay et al. (2020)</a>: Large wound fibrotic (hairless) and regenerative (hair follicle neogenesis) day 18 mice</li></ol><p>The unwounded_* files were used to briefly integrated unwounded skin scRNA-seq from mouse models of different ages that have been used to analyze wound healing in <a href="https://doi.org/10.1016/j.celrep.2020.02.091">Haensel et al. (2020),</a> <a href=" https://doi.org/10.7554/eLife.60066">Phan et al. (2021)</a>, and <a href="https://doi.org/10.1016/j.celrep.2022.111155">Vu et al. (2022)</a>, which generated scRNA-seq for unwounded skin from mice aged P21, P49, and P616, respectively. </p><p>The data can be loaded using the Python package Scanpy or AnnData, but you can also load it in R if you use zellkonverter. </p>
Sequence/simulation data for Direct Prediction of Intrinsically Disordered Protein Conformational Properties From Sequence
<p>This is a DOI-linked deposition of sequence/biophysical properties pairs used in the associated paper by Lotthammer et al:</p><p>Lotthammer, J. M.<strong>*</strong>, Ginell, G. M.<strong>*</strong>, Griffith, D.<strong>*</strong>, Emenecker, R. J. & Holehouse, A. S. <br>Direct Prediction of Intrinsically Disordered Protein Conformational Properties From Sequence.<br><i><strong>Nature Methods</strong></i> (<i>in press</i>), (2023).</p><p> </p>
Supplementary Data: Uncovering DFG-out sequence propensity determinants of kinases with machine learning
<h2>General description</h2><p>This submission accompanies the paper "Uncovering DFG-out sequence propensity determinants of kinases with machine learning" and covers:</p><ul><li><strong>Trained models (train/models_latest)</strong>. A subset of these is also published as a part of the <a href="https://github.com/edikedik/kinactive">KinActive tool</a>; each model is of `KinactiveClassifier` type and can be loaded using this tool as well via `kinactive.io.load()` function).</li><li><strong>Patched sequences (patched_seqs)</strong>. PDB sequences, with missing regions patched by UniProt sequences. This is a collection of <a href="https://github.com/edikedik/lXtractor">lXtractor</a> ChainSequence objects.</li><li><strong>All labels, variables, and datasets</strong> (datasets and labels).</li><li><strong>Initial chains and predictions for SwissProt proteins</strong> (SP_predictions).</li></ul><h2>Note on model abbreviations</h2><p>The main text emphasized datasets that yielded more interpretable results. These were constructed from domain sequences, labeled as apo, inactive, DFG-in, or DFG-out, and further divided into TK and STK subsets. We refer to these datasets with the abbreviation <strong>AAIO</strong> (Apo All In or Out) to distinguish them from additional datasets.</p><p>In addition to <strong>AAIO</strong>, we explored two alternative labeling strategies:</p><ul><li><strong>AHAO</strong> (Apo/Holo Any Out): This dataset includes sequences from ligand-bound entries. All sequences within 95\% identity clusters are labeled as DFG-out if the cluster contains at least one sequence in this state. All others are labeled as DFG-in.</li><li><strong>AAO</strong> (Apo Any Out): This dataset excludes sequences corresponding to ligand-bound entries but includes those with conflicting conformational tendencies within 95\% identity clusters.</li></ul><p>Each of these datasets had two versions:</p><ol><li>A seed version, denoted by a "*" symbol.</li><li>A version enriched with orthologous sequences (no special designation).</li></ol><p>Together with the <strong>TkST</strong> datasets (encompassing TK and STK labels) used for testing the methodology, a total of 14 datasets were used, and both <i>RF</i> and <i>XGB</i> models were applied to each, using the same initial settings, resulting in 28 different models. This additional information is provided for completeness.</p>
VCF file of whole genome sequencing data of 163 rats mapped jointly to mRatBN7.2
<p>We analyzed whole genome sequencing data of 163 rats. These data were first mapped to mRatBN7.2, followed by variant calling using deepvariant and joint analysis using GLNexus. Sites that are likely called due to base-level errors in mRatBN7.2 are removed. </p>
Genotype likelihoods for low-coverage whole-genome sequencing data of yellow warblers
<p>The following datasets include the required input files used to empirically test population assignment in WGSassign on Yellow Warbler data. The file "yewa.known.ind105.ds_2x.beagle.gz" includes the filtered variants of 105 Yellow Warbler individuals output as genotype likelihoods and stored in a Beagle-formatted file. The ID file, "yewa.known.ind105.reference.IDs.txt", is a tab-delimited file with 2 columns, the first being the sample ID, and the second being the known reference population. The sample order in the ID file should match that of the input beagle file. To measure the assignment accuracy of WGSassign, we used leave-one-out cross validation using the input beagle file and our ID file.</p>
Using big sequencing data to identify chronic SARS-Coronavirus-2 infections
<p>This dataset supports the "Using big sequencing data to identify chronic SARS-Coronavirus-2 infections" publication in Nature Communications. </p><p>It contains the following files and folders:<br> </p><ul><li>Readme.txt - detailed information on all folders and files.</li><li>Supplementary Dataset 1</li><li>Supplementary Dataset 2</li><li>Supplementary Dataset 3</li><li>Supplementary Dataset 4</li><li>Supplementary Dataset 5</li><li>Supplementary Dataset 6</li><li>Supplementary Dataset 7</li></ul>
cDNA sequence of E2 gene family in Arabidopsis thaliana and data of statistical analysis
<p>E2 ubiquitin-conjugating enzymes act as a heart role in the ubiquitination process and are responsible for catalysis ubiquitin transfer. Although the function of ubiquitin-protein ligases (E3s) in plant response to diverse abiotic stress by targeting specific substrates has been well studied, the E2s' involvement in environmental responses and their downstream targets are not well understood. Here, we demonstrated that the E2 ubiquitin-conjugating enzyme 18 (UBC18) regulates the stability of FREE1 to modulate iron deficiency stress. UBC18 affects the ubiquitination of FREE1 and promotes its degradation, overexpression of<em> UBC18</em> in plants decreases their sensitivity to iron deficiency by reducing the level of FREE1, and high accumulation of FREE1 in<em> </em>the<em> ubc18</em> mutant resulted in sensitivity to iron deficiency. In addition, we demonstrated the lysine residues K227, K295, K315, and K540 are required for FREE1 ubiquitination and stability regulation, and mutation of these lysines of FREE1 residues resulted in sensitivity to iron starvation in plants. Taken together, our findings reveal a mechanism of UBC18 in response to iron deficiency stress by altering the abundance of FREE1, and further elucidate the role of ubiquitination sites in FREE1 stability regulation and the plant iron deficiency response.</p>
RNA sequencing data from the guts of Drosophila wildtype and CG3740/nazo double mutant 20-day-old males
<p>Lipid dyshomeostasis has been implicated in a variety of diseases ranging from obesity to neurodegenerative disorders such as Neurodegeneration with Brain Iron Accumulation (NBIA). Here, we uncover the physiological role of Nazo, the Drosophilamelanogaster homolog of the NBIA-mutated protein – c19orf12, whose function has been elusive. Ablation of Drosophila c19orf12 homologs leads to dysregulation of multiple lipid metabolism genes. nazo mutants exhibit markedly reduced gut lipid droplet and whole-body triglyceride contents. Consequently, they are sensitive to starvation and oxidative stress. Nazo is required for maintaining normal levels of Perilipin-2, an inhibitor of the lipase – Brummer. Concurrent knockdown of Brummer or overexpression of Perilipin-2 rescues the nazo phenotype, suggesting that this defect, at least in part, may arise from diminished Perilipin-2 on lipid droplets leading to aberrant Brummer-mediated lipolysis. Our findings potentially provide novel insights into the role of c19orf12 as a possible link between lipid dyshomeostasis and neurodegeneration, particularly in the context of NBIA.</p>
MRI Raw Data of Cardiac Radial Real-Time BSSFP Sequence (Part 1/5)
<p>Other parts of this dataset are available at:</p> <div> <div>* Part 1: 10.5281/zenodo.10492333</div> <div>* Part 2: 10.5281/zenodo.10912299</div> <div>* Part 3: 10.5281/zenodo.10492343</div> <div>* Part 4: 10.5281/zenodo.10492455</div> <div>* Part 5: 10.5281/zenodo.10493095</div> </div>
MRI Raw Data of Cardiac Radial Real-Time BSSFP Sequence (Part 3/5)
<p>Other parts of this dataset are available at:</p> <div> <div>* Part 1: 10.5281/zenodo.10492333</div> <div>* Part 2: 10.5281/zenodo.10912299</div> <div>* Part 3: 10.5281/zenodo.10492343</div> <div>* Part 4: 10.5281/zenodo.10492455</div> <div>* Part 5: 10.5281/zenodo.10493095</div> </div>
MRI Raw Data of Cardiac Radial Real-Time BSSFP Sequence (Part 4/5)
<p>Other parts of this dataset are available at:</p> <div> <div>* Part 1: 10.5281/zenodo.10492333</div> <div>* Part 2: 10.5281/zenodo.10912299</div> <div>* Part 3: 10.5281/zenodo.10492343</div> <div>* Part 4: 10.5281/zenodo.10492455</div> <div>* Part 5: 10.5281/zenodo.10493095</div> </div>
STAMINA project 883441 related raw sequencing data of RTPCR positive SARS-CoV-2 amplicons
Open the record for dataset details and reuse information.
The 3'-RACE data (InPACT: A computational method for accurate characterization of intronic polyadenylation from RNA sequencing data)
Open the record for dataset details and reuse information.
Raw k-space data for free-breathing diffusion weighted sequence
Open the record for dataset details and reuse information.
Supplementary data: Predicting grid frequency short-term dynamics with Gaussian processes and sequence modeling
<p>This repository contains data and result files for the paper "Predicting grid frequency short-term dynamics with Gaussian processes and sequence modelling". The code to generate the models and reproduce the results of the comparative study in the above paper is available on this <a href="https://github.com/bolin-liu/sequence-model-and-gaussian-process-for-frequency-prediction">github repository</a></p> <p><strong>Supplementary data</strong>:</p> <p>- The <strong>trained_models</strong> folder contains the results of the trained models.</p> <p>- The folder <strong>data</strong> contains data needed for for the comparative study for the year 2019 in the paper above. This data set (except knn_point_predictions.npy) is generated with the code in this <a href="https://github.com/johkruse/PIML-for-grid-frequency-modelling">github repository</a>. knn_point_predictions.npy is generated with the code in this <a href="https://github.com/bolin-liu/sequence-model-and-gaussian-process-for-frequency-prediction">github repository </a>.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.