Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
18,921
datasets available to search
ShareScore release 0.7.1
Dataset results
18,921 results for “Chip”
DCSsim (simulated) and DCSsub (sub-sampled) ChIP-seq data from different chromosomes.
<p>These data are the results from three independent runs of DCSsim and DCSsub for TF, sharp and broad mark signals in 50:50 regulation scenarios for mm10 chr1, chr8, chr11, chr19 and chrX.</p> <p>Simulated data from DCSsim: simulated_ChIP-seq_data.zip</p> <p>set1: TF 50:50 chr11<br> set4: TF 50:50 chr8<br> set7: TF 50:50 chrX<br> set10: TF 50:50 chr1<br> set22: TF 50:50 chr19</p> <p>set2: Sharp mark 50:50 chr11<br> set5: Sharp mark 50:50 chr8<br> set8: Sharp mark 50:50 chrX<br> set11: Sharp mark 50:50 chr1<br> set23: Sharp mark 50:50 chr19</p> <p>set3: Broad mark 50:50 chr11<br> set6: Broad mark 50:50 chr8<br> set9: Broad mark 50:50 chrX<br> set12: Broad mark 50:50 chr1<br> set24: Broad mark 50:50 chr19</p> <p><br> Sub-sampled data from DCSsub: sub-sampled_ChIP-seq_data.zip</p> <p>Set1: C/EBPa-ChIP-seq 50:50 chr11<br> Set2: C/EBPa-ChIP-seq 50:50 chr8<br> Set3: C/EBPa-ChIP-seq 50:50 chrX<br> Set4: C/EBPa-ChIP-seq 50:50 chr1</p> <p>Set5: H3K27ac-ChIP-seq 50:50 chr11<br> Set6: H3K27ac-ChIP-seq 50:50 chr8<br> Set7: H3K27ac-ChIP-seq 50:50 chrX<br> Set8: H3K27ac-ChIP-seq 50:50 chr1</p> <p>Set9: H3K36me3-ChIP-seq 50:50 chr11<br> Set10: H3K36me3-ChIP-seq 50:50 chr8<br> Set11: H3K36me3-ChIP-seq 50:50 chrX<br> Set12: H3K36me3-ChIP-seq 50:50 chr1</p> <p><br> C/EBPa-ChIP-seq 50:50 chr19 can be found in sub-sampled_ChIP-seq_data.zip of the FRIP data set (DOI: 10.5281/zenodo.6042902 set8)<br> H3K27ac-ChIP-seq 50:50 chr19 can be found in sub-sampled_ChIP-seq_data.zip of the FRIP data set (DOI: 10.5281/zenodo.6042902 set9)<br> H3K36me3-ChIP-seq 50:50 chr19 can be found in sub-sampled_ChIP-seq_data.zip of the FRIP data set (DOI: 10.5281/zenodo.6042902 set10)</p>
Data for: Comparison of Friction Extrusion Processing from Bulk and Chips of Aluminum-Copper Alloys
<p>This dataset contains measurement data, machine logs as well as microstructure and overview images for the publication "Comparison of Friction Extrusion Processing from Bulk and Chips of Aluminum-Copper Alloys".</p>
Datasets for predicting TF binding using Virtual ChIP-seq
<p>This repository contains datasets necessary for using the Virtual ChIP-seq software.</p> <p>Virtual ChIP-seq requires the following datasets to predict transcription factor binding:</p> <ul> <li> <p>chipExpDir_AtoH_V1.0.0.tar.gz: Reference matrices of correlation between TF binding and gene expression for TFs starting with letters A-H.</p> </li> <li> <p>chipExpDir_ItoZ_V1.0.0.tar.gz: Reference matrices of correlation between TF binding and gene expression for TFs starting with letters I-Z.</p> </li> <li> <p>refTables_V1.1.0.tar.gz: PhastCons genomic conservation, FIMO PWM scores for JASPAR motifs, and ChIP-seq data of ENCODE and Cistrome database.</p> </li> <li> <p>hg38_chrsize.tsv: Length of chromosomes in hg38</p> </li> <li> <p>trainedModels_V1.0.0.tar.gz: Virtual ChIP-seq scikit-learn trained models saved in joblib format</p> </li> <li> <p><CellType>.tar.gz: Pre-calculated matrices suitable for training with other algorithms or re-training with Virtual ChIP-seq.</p> </li> </ul> <p>Some predictive features of TF binding are the same in each cell type and are stored together for simplicity in refTables_V1.0.0.tar.gz. You can use datasets from other cell types (named here as <CellType>.tar.gz) for the purpose of re-training the model. The <CellType>.tar.gz files contain pre-calculated predictive features of transcription factor binding in 4 chromosomes (5, 10, 15, 20).</p> <p>These features include:</p> <ul> <li> <p>PhastCons genomic conservation</p> </li> <li> <p>FIMO score for sequence motifs of TF in the JASPAR database</p> </li> <li> <p>Chromatin accessibility</p> </li> <li> <p>TF binding in ENCODE + Cistrome DB datasets</p> </li> <li> <p>Virtual ChIP-seq expression score</p> </li> </ul> <p> </p>
Virtual ChIP-seq predictions of binding of 36 transcription factor in Roadmap Epigenomics Project tissues
<p>This dataset contains predictions of Virtual ChIP-seq for binding of 36 transcription factors in Roadmap Epigenomics dataset tissues with matched DNase-seq and RNA-seq data.</p> <p>Tarball contains subfolders for each of the 36 TFs where Virtual ChIP-seq median MCC in validation cell types was > 0.3.</p> <p>Each subfolder contains gzipped BED files. Each file is named as <Tissue>_<Age>_<TF>_<Accession>_Predictions.bed.gz. Columns correspond to Chromosome, Start, End, <Tissue>_<Age>_<TF>_<Accession>, Posterior probability</p> <p>You can use the posterior probabilities provided in Virchip_PosteriorCutoffs_V3.0.0.tsv. These are posterior probability cutoffs which maximized MCC in H1-hESC cell type, or are set to 0.4 if there was no ChIP-seq data of that TF in H1-hESC (0.4 is the mode of all optimal posterior probability cutoffs in H1-hESC).</p>
Collection of Schistosoma mansoni ChIP-Seq input fastq files
<p>These are fastq files of ChIP-Seq input files for different life cycle stages of <em>Schistosoma mansoni</em>.</p> <ul> <li>adult female worms</li> <li>pairs of adults</li> <li>female cercariae</li> <li>miracidia</li> <li>primary sporocysts (sp1)</li> </ul> <p>Produced at IHPE (http://ihpe.univ-perp.fr/)</p>
Experimental Results of "Bioprinting Cell- and Spheroid-Laden Protein-Engineered Hydrogels as Tissue-on-Chip Platforms"
<p>This repository contains the experimental results of the article "Bioprinting Cell- and Spheroid-Laden Protein-Engineered Hydrogels as Tissue-on-Chip Platforms" by Duarte Campos, D., Lindsay, C., Roth, J., LeSavage, B., Seymour, A., Krajina, B., Ribeiro, R., Costa, P., Heilshorn, S., published in <em>Front. bioeng. biotechnol. </em><strong>8, 374</strong> (2020). https://doi.org/10.3389/fbioe.2020.00374</p>
Convolutional Neural Net (CNN) models for ENCODE-Roadmap DNase-seq peaks and Transcription Factor ChIP-seq peaks - Basset architecture
<p>Deep learning models trained on epigenomic landscapes from ENCODE and Roadmap Epigenomics. The models are Basset convolutional neural networks (Kelley, et al 2016). The dataset used to train these models can be found at https://doi.org/10.5281/zenodo.4059038. The file `nn.encode-roadmap.models.basset.clf.tar.gz` contains 10 cross-validated models in Tensorflow framework files as well as details on the architecture, cross-validation scheme, and training of these models. The file `nn.encode-roadmap.models.basset.clf.np_weights.tar.gz` contains the 10 cross-validated models' weights extracted to numpy array files (.npz).</p>
Dataset - A lung-on-chip model reveals an essential role for alveolar epithelial cells in controlling bacterial growth during early M. tuberculosis infection
<p>Description of the sub-folders<br> Name, type of data, corresponding Figure in the manuscript<br> 3D view of the LoC model - .tiff image stack, Figure 1.</p> <p>Bacterial Growth Rate Data - .tiff image stacks, .csv files and MATLAB code to extract the fluorescence intensity over time, Figure 2, Figure 2 - figure supplement 2, Figure 2 - figure supplement 4, Figure 3, Figure 3 - figure supplement 2, Figure 4.</p> <p>AT Characterization - .tiff image stacks and MATLAB code to extract the number and volume of lamellar bodies from the stack of confocal images, Figure 1, Figure 1 - figure supplement 1, Figue 1 - figure supplement 2.</p> <p>AT Infection in LoC model - .tiff image stacks, Figure 2 - figure supplement 1.</p> <p>AT Infection in vivo - .tiff image stacks, Figure 1 - figure supplement 3.</p> <p>Simulations of in vivo infections - .dat files of growth rates in macrophages for the WT and ESX-1 deficient populations and MATLAB code to simulate an infection from this data, Figure 4.</p> <p> </p>
A Two Level Neural Approach Combining Off-Chip Prediction with Adaptive Prefetch Filtering
<p>To alleviate the performance and energy overheads of contemporary applications with large data footprints, we propose the Two Level Perceptron (TLP) predictor, a neural mechanism that effectively combines predicting whether an access will be off-chip with adaptive prefetch filtering at the first-level data cache (L1D). TLP is composed of two connected microarchitectural perception predictors, named First Level Predictor (FLP) and Second Level Predictor (SLP). FLP performs accurate off-chip prediction by using several program features based on virtual addresses and a novel selective delay component. The novelty of SLP relies on leveraging off-chip prediction to drive L1D prefetch filtering by using physical addresses and the FLP prediction as features. TLP constitutes the first hardware proposal targeting both off-chip prediction and prefetch filtering using a multi-level perception hardware approach. TLP only requires 7KB of storage. To demonstrate the benefits of TLP we compare its performance with state-of-the-art approaches using off-chip prediction and prefetch filtering on a wide range of single-core and multi-core workloads. Our experiments show that TLP reduces the average DRAM transactions by 30.7% and 17.7%, as compared to a baseline using state-of-the-art cache prefetchers but no off-chip prediction mechanism, across the single-core and multi-core workloads, respectively, while recent work significantly increases DRAM transactions. As a result, TLP achieves geometric mean performance speedups of 6.2% and 11.8% across single-core and multi-core workloads, respectively. In addition, our evaluation demonstrates that TLP is effective independently of the L1D prefetching logic.</p>
A Two Level Neural Approach Combining Off-Chip Prediction with Adaptive Prefetch Filtering
<p>To alleviate the performance and energy overheads of contemporary applications with large data footprints, we propose the Two Level Perceptron (TLP) predictor, a neural mechanism that effectively combines predicting whether an access will be off-chip with adaptive prefetch filtering at the first-level data cache (L1D). TLP is composed of two connected microarchitectural perception predictors, named First Level Predictor (FLP) and Second Level Predictor (SLP). FLP performs accurate off-chip prediction by using several program features based on virtual addresses and a novel selective delay component. The novelty of SLP relies on leveraging off-chip prediction to drive L1D prefetch filtering by using physical addresses and the FLP prediction as features. TLP constitutes the first hardware proposal targeting both off-chip prediction and prefetch filtering using a multi-level perception hardware approach. TLP only requires 7KB of storage. To demonstrate the benefits of TLP we compare its performance with state-of-the-art approaches using off-chip prediction and prefetch filtering on a wide range of single-core and multi-core workloads. Our experiments show that TLP reduces the average DRAM transactions by 30.7% and 17.7%, as compared to a baseline using state-of-the-art cache prefetchers but no off-chip prediction mechanism, across the single-core and multi-core workloads, respectively, while recent work significantly increases DRAM transactions. As a result, TLP achieves geometric mean performance speedups of 6.2% and 11.8% across single-core and multi-core workloads, respectively. In addition, our evaluation demonstrates that TLP is effective independently of the L1D prefetching logic.</p>
A Two Level Neural Approach Combining Off-Chip Prediction with Adaptive Prefetch Filtering
<p>To alleviate the performance and energy overheads of contemporary applications with large data footprints, we propose the Two Level Perceptron (TLP) predictor, a neural mechanism that effectively combines predicting whether an access will be off-chip with adaptive prefetch filtering at the first-level data cache (L1D). TLP is composed of two connected microarchitectural perception predictors, named First Level Predictor (FLP) and Second Level Predictor (SLP). FLP performs accurate off-chip prediction by using several program features based on virtual addresses and a novel selective delay component. The novelty of SLP relies on leveraging off-chip prediction to drive L1D prefetch filtering by using physical addresses and the FLP prediction as features. TLP constitutes the first hardware proposal targeting both off-chip prediction and prefetch filtering using a multi-level perception hardware approach. TLP only requires 7KB of storage. To demonstrate the benefits of TLP we compare its performance with state-of-the-art approaches using off-chip prediction and prefetch filtering on a wide range of single-core and multi-core workloads. Our experiments show that TLP reduces the average DRAM transactions by 30.7% and 17.7%, as compared to a baseline using state-of-the-art cache prefetchers but no off-chip prediction mechanism, across the single-core and multi-core workloads, respectively, while recent work significantly increases DRAM transactions. As a result, TLP achieves geometric mean performance speedups of 6.2% and 11.8% across single-core and multi-core workloads, respectively. In addition, our evaluation demonstrates that TLP is effective independently of the L1D prefetching logic.</p>
Organ-on-a-Chip (OOC) Image Dataset
<p><strong>Overview: </strong>This dataset contains 3000+ images generated from OOC (organ-on-a-chip) setup with different cell types. The images were generated by an automated brightfield microscopy setup; for each image, such parameters as cell type, time after seeding, and class label ('good' or 'bad' sample quality as assessed by a biology expert) are provided. Furthermore, for some images, seeding density and flow rate are given as well. The dataset can be used for training machine learning classifiers for the automated analysis of the data generated with OOC setup, allowing to create more reliable tissue models and automate decision making processes for growing OOC.</p><p>The dataset comprises images of OOC samples from the following cell lines:</p><ul><li>A549 (human lung adenocarcinoma alveolar basal epithelial cells, CCL-185, ATTC)</li><li>Caco-2 (colorectal adenocarcinoma epithelial cells, HTB-37, ATCC)</li><li>HPMEC (human pulmonary microvascular endothelial cells; 3000, ScienCell)</li><li>HUVEC (human umbilical vein endothelial cells, CRL-1730, ATCC)</li><li>NHBE (normal human bronchial epithelial cells, CC-2541, Lonza)</li><li>HSAEC (human small airway epithelial cells, PCS-301-010, ATCC)</li></ul><p><strong>Structure of the dataset:</strong> The dataset is split into three main folders that correspond to the data split for training machine learning models, i.e., 'train', 'val', and 'test'. The train/val/test split is done proportionally with respect to the class labels, cell lines, and time after seeding (see below), yet the data can be split or merged in other ways to suit the needs of prospective users of the dataset. Within each of the main folders, there are a 'bad' and a 'good' folder with the images corresponding to the respective class labels (see 'Overview' above). The images in 'bad' / 'ģood' folders are further subdivided into folders corresponding to respective cell lines, which are in their turn subdivided into folders corresponding to the different times after seeding. Therefore, it is easy to find images of interest, e.g., '4+ days' 'good' images of the cell line A549 from the 'train' dataset. Further information about the images is available in the file 'OOC_datasheet.xlsx'. </p><p><strong>Acknowledgement:</strong> The work presented in this paper was supported by the project 'AI-improved organ on chip cultivation for personalised medicine (AimOOC)' (contract with Central Finance and Contracting Agency of Republic of Latvia no. 1.1.1.1/21/A/079; the project is co-financed by REACT-EU funding for mitigating the consequences of the pandemic crisis).</p>
Data for: Upcycling Chips‐Bags for Passive Daytime Cooling
<p>The data include all raw data regarding the employed characterization techniques as discussed in the affected publication. These comprise: optical spectroscopy, indoor and field test measurements for passive cooling characterization, RAMAN spectroscopy and scanning electron microscopy</p>
Dataset: Mirror symmetric on-chip frequency circulation of light
<p>The calibrated dataset comprising main text figure 4 and supplementary figure S8 for the paper "Mirror symmetric on-chip frequency circulation of light." The CSV files contain isolation and insertion loss data. The ".m" file contain scripts for plotting the CSV content in MATLAB as heatmaps. Additional details describing the dataset and usage instructions are described in the "readme.txt" file.</p>
Broadband microwave detection using electron spins in a hybrid diamond-magnet sensor chip
<p>Dataset accompanying "Broadband microwave detection using electron spins in a hybrid diamond-magnet sensor chip". </p>
Datasets for 'Label-free imaging of 3D pluripotent stem cell differentiation dynamics on chip'
<p>Here we publish the datasets associated with our publication ‘Label-free imaging of 3D pluripotent stem cell differentiation dynamics<br> on chip’. 3D cultures of human induced pluripotent stem cells (hiPSCs) were imaged while undergoing definitive endoderm (DE) differentiation.<br> Data were acquired either on the live 3D cultures at various time points during the 3 days DE differentiation, or after fixation and immunostaining, with the 3D cultures being fixed at regular 24 h intervals during differentiation, as to form a timeline. Images were recorded with a standard confocal microscope using a xy-resolution of 0.25 μm/px and a z-resolution of 1 μm/plane.</p>
Training data for ChIP-seq data analysis (Galaxy Training Material): Identification of the binding sites of the Estrogen receptor
<p>The data provided here are part of a Galaxy Training Network tutorial that analyzes ChIP-seq data from a study published by Ross-Inness et al., 2012 (DOI:10.1038/nature10730) to identify the binding sites of the Estrogen receptor, a transcription factor known to be associated with different types of breast cancer.</p>
The raw data of Chip-seq in ESM-DBP
<p>The raw sequencing data in fastq format of Chip-seq using GTF2E2 and SUPT6H antibodies on Hale cell.</p>
Data and codes for "Decay-protected superconducting qubit with fast control enabled by integrated on-chip filters"
<p>Data and codes for "Decay-protected superconducting qubit with fast control enabled by integrated on-chip filters".</p>
Virtual ChIP-seq predictions for TF binding in Cistrome and ENCODE-DREAM datasets
<p>Each gzipped tarball contains BED files of Virtual ChIP-seq posterior probabilities.</p> <p>Each BED file corresponds to binding of a TF in one chromosome (one of chr5, chr10, chr15, and chr20) in a validation cell type of Cistrome (virchipCistromePredictions_V2.0.0.tar.gz) or ENCODE-DREAM Challenge datasets (virchipDreamPredictions_V1.0.0.tar.gz).</p> <p> </p> <p>V2.0.0 update:</p> <ul> <li>We provided our predictions on ENCODE-DREAM Challenge validation chromosomes (chr1, chr8, and chr21) in virchipDreamPredictions_V2.0.0.tar.gz</li> <li>FOXA1 predictions for binding in ch5 and chr10 of liver, and CTCF predictions for binding in chr10 of MCF-7 were missing from virchipCistromePredictions_V1.0.0.tar.gz</li> </ul> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.