Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

3,481

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

3,481 results for “data set”

Learn how ShareScore rates datasets ↗
geo16/100

UCSD GBM Data Set

GEO Series GSE60184. Homo sapiens. 23 samples. Type: Expression profiling by array.

openGEO-OpenAug 2014View details →
geo16/100

Intertwining threshold settings, biological data and database knowledge to optimize the selection of differentially expressed genes

GEO Series GSE22858. Homo sapiens. 6 samples. Type: Expression profiling by array.

openGEO-OpenJun 2011View details →
geo16/100

Renal allograft dysfunction (kidney tissue data set)

GEO Series GSE26578. Homo sapiens. 113 samples. Type: Expression profiling by array.

openGEO-OpenFeb 2011View details →
geo16/100

DNA methylation dynamics of oocyte in follicular maturation (RRBS-Seq data set)

GEO Series GSE111684. Mus musculus. 31 samples. Type: Methylation profiling by high throughput sequencing.

openGEO-OpenJun 2021View details →
geo16/100

Dynamic rewiring of transcription factor networks during smooth muscle cell phenotypic modulation (ChIP-Rx data sets)

GEO Series GSE111710. Rattus norvegicus. 20 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.

openGEO-OpenJan 2019View details →
geo16/100

Mutating Zta(N182) to S, Q, T, I, and V changes sequence specific DNA binding to four types of DNA (32k data set)

GEO Series GSE126589. synthetic construct; Mus musculus. 24 samples. Type: Other.

openGEO-OpenJan 2020View details →
geo16/100

Expression data from engorged Haemaphysalis flava ticks (ET) salivary gland and semi-engorged Haemaphysalis flava ticks (SET) salivary gland (SGs)

GEO Series GSE67247. Haemaphysalis flava. 2 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenJan 2018View details →
geo16/100

Id2-deficient NK cells acquire a naïve-like fate (ATAC-seq data set)

GEO Series GSE109517. Mus musculus. 6 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.

openGEO-OpenFeb 2018View details →
geo16/100

Mutations in Bcl9 and Pygo genes cause congenital heart defects by tissue-specific perturbation of Wnt/β-catenin signaling (ChIP-seq data set)

GEO Series GSE110781. Mus musculus. 7 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.

openGEO-OpenFeb 2021View details →
geo16/100

Identification of a robust gene signature that predicts breast cancer outcome in independent data sets [UCSF]

GEO Series GSE123833. Homo sapiens. 0 samples. Type: Expression profiling by array; Third-party reanalysis.

openGEO-OpenDec 2018View details →
geo16/100

Transcriptome Cappable-Seq Sequencing Data Set of Gene Expression in Salmonella enterica serovar Typhimurium 14028S inside Acanthamoeba castellanii, and under Oxidative and Starvation Stress Condition

GEO Series GSE271311. Salmonella enterica subsp. enterica serovar Typhimurium str. 14028S. 20 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenOct 2024View details →
geo16/100

Mesenchymal transition and acquired therapy resistance of glioblastoma cells is connected to astrocyte reactivity (DNA methylation data set)

GEO Series GSE122808. Homo sapiens. 14 samples. Type: Methylation profiling by genome tiling array.

openGEO-OpenNov 2020View details →
geo16/100

Mutating Zta(N182) to S, Q, T, I, and V changes sequence specific DNA binding to four types of DNA (65k data set)

GEO Series GSE126588. synthetic construct; Mus musculus. 36 samples. Type: Other.

openGEO-OpenJan 2020View details →
zenodo16/100

MLCQProjects: Towards evolving data set of industry-relevant software projects

<p>Short: This data set contains three snapshots from an evolving software project data set and a script that generates those snapshots. Projects in snapshots are annotated with features used to assess industry-relevance - documentation links, installation means, support channels and whether the project delivers functionalities or samples.</p> <p>Context:&nbsp;Researchers involved in mining software repositories face a challenge that many of existing data sets reflect how software was developed in the past (typically many years ago), instead of in the present. Another challenge is that data sets based on open-source software projects include projects that are not necessarily industry-relevant.</p> <p>Aim:&nbsp;The aim of this paper is to address the aforementioned challenges and provide both: 1) a snapshot-based evolving data set of software projects (thus reflecting their present, as well as previous states), manually enriched with data considered important for industrial relevance assessment, and 2) a method of assessing industrial relevance of software projects.</p> <p>Method:&nbsp;We present a systematic method of selecting data sets of software projects for the purposes of mining software repositories of potentially industry-relevant projects and a semi-systematic method of assessing the industrial relevance of those projects.</p> <p>Data set:&nbsp;The data set contains three snapshots (spanning over 10 months) of popular Java projects from GitHub, manually enriched with industrial relevance and maintenance-related information. The presented acquisition method is sufficient to generate further snapshots and one should be able to assess the industrial relevance of the projects using the presented assessment method. Provided data set open directions of further research, e.g, 1) evaluation of code smells or defect prediction models on industry-relevant software projects prepared by independent authors and not used to build models, 2) analysis of some social aspects of open source projects (such as tooling used for providing support) and how they evolve in time.</p>

restrictedFeb 2020View details →
zenodo16/100

Introduction to curating and managing research data for re-use: natural gas infrastructure case study data set

<p>This data set underpins the report and data provided between 2004-2006 for registration in the restrictions application of Estonian Land Board.</p>

restrictedFeb 2020View details →
zenodo16/100

Data sets for Perspective: SARS-CoV-2's potential mechanism of regulating cellular responses through depletion of specific host miRNAs

<p>Using a bioinformatic analysis of human miRNA potential interactions with the SARS-CoV-2&rsquo;s genome, we examined the potential miRNA target sites in 7 coronavirus genomes that include SARS-CoV-2, MERS-CoV, SARS-CoV, and 4 non-pathogenic coronaviruses. (<strong>Data Set 1</strong>).</p> <p><em>Our approach was to examine and compare 3 pathogenic and 4 non-pathogenic strains of HCoVs. The HCoVs&#39; RNA genomes of pathogenic strains were SARS-CoV-2 (NC_045512.2), SARS-CoV (NC_004718.3), MERS-CoV (NC_019843.3). The non-pathogenic strains were HCoV-OC43 (KU131570.1), HCoV-229E (NC_002645.1), HCoV-HKU1 (KF686346.1), and HCoV-NL63 (NC_005831.2). These coronaviruses were tested against the set of 896 confident mature human miRNA sequences that were obtained from the miRBbase v2.21 using the RNA22 v2 microRNA target discovery tool web-server. In order to reduce the false discovery rate of the MTS predictions, the most strict parameters were applied to the default computation workflow using a specificity of 92% versus a sensitivity of 22%.</em></p> <p><em>In <strong>Data set 2</strong>, using the miRDIP database with only top 1% of the most probable targets considered, we analyzed the potential targets of miRNA that could be bound to either the pathogenic, the non-pathogenic or both groups of HCoVs. </em></p> <p>In <strong>Data set 3</strong>, using miRNAFold webserver&nbsp; we identified 10 pre-miRNA sequences in the <em>SARS-CoV-2</em> RNA sequence that could potentially enter the human RNAi pathway.</p> <p>The graphical summary of our working hypothesis is provided in the<strong> graphical abstract</strong>.</p> <p>&nbsp;</p> <p>&nbsp;</p>

restrictedMay 2020View details →
zenodo16/100

Data set and results of regression analysis of gene variants in Prostate Cancer patients from Mexico-mestizo population.

<p>Genotyping assays were performed using TaqMan SNP Genotyping Assays probes C_27532228_20 and C_2362601_10 (Life Technologies, Carlsbad, CA, USA).</p> <p>Thermalcycler:&nbsp;StepOnePlus&trade; Real-Time PCR (Thermo Fisher Scientific, Waltham, MA, USA).</p> <p>Sample: genomic DNA from blood.</p> <p>Subjects: Prostate cancer patients and controls (healthy and benign prostatic hyperplasia).&nbsp;</p> <p>Results of Regression Anlysis for clinical features of gene SRD5A2 variants (rs9282858 and<strong>&nbsp;</strong>rs523349) in prostate cancer patients from Mexican-mestizo population.</p> <p>Starting with the full model, successive models are created, each one using one less regressor (or covariate) than the previous model.&nbsp;Each of the regressors currently in the model is removed to create a &quot;trial&quot; model excluding that regressor. The p-value of the current model (or full model) versus the trial model (or reduced model) is calculated, and the model with the smallest p-value is used as the next model. This method removes the least significant variable from the current model. If every p-value is smaller than the p-value cut off specified, the backward elimination method stops. The method also stops if all variables have been removed from the model, or if all variables left are included in the original reduced model. From the standpoint of further analysis, the final model becomes the &quot;full model&quot; for this set of potential regressors</p> <p>Ethics approval and consent to participate in the study: The protocol and informed consent were approved by the Ethics and Research Committee of the School of Medicine (Universidad Autonoma de Nuevo Leon), with registration number UR16-00007. In the data set, no data was included that compromises the confidentiality of the participating subjetcs.</p> <pre> &nbsp;</pre>

restrictedJul 2020View details →
zenodo16/100

Data set from "Combining fiber Brillouin amplification with a repeater laser station for fiber-based optical frequency dissemination over 1400 km"

<p>The data set contains the data underlying the fiber link performance evaluation published in Koke et al 2019 New J. Phys. 21 123017 (https://doi.org/10.1088/1367-2630/ab5d95). The experimental setup and the methodology used is explained in this publication.</p> <p>The data is stored in the Matlab(R)-native file format. Although proprietary, import of data from this file format is supported by other numerical computing environments, too. An example script for importing into Python is included.</p> <p>The files &#39;data_export_campaign_*.mat&#39; contain the timeseries data shown in Figures 2, 3 and 4. Each of these files include variables with the following meaning:<br> &#39;date_year&#39;, &#39;date_month&#39;, &#39;date_day&#39;, &#39;time_hour&#39;, &#39;time_minute&#39;, &#39;time_second&#39;: Timestamp of the data sample in UTC time<br> &#39;inloop1&#39; and &#39;inloop2&#39;: $\Lambda_{1s}$ frequency offset from nominal value in Hz of the beat signal used for stabilizing the uplink<br> &#39;aom&#39;: $\Lambda_{1s}$ frequency offset from nominal value in Hz of the uplink servo acoustic-optical modulator&#39;s drive frequency<br> &#39;remote1&#39; and &#39;remote2&#39;: $\Lambda_{1s}$ frequency offset from nominal value in Hz of the roundtrip beat signal used for out-of-loop characterization of the frequency transfer error<br> &#39;no_of_run_in_NJP_21_123017&#39;: A flag indicating the measurement runs discussed in the paper. Samples with the same non-zero integer values belong to the same measurement run. Samples marked with 0 did not enter the publication.<br> &#39;DataCollected_FBAPTBLocked_CavityLocked&#39;: Result of our current monitoring of the fulfillment of the prerequisites for fiber link performance evaluation; values of 1/True indicate that data logging was active, the transfer laser was locked to the signal of the ultra-stable cavity, and the lock of the FBA(PTB) pump laser was active.</p> <p>The file &#39;data_export_uplink_downlink_aom.mat&#39; contains the data underlying the fiber phase noise correlation analysis in Fig. 5.<br> &#39;aom_uplink&#39;: $\Lambda_{1s}$ frequency offset from nominal value in Hz of the uplink servo acoustic-optical modulator&#39;s drive frequency<br> &#39;aom_downlink&#39;: $\Lambda_{1s}$ frequency offset from nominal value in Hz of the downlink servo acoustic-optical modulator&#39;s drive frequency</p> <p>Since publication of the paper, we discovered a slight inconsistency of the nominal frequencies used in our analysis. Hence, remote fractional frequency offsets published in Koke et al 2019 New J. Phys. 21 123017 have to be corrected by subtracting fractional frequency values of 4.2E-22 (campaigns 2015-06, 2016-03, 2018-03) and 3.3E-22 (campaign 2018-12). This inconsistency does not change the conclusions drawn in the paper as these corrections are well below the associated statistical uncertainties. The correct nominal frequency values have been employed for the uploaded data set.</p>

restrictedSep 2020View details →
zenodo16/100

Eddy Covariance and Agronomical Meta Data Set from a Three Year Agroecosystem Site (RNG2)

<p>Dataset of soil, plant, forage, agronomical and eddy flux data from three years of commercial bioenergy feedstock production in Arkansas funded under the DOE ARPA-E SMARTFARM program in Phase I. Data set is provided to restricted users only. Data sets include three years of all 30 minute interval data. Full data sets will be published through Ameriflux. Some data, such as methane and nitrous oxide are embargoed.</p> <p>This work was partially funded through the U.S. Department of Energy, ARPA-E, under Cooperative Agreement DE-AR0001228 led by ARVA Intelligence Corp. (https://www.arvaintelligence.com) in collaboration with Lawrence Berkeley National Laboratory (LBNL).</p>

restrictedJul 2023View details →
zenodo16/100

Data set for Dermal Biomimicry: Human dermal decellularized ECM Hydrogels as Fidelity-Rich Scaffolds for In Vitro Skin Models.

<p>Data set for the article <strong>Dermal Biomimicry: Human dermal decellularized ECM Hydrogels as Fidelity-Rich Scaffolds for In Vitro Skin Models.</strong></p>

restrictedcc-by-4.0Nov 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record