Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,481

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,481 results for “data processing”

Learn how ShareScore rates datasets ↗
zenodo36/100

Multi-omic integration of DNA methylation and gene expression data reveals molecular vulnerabilities in glioblastoma (processed data)

<p>Glioblastoma multiforme (GBM) is one of the most aggressive types of cancer and exhibits profound genetic and epigenetic heterogeneity, making the development of an effective treatment a major challenge. The recent incorporation of molecular features into the diagnosis of GBM patients has led to an improved categorisation into various tumour subtypes with different prognoses and disease management. In this work, we have exploited the benefits of genome-wide multi-omic approaches to identify potential molecular vulnerabilities existing in GBM patients. Integration of gene expression and DNA methylation data from both bulk GBM and patient-derived GBM stem cell lines has revealed the presence of major sources of GBM variability, pinpointing subtype-specific tumour vulnerabilities amenable to pharmacological interventions. In this sense, inhibition of the AP1, SMAD3 and RUNX1 / RUNX2 pathways, in combination or not with the chemotherapeutic agent temozolomide, led to the subtype-specific impairment of tumour growth, particularly in the context of the aggressive, mesenchymal-like subtype. These results emphasize the involvement of these molecular pathways in the development of GBM and have potential implications for the development of personalized therapeutic approaches.</p>

opencc-by-4.0Dec 2023View details →
zenodo36/100

Primary collected data for modelling the additive MAR/R process by means of the PBF-LB process based on the example of tool steel 1.2709 powder

<p>For the evaluation of a MAR/R process, not only process-, material- and demonstrator-specific correlations and data must be combined. In addition to secondary data (e.g. databases, publications, etc.), primary data (e.g. process times, volume flows, etc.) must also be collected for the specific application.</p> <p>The attached table shows the primary data to be collected for the cradle-to-gate process depending on the process phases and steps.&nbsp;This data is used as support for ecological as well as economic process and component evaluations.</p>

opencc-by-4.0May 2023View details →
zenodo36/100

Processed GPR data of Mosvatnet, Norway

<p>SEG-Y files corresponding to the processed GPR data collected on- and offshore Mosvatnet, Stavanger, Norway. In the GPR radargram is possible to identify and map Quaternary glacial and post-glacial sediments and non-natural&nbsp;infill materials.&nbsp;</p>

opencc-by-4.0May 2023View details →
zenodo36/100

MMCFSv2 and MMCFSv1 Processed data for GMD Manuscript

<p>This dataset was used to make all the plots in the manuscript titled</p> <p>&quot;Monsoon Mission Coupled Forecast System Version 2.0: Model Description and Indian Monsoon Simulations&quot;</p> <p>submitted to GMD</p>

opencc-by-4.0May 2023View details →
zenodo36/100

TFG Systematization process of generating semantic data and ontology

<pre>Set of TALIS files, the main source of data for the investigation, in CSV format.General ontology manually and by new software. Semantic data, as well as the DSL code. Finally, the web design of the new tool.</pre>

opencc-by-4.0May 2023View details →
zenodo36/100

Data from: Integrated survey methodologies provide process-driven framework for marine renewable energy environmental impacts

<p>Data from: Integrated survey methodologies provide process-driven framework for marine renewable energy environmental impacts</p>

opencc-by-4.0May 2023View details →
dryad36/100

Example data and scripts for: Processing IMU signals to recreate sacral trajectory during treadmill walking

<p>Example IMU data and scripts for reconstructing the trajectory of a sacral IMU and validation with motion capture data. Also contains example code of gait event detection and synchronization.</p>

opencc-zeroJun 2023View details →
zenodo36/100

LTEE-INSeq-processed-data

<p>This repository contains processed data from transposon sequencing (INSeq) of the ancestral strain (REL606) and clones sampled at 2,000 and 15,000 generations for the Ara+2 and Ara-1 lineages from the Lenski&#39;s Long-Term Evolution Experiment (LTEE). Files contain the number of reads per data point for genome-wide insertion mutants that were subjected to bulk competition under similar conditions to those of the LTEE. Raw sequencing reads are available from the NCBI BioProject database (PRJNA979973). Scripts for analyses are available from GitHub (https://github.com/ACouce/LTEE2022).</p>

opencc-by-4.0Dec 2022View details →
dryad36/100

Data from: Complex evolutionary processes maintain an ancient chromosomal inversion

<p>Genome re-arrangements such as chromosomal inversions are often involved in adaptation. As such, they experience natural selection, which can erode genetic variation. Thus, whether and how inversions can remain polymorphic for extended periods of time remains debated. Here we combine genomics, experiments, and evolutionary modeling to elucidate the processes maintaining an inversion polymorphism associated with the use of a challenging host plant (Redwood trees) in <em>Timema </em>stick insects. We show that the inversion is maintained by a combination of processes, finding roles for life-history trade-offs, heterozygote advantage, local adaptation to different hosts, and gene flow. We use models to show how such multi-layered regimes of balancing selection and gene flow provide resilience to help buffer populations against the loss of genetic variation, maintaining the potential for future evolution. We further show that the inversion polymorphism has persisted for millions of years and is not a result of recent introgression. We thus find that rather than being a nuisance, the complex interplay of evolutionary processes provides a mechanism for the long-term maintenance of genetic variation.</p>

opencc-zeroJun 2023View details →
dryad36/100

Data from: Comparative LCA studies of simulated HMF-biorefineries from maize and miscanthus as an example of first- and second-generation biomass as a tool for process development

<p class="MsoNormal"><span>5-Hydroxymethylfurfural (HMF) is a versatile platform chemical for a fossil free, bio-based chemical industry. HMF can be produced by using fructose as a feedstock. Using edible, first-generation biomass to produce chemicals has been questioned in terms of potential competition with food supply. Second-generation biomass like miscanthus could be an alternative. However, there is a lack of information if second-generation lignocellulosic biomass is a more sustainable feedstock to produce HMF. Therefore, a life cycle assessment was performed in this study to determine the environmental impacts of HMF production from miscanthus and to compare it with HMF from high fructose corn syrup (HFCS). HFCS from either Hungary or Baden-Württemberg (Germany) was considered. Compared to the HFCS biorefineries the miscanthus concept is producing less emissions in all impact categories studied, except land occupation. Overall, the production and usage of second-generation biomass could be especially beneficial in areas where the use of N-fertilizers is restricted. Besides, conclusions for the further development of the on-farm-biorefinery concept were elaborated. For this purpose, process simulations from a previous study were used. Results of the previous study in terms of TEA and the current LCA study in terms of environmental sustainability indicate that the lignin depolymerization unit in the miscanthus biorefinery has to be improved. The scenario without lignin depolymerization performs better in all impact categories. The authors recommend to not further convert the lignin to products like phenol and other aromatic compounds. The results of the contribution analyses show that the major impact in the HMF production is caused by the auxiliary materials in the separation units and the required heat. Further technical development should focus on efficient heat as well as solvent use and solvent recovery. At this point further optimizations will lead to reduced emissions and costs at the same time. The presented data set is the used inventory to model the environmental impacts. </span></p>

opencc-zeroJun 2023View details →
zenodo36/100

BIG DATA ANALYTICS IN DIGITAL HUMAN RESOURCES MANAGEMENT: IMPACT ON THE RECRUITMENT PROCESS

<p>This study investigates how HR employees experience the big data phenomenon in the recruitment function of HRM and how their perceptions of the phenomenon have evolved. This study also examines how BD will affect organizational and HRM and how it can be improved in other functions of HR. In this exploratory study, which comprehensively addresses the BD phenomenon in HRM, the phenomenological design approach, one of the qualitative research methods, was applied to test the research questions and a semi-structured interview form was used for research data. Using the snowball sampling method, in-depth interviews were conducted with 10 HR employees working in large and semi-structured organizations in Turkey and the interviews were analyzed with MAXQDA 20. The findings show that HRM employees are aware of BD. On the other hand, it is understood that BD technologies provide easy accessibility in recruitment, offer a strategic competitive advantage, and enable more effective management of information management, which saves HR employees&#39; work in a facilitating way. They benefit from technology as a decision support assistant. Finally, the research results provide theoretical and practical implications for future researchers and practitioners for the development and effective use of BD technology in the field of HRM.</p>

opencc-by-4.0Jun 2023View details →
zenodo36/100

processed single-cell data from "Cancer cell non-autonomous tumor progression from chromosomal instability"

<p>The h5ad files can be used as processed scRNA-seq input to the ContactTracing code (https://zenodo.org/badge/latestdoi/625036312).</p> <p>There is one file for the highCIN/lowCIN comparison, and another for the highCIN/noSTING comparison.</p>

opencc-by-4.0Jun 2023View details →
dryad36/100

Data from: Diversification processes in Gerp's mouse lemur demonstrate the importance of rivers and altitude as biogeographic barriers in Madagascar's humid rainforests

<p><span>Madagascar exhibits exceptionally high levels of biodiversity and endemism. Models to explain the diversification and distribution of species in Madagascar stress the importance of historical variability in climate conditions which may have led to the formation of geographic barriers by changing water and habitat availability. The relative importance of these models for the diversification of the various forest-adapted taxa of Madagascar has yet to be understood. Here, we reconstructed the phylogeographic history of Gerp's mouse lemur (<em>Microcebus gerpi</em>) to identify relevant mechanisms and drivers of diversification in Madagascar's humid rainforests. We used restriction site associated DNA (RAD) markers and applied population genomic and coalescent-based techniques to estimate genetic diversity, population structure, gene flow and divergence times among <em>M. gerpi</em> populations and its two sister species <em>M. jollyae</em> and <em>M. marohita</em>. Genomic results were complemented with ecological niche models to better understand the relative barrier function of rivers and altitude. We show that <em>M. gerpi</em> diversified during the late Pleistocene. The inferred ecological niche, patterns of gene flow and genetic differentiation in <em>M. gerpi</em> suggest that the potential for rivers to act as biogeographic barriers depended on both size and elevation of headwaters. Populations on opposite sides of the largest river in the area with headwaters that extend far into the highlands show particularly high genetic differentiation, whereas rivers with lower elevation headwaters have weaker barrier functions, indicated by higher migration rates and admixture. We conclude that<em> M. gerpi</em> likely diversified through repeated cycles of dispersal punctuated by isolation to refugia as a result of paleoclimatic fluctuations during the Pleistocene. We argue that this diversification scenario serves as a model of diversification for other rainforest taxa that are similarly limited by geographic factors. In addition, we highlight conservation implications for this critically endangered species, which faces extreme habitat loss and fragmentation. </span></p>

opencc-zeroJul 2023View details →
zenodo36/100

Data files for 'Tan et al., (2023). Structural heterogeneity-controlled rupture process of the 2021 Mw 7.1 Fukushima, Japan earthquake revealed by joint inversion of seismic and geodetic data'

<p>slip model.dat: rupture model of the&nbsp;2021 Mw 7.1 Fukushima earthquake</p> <p>In &#39;slip model.dat&#39;, each row contains the moment rate function of each sub-fault.&nbsp;The numbers of the sub-faults are given in the first two columns.</p>

opencc-by-4.0Jul 2023View details →
zenodo36/100

Data to Deliverable 3.4 - Co-creation process for innovative planning and governance able to facilitate the transition to a water circular economy

<p>Transcription of co-creation workshops; Worksheets filled by groups of the co-creation workshops; and final worksheets filled of the co-creation workshops.<br> &nbsp;</p>

opencc-by-4.0Jul 2023View details →
dryad36/100

Data from: Remotely sensed environmental measurements detect decoupled processes driving population dynamics at contrasting scales

<p class="MsoNormal">The increasing availability of satellite imagery has supported a rapid expansion in forward-looking studies seeking to track and predict how climate change will influence wild population dynamics. However, these data can also be used in retrospect to provide additional context for historical data in the absence of contemporaneous environmental measurements. We used 167 Landsat-5 Thematic Mapper (TM) images spanning 13 years to identify environmental drivers of fitness and population size in a well-characterized population of banner-tailed kangaroo rats (<em>Dipodomys spectabilis</em>) in the southwestern United States. We found evidence of two decoupled processes that may be driving population dynamics in opposing directions over distinct time frames. Specifically, increasing mean surface temperature corresponded to increased individual fitness, where fitness is defined as the number of offspring produced by a single individual. This result contrasts with our findings for population size, where increasing surface temperature led to decreased numbers of active mounds. These relationships between surface temperature and (i) individual fitness and (ii) population size would not have been identified in the absence of remotely sensed data, indicating that such information can be used to test existing hypotheses and generate new ecological predictions regarding fitness at multiple spatial scales and degrees of sampling effort. To our knowledge, this study is the first to directly link remotely sensed environmental data to individual fitness in a nearly exhaustively sampled population, opening a new avenue for incorporating remote sensing data into eco-evolutionary studies.</p>

opencc-zeroJul 2023View details →
zenodo36/100

A multi-model ensemble of baseline and process-based models improves the predictive skill of near-term lake forecasts: data, forecasts, and scores

<p>This data publication contains zipped parquet from the Falling Creek Reservoir multi-model ensemble (MME) forecasting work using the FLARE (Forecasting Lake And Reservoir Ecosystems) system and baseline models:&nbsp;drivers.zip contains NOAA driver forecast files, targets.zip contains in-situ water temperature observations, forecasts.zip contains forecast parquet files generated from the MME&nbsp;workflow (FLARE&nbsp;&amp; baseline models), and scores.zip contains forecast skill metrics required for analysis.</p>

opencc-by-4.0Jul 2023View details →
zenodo36/100

Raw data for " Unnatural evolutionary processes of SARS-CoV-2 variants and possibility of deliberate natural selection" DOI 10.5281/zenodo.8248320

<p>Compressed raw data for&nbsp; &quot; Unnatural evolutionary processes of SARS-CoV-2 variants and possibility of deliberate natural selection&quot;&nbsp;&nbsp;&nbsp;</p> <p>DOI&nbsp;&nbsp; 10.5281/zenodo.8248320</p>

opencc-by-4.0Aug 2023View details →
zenodo36/100

Processed RNA expression count data and metadata from Gupta et al.:Systems genomics of salinity stress response in rice

<p>We assessed gene expression variation in a population of 130 accessions of rice (Oryza sativa) belonging to the major varietal group indica. The field experiment was conducted in the dry season of 2017 at IRRI in Los Banos, Laguna, Philippines. Seeds from each accession were sown on December 16, 2016, and seedlings were then transplanted into the experimental fields at 17 days after sowing (DAS), on January 5, 2017. The field experiments was conducted across two locations close-by: one non-salinized normal field and the other salinized field. Both field environments used a randomized complete block experimental design with each accession planted in three replicates. Each experimental plot included the accessions NSICRC 222 and NSICRC 182 serving as border rows. The application of salt in Block L5 started on January 19, 2017 when the plants were 31 days old. The salinity level was monitored by recording electrical conductivity (EC), using EC meters installed in each of the parcels at a depth of 30 cm. The salinity levels were then maintained at 6 dS/m (considered mild to moderate salinity stress) until maturity. Management and maintenance of the fields included the application of basal fertilizer, spraying of insecticides against thrips and removal of plants potentially infected with the rice tungro virus disease. Tissue collection for transcriptome. Briefly, leaf collection was done at 38 DAS (8 days after the beginning of the salt treatment) in the non-saline and saline field. Transcript levels were measured using a liquid automation-based 3 prime mRNA-seq quantification approach. Samples were multiplexed in batches of 96 per library. Raw sequencing data are available at the SRA in BioProject accession number PRJNA1010833. A key to the raw sequencing data in this BioProject can be found in the metadata of the processed RNA expression count data here</p>

opencc-by-4.0Aug 2023View details →
zenodo36/100

BriTROC-1 study: genomic landscape of recurrent ovarian high grade serous carcinoma - pre processed data

<p>Dataset containing&nbsp;pre-processed data files required to replicate analysis performed in the publication &quot;<strong>The genomic landscape of recurrent ovarian high grade serous carcinoma: the BriTROC-1 study</strong>&quot; (<a href="https://doi.org/10.1038/s41467-023-39867-7">Smith &amp; Bradley&nbsp;et al. 2023</a>).</p>

opencc-by-4.0Jul 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record