Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
67
datasets available to search
ShareScore release 0.9.0
Dataset results
67 results for “data cleaning”
Dataset of "Denoising Image-based Experimental Data without Clean Targets based on Deep Autoencoders"
<p>Dataset of the paper "Denoising Image-based Experimental Data without Clean Targets based on Deep Autoencoders", published in Experimental Thermal and Fluid Science (<a title="Persistent link using digital object identifier" href="https://doi.org/10.1016/j.expthermflusci.2024.111195" target="_blank" rel="noreferrer noopener">https://doi.org/10.1016/j.expthermflusci.2024.111195</a>)</p> <p>The project received funding from: the European Research Council (ERC) under the European Union’s Horizon 2020 research and innovation program (grant agreement No 949085); the National Natural Science Foundation of China (NSFC No 12227803 and No 12372276).</p>
Data and code for "Meeting U.S. Greenhouse Gas Emissions Goals with the International Air Pollution Provision of the Clean Air Act"
<p>For the files and data associated with the Yuan et al. 2022 "Meeting U.S. Greenhouse Gas Emissions Goals with the International Air Pollution Provision of the Clean Air Act"</p> <p>Description: Data/code used in energy-economic impacts and health impacts analysis.</p> <p>Directory contents:</p> <p><strong>Energy Economic Impacts</strong></p> <ul> <li><strong>Code </strong>used for producing figures and data tables <ul> <li>'paperFigs_March2022.Rmd' contains the R code used for data analysis and visualization in the paper. (<em>The code runs with R v4.0.0, RStudio v1.4.1106, and the following packages: scales_1.1.1, ggpubr_0.4.0, cowplot_1.1.0, readxl_1.3.1, here_0.1, forcats_0.5.0, stringr_1.4.0, dplyr_1.0.4, purrr_0.3.4, readr_1.3.1, tidyr_1.1.0, tibble_3.0.6, ggplot2_3.3.4, and tidyverse_1.3.0.</em>)</li> <li>'ERL_Figure4.py' contains the Python code used for generating Figure 4 in the paper</li> </ul> </li> <li><strong>Table</strong>: data tables for figures in the paper and supplementary materials</li> <li><strong>Figure</strong>: figures in the paper and supplementary materials</li> <li><strong>Data</strong>: USREP-ReEDS results and data from other sources <ul> <li>'rrpt_subset.csv' contains the portions of the ReEDS output from February 26, 2021 that are necessary to create the figures in the paper.</li> <li>'urpt_subset.csv' contains the portions of the USREP output from February 26, 2021 that are necessary to create the figures in the paper.</li> <li>'urpt_welfare_subset.csv' contains more detailed USREP welfare output from February 26, 2021.</li> <li>'cooper_pop_proj.csv' contains U.S. population projections from the University of Virginia Weldon Cooper Center for Public Service published in 2018.</li> <li>'carbon_price_comparison.csv' contains data from other recent carbon pricing studies, as described in supplementary materials G.</li> </ul> </li> </ul> <p><strong>Health Impacts</strong></p> <ul> <li><strong>analysis</strong>: <ul> <li><strong>lib</strong>: annotated code library, which loads raw data from the root data folder and conducts health impacts analysis</li> <li><strong>data</strong>: outputs <ul> <li><strong>inmap</strong>: spatial inputs/outputs for inmap</li> <li><strong>working</strong>: intermediate procssed output files</li> <li><strong>final</strong>: final health impacts results</li> </ul> </li> </ul> </li> <li><strong>data</strong>: raw data used in analysis <ul> <li><strong>working</strong>: processed intermediate raw data for faster loading in R</li> </ul> </li> </ul>
Data from: Influence of different data cleaning solutions of point-occurrence records on downstream macroecological diversity models
<p><span>Digital point-occurrence records from the Global Biodiversity Information Facility (GBIF) and other data providers enable a wide range of research in macroecology and biogeography. However, data errors may hamper immediate use. Manual data cleaning is time-consuming and often unfeasible, given that the databases may contain thousands or millions of records. Automated data cleaning pipelines are therefore of high importance. This study examined the extent to which cleaned data from six pipelines using data cleaning tools (e.g., the GBIF web application, different R packages) affect downstream species distribution models. In addition, we assessed how the pipeline data differ from expert data. From 13,889 North American <i>Ephedra</i> observations in GBIF, the pipelines removed 31.7% to 62.7% false-positives, invalid coordinates, and duplicates, leading to data sets that included between 9,484 (GBIF application) and 5,196 records (manual-guided filtering). The expert data consisted of 703 thoroughly handpicked records, comparable to data from field studies. Although differences in the record numbers were relatively large, stacked species distribution models (sSDM) from the pipelines and the expert data were strongly related (mean Pearson's <i>r</i> across the pipelines: 0.9986, versus the expert data: 0.9173). The ever-stronger correlations resulted from occurrence information that became increasingly condensed in the course of the workflow (from individual occurrences to collectivized occurrences in grid cells to predicted probabilities in the sSDMs). In sum, our results suggest that the <i>R</i> package-based pipelines reliably identified invalid coordinates. In contrast, the GBIF-filtered data still contained both spatial and taxonomic errors. However, major drawbacks emerge from the fact that no pipeline fully discovered misidentified specimens without the assistance of expert taxonomic knowledge. We conclude that application-filtered GBIF data will still need additional review to achieve higher spatial data quality. Achieving high-quality taxonomic data will require extra effort, probably by thoroughly analyzing the data for misidentified taxa, supported by experts.</span></p>
Cacatoblastsis_cactorum_cleaned_data
<p>Cleaned database for <em>Cactoblastis cactorum</em> occurrences downloaded from the GBIF repository on 09/25/2021</p>
Cooking demand data for clean cooking simulation.
<p>This will be a dataset of clean cooking data collected hourly to allow for clean cooking simulation. Currently this is dummy data but it will soon be uploaded with live working data.</p>
p-IgGen Dataset: Cleaned paired and unpaired antibody sequence data for machine learning applications.
<p>This data is released alongside "p-IgGen: A Paired Antibody Generative Language Model", which contains full details on the data processing and cleaning.</p> <p>p-IgGen Paper: https://www.biorxiv.org/content/10.1101/2024.08.06.606780v1 .</p> <p>OAS: https://opig.stats.ox.ac.uk/webapps/oas/</p> <p> </p>
Cleaned data from Supplementary Table S1 from "Temperature-Dependent Estimation of Gibbs Energies Using an Updated Group-Contribution Method"
<p>This data is derived from <a href="https://doi.org/10.5281/zenodo.5277805">https://doi.org/10.5281/zenodo.5277805</a>. As of 2021-09-08, <a href="https://doi.org/10.5281/zenodo.5277805">https://doi.org/10.5281/zenodo.5277805</a> contained an Excel file with md5:8907e7444f49d02d58fdef9fb2a1089a and filename "TableS1.xlsx"; the file was licensed under <a href="https://creativecommons.org/licenses/by/4.0/legalcode">Creative Commons Attribution 4.0 International</a>.</p> <p>In "TableS1.xlsx" there were, among others, worksheets with the names "Table S1. TECRDB Keqs" and "Table S2. TECRDB ΔrH data".</p> <p>Worksheet "Table S1. TECRDB Keqs" was exported as a csv with filename "TableS1_Keq.csv". A column "id" was added with a persistent identifier created in w3id.org.</p> <p>Worksheet "Table S2. TECRDB ΔrH data"" was exported as a csv with filename "TableS1_deltaH.csv". A column "id" was added with a persistent identifier created in w3id.org.</p> <p>No other corrections were made.</p>
Data from: Hazard and catch composition of ghost fishing gear revealed by a citizen science clean-up initiative
<p><span>Ghost fishing, the continued catch of fishes and invertebrates by lost fishing gear, represents an animal welfare issue as well as a waste of both potential food and ecosystem resources. Fishing gear is lost by both commercial and recreational fishers, and management authorities often lack an overview of gear loss and subsequently potential impact on coastal populations. </span><span>To investigate the hazard and catch composition of lost fishing gear along the Norwegian coast</span><span>, recreational divers in collaboration with scientists conducted systematic reporting of retrieved lost fishing gear. </span><span>Through this citizen science project,</span><span> a total of 12,101 gear items were retrieved and reported, including traps, gillnets and fyke nets. Combining both data on the catch ratio of the gear and its relative quantity, we identified the five most hazardous gear types to be parlor traps, gillnets, fyke nets, wrasse traps and square collapsible traps. The parlour trap was the most hazardous trap, due to high catchability and quantity. The correct classification of gear type could not be confirmed in 2.8 – 6.1 % of the pictures taken by divers, depending on reporting format, and divers reported the wrong gear type in 1.4 % of the reports. Brown crab (<em>Cancer</em> <em>pagurus</em>) was the species most often found in retrieved gear. Furthermore, the vulnerable species European lobster (<em>Homarus</em> <em>gammarus</em>) and Atlantic cod (<em>Gadus</em> <em>morhua</em>) were also common. These results can inform future clean up-initiatives and management responses to ghost fishing, including preventive measures against gear loss and gear restrictions and customization. </span></p>
Clean Sky 2 - SALUTE H2020 Project : 2D smart liner data
<p>This data concerns all obtained measurment in paper :In flow acoustic characterisation of a 2D active liner with local and non<br> local strategies. K. Billon, E. De Bono, M. Perez , E. Salze, G. Mattenand al, Applied Acoustics 191 (2022) 108655, https://doi.org/10.1016/j.apacoust.2022.108655</p>
Data on technology transfer in Scopus (without cleaning)
<p>The project presented has the objective of developing a methodology for technology transfer (TT), based on descriptive research, with a longitudinal design by collecting data over time in an express period to make inferences regarding change, on the characteristics of scientific production on TT. With the trend obtained, the dimensions, variables and indicators, associated with the context, will be determined for the subsequent validation of experts and the use of other resources from the information sciences.</p>
Data on technology transfer in WoS (without cleaning)
<p>The project presented has the objective of developing a methodology for technology transfer (TT), based on descriptive research, with a longitudinal design by collecting data over time in an express period to make inferences regarding change, on the characteristics of scientific production on TT. With the trend obtained, the dimensions, variables and indicators, associated with the context, will be determined for the subsequent validation of experts and the use of other resources from the information sciences.</p>
Cleaned Data from Behavioural Measure of Learning agility Pilot
<p>SPSS file with cleaned and processed data from the Behavioural Measure of Learning agility Pilot </p>
Data from: Influence of different data cleaning solutions of point-occurrence records on downstream macroecological diversity models
Open the record for dataset details and reuse information.
Data from: Hazard and catch composition of ghost fishing gear revealed by a citizen science clean-up initiative
Open the record for dataset details and reuse information.
Data and code from: Competitive cleaning: Behavioural variation supports coexistence of two juvenile sympatric cleanerfishes
Open the record for dataset details and reuse information.
Data cleaning and enrichment through data integration: networking the Italian academia
Open the record for dataset details and reuse information.
Data from: Cost of an elaborate trait: a tradeoff between attracting females and maintaining a clean ornament
Open the record for dataset details and reuse information.
Data from: Endovascular treatment in older adults with acute ischemic stroke in the MR CLEAN Registry
<p><strong>Objective</strong>: To explore clinical outcomes in older adults with acute ischemic stroke treated with endovascular thrombectomy (EVT).</p> <p><strong>Methods</strong>: We included consecutive patients (2014–2016) with an anterior circulation occlusion undergoing EVT from the MR CLEAN Registry. We assessed the effect of age (dichotomized at ≥80 years, and as continuous variable) on the modified Rankin Scale [mRS] score at 90 days, symptomatic intracranial hemorrhage (sICH), and reperfusion rate. The association between age and mRS was assessed with multivariable ordinal logistic regression, and a multiplicative interaction term was added to the model to assess modification of reperfusion by age on outcome.</p> <p><strong>Results</strong>: 380/1526 (25%) of patients were 80 years or older (=older adults). Older adults had a worse functional outcome than younger patients (adjusted common OR for an mRS shift towards better outcome: 0.31, 95%CI 0.24–0.39). Mortality was also higher in older adults (51% vs. 22%, aOR 3.12, 95%CI 2.33–4.19). There were no differences in proportion of patients with mRS 4-5, sICH, or reperfusion rates. Successful reperfusion was more strongly associated with a shift towards good functional outcome in older adults than in younger patients (acOR 3.22, 95%CI 2.04–5.10 vs. 2.00, 95%CI 1.56–2.57, P<sub>interaction</sub>=0.026).</p> <p><strong>Conclusion</strong>: Older age is associated with an increased absolute risk of poor clinical outcome, while the relative benefit of successful reperfusion seems to be higher in these patients. These results should be taken into consideration when selecting older adults for EVT.</p>
MME-only models trained with clean data for JAMES paper "Machine-learned uncertainty quantification is not magic"
<p>This tar file contains all 100 trained models in the MME-only ensemble from Experiment 1 (i.e., those trained with clean data, not with lightly perturbed data). To read one of the models into Python, you can use the method neural_net.read_model in the ml4rt library.</p>
MME/CRPS models trained with clean data for JAMES paper "Machine-learned uncertainty quantification is not magic"
<p>This tar file contains all 100 trained models in the MME/CRPS ensemble from Experiment 1 (i.e., those trained with clean data, not with lightly perturbed data). To pare the ensemble down to 50 models, we randomly select 50. To read one of the models into Python, you can use the method neural_net.read_model in the ml4rt library.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.