Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
67
datasets available to search
ShareScore release 0.9.0
Dataset results
67 results for “data cleaning”
Data cleaning and analysis for the Master's thesis: DIFFERENCES IN CONSUMER PREFERENCES FOR UNWEATHERED AND WEATHERED WOOD
<p>The data and analytical support the Master's thesis submitted by Hana Remesova at the University of Primorska<br> Faculty of Mathematics, Natural Sciences, and Information Technologies. The .csv files are data files, the .Rmd file is an R markdown which can be run. The product of knitting the .Rmd file is the .html.</p>
Cleaned spawner survey data
<p>This dataset is a version of the New Salmon Escapement Database (NuSEDS) downloaded from Open Data Canada on January 21, 2025 at <a href="https://open.canada.ca/data/en/dataset/c48669a3-045b-400d-b730-48aafe8c5ee6">https://open.canada.ca/data/en/dataset/c48669a3-045b-400d-b730-48aafe8c5ee6</a>.</p> <p>These data have been cleaned and matched to Conservation Units (CUs) by staff at the Pacific Salmon Foundation. </p> <p>Details of the data cleaning procedure are outlined in the Pacific Salmon Explorer Technical Report (Appendix 2) available online at <a href="https://www.salmonexplorer.ca/methods/appendix-2.html">https://www.salmonexplorer.ca/methods/appendix-2.html</a>.</p> <p>The entire cleaning procedure is available online at <a href="https://bookdown.org/salmonwatersheds/1_nuseds_collation_atkinson/1_nuseds_collation.html">a_nuseds_collation</a> and corresponding code <a href="https://github.com/salmonwatersheds/salmon_data_demise/blob/main/code/a_nuseds_collation.Rmd">a_nuseds_collation.Rmd</a>.</p> <p>The procedure to match the cleaned data to CUs is available online at <a href="https://bookdown.org/salmonwatersheds/2_nuseds_cuid_pse_atkinson/2_nuseds_cuid_pse.html">b_nuseds_cuid_pse</a> and corresponding code <a href="https://github.com/salmonwatersheds/salmon_data_demise/blob/main/code/b_nuseds_cuid_pse.Rmd">b_nuseds_cuid_pse.Rmd</a>. </p> <p>This revised dataset (v3, 2025-04-15) includes observed counts of zero.</p> <p>This is the dataset used in <a href="https://doi.org/10.1139/cjfas-2024-0387">Atkinson et al. 2025. Monitoring for fisheries or for fish? Declines in monitoring of salmon spawners continue despite a conservation crisis</a>, published in the CJFAS.</p> <p>The definition of the fields/columns is available at: <a href="https://github.com/salmonwatersheds/population-indicators/blob/f33058a4aaa0e5fc30354588b5dab469cc9d2a4a/spawner-surveys/output/METADATA_2_nuseds_cuid_streamid.csv">METADATA_2_nuseds_cuid_streamid.csv</a></p>
Data and Code for: Real-Time Pricing and the Cost of Clean Power
<p>Solar and wind power are now cheaper than fossil fuels but are intermittent. The extra supply-side variability implies growing benefits of using real-time retail pricing (RTP). We evaluate the potential gains of RTP using a model that jointly solves investment, supply, storage, and demand to obtain a chronologically detailed dynamic equilibrium for the island of Oahu, Hawai'i. We find that RTP reduces costs in high-renewable systems by roughly 6 to 12 times as much as in fossil systems holding demand assumptions fixed, markedly lowering the cost of clean energy integration.</p>
SCALE-WIN19 Trace Metal Clean Rosette Data
<p>The files here contain the trace metal clean CTD bottle and sensor files for the SCALE Winter Cruise.</p> <p>Oxygen sensor data is now included - except for stations PUZ, SAZ2, GT1, GTE, GT1, MIZ1 and MIZ2.</p> <p>For notes on how the data was processed please refer to the SCALE CTD Processing Report.</p> <p> </p> <table><colgroup><col><col></colgroup> <tbody> <tr> <td><strong>Variable</strong></td> <td><strong>Units</strong></td> </tr> <tr> <td>Temperature</td> <td>degrees C</td> </tr> <tr> <td>Conductivity</td> <td>S/m</td> </tr> <tr> <td>Salinity</td> <td>PSU</td> </tr> <tr> <td>Oxygen in situ/Oxygen theoretical/AOU</td> <td>mL/L</td> </tr> <tr> <td>Oxygen Saturation</td> <td>%</td> </tr> <tr> <td>Density</td> <td>kg/m3</td> </tr> <tr> <td>Chlorophyll</td> <td>mg/m3</td> </tr> <tr> <td>Beam Transmission</td> <td>%</td> </tr> <tr> <td>Beam Attenuation</td> <td>m</td> </tr> </tbody> </table>
SCALE-SPR19 Trace Metal Clean CTD Rosette Data
<p>The files here contain the trace metal clean CTD bottle and sensor files for the SCALE Spring Cruise. </p> <p>Oxygen sensor data is now included.</p> <p>For notes on how the data was processed please refer to the SCALE CTD Processing Report.</p> <table><colgroup><col><col></colgroup> <tbody> <tr> <td><strong>Variable</strong></td> <td><strong>Units</strong></td> </tr> <tr> <td>Temperature</td> <td>degrees C</td> </tr> <tr> <td>Conductivity</td> <td>S/m</td> </tr> <tr> <td>Salinity</td> <td>PSU</td> </tr> <tr> <td>Oxygen in situ/Oxygen theoretical/AOU</td> <td>mL/L</td> </tr> <tr> <td>Oxygen Saturation</td> <td>%</td> </tr> <tr> <td>Density</td> <td>kg/m3</td> </tr> <tr> <td>Chlorophyll</td> <td>mg/m3</td> </tr> <tr> <td>Beam Transmission</td> <td>%</td> </tr> <tr> <td>Beam Attenuation</td> <td>m</td> </tr> </tbody> </table>
Input data for the OnStove Nepal model "AAchieving Nepal's clean cooking ambitions: an open source and geospatial cost–benefit analysis"
<p>This repository includes input data to run the OnStove Nepal model presented in the paper "<strong>Achieving Nepal's clean cooking ambitions: an open source and geospatial cost–benefit analysis</strong>" DOI: <a href="https://doi.org/10.1016/S2542-5196(24)00209-2">https://doi.org/10.1016/S2542-5196(24)00209-2</a>.</p> <p>The code and automated workflow to run the model can be found in the Github repository <a href="https://github.com/Open-Source-Spatial-Clean-Cooking-Tool/OnStove-Nepal">https://github.com/Open-Source-Spatial-Clean-Cooking-Tool/OnStove-Nepal</a>. All result files and figures can be downloaded from the permanent repository <a href="https://doi.org/10.5281/zenodo.10643983">https://doi.org/10.5281/zenodo.10643983</a>.</p> <p>The "<strong>GIS_input_data/</strong>" directory includes all the geospatial datasets needed to run the model. Each dataset folder contains a Source.md file describing the dataset, source, attribution, and license. To run the model extract the data inside your "<strong>1. Data</strong>"<strong> </strong>folder in your project. </p> <p>The "<strong>Scenario_inputs/</strong>" directory includes the CSV files with the input socio- and techno-economic data for the different scenarios. Sources for the socio- and techno-economic data can be found in the <strong>supplementary material</strong> of the related publication in the link <a href="https://doi.org/10.1016/S2542-5196(24)00209-2">https://doi.org/10.1016/S2542-5196(24)00209-2</a>. To run the model extract the scenario data inside your "<strong>2. Scenario inputs</strong>"<strong> </strong>folder in your project. </p>
The processed clean data of 16S rRNA V4 amplicon sequecnces for the six stage of phenolic microbiome domestication
Open the record for dataset details and reuse information.
EIA_Cleaned_Hourly_Electricity_Demand_Data: 2015 - 2024
<div> <p>Cleaned hourly electricity demand data for electric balancing authorities within the contiguous US. Raw data is based on the U.S. Energy Information Administration's collected data <a href="http://www.eia.gov/opendata/qb.php?category=2122628">here</a>.</p> <p>Cleaned hourly data spans July 2, 2015 - Dec 31, 2024 (9.5 continuous year).</p> <p> </p> <p>Please consider citing:</p> <p>Ruggles, T.H., Farnham, D.J., Tong, D. <em>et al.</em> Developing reliable hourly electricity demand data through screening and imputation. <em>Sci Data</em> <strong>7, </strong>155 (2020). <a href="https://doi.org/10.1038/s41597-020-0483-x">https://doi.org/10.1038/s41597-020-0483-x</a></p> </div> <p> </p> <p>Since the prior releases of this data, the EIA has limited the range of historical data available via their API. This release focuses on cleaning calendar years 2020 through 2024 and including the many balancing area subregions.</p> <p>Users can combine prior data releases with this latest release to have data extending from from 2 July 2015 through 31 Dec 2024.</p>
Data Cleaning, Translation & Split of the Dataset for the Automatic Classification of Documents for the Classification System for the Berliner Handreichungen zur Bibliotheks- und Informationswissenschaft
<ul> <li>Cleaned_Dataset.csv – The combined CSV files of all scraped documents from DABI, e-LiS, o-bib and Springer.</li> <li>Data_Cleaning.ipynb – The Jupyter Notebook with python code for the analysis and cleaning of the original dataset.</li> <li>ger_train.csv – The German training set as CSV file.</li> <li>ger_validation.csv – The German validation set as CSV file.</li> <li>en_test.csv – The English test set as CSV file.</li> <li>en_train.csv – The English training set as CSV file.</li> <li>en_validation.csv – The English validation set as CSV file.</li> <li>splitting.py – The python code for splitting a dataset into train, test and validation set.</li> <li>DataSetTrans_de.csv – The final German dataset as a CSV file.</li> <li>DataSetTrans_en.csv – The final English dataset as a CSV file.</li> <li>translation.py – The python code for translating the cleaned dataset.</li> </ul>
MULTIPLIERS_WP3/4/5_Science learning project on Clean water and sanitation_Public data_IREN_20241021_v1
<p><span>This dataset contains the following data related to</span><span> the science learning project on <em>Clean water and sanitation</em></span><span>:</span></p> <ul> <li><span>S</span><span>ummary of transcripts, quotes and translations, from interviews with teachers, OSC members, and students (</span><span>Pseudo-/Anonymised)</span></li> <li>Summary<span> of transcripts, quotes and translations, from focus group discussions with students (</span><span>Pseudo-/Anonymised)</span></li> </ul>
Yield Prediction Through Integration of Genetic, Environment, and Management Data Through Deep Learning: Cleaned Data
<p>The included files and script are to allow for reconstruction of the data directory and cleaned data used in "Yield Prediction Through Integration of Genetic, Environment, and Management Data Through Deep Learning" ( https://doi.org/10.1101/2022.07.29.502051 ). Code used is available at 10.5281/zenodo.7401113 .</p> <table> <tbody> <tr> <th>Filename</th> <th>Description</th> </tr> <tr> <td>interim.tar.gz</td> <td>Contains site grouping dictonary</td> </tr> <tr> <td>processed.tar.gz</td> <td>Processed data</td> </tr> <tr> <td>raw.tar.gz</td> <td>Input data</td> </tr> <tr> <td>SetupInstructions.sh</td> <td>Bash script to prepare folders and unzipped data expected by code in 10.5281/zenodo.7401113</td> </tr> <tr> <td>SetupInstructions.txt</td> <td>Instructions for unzipping the data</td> </tr> <tr> <td>Train_Test_Split_Reference_Phenotypes.csv</td> <td>Reference spreadsheet to allow for easily exploring training and test set groupings</td> </tr> </tbody> </table> <ul> </ul> <p>This work was supported through funding from the USDA Agricultural Research Service, ARS project number 5070-21000-041-000-D. Raw data provided by the [Genomes to Field Initiative](https://www.genomes2fields.org/) and the [Daymet database](https://daymet.ornl.gov/).</p>
Training data from: Machine learning predicts which rivers, streams, and wetlands the Clean Water Act regulates
Open the record for dataset details and reuse information.
Prediction data from: Machine learning predicts which rivers, streams, and wetlands the Clean Water Act regulates
Open the record for dataset details and reuse information.
Data From: Emissions redistribution and environmental justice implications of California's Clean Vehicle Rebate Project
Open the record for dataset details and reuse information.
temporalNEON: Repository containing raw and cleaned-up organismal data from the National Ecological Observatory Network (NEON) useful for evaluating the links between change in biodiversity and ecosystem stability
Organismal data include the following taxonomic groups: small mammals, fish, ground beetles, and aquatic macroinvertebrates. Data were retrieved from the National Ecological Observatory Network (NEON) database in November 2020. We submit both raw data retrieved from NEON as .rds files, R code used to process these data, as well as processed data as .csv files.
Data from: Cost of an elaborate trait: a tradeoff between attracting females and maintaining a clean ornament
<p><span><span><span><span><span><span><span><span><span><span><span>Many sexually selected ornaments and weapons are elaborations of an animal's outer body surface, including long feathers, colorful skin, and rigid outgrowths. The time and energy required to keep these traits clean, attractive, and in good condition for signaling may represent an important, but understudied cost of bearing a sexually selected trait. Male fiddler crabs possess an enlarged and brightly colored claw that is used both as a weapon to fight with rival males and also as an ornament to court females. Here, we demonstrate that males benefit from grooming because females prefer males with clean claws over dirty claws, but also that the time spent grooming detracts from the amount of time available for courting females. Males therefore face a temporal tradeoff between attracting the attention of females and maintaining a clean claw. Our study provides rare evidence of the importance of grooming for mediating sexual interactions in an invertebrate, indicating that sexual selection has likely shaped the evolution of self-maintenance behaviors across a broad range of taxa.</span></span></span></span></span></span></span></span></span></span></span></p>
Tripsacum isoseq clean data
<p>The final set of unique tripsacum isoforms using Pacbio full-length sequencing which mapped to a total of 13,089 maize genes</p>
Practice cleaning data
Open the record for dataset details and reuse information.
Air phyto-cleaning by an urban meadow – Filling the winter gap - DATA
<p>Database of article: Air phyto-cleaning by an urban meadow – Filling the winter gap.</p>
Data for "From Kitchen Garden to Multifunctionality: Leek-inspired Surface Structures Introduce Optical and Self-cleaning Properties to Cellulose-based Films"
<p>UV-Vis:<br>The optical properties were measured from 300 nm to 800 nm. The transmittance and haze were calculated using the following equations, and the results were reported for three sets of measurements:<br>Transmittance (%) = T2/T1 × 100 <br>Haze (%) = (T4/T2 - T3/T1) × 100 <br>where T1 is the reference transmitted light without the sample, T2 is the total light transmitted with the presence of the sample, T3 is light beam scattering by the UV-Vis device, and T4 is the diffusive transmittance, referring to the light transmitted by both the sample and the device. <br>The zip files are named by noting 'UV-Vis' followed by the type of sample.</p> <p>Current density-Voltage:<br>Seven sets of measurements were conducted on perovskite solar cell (PSC) devices, comparing the performance of uncoated (noted as pristine) devices to those coated with the replica. Additionally, another seven sets of measurements compared the performance of devices with and without the replica+2%CW coating. The data are recorded in .txt files noted by 'Current density-Voltage spectra.'</p> <p>Light scattering with halogen light beam:<br>The intensity of the illuminated light is determined for the length of a horizontal line passing through the center by using image analysis. The intensities are provided for all the cases, i,e,. with no film, with plain CA film, and with the replica. The .txt file is named 'Light scattering with halogen light beam.'</p> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.