Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
67
datasets available to search
ShareScore release 0.9.0
Dataset results
67 results for “data cleaning”
Exhaust gas cleaning systems (scrubbers): characterisation of waste streams and supporting operational data
<p>Compilation of ship scrubber water constituents (e.g. metals, polycyclic aromatic hydrocarbons) and properties (e.g. pH, temperature). Including ship operational data (e.g. discharge flow rate, engine load, fuel sulphur content) where available. </p>
Marambio DMPS clean data
<p>DMPS data inverted with Matlab</p>
Data repository Moringa filber filters npJ clean water
<p>The file contains the experimental data and analysis presented in the manuscript on Moringa functionalized natural fiber filters.</p>
Data cleaning using unstructured data
<p>In this project, we work on repairing three datasets:</p> <ul> <li>Trials design: This dataset was obtained from the European Union Drug Regulating Authorities Clinical Trials Database (<a href="https://eudract.ema.europa.eu/" target="_blank" rel="noopener">EudraCT</a>) register and the ground truth was created from external registries. In the dataset, multiple countries, identified by the attribute <code>country_protocol_code</code>, conduct the same clinical trials which is identified by <code>eudract_number</code>. Each clinical trial has a <code>title</code> that can help find informative details about the design of the trial.</li> <li>Trials population: This dataset delineates the demographic origins of participants in <a href="https://eudract.ema.europa.eu/" target="_blank" rel="noopener">clinical trials</a> primarily conducted across European countries. This dataset include structured attributes indicating whether the trial pertains to a specific gender, age group or healthy volunteers. Each of these categories is labeled as (`1') or (`0') respectively denoting whether it is included in the trials or not. It is important to note that the population category should remain consistent across all countries conducting the same clinical trial identified by an <code>eudract_number</code>. The ground truth samples in the dataset were established by aligning information about the trial populations provided by external registries, specifically the <a href="https://clinicaltrials.gov/" target="_blank" rel="noopener">CT.gov</a> database and the <a href="https://drks.de/search/en" target="_blank" rel="noopener">German Trials</a> database. Additionally, the dataset comprises other unstructured attributes that categorize the inclusion criteria for trial participants such as <code>inclusion</code>.</li> <li>Allergens: This dataset contains information about products and their allergens. The data was collected from the German version of the `<a href="https://www.alnatura.de/de-de/maerkte/" target="_blank" rel="noopener">Alnatura</a>' (Access date: 24 November, 2020), a free database of food products from around the world `<a href="https://world.openfoodfacts.org/data" target="_blank" rel="noopener">Open Food Facts</a>', and the websites: `<a href="https://migipedia.migros.ch/en">Migipedia</a>', '<a href="https://www.piccantino.com/">Piccantino</a>', and `<a href="http://das-ist-drin.de/">Das Ist Drin</a>'. There may be overlapping products across these websites. Each product in the dataset is identified by a unique <code><em>code</em></code>. Samples with the same <code><em>code</em></code> represent the same product but are extracted from a differentb <code>source</code>. The allergens are indicated by (‘2’) if present, or (‘1’) if there are traces of it, and (‘0’) if it is absent in a product. The dataset also includes information on <code>ingredients</code> in the products. Overall, the dataset comprises categorical structured data describing the presence, trace, or absence of specific allergens, and unstructured text describing ingredients. </li> </ul> <p>N.B: Each '.zip' file contains a set of 5 '.csv' files which are part of the afro-mentioned datasets:</p> <ul> <li>"{dataset_name}_train.csv": samples used for the ML-model training. (e.g "allergens_train.csv")</li> <li>"{dataset_name}_test.csv": samples used to test the the ML-model performance. (e.g "allergens_test.csv")</li> <li>"{dataset_name}_golden_standard.csv": samples represent the ground truth of the test samples. (e.g "allergens_golden_standard.csv")</li> <li>"{dataset_name}_parker_train.csv": samples repaired using <a href="https://gitlab.com/ledc/ledc-sigma/-/blob/master/docs/repair.md#parker-repair" target="_blank" rel="noopener">Parker Engine</a> used for the ML-model training. (e.g "allergens_parker_train.csv")</li> <li>"{dataset_name}_parker_train.csv": samples repaired using Parker Engine used to test the the ML-model performance. (e.g "allergens_parker_test.csv")</li> </ul>
Data from: History cleans up messes: the impact of time in driving divergence and introgression in a tropical suture zone
Contact zones provide an excellent arena in which to address questions about how genomic divergence evolves during lineage divergence. They allow us to both infer patterns of genomic divergence in allopatric populations isolated from introgression and to characterize patterns of introgression after lineages meet. Thusly motivated, we analyze genome-wide introgression data from four contact zones in three genera of lizards endemic to the Australian Wet Tropics. These contact zones all formed between morphologically cryptic lineage-pairs within morphologically defined species, and the lineage-pairs meeting in the contact zones diverged anywhere from 3.1 to 5.8 million years ago. By characterizing patterns of molecular divergence across an average of 11K genes and fitting geographic clines to an average of 7.5K variants, we characterize how patterns of genomic differentiation and introgression change through time. Across this range of divergences, we find that genome-wide differentiation increases but becomes no less heterogeneous. In contrast, we find that introgression heterogeneity decreases dramatically, suggesting that time helps isolated genomes "congeal". Thus, this work emphasizes the pivotal role that history plays in driving lineage divergence.
Data for "Clean ballistic quantum point contact in SrTiO3"
<p>Spreadsheets containing raw data, axis and trace labels from submitted manuscript "Clean ballistic quantum point contact in SrTiO<sub>3</sub>".</p>
complete_allsides_data_cleaned
<p>cleaned complete allsides data, without citations and special characters in the fulltexts</p>
Data from: Cleaning interactions by gobies on a Tropical Eastern Pacific coral reef
Open the record for dataset details and reuse information.
Data from: Cleaner personality and client identity have joint consequences on cleaning interaction dynamics
Open the record for dataset details and reuse information.
Data from: History cleans up messes: the impact of time in driving divergence and introgression in a tropical suture zone
Open the record for dataset details and reuse information.
Data from: Endovascular treatment in older adults with acute ischemic stroke in the MR CLEAN Registry
Open the record for dataset details and reuse information.
Raw sequencing data of Anaplamsa phagocytophilum loci (ankA, msp4, groEL) obtained from 454 and parameter files to clean these data using MOTHUR
<p>A compressed archive including: i) raw sequences in ssf file; ii) Mothur oligo files to sort out sequences among loci and individual samples.</p>
Supporting data for "Health benefits of US light-duty vehicle electrification: roles of fleet dynamics, clean electricity, and policy timing"
Open the record for dataset details and reuse information.
Cleaned data, cleaning code and analysis code for 'Feedback timing affects L2+ perceptual vowel acquisition'
<p>This dataset uses 4.3.1 and the analysis code requires use of the groundhog package (Simonsohn & Gruson, 2021) to aid reproducibility.</p>
137_Clean-data_EEG
<p>EEG data for project 137_2020_GVA_EEG (decoding of children's working memory content) that has been cleaned according to the preprocessing pipeline described in the document Preprocessing_steps.docx in the related OSF repository (https://osf.io/hrwtb/). Data are in .set format, which can be read by EEGLAB (Matlab).</p>
Data design thinking: data cleaning improvements using tableau prep
Open the record for dataset details and reuse information.
Apollo 17 ALSEP ARCSAV Lunar Ejecta And Meteorites Experiment Raw Cleaned ASCII Data Bundle
This bundle contains fixed-width ASCII files of daily, raw cleaned measurements acquired by the Lunar Ejecta And Meteorites (LEAM) Experiment at the Apollo 17 landing site for the time span of 02 April through 30 June 1975. These data were extracted from NASA's original Apollo Lunar Surface Experiments Package (ALSEP) archive tapes, also known as ARCSAV tapes.
Apollo 16 ALSEP ARCSAV Lunar Surface Magnetometer Raw Cleaned ASCII Data Bundle
This bundle contains fixed-width ASCII files of daily, raw cleaned measurements acquired by the Lunar Surface Magnetometer (LSM) Experiment at the Apollo 16 landing site for the time span of 02 April through 30 June 1975. These data were extracted from NASA's original Apollo Lunar Surface Experiments Package (ALSEP) archive tapes, also known as ARCSAV tapes.
Apollo 17 ALSEP ARCSAV Lunar Surface Gravimeter Raw Cleaned ASCII Data Bundle
This bundle contains fixed-width ASCII files of daily, raw cleaned measurements acquired by the Lunar Surface Gravimeter (LSG) at the Apollo 17 landing site for the time span of 02 April through 30 June 1975. These data were extracted from NASA's original Apollo Lunar Surface Experiments Package (ALSEP) archive tapes, also known as ARCSAV tapes.
Apollo 12 ALSEP ARCSAV Solar Wind Spectrometer Raw Cleaned ASCII Data Bundle
This bundle contains fixed-width ASCII files of daily, raw cleaned measurements acquired by the Solar Wind Spectrometer (SWS) at the Apollo 12 landing site for the time span of 03 April through 01 July 1975. These data were extracted from NASA's original Apollo Lunar Surface Experiments Package (ALSEP) archive tapes, also known as ARCSAV tapes.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.