Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

958

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

958 results for “Data Quality”

Learn how ShareScore rates datasets ↗
edi48/100

Missouri reservoir water quality data (2022 - current) from the Statewide Lake Assessment Program (SLAP)

This dataset of limnological water quality data continues from North et al., 2025, starting in 2022 until present. The data is from reservoirs, primarily within the state of Missouri (MO) in the USA collected by the University of Missouri Limnology Lab. Water quality parameters analyzed in the MU Limnology Lab during this time frame include: ammonium (NH4), anatoxin, chlorophyll a (corrected and uncorrected for pheophytins), chloride, cylindrospermopsin, dissolved organic carbon, microcystin, nitrate & nitrite (NO3), pheophytin, particulate inorganic matter, particulate organic matter, phycocyanin, saxitoxin, Secchi disk depth, total dissolved nitrogen (TDN), total dissolved phosphorus (TDP), total nitrogen (TN), total phosphorus (TP), total suspended solids, and urea. Most of the samples were collected during the summer months (May-September) when the reservoirs were thermally stratified, but a few were taken during the rest of the year (October-April). The majority of samples were taken at the deepest point in the reservoir directly up-reservoir of the dam. Sampling was conducted from a boat most of the time, but a few samples were taken from shorelines and drinking water treatment intake pipes. The bulk of the data come from the Statewide Lake Assessment Project (SLAP), funded by the Missouri Department of Natural Resources. This data represents duplicate or triplicate water samples collected from either the water surface, integrated over the depth of the epilimnion, or from discrete depths in the hypolimnion.

openCC (other)Jun 2025View details →
edi48/100

The Jefferson Project 2022 hydrologic, water quality, and soil quality data from 11 Tributary Stations within the Lake George basin, NY, USA.

The Jefferson Project at Lake George -- a partnership between Rensselaer Polytechnic Institute, IBM Research, and Lake George Association -- combines Internet of Things technology and powerful analytics with science to create a new model for environmental monitoring and prediction. The project is building a computing platform that captures and analyzes data from a network of sensors tracking water quality and movement. These sensor data are combined with other monitoring and experimental data to create a thorough understanding of the factors that drive the lake's food web, hydrology, and water quality. More information about The Jefferson Project is available at https://jeffersonproject.rpi.edu/ In 2022, The Jefferson Project had eleven tributary monitoring stations around the lake collecting data on water quality, soil quality, and hydrology. These stations are TS_Finkle, TS_Hague, TS_Indian, TS_NorthwestBay, TS_Outlet, TS_PoleHill, TS_English, TS_Sunset, TS_ShelvingRock, TS_East, and TS_West. The stations have a sensor payload that may include some or all of the following sensors: YSI EXO2 Multi-parameter sonde, Campbell Scientific CS451 pressure transducer, SonTek-IQ+ multi-beam acoustic flow meter, Sontek-SL Doppler current meter, YSI WaterLOG® H-3123 submersible pressure transducer, and Stevens HydraProbe soil moisture sensor. The sensors collect data at high-frequency (~1 sample per minute) and the data are transferred in near real-time to off-site databases for monitoring and review by Jefferson Project researchers. The data provided here are level 4 data which underwent data correction and downsampling to an hourly frequency.

openCC (other)Jul 2025View details →
edi48/100

The Jefferson Project 2022 water quality data from two vertical profiler stations in Lake George, NY, USA.

The Jefferson Project at Lake George -- a partnership between Rensselaer Polytechnic Institute, IBM Research, and Lake George Association -- combines Internet of Things technology and powerful analytics with science to create a new model for environmental monitoring and prediction. The project is building a computing platform that captures and analyzes data from a network of sensors tracking water quality and movement. These sensor data are combined with other monitoring and experimental data to create a thorough understanding of the factors that drive the lake's food web, hydrology, and water quality. More information about The Jefferson Project is available at https://jeffersonproject.rpi.edu/ In 2022, The Jefferson Project deployed two vertical profiler stations on the lake, collecting data on water quality and weather. Meteorological data have been included with the Jefferson Project Weather Station dataset for 2022. These vertical profiler stations are named VP_HarrisBay and VP_TeaIsland. The water quality data are collected by a YSI EXO2 Multi-parameter sonde sensors. The sensors collect data at 1 meter or less depth increments, starting at 1 meter and proceeding to 2 meters off bottom. The data are transferred in near real-time to an off-site database for monitoring and review. The data provided here have undergone data correction by Jefferson Project researchers.

openCC (other)Jul 2025View details →
edi48/100

Water Quality Data (Extensive) from the Taylor Slough, just outside Everglades National Park (FCE), from August 1998 to December 2006

Water quality samples are being collected using ISCO autosamplers at all wetland sites (that is, all sites except TS/Ph-9, 10, and 11). The autosamplers contain 24 1L bottles. Water is sampled by programming the autosamplers to take composite samples once every 3 days. These samples are a composite of four 250mL subsamples drawn every 18 hours (a sampling scheme that captures a dawn, noon, dusk, and midnight sample in every three day composite). Starting in December 2006 for Sites SRS1d, SRS2, SRS3, and June 2007 for TS/Ph1a, TS/Ph2, and TS/Ph3 - the 3 day composite samples are now combined in a 2 liter bottle upon returning to the lab to form 6 day composite sample. Samples are retrieved every 3-4 weeks. Retrieval of the composite samples may result in a composite sample which is less than three or six days. The recorded date for each composite sample indicates the end date of the sample interval. Samples are analyzed for total phosphorus (TP), total nitrogen (TN), and salinity. When sites are visited to collect these samples, we also collect a grab sample that is immediately put on ice. A portion of these grab samples are filtered through a Whatman GF/F filters (0.7 um) immediately upon return to the lab, and the filtered samples are analyzed for inorganic nutrients such as NO2-, NO3-, NH4+, SRP, and DOC. The unfiltered fraction of these grab samples is analyzed for TP, TN, and TOC (TOC is no longer analyzed starting in August 2005 for all sites). We use these montly grab samples to generate relationships between TP and SRP, and between TN and NO2- + NO3- + NH4+. Dissolved nutrients are measured using standard rapid flow analyzer (RFA) techniques. TP is analyzed with a modified Solorzano and Sharp (1980) technique. TN is measured with an Antec TN analyzer, TOC and DOC are quantified on a Shimadzu TOC Analyzer, and salinity is measured with a YSI conductivity meter. In addition to the regular water quality monitoring, we use the rain level actuators at all freshwat

openCC (other)Mar 2019View details →
edi48/100

North Temperate Lakes LTER: Multiparameter Water Quality Data -- CFL Pier, Lake Mendota.

This is data from a SUNA V2 nitrate sensor and a YSI EXO 2 sonde instrumented with water temperature, dissolved oxygen, pH, chlorophyll, phycocyanin, conductivity, turbidity, and fDOM sensors. The sensors are located at the lake end of the pier serving the Center for Limnology on the UW-Madison campus. The YSI sonde is fixed on the pier with sensors nominally at 0.5 meters depth, while the SUNA is suspended below the pier at one meter depth in the open water season. In the winter season, the SUNA is placed in a cage on the lake bottom close to shore with the sensor 18cm off the bottom. The depth of the sensors will vary with lake level over the season. The water depth at the lake end of the pier is normally about 3 meters. YSI sonde data are sampled once per minute. Hourly and daily averages are provided as separate CSV files. The SUNA sample rate varies. In the winter (under the ice) it relies on single battery charge, so the wiper is deactivated and the sensor samples every 1-2 hours. Daily averages of SUNA are also provided as a separate CSV. The YSI sonde is deployed only during the ice-free season coinciding with the placement of the pier. Sensors are cleaned and maintained roughly every two weeks. Number of sites: 1. Location lat/long: 43.07758, -89.40297

openCC (other)Aug 2025View details →
zenodo44/100

MACREL software benchmark data set: Simulated metagenomes with sequencing quality, errors profile and abundance distributions derived from real samples

<p>These metagenomes were used in the benchmarking of FACS pipeline, and were designed after NGLess benchmark dataset (doi.org/10.5281/zenodo.2560288).&nbsp; Metagenomes were simulated with <a href="https://www.niehs.nih.gov/research/resources/software/biostatistics/art/index.cfm">ART-bin-MountRainier-2016.06.05</a> using real abundance profiles (.abund files) available <a href="https://doi.org/10.5281/zenodo.2560288">elsewhere</a>, and <a href="http://progenomes1.embl.de/data/repGenomes/representatives.contigs.fasta.gz">proGenomes&#39; representative contigs</a> as reference genomes. There are available metagenomes with 40, 60 and 80 M (million of reads) based in the reference genomes and abundances of the following samples:</p> <pre><code>SAMEA2466916 SAMEA2466953 SAMEA2466965 SAMEA2621107 SAMEA2621229 SAMEA2621247</code></pre> <p>To convert them from the CRAM format back to fastq files:</p> <pre><code> ## 1. converting from cram to bam format: samtools view -b -T refgenome.fa -o file.bam file.cram ## 2. sorting the bam file: samtools sort -n file.bam -o input_sorted.bam # sort reads by identifier-name (-n) ## 3. converting from bam to fastq format: bedtools bamtofastq -i input_sorted.bam -fq output_r1.fastq -fq2 output_r2.fastq </code></pre> <p>&nbsp;</p>

opencc-by-4.0Nov 2019View details →
zenodo44/100

Data relating to Clyne et al. Quality, scope and reporting standards of randomised controlled trials in Irish Health Research: an observational study

<p>Data relating to the study reported in the paper &quot;Quality, scope and reporting standards of randomised controlled trials in Irish Health Research: an observational study&quot;.</p>

opencc-by-4.0Apr 2020View details →
zenodo44/100

Data and R-Scripts for "Quality and timing of crowd-based water level class observations"

<p>This are the data and the R-scripts used for the manuscript &quot;Quality and timing of crowd-based water level class observations&quot; accepted for publication in the journal Hydrological Processes in July 2020 as a Scientific Briefing. To run the code, just run the R-script with the name &quot;RunThisForResults.R&quot;. Results will be written to the &quot;Figures&quot; and the &quot;Results&quot; folder.</p>

opencc-by-4.0Feb 2020View details →
zenodo44/100

Water quality and diarrhoea bibliometric data and visualisation

<p>This repository contains bibliometric data and its visualisations:</p> <p>Bibliometric data</p> <ul> <li>Keyword: "water quality" AND diarrhoea</li> <li>Database: Scopus</li> <li>Date taken: 28 June 2017</li> <li>Formats: bib, csv, ris</li> <li>Reference manager: Jabref and Zotero</li> </ul> <p>Visualisations</p> <ul> <li>Tools: VosViewer (http://VosViewer.com)</li> <li>Tool's citation: Van Eck, N.J., &amp; Waltman, L. (2010). Software survey: VOSviewer, a computer program for bibliometric mapping. Scientometrics, 84(2), 523-538. (paper, preprint, supplementary material) (http://dx.doi.org/10.1007/s11192-009-0146-3)</li> <li>Mindmap of analysis procedures using Freeplane https://www.freeplane.org/wiki/index.php/Main_Page</li> </ul>

opencc-by-4.0Jun 2017View details →
zenodo44/100

Weekly county-level pollution data for China from Zhang, Carleton, Lin, and Zhou (accepted, Nature Sustainability), "Estimating the role of air quality improvements in the decline of suicide rates in China"

<p>This dataset contains weekly, county-level air pollution data for 2,839 counties from 2013 to early 2018. These data are used and described in Zhang, Carleton, Lin, and Zhou (accepted,&nbsp;<em>Nature Sustainability</em>), "Estimating the role of air quality improvements in the decline of suicide rates in China". When the paper is published a link to the manuscript will be added here.&nbsp;</p> <p>The manuscript Methods section details data construction. In summary, these county-level observations are obtained from monitoring stations maintained by the China National Environmental Monitoring Center (CNEMC), which is affiliated with the Ministry of Ecology and Environment of China. CNEMC began publishing hourly air pollution data in 2013, including the Air Quality Index, PM2.5, PM10, ozone, sulfur dioxide, nitrogen dioxide, and carbon monoxide. We average hourly data to the station-day level and use inverse-distance weighting with a radius of 200km to convert data from station to the county level. We average across days to generate county-level weekly values. Any missing station-hour observations in the raw data are omitted in this spatial and temporal aggregation. Our main analysis relies on PM2.5, but all pollutants are released here.</p>

opencc-by-4.0Dec 2023View details →
zenodo44/100

Quality controlled observations of hourly incoming shortwave radiation data at the surface for solar resource mapping in Norway (2016-2020).

<p>Observed hourly incoming shortwave radiation data at the surface of Norway for the years 2016-2020 along with quality control flags, visualization plots and a descriptive report. The data has been collected, visually inspected and quality controlled within the SunPoint project (SUn in Norway - POtential and INTegration of the solar energy resource, Norwegian Research Council project 320750). The main data source is frost.met.no but some gaps were filled with data directly obtained by the station holders.</p> <p>There are three NetDCF files for 47 stations selected after quality control:</p> <ul> <li>rsds_1hr_selection_v5_2016-2020.nc: Raw data</li> <li>rsds_flagged_1hr_selection_v5_2016-2020.nc: Raw data with flags</li> <li>rsds_cleaned_1hr_selection_v5_2016-2020.nc: Filtered data (i.e. all flagged data has been removed)</li> </ul> <p>and one NetCDF file for all available stations (106 stations)</p> <ul> <li>rsds_1hr_frost_and_more_2016-2020.nc</li> </ul> <p>Version 3.5 of the McClear clear-sky model is used for flagging which reduces the bias to ground measurements compared to earlier versions.&nbsp;</p> <p>The visualization and automated quality control routines are available in the Scripts.zip file (python).</p>

opencc-by-4.0Mar 2024View details →
zenodo44/100

Machine learning-based quality assessment of Antarctic margins salinity - code, data and figures

<p>The submission contains the data, functions and code needed to reproduce the figures in Sohail et al., 2025</p>

opencc-by-4.0Nov 2024View details →
zenodo44/100

Supplementary data for Willemsen et al., 2024 "Novel high-quality amoeba genomes reveal widespread codon usage mismatch between giant viruses and their hosts".

<p>Supplementary data for Willemsen et al., 2024 "Novel high-quality amoeba genomes reveal widespread codon usage mismatch between giant viruses and their hosts". The data set consists of five folders: &ldquo;Codon_usage_amoebae_and_viruses&rdquo;, "Genome_annotations_amoebae", "Phylogenetic_trees_18S_amoebae", &ldquo;Phylogenomic_trees_amoebae&rdquo;, and "Viral_integration_detection_amoebae". The &ldquo;Codon_usage_amoebae_and_viruses&rdquo; folder contains for each amoeba host the calculated codon usage tables in the subfolder "codon_usage_table_host", the calculated codon usage preferences using different scores in the subfolder "codon_usage_scores_host", and the calculated codon usage preferences of giant viruses versus each host in the subfolder "codon_usage_scores_viruses_vs_host". The giant viruses in the subfolder "codon_usage_scores_viruses_vs_host" are organised by viral family and genus in separate sub-subfolders. The "Genome_annotations_amoebae" folder contains the generated genome annotations in different formats and the manually curated mitochondrial genome annotations for each amoeba host.&nbsp; The "Phylogenetic_trees_18S_amoebae" contains for the eukaryotic phyla <em>Discosea</em>, <em>Heterolobosea</em>, and <em>Tubulinea,&nbsp;</em>the 18S rRNA&nbsp;nucleotide alignments, distance matrices, and computed phylogenetic trees. The folder "Phylogenomic_trees_amoebae" contains for the eukaryotic clades <em>Amoebozoa</em> and <em>Discoba,&nbsp;</em>the protein alignment matrices and computed phylogenomic trees. The folder "Viral_integration_detection_amoebae" contains the MCP databases used (fasta file, alignment file, HMM profile and DIAMOND BLASTX database) and the MCP sequences detected in this study and the blast results of these.&nbsp;&nbsp;</p>

opencc-by-4.0Jul 2024View details →
zenodo44/100

Data to "Phantom-based quality assurance for multicenter quantitative MRI in locally advanced cervical cancer"

<p>This record includes the DICOM images and analysed data that were used in the multicenter QA program for quantitative MRI in cervical cancer as published (<a href="https://www.sciencedirect.com/science/article/pii/S0167814020307854?via%3Dihub">https://doi.org/10.1016/j.radonc.2020.09.013</a> ).</p> <p>The DICOM data includes the acquired DICOM data for each institute selected to those that were used in the publication. Acquisitions that were not used were removed. Data was anonymized with conquest dicom server tools.</p> <p>The analyzed data files are included giving per measurement the estimated quantitative parameter values as well as the position of the ROIs and extracted signal intensity values per phantom sample. An explanation of the structure of the files is added in the readme file. The analysis was done with in-house written code in matlab.</p> <p>Included are a description of the sequence parameters for each institute (IQEMBRACE_PhantomQA_OverviewInstitutionalSequenceParameters_20241114) and details on the choices in the analysis of the data (IQEMBRACE_PhantomQA_OverviewPhantomData_20241114). As background also the description of the measurements was added, giving more information on how the measurements were performed.</p> <p>This work was in preparation for the IQ-EMBRACE trial (clinicaltrials.gov NCT03210428)</p>

opencc-by-4.0Nov 2024View details →
zenodo44/100

GWTC-3: Compact Binary Coalescences Observed by LIGO and Virgo During the Second Part of the Third Observing Run — Data Quality Products for GW Searches

<p>This material is part of several data products associated with GWTC-3, the third Gravitational-Wave Transient Catalog from the <a href="https://www.ligo.org/">LIGO</a> Scientific Collaboration, the <a href="https://www.virgo-gw.eu/">Virgo</a> Collaboration, and the <a href="https://gwcenter.icrr.u-tokyo.ac.jp/en/">KAGRA</a> Collaboration. For more information, see the paper (<a href="https://dcc.ligo.org/LIGO-P2000318/public">dcc.ligo.org/LIGO-P2000318/public</a>), the related material linked from this page, and the GWTC-3 data release documentation (<a href="https://www.gw-openscience.org/GWTC-3/">www.gw-openscience.org/GWTC-3/</a>).</p> <p>This release contains two data-quality products that are used by search analyses to help mitigate non-Gaussian noise in the detector data. Gating removes&nbsp;short-duration artifacts from the data by smoothly rolling the affected data&nbsp;to zero. The&nbsp;<a href="https://doi.org/10.1088/2632-2153/abab5f">iDQ glitch likelihood</a> uses machine learning to predict the probability that a non-Gaussian transient is present using&nbsp;information from auxiliary channels.</p> <p><br> <strong>Gating files used in analyses of O3 LIGO data</strong></p> <p>As a pre-processing step, the <a href="https://pycbc.org/">PyCBC</a> search pipeline uses an inverted-Tukey window to mitigate the effect of loud, non-Gaussian features in the data. This is further described in&nbsp;<a href="https://dx.doi.org/10.1088/0264-9381/33/21/215004">Usman <em>et al.</em> 2016</a>.</p> <p>A subset of these times are the times listed in the txt files</p> <ul> <li>H1-O3_GATES_1238166018-31197600.txt</li> <li>L1-O3_GATES_1238166018-31197600.txt</li> </ul> <p>These times in these files were chosen based on auxiliary monitors of overflows in the digital-to-analog converters used to control the positions of the test masses. The gated times (i.e. the time period where the data is zeroed) are time segments where these monitors recorded an overflow were. The central time&nbsp;and suggested half-width of zero time&nbsp;were chosen to fully cover these time seconds. The final gating parameter, the suggested taper time&nbsp;was chosen to be 0.5 to balance the cost of impacting more data with the window function versus introducing additional artifacts into the data.</p> <p>The syntax of the files themselves is</p> <p>{central time} {suggested half-width of zero time} {suggested taper time}</p> <p>with each row containing the parameters of a single gate.</p> <p>The included notebook provides an example of how to read in and apply one of the suggested gates.</p> <p><br> <strong>Renormalized iDQ timeseries</strong></p> <p>The renormalized iDQ timeseries data-quality product is used within the GstLAL search pipeline to generate results for GWTC-3. This data product was found to be statistically helpful in improving data quality within the <a href="https://lscsoft.docs.ligo.org/gstlal/">GstLAL</a> search pipeline. This is further described in <a href="https://arxiv.org/abs/2010.15282">Godwin <em>et al</em>. 2020</a>.</p> <p>This file contains a time series for each LIGO detector related to&nbsp;the probability of a glitch in the&nbsp;strain data given the behavior in analyzed auxiliary channels monitoring the behavior of the detectors and their environment.</p> <ul> <li>H1L1-IDQ_TIMESERIES-1256655642-12905976.h5</li> </ul> <p>The HDF5-formatted file contains two groups, H1 and L1, corresponding to LIGO Hanford and LIGO Livingston, respectively. Each group contains several datasets; the data dataset corresponds to the renormalized iDQ log-likelihoods, as described in <a href="http://doi.org/10.1088/2632-2153/abab5f">Godwin <em>et al</em>. 2020</a>, and the time dataset corresponds to the times associated with the renormalized iDQ log-likelihoods in the data&nbsp;dataset.</p> <p>&nbsp;</p> <p><strong>How to download all files from this page</strong></p> <p>If you would like to download all files on this page, we recommend <a href="https://gitlab.com/dvolgyes/zenodo_get">zenodo_get</a>:</p> <pre><code class="language-bash">pip install zenodo_get zenodo-get RECORD_ID_OR_DOI </code></pre> <p>where the record ID for the most recent version of this page is&nbsp;5636795 and IDs for other versions can be found in the Versions section at the side of this page.</p> <p>&nbsp;</p> <p>For more general background on gravitational-wave data quality, try the materials from a <a href="https://www.gw-openscience.org/workshops/">GW Open Data Workshop</a> or the <a href="https://doi.org/10.1088/1361-6382/ab685e">guide to LIGO&ndash;Virgo data analysis</a>.&nbsp;</p>

opencc-by-4.0Nov 2021View details →
zenodo44/100

Data used to create figures and tables in the ACP manuscript "Two-way coupled meteorology and air quality models in Asia: a systematic review and meta-analysis of impacts of aerosol feedbacks on meteorology and air quality" by Gao et al. (2022)

<p>This dataset contains the original data that extracted from all collected papers refering applications of two-way coupled&nbsp;models in Asia. It is supplied to the review paper, which titled as &quot;Review&nbsp;on&nbsp;two-way coupled meteorology and air quality models in Asia: impacts of aerosol feedbacks on meteorology and air quality&quot;. The dataset includes three excel files (in the format of xlsx) as follows:</p> <p>1. Basic information of literatures&nbsp;(Table S1.xlsx)</p> <p>2. Model performance metrics (Table S2.xlsx)</p> <p>3. Quantitative results of aerosol effects on meteorological and air quality variables (Table S3.xlsx)</p> <p>4.&nbsp;Basic information of model setup for two-way coupled model applications in Asia (Table S4.xlsx)</p> <p>5.&nbsp;Summary of aerosol-induced variations of simulated shortwave and longwave radiative forcing at the bottom and top of atmosphere and in the atmosphere in Asia (Table S5.xlsx)</p> <p>.</p>

opencc-by-4.0Feb 2022View details →
zenodo44/100

Effect of spatial input data quality on SWAT modelling in the Porijõgi catchment

<p>The Porij&otilde;gi Catchment near Tartu, Estonia is the study area for this research. Four model setups were created using global/regional level data (HWSD soil, CORINE), and local high-resolution spatial data including the new Estonian high-resolution EstSoil-EH soil dataset and the Estonian Topographic Database (ETAK). The study employed statistical criteria to assess SWAT model performance for monthly simulated stream flows from 2007 to 2019.</p> <p>Data deposit in preparation for article:</p> <p>Effect of spatial input data quality on the uncertainty of the<br> SWAT model, submitted 2022</p> <p>Alexander Kmoch, Desalew Meseret Moges, Mahdiyeh Sepehrar, Balaji Narasimhan and Evelyn<br> Uuemaa</p> <p>contact: alexander.kmoch@ut.ee</p>

opencc-by-4.0Mar 2022View details →
zenodo44/100

GWTC-2.1: Deep Extended Catalog of Compact Binary Coalescences Observed by LIGO and Virgo During the First Half of the Third Observing Run - Data Quality Products for GW Searches

<p>This material is part of several data products associated with GWTC-2.1, the deep extended catalog of compact binary coalescences observed by the <a href="https://www.ligo.org/">LIGO</a> Scientific Collaboration and the <a href="https://www.virgo-gw.eu/">Virgo</a> Collaboration during the first half of the third observing run. For further information, see the paper (<a href="https://dcc.ligo.org/LIGO-P2100063/public">dcc.ligo.org/LIGO-P2100063/public</a>), the related material linked from this page, and the GWTC-2.1&nbsp;data release documentation (<a href="https://www.gw-openscience.org/GWTC-2.1/">www.gw-openscience.org/GWTC-2.1/</a>).</p> <p>This release contains data quality products that are used by search analyses to help mitigate non-Gaussian noise in the detector data.</p> <p><strong>Renormalized iDQ timeseries</strong></p> <p>This release contains the renormalized iDQ timeseries data quality product used within the GstLAL search to generate results for GWTC-2.1 as described in <a href="https://arxiv.org/abs/2010.15282">Goodwin <em>et al</em>. 2020</a>. This data product was found to be statistically helpful in improving data quality within the GstLAL search. For further information about iDQ see <a href="https://iopscience.iop.org/article/10.1088/2632-2153/abab5f">Essick <em>et al</em>. 2020</a>.</p> <p>The file&nbsp;</p> <ul> <li>H1L1-IDQ_TIMESERIES-1238166018-15843600.h5</li> </ul> <p>contains a time series for each LIGO detector related to the probability of a glitch in the strain data given the behavior in the analyzed auxiliary channels which monitor the behavior of the detectors and their environment.</p> <p>The HDF5-formatted file contains two groups, H1 and L1, corresponding to LIGO Hanford and LIGO Livingston, respectively. Each group contains several datasets; the data dataset corresponds to the renormalized iDQ log-likelihoods, as described in Godwin <em>et al</em>. 2020, and the time dataset corresponds to the times associated with the renormalized iDQ log-likelihoods in the data&nbsp;dataset.</p> <p>&nbsp;</p> <p>For more general background on gravitational-wave data quality, try the materials from a <a href="https://www.gw-openscience.org/workshops/">GW Open Data Workshop</a> or the guide to <a href="https://doi.org/10.1088/1361-6382/ab685e">LIGO-Virgo data analysis</a>.</p>

opencc-by-4.0Apr 2022View details →
zenodo44/100

Underlying Data for Manuscript titled "Hydrothermal Carbonization (HTC) of Dairy Waste: Effect of Temperature and Initial Acidity on the composition and quality of solid and liquid products"

<p>The embodied files include the raw data and initial calculations used to generate the extended data for Manuscript titled &quot;Hydrothermal Carbonization (HTC) of Dairy Waste: Effect of Temperature and Initial Acidity on the composition and quality of solid and liquid products&quot;. The files include calculations for Phosphorus Recovery from Hydrochar, as well as Heavy Metals Fractions retrieved by the hydrochar (solid product of Hydrothermal Carbonization).</p>

opencc-by-4.0Jun 2022View details →
zenodo44/100

Adapting the Harmonized Data Quality Framework for Ontology Quality Assessment

<p>Ontologies play an important role in the representation, standardization, and integration of biomedical data, but are known to have data quality (DQ) issues. We aimed to understand if the Harmonized Data Quality Framework (HDQF), developed to standardize electronic health record DQ assessment strategies, could be used to improve ontology quality assessment. A novel set of 14 ontology checks was developed. These DQ checks were aligned to the HDQF and examined by HDQF developers. The ontology checks were evaluated using 11 Open Biomedical Ontology Foundry ontologies. 85.7% of the ontology checks were successfully aligned to at least 1 HDQF category. Accommodating the unmapped DQ checks (n=2), required modifying an original HDQF category and adding a new Data Dependency category. While all of the ontology checks were mapped to an HDQF category, not all HDQF categories were represented by an ontology check presenting opportunities to strategically develop new ontology checks. The HDQF is a valuable resource and this work demonstrates its ability to categorize ontology quality assessment strategies.</p>

opencc-by-4.0Jul 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record