Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

419

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

419 results for “large dataset”

Learn how ShareScore rates datasets ↗
zenodo44/100

Dataset: High emission rates and strong temperature response make boreal wetlands a large source of isoprene and terpenes

<p>Dataset used in the article &quot;High emission rates and strong temperature response make boreal wetlands a large source of isoprene and terpenes&quot;</p> <p>The tab-delimited file contains direct surface-atmosphere Volatile Organic Compound fluxes, measured by Eddy Covariance with a Vocus- proton transfer reaction mass spectrometer (Vocus-PTR) at a subarctic fen&nbsp;during 2021. It also contains PAR (Photosynthetic Active Radiation), temperature and flux&nbsp;quality criteria.</p>

opencc-by-4.0Aug 2022View details →
zenodo44/100

The COUGHVID crowdsourcing dataset: A corpus for the study of large-scale cough analysis algorithms

<p><strong>Overview</strong></p> <p>Cough audio signal classification has been successfully used to diagnose a variety of respiratory conditions, and there has been significant interest in leveraging Machine Learning (ML) to provide widespread COVID-19 screening. The COUGHVID dataset provides over 30,000 crowdsourced cough recordings representing a wide range of subject ages, genders, geographic locations, and COVID-19 statuses. Furthermore, experienced pulmonologists labeled more than 2,000 recordings to diagnose medical abnormalities present in the coughs, thereby contributing one of the largest expert-labeled cough datasets in existence that can be used for a plethora of cough audio classification tasks.&nbsp;As a result, the COUGHVID dataset contributes a wealth of cough recordings for training ML models to address the world&rsquo;s most urgent health crises.</p> <p><strong>Private Set and Testing Protocol</strong></p> <p>Researchers interested in testing their models on the private test dataset should contact us at coughvid@epfl.ch, briefly explaining the type of validation they wish&nbsp;to make, and their obtained results obtained through&nbsp;cross-validation with the public data. Then, access to the unlabeled recordings will be provided, and&nbsp;the researchers should&nbsp;send the predictions of their models on these recordings. Finally,&nbsp;the&nbsp;performance metrics of the predictions will be sent to the researchers. The private testing data is not included in any file within our Zenodo record, and it can only be accessed by contacting the COUGHVID team at the aforementioned e-mail address.</p> <p><strong>New Semi-Supervised Labeling</strong></p> <p>The third version of the COUGHVID dataset contains thousands of additional recordings obtained through October 2021. Additionally, the recordings containing coughs were re-labeled according to a semi-supervised learning algorithm that combined the user labels with those of the expert physicians, which were&nbsp;modeled using ML and expanded on the previously unlabeled data. These labels can be found in the &quot;status_SSL&quot; column of the &quot;metadata_compiled.csv&quot; file.</p>

opencc-by-4.0Sep 2020View details →
zenodo44/100

Large-Eddy Simulation of Wind Turbine Flows: A New Evaluation of Actuator Disk Models - Dataset

<p>Main data used in the following paper: Revaz, T.; Port&eacute;-Agel, F. Large-Eddy Simulation of Wind Turbine Flows: A New Evaluation of Actuator Disk Models. <em>Energies</em> <strong>2021</strong>, <em>14</em>, 3745. https://doi.org/10.3390/en14133745</p>

opencc-by-4.0Sep 2022View details →
zenodo44/100

Rafael et al., 2018: Deep Learning-Based Culture-Free Bacteria Detection in Urine Using Large-Volume Microscopy (Dataset)

<p>Dataset containing 1um polystyrene beads, urine samples, urine samples mixed with ecoli and homogenous ecoli. Dataset is the post-processing version of the images to remove static background. Original model was trained on the post-processed images exclusiviely.</p>

opencc-by-4.0Feb 2024View details →
zenodo44/100

LSPO: A Large-Scale Physics ORCiD-Linked Dataset for Author Name Disambiguation

<p>The LSPO dataset, a Large-Scale Physics ORCiD-Linked Dataset for Author Name Disambiguation is comprised of 554,962 NASA/ADS publications linked to 125,486 unique researchers through ORCiD identifiers. The available meta-data fields are: ORCiD identifier, author name, affiliation, title, asbtract, and name block. The dataset can be utilized to make pairs or triplets for training a author name disambiguation model.&nbsp;</p>

opencc-by-4.0Jun 2024View details →
zenodo44/100

Datasets for "Advancing Drug-Target Interactions Prediction: Leveraging a Large-Scale Dataset with a Rapid and Robust Chemogenomic Algorithm"

<p>All datasets required to reproduce the results of publication "Drug-Target Interactions Prediction at Scale: the Komet Algorithm with the LCIdb Dataset"</p>

opencc-by-4.0Feb 2024View details →
zenodo44/100

Dataset for Paper Titled "First assessment of cloud-land coupling in LASSO Large-Eddy Simulations"

<div>The attached two files were used for analysis in the paper "First assessment of cloud-land coupling in LASSO Large-Eddy Simulations." The NetCDF file included planetary boundary layer heights derived from lidar and radiosondes. The CSV file detailed the model configurations for selected case days in simulation sets ID1-5.</div> <div> <div> <p>&nbsp;</p> </div> </div>

opencc-by-4.0Jun 2024View details →
zenodo44/100

UnientrezDB: Large-scale Gene Ontology Annotation Dataset and Evaluation Benchmarks with Unified Entrez Gene Identifiers

<p>Our work focuses on providing a comprehensive dataset and benchmarks for evaluating gene ontology annotations using a unified system of Entrez Gene Identifiers.</p>

opencc-by-4.0Aug 2024View details →
zenodo44/100

Dataset - Papyrus 2024 - A large scale curated dataset aimed at bioactivity predictions

<p><strong>This update of release 2024.1 fixes the following:</strong></p> <ul> <li>Metadata in the columns <em>type_IC50</em>, <em>type_EC50</em>, <em>type_KD</em>, <em>type_Ki</em>, and <em>type_other</em> did not contain multiple values when multiple pChEMBL values where available but reported only a single value. This fix ensures all values are reported.</li> <li>Molecules were incorrectly standardized and mixtures were included in the dataset. Standardization (using the&nbsp;<a href="https://github.com/OlivierBeq/papyrus_structure_pipeline" target="_blank" rel="noopener">papyrus_structure_pipeline</a>) is now correctly enforced and mixtures have been removed.</li> </ul> <p><strong>Changes since version 05.6</strong></p> <ul> <li>ChEMBL data was updated to ChEMBL version 34</li> <li>data from the IUPHAR/BPS Guide to PHARMACOLOGY has been included</li> <li>data from Pickett et al.'s publication on MMP-12 has been included (<a href="https://doi.org/10.1021/ml100191f">ACS Med Chem Lett. 2011 Jan 13; 2(1): 28&ndash;33. DOI: 10.1021/ml100191f</a>)</li> </ul> <p><strong>Papyrus++:</strong></p> <p>Previous versions mistakenly considered a deviation of 2 log units around compound-target pairs to determine the reproducibility of assays (see published article for more details). This has been fixed to 0.5 log units to ensure data points fall within a maximum range of 1 log unit. As a result, the number of entries in the Papyrus++ set from this release has drastically reduced compared to previous releases.</p>

opencc-by-sa-4.0Dec 2023View details →
zenodo44/100

MultiSubs: A Large-scale Multimodal and Multilingual Dataset

<p>MultiSubs&nbsp;is a dataset of multilingual subtitles gathered from&nbsp;<a href="https://opus.nlpl.eu/OpenSubtitles.php">the OPUS OpenSubtitles dataset</a>,&nbsp;which in turn was sourced from <a href="http://www.opensubtitles.org/">opensubtitles.org</a>. We have supplemented some text fragments (visually salient nouns in this release) within the subtitles with web images, where the word sense of the fragment has been disambiguated using a cross-lingual approach.&nbsp;</p> <p>Please refer to our&nbsp;paper for a more detailed description of the dataset:</p> <p>Josiah Wang, Pranava Madhyastha, Josiel Figueiredo, Chiraag Lala, Lucia Specia (2021). <a href="https://arxiv.org/abs/2103.01910">MultiSubs: A Large-scale Multimodal and Multilingual Dataset</a>. CoRR, abs/2103.01910. Available at: <a href="https://arxiv.org/abs/2103.01910">https://arxiv.org/abs/2103.01910</a></p>

opencc-by-4.0Jun 2021View details →
zenodo44/100

Single objective light-sheet acquired large-scale 3D dataset

<p>This dataset covers Fig. 5 of the following publication:</p> <p>Title: Tilt-invariant scanned oblique plane illumination microscopy for large-scale volumetric imaging<br> Authors: Manish Kumar and Yevgenia Kozorovitskiy&nbsp;<br> Optics Letters Vol. 44, Issue 7, pp. 1706-1709 (2019)<br> https://doi.org/10.1364/OL.44.001706</p> <p>Briefly: The sample imaged is a Thy1GFP expressing transgenic&nbsp;mice brain slice - fixed and coverslipped. No clearing was performed for these.</p> <p>See &quot;readme.txt&quot; for additional info and details about how to use &quot;shearNscaleObliq&quot; file.</p>

opencc-by-4.0Jul 2021View details →
zenodo44/100

Building Large-Scale Gene-Disease Association Datasets for Biomedical Relation Extraction

<p>This repository contains the GDAb and GDAt datasets. GDAb and GDAt are large-scale, distantly supervised, and manually enhanced datasets for Gene-Disease Association (GDA) extraction. Each dataset consists of three text files, corresponding to train, validation, and test sets, plus an additional JSON file containing the mapping between relation names and IDs. Each record in train, validation, or test files&nbsp;corresponds to a single GDA extracted from a sentence. Records are represented as JSON objects with the following structure:</p> <ul> <li><strong>text:</strong>&nbsp;sentence from which the GDA was extracted.</li> <li><strong>relation:</strong>&nbsp;relation name associated to the given GDA.</li> <li><strong>h:&nbsp;</strong>JSON object representing the gene entity, composed of: <ul> <li><strong>id:&nbsp;</strong>UMLS CUI associated to the gene entity.</li> <li><strong>name:</strong>&nbsp;UMLS preferred name associated to the gene entity.</li> <li><strong>pos:&nbsp;</strong>list consisting of starting position and length of the gene mention within text.</li> </ul> </li> <li><strong>t:</strong>&nbsp;JSON object representing the disease entity, composed of: <ul> <li><strong>id:&nbsp;</strong>UMLS CUI associated to the disease entity.</li> <li><strong>name:</strong>&nbsp;UMLS preferred name associated to the disease entity.</li> <li><strong>pos:</strong>&nbsp;list consisting of starting position and length of the disease mention within text.</li> </ul> </li> </ul> <p>Both datasets contain over 2,500,000 sentences and 500,000 bags.<br> The zip file consists of two folders, GDAb and GDAt,&nbsp;&nbsp;containing the files corresponding to the two datasets, respectively.</p> <p>&nbsp;</p>

opencc-by-4.0Jul 2021View details →
zenodo44/100

EA-MD-QD: Large Euro Area and Euro Member Countries Datasets for Macroeconomic Research

<p>EA-MD-QD is a collection of large monthly and quarterly EA and EA member countries datasets for macroeconomic analysis.<br>The EA member countries covered are: AT, BE, DE, EL, ES, FR, IE, IT, NL, PT.</p> <p>The formal reference to this dataset is:&nbsp;</p> <p><strong>Barigozzi, M. and Lissona, C. (2024) "EA-MD-QD: Large Euro Area and Euro Member Countries Datasets for Macroeconomic Research". Zenodo.</strong></p> <p>Please refer to it when using the data.</p> <p>Each zip file contains:<br><br>- Excel files for the EA and the countries covered, each containing an unbalanced panel of raw de-seasonalized data.<br><br>- A Matlab code taking as input the raw data and allowing to perform various operations such as:<br>choose the frequency, fill-in missing values, transform data to stationarity, and control for covid outliers.<br><br>- A pdf file with all informations about the series names, sources, and transformation codes.</p> <p><strong>This version (10.2025):</strong></p> <p>Updated data as of 31-October-2025.&nbsp;</p>

opencc-by-nc-4.0Jan 2024View details →
zenodo44/100

TweetDIS: A Large Twitter Dataset for Natural Disasters Built using Weak Supervision

<p>This repository contains the silver standard dataset and code for the paper &quot;TweetDIS: A Large Twitter Dataset for Natural Disasters Built using Weak Supervision&quot;.</p> <p>The file &quot;heuristic_uniq_terms_nd.txt&quot; contains the list of terms used as the heuristic and the file &quot;natural_disasters_ssd_tweetids.tsv&quot; contains the tweet ids in&nbsp;the silver standard dataset.&nbsp;</p> <p>To hydrate the tweets, you can use tools like twarc or Social Media Mining toolkit - https://www.ncbi.nlm.nih.gov/pmc/articles/PMC7362951/</p>

opencc-by-4.0Jun 2022View details →
zenodo44/100

CABra: a novel large-sample dataset for Brazilian catchments

<p>Hydrometeorological time series and catchment&nbsp;attributes from the CABra dataset. The manuscript of &quot;CABra: a novel large-sample dataset for Brazilian catchments&quot; is under review in&nbsp;Hydrology and Earth System Sciences (HESS) journal.</p> <p>Here we present the Catchments Attributes for Brazil (CABra), which is a large-sample dataset for Brazilian catchments that includes long-term data (30 years) for 735 catchments in eight main catchment attribute classes (climate, streamflow, groundwater, geology, soil, topography, land-use and land-cover, and hydrologic disturbance). We have collected and synthesized data from multiple sources (ground stations, remote sensing, and gridded datasets). To prepare the dataset, we delineated all the catchments using the Multi-Error-Removed Improved-Terrain Digital Elevation Model and the coordinates of the streamflow stations provided by the Brazilian Water Agency (ANA), where only the stations with 30 years (1980-2010) of data and less than 10% of missing records were included. Catchment areas range from 9 to 4,800,000 km&sup2; and the mean daily streamflow varies from 0.02 to 9 mm day<sup>-1</sup>. Several signatures and indices were calculated based on the climate and streamflow data. Additionally, our dataset includes boundary shapefiles, geographic coordinates, and drainage areas for each catchment, aside from more than 100 attributes within the attribute classes.</p> <p>Data can also be accessed at: thecabradataset.shinyapps.io/CABra&nbsp;</p> <p>&nbsp;</p> <p><em><strong>* This version includes water demand in CABra catchments for 2020 and 2040 (projection).</strong></em></p>

opencc-by-4.0Oct 2020View details →
zenodo44/100

[Dataset] One year of high-precision operational data including measurement uncertainties from a large-scale solar thermal collector array with flat plate collectors, located in Graz, Austria

<p><strong>Highlights:</strong></p> <ul> <li>High-precision measurement data acquired within a scientific research project, using high-quality measurement equipment and implementing extensive data quality assurance measures.</li> <li>The dataset includes data from one full operational year in a 1-minute sampling rate, covering all seasons.</li> <li>Measured data channels include global, beam and diffuse irradiances in horizontal and collector plane. Heat transfer fluid properties were determined in a dedicated laboratory test.</li> <li>In addition to the measured data channels, calculated data channels, such as thermal power output, mass flow, fluid properties, solar incidence angle and shadowing masks are provided to facilitate further analysis.</li> <li>Uncertainties of data channels are provided based on data sheet specifications and GUM error propagation.</li> <li>The dataset refers to a real-scale application which is representative of typical large-scale solar thermal plant designs (flat plate collectors, common hydraulic layout).</li> <li>Additional information is provided in a &quot;Data in Brief&quot; journal article: <a href="https://doi.org/10.1016/j.dib.2023.109224">https://doi.org/10.1016/j.dib.2023.109224</a></li> </ul> <p>&nbsp;</p> <p><strong>Collector array description: </strong>The data is from a flat&nbsp;plate collector array with a total gross collector area of 516&nbsp;m<sup>2</sup> (361&nbsp;kW nominal thermal power). The array consists of four parallel collector rows with a common inlet and outlet manifold. Large-area flat-plate collectors from Arcon-Sunmark A/S are used in the plant. Collectors are all oriented towards the south (180&deg;), have a tilt angle of 30&deg; and a row spacing of 3.1&nbsp;m. The collector array is part of a large-scale solar thermal plant located at Fernheizwerk Graz, Austria (latitude: 47.047294 N, longitude: 15.436366 E). The plant feeds into the local district heating network and is one of the largest Solar District Heating installations in Central Europe.</p> <p>&nbsp;</p> <p><strong>Data files:</strong></p> <ul> <li><strong>FHW_ArcS__main__2017.csv</strong> &ndash; This is the main dataset. It is advised to use this file for further analysis. The file contains the full time series of all measured and all calculated data channels and their (propagated) measurement uncertainty (53 data channels in total). Calculated data channels are derived from measured channels (see script make_data.py below) and have the suffix __calc in their channel names. Uncertainty information is given in terms of standard deviation of a normal distribution (suffix __std); some data channels are assumed to have no uncertainty (e.g., sun azimuth or shadowing).</li> <li><strong>FHW_ArcS__main__2017.parquet</strong> &ndash; Same as FHW_ArcS__main__2017.csv, but in parquet file format for smaller file size and improved performance when loading the dataset in software.</li> <li><strong>FHW_ArcS__parameters.json</strong> &ndash; Contains various metadata about the dataset, in both human and machine-readable format. Includes plant parameters, data channel descriptions, physical units, etc.</li> <li><strong>FHW_ArcS__raw__2017.csv </strong>&ndash; Dataset with time series of all measured data channels and their measurement uncertainty. The main dataset FHW_ArcS__main__2017.csv, which includes all calculated data channels, is a superset of this file.</li> </ul> <p>&nbsp;</p> <p><strong>Scripts: </strong></p> <ul> <li><strong>make_data.py</strong> &ndash; This Python script exposes the calculation process of the calculated data channels (suffix __calc), including error propagation. The main calculations are defined as functions in the module utils_data.py.</li> <li><strong>make_plots.py</strong> &ndash; This Python script, together with utils_plots.py, generates several figures based on the main dataset.</li> </ul> <p>&nbsp;</p> <p><strong>Data collection and preparation</strong>: AEE &mdash; Institute for Sustainable Technologies (AEE INTEC), Feldgasse 19, 8200 Gleisdorf, Austria; and SOLID Solar Energy Systems GmbH (SOLID), Am Pfangberg 117, 8045 Graz, Austria</p> <p>&nbsp;</p> <p><strong>Data owner</strong>: solar.nahwaerme.at Energiecontracting GmbH, Puchstrasse 85, 8020 Graz, Austria</p> <p>&nbsp;</p> <p><strong>Additional information</strong> is provided in a journal article in &quot;Data in Brief&quot;, titled <a href="https://doi.org/10.1016/j.dib.2023.109224">&quot;One year of high-precision operational data including measurement uncertainties from a large-scale solar thermal collector array with flat plate collectors in Graz, Austria&quot;</a>.</p> <p>&nbsp;</p> <p><strong>Note: </strong>A Gitlab repository is associated with this dataset, intended as a companion to facilitate maintenance of the Python code that is provided along with the data. If you want to use or contribute to the code, please do so using the Gitlab project: <a href="https://gitlab.com/sunpeek/zenodo-fhw-arconsouth-dataset-2017">https://gitlab.com/sunpeek/zenodo-fhw-arconsouth-dataset-2017</a></p> <p>&nbsp;</p>

opencc-by-4.0Mar 2023View details →
zenodo44/100

A non-intrusive reduced order model for the characterisation of the spatial power distribution in large thermal reactors (dataset)

<p>This repository contains the software and datasets needed to reproduce the results presented in the article &quot;<a href="https://doi.org/10.1016/j.anucene.2022.109674">A non-intrusive reduced order model for the characterisation of the spatial power distribution in large thermal reactors</a>&quot;, published in Annals of Nuclear Energy.</p>

opencc-by-4.0Dec 2022View details →
zenodo44/100

Dataset Worldwide Survey on the Impact of AI Chatbots and Large Language Models in Dental Education: Insights from Dental Educators

<p><strong>This dataset contains responses from participants regarding their awareness, knowledge, and perceptions of AI-powered tools in dental education. The data was collected during May-June 2023 to investigate the potential enhancement that AI can bring to dental education. The dataset includes variables related to participants&#39; demographics, experiences, perceptions, and opinions.</strong></p> <p><strong>Details in the published protocol by Uribe, S. E., &amp; Maldupa, I. (2023, June 2). Chatbots In Dental Education - Research Protocol. https://doi.org/10.17605/OSF.IO/3BSG2</strong></p>

opencc-by-4.0Jun 2023View details →
zenodo44/100

Dataset supporting the paper "Large Orbital Moment of Two Coupled Spin‑Half Co Ions in a Complex on Gold. ACS Nano 17, 10608 (2023)"

<p>Dataset corresponding to theoretical calculations in the paper &quot;Large Orbital Moment of Two Coupled Spin‑Half Co Ions in a Complex on Gold&quot; ACS Nano 17, 10608 (2023), https://pubs.acs.org/doi/10.1021/acsnano.3c01595</p> <p>List of files:</p> <p>Several folders corresponding to the figures of the paper. They contain:<br> .siesta files: STM images in WsXM format (http://www.wsxm.eu/) simulated using STMpw (https://doi.org/10.5281/zenodo.3581159).<br> CONTCAR files: relaxed structures in VASP format. They can be visualized with VESTA (https://jp-minerals.org/vesta/en/).<br> .agr: grace files (https://plasma-gate.weizmann.ac.il/Grace/).</p>

opencc-by-4.0Jun 2023View details →
zenodo44/100

NPM3D dataset with instance labels used in paper "Toward Accurate Instance Segmentation in Large-scale LiDAR Point Clouds"

<p>NPM3D (https://npm3d.fr/paris-carla-3d) consists of mobile laser scanning (MLS) point clouds collected in four different regions in the French cities of Paris and Lille, where each point has been annotated with two labels: one that assigns it to one out of 10 semantic categories and another one that assigns it to an object instance. When inspecting the data, we found 9 cases where multiple tree instances had not been separated correctly (i.e., they had the same ground truth instance label). These cases were manually corrected using the CloudCompare software (https://www.cloudcompare.org), and 35 individual tree instances were obtained. Our variant of the dataset with 10 semantic categories and enhanced instance labels is publicly available.</p>

opencc-by-4.0Jul 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record