Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
85
datasets available to search
ShareScore release 0.9.0
Dataset results
85 results for “standard dataset”
LMAS Test Dataset - BMock12 Community Standards
<p>The twelve bacterial replicons of the <a href="https://www.nature.com/articles/s41597-019-0287-z#Abs1">BMock12 Community Standards</a> were used as reference. It includes the following strains:</p> <table> <thead> <tr> <th>Species</th> <th>Sample ID</th> <th>Depth of Coverage (x)</th> </tr> </thead> <tbody> <tr> <td>Muricauda sp. ES.050</td> <td>2615840527</td> <td>618.76</td> </tr> <tr> <td>Thioclava sp. ES.032</td> <td>2615840533</td> <td>78.32</td> </tr> <tr> <td>Cohaesibacter sp. ES.047</td> <td>2615840601</td> <td>170.59</td> </tr> <tr> <td>Propionibacteriaceae bacterium</td> <td>2615840646</td> <td>31.90</td> </tr> <tr> <td>Marinobacter sp. LV10R510-8</td> <td>2615840697</td> <td>447.83</td> </tr> <tr> <td>Marinobacter sp. LV10MA510-1</td> <td>2616644829</td> <td>135.05</td> </tr> <tr> <td>Psychrobacter sp. LV10R520-6</td> <td>2617270709</td> <td>425.47</td> </tr> <tr> <td>Micromonospora echinaurantiaca</td> <td>2623620557</td> <td>14.91</td> </tr> <tr> <td>Micromonospora echinofusca</td> <td>2623620567</td> <td>18.19</td> </tr> <tr> <td>Micromonospora coxensis</td> <td>2623620609</td> <td>0.02</td> </tr> <tr> <td>Halomonas sp. HL-4</td> <td>2623620617</td> <td>507.08</td> </tr> <tr> <td>Halomonas sp. HL-93</td> <td>2623620618</td> <td>579.87</td> </tr> </tbody> </table> <p>The raw sequence data of the mock communities, with an even distribution of species, is available at:</p> <ul> <li><a href="https://www.ebi.ac.uk/ena/browser/view/SRR8073716?show=reads">SRR8073716</a></li> </ul> <p>For the sake of a less resource-intense evaluation we downsampled to have 20% of the reads available in the original sample and further processed only these. The main challenges of this data set are possibly to assemble the genomes of the three <em>Micromonospora</em> spp. and the two <em>Halomonas</em> spp. strains, which have an ANIb >0.84 and 0.98, respectively. The two <em>Marinobacter</em> spp. have an ANIb of 0.78.</p>
LMAS Test Dataset - ZymoBIOMICS Microbial Community Standards
<p>The <a href="https://zenodo.org/record/4588970#.YEeA83X7RhE">eight bacterial genomes and four plasmids of the ZymoBIOMICS Microbial Community Standards</a> were used as reference. It contains tripled complete sequences for the following species:</p> <ul> <li><em>Bacillus subtilis</em></li> <li><em>Enterococcus faecalis</em></li> <li><em>Escherichia coli</em> <ul> <li><em>Escherichia coli</em> plasmid</li> </ul> </li> <li><em>Lactobacillus fermentum</em></li> <li><em>Listeria monocytogenes</em></li> <li><em>Pseudomonas aeruginosa</em></li> <li><em>Salmonella enterica</em></li> <li><em>Staphylococcus aureus</em> <ul> <li><em>Staphylococcus aureus</em> plasmid 1</li> <li><em>Staphylococcus aureus</em> plasmid 2</li> <li><em>Staphylococcus aureus</em> plasmid 3</li> </ul> </li> </ul> <p>It also downloads the raw sequence data of the mock communities, with an even and logarithmic distribution of species:</p> <ul> <li><a href="https://www.ebi.ac.uk/ena/browser/view/ERR2984773">ERR2984773</a> - Evenly distributed real sample</li> <li><a href="https://www.ebi.ac.uk/ena/browser/view/ERR2935805">ERR2935805</a> - Log distributed real sample</li> </ul> <p>A set of simulated samples were generated from the genomes in the ZymoBIOMICS standard though the <a href="https://github.com/HadrienG/InSilicoSeq">InSilicoSeq sequence simulator</a> (version 1.5.2), including both even and logarithmic distribution, with and without Illumina error model. The number of read pairs generated matches the number of read pairs in the real data for each distribution. The following samples are available in Zenodo:</p> <ul> <li><strong>ENN</strong> - Evenly distributed sample with no error model</li> <li><strong>EMS</strong> - Evenly distributed sample with Illumina MiSeq error model</li> <li><strong>LNN</strong> - Log distributed sample with no error model</li> <li><strong>LHS</strong> - Log distributed sample with Illumina HiSeq error model</li> </ul> <p><a href="https://doi.org/10.5281/zenodo.4588970">DOI Dataset</a></p>
Dataset for use with "Technology standards vs. carbon taxes: comparing emissions reductions and system costs for US decarbonization"
<p>This sqlite file contains the input data and results used in "Technology standards vs. carbon taxes: comparing emissions reductions and system costs for US decarbonization." The file is a version of the nine region Open Energy Outlook database for use with the Temoa model. The database is continually updated. Updated versions can be found in the Temoa GitHub repository, https://github.com/TemoaProject/oeo.</p>
Molecular surface coverage standards by reference-free GIXRF supporting SERS and SEIRA substrate benchmarking - Dataset
<p>This is the dataset of "Molecular surface coverage standards by reference-free GIXRF supporting SERS and SEIRA substrate benchmarking". </p> <p><a href="https://doi.org/10.1515/nanoph-2024-0222" target="_blank" rel="noopener">https://doi.org/10.1515/nanoph-2024-0222</a></p> <p>Part of this work was supported by the European project OpMetBat, code 21GRD01. The project has received funding from the European Partnership on Metrology, cofinanced from the the European Union's Horizon Europe Research and Innovation Programme, and by Participating States.</p>
NLM-Gene, a richly annotated gold standard dataset for gene entities that addresses ambiguity and multi-species gene recognition
<p>The automatic recognition of gene names and their corresponding database identifiers in biomedical text is an important first step for many downstream text-mining applications. The NLM-Gene corpus is a high-quality manually annotated corpus for genes, covering ambiguous gene names, with an average of 29 gene mentions (10 unique identifiers) per article, and a broader representation of different species (including <i>Homo sapiens, Mus musculus, Rattus norvegicus, Drosophila melanogaster, Arabidopsis thaliana, Danio rerio,</i> etc.) when compared to previous gene annotation corpora. NLM-Gene consists of 550 PubMed articles from 156 biomedical journals, doubly annotated by six experienced NLM indexers, randomly paired for each article to control for bias. The annotators worked in three annotation rounds until they reached a complete agreement. Using the new resource, we developed a new gene finding algorithm based on deep learning which improved both on precision and recall from existing tools. The NLM-Gene annotated corpus is freely available at Dryad and at <a href="https://www.ncbi.nlm.nih.gov/research/bionlp/">https://www.ncbi.nlm.nih.gov/research/bionlp/</a>. The gene finding results of applying this tool to the entire PubMed/PMC are freely accessible through our web-based tool PubTator.</p>
Supplementary Dataset for "Synthetic Witherite for Standardization of Clumped Isotope (Delta 47) Analyses"
<p>This supplementary dataset contains raw clumped isotopic data used to generate figures in the article "Synthetic Witherite for Standardization of Clumped Isotope (Delta 47) Analyses" by Kong et al. </p>
A multi-sensor gait dataset collected under non-standardized dual-task conditions
Open the record for dataset details and reuse information.
Standardized incidence ratio dataset of human West Nile Virus in Italy (2012-2024)
Open the record for dataset details and reuse information.
NLM-Gene, a richly annotated gold standard dataset for gene entities that addresses ambiguity and multi-species gene recognition
Open the record for dataset details and reuse information.
MESINESP: Post-workshop datasets. Silver Standard and annotator records
<p><em><strong>Please use the <a href="https://doi.org/10.5281/zenodo.4612274">MESINESP2 corpus (the second edition of the shared-task)</a> since it has a higher level of curation, quality and is organized by document type (scientific articles, patents and clinical trials).</strong></em></p> <p>The MESINESP (Spanish BioASQ track, see https://temu.bsc.es/mesinesp) Challenge was held in May-June 2020, and as a result of a strong participation and the manual annotation of an evaluation dataset, two additional datasets are released now:</p> <p>1) "all_annotations_withIDsv3.tsv" contains a tab-separated file with all manual annotations (both validated and non-validated) of the evaluation dataset prepared for the competition. It contains the following fields:</p> <ul> <li>annotatorName: Human annotator id</li> <li>documentId: Document ID in the source database</li> <li> decsCode: A DeCS code added to it or validated</li> <li>timestamp: When it was added</li> <li>validated: if it was validated at that point by another annotator, or not yet</li> <li>SpanishTerm: The Spanish descriptor corresponding to the DeCS code</li> <li>mesinespId: The internal document id in the distributed evaluation file</li> <li>dataset: if part of the evaluation or the test sets</li> <li>source: which database it was taken from</li> </ul> <p><em><strong>Example:</strong></em></p> <p><strong>annotatorName documentId decsCode timestamp validated SpanishTerm mesinespId dataset source</strong><br> A7 biblio-1001069 6893 2020-01-17T11:27:07.000Z false caballos mesinesp-dev-671 dev LILACS<br> A7 biblio-1001069 4345 2020-01-17T11:27:12.000Z false perros mesinesp-dev-671 dev LILACS</p> <p> </p> <p>2) A "Silver Standard" created from the 24 system runs submitted by 6 participating teams. It contains each of the submitted DeCS code for each document in the test set, as well as other information that can help ascertain reliability and source for anyone that wants to use this dataset to enrich their training data. It contains more that 5.8 million datapoints, and is structured as follows</p> <ul> <li>SubmissionName: Alias of the team that submitted the run</li> <li>REALdocumentId: The real id of the document</li> <li>mesinespId: The mesinesp assigned id in the evaluation dataset</li> <li>docSource: The source database</li> <li>decsCode: the DeCS code assigned to it by the team's system</li> <li>SpanishTerm: The Spanish descriptor of the DeCS code</li> <li>MiF: The Micro-f1 scored by that system's run</li> <li>MiR: The Micro-Recall scored by that system's run</li> <li>MiP: The Micro-Precision scored by that system's run </li> <li>Acc: The Accuracy scored by that system's run</li> <li>consensus: The number of runs where that DeCS code was assigned to this document by the participating teams (max. is 24)</li> </ul> <p><strong><em>Example:</em></strong></p> <p><strong>SubmissionName REALdocumentId mesinespId docSource decsCode SpanishTerm MiF MiR MiP Acc consensus</strong><br> AN ibc-177565 mesinesp-evaluation-00001 IBECS 28567 riesgo 0.2054 0.1930 0.2196 0.1198 4<br> AN ibc-177565 mesinesp-evaluation-00001 IBECS 15335 trabajo 0.2054 0.1930 0.2196 0.1198 4<br> AN ibc-177565 mesinesp-evaluation-00001 IBECS 33182 conocimiento 0.2054 0.1930 0.2196 0.1198 7<br> </p> <p><strong>For citation and a detailed description of the Challenge, please cite:</strong><br> <em>Anastasios, Nentidis and Anastasia, Krithara and Konstantinos, Bougiatiotis and Martin, Krallinger and Carlos, Rodriguez-Penagos and Marta, Villegas and Georgios, Paliouras.</em> <strong>Overview of BioASQ 2020: The eighth BioASQ challenge on Large-Scale Biomedical Semantic Indexing and Question Answering</strong> (2020). Proceedings of the Eleventh International Conference of the CLEF Association (CLEF 2020). Thessaloniki, Greece, September 22--25</p> <p><strong>Citation</strong></p> <p>@inproceedings{durusan2019overview,<br> title={Overview of BioASQ 2020: The eighth BioASQ challenge on Large-Scale Biomedical Semantic Indexing and Question Answering},<br> author={Anastasios, Nentidis and Anastasia, Krithara and Konstantinos, Bougiatiotis and Martin, Krallinger and Carlos, Rodriguez-Penagos and Marta, Villegas and Georgios, Paliouras},<br> booktitle={Experimental IR Meets Multilinguality, Multimodality, and Interaction Proceedings of the Eleventh International Conference of the CLEF Association (CLEF 2020), Thessaloniki, Greece, September 22--25, 2020, Proceedings},<br> volume={12260},<br> year={2020},<br> organization={Springer}<br> }</p> <p> </p> <p>Copyright (c) 2020 Secretaría de Estado de Digitalización e Inteligencia Artificial</p>
Gold Standard and Annotation Dataset for CO2 Emissions Annotation
<p>This repository contains the results of a research project which provides a benchmark dataset for extracting greenhouse gas emissions from corporate annual and sustainability reports. </p> <p>The zipped <code>datasets</code> file contains two datasets, <code>gold_standard</code> and <code>annotation_dataset</code>(inside the outer zip file there is a password-protected zip file containing the two datasets. To unpack, use the password is provided in the outer zip file).</p> <h3>Data collection</h3> <ol> <li>A Large Language Model (LLM) based pipeline was used to extract the greenhouse gas emissions from the reports (see columns prefixed with <code>llm_</code> in <code>annotation_dataset</code>). The extracted emissions follow the categories Scope 1, 2 (market-based) and 2 (location-based) and 3, as defined in the GHGP protocol (see variables <code>scope</code>).</li> <li>Annotation of the pipeline output was done in 3 phases: first by non-experts (see columns prefixed with <code>non_expert_</code> in <code>annotation_dataset</code>), then by expert groups (columns prefixed with <code>exp_group_</code> in <code>annotation_dataset</code>) in case of disagreement of non-experts and finally in a discussion of all experts (columns prefixed with <code>exp__disc</code> in <code>annotation_dataset</code>) in case of disagreement between expert groups. The annotation guidelines for the <a href="https://zenodo.org/api/records/15124118/draft/files/Non-Expert%20Annotation%20Guidelines.pdf/content">non-experts</a> and <a href="https://zenodo.org/api/records/15124118/draft/files/Expert%20Annotation%20Guidelines.pdf/content">experts</a> are also included in this repository.</li> <li>The annotation results from all three phases are combined to form the final benchmark dataset: <code>gold_standard</code>. Codebooks detailing each variable of each of the two datasets are also provided. More details about the annotation template or the data wrangling scripts can be found in the <a href="https://github.com/soda-lmu/gist-data-descriptor/" target="_blank" rel="noopener">GitHub repository</a>. </li> </ol> <h3>Merging of datasets</h3> <p>Users can match the two datasets (<code>gold_standard</code> and <code>annotation_dataset</code>) using the variable combination of <code>company_name</code>, <code>report_year</code> and <code>merge_id</code> (index column). The <code>merge_id</code> already includes the company name and report year implicitly, but to avoid column duplication in the join operation, it should be included as join variables. For example this is useful when comparing LLM extractions to gold standard data.</p>
LegacyPollen 1.0: A taxonomically harmonized global Late Quaternary pollen dataset of 2831 records with standardized chronologies
<p>This repository consists of code for downloading pollen data from the Neotoma Paleoecology Database, the harmonization of pollen taxa, and the assignment of age-depth data so that datasets for customized harmonization levels can be easily established. The input data includes a harmonization table and example data, stored in machine-readable data format (.CSV).</p>
"Play by Play": a dataset of handball and basketball game situations in a standardized space
<p>Synthetic dataset of labeled game situations present in recordings of federated handball and basketball matches played in Galicia, Spain. The dataset consists of synthetic data generated from real video frames, including 308,805 labeled handball frames and 56,578 labeled basketball frames extracted from 2,105 handball and 383 basketball 5-second video clips.</p> <p>For each frame, the dataset includes player positions and directions (x, y, vx, vy), and ball position (referees are also considered as players, we do not distinguish teams or players from referees). For each tuple there is an associated game situation: right/left attack, right/left counterattack, right/left penalty and time out.</p>
IVD-SEG:Standardized Datasets for Industrial Vision Defect Segmentation
<h1> </h1> <h1>Code[<a href="https://github.com/KLIVIS/IVD-SEG">Github</a>]</h1> <h1>abstract</h1> <p>We introduce IVD-SEG, a dataset encompassing defect images from 43 different industrial products, totaling 5686 images, spanning tasks that include binary and multiclass segmentation. Within IVD-SEG, we meticulously propose 12 sub-datasets, including two newly developed datasets by our team. Serving as a large-scale standardized industrial defect image dataset, all images are unified to a 256 × 256 size, accompanied by semantic segmentation annotations. This standardization facilitates users unfamiliar with industrial product defects to utilize the dataset, allowing them to focus on exploring algorithmic performance on the IVD-SEG dataset. To our knowledge, the dataset we propose is currently the most comprehensive and voluminous compilation of industrial defect images, thereby contributing to the advancement of relevant research in industrial image analysis. Simultaneously, our dataset supports research and education in various fields, including computer vision and machine learning. We conducted benchmark tests on IVD-SEG using several baseline methods, including representative CNN and ViT networks.<br><br></p>
Dataset: Gold standard dataset for explainability need detection in app reviews.
<p>We crawled 90,000 app reviews from both Google Play Store and Apple App Store, including reviews from both free and paid apps. These reviews were filtered for explainability needs, and after this process, 4,495 reviews remained. Among them, 2,185 reviews indicated an explanation need, while 2,310 did not. This resulting gold standard dataset was used to train and evaluate several machine learning models and rule-based approaches for detecting explanation needs in app reviews.</p> <p>The dataset includes both balanced and unbalanced evaluation sets, as well as the original crawled data from October 2023. In addition to machine learning approaches, rule-based methods optimized for F1 score, precision, and recall are also included.</p> <p>We provide several pre-trained machine learning models (including BERT, SetFit, AdaBoost, K-Nearest Neighbor, Logistic Regression, Naive Bayes, Random Forest, and SVM) along with training scripts and evaluation notebooks. These models can be applied directly or retrained using the included datasets.</p> <p>For further details on the structure and usage of the dataset, please refer to the README.md file within the provided ZIP archive.</p>
Dataset for Robot-mounted digital 3D-microscope versus standard microscope for spinal decompression surgery
<p>Compilation of all data collected in the study.</p> <p>- Depth Perception Robotc: Data collected during depth perception experiments with the BHS RoboticScope</p> <p>- Depth Perception Cons: Data collected during depth perception experiments with the OM</p> <p>- Both Scopes Combined: Combined data regarding depth perception experiments</p> <p>- Questionaire Combined: Combined data regarding cadaveric experiments</p> <p>- Questionaire expsurg1: completed questionnaire by the experienced surgeon 1</p> <p>- Questionaire expsurg3: completed questionnaire by the experienced surgeon 3</p>
Dataset related to article "Standardized Diagnostic Reference Levels for Paediatric Interventional Cardiology: data from an Italian referral centre"
Open the record for dataset details and reuse information.
The Beyond the Standard Model dataset
<p>The record contains LHC simulations of several Beyond the Standard Model theories.</p>
LAGOS-AND-PM: A Large Gold Standard Dataset for PubMed Author Name Disambiguation
<p>LAGOS-AND-PM is an author name disambiguation (AND) dataset for the PubMed database, containing several versions, and they are built based on the ORCID database and the PubMed literature database.</p> <p>Note that we have previously created another dataset named <a href="https://zenodo.org/record/7313380">LAGOS-AND</a>, which refers to a series of AND datasets created for the MAG/OpenAlex database.</p>
RSS SMAP Level 3 Sea Surface Salinity Standard Mapped Image Monthly V6.0 Validated Dataset
The RSS SMAP Level 3 Sea Surface Salinity Standard Mapped Image Monthly V6.0 Validated Dataset produced by the Remote Sensing Systems (RSS) and sponsored by the NASA Ocean Salinity Science Team, is a validated product that provides orbital/swath data on sea surface salinity (SSS) derived from the NASA's Soil Moisture Active Passive (SMAP) mission. The SMAP satellite was launched on 31 January 2015 with a near-polar orbit at an inclination of 98 degrees and an altitude of 685 km. It has an ascending node time of 6 pm and is sun-synchronous. With its 1000km swath, SMAP achieves global coverage in approximately 3 days, but has an exact orbit repeat cycle of 8 days. Malfunction of the SMAP scatterometer on 7 July, 2015, has necessitated the use of collocated wind speed, primarily from WindSat, for the surface roughness correction required for the surface salinity retrieval.<br><br>The major changes in Version 6.0 from Version 5.0 are: (1) Removal of biases during the first few months of the SMAP mission that are related to the operation of the SMAP radar during that time. (2) Mitigation of biases that depend on the SMAP look angle. (3) Mitigation of salty biases at high Northern latitudes. (4) Revised sun-glint flag. The RSS SMAP L3 monthly product includes data for a range of parameters: derived sea surface salinity (SSS) with SSS-uncertainty, rain filtered SMAP sea surface salinity, collocated wind speed, data and ancillary reference surface salinity data from HYCOM. Each data file is available in netCDF-4 file format and is averaged over one-month time intervals with about 7-day latency (after the end of the averaging period). Data begins on April 1,2015 and is ongoing. Observations are global in extent with an approximate spatial resolution of 40KM. Note that while a SSS 40KM variable is also included in the product for most open ocean applications, The standard product of the SMAP Version 6.0 release is the smoothed salinity product with a spatial resolution of approximately 70 km.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.