Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
240
datasets available to search
ShareScore release 0.7.1
Dataset results
240 results for “research evaluation”
Expert judgements for evaluating deduplication of OpenAIRE Research Graph
<p>Expert judgments used to evaluate the deduplication algorithm used for constructing the OpenAIRE Research Graph.</p> <p>Each expert assigned each group with one of the following predetermined classes (that also indicate whether contained entities are equivalent or not):</p> <ul> <li>AMBIGUOUS: At least one DOIs is invalid (no metadata are available) (N/A)</li> <li>DELETED-DUPLICATES: DOIs once pointing to the same research object, currently deleted. (TRUE)</li> <li>MULTI-PUBLISHED: Article published in more than one locations (full or abstract) (TRUE)</li> <li>VERSIONS: Multiple versions of the same research object (e.g. pre-prints, post-prints etc). (TRUE)</li> <li>ERRONEOUS: Unrelated set of objects. (FALSE)</li> <li>PAPER-EXTENSIONS: Extended version of a conference paper in a journal. (FALSE)</li> <li>PART-OF-A-GROUP: Multiple parts of the same research object (e.g. multi-part publication, photos of the same collection etc). (FALSE)</li> <li>SUPPLEMENTARY: A publication and its supplementary material (including errata). (FALSE)</li> </ul> <p>The following are provided:</p> <ul> <li>OpenAIRE identifier</li> <li>Judgement</li> <li>Count of numbers in group</li> <li>DOIs in group</li> </ul>
Evaluation Set - Contributions Similarity in the Open Research Knowledge Graph
<p>This evaluation set has been created for evaluating a content-based recommender system in the context of the Open Research Knowledge Graph (ORKG). The recommender system accepts structured ORKG contribution as input and recommends existing contributions in the ORKG semantically relevant to the given one.</p> <p> </p> <p>The evaluation set is manually annotated based on the <a href="https://www.orkg.org/orkg/featured-comparisons">featured comparisons</a> in the ORKG. In the course of this, it has been distinguished between homogeneous (those who are dissimilar in 2-3 properties) and heterogeneous (otherwise) instances. Multiple annotations have been obtained for the former and exactly one for the latter.</p> <p> </p> <p>It has been also distinguished between "with_response" and "without_response" instances (50 instances for each). The former are those contributions for them the initial version of the contributions similarity service has found similarities and the latter are the opposite case.</p> <p> </p> <p>This evaluation set has been created and applied on a modified version of the contributions similarity service in the context of <a href="https://doi.org/10.15488/11834">this master's thesis</a>. The modified version of the service has simplified the document representation of contributions that are stored in an ElasticSearch index by omitting redundant terms.</p> <p>The evaluation set has the following schema:</p> <pre><code class="language-json">{ "with_response": [ { "contribution_id": "some_id", "comparison_id": "some_id", "comparison_label": "some_label", "contribution_label": "some_label", "paper": "some_id", "research_field": "some_id", "research_problems": [ "some_id" ], "annotations": [ "some_id of a similar contribution", ... ] }, ... ], "without_response": [ ... ] }</code></pre> <p> </p>
Data and Statistical analysis for: "Predator in the pool? A quantitative evaluation of non-indexed open access journals in aquaculture research"
<p>Data and Statistical analysis for: "Predator in the pool? A quantitative evaluation of non-indexed open access journals in aquaculture research" published in <em>Frontiers in Marine Science</em></p>
Evaluating Open Science Practices in Indoor Positioning and Indoor Navigation Research (Supplementary Material: Full Paper Listing and Analysis)
<p>Supplementary material of the paper:</p> <p>Title: "Evaluating Open Science Practices in Indoor Positioning and Indoor Navigation Research"<br>Subtitle: "A Survey of the IPIN's Reference Papers of 2022 and 2023 Editions"</p> <p>The paper is accepted to the "14th International Conference on Indoor Positioning and Indoor Navigation, IPIN 2024, Hong Kong, October 14-17, 2024, IEEE, 2024.</p> <p>An Author's accepted version of the manuscript is available here: <a href="../records/13684170" target="_blank" rel="noopener">https://zenodo.org/records/13684170</a> </p> <p>If you want to refer to this work, please cite this Zenodo entry as well as the published conference version.</p> <p> </p> <p>---------------------------------------</p> <p>This entry contains two files:</p> <ul> <li>"Paper Characterization Spreadsheet.xlsx": <strong>The spreadsheet of the full analysis of this work</strong>, as described in the paper. It characterizes various features of the analyzed papers and forms the raw data on which the analyses of our work were based.</li> <li>"Main features of the manuscripts analysed in Zenodo Record #12088175.pdf": A document summarizing the main features of the IPIN's Reference Papers of the 2022 and 2023 Editions, that contain some form of open resources (Open Data, Code, or Material).</li> </ul> <p> </p> <p> </p> <p> </p>
Dataset: Comparative evaluation of a keyword based search and semantic search in a data portal for biodiversity research.
<p>Supplementary material for a comparative evaluation of a keyword based search and semantic search in a data portal for biodiversity research. We conducted a relevance evaluation with 6 users over 19 search questions in two search interfaces.</p> <p>The users provided up to five search questions and relevant keywords from their research background. We setup a dataset search over a corpus of ~92,000 randomly selected metadata files from GFBio (<a href="https://www.gfbio.org">https://www.gfbio.org</a>). For each of their own search queries, the users got two result sets presented. The first one displayed results obtained from a keyword search. The second panel contained dataset results from a prototypical semantic search. Instead of results with exact mentions of the query terms, the semantic search also presented related results with synonyms and more specific terms or terms obtained from concept nodes of a higher hierarchy level.</p> <p>Each user rated the relevance of his/her own search queries on a 7-point Likert scale for both search results.<br> In addition, users also assessed the expanded keywords for each question.</p> <p>More information can be found in our publication:</p> <p>Löffler, F. and Klan, F. (2016): Does Term Expansion Matter for the Retrieval of Biodiversity Data? in Joint Proceedings of the Posters and Demos Track of the 12th International Conference on Semantic Systems - SEMANTiCS2016 and the 1st International Workshop on Semantic Change & Evolving Semantics (SuCCESS'16), co-located with the 12th International Conference on Semantic Systems (SEMANTiCS 2016),2016, <a href="http://ceur-ws.org/Vol-1695/paper2.pdf">http://ceur-ws.org/Vol-1695/paper2.pdf</a></p> <p> </p>
Extensive crowdsourced dataset of in-situ evaluated binaural soundscapes of private dwellings containing subjective sound-related and situational ratings along with person factors to study time-varying influences on sound perception — research data
<p><strong>Abstract:</strong></p> <p>The soundscape approach highlights the role of situational factors in sound evaluations; however, only a few studies have applied a multi‐domain approach including sound‐related, person‐related, and time‐varying situational variables. Therefore, we conducted a study based on the Experience Sampling Method to measure the relative contribution of a broad range of potentially relevant acoustic and non‐auditory variables in predicting indoor soundscape evaluations. Here we present the comprehensive dataset for which 105 participants reported temporally (rather) stable trait variables such as noise sensitivity, trait affect, and quality of life. They rated 6.594 situations regarding the soundscape standard dimensions, perceived loudness, and the saliency of its sound components and evaluated situational variables such as state affect, perceived control, activity, and location. To complement these subject‐centered data, we additionally crowdsourced object‐centered data by having participants make binaural measurements of each indoor soundscape at their homes using a low‐(self‐)noise recorder. These recordings were used to compute (psycho‐)acoustical indices such as the energetically averaged loudness level, the A‐weighted energetically averaged equivalent continuous sound pressure level, and the A‐weighted five‐percent exceedance level. This complex hierarchical data can be used to investigate time‐varying non‐auditory influences on sound perception and to develop soundscape indicators based on the binaural recordings to predict soundscape evaluations.</p> <p><strong>Content:</strong></p> <ul> <li><a href="https://zenodo.org/record/7858848/files/01%20StudyDescription.pdf">01 StudyDescription.pdf </a> <ul> <li>Description of the field study.</li> <li>Information about the methods and materials used.</li> </ul> </li> <li><a href="https://zenodo.org/record/7858848/files/02%20Dataset.csv">02 Dataset.csv</a> <ul> <li>The dataset, consisting of 93 variables describing 6594 observations taken by 105 participants.</li> </ul> </li> <li><a href="https://zenodo.org/record/7858848/files/03%20VariableDescriptions_EnglishPersonQuestionnaire.pdf">03 VariableDescriptions_EnglishPersonQuestionnaire.pdf</a> <ul> <li>Descriptions of all variables, their measurement scale, scale ranges and levels.</li> <li>Questions and task descriptions of the Experience Sampling Method questionnaire in German language with an English translation.</li> <li>English translations of questions asked in the person questionnaire.</li> </ul> </li> <li><a href="https://zenodo.org/record/7858848/files/04%20ESM-Questionnaire.pdf">04 ESM-Questionnaire.pdf</a> <ul> <li>Screenshots of the original Experience Sampling Method questionnaire with English translations.</li> </ul> </li> <li><a href="https://zenodo.org/record/7858848/files/05%20PersonQuestionnaire_OriginalGermanVersion.pdf">05 PersonQuestionnaire_OriginalGermanVersion.pdf</a> <ul> <li>Original version of the person questionnaire in German language.</li> </ul> </li> <li><a href="https://zenodo.org/record/7858848/files/06%20HelpTexts.pdf">06 HelpTexts.pdf</a> <ul> <li>Descriptions of the study task.</li> <li>Explanations of the scales used in the questionnaire.</li> <li>Explanations of the sound categories and the soundscape composition.</li> <li>Explanation of the operation of the recording device.</li> </ul> </li> <li><a href="https://zenodo.org/record/7858848/files/AcousticFeatures_README.md">AcousticFeatures_README.md</a> <a href="https://zenodo.org/api/files/3d784540-c0f4-412f-8742-df1db6f5401d/TimeSeries_and_Spectrograms_README.md?versionId=9291496c-d2c6-4151-96f1-a2ad99e1a540"> </a> <ul> <li>Descriptions of the structure of the AcousticFeatures_xxx.csv and .zip files.</li> <li>Analyis settings used in Artemis Suite to generate the acoustic features.</li> </ul> </li> <li><a href="https://zenodo.org/record/7858848/files/AcousticFeatures_SingleValues.csv">AcousticFeatures_SingleValues.csv</a> <ul> <li>All acoustic features, aggregated to single values per feature, recording, and channel.</li> </ul> </li> <li><a href="https://zenodo.org/record/7858848/files/AcousticFeatures_Spectra.csv">AcousticFeatures_Spectra.csv</a> <ul> <li>Time-averaged 1/3 octave spectra of each channel of each recording, A-weichted and un-weighted.</li> </ul> </li> <li><a href="https://zenodo.org/record/7858848/files/AcousticFeatures_Spectrograms.zip">AcousticFeatures_Spectrograms.zip</a> <ul> <li>13188 .csv files with un-weighted spetrograms of each channel of each recording.</li> </ul> </li> <li><a href="https://zenodo.org/record/7858848/files/AcousticFeatures_TimeSeries.zip">AcousticFeatures_TimeSeries.zip</a> <ul> <li>A .csv file containing LAeq and LZeq time series of each channel of each recording.</li> </ul> </li> </ul> <p><strong>Publications refering to this dataset:</strong></p> <p>Versümer, Siegbert; Steffens, Jochen; Weinzierl, Stefan (currently under review): "The role of loudness predictions, personal and situational factors in day-to-day loudness assessments of indoor soundscapes."</p> <p><strong>Funding:</strong></p> <p>This study was sponsored by the German Federal Ministry of Education and Research. “FHprofUnt” funding code: 13FH729IX6. </p> <p><strong>License: </strong></p> <p>CC 4.0 BY, <a href="https://creativecommons.org/licenses/by/4.0/legalcode">https://creativecommons.org/licenses/by/4.0/legalcode</a></p> <p><strong>Version history:</strong></p> <p>Details can be found in the <a href="https://zenodo.org/api/files/a15d6a91-1a35-4b5e-a7ec-da8a9bcbee2b/Changelog.md">Changelog.md</a> file.</p> <ul> <li> V.01.0. March 7, 2023: Initial publication. <a href="https://doi.org/10.5281/zenodo.7193938">https://doi.org/10.5281/zenodo.7193938</a></li> <li> V.01.1. April 25, 2023. <a href="https://doi.org/10.5281/zenodo.7858848">https://doi.org/10.5281/zenodo.7858848</a></li> </ul>
Research Data for Comparative Evaluation of RT-PCR and Antigen-based Rapid Diagnostic Tests (Ag-RDTs) for SARS-CoV-2 Detection: Performance, Variant Specificity, and Clinical Implications
<p>This dataset represents laboratory findings for the comparative evaluation of the diagnostic performance of Ag-RDTs (Flourescence Immunoassay and Lateral Flow Immunoassay) with RT-PCR</p>
Research Management Systems: Systematic Mapping of Literature (2007-2017) - Number of articles included during the search and qualitative evaluation process of the study
<p>This image is uploaded as an integrated part of systematic mapping of literature "Research Management Systems: Systematic Mapping of Literature (2007-2017)". This image will be cited across all future publications related to this project as Attribution-NonCommercial-NoDerivatives 4.0 International image.</p>
Reproducibility Package for "Reproducible research and GIScience: an evaluation using AGILE conference papers"
<p>Data and code for analysis and plots used in the manuscript "Reproducible research and GIScience: an evaluation using AGILE conference papers": <a href="https://doi.org/10.7287/peerj.preprints.26561v1">https://doi.org/10.7287/peerj.preprints.26561v1</a></p> <p>The deposited archived includes a <a href="https://en.wikipedia.org/wiki/Docker_(software)">Dockerfile</a> and an <a href="http://rmarkdown.rstudio.com/">R Markdown</a> document suitable for use with <a href="http://mybinder.org/">Binder</a>: <a href="https://mybinder.org/v2/gh/nuest/reproducible-research-and-giscience/6">https://mybinder.org/v2/gh/nuest/reproducible-research-and-giscience/6</a></p> <p>The version tag of this repository matches the <a href="https://git-scm.com/book/en/v2/Git-Basics-Tagging">git tag</a> on the code repository at <a href="https://github.com/nuest/reproducible-research-and-giscience">https://github.com/nuest/reproducible-research-and-giscience</a>, except version <code>6-fixed</code> which matches the tag <code>6</code>.</p> <p> </p>
Evaluating an instrument of the research software related to software use and disclosure - Dimension 1 - Dataset of Focus Groups
<p>Artifacts used for data collection and analysis of the focus groups sessions during the evaluation of an instrument for research software related to software use and disclosure - dimension 1.</p>
Research project on field data collection for honey bee colony model evaluation - datasets
<p><strong>Description of the datasets</strong></p> <p>The file 00_MUSTB_field_data_model.docx contains the data model according to which the data collected in the context of the MUSTB field data collection were reported to EFSA. The current data model description includes some modifications with respect to the specifications published before the beginning of the project (EFSA, 2017, https://doi.org/10.2903/sp.efsa.2017.EN-1234). All the tables included in the data model are published here in csv format. The underlying schemas are also published in xsd format.</p> <p>Sites: General information about the sites where the data collection took place;</p> <p>Polygons: General information about the polygons where the botanical survey took place.</p> <p>Table I: Pesticide application, reporting data on experimental spraying events;</p> <p>Table II: Resource providing unit and landscape fitness, reporting data on abundance of flowering plants in polygons mostly within 1.5 km, but in some cases up to 3 km of the experimental colony;</p> <p>Table III: Master list of all hives included in the study;</p> <p>Table IV: Colony management, reporting the log of the beekeeper regarding input (if material was added to the hive: e.g. empty frames, chemicals for varroa treatment, sugar), output (if material was removed from the hive, e.g. honey combs, supers), queen loss, swarming, or clinical signs observed in the experimental hives;</p> <p>Table V: Hive inspection, reporting data on in-hive measurements in the experimental colonies. This table contained several types of data, including:</p> <ul> <li>Data on brood development and food provision (“cell utilization”) obtained from image analysis of combs;</li> <li>Data on forager activity obtained from automatic video recordings and image analysis by a bee counter;</li> <li>Data on hive weight obtained from automatic logging by a hive scale;</li> <li>Data on adult bee strength, obtained by weight assessment of combs with and without adult bees (“bees per comb data”);</li> </ul> <p>Table VI: SSD2, reporting data on results of laboratory analyses of pollen, pesticide residues and parasites/pathogens. These four types of laboratory analyses involved different methods, and were reported according to different standards. Therefore, a number of the fields in the technical specifications for the SSD2 table (EFSA, 2017) were not applicable for records reporting results of some analyses, in particular palynological, parasite and pathogen analyses. These fields were left empty;</p> <p>Table VII: Colony observation, reporting observations of honey bee waggle dances from observation hives. Orientation denotes the angle of the waggling phase relative to the vertical axis on the comb. Direction denotes the actual direction in the landscape, as calculated from the orientation of the waggle dance.</p> <p>In all the csv files, columns with the suffix "_desc" have been included, where relevant, to include the name corresponding to the EFSA controlled terminology used in the previous column (e.g. resUnit contains EFSA term codes while resUnit_desc contains the term names).</p> <p><strong>Data storage</strong></p> <p>All data collected during the project was stored in a relational database. The database was developed in .NET Entity Framework Core, ran on a PostgreSQL, and was hosted by Amazon Web Service during the whole duration of the project development. Data could be imported or entered manually in the database through a web form. Administrators could create new users and administrators, new sites, and new colonies, i.e., administrators were allowed to enter or change data of all tables. Users were allowed to enter data, and could view, retrieve, and modify their own data of all tables, except for Table III (description of experimental colonies). Administrators could view and retrieve all data. Data was retrieved in CSV and XML formats, and were structured to secure a smooth transmission of data to the Data Collection Framework of EFSA. Furthermore, data flow from the field data collection to the development of ApisRAM was secured by direct communication between the field and modelling teams.</p> <p> </p> <p><strong>Version 2</strong> contains the UTM coordinates in tables Sites, Polygons and Resource providing unit.</p>
Dataset: Relevance and usability evaluation in a data portal for biodiversity research
<p>Supplementary material for a relevance and usability evaluation in a data portal for biodiversity research.</p> <p>Data portal: GFBio (<a href="https://www.gfbio.org">https://www.gfbio.org</a>)</p> <p>Evaluation time:<span> February 2016 (at that time the search index consisted of ~ 2 Mio datasets)</span></p> <p>Eight domain experts rated the Top25 search results of 16 provided search questions on a 7-point-Likert scale from 0 (irrelevant) to 6 (highly relevant). Afterwards, we asked the users to provide and rate up to two own queries.</p> <p><span>The users also rated 28 statements in a subsequent usability evaluation on a 5-point Likert scale from 'completely disagree' to 'highly agree'. For some statements, only binary ratings were given.</span></p>
Data files for manuscript "Re-evaluation and Re-analysis of 152 research exomes five years after the initial report reveals clinically relevant changes in 18%"
<p>#2023-06-16<br> #Summary<br> This ZIP-file contains the data files used for all analyses for the manuscript "Re-evaluation and Re-analysis of 152 research exomes five years after the initial report reveals clinically relevant changes in 18%".</p> <p><br> #File structure<br> README.txt This README file.<br> File S02 ("FileS2_conNDD-cohort.xlsx") All variants identified by Reuter et al. previously with reevaluated variants and addition variants identified in this <br> project togetehr with information about the families, individuals, samplesand the BAM files assessed in this project.<br> File S03 ("FileS3_conNDD-variants.xlsx") All variant data analyzed from the cohort. Including a sheet with thresholdes for in silico predictions tools used to predict effect of variants, <br> a table with exome wide homozygous variants in 4 categories (A45, LGD, Missense, Splice), a table with exome wide variants in 4 categories (A45, LGD, Missense, Splice)<br> filtered for domiant genes associated with neurodevelopmental disorders in SysID (Prime and Candidate list), a table with exome wide variants in 4 categories (A45, LGD, Missense, Splice) filtered for recessive genes associated with neurodevelopmental disorders in SysID (Prime and Candidate list), a table withcopy number (CN) calls for the cohort and a table withcalls for runs of homozygosity (RoH) regions.</p> <p>#Files and checksums<br> 29c4b2f3dd8985d268f50dd3e0265798 ./FileS2_conNDD-cohort.xlsx<br> a054334637b8b22a9bf743db1e348663 ./FileS3_conNDD-variants.xlsx<br> </p>
The dataset for the research "Evaluation of Digital Supply Chain Technology’s Impact on Sustainability Under the Moderate Effect of Supply Chain Dynamism: An Empirical Research in the Chinese Energy Supply Chain"
In recent years, the topic of digitalisation and sustainability of supply chains has become increasingly important. In addition, as the environmental dynamism becomes more complex, it is essential to explore how technologies impacts on sustainability under the supply chain dynamism. Hence, there is a study to explore the relationship between technologies and sustainability under the supply chain dynamism in the energy supply chain. In this study, the author collects quantitative data from two Chinese companies, including China Resources Power Zhejiang Company and Hunan HuaDian Changsha Electric Co., Ltd. This is a questionnaire survey and it has 24 questions, including 3 general questions, 5 technologies dimension questions, 12 sustainability dimension questions and 4 supply chain dynamism questions. The author collected data from 30 May 2024 to 6 June May 2024, and there are totally 316 answers.
Statistics and Evaluation Data for Publication "Using Supervised Learning to Classify Metadata of Research Data by Field of Study"
<p>Automated classification of metadata of research data by their discipline(s) of research can be used in scientometric research, by repository service providers, and in the context of research data aggregation services. Openly available metadata of the DataCite index for research data were used to compile a large training and evaluation set comprised of 609,524 records. This publication contains aggregated data for the paper. It also contains the evaluation data of all model/hyper-parameter training and test runs.</p>
Contextualized Adaptive Research Description INterfaces Applying LinkedData Evaluation Survey Responses
<p>Building adaptive web forms for publishing research datasets based on contextual information and established ontologies allows the acquisition of fine-grained, structured, semantic metadata which facilitation interdisciplinary findability and reuse.</p> <p>The provided dataset contains the responses of 74 participants in a Turtle (ttl) format from an online survey experiment where they had to use a prototypical web application (CARDINAL) as a proof-of-concept in practice based on a fictious scenario about a political election poll. The evaluation was conducted in July 2020.</p>
Digital research data from: Evaluation of a pH- and time-dependent model for the sorption of heavy metal cations by poultry litter-derived biochar
<p>This is digital research data corresponding to a published manuscript, Evaluation of a pH- and time-dependent model for the sorption of heavy metal cations by poultry litter-derived biochar. Chemosphere (2024), 347, 140688. https://doi.org/10.1016/j.chemosphere.2023.140688. </p> <p>Common isotherm and kinetic models cannot describe the pH-dependent sorption of heavy metal cations by biochar. In this paper, we evaluated a pH-dependent, equilibrium/kinetic model for describing the sorption of cadmium (Cd), copper (Cu), nickel (Ni), lead (Pb), and zinc (Zn) by poultry litter-derived biochar (PLB). We performed sorption experiments across a range of solution pH, initial metal concentration, and reaction time. </p>
Dataset for the research article titled "Evaluation of Reanalysis and Satellite Products against Ground-based Observations in a Desert Environment "
Open the record for dataset details and reuse information.
Supplementary Material for "Evaluating Classifiers in SE Research: The ECSER Pipeline and Two Replication Studies"
<p>This supplementary material for the article"Evaluating Classifiers in SE Research: The ECSER Pipeline and Two Replication Studies. Dell’Anna, D.; Aydemir, F. B.; and Dalpiaz, F. Empirical Software Engineering. 2022" includes</p> <ul> <li>ECSER-ExploratoryStudy.csv: The annotated meta-data of the papers that have been published in ICSE between 2019 and 2021.</li> <li>ECSER_ROCplots+StatTest.ipynb: A python notebook that compares classifiers adn checks the statistical significance of the comparison results.</li> <li>ECSER_SummaryOfReplicationSteps.pdf: This table presents a summary of ECSER steps for the two original studies and our applications on ECSER.</li> <li>ECSER_RE: The directory that holds the datasets and code for the replication of Hay et al. [1] and additional runs on the new data sets.</li> <li>ECSER_FF: The directory that holds the code and data for the replication of Alshammari et al. [2]</li> <li>README.md presents the structure of the supplementary materials.</li> <li>requirements.txt lists the dependencies needed to run the code</li> </ul> <p>In the ECSER_RE directory, the code for multiple classifiers that are compared are kept in the "Classifiers" directory. The public data sets are shared in the Datasets directory. "ECSER_RE_Compare_Classifiers.ipynb" python notebook includes the code that runs each classifier. The results are presented in "ECSER_RE_results-Promise-vs-all.csv".</p> <p>In the ECSER_FF directory, the data sets are presented directly under the main directory. The notebook "ECSER-FF-Compare_Classifiers.ipynb" compares the classifiers of the original study and the results are kept under the "ECSER_FF_results" directory.</p> <p><strong>How to cite this repository </strong><br>If you use this repository, please cite the reference paper, and the repository, as below:</p> <p>Dell’Anna, Davide, Fatma Başak Aydemir, and Fabiano Dalpiaz. "Evaluating classifiers in SE research: the ECSER pipeline and two replication studies." Empirical Software Engineering 28.1 (2023): 3.</p> <p>Davide Dell'Anna, Fatma Başak Aydemir, & Fabiano Dalpiaz. (2021). Supplementary Material for "Evaluating Classifiers in SE Research: The ECSER Pipeline and Two Replication Studies" [Data set]. Zenodo. https://doi.org/10.5281/zenodo.6266675</p> <p> </p> <p> </p> <p> </p> <ol> <li>Tobias Hey, Jan Keim, Anne Koziolek, and Walter F. Tichy. 2020. SupplementaryMaterial of "NoRBERT: Transfer Learning for Requirements Classification". https://doi.org/10.5281/zenodo.3874137</li> <li>Abdulrahman Alshammari, Christopher Morris, Michael Hilton, and JonathanBell. 2021. FlakeFlagger: Predicting Flakiness Without Rerunning Tests. In43rdIEEE/ACM International Conference on Software Engineering, ICSE 2021, Madrid,Spain, 22-30 May 2021. IEEE, 1572–1584. https://doi.org/10.1109/ICSE43902.2021.00140</li> </ol>
l-sized Training and Evaluation Data for Publication "Using Supervised Learning to Classify Metadata of Research Data by Field of Study"
<p>Automated classification of metadata of research data by their discipline(s) of research can be used in scientometric research, by repository service providers, and in the context of research data aggregation services. Openly available metadata of the DataCite index for research data were used to compile a large training and evaluation set comprised of 609,524 records. This is the cleaned and vectorized version with a feature selection of large size.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.