Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
20,299
datasets available to search
ShareScore release 0.7.1
Dataset results
20,299 results for “systematics”
MiRoR15-P1-Tools used to assess the quality of peer review reports: a methodological systematic review
<p>Database, data extraction form, R codes and protocol related to: Superchi C, González JA, Solà I, Cobo E, Hren D, Boutron I. <em>Tools used to assess the quality of peer review reports: a methodological systematic review</em>. BMC Med Res Methodol. 2019;19(48):1–14. DOI: <a href="https://doi.org/10.1186/s12874-019-0688-x">https://doi.org/10.1186/s12874-019-0688-x</a></p> <p> </p>
Dataset related to the manuscript: "An open-source integrated framework for the automation of citation collection and screening in systematic reviews"
<p>Dataset related to the manuscript: “An open-source integrated framework for the automation of citation collection and screening in systematic reviews”, to be used together with the code stored at https://github.com/AD-Papers-Material/BART_SystReviewClassifier to reproduce the results.</p> <p>There are three datasets:<br> - The Record data collected from the online scientific databases;<br> - The session journal which describes the search session, i.e., how many records were collected and from which source, for each query/session pairs.<br> - The session data which is the outcome of the classification and review tasks;</p>
Cross cultural tears: A systematic investigation of the interpersonal effects of emotional crying across different cultural backgrounds
<p>The present project wants to examine the importance of emotional crying as an attachment behaviour and its fundamental role across a number of diverse cultures.</p> <p>Emotional tears are uniquely human and have fascinated scholars across several decades (Vingerhoets, 2013). Some researchers argue that tearful crying played a significant role in the evolution of humankind with regard to social development and solidarity (Walter, 2006). Recent years have seen an increased interest in exploring the interpersonal effects of human tears (see Gračanin, Bylsma, & Vingerhoets, 2018 for a review), with findings that emotional tears foster approach or support behavior (Gračanin, Krahmer, Rinck, & Vingerhoets, 2018) and crying individuals being evaluated as more communal (e.g., Zickfeld, van de Ven, Schubert, & Vingerhoets, 2018). These findings generally fit the hypothesis that emotional tears constitute a social act, promote social bonding and fulfill an attachment function (Nelson, 2005; Bowlby, 1982; Gračanin, Bylsma, et al., 2018; Murube, Murube, & Murube, 1999; Radcliffe-Brown, 1922; Vingerhoets, 2013). The present projects aims to answer the question whether emotional tears present a fundamental form of solidarity and bonding and whether the findings on increased attributions of warmth and higher approach intentions for tearful individuals replicate across a number of diverse contexts.</p> <p><a href="https://osf.io/fj9bd">Published in OSF: https://osf.io/fj9bd</a>.</p> <p>The OPen SCience FRamework also includes:</p> <ul> <li>Data from the pilot study: https://osf.io/txcw3/</li> <li>General information about the translation process (including the Portuguese version https://osf.io/t4cas/),</li> <li>Data management and research protocol (https://osf.io/5bh7m/),</li> <li>Approvals from ethical committees (https://osf.io/v8rqh/),</li> <li>Data and descriptive document of supplementary analyses (https://osf.io/s8ack/),</li> <li>Data and syntax (https://osf.io/x2pks/),</li> <li>Information about stimuli (https://osf.io/x2pks/) </li> </ul>
MiRoR2 - P1 - Shortcomings in the evaluation of biomarkers in ovarian cancer: a systematic review
<p>Data set for the study “Shortcomings in the evaluation of biomarkers in ovarian cancer: a systematic review”, including search strategy, extraction form, extracted data with summary of results, and protocol</p>
AIOps Systematic Mapping Study - Results
<p>Results from our work "A Systematic Mapping Study in AIOps" (https://arxiv.org/abs/2012.09108).</p> <p>The file 'papers.csv' contains the complete list of AIOps papers identified, with corresponding metadata (author, year, citations, venue) and indexing annotations (macro-area, category, data sources, etc.) as columns. This file enables to explore and reproduce our results from scratch. For a quick and interactive exploration of the results without coding, you can check out the same dataset on Exploratory.io (https://exploratory.io/project/EDK0DNx1Qe/AIOps___Mapping_Study_pXe9pUn2). </p>
Supplementary data for the article: Future environmental impacts of metals: a systematic review of impact trends, modelling approaches, and challenges
<p>This repository provides the supplementary data to the paper titled <a href="https://doi.org/10.1016/j.resconrec.2024.107572" target="_blank" rel="noopener"><em>"Future environmental impacts of metals: a systematic review of impact trends, modelling approaches, and challenges"</em></a>, published 2024 in <em>Resources, Conservation and Recycling</em>.</p> <h4><strong>Contents</strong></h4> <p>The repository is split in 3 parts and comprises the following files (more details are provided in the <em>README.md</em>):</p> <p><strong>A_Database of reviewed studies:</strong></p> <ul> <li>contains the detailed review data, meant for readers to use as an overview file to gather studies relevant to them. It also includes an overview of all data sources that the reviewed studies used.</li> </ul> <p><strong>B_Scientific supplement to paper:</strong></p> <ul> <li>Contains all data relevant to the related publication Harpprecht et al. (2024), such as studies screened , FAIR data analysis, or analyzed impact trends.</li> </ul> <p><strong>C_Data for figures in paper:</strong></p> <ul> <li>This file contains all the data for Figures 3, 4 and 5 in tabular form, representing impact trends, scenario variables, scenario modelling approaches and data sources used.</li> </ul> <h4><strong>Summary</strong></h4> <p>These files allow to reproduce the results of our study. In this work, we systematically reviewed studies which assessed future environmental impacts of metal supply chains. Our review yielded 40 publications covering 15 metals: copper, iron, aluminium, nickel, zinc, lead, cobalt, lithium, gold, manganese, neodymium, dysprosium, praseodymium, terbium, and titanium. We evaluated their results regarding future impact trends, and their methods, i.e., modelling approaches, scenario variables, and data sources of scenario variables. We identified 15 scenario variables. The most common variables are background electricity mix, ore grade, recycling shares, demand, and energy efficiency. We identified 229 unique data sources for the reviewed scenario variables.</p> <h4><strong>Related publication</strong></h4> <p>More details on the data and its interpretation as well as the scientific context are provided in the publication itself:</p> <p><a href="https://doi.org/10.1016/j.resconrec.2024.107572" target="_blank" rel="noopener">Harpprecht, C., Miranda Xicotencatl, B., van Nielen, S., van der Meide, M., Li, C. , Li, Z., Tukker, A., Steubing, B. (2024). <em>Future environmental impacts of metals: a systematic review of impact trends, modelling approaches, and challenges.</em> Resources, Conservation and Recycling.</a></p> <h4><strong>Funding </strong></h4> <p>Carina Harpprecht received funding from the Energy Program of the German Aerospace Center in 2022. Zhijie Li received funding from the European Institute of Innovation and Technology (EIT) under the project Valomag (Project No. 14049).</p> <h4><strong>License</strong></h4> <p>CC-BY 4.0 license for DLR (German Aerospace Center)</p>
Dataset for: A systematic literature review on user factors to support the sense of presence
<p>This dataset was created for a publication of Wiepke, Axel and Heinemann, Birte called "A systematic literature review on user factors to support the sense of presence". In this paper we used the PRISMA-method to collect Papers via Google Scholar on the third of April 2023 with the search term:<br>(framework OR model OR frameworks OR models OR processes OR ontologies) AND ((“personality traits” OR “personality variables” OR “personality factors”) AND “spatial presence”) AND (“virtual reality”) AND (learn OR edu\*)</p> <p>The results were pictured in "agreed" findings, where more than 50% of found studies supported a category of results and in "controversial", where there were significant findings, but less than 50% of the studies reported significance.</p> <p>This dataset contains:</p> <ul> <li>raw data for our literature review in .bib</li> <li>our main findings with categories in .csv</li> <li>a short Jupyter notebook script for one graphic in ipynb</li> <li>other graphics as .png</li> </ul>
Systematic Literature Review on Tourism Marketing in the Metaverse
<p>The aim of this research is to investigate tourist marketing within the embryonic context of the metaverse in order to comprehend the building blocks and the primary technologies employed in the sector. For this purpose, a systematic literature review is conducted. The references are extracted in January 2023. The data in this document correspond to the articles finally included in the systematic literature review after the article screening phase.</p> <p>Keywords: tourism marketing, metaverse, technologies, building blocks, SLR (Systematic Literature Review), PRISMA</p> <p> </p>
New Challenges in Point Cloud Visual Quality Assessment: A Systematic Review (Dataset)
<p>This dataset is a collection of annotated information on the scientific papers screened and analyzed for the systematic review of the literature in Point Cloud Visual Quality Assessment. </p> <p>The data is structured as follows:</p> <ul> <li>General information <ul> <li>Document title</li> <li>Authors</li> <li>Year of publication</li> <li>Venue (Conference or Journal title)</li> <li>Citations (number)</li> <li>URL/DOI</li> </ul> </li> </ul> <ul> <li>About the content <br> <ul> <li>Content Type: Point clouds (PC), Colored Point clouds (CPC), Meshes, Dynamic Point Clouds (DPC)</li> <li>Content source: Source of the content used in a subjective QA test or the evaluation of one or more QA metrics</li> </ul> </li> </ul> <ul> <li>About metric benchmarks <ul> <li>Subjective Ground-truth Data: Dataset(s) Source of the subjective scores used as ground-truth in a QA metric benchmark</li> <li>Assessed Metrics: Types of metrics assessed in a benchmark (JPEG standards, IQM, NR, State-of-the-art, others)</li> <li>Performance Measures: PLCC, SROCC, KRCC, RMSE, OR, others</li> </ul> </li> </ul> <ul> <li>About Objective QA metrics <ul> <li>Metric: Name given to the metric introduced in this paper</li> <li>Base: 3D-based or Projection-based</li> <li>Categories: Categories that characterize the approach of the proposed metric (Feature-based, Learning-Based, Perceptual-based, IQM, others) </li> <li>Reference: Full-Reference (FR), Reduced-Reference (RR) or No-Reference (NR)</li> </ul> </li> </ul> <ul> <li>About Subjective QA experiments <ul> <li>Display: Type of display (2D, 3D, AR, MR, VR) and interaction approach (passive, interactive, 3DoF, 6DoF) used in the described experiment.</li> <li>Rendering: Type of rendering used to display the stimuli (Points, Squares, Cubes, Surface)</li> <li>Lab/Remote: The experiment was run in one or more lab environments, or remotely (Lab, Cross-Lab, Remote)</li> <li>Rating: Subjective rating methodology used in the experiment (ACR, DSIS, PWC, others)</li> <li>Dataset: Name of the new subjective dataset if the experiment's results were published.</li> <li>Observers: Number of observers </li> <li>Distortion type: Types of distortions applied to the stimuli and assessed in the experiment</li> </ul> </li> </ul>
Automated Literature Screening for Systematic Reviews: Dataset for Evaluation Against Human Title and Abstract and Full-Text Screening Decisions
<p>This Zenodo entry contains the supplementary material associated with the manuscript titled <em>Automated Literature Screening for Systematic Reviews: A 5-Tier Prompting Approach Meeting Cochrane’s Sensitivity Requirement of Greater Than 0.99.</em> The paper will be presented at <a href="https://dbis.rwth-aachen.de/LLMs4MI2024/">LLMsMI 2024</a> in November 2024.</p> <p>A script is provided for replicating the executed experiments, along with a comprehensive evaluation file that reports all the experiment results. Provided data files represent an extension to the original datasets as provided by [1]. For associated systematic review manuscripts and eligibility criteria, please refer to [1] as well. </p> <p>[1] Guo, Eddie; Gupta, Mehul; Deng, Jiawen; Park, Ye-Jean; Paget, Mike; Naugler, Christopher (2023). "Automated Paper Screening for Clinical Reviews Using Large Language Models." <em>Mendeley Data</em>, V1, doi: 10.17632/np79tmhkh5.1. Accessed from: <a href="https://data.mendeley.com/datasets/np79tmhkh5/1" target="_new" rel="noopener">https://data.mendeley.com/datasets/np79tmhkh5/1</a>.</p>
Data supporting 'Empirical correction of systematic orthorectification error in Sentinel-2 velocity fields for Greenlandic outlet glaciers'
<p><strong>Note: An updated dataset covering the majority of Greenland's marine-terminating glaciers is available as part of the NASA Making Earth System Data Records for Use in Research Environments (MEaSUREs) project through the National Snow and Ice Data Center (NSIDC) at <a href="https://doi.org/10.5067/B28FM2QVVYWY">https://doi.org/10.5067/B28FM2QVVYWY</a>. </strong></p> <p>Data supporting the paper:</p> <blockquote> <p>Chudley, T. R., Howat, I. M., Yadav, B. N., & Noh, M. J. (2022). Empirical correction of systematic orthorectification error in Sentinel-2 velocity fields for Greenlandic outlet glaciers. <em>The Cryosphere. </em>16, 2629–2642, https://doi.org/10.5194/tc-16-2629-2022</p> </blockquote> <p>Dataset consists of four netCDF files containing stacked Sentinel-2 velocity data of four Greenlandic outlet glaciers (Helheim Glacier, Jakobshavn Isbræ, Store Glacier, and Kangerlussuaq) between 2017 and 2021. Velocity data are derived and corrected following the methods outlined in Chudley <em>et al.</em> (2022). </p> <p>NetCDF files are created by, and tested to be readable by, Python's xarray package.</p> <p>The dimensions of the netCDF file are as follows:</p> <ul> <li><strong>X</strong> - <em>x </em>coordinates in NSDIC Sea Ice Polar Stereographic North (EPSG:3413).</li> <li><strong>Y</strong> - <em>y</em> coordinates in NSDIC Sea Ice Polar Stereographic North (EPSG:3413).</li> <li><strong>time</strong> - temporal midpoint of velocity field.</li> </ul> <p>The variables of the netCDF file are as follows:</p> <ul> <li><strong>dmag</strong> - the absolute magnitude of the velocity, in metres per day.</li> <li><strong>dx</strong> - the velocity in the <em>x</em> direction, in metres per day.</li> <li><strong>dy</strong> - the velocity in the <em>y</em> direction, in metres per day.</li> <li><strong>date1</strong> - the date and time of the first scene acquisition.</li> <li><strong>date2</strong> - the date and time of the second scene acquisition.</li> <li><strong>baseline</strong> - the temporal baseline, in days, between scene acquisitions.</li> <li><strong>orbit_pair</strong> - the combination of orbital pathways in the string format 'RXXX_RYYY', where XXX is relative orbit number of the first scene and YYY the relative orbit number of the second scene.</li> <li><strong>mag_rmse</strong> - the root mean square error of the absolute velocity of the off-ice area. </li> <li><strong>dx_mean</strong> - the mean velocity of the off-ice area in the <em>x</em> direction.</li> <li><strong>dx_sd</strong> - the standard deviation of the velocity of the off-ice area in the <em>x</em> direction.</li> <li><strong>dy_mean</strong> - the mean velocity of the off-ice area in the <em>x</em> direction.</li> <li><strong>dy_sd</strong> - the standard deviation of the velocity of the off-ice area in the <em>y</em> direction.</li> </ul>
Architectural Languages for the Microservices Architecture: A systematic mapping study [Data set]
<p>This repository contains all artifacts related to the study: Architectural Languages for the Microservices Architecture: A systematic mapping study.</p>
Cancer screening attendance rates in transgender and gender-diverse patients: a systematic review and meta-analysis
<p>Supplementary Data to support the findings of a systematic review investigating cancer screening rates in transgender and gender-diverse individuals.</p>
Efficacy and safety of subcutaneous vs. sublingual immunotherapy in allergic rhinitis: a systematic review and meta-analysis
<p>Allergic rhinitis significantly impacts patients' quality of life, and allergen immunotherapy (AIT) offers an alternative to conventional treatments. This study compares the efficacy and safety of subcutaneous immunotherapy (SCIT) and sublingual immunotherapy (SLIT) for allergic rhinitis. A comprehensive search of PubMed, Embase, and ClinicalTrials.gov identified nine randomized controlled trials involving 780 patients (427 SCIT, 353 SLIT). The primary outcome was symptom score; secondary outcomes included medication score, symptom medication score, and local and systemic reactions. Results showed SCIT significantly reduced symptom scores compared to SLIT (Pooled SMD: -0.52, 95% CI: -0.60, -0.03, I2 =83%, P<0.05). However, SCIT patients experienced more severe systemic reactions (grade 3&4) than SLIT patients (Pooled SMD: 6.27, 95% CI: 1.47, 26.73, I2 =0%, P=0.01). Other outcomes were comparable between both groups. In conclusion, SCIT is slightly more effective than SLIT but is associated with a higher frequency of severe systemic reactions, guiding clinicians to tailor treatments to individual patient needs.</p>
Is there a non-invasive biomarker for the early detection of ovarian torsion? A systematic review and meta-analysis
<p>We have performed a systematic review and meta-analysis and identified multiple biomarkers that warrant further study as part of a broader diagnostic panel for ovarian torsion. These include SCUBE1, s-DD, IL-6, IMA and TNF-a. </p>
Datasets for phylogenetic analyses and phylogenetic trees for: Genetic barcodes for species identification and phylogenetic estimation in ghost spiders (Araneae: Anyphaenidae: Amaurobioidinae). Invertebrate Systematics, 2024
<p>We combined the COI sequence data with legacy multigene sequence data to create a new, taxon-rich phylogeny for the Amaurobioidinae. We used sequences for four loci that have been used in previous studies on the subfamily: two mitochondrial loci, COI (658bp) and ribosomal subunit 16S (16S, 410bp); and two nuclear loci, Histone H3 (H3, 327bp) and ribosomal subunit 28S (28S, 839bp). We complemented the Amaurobioidinae data with sequences from several non-amaurobioidine anyphaenids and two clubionids as outgroups. Sequence alignment was performed using the MAFFT (ver. 7.308) plugin in Geneious, allowing MAFFT to automatically select an appropriate alignment strategy based on the properties of each locus, or with the online MAFFT server (https://mafft.cbrc.jp), which consistently selected the L-INS-i algorithm. Finally, alignments of the four loci were concatenated to construct a 2234 bp multigene sequence matrix containing 692 taxa, with about 55% missing/gap data (“full” matrix henceforth). To ensure that excessive missing data did not affect the resulting topology, we also constructed a reduced matrix by removing additional COI-only specimens so that each species and morphotype was represented by just one or two specimens for which all loci were available (where possible). After realignment, this reduced matrix was 2235 bp long, included 167 taxa, and had about 22% missing/gap data (“reduced” matrix henceforth). Phylogenetic analyses under maximum likelihood, including model selection, were then conducted with IQ-TREE 2. We performed phylogenetic analyses on both concatenated matrices (the full matrix and the reduced matrix) and on each individual locus. For model selection, we provided an initial scheme that partitioned the matrix by locus, and further partitioned the protein-coding loci (COI and H3) by codon position. We used ModelFinder and searched for the best partition scheme, all in IQ-TREE. The best models (partitions) for the full dataset were: GTR+F+I+G4 (16S), GTR+F+I+I+R4 (28S), TVM+F+I+I+R2 (COI-1), TIM2+F+R4 (COI-2), GTR+F+R5 (COI-3), TVMe+G4 (H3-1-H3-2), SYM+G4 (H3-3); and for the reduced dataset: GTR+F+I+G4 (16S), GTR+F+I+G4: (28S), GTR+F+I+G4: (COI-2), GTR+F+I+G4: (COI-3), TVM+F+I+G4: (COI-1, H3-2), GTR+F+I+G4: (H3-1), GTR+F+I+G4: (H3-3). For each dataset, once the best models and partitions were defined, we executed 10 independent replicates of tree calculations followed by 1000 ultrafast bootstrap replicates, and the replicate reaching the maximum likelihood was chosen. Phylogenetic analyses under parsimony were made with TNT, under equal weights, using the “new technology” search with default values, asking for 10 independent hits to the minimal length, and submitting the resulting trees to a round of TBR branch swapping. </p>
Dataset: Systematics of the color-polymorphic spider genus Cybaeolus, with comments on the phylogeny of the family Hahniidae (Araneae)
<p>Phylogenetic analysis of the spiders of the genus Cybaeulus, with outgroups in the marronoid clade. Data from six DNA markers, analyzed with maximum likelihood and parsimony.</p> <p><br>PHYLOGENETIC ANALYSIS</p> <p>We obtained sequences from 26 samples of the three known species of Cybaeolus, and of five additional species of Hahniidae. To these, we added legacy sequences of Cybaeolus and of other genera of Hahniidae, as well as representatives of the remaining families in the marronoid clade. For the new sequences, the extraction and amplification of DNA was made in the Laboratory of Molecular Tools at Museo Argentino de Ciencias Naturales (MACN), from tissues preserved in absolute alcohol at -18ºC. We targeted the markers histone H3 (H3), cytochrome oxidase subunit I (CO1), 28S ribosomal RNA (28S) and 16S ribosomal RNA (16S), previously used to estimate relationships of marronoid spiders (Wheeler et al., 2017). Details of extraction, primers and PCR protocols are the same as in Magalhaes & Ramírez (2022). Sequencing was outsourced to Macrogen Inc., South Korea. The resulting chromatograms were analyzed individually to detect contaminated sequences or ambiguous portions. In addition to these sequences obtained in the laboratory, we combined our data with additional sequences from previous work (Wheeler et al., 2017; Rivera-Quiroz et al., 2020), using the markers mentioned above plus 12S ribosomal RNA (12S) and 18S ribosomal RNA (18S). For the CO1 marker, additional sequences obtained by the Arachnology Division at MACN and deposited in the BOLDSYSTEMS platform (https://www.boldsystems.org/) were also used. Sequences were aligned with MAFFT Online v.7.463 (Katoh & Standley, 2013), using the L-INS-I algorithm. See Table 1 for list of vouchers and sequence identifiers.</p> <p>Maximum likelihood<br>For the maximum likelihood analyses we used the program IQ-TREE 2.2.0 (Minh et al., 2020), partitioning the data by marker, and selecting the best combination of partitions and evolution models by Bayesian information criterion (best fitting models were TPM2+I+G4 for H3, GTR+F+I+G4 for 18S, GTR+F+I+G4 for 16S and 12S together, GTR+F+I+G4 for CO1, and GTR+F+I+G4 for 28S). Since the relationships of outgroup taxa in the resulting trees were slightly different to that found in recent phylogenomic studies, we used the study of Gorneau et al. (2023) based on ultraconserved elements as a backbone topology to constrain our tree search, considering only the taxa in common with our analysis (see supplementary Fig. S1); this means that all the rest of the taxa are free to move anywhere during tree search. Support for groups (branches) was estimated by 1000 cycles of ultrafast bootstrapping. Ten independent runs were performed; of those, six converged into nearly identical log likelihood values (-57417.7725 to -57417.9604) and identical topologies; the tree with top-ranking log likelihood is presented in Results, after collapsing branches with bootstrap below 0.5. To estimate the support of an alternative topology with Cybaeolus as sister to the rest of the hahniids, we used TNT 1.6 (Goloboff & Morales, 2023) to modify the optimal tree placing Cybaeolus in such position, and asked for the frequency of the branch of interest (all hahniids except Cybaeolus) in the 1000 bootstrapped trees previously saved by IQTREE.<br>Ancestral character states for the arrangement of spinnerets (grouped; separated in a transversal line) were estimated by maximum likelihood on the optimal tree, using the R packages phytools and ape, under the models ER and ARD, and the best fitting model selected by the Akaike information criterion. </p> <p>Parsimony<br>For the parsimony analyses we used TNT 1.6. For the equal weights analysis, a heuristic search was made using a driven search with the default parameters of the “new technologies”, aiming for 10 independent hits to minimum length. The resulting trees were then submitted to an additional round of tree-bisection reconnection (TBR) branch swapping. These results were compared to a simpler search strategy of 300 random addition sequences, each followed by TBR, which produced 20 hits to minimal length. As both strategies reached the same trees with multiple independent hits, it is likely that the optimal trees were found. Finally, the strict consensus of all the optimal trees was obtained, and on this consensus the support values were calculated by means of 1000 bootstrap pseudoreplicates. </p>
Data from systematic audit for paper: Insights into the quantification and reporting of model-related uncertainty across different disciplines
<p>This upload contains 7 data files (each contains cleaned and compiled data for a given scientific field) and 2 R scripts. These files support the paper: Insights into the quantification and reporting of model-related uncertainty across different disciplines.</p> <p> </p> <p><strong>Description of the data</strong></p> <p>Compiled data files for each field contain all reviewers audit answers for eligible papers. All papers that met exclusion criteria have been removed.</p> <p>Data checks have been performed and formatting errors corrected either in R or manually, following steps detailed in the STAR methods.</p> <p>Column names and description:</p> <ul> <li>Number: number of question from 1 to 9</li> <li>Questions: question text – question to be answered by the reviewer</li> <li>QuestionCode: shortened code for each question</li> <li>Paper: paper code - first author surname/initial and surname and year</li> <li>Initials: initials of reviewer</li> <li>Answer: answer to the question</li> <li>Details: extra details to support the answer</li> <li>Location: where in the text the uncertainty was presented</li> <li>Presentation: how the uncertainty was presented</li> <li>ModelType: type of model (focal model)</li> <li>Comments: any other comments from the reviewer</li> <li>Checks: checks of whether NA or no have been included in correct places e.g. if answers to questions 1:4 are no then question 9 is NA, if question 7 is no then 8 is NA</li> <li>Check 1 = when Answer = No, Location is NA</li> <li>Check 2 = when Answer to Number 1-4, 6 or 8-9 is Yes that Details are not NA</li> <li>Check 3 = when Answer = No, Presentation = NA</li> <li>Check 4 = when Location is not NA, presentation is not NA</li> <li>Check 5 = if the Answer to 5 or 7 is "No" then Answer to 6 and 8 = "NA"</li> <li>Check 6 = if Answer for 1-4 is "No", then Answer for 9 = "NA"</li> </ul> <p><strong>Code description</strong></p> <p>Two scripts are included, the first is theme_script.R, this includes code to set up a ggplot theme for the figures. The second is Figure_code.R, this script contains all code to plot and save the three figures from the paper.</p>
Dataset for paper: A Systematic Literature Review and Recommendations for Ontology-based Support of Digital Forensics
<p>PLEASE, READ THE README.TXT FILE</p> <p>This document describes how to interpret the data and metadata files, and it is licensed under Creative Commons CC BY-NC-AS (https://creativecommons.org/licenses).</p> <p>The file "primary_studies_final_set-DATA.csv" is a CSV file format and contains the raw data extracted from our systematic literature review primary studies. Such data were extracted based on the research questions defined for our study.<br> The file "primary_studies_final_set-METADATA.csv" is a CSV file format and contains the following:<br> - the first row contains two pieces of information: the data type, which might be original or reused;<br> - the second row contains the reused data URL/DOI, which should inform the URL or DOI from which the data was reused, or n/a if the data is original;<br> - the third row contains the date of data generation in the format mm/dd/yyyy;<br> - the fourth row contains 11 elements describing each of the fields of the file "primary_studies_final_set-DATA.csv": the study ID, title, objective, six research questions, and an observation field; and<br> - the fifth row describes the data type of each field of the file "primary_studies_final_set-DATA.csv".<br> The .bib files contain the bibtex entry for the final set of studies.<br> The license.txt file describes the Creative Commons license for this material.</p> <p>We hope you have an excellent read!!</p> <p>Cheers!<br> Thiago, Edson, and Avelino</p>
Not So Weak-PICO: Leveraging weak supervision for Participants, Interventions, and Outcomes recognition for systematic review automation
<p>EBM-PICO is a widely used dataset with PICO annotations at two levels: span-level or coarse-grained and entity-level or fine-grained. Span-level annotations encompass the full information about each class. Entity-level annotations cover the more fine-grained information at the entity level, with PICO classes further divided into fine-grained subclasses. For example, the coarse-grained Participant span is further divided into participant age, gender, condition and sample size in the randomised controlled trial. This dataset comes pre-divided into a training set (n=4,933) annotated through crowd-sourcing and an expert annotated gold test set (n=191) for evaluation.</p> <p>The <a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC6174533/bin/NIHMS988059-supplement-Appendix.pdf">EBM-PICO annotation guidelines</a> caution about variable annotation quality. <a href="http://ceur-ws.org/Vol-2429/paper1.pdf">Abaho et al.</a> developed a framework to post-hoc correct EBM-PICO outcomes annotation inconsistencies. <a href="https://arxiv.org/pdf/1904.09557.pdf">Lee et al.</a> studied annotation span disagreements suggesting variability across the annotators. Low annotation quality in the training dataset is excusable, but the errors in the test set can lead to faulty evaluation of the downstream ML methods. We evaluate 1% of the EBM-PICO training set tokens to gauge the possible reasons for the fine-grained labelling errors and use this exercise to conduct an error-focused PICO re-annotation for the EBM-PICO gold test set. The file 'test_ebm_correctedlabels.tsv' has error corrected EBM-PICO gold test set.</p> <p> </p> <p>The upload also contains two zip files containing labelling sources mentioned in the Distant-PICO paper. </p> <ol> <li>ds_cto_dict.zip: contains the four distant supervision dictionaries (P: participant.txt, I = intervention.txt, intervetion_syn.txt, O: outcome.txt) generated from clinicaltrials.gov using the methodology described in Distant-CTO. </li> <li>handcrafted_dictionaries.zip: contains three files <ul> <li>gender_sexuality.txt: contains a list of possible genders and sexual orientations found across the web. The list is not comprehensive.</li> <li>endpoints_dict.txt: contains outcome names and the names of questionnaires used to measure outcomes assembled from PROM questionnaires and PROMs.</li> <li>comparator_dict: contains a list of idiosyncratic comparator terms like a sham, saline, placebo, etc., compiled from the literature search. The list is not comprehensive.</li> </ul> </li> </ol>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.