Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

50

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

50 results for “Dataset construction”

Learn how ShareScore rates datasets ↗
zenodo36/100

Dataset for paper 'McAN: a novel computational algorithm and platform for constructing and visualizing haplotype networks'

<p>The .zip file includes four datasets for testing the performance of McAN (doi: https://doi.org/10.1093/bib/bbad174).</p>

opencc-by-4.0Oct 2023View details →
zenodo36/100

Dataset for 'Beyond bidirectional association: Distinguishing light verb constructions from other conventionalised noun-verb combinations in modern Tibetan'

Open the record for dataset details and reuse information.

opencc-by-4.0Mar 2024View details →
zenodo36/100

Dataset of the paper "An Empirical Study on the Fault-Inducing Effect of Functional Constructs in Python"

<p>This package contains the dataset of the manuscript&nbsp;&quot;An Empirical Study on the Fault-Inducing Effect of Functional Constructs in Python&quot;</p>

opencc-by-4.0Jun 2022View details →
zenodo36/100

QAngaroo (MedHop + WikiHop) - Constructing Datasets for Multi-hop Reading Comprehension Across Documents

<p>Most Reading Comprehension methods limit themselves to queries which can be answered using a single sentence, paragraph, or document. Enabling models to combine disjoint pieces of textual evidence would extend the scope of machine comprehension methods, but currently no resources exist to train and test this capability. We propose a novel task to encourage the development of models for text understanding across multiple documents and to investigate the limits of existing methods. In our task, a model learns to seek and combine evidence &mdash; effectively performing multihop, alias multi-step, inference. We devise a methodology to produce datasets for this task, given a collection of query-answer pairs and thematically linked documents. Two datasets from different domains are induced, and we identify potential pitfalls and devise circumvention strategies. We evaluate two previously proposed competitive models and find that one can integrate information across documents. However, both models struggle to select relevant information; and providing documents guaranteed to be relevant greatly improves their performance. While the models outperform several strong baselines, their best accuracy reaches 54.5% on an annotated test set, compared to human performance at 85.0%, leaving ample room for improvement.</p>

opencc-by-sa-3.0Jun 2018View details →
zenodo36/100

Dataset obtained from DoE for the construction of lightweight concrete blocks using expanded polystyrene

<p>The dataset used in this study was generated through a Design of Experiments (DOE) approach to predict the material mixture proportions using machine learning techniques. It contains 36 experiments, where the independent variables include the proportions of cement, sand, gravel, and expanded polystyrene (Styrofoam). The dependent variables are the compressive strength values and water absorption rates, calculated across multiple components and used to classify the mixtures.</p> <p>The dataset columns include:</p> <ul> <li><strong>Cement:</strong> The proportion of cement used in the mixture.</li> <li><strong>Sand:</strong> The proportion of sand used.</li> <li><strong>Gravel:</strong> The proportion of gravel.</li> <li><strong>Styrofoam (Sty):</strong> The proportion of expanded polystyrene.</li> <li><strong>Comp. 1-6:</strong> Compressive strength values for different components.</li> <li><strong>Comp. Mean:</strong> The mean compressive strength value.</li> <li><strong>Abs. 1-3:</strong> Water absorption values for different components.</li> <li><strong>Abs. Mean:</strong> The mean water absorption rate.</li> <li><strong>Classification:</strong> The final classification of the mixture based on the obtained results, indicating if the mixture belongs to categories such as "B," "C," "D," or "No classification."</li> </ul> <p>This dataset was used to train and validate machine learning models, aiming to predict the mixture properties based on the material proportions.</p>

opencc-by-4.0Oct 2024View details →
zenodo36/100

Dataset of construction and demolition waste images: aerated autoclaved concrete (AAC), asphalt, ceramics, and concrete

<p>Image subsets: the dataset of images (RGB) of CDW materials (aerated autoclaved concrete (AAC), asphalt, ceramics, and concrete) cropped to 200x200 px. The images are annotated and split into testing and training datasets for the purposes of machine-learning models&#39; training.</p> <p>Whole CDW fragments: images used for validation of algorithms - whole fragments placed on contrast background.</p>

opencc-by-4.0Feb 2023View details →
zenodo36/100

Dataset for "Non-canonical possessive constructions in Negidal and other Tungusic languages: a new analysis of the so-called 'alienable possession' suffix"

<p>This is the dataset used in the paper: Aralova, N. &amp; Pakendorf, B. (2023). Non-canonical possessive constructions in Negidal and other Tungusic languages: a new analysis of the so-called &ldquo;alienable possession&rdquo; suffix.&nbsp;Special Issue &ldquo;Re-assessing the explanatory potential of alienability contrasts&rdquo;, guest-edited by Fran&ccedil;oise Rose &amp; An Van linden. <em>Linguistics</em>. <a href="https://doi.org/10.1515/ling-2022-0030">https://doi.org/10.1515/ling-2022-0030</a></p> <p>For more details, see the ReadMe file.</p>

opencc-by-4.0May 2023View details →
zenodo36/100

DIB _Dataset for the psychometric properties of the construction of a verbal aptitude test instrument to assess prospective high school students' majors

<p>The dataset consists of raw data from the response to the construct resulting from the development of a verbal aptitude test instrument and the results of data analysis using the Rasch model analysis approach.</p>

opencc-by-4.0May 2023View details →
zenodo36/100

DIB_Dataset for the psychometric properties of the construction of a verbal aptitude test instrument to assess prospective high school students' majors

<p>The dataset consists of raw data from the response to the construct resulting from the development of a verbal aptitude test instrument and the results of data analysis using the Rasch model analysis approach.</p>

opencc-by-4.0May 2023View details →
zenodo36/100

Dataset related to paper "Preliminary Analysis of Non-destructive Test Methods to Evaluate the Self-healing Efficiency on the Construction Site"

<p>Dataset of the study reported in the conference paper &quot;&quot;Preliminary Analysis of Non-destructive Test Methods to Evaluate the Self-healing Efficiency on the Construction Site&quot; presented at the SynerCrete&#39;23 Conference (14-16 June 2023, Milos Island, Greece).</p>

opencc-by-4.0Jun 2023View details →
zenodo36/100

R Notebook and Dataset for "Usage-based perspective on argument realisation: A corpus study of Indonesian BUY verbs in applicative construction with -kan" (1.0.0)

<p>This repository contains the dataset and R codes for our paper that has been published in <a href="http://www.aa.tufs.ac.jp/en/publications/nusa">NUSA</a> (<i>Linguistic studies of languages in and around Indonesia</i>) special volume (74) on "Applicatives in Austronesian Languages".</p><h4>How to cite the paper</h4><p>Rajeg, Gede Primahadi Wijaya &amp; I Wayan Arka. 2023. Usage-based perspective on argument realisation: A corpus study of Indonesian BUY verbs in applicative construction with -<i>kan</i>. In Jocelyn Aznar, Christian Döhler &amp; Jozina Vander Klok (eds.), <i>NUSA (special issue on "Applicatives in Austronesian Languages")</i>, vol. 74, 83–114. <a href="https://tufs.repo.nii.ac.jp/records/2000019">https://tufs.repo.nii.ac.jp/records/2000019</a>.</p><h4>Description of the repository</h4><p>The .qmd file contains the R codes used to produce the quantitative analyses in the paper, including the statistical figures. This .qmd file also interweaves some text narratives with the codes. The file is published as a webpage at: <a href="https://gederajeg.github.io/applicative-buy/">https://gederajeg.github.io/applicative-buy/</a></p><p>The raw, annotated concordance data is located in the <a href="https://github.com/gederajeg/applicative-buy/tree/main/data">data</a> directory.</p><p>The statistical figures can also be accessed individually <a href="https://github.com/gederajeg/applicative-buy/tree/main/nusa-applicative-code_files/figure-html">here</a>.</p>

opencc-by-sa-4.0Jun 2023View details →
zenodo36/100

Dataset for Evaluating the Construct Validity of the Charité Alarm Fatigue Questionnaire

<p>These are the datasets that we used for evaluating the construct validity of the Charit&eacute; Alarm Fatigue Questionnaire (CAFQa) in a forthcoming publication. All items were answered on a 5-point Likert scale and were scored by us as follows: -2/&ldquo;I do not agree at all&rdquo;, -1/&ldquo;I do not agree&rdquo;, 0/&ldquo;I agree in part&rdquo;, 1/&ldquo;I agree&rdquo;, 2/&quot;I very much agree&quot;.</p> <p>A previous version of this upload included only the data of Study 1. A new version provides the data of Study 2. Please refer to the methods section of the forthcoming publication for more details.</p> Variable names and their corresponding item. Items marked with <table><tbody><tr> <th>Variable Name</th> <th>CAFQa Item</th> </tr> </tbody><tbody> <tr> <td>procedural_instruction</td> <td>In my ward, procedural instruction on how to deal with alarms is regularly updated and shared with all staff.<sup>a</sup></td> </tr> <tr> <td>respond_quickly</td> <td>Responsible personnel respond quickly and appropriately to alarms.<sup>a</sup></td> </tr> <tr> <td>motivation_decrease</td> <td>With too many alarms on my ward, my work performance, and motivation decrease.</td> </tr> <tr> <td>physical_symptoms</td> <td>Too many alarms trigger physical symptoms for me, e.g., nervousness, headaches, and sleep disturbances.</td> </tr> <tr> <td>ward_floor</td> <td>The acoustic and visual monitor alarms used on my ward floor and in my nurse station allow me to assign the patient, the device, and the situation clearly.<sup>a</sup></td> </tr> <tr> <td>reduce_concentration</td> <td>Alarms reduce my concentration and attention.</td> </tr> <tr> <td>alarm_limits</td> <td>Alarm limits are regularly adjusted based on patients&#39; clinical pictures (e.g., blood pressure limits for conditions after bypass surgery).<sup>a</sup></td> </tr> <tr> <td>interrupt_workflow</td> <td>My or neighboring patients&#39; alarms or crisis alarms frequently interrupt my workflow.</td> </tr> <tr> <td>alarms_confuse</td> <td>There are situations when alarms confuse me.</td> </tr> </tbody> </table> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> Other variables in the data set. <table><tbody><tr> <th>Variable Name</th> <th>Explanation</th> </tr> </tbody><tbody> <tr> <td>self_reported_AF</td> <td>self-estimated alarm fatigue in percent</td> </tr> <tr> <td>estimated_false_alarms</td> <td>perceived rate of false alarms in the participant&#39;s ICU</td> </tr> <tr> <td>monthly_time_on_ICU</td> <td>the average number of workdays per month in an intensive care or monitoring area</td> </tr> <tr> <td>ICU_experience</td> <td>number of years/months of ICU experience</td> </tr> <tr> <td>profession</td> <td>physician, nurse, or supporting nurse</td> </tr> </tbody> </table> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> <p><strong>Members of the Study Group</strong> <strong>in alphabetical order</strong>: <em>Dr. med. Mirza Aghamov</em><sup><em>1</em></sup><em>, Prof. Dr. med. Manfred Blobner<sup>2</sup>, Prof. Dr. med. Ulrich Frey<sup>3</sup></em><em>, Prof. Dr. Christian von Heymann<sup>4</sup></em><em>, Prof. Dr. med. Bettina Jungwirth</em><sup><em>1</em></sup><em>, Dr. med. Dragutin Popovic<sup>4</sup></em><em>, Prof. Dr. med. Michael Sander<sup>5</sup></em><em>. </em></p> <p>1: <em>Department of Anesthesiology and Intensive Care Medicine, University Hospital Ulm, Ulm University, Ulm, Germany</em></p> <p>2: <em>Technical University Munich, School of Medicine, Klinikum Rechts der Isar, Department of Anaesthesiology &amp; Intensive Care Medicine, Munich, Germany</em></p> <p>3: <em>Department for Anesthesiology, Surgical Intensive Care, Pain and Palliative Medicine, Marien Hospital Herne &ndash; Universit&auml;tsklinikum der Ruhr-Universit&auml;t Bochum, Herne, Germany</em></p> <p>4: <em>Department for Anaesthesiology, Intensive Care Medicine and Pain Therapy, Vivantes Klinikum im Friedrichshain, Berlin, Germany</em></p> <p>5: <em>Department for Anaesthesiology, Intensive Care Medicine and Pain Therapy, Justus Liebig University, Giessen, Germany</em></p> <p>&nbsp;</p>

openMar 2023View details →
zenodo32/100

FIGURE 5. Maximum likelihood phylogenetic tree constructed with UCEs and exon loci dataset for the novel species P in A new species of Plumarella (Octocorallia: Calcaxonia: Primnoidae) from the Northeast Pacific, and the redescription of Plumarella longispina Kinoshita, 1908

FIGURE 5. Maximum likelihood phylogenetic tree constructed with UCEs and exon loci dataset for the novel species P. williamsi (in bold), the redescribed species P. longispina (in red), the related taxa and rooted to outgroup genera. ML bootstrap support values&gt;70% are shown above branches.

opennotspecifiedJul 2024View details →
zenodo32/100

The SNP dataset tested for constructing HITSNP algorithm

<p>The SNP dataset tested for constructing HITSNP algorithm</p>

opencc-by-4.0Sep 2024View details →
zenodo32/100

Dataset for "Constructing many-body dissipative particle dynamics models of fluids from bottom-up coarse-graining"

<ul> <li>The trajectory files used for processing time correlation functions, as reported in the original paper, are provided here.</li> <li>The analysis tools can be found in the following GitHub repository: <a href="https://github.com/jaehyeokjin/ManyBodyDPD/tree/main/Time-Correlation" target="_new" rel="noopener">ManyBodyDPD/Time-Correlation</a>. These tools are designed to work with the two trajectory files included in this repository.</li> </ul>

opencc-by-4.0Sep 2024View details →
zenodo32/100

Dataset related to the article "LUMINAL ENDOTHELIALIZATION OF SMALL CALIBER SILK TUBULAR GRAFT FOR VASCULAR CONSTRUCTS ENGINEERING"

<p>This record contains raw data related to the article&nbsp;&ldquo;LUMINAL ENDOTHELIALIZATION OF SMALL CALIBER SILK TUBULAR GRAFT FOR VASCULAR CONSTRUCTS ENGINEERING&quot;</p> <p>The constantly increasing incidence of coronary artery disease worldwide makes necessary to set advanced therapies and tools such as tissue engineered vessel grafts (TEVGs) to surpass the autologous grafts [(i.e., mammary and internal thoracic arteries, saphenous vein (SV)] currently employed in coronary artery and vascular surgery. To this aim,&nbsp;<em>in vitro</em>&nbsp;cellularization of artificial tubular scaffolds still holds a good potential to overcome the unresolved problem of vessel conduits availability and the issues resulting from thrombosis, intima hyperplasia and matrix remodeling, occurring in autologous grafts especially with small caliber (&lt;6 mm). The employment of silk-based tubular scaffolds has been proposed as a promising approach to engineer small caliber cellularized vascular constructs. The advantage of the silk material is the excellent manufacturability and the easiness of fiber deposition, mechanical properties, low immunogenicity and the extremely high&nbsp;<em>in vivo</em>&nbsp;biocompatibility. In the present work, we propose a method to optimize coverage of the luminal surface of silk electrospun tubular scaffold with endothelial cells. Our strategy is based on seeding endothelial cells (ECs) on the luminal surface of the scaffolds using a low-speed rolling. We show that this procedure allows the formation of a nearly complete EC monolayer suitable for flow-dependent studies and vascular maturation, as a step toward derivation of complete vascular constructs for transplantation and disease modeling.</p>

opencc-by-4.0Dec 2022View details →
zenodo28/100

Intelligent Java Dataset Construction and Visualization Evaluation for Reliable Software Development

<p><span>These datasets include LiuKui, New-100, New-400, and the dataset utilized in RQ5. </span></p>

opencc-by-4.0Apr 2024View details →
zenodo28/100

Intelligent Java Dataset Construction and Visualization Evaluation for Reliable Software Development (Sci)

Open the record for dataset details and reuse information.

opencc-by-4.0Apr 2024View details →
zenodo28/100

Dataset construction method of cross-lingual summarization based on filtering and text augmentation

<p>The NCLS dataset is provided by its authors (Zhu et al.): https://drive.google.com/file/d/1GZpKkHnTH_1Wxiti0BrrxPm18y9rTQRL/view. We work on the train set, validation set, and manually corrected test set.</p>

opencc-by-4.0Mar 2023View details →
dryad28/100

The supplementary datasets of the study of free moment Induced by oblique transverse tarsal joint: investigation by constructive approach

Open the record for dataset details and reuse information.

publicMar 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record