Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
3,481
datasets available to search
ShareScore release 0.9.0
Dataset results
3,481 results for “data set”
Epistemic Insight Initiative Data set: Findings from a hands-on science intervention to build school students' agency and epistemic insight into the power and limitations of science.
<p>Data set that reports on the Findings from a hands-on science intervention to build school students’ agency and epistemic insight into the power and limitations of science.</p> <p>We report on an empirical study that sought to address three questions:</p> <ul> <li>A study to extend the cognitive science of insight into the field of children’s developing understanding of the power and limitations of science</li> <li>A study to examine the effectiveness of the Epistemic Insight (EI) strategy which is designed to give each child hands-on, minds-on science</li> <li>A study to examine the effectiveness of the EI strategy to sustain science activity in schools during the challenging conditions of the pandemic</li> </ul> <p>1919 children and 63 teachers from 24 schools in England took part in the project. This includes both primary and secondary schools. The data reported here are taken from ten primary schools representing 1045 children aged 8-12 in years 4 to 7.</p> <p> </p>
Data sets from "How amenable is type 2 diabetes treatment for precision diabetology? A meta-regression of glycaemic control data from 174 randomised trials"
<p>The two data sets (in CSV format) contain the complete infomation to reproduce the analyses from the paper "How amenable is type 2 diabetes treatment for precision diabetology? A meta-regression of glycaemic control data from 174 randomised trials" which will be published in Diabetologia.</p> <p>The data set "Data_LogSD.csv" contains the data for the primary analysis of the Log(SD) as given in the main paper, the data set "Data_LogSD_BL_CORR.csv" contains the data for the analysis of the baseline-corrected Log(SD) as given in the electronic supplementary material.</p>
Training data and test data sets for simultaneous inversion of velocity density based on U-T
<p>Here are the training and testing data sets involved in the numerical experiments in the article that has been submitted to the journal “Journal of Geophysical Research: Solid Earth”, named “Joint Model and Data-Driven Simultaneous Inversion of Velocity and Density”: Marmousi model. Each dataset consists of two parts: a training dataset and a testing dataset. Both training and testing data sets contain three parts: seismic data, velocity model and density model.</p>
Data set for the evaluation of the pilot course at HTW Berlin on Green Coding by Prof. Dr. Volker Wohlgemuth
<p>Data set for the evaluation of the pilot course at HTW Berlin on Green Coding by Prof. Dr. Volker Wohlgemuth</p>
Data set related to the article: "The Role of BPIFB4 in Immune System and Cardiovascular Disease: The Lesson from Centenarians."
<p>Abstract:</p> <p>Recent discoveries have shed light on the participation of the immune system in the physio pathology of the car- diovascular system underpinning the importance of keeping the balance of the first to preserve the latter. Aging, along with other risk factors, can challenge such balance triggering the onset of cardiovascular diseases.</p> <p>Among several mediators ensuring the proper cross-talk between the two systems, bactericidal/permeability- increasing fold-containing family B member 4 (BPIFB4) has been shown to have a pivotal role, also by sustaining important signals such as eNOS and PKC-alpha.</p> <p>In addition, the Longevity-associated variant (LAV), which is an haplotype allele in BPIFB4 characterized by 4 missense polymorphisms, enriched in homozygosity in Long Living Individuals (LLIs), has been shown to be efficient, if admin- istered systemically through gene therapy, in improving many aspects of cardiovascular diseases (CVDs). This occurs mainly through a fine immune system remodeling across: 1) a M2 macrophage polarizing effect, 2) a favorable redistri- bution of the circulating monocyte cell subsets and 3) the reduction of T-cell activation. Furthermore, LAV-BPIFB4 treat- ment induced a desirable recovery of the inflammatory balance by mitigating the pro-inflammatory factor levels and enhancing the anti-inflammatory boost through a mechanism that is partially dependent on SDF-1/CXCR4 axis.</p> <p>Importantly, the remarkable effects of LAV-BPIFB4 treatment, which translates in increased BPIFB4 circulating levels, mirror what occurs in long-living individuals (LLIs) in whom the high circulating levels of BPIFB4 are protective from age-related and CVDs and emphasize the reason why LLIs are considered a model of successful aging. Here, we review the mechanisms by which LAV-BPIFB4 exerts its immunomodulatory activity in improving the cardiovascular-immune system dialogue that might strengthen its role as a key mediator in CVDs.</p>
Anndata object of 10x Mouse Brain 5k data set for scDGD training
<p>This data from 10x (5k Adult Mouse Brain Nuclei Isolated with Chromium Nuclei Isolation Kit, Single Cell Gene Expression Dataset by Cell Ranger 7.0.0, 2022) is comprised of 7377 cells from the adult mouse brain with 32285 features. Cell type annotations were approximated using CellTpyist with the `Developing_Mouse_Brain` reference model and majority voting. This resulted in 7 distinct cell types.</p>
Data Set "Protein network centralities as descriptor for QM region construction in QM/MM simulations of enzymes"
<p>This data set accompanies the publication "Efficient automatic construction of atom-economical QM regions with point-charge variation analysis" by Felix Brandt and Christoph R. Jacob (TU Braunschweig, Germany) </p> <p>It contains:</p> <p>- PDB files of the starting structures</p> <p>- modified AMBER95 force field file</p> <p>- AMS fragment files for the substrates and ions</p> <p>- AMS input files for all geometry optimizations and single point calculations</p> <p>- Python script for WISP and centrality analysis</p>
Individual and collective school-students' self-efficacy on climate change. Raw data set.
<p>The dataset comprises 12 items on the topic of "self-efficacy on climate change" and additionally some items on person-related information. The items were developed with individual self-efficacy (8 items, iSE1 to iSE8) and collective self-efficacy (4 items, coSE1 to coSE4) in mind. The instruction was: "How do you think about yourself? Please tick to what extent the following statements apply to you" [translated from German]. A six-point Likert-type scale was used as response scale, headed with numbers from 1 to 6, only the ends of the response scale were verbally labelled (1 = "strongly disagree" to 6 = "strongly agree" [translated from German]). The items were used in German. The English translations were added to the dataset in square brackets. Data collection took place at German "Gymnasien" (equivalent to grammar schools) in 2022. The data set includes N = 163 school-students.</p>
Individual and collective university students' self-efficacy on climate change. Raw data set.
<p>The dataset comprises 12 items on the topic of "self-efficacy on climate change" and additionally some items on person-related information. The items were developed with individual self-efficacy (8 items, iSE1 to iSE8) and collective self-efficacy (4 items, coSE1 to coSE4) in mind. The instruction was: "How do you think about yourself? Please tick to what extent the following statements apply to you" [translated from German]. A six-point Likert-type scale was used as response scale, headed with numbers from 1 to 6, only the ends of the response scale were verbally labelled (1 = "strongly disagree" to 6 = "strongly agree" [translated from German]). The items were used in German. The English translations were added to the dataset in square brackets. Data collection took place at a German university in 2022. The data set includes N = 145 university students studying to become geography teachers.</p>
Data set for Dynamic Service Restoration of Distribution Networks with Volt-Var Devices, Distributed Energy Resources, and Energy Storage Systems
<p>Two power distribution systems are presented. The first system consists of 53 nodes and 61 branches, while the second system consists of 404 buses and 430 branches. Both distribution systems offer extensive applications in problems related to multi-time service restoration, Volt/Var devices, and distributed energy resource operation.</p>
Data sets and machine learning models for: Predicting critical properties and acentric factor of fluids using multi-task machine learning
<p>The experimental data sets, data splits, additional features, QM calculations, model predictions, and final machine learning models for the manuscript "Predicting Critical Properties and Acentric Factor of Fluids Using Multi-Task Machine Learning". <strong>Citation should refer directly to the manuscript:</strong></p> <ul> <li> <p>Biswas, S.; Chung, Y.; Ramirez, J.; Wu, H.; Green, W. H. Predicting Critical Properties and Acentric Factors of Fluids Using Multitask Machine Learning. <em>Journal of Chemical Information and Modeling.</em> <strong>2023</strong> <em>63</em> (15), 4574-4588. DOI: <a href="https://doi.org/10.1021/acs.jcim.3c00546">10.1021/acs.jcim.3c00546</a></p> </li> </ul> <p>To use the machine learning models, please refer to the sample files and instructions on <a href="https://github.com/yunsiechung/chemprop/tree/crit_prop">https://github.com/yunsiechung/chemprop/tree/crit_prop</a>. </p> <p>Detailed information can be found in README.md file.</p> <p> </p> <p><strong>Details on the properties considered</strong></p> <p>The data set includes the following 8 properties:</p> <ul> <li>Tc: critical temperature, in K</li> <li>Pc: critical pressure, in bar</li> <li>rhoc: critical density, in mol/L</li> <li>omega: acentric factor, unitless</li> <li>Tb: boiling point, in K</li> <li>Tm: melting point, in K</li> <li>dHvap: enthalpy of vaporization at boiling point, in kJ/mol</li> <li>dHfus: enthalpy of fusion at melting point, in kJ/mol</li> </ul> <p><strong>Details on the files</strong></p> <p>1. Data sets under CritProp_v1.1.0:</p> <ul> <li>all_data: includes the data sets used in this work. All data points are listed for each chemical compound as well as its corresponding data source. The details of the data sources can be found in the README.md file. The distribution of the data set is included in each folder. <ul> <li>estimated_data_for_pretraining: contains the estimated data from Yaws' handbook that are used to pre-train our machine learning (ML) model.</li> <li>experimental_data: contains the experimental data (references 1 - 15) used to fine-tune our final ML model.</li> </ul> </li> <li>additional_features: includes the additional features tested for the ML model. The Abraham features are generated for all data (references 1 - 15) while the acsf, qm, and rdkit features are only generated for the data from references 1 - 9. <ul> <li>abraham: Abraham solute parameters (E, S, A, B, L). Molecular features.</li> <li>acsf: ACSF (atom-centered symmetry functions). Atomic features that are coverted from the 3D coordinates of the compound</li> <li>qm_atom: QM (quantum chemical) atomic feature. </li> <li>qm_mol: QM molecular feature.</li> <li>rdkit: Selected RDKit 2D molecular features.</li> </ul> </li> <li>data_splits_and_model_predictions: contains the training and test sets used to evaluate the model. It also contains the predicted values from our final ML model for each test set. <ul> <li>random and scaffold splits: training and test sets that include the data from references 1 - 9.</li> <li>external test set: a test set that includes the data from only references 10 - 15.</li> </ul> </li> </ul> <p>2. Machine learning (ML) model files:</p> <ul> <li>CritProp_ML_model_files_with_abraham_feat.zip: contains the Chemprop ML model files that are trained using Abraham features as additional molecular features. This gives the best results.</li> <li>CritProp_ML_model_files_without_additional_feat.zip: contains the Chemprop ML model files that are trained without any additional features. This gives the second best results.</li> </ul> <p>To use these ML models, please refer to the sample files and instructions on <a href="https://github.com/yunsiechung/chemprop/tree/crit_prop">https://github.com/yunsiechung/chemprop/tree/crit_prop</a></p> <p>3. QM (quantum chemical) calculations:</p> <ul> <li>QM_calculations.zip: contains the results of the QM calculations that are performed to compute QM features.</li> </ul> <p> </p> <p> </p>
Data from: Axial conduit widening, tree height and height growth rate set the hydraulic transition of sapwood into heartwood
<p><span>The size-related xylem adjustments required to maintain</span><span> a constant leaf-specific sapwood conductance (<em>K<sub>LEAF</sub></em>) with increasing height (<em>H</em>) are still under discussion. Alternative hypotheses are that: (i) the conduit hydraulic diameter (<em>Dh</em>) at any position in the stem and/or (ii) the number of sapwood rings at stem base (<em>NSWr</em>) increase with <em>H.</em> In addition, (iii) lower stem elongation (</span><em>Δ<span>H</span></em><span>) increases the tip-to-base conductance through inner xylem rings, thus possibly the <em>NSWr</em> contributing to <em>K<sub>LEAF</sub></em>.</span></p> <p><span>A detailed stem analysis showed that </span><em><span>Dh</span></em><span><em> </em>increased with the distance from the apex (<em>DCA</em>) in all rings of a <em>P. abies</em> and a <em>F. sylvatica</em> tree. Net of <em>DCA</em> effect, <em>Dh</em> did not increase with <em>H</em>. Using sapwood traits from a global dataset, <em>NSWr</em> increased with <em>H</em> and decreased with </span><em>Δ<span>H</span></em><span>, and the mean sapwood ring width (<em>SWrw</em>) increased with </span><em>Δ<span>H</span></em><span>. A numerical model based on anatomical patterns predicted the effects of <em>H</em> and </span><em>Δ<span>H</span></em><span> on the conductance of inner xylem rings.</span></p> <p><span>Results suggested the sapwood/heartwood transition depends on both <em>H</em> and </span><em>Δ<span>H</span></em><span>, and is set when the C allocation to maintenance respiration of living cells in inner sapwood rings produces a lower gain in total conductance than investing the same C in new vascular conduits.</span></p>
Data set supplementing "Characteristics of Users and Nonusers of Symptom Checkers in Germany: Cross-Sectional Survey Study"
<p>This is the data set used to conduct the analyses in the article published by the Journal of Medical Internet Research under the title "Characteristics of Users and Nonusers of Symptom Checkers in Germany: Cross-Sectional Survey Study" (<a href="https://doi.org/10.2196/46231">https://doi.org/10.2196/46231</a>)</p> <p>This data set contains information collected of 1,084 respondents about their awareness, use, and usefulness of symptom checkers and several individual characteristics. </p>
Data set for "Electrical control of hybrid exciton transport in a van der Waals heterostructure"
<p>Data set for "Electrical control of hybrid exciton transport in a van der Waals heterostructure"</p>
Car Data Set
<p>Car data set </p>
Velocity data from GETM GBA set up
<p>Sample surface velocity data from a GETM model of the Greater Bay Area region encompassing Hong Kong. Data for use with https://github.com/julianmak/ParcelsGETM_demo for a summer school demonstrating the use of the Parcels package (https://oceanparcels.org/)</p>
Data Set for Copper-Catalyzed Benzylic Functionalization of Lignin-Derived Monomers
<p>Raw NMR and HRMS data for the article entitled Copper-Catalyzed Benzylic Functionalization of Lignin-Derived Monomers. The folder names corrspond to the compound names. </p>
Supporting Data Set for Paper "Can Videos as a By-Product of GUI Testing Help Developers Understand GUI Tests?"
<p>This data set is a supporting material for an accepted paper "Can Videos as a By-Product of GUI Testing Help Developers Understand GUI Tests?" on 2023 IEEE 31st International Requirements Engineering Conference Workshops (REW 2023).</p> <p>This dataset consists of</p> <ul> <li>a consent form of the study in English;</li> <li>a tutorial video for TakeNote App (see <em>TakeNote Tutorial-v02</em>);</li> <li>a questionnaire in HTML format;</li> <li>used videos (in <em>HTML Video Player with Videos and VTT files</em>) and screenshots;</li> <li>the source code of the HTML Video Player (in <em>HTML Video Player with Videos and VTT files</em>);</li> <li>obtained and coded results from the questionnaire (see <em>Study-Data-4EmpiRE-v22</em>);</li> <li>calculation steps of the Mann-Whitney U Test;</li> <li>the source code of the TakeNote App.</li> </ul>
TEAMx-PC22 (TEAMx pre-campaing 2022) - ACINN automatic weather station data set from Nafingalm
<p><strong>ABSTRACT</strong></p> <p>This data set was collected with an automatic weather station (AWS) of <a href="http://acinn.uibk.ac.at/">ACINN</a> at the Nafingalm, Austria, in summer 2022 in the framework of the TEAMx pre-campaign 2022 (TEAMx-PC22). The aim of TEAMx-PC22 was to test new instruments, new instrument configurations and new measurement sites to support the planning of the main TEAMx observational campaign (TOC) in 2024/2025. More details about TEAMx can be found at <a href="http://www.teamx-programme.org">http://www.teamx-programme.org</a> as well as in Serafin et al. (2020) and in Rotach et al. (2022).</p> <p><strong>DATA SET DESCRIPTION</strong></p> <p><strong>1. Location</strong></p> <p>The AWS was located close to a small lake at the Nafinglam in the Weer Valley, Tyrol, Austria. The exact location is: 47.2151419°N / 11.712628°E / 1928 m MSL</p> <p><strong>2. Period</strong></p> <p>The TEAMx-PC22 lasted from mid-May 2022 to early October 2022. However, the AWS data set provided here contains the period from 15 June to 12 September 2022. The time series has a measurement interval of 5 minutes and, depending on the parameter, contains both mean values and instantaneous values.</p> <p><strong>3. Instrument details</strong></p> <p>A detailed description of the AWS, its sensors, parameters, calibration and correction procedures is provided as part of the netCDF file metadata, as well as in a PDF file containing the netCDF header extracted with the Linux command ncdump -h.</p> <p><strong>4. Data file</strong></p> <p>The data are provided in a single netCDF file teamx_pc22_aws_nafingalm.nc together with a description of its content in the PDF file teamx_pc22_aws_nafingalm_ncdump_output.pdf.</p> <p><strong>5. Analysis</strong></p> <p>A first analysis of the data was performed in a Bachor thesis (Viebahn, 2023), which is available upon request from the first author of this data set.</p> <p><strong>6. Contact</strong></p> <p>Contact alexander.gohm(at)uibk.ac.at for any questions regarding the data set.</p> <p><strong>7. References</strong></p> <p>Rotach, M. W., S. Serafin, H. C. Ward, M. Arpagaus, I. Colfescu, J. Cuxart, S. F. J. D. Wekker, V. Grubišic, N. Kalthoff, T. Karl, D. J. Kirshbaum, M. Lehner, S. Mobbs, A. Paci, E. Palazzi, A. Bailey, J. Schmidli, C. Wittmann, G. Wohlfahrt, D. Zardi, 2022: A collaborative effort to better understand, measure, and model atmospheric exchange processes over mountains. <em>Bulletin of the American Meteorological Society</em>, <strong>103</strong>, E1282–E1295. <a href="https://doi.org/10.1175/bams-d-21-0232.1">https://doi.org/10.1175/bams-d-21-0232.1</a></p> <p>Serafin, S., M. W. Rotach, M. Arpagaus, I. Colfescu, J. Cuxart, S. F. J. De Wekker, M. Evans, V. Grubišić, N. Kalthoff, T. Karl, D. J. Kirshbaum, M. Lehner, S. Mobbs, A. Paci, E. Palazzi, A. Raudzens Bailey, J. Schmidli, G. Wohlfahrt, B. Zardi, 2020: <em>Multi-scale transport and exchange processes in the atmosphere over mountains: Programme and experiment</em>. Innsbruck University Press. <a href="https://doi.org/10.15203/99106-003-1">https://doi.org/10.15203/99106-003-1</a></p> <p>Viebahn, T., 2023: <em>Windregime und Stabilität in einem alpinen Seitental im Sommer: Eine Standortcharakterisierung im Rahmen der TEAMx Vorkampagne 2022.</em> Bachlor thesis, University of Innsbruck, 75 pp.</p>
Blommersia wittei complex revision - data sets
<p>Data set containing original sequence alignments, call recordings and a table with sequences and sequence metadata for the paper: Vences et al., "<strong>Integrative revision of the <em>Blommersia wittei</em> complex, with description of a new species of frog from western and north-western Madagascar". </strong>Zootaxa 5319 (2): 178–198.</p> <p>The paper can be found at DOI: https://doi.org/10.11646/zootaxa.5319.2.2</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.