Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
566
datasets available to search
ShareScore release 0.9.0
Dataset results
566 results for “analytics”
Analytical Framework for Precise Relative Motion in Low Earth Orbits
<p>The data sets provided here can be used to recreate the plots of the paper “Analytical Framework for Precise Relative Motion in Low Earth Orbits” available at this <a href="https://arc.aiaa.org/doi/10.2514/1.G004716">link</a>.</p> <p>That paper presents a practical and efficient analytical framework for the precise modelling of the relative motion in low Earth orbits.</p>
Dataset of reports about MOF-based SERS substrates since 2011 until March 2023. Structure, characteristics, analytes, and performances.
<p>This dataset was generated to aid the creation of a review article addressing the use of Metal-Organic Frameworks (MOF)-based Surface Enhanced Raman Spectroscopy (SERS) platforms for the detection of Volatile Organic Compounds (VOCs).</p> <p>This dataset was generated employing the Web of Science database, encompassing manuscripts published up to March 2023. A literature search was initially conducted using a combination of keywords, including "MOF," "Metal-Organic Framework," "SERS," "Surface Enhanced Raman Spectroscopy," and "Surface Enhanced Raman Scattering." This search spanned the "Topic" category, enabling exploration across title, abstract, author keywords, and keyword-plus fields.</p> <p>From the initial pool of 238 documents, review articles and duplicates were systematically excluded, resulting in a refined collection of 182 articles. Subsequently, articles not concurrently addressing MOF and SERS or those utilizing MOF as sacrificial templates were further excluded, resulting in a final subset of 72 articles. From this curated set, relevant parameters were extracted, resulting in 229 entries for the dataset. </p> <p>Characteristics about the structure (in terms of MOF type and configuration; Plasmonic element type and configuration), target analyte (including type, phase, and incubation time), measurement specifications (in terms of laser, laser power, exposure time), and performance of the MOF-based SERS substrates were collected.</p> <p>Listed references 1-72 correspond with the manuscript number in the dataset.</p> <p>Listed references 73-80 correspond with references for selected examples of MOF pore diameters.</p>
Data for Reducing leakage of single-qubit gates for superconducting quantum processors using analytical control pulse envelopes
<p>This dataset contains the experimental data used in the figures of the paper "Reducing leakage of single-qubit gates for superconducting quantum processors using analytical control pulse envelopes" by E. Hyyppä, A. Vepsäläinen, ..., and J. Heinsoo published in PRX Quantum 5, 030353 (2024): https://doi.org/10.1103/PRXQuantum.5.030353.</p> <p>The data is stored mostly as csv-files, the contents of which are explained in the readme-files. Each subfolder corresponds to one figure of the paper and also contains a Jupyter Notebook for plotting the data. The subfolders S1-S10 correspond to the supplementary figures, i.e., figures 6-15 in the Appendix of the paper.</p> <p>Furthermore, we provide a Jupyter notebook in the folder Code_to_plot_FAST_and_HD_DRAG_pulses/ that provides Python functions for evaluating and plotting the proposed FAST DRAG and HD DRAG pulses in time domain and frequency domain. Please cite our paper if you use the Python code for your published research.</p> <p>The notebooks have been tested using the following Python package versions<br>Python 3.11<br>scipy 1.14.1<br>numpy 2.1.0<br>matplotlib 3.9.2</p>
Dataset: An Analytic Hierarchy Process-Based Multicriteria Model for Component Selection in a Computational Numerical Control (CNC) Machine
<p><i><strong>"An Analytic Hierarchy Process-Based Multicriteria Model for Component Selection in a Computational Numerical Control (CNC) Machine"</strong></i></p><p><i>CHILECON 2023 - </i><a href="https://site.ieee.org/chilesur/ieee-chilecon-2023/"><i>https://site.ieee.org/chilesur/ieee-chilecon-2023/</i></a><i> </i></p><p>---</p><p>En el marco del trabajo de referencia, los autores ponemos a disposición de los lectores la base de datos utilizada para el proceso de toma de decisión multicriterio para la selección del software y del MCU de una maquina CNC. </p><p>En el repositorio podrán encontrar los datos referentes a los criterios, subcriterios, indicadores, datos, fuentes de los datos extraídos, política de decisión, cálculos de las evaluaciones de los modelos AHP aplicados y el análisis de sensibilidad de estos. Además, podrán encontrar las gráficas utilizadas en el estudio en la mejor calidad posible. </p><p>El material fue puesto a disposición de todos los interesados para fines académicos y científicos. </p><p>Atte. </p><p>Los autores. </p><p>---</p>
Synthetic time series data generation for edge analytics
<p>In this research, we create synthetic data with features that are like data from IoT devices. We use an existing air quality dataset that includes temperature and gas sensor measurements. This real-time dataset includes component values for the Air Quality Index (AQI) and ppm concentrations for various polluting gas concentrations. We build a JavaScript Object Notation (JSON) model to capture the distribution of variables and structure of this real dataset to generate the synthetic data. Based on the synthetic dataset and original dataset, we create a comparative predictive model. Analysis of synthetic dataset predictive model shows that it can be successfully used for edge analytics purposes, replacing real-world datasets. There is no significant difference between the real-world dataset compared the synthetic dataset. The generated synthetic data requires no modification to suit the edge computing requirements. The framework can generate correct synthetic datasets based on JSON schema attributes. The accuracy, precision, and recall values for the real and synthetic datasets indicate that the logistic regression model is capable of successfully classifying data</p>
AutoML for Video Analytics with Edge Computing - Dataset
<p>Latency and confidence measurements obtained from an edge-assisted object recognition system.</p> <p>The records are obtained by tuning the image encoding rate and Neural Network input layer size and measuring the latency on completing each system operation, i.e. encoding the image, transmiting it wirelessly, decoding and rotating at the server, and performing object recognition with YOLO on the server's GPU. Moreover, we document the achievable frame rate as a result of the total latency, as well as the object recognition confidence and cumulative confidence for all identified objects of each image.</p>
S95 | PFASANEXCH | PFAS List from the NORMAN PFAS Analytical Exchange Activity
<p>This is the collection associated with list S95 PFASANEXCH on the NORMAN Suspect List Exchange.</p> <p><a href="https://www.norman-network.com/nds/SLE/">https://www.norman-network.com/nds/SLE/</a></p> <p>This is a list from the <a href="https://www.norman-network.net/sites/default/files/files/QA-QC%20Issues/2021%20NORMAN%20network%20PFAS%20Analytical%20Exchange%20Final%20Report%2014022022.pdf">PFAS Analytical Exchange Activity</a>, part of NORMAN Joint Programme of Activities (JPA) 2021 coordinated by UK Environment Agency. This activity aimed to gain an understanding of the current analytical capability of PFAS as Limit of Detection (LOD) in participating international laboratories.</p>
SCG Dataset from Graph Neural Networks in Supply Chain Analytics and Optimization: Concepts, Perspectives, Dataset and Benchmarks
<p><strong>Abstract:</strong> Graph Neural Networks (GNNs) have recently gained traction in transportation, bioinformatics, language and image processing, but research on their application to supply chain management remains limited. Supply chains are inherently graph-like, making them ideal for GNN methodologies, which can optimize and solve complex problems. The barriers include a lack of proper conceptual foundations, familiarity with graph applications in SCM, and real-world benchmark datasets for GNN-based supply chain research. To address this, we discuss and connect supply chains with graph structures for effective GNN application, providing detailed formulations, examples, mathematical definitions, and task guidelines. Additionally, we present a multi-perspective real-world benchmark dataset from a leading FMCG company in Bangladesh, focusing on supply chain planning. We discuss various supply chain tasks using GNNs and benchmark several state-of-the-art models on homogeneous and heterogeneous graphs across six supply chain analytics tasks. Our analysis shows that GNN-based models consistently outperform statistical ML and other deep learning models by around 10-30% in regression, 10-30% in classification and detection tasks, and 15-40% in anomaly detection tasks on designated metrics. With this work, we lay the groundwork for solving supply chain problems using GNNs, supported by conceptual discussions, methodological insights, and a comprehensive dataset.</p>
Attributes: A Curriculum Analytics System for measuring learning outcomes - Overview
<p><span><strong>Link to video </strong><a href="https://vimeo.com/1015456231?share=copy#t=0"><strong>https://vimeo.com/1015456231?share=copy - t=0</strong></a><br><br>The Curriculum Analytics System at the Instituto Tecnológico de Costa Rica, integrated into TEC Digital, supports faculty, coordinators, and students in assessing engineering learning outcomes during accreditation processes. The system offers two key user modules: one for coordinators to map and manage learning outcomes, and another for instructors to conduct assessments through the course portal. Coordinators oversee course and attribute mapping using visual representations of study plans, control points, and outcome visualizations. Instructors configure assignments and evaluate student submissions with standardized rating scales. The system tracks progress in real-time and generates</span> <span>comprehensive reports with performance metrics, facilitating continuous improvement in academic programs.</span></p> <p><strong><span>Key words: </span></strong><span>attributes, learning outcomes, TEC Digital, curriculum analytics, continuous improvement. </span></p>
Data, Analytical Code, and Model Outputs From: Restoration Treatments Enhance Tree Growth and Alter Climatic Constraints During Extreme Drought
<p>This archive includes data (forest inventories, tree ring measurements, climate variables), statistical code, model outputs, and a preprint copy of Rodman et al. (2024). For more information on specific information, processing methods, and data formats, see "README.md" or "README.html" files associated with this archive</p>
Dataset for 'Room-temperature monitoring of CH4 and CO2 using a metal-organic framework-based QCM sensor showing inherent analyte discrimination'
<p>Associated data for the manuscript 'Room-temperature monitoring of CH4 and CO2 using a metal-organic framework-based QCM sensor showing inherent analyte discrimination' (doi://10.26434/chemrxiv-2023-djhp2)</p> <p> </p> <p> </p>
Bayesian Symbolic Learning to Build Analytical Correlations from Rigorous Process Simulations: Application to CO2 Capture Technologies
<p>Dataset of process simulations results of the natural gas sweetening and flue gas treatment (first and second sheet, respectively as indicated by the sheet name in the .xlsx file). The dataset refers to the publication <em>Bayesian Symbolic Learning to Build Analytical Correlations from Rigorous Process Simulations: Application to CO<sub>2</sub> Capture Technologies </em>by V. Negri, Vàzquey D., Sales-Pardo, Marta, Guimerà, R. and Guillén-Gosàlbez, G. The training and testing dataset are used to generate the figures in the main manuscript and supplementary information. </p> <p> </p>
Mobility analytic results
<p>The dataset contains information about the trips made by the participants with the MyCorridor app within the context of the second iteration phase in the MyCorridor project. The collected information is related to certain trip characteristics as trip length, trip distance, number of transfers and distribution of service clusters. Moreover, the data shows the number of users and trips.</p>
A scoping review on bovine tuberculosis highlights the need for novel data streams and analytical approaches to curb zoonotic diseases
<p>The following data and scripts are part of the manuscript titled 'A scoping review on bovine tuberculosis highlights the need for novel data streams and analytical approaches to curb zoonotic diseases' which is currently going through the peer-review process and has already been published as a preprint. Please read the README.txt file for information on the files uploaded.</p>
Fetal exposure to the Ukraine famine of 1932-1933 and adult Type 2 Diabetes Mellitus (Public data and analytical code)
<p><strong>Abstract</strong></p> <p>The short-term impact of famines on death and disease is well documented but it is difficult to estimate their potential long-term impact. We used the setting of the man-made Ukrainian Holodomor famine of 1932-1933 to examine the relationship between prenatal famine and adult Type 2 diabetes mellitus (T2DM). This ecological study included 128,225 T2DM cases diagnosed between 2000-2008 among 10,186,016 male and female Ukrainians born between 1930 and 1938. Individuals who were born in the first half-year of 1934, and hence exposed in early gestation to the mid-1933 peak famine period, had a larger than two-fold likelihood of T2DM (OR 2.21; 95% CI 2.00-2.45) compared to unexposed controls. There was a dose-response relationship between severity of famine exposure and adult T2DM risk comparing individuals born in regions with severe, very severe, and extreme famine to births in the no-famine region.</p> <p> </p> <p><strong>Description of the data and analytical code</strong></p> <p>In exploratory analyses we first examined whether the odds for T2DM were elevated for any month of birth in the period January 1930 to December 1938 in any of the four regions of varying famine intensity. This was achieved by comparing, within each region, the T2DM odds for births in any month and year of birth relative to the T2DM odds for births in the same month combining all other years of birth. The analysis served to identify potential relations of famine with specific months and years of birth, controlling for month of birth effects. We observed increased T2DM odds ratios for births between January and June 1934 in famine-exposed oblasts, with smaller increases for births in 1935 and 1936 in these months. Our findings suggested that in multivariate modelling statistical control for month of birth effects could be accomplished by adjusting for the January-June period. Our findings are presented in the data file '01 Odds Ratio for T2DM Over Time' and show the odds ratios (ORs) for Type 2 Diabetes Mellitus (T2DM) comparing the region-specific T2DM odds for each birth year and month relative to births in the same months but combining all other years of birth. The R syntax file '01 Odds of T2DM Over Time Figure' provides the code necessary to reproduce the figure.</p> <p> </p> <p>For confirmatory analyses we employed a Difference-in-Differences approach to quantify associations between prenatal exposure to famine and T2DM, taking into account year of birth, half-year of birth (Jan-Jun vs Jul-Dec), region, and their interactions. This analysis was conducted initially for each gender separately and then for both genders combined, adjusting for We carried out sensitivity analyses to assess potential changes in T2DM odds arising from the use of pre-famine births vs post-famine births as controls. Our findings are presented in the data file '02 Ukraine Famine 1932-33 Main Data'. Information on the number of T2DM cases by gender, region of residence, and year and month of birth 1930-1938 in Ukraine was collected by the national Ukraine Diabetes Register (Komisarenko Institute of Endocrinology and Metabolism, Kyiv) between 2000-2008. The number of births in the same subgroups, representing the populations at risk for T2DM, was estimated by demographic population reconstruction methods as reported in the publication. We classified the birth counts by year of birth, the semi-annual birth period (January-June vs. July-December), region of birth, and gender. The SPSS syntax file titled '02 Ukraine Famine 1932-33 Main Analysis' provides the code to replicate our main findings as presented in the publication.</p> <p> </p> <p>In a separate analysis we visualized by a meta-regression approach the relation between famine intensity at the oblast level in 1933 and the odds for adult T2DM. The data required for the replication of our findings are included in the file '03 Odds Ratio for T2DM and Famine Intensity at Oblast Level'. The R syntax file titled '03 Ukraine Famine 1932-33 Meta-regression' provides details on conducting the meta-regression using the R package ‘metafor’.</p> <p> </p> <p><strong>Funding</strong></p> <p>Ukraine State complex program Diabetes Mellitus, project number 0106U000844 (M.K.). Holodomor Research and Education Consortium in Canada (L.H.L., O.W.). NIDI-NIAS Fellowship of the Royal Netherlands Academy of Sciences (L.H.L.). National Institute of Aging R01 AG028593 (L.H.L.). National Institute of Aging R01 AG06687 (L.H.L.).</p> <p> </p> <p><strong>Sharing/Access information</strong></p> <p>Data sharing and use are unrestricted with acknowledgement of the original publication and listing of the funding sources as per the above. Researchers are encouraged to contact the Principal Investigators (PIs) for consultations on data structure and use as needed (L.H. Lumey, <a href="mailto:lumey@columbia.edu">lumey@columbia.edu</a>; Oleh Wolowyna, <a href="mailto:olehw@aol.com">olehw@aol.com</a>).</p>
Antisemitism on Twitter: A Dataset for Machine Learning and Text Analytics
<h1><strong><span><span>Dataset from the Institute for the Study of Contemporary Antisemitism (ISCA) at Indiana University: </span></span></strong></h1> <p> </p> <div> <div> <p><span><span>The </span><span>Social Media</span><span> & Hate research lab at the Institute for the Study of Contemporary Antisemitism compiled this dataset using an annotation portal (Jikeli, Soemer, and Karali 2024), which was used to label tweets as either antisemitic or non-antisemitic, among other labels. Note that annotation was done on live data, including images and context, such as threads. All data was annotated by two experts, and all discrepancies were discussed</span><span> (Jikeli et al. 2023)</span><span>.</span></span><span> </span></p> </div> </div> <p><br><strong>Content: </strong></p> <p><span><span>This dataset </span><span>contains</span> <span>1</span><span>1</span><span>311</span><span> tweets </span><span>covering</span><span> a wide range of topics common in conversations about Jews, Israel, and antisemitism between January 2019 and </span><span>April 2023</span><span>. </span><span>The dataset consists of random samples of relevant keywords during this </span><span>time period</span><span>.</span><span> 1,</span><span>953</span><span> tweets (1</span><span>7</span><span>%) </span><span>are antisemitic </span><span>according to </span><span>the IHRA definition of antisemitism.</span><span> </span></span><span> </span></p> <div> <p><span><span>The distribution of tweets by year is as follows:</span><span> 1499 (</span><span>13</span><span>%) from 2019, 371</span><span>2</span><span> (</span><span>33</span><span>%) from 2020, </span><span>2591</span><span> (2</span><span>3</span><span>%) from 2021</span><span>, 2644 from 2022 </span><span>(23%)</span> <span>and 865 </span><span>(8%)</span> <span>f</span><span>rom 2023</span><span>. </span><span>6365</span><span> (</span><span>56</span><span>%) </span><span>contain</span><span> the keyword "Jews,"</span><span> 4134 </span><span>(</span><span>3</span><span>7</span><span>%) include "Israel," 529 (</span><span>5</span><span>%) feature the derogatory term "</span><span>ZioNazi</span><span>*," and 283 (</span><span>3</span><span>%) use the slur "K---s." Some tweets may </span><span>contain</span><span> multiple keywords. </span></span><span> </span></p> </div> <div> <p><span><span>725</span><span> out of the </span><span>6365</span><span> tweets with the keyword "Jews" (11%) and </span><span>664</span><span> out of the </span><span>4134</span><span> tweets with the keyword "Israel" (1</span><span>6</span><span>%) were classified as antisemitic. 97 out of the 283 tweets using the antisemitic slur "K---s" (34%) are antisemitic.</span> <span>Interestingly, many tweets featuring the slur "K---s" actually </span><span>call out</span><span> its u</span><span>s</span><span>e.</span><span> In contrast, </span><span>the majority of</span><span> tweets </span><span>using</span><span> the derogatory term "</span><span>ZioNazi</span><span>*" are antisemitic, with 467 out of 529 (88%) being classified as such. </span></span><span> </span></p> </div> <p> </p> <p><strong>File Description: </strong></p> <div> <div> <p><span><span>The dataset is provided in a csv file format, with each row </span><span>representing</span><span> a single message, including replies, quotes, and retweets. The file </span><span>contains</span><span> the following columns: </span></span><span> </span></p> </div> <div> <p><span><span> </span></span><span><span> </span><br></span><span><span>‘ID’:</span> <span>Represents</span><span> the tweet ID. </span></span><span> </span></p> </div> <div> <p><span><span>‘Username’: </span><span>Represents</span><span> the username </span><span>that posted </span><span>the tweet</span><span>. </span></span><span> </span></p> </div> <div> <p><span><span>‘Text’: </span><span>Represents</span><span> the full text of the tweet (not pre-processed).</span></span><span> </span></p> </div> <div> <p><span><span>‘</span><span>CreateDate</span><span>’: </span><span>Represents</span><span> the date </span><span>on which </span><span>the tweet was created</span><span>. </span></span><span> </span></p> </div> <div> <p><span><span>‘Biased’: </span><span>Represents</span><span> the label given by our annotations as to whether the tweet </span><span>is antisemitic or no</span><span>t</span><span>.</span></span><span> </span></p> </div> <div> <p><span><span>‘Keyword’: </span><span>Represents</span><span> the keyword that was used in the query. The keyword can be in the text, including </span><span>hashtags, </span><span>mentioned </span><span>users</span><span>, or the username</span><span> itself.</span><span> </span></span><span> </span></p> </div> </div> <p> </p> <p>Licences </p> <p>Data is published under the terms of the "Creative Commons Attribution 4.0 International" licence (https://creativecommons.org/licenses/by/4.0) </p> <p> </p> <p>Acknowledgements </p> <p>We are grateful for the support of Indiana University’s Observatory on Social Media (OSoMe) (Davis et al. 2016) and the contributions and annotations of all team members in our Social Media & Hate Research Lab at Indiana University’s Institute for the Study of Contemporary Antisemitism, especially Grace Bland, Elisha S. Breton, Kathryn Cooper, Robin Forstenhäusler, Sophie von Máriássy, Mabel Poindexter, Jenna Solomon, Clara Schilling, and Victor Tschiskale. </p> <p>This work used Jetstream2 at Indiana University through allocation HUM200003 from the Advanced Cyberinfrastructure Coordination Ecosystem: Services & Support (ACCESS) program, which is supported by National Science Foundation grants #2138259, #2138286, #2138307, #2137603, and #2138296.</p>
Web requests analysis of Italy websites which use Google Analytics
<p>List of 504,038 domains of Italy found to contain Google Analytics.</p> <p>The front page for Italy-related domain names has been accessed through HTTPS or HTTP and analysed with webbkoll and jq to gather data about third-party requests, cookies and other privacy-invasive features. Together with the actual URL visited, the user/property ID is provided for 495,663 domains (extracted either from the cookies deposited or the URL of requests to Google Analytics). MX and TXT records for the domains are also provided.</p> <p>The most common ID found was 23LNSPS7Q6, with over 35k domains calling it (seemingly associated with italiaonline.it). The most common responding IP addresses were 3 AWS IPv4 addresses (over 40k domains) and 2 CloudFlare IPv6 addresses (over 12k domains).</p>
Data associated to: Analytical Physical Model for Organic Metal-Electrolyte-Semiconductor Capacitors
<p>Data associated to the manuscript entitled: Analytical Physical Model for Organic Metal-Electrolyte-Semiconductor Capacitors by Larissa Huetter, Adrica Kyndiah and Gabriel Gomila</p>
Dataset: Analytical Physical Model for Electrolyte Gated Organic Field Effect Transistors in the Helmholtz Approximation
<p>Data corresponding to the figures of the manuscript "Analytical Physical Model for Electrolyte Gated Organic Field Effect Transistors in the Helmholtz Approximation" by Larissa Huetter, Adrica Kyndiah and Gabriel Gomila</p>
myExperiment Workflows, "abstracted" (all non-analytical nodes removed)
<p>To do the SCOFF analysis (detecting highly similar workflow fragments) we took all bioinformatics-related workflows from myExperiment and removed all non-analytical nodes. These included nodes referred-to as "shims" - those that do data structure/type transformations, but not any "semantic" transformation. This deposit contains all such abstracted workflows.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.