Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,782
datasets available to search
ShareScore release 0.7.1
Dataset results
1,782 results for “algorithms”
Ergebnisse der Evaluation - Implementing Layout Algorithms Using Cytoscape in ExplorViz
<p>Diese Datei enthält die Ergebnisse der 20 Teilnehmer der Evaluation, die für die Bachelorarbeit durchgeführt wurde</p>
MLuqe algorithm outputs for two unplanned hydroacoustic transects in Zóñar Lake nature reserve (southern Iberian Peninsula).
<p>The outputs generated by the MLuqe algorithm (Gutiérrez-Estrada and Pulido Calvo, 2024; DOI: 10.5281/zenodo.13919472) are organised in two Excel files, one for each hydroacoustic transect. Each file is structured as follows:<br>1. Each file contains a total of 15 sheets, with each sheet corresponding to a set of theoretical fish density (FDS).<br>2. Column A presents details of each simulation. Cell A3 indicates the minimum Hamming similarity between the real and simulated data. Cells A5 and A7 show the simulation configuration and repetition with the minimum Euclidean distance. Cell A11 displays the maximum target density.<br>3. Each cell represents the Euclidean distance between the actual densities observed in the echograms for each transect and the densities simulated by MLuqe.</p> <p>Further details of the MLuqe algorithm and the results obtained can be found in the paper entitled 'Abundance estimation of introduced Goldfish in a protected lake from unplanned hydroacoustic transects' by Gutiérrez-Estrada, J.C., Pulido-Calvo, I., de la Cruz, J., Sánchez-Polaina, F.</p>
Dataset used in "Deamidation analysis of Therapeutic Drugs using Matrix-assisted Laser Desorption Ionization Mass Spectrometry and A Novel Algorithm QuanDA"
<p><span>Overview</span></p> <p><span>Theoretical MALDI spectra of peptides, both before and after deamidation, are crucial for the analysis of real samples. QuanDA is a method utilizing non-negative least squares to calculate the percentage of deamidation during the aging process of standard peptides and therapeutic drugs.</span></p> <p><span>Dataset Description</span></p> <p><span>The `data` directory contains example datasets specifically designed for analyzing peptide deamidation:</span></p> <p><span>D.txt: Simulated spectra of the pure synthetic peptide Pep-D.</span></p> <p><span>N.txt: Simulated spectra of the pure synthetic peptide Pep-N.</span></p> <p><span>Asn2Asp18: Contains 10 measurements from a sample comprising 10% Pep-N and 90% Pep-D.</span></p> <p><span>Asn18Asp2: Comprises 10 measurements from a sample with 90% Pep-N and 10% Pep-D.</span></p> <p><span>Peptide Sequences: Pep-N: YTHQGLSSPVTKSFNRGE; Pep-D: YTHQGLSSPVTKSFDRGE</span></p> <p><span>These synthetic peptides are utilized to investigate deamidation behavior.</span></p> <p><span>Somatostatin Dataset: Contains data from actual samples analyzed over 0-13 days under physiological conditions (pH 7.4, 37°C) and storage conditions (pH 4, 4°C). This dataset includes theoretical values before and after deamidation.</span></p>
Parareal Algorithm Cycles Illustration
<p>Please refer to journal paper from S.J.P.Pamela, titled</p> <p>"Neural-Parareal: Self-improving acceleration of fusion MHD simulations using time-parallelisation and neural operators"</p> <p>Available on ArXiV and on Comp.Phys.Comm.: https://doi.org/10.1016/j.cpc.2024.109391</p>
Using AI Algorithms for Predictive Analysis in Personalized Medicine
<p><strong><span>This study explored the factors influencing patients' willingness to adopt AI-powered personalized medicine. This research found the problems. Integrating AI and personalized medicine has the potential to revolutionize healthcare. However, public trust in AI for healthcare applications remains a challenge. This research examines the factors determining people's views toward using artificial intelligence for predictive analytics in personalized medicine. A cross-sectional design was employed through a survey distributed via Google Forms in April 2024 using purposive sampling. The target respondents included residents of the Jabodetabek area (Jakarta, Bogor, Depok, Tangerang, Bekasi- cities in Indonesia) with prior experience seeking medical consultation or checkups. A total of 267 responses were collected after removing outliers. The study used a Partial Least Squares Structural Equation Modeling (PLS-SEM) approach to analyze the data. </span></strong><strong><span>The study considered six independent variables: AI knowledge, trust in AI, attitude towards data privacy, personalized medicine expectations, personalized medicine understanding, and perceived risk of discrimination in AI. The dependent variable was the intention to use AI in personalized medicine. It found five of six hypotheses have significant impact. </span></strong></p>
The data of the mesh used in: Pan M, Zou R, Jüttler B. Algorithms and Data Structures for Cs-smooth RMB-splines of Degree 2s+ 1. Computer Aided Geometric Design, 2024: 102389.
Open the record for dataset details and reuse information.
Qiskit-Algorithms
<p>Code with auxiliary scripts for data treatment as well as excel file with test classification for the paper "..."</p>
Data from: Wide range screening of algorithmic bias in word embedding models using large sentiment lexicons reveals underreported bias types
<p>Concerns about gender bias in word embedding models have captured substantial attention in the algorithmic bias research literature. Other bias types however have received lesser amounts of scrutiny. This work describes a large-scale analysis of sentiment associations in popular word embedding models along the lines of gender and ethnicity but also along the less frequently studied dimensions of socioeconomic status, age, physical appearance, sexual orientation, religious sentiment and political leanings. Consistent with previous scholarly literature, this work has found systemic bias against given names popular among African-Americans in most embedding models examined. Gender bias in embedding models however appears to be multifaceted and often reversed in polarity to what has been regularly reported. Interestingly, using the common operationalization of the term bias in the fairness literature, novel types of so far unreported bias types in word embedding models have also been identified. Specifically, the popular embedding models analyzed here display negative biases against middle and working-class socioeconomic status, male children, senior citizens, plain physical appearance and intellectual phenomena such as Islamic religious faith, non-religiosity and conservative political orientation. Reasons for the paradoxical underreporting of these bias types in the relevant literature are probably manifold but widely held blind spots when searching for algorithmic bias and a lack of widespread technical jargon to unambiguously describe a variety of algorithmic associations could conceivably be playing a role. The causal origins for the multiplicity of loaded associations attached to distinct demographic groups within embedding models are often unclear but the heterogeneity of said associations and their potential multifactorial roots raises doubts about the validity of grouping them all under the umbrella term bias. Richer and more fine-grained terminology as well as a more comprehensive exploration of the bias landscape could help the fairness epistemic community to characterize and neutralize algorithmic discrimination more efficiently.</p>
Beating classical heuristics for the binary paint shop problem with the quantum approximate optimization algorithm
<p>Data repository for the paper "Beating classical heuristics for the binary paint shop problem with the quantum" by Michael Streif, Sheir Yarkoni, Andrea Skolik, Florian Neukart, and Martin Leib.</p> <p> </p>
Comparison and assessment of different object-based classifications using machine learning algorithms and UAVs multispectral imagery in the framework of precision agriculture
<p>Supplementary material of the paper</p>
Velocity, density and energy budget statistics from the article: Validation and application of the lattice Boltzmann algorithm for a turbulent immiscible Rayleigh-Taylor system
<p>We develop a multicomponent lattice Boltzmann (LB) model for the 2D Rayleigh--Taylor turbulence with a Shan-Chen pseudopotential implemented on GPUs. Accuracy of the LB model is tested both for early and late stages of instability. For the developed turbulent motion we analyze the balance between different terms describing variations of the kinetic and potential energies. Then, we analyze the role of interface in the energy balance, and also the effects of the vorticity induced by the interface in the energy dissipation. Statistical properties are compared for miscible and immiscible flows.</p> <p>In this Dataset, we have included the files for the statistics of the energy budget for the miscible and immiscible Rayleigh-Taylor flows studied in the referred article.</p> <p>We also included the examples of the density and velocity fields showed in the Figures 1 and 2 of the article.</p>
Clinical Categorization Algorithm (Clical) and Machine-Learning Approach (Srf-clical) to Predict Clinical Benefit to Immunotherapy in Metastatic Melanoma Patients: Real-world Evidence from Istituto Nazionale Tumori Irccs Fondazione Pascale, Napoli, Italy.
<p>Raw-data related to a manuscript submitted to "Cancers" journal - MDPI - https://www.mdpi.com/journal/cancers</p> <p><strong>Manuscript Title</strong>: Clinical Categorization Algorithm (Clical) and Machine-Learning Approach (Srf-clical) to Predict Clinical Benefit to Immunotherapy in Metastatic Melanoma Patients: Real-world Evidence from Istituto Nazionale Tumori Irccs Fondazione Pascale, Napoli, Italy.</p> <p><strong>Authors:</strong> Gabriele Madonna1,#, Giuseppe V. Masucci2,3,#, Mariaelena Capone1, Domenico Mallardo1, Antonio Maria Grimaldi1, Ester Simeone1, Vito Vanella1, Lucia Festino1, Marco Palla1, Luigi Scarpato1, Marilena Tuffanelli1, Grazia D’angelo1, Lisa Villabona2, Isabelle Krakowski2,4, Hanna Eriksson2,3, Felipe Simao5, Rolf Lewensohn2,3, Paolo Antonio Ascierto1,+</p> <p><strong>Affiliations</strong>:</p> <p>1 Cancer Immunotherapy and Development Therapeutics Unit, Istituto Nazionale Tumori IRCCS Fondazione "G. Pascale", Napoli, Italy</p> <p>2 Theme Cancer, Karolinska University Hospital, Stockholm, Sweden</p> <p>3 Department of Oncology-Pathology, Karolinska Institutet, Stockholm, Sweden</p> <p>4 Theme Inflammation, Karolinska University Hospital Stockholm, Sweden</p> <p>5 Genevia technologies OY, Tampere, Finland</p> <p># these authors equally contributed</p> <p>+ Corresponding author</p> <p><strong>Abstract of submitted Manuscript:</strong> The real-life application of immune checkpoint inhibitors (ICI) may yield different outcomes compared to the benefit presented in clinical trials. For this reason, there is a need to define the group of patients that may benefit from treatment. We retrospectively investigated 578 metastatic melanoma patients treated with ICI at Istituto Nazionale Tumori IRCCS Fondazione “G. Pascale” of Napoli Italy (INT-NA). To compare patients’ clinical variables (age, Lactate Dehydrogenase (LDH), Neutrophil-Lymphocyte Ratio (NLR), eosinophil, BRAF status, previous treatment) and their predictive and prognostic power in a comprehensive non-hierarchical way, a Clinical Categorization Algorithm (CLICAL) was defined and validated by the application of machine learning, Survival Random Forest (SRF-CLICAL). The comprehensive analysis of the clinical parameters by log risk-based algorithms convened into predictive signatures that could identify groups of patients with great benefit or not, regardless of the ICI received. From a real-life retrospective analysis of metastatic melanoma patients, we generated and validated an algorithm based on machine learning that could assist with the clinical decision of whether or not to apply ICI therapy by defining five signatures of predictability with a 95% accuracy.</p> <p><strong>Funding: </strong>This research was funded by Italian Ministry of Health (IT-MOH) through “Ricerca Corrente”, grants number M2-2. Additional funding [N#184093) from the Stockholm Cancer Society and King Gustav V’s Jubilee foundation Stockholm.</p>
Sling load - Dataset of software tests on column detection algorithm with laserscanner data
<p>This dataset contains data and scripts for plotting results of testing of the algorithm of column detection, using laserscanner data.</p>
Step-by-step simulation of the ROUTR algorithm on a road-network graph.
<p>This video simulates, step-by-step, the application of the ROUTR algorithm on a small road-network graph.</p>
Evaluation of the International Society for Cutaneous Lymphoma Algorithm for the Diagnosis of Early Mycosis Fungoides
<p><strong>Supplementary Table 1.</strong> Clinicopathologic features and ISCL scores of cases included in this study</p> <p><strong>Supplementary Table 2.</strong> Curvilinear coordinates of ROC, sensitivity, specificity, and Youden’s index for CD2, CD3, CD5, and CD7 expression</p>
Figure 5. Agreement subtree cladogram obtained with the Ratchet algorithm for parsimonious analyses using the complete morphological matrix without gamete-related characters. Values above branches are bootstrap supports after 1000 in High level of phenotypic homoplasy amongst eutardigrades (Tardigrada) based on morphological and total evidence phylogenetic analyses
Figure 5. Agreement subtree cladogram obtained with the Ratchet algorithm for parsimonious analyses using the complete morphological matrix without gamete-related characters. Values above branches are bootstrap supports after 1000 replicates; values under branches are Bremer relative supports.
Extension of Towards Large Scale Automated Algorithm Design by Integrating Modular Benchmarking Frameworks
<p>This provides the dataset and the corresponding parsing scripts that extend the "Towards Large Scale Automated Algorithm Design by Integrating Modular Benchmarking Frameworks" study.</p>
Synthetic data set to evaluate and benchmark the performance of multiple linear regression algorithms in Scikit-Learn and SANElib
<p>The datasets respresent different numbers of columns and rows to measure the scalability of linear regression algorihms in terms of columns and rows.</p>
Multiple Sclerosis lesions detection by a hybrid Watershed-Clustering algorithm
<p>Computer Aided Diagnosis (CAD) systems have been developing in the last years with the aim of helping the diagnosis and monitoring of several diseases. We present a novel CAD system based on a hybrid Watershed-Clustering algorithm for the detection of lesions in Multiple Sclerosis. Magnetic Resonance Imaging scans (FLAIR sequences without gadolinium) of 20 patients affected by Multiple Sclerosis with hyperintense lesions were studied. The CAD system consisted of the following automated processing steps: images recording, automated segmentation based on the Watershed algorithm, detection of lesions, extraction of both dynamic and morphological features, and classification of lesions by Cluster Analysis. The investigation was performed on 316 suspect regions including 255 lesion and 61 non-lesion cases. The Receiver Operating Characteristic analysis revealed a highly significant difference between lesions and non-lesions; the diagnostic accuracy was 87% (95% CI: 0.83–0.90), with an appropriate cut-off of 192.8; the sensitivity was 77% and the specificity was 87%. In conclusion, we developed a CAD system by using a modified algorithm for automated image segmentation which may discriminate MS lesions from non-lesions. The proposed method generates a detection out-put that may be support the clinical evaluation.</p>
MIRA-Datasets: Datasets from Metrics for Intercomparison of Remapping Algorithms
<p>The Metrics for Intercomparison of Remapping Algorithms (<a href="https://github.com/CANGA/MIRA">MIRA</a>) project provides the Python drivers for the intercomparison study to enable the computation of metrics for different remapping algorithms of interest in ESM.</p> <p>The dataset repository contains three groups of artifacts: the original test cases used in the study and the output metrics data from four different remapping algorithms, along with some helpful scripts to compare the metrics data. Details are provided below.</p> <ol> <li> <p>All of the input meshes, sampled reference data on the meshes for several uniformly refined resolutions, and regionally refined cases are contained within the <code>Meshes</code> directory.</p> <ul> <li>The uniformly refined meshes for Cubed-Sphere (CS), polygonal quasi-uniform MPAS and Regular Latitude-Longitude (RLL) meshes along with sampled field data for five different fields are provided in <code>Meshes/UniformlyRefined/</code> directory.</li> <li>The regionally refined meshes for CS and MPAS meshes around continental-US (CONUS) region with the sampled reference field data is available under <code>Meshes/RegionallyRefined</code> directory.</li> </ul> </li> <li> <p>The input meshes provided under <code>Meshes</code> directory were used to perform a remapping intercomparison study that analyzed the key numerical metrics to gain better insight into the behavior of remapping algorithms, and to compare several key properties under a unified framework. Four different remapping algorithms were considered in this study.</p> <ul> <li> <p>Earth System Modeling Framework (ESMF) Regrid</p> </li> <li> <p>TempestRemap high-order conservative maps</p> </li> <li> <p>Generalized Moving-Least-Squares (GMLS) algorithm</p> <ul> <li>A variation with the Clip-And-Assured-Sum (CAAS) algorithm to enforce bounds preservation</li> </ul> </li> <li> <p>Weighted-Least-Squares Essentially Non-oscillatory Remap (WLS-ENOR) scheme</p> <p>The metrics data collected for each of the cases and remapping algorithms are stored under the <code>MetricsData</code> directory. The metrics CSV files include details about:</p> <ul> <li>Error convergence data in global norms $L_1, L_2, L_{\inf}, H_1$ and $\left|H_1\right|$</li> <li>Global bounds preservation for determining monotonicity</li> <li>Local feature preservation through repeated remapping cycles</li> <li>Grid independence by using test cases with different mesh types and (uniformly refined/regionally refined) resolutions</li> </ul> </li> </ul> </li> <li> <p>A set of helpful Python scripts have also been provided to easily compare different aspects of the metrics data to gain more insight into the behavior of the remapping algorithms. These are under <code>Scripts</code> directory.</p> </li> </ol>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.