Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,782

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,782 results for “Algorithm”

Learn how ShareScore rates datasets ↗
nasa20/100

Scalable Distributed Change Detection from Astronomy Data Streams using Local, Asynchronous Eigen Monitoring Algorithms

This paper considers the problem of change detection using local distributed eigen monitoring algorithms for next generation of astronomy petascale data pipelines such as the Large Synoptic Survey Telescopes (LSST). This telescope will take repeat images of the night sky every 20 seconds, thereby generating 30 terabytes of calibrated imagery every night that will need to be coanalyzed with other astronomical data stored at different locations around the world. Change point detection and event classification in such data sets may provide useful insights to unique astronomical phenomenon displaying astrophysically significant variations: quasars, supernovae, variable stars, and potentially hazardous asteroids. However, performing such data mining tasks is a challenging problem for such high-throughput distributed data streams. In this paper we propose a highly scalable and distributed asynchronous algorithm for monitoring the principal components (PC) of such dynamic data streams. We demonstrate the algorithm on a large set of distributed astronomical data to accomplish well-known astronomy tasks such as measuring variations in the fundamental plane of galaxy parameters. The proposed algorithm is provably correct (i.e. converges to the correct PCs without centralizing any data) and can seamlessly handle changes to the data or the network. Real experiments performed on Sloan Digital Sky Survey (SDSS) catalogue data show the effectiveness of the algorithm.

restrictednotspecifiedMar 2025View details →
nasa20/100

Discovering Anomalous Aviation Safety Events Using Scalable Data Mining Algorithms

The worldwide civilian aviation system is one of the most complex dynamical systems created. Most modern commercial aircraft have onboard flight data recorders that record several hundred discrete and continuous parameters at approximately 1Hz for the entire duration of the flight. These data contain information about the flight control systems, actuators, engines, landing gear, avionics, and pilot commands. In this paper, recent advances in the development of a novel knowledge discovery process consisting of a suite of data mining techniques for identifying precursors to aviation safety incidents are discussed. The data mining techniques include scalable multiple-kernel learning for large-scale distributed anomaly detection. A novel multivariate time-series search algorithm is used to search for signatures of discovered anomalies on massive datasets. The process can identify operationally significant events due to environmental, mechanical, and human factors issues in the high-dimensional flight operations quality assurance data. All discovered anomalies are validated by a team of independent domain experts. This novel automated knowledge discovery process is aimed at complementing the state-of-the-art human-generated exceedance-based analysis that fails to discover previously unknown aviation safety incidents. In this paper, the discovery pipeline, the methods used, and some of the significant anomalies detected on real-world commercial aviation data are discussed.

restrictednotspecifiedMar 2025View details →
nasa20/100

An Efficient Local Algorithm for Distributed Multivariate Regression

This paper offers a local distributed algorithm for multivariate regression in large peer-to-peer environments. The algorithm is designed for distributed inferencing, data compaction, data modeling and classification tasks in many emerging peer-to-peer applications for bioinformatics, astronomy, social networking, sensor networks and web mining. Computing a global regression model from data available at the different peer-nodes using a traditional centralized algorithm for regression can be very costly and impractical because of the large number of data sources, the asynchronous nature of the peer-to-peer networks, and dynamic nature of the data/network. This paper proposes a two-step approach to deal with this problem. First, it offers an efficient local distributed algorithm that monitors the “quality ” of the current regression model. If the model is outdated, it uses this algorithm as a feedback mechanism for rebuilding the model. The local nature of the monitoring algorithm guarantees low monitoring cost. Experimental results presented in this paper strongly support the theoretical claims.

restrictednotspecifiedMar 2025View details →
nasa20/100

Spectral Decomposition Algorithm (SDA)

Spectral Decomposition Algorithm (SDA) is an unsupervised feature extraction technique similar to PCA that was developed to better distinguish spectral features in the space shuttle main engine's optical plume. See paper below: Code is not open sourced and therefore it is not available. See paper for sample pseudo code.

restrictednotspecifiedMar 2025View details →
nasa20/100

AMSR-E/Aqua L2B Global Swath Ocean Products derived from Wentz Algorithm V002

This daily Level-2B swath data set includes Sea Surface Temperature (SST), Near-Surface Wind Speed, Columnar Water Vapor, and Cloud liquid Water data arrays, and was used as input to generate the following daily, weekly, and monthly Level-3 gridded ocean products; AE_DyOcn, AE_WkOcn, and AE_MoOcn.

restrictednotspecifiedApr 2025View details →
geo16/100

The significance of different algorithms in transcriptome analysis of leukemic cells with rearranged MLL (RNA-seq)

GEO Series GSE54641. Homo sapiens. 4 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenFeb 2017View details →
geo16/100

SymMap database and TMNP algorithm reveal Huanggui Tongqiao granules for Allergic rhinitis through IFN-mediated neuroimmuno-modulation

GEO Series GSE212198. Mus musculus. 6 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenAug 2022View details →
zenodo16/100

Dataset related to article "An individualized algorithm to predict mortality in COVID-19 pneumonia: a machine learning based study "

<p>This record contains raw data related to article &ldquo;An individualized algorithm to predict mortality in COVID-19 pneumonia: a machine learning based study&quot;</p> <p>Abstract:</p> <p><strong>Introduction: </strong> Identifying SARS-CoV-2 patients at higher risk of mortality is crucial in the management of a pandemic. Artificial intelligence techniques allow one to analyze large amounts of data to find hidden patterns. We aimed to develop and validate a mortality score at admission for COVID-19 based on high-level machine learning.</p> <p><strong>Material and methods: </strong> We conducted a retrospective cohort study on hospitalized adult COVID-19 patients between March and December 2020. The primary outcome was in-hospital mortality. A machine learning approach based on vital parameters, laboratory values and demographic features was applied to develop different models. Then, a feature importance analysis was performed to reduce the number of variables included in the model, to develop a risk score with good overall performance, that was finally evaluated in terms of discrimination and calibration capabilities. All results underwent cross-validation.</p> <p><strong>Results: </strong> 1,135 consecutive patients (median age 70 years, 64% male) were enrolled, 48 patients were excluded, and the cohort was randomly divided into training (760) and test (327) groups. During hospitalization, 251 (22%) patients died. After feature selection, the best performing classifier was random forest (AUC 0.88 &plusmn;0.03). Based on the relative importance of each variable, a pragmatic score was developed, showing good performances (AUC 0.85 &plusmn;0.025), and three levels were defined that correlated well with in-hospital mortality.</p> <p><strong>Conclusions: </strong> Machine learning techniques were applied in order to develop an accurate in-hospital mortality risk score for COVID-19 based on ten variables. The application of the proposed score has utility in clinical settings to guide the management and prognostication of COVID-19 patients.</p>

restrictedOct 2022View details →
zenodo16/100

Surface precipitation type based on new decision algorithm about 131 precipitation events in 5 ICE-POP 2018 sites

<p>The documents are decision results (+simple event description) of surface precipitation type based on new decision algorithm about 131 precipitation events in 5 ICE-POP 2018 sites.</p> <p>Simple event description include PARSIVEL V-D scatterplots and rawinsonde profiles.</p> <p>Please see Description_plot.pdf.</p> <p>and other pdf files show figures relating events belong to each types (RA[Rain], SN[Snow], RASN).</p>

restrictedcc-by-4.0Aug 2024View details →
zenodo16/100

STAVER: A Standardized Dataset-Based Algorithm for Efficient Variation Reduction in Large-Scale DIA MS Data

<p>This project focuses on the development and application of STAVER, a standardized dataset-based algorithm designed to reduce variation in large-scale data-independent acquisition (DIA) mass spectrometry data. By using a reference dataset to standardize mass spectrometry signals, STAVER effectively reduces noise and enhances protein quantification accuracy, especially in the context of multi-library search. The effectiveness of STAVER is demonstrated in several large-scale DIA datasets, showing improved identification and quantification of thousands of proteins. The project aims to promote the adoption of multi-library search and improve the quality of DIA proteomics data through the open-source STAVER software package.</p>

restrictedApr 2023View details →
ClinicalTrials.gov16/100

Predictive Clinical Diagnosis of Rheumatoid Arthritis Flares Using Non-Invasive Infra-red Thermal Imaging and an AI/ML Algorithm

ClinicalTrials.gov study NCT05124990. IPD Sharing: Not stated. Countries: 0. Publications: 0.

restrictedIPD-UNDECIDEDFeb 2026View details →
ClinicalTrials.gov16/100

Validation of Sleepware G3 Autoscoring Algorithm

ClinicalTrials.gov study NCT01564472. IPD Sharing: Not stated. Countries: 0. Publications: 0.

restrictedIPD-UNDECIDEDFeb 2026View details →
geo16/100

Constructing a rat model of stress cardiomyopathy and screening for diagnostic markers using machine learning algorithms

GEO Series GSE223385. Rattus norvegicus. 20 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenJan 2023View details →
geo16/100

Statistical algorithm enabled high precision tumor biomarker discovery for circulating extracellular vesicle-based cancer liquid biopsy

GEO Series GSE246925. Homo sapiens. 28 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenJan 2024View details →
geo16/100

OncoSig: a new algorithm for the accurate and systematic de novo prediction of Oncoprotein-centric Signaling networks

GEO Series GSE107042. Mus musculus. 174 samples. Type: Other.

openGEO-OpenNov 2020View details →
dryad16/100

Testing J48 and KNN Algorithms with Hypothyroid dataset

Open the record for dataset details and reuse information.

publicMay 2017View details →
geo12/100

Network Analysis of Breast Cancer Progression and Reversal with a Tree-Evolving Network Algorithm

GEO Series GSE42125. Homo sapiens. 15 samples. Type: Expression profiling by array.

openGEO-OpenJul 2014View details →
geo12/100

The significance of different algorithms in transcriptome analysis of leukemic cells with rearranged MLL

GEO Series GSE54654. Homo sapiens. 8 samples. Type: Expression profiling by high throughput sequencing; Expression profiling by array.

openGEO-OpenFeb 2017View details →
geo12/100

A systematic evaluation of pattern discovery algorithms

GEO Series GSE15370. Homo sapiens. 23 samples. Type: Genome binding/occupancy profiling by genome tiling array.

openGEO-OpenNov 2009View details →
zenodo12/100

Investigating the Combined Impact of Climate Change and Land Use/Land Cover on Flood Vulnerability Using a Machine Learning Algorithm

<p>Using the uploaded code in preparation of the Flood vulnerability maps.</p><p>&nbsp;</p>

restrictedcc-by-4.0Aug 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record