Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,075
datasets available to search
ShareScore release 0.7.1
Dataset results
1,075 results for “ML”
PDB ML Dataset
<p>Included are machine learning-ready datasets of experimental PDB protein structures.</p>
Fibertools: fast and accurate m6A calling using single-molecule long-read sequencing (ML data)
<p>Fibertools is a convolutional neural network that permits the fast and accurate identification of endogenous and exogenous N6-methyladenine (m6A)-marked bases using single-molecule long-read sequencing.<strong> </strong>This dataset (ML data) provides training and validation data for training fibertools supervised and semi-supervised CNN models for three long-read chemistries.</p>
IRAS ML Predictions
<p>Contained in this repository are the large files required for running the code in the following GitHub repository: https://github.com/zfried/IRAS_ML_Predictions</p> <p>fullDataset.csv contains all of the SMILES strings and length 70 feature vectors of the molecules used to train the mol2vec model.</p> <p>mol2vec_model_final_70.pkl is the trained mol2vec model that produces length 70 molecular feature vectors. </p> <p>prediction_molecules.csv contains all of the SMILES strings and length 89 feature vectors (including isotopic encoding) of the 84,863 molecules that are inputted into the trained model. </p>
Weather-based ML Datasets
<p>Included are three datasets and machine learning models from Handler et al. (2020), Flora et al. (2021), and Chase et al. (2022. 2023). Handler et al. (2020) is a 1-hr HRRR-based nowcasting dataset predicting frozen road surfaces. Flora et al. (2021) is a 0-150 min severe weather dataset derived from the convection-allowing ensemble system known as the Warn-on-Forecast System. Lastly, the Chase et al. dataset is a lower resolution of the Storm EVent ImagRy (SEVIR; Veillette et al. 2020) dataset where the goal is to predict lightning.</p>
Modeled waver data for ML
<p>The uploaded data are numerical model generated daily wave height, period, and wind frocings in the Chesapeake Bay. Data are used for the publication "Machine Learning-based Wave Model with High Spatial Resolution in Chesapeake Bay" submitted to the Journal for review.</p>
How Do ML Practitioners Perceive Explainability? An Interview Study of Practices and Challenges
<p>This repository contains supplementary materials for our interview Study.</p>
A Clinical Trial to Evaluate the Safety and Efficacy of 20 ml Cerebrolysin in Patients With Vascular Dementia
ClinicalTrials.gov study NCT00947531. IPD Sharing: Not stated. Countries: 1. Publications: 1.
Encouraging Flu Vaccination Among High-Risk Patients Identified by ML
ClinicalTrials.gov study NCT04323137. IPD Sharing: YES. Countries: 1. Publications: 9.
Clinical Performance Evaluation of the Artificial Intelligence (AI)/ Machine Learning (ML) Technologies Utilized by the Origin Medical EXAM ASSISTANT
ClinicalTrials.gov study NCT06952439. IPD Sharing: NO. Countries: 1. Publications: 15.
Non-interventional Study on the Monthly Administration of 300 mg AliRocumab (PRALUENT®) With the 2 ml SYDNEY Auto-injector
ClinicalTrials.gov study NCT05129241. IPD Sharing: YES. Countries: 1. Publications: 1.
12 Versus 20 mL PCB for D&E Cervical Prep
ClinicalTrials.gov study NCT03356145. IPD Sharing: NO. Countries: 1. Publications: 1.
TVEC and Preop Radiation for Sarcoma (4 ml Dose)
ClinicalTrials.gov study NCT02453191. IPD Sharing: NO. Countries: 1. Publications: 1.
Dataset for "The State of the ML-universe: 10 Years of Artificial Intelligence & Machine Learning Software Development on GitHub"
<p>Supplementary data to "The State of the ML-universe: 10 Years of Artificial Intelligence & Machine Learning Software Development on GitHub" accepted for publication at MSR 2020.</p> <p>The data included in this package were used to conduct analyses to characterize the AI & ML software development community hosted on GitHub. Please read the paper for a full understanding of what data was collected and how it was used.</p> <p>Questions and comments can be directed to Danielle Gonzalez dng2551@rit.edu</p>
Mining For Ligandable Cavities in RNA [Dataset and ML-Code]
<p>Data and code to train models described in the manuscript: Mining For Ligandable Cavities in RNA (DOI: 10.1021/acsmedchemlett.1c00068)</p> <p> </p>
Shotgun MVP: Reducing risk in highly unpredictable ML products
<p>Machine learning products do not fit neatly into the agile methodology. There are different reasons for this, the primary one being the lack of certainty behind which directions will be successful (due to data quality and similar issues). Thus, what often happens in organizations is <em>pilotitis</em> - the disease of launching MVP after MVP without tangible progress and commitment.</p> <p>There is a method to address this issue while still alleviating the uncertainty behind planning MVPs - the Shotgun MVP. As a first step, several MVP initiatives are launched simultaneously (staffed to the bare minimum needed). Thus, all possible approaches are covered. This process continues for several sprints until a clear frontrunner becomes apparent - based on a "success" metric. There is flexibility in defining this metric, and possible examples include performance metrics (i.e., model accuracy) and solution complexity (such as the workforce needed to complete further development iterations). After this MVP (MVP II in our case) is selected, resources are committed fully to it for the following sprints while keeping the other options as backup plans or possible enhancements for downstream work.</p> <p>This simple procedure is not easy to perform and requires careful and restrained execution but can provide the shortest way to delivering a successful ML product.</p>
TC-MYNN-ML
<p>Due to the large amount of data, this paper only provides part of the model output.</p> <p>2 sets of files:</p> <p>Intensity has the data to make Fig.1</p> <p>Energy_Spectr has the data to make Fig.6</p>
Pre-Islamic Alabaster Jar, Mleiha ML 5, Sharjah
Pre-Islamic Alabaster Jar, Mleiha ML 5, Sharjah, UAE. In storage at the Sharjah Archaeology Authority. 250-150 BCE. Catalog number unk. Processed in Reality Capture from 245 images. Source: Objaverse 1.0 / Sketchfab
Pre-Islamic Alabaster Jar, Mleiha ML 5, Sharjah
Pre-Islamic Alabaster Jar, Mleiha ML 5, Sharjah, UAE. In storage at the Sharjah Archaeology Authority. 250-150 BCE. Catalog number unk. Processed in Reality Capture from 480 images. Source: Objaverse 1.0 / Sketchfab
TimeSeries_Clustering_ML_data
<p>extended database upload for Time Series Clustering Machine Learning framework</p> <p>(this version includes all the initial large dataset files >10 GB that were created with Time Series Clustering Machine Learning)</p> <p> </p> <p> </p>
Replication Package for ML-EUP Conversational Agent Study
<p>This is the replication package of the paper <a href="https://conf.researchr.org/details/icse-2024/icse-2024-research-track/5/How-to-Support-ML-End-User-Programmers-through-a-Conversational-Agent">How to Support ML End-User Programmers through a Conversational Agent</a>, published at ICSE 2024.</p> <p><strong>Replication Package Files</strong></p> <ul> <li><strong>Readme.pdf:</strong> document that describes the replication package and indicates how to use it. </li> <li> <p><strong>1. Forms.zip: </strong>contains the forms used to collect data for the experiment.</p> </li> <li> <p><strong>2. Experiments.zip: </strong>contains the participants’ and sandboxers’ experimental task workflow with Newton.</p> </li> <li> <p><strong>3. Responses.zip: </strong>contains the responses collected from participants during the experiments.</p> </li> <li> <p><strong>4. Analysis.zip:</strong> contains the data analysis scripts and results of the experiments.</p> </li> <li> <p><strong>5. newton.zip: </strong>contains the tool we used for the WoZ experiment.</p> </li> <li> <p><strong>Interactions.pdf: </strong>explains Figure 4 of the paper in detail by depicting the interactions of P4.</p> </li> <li> <p><strong>TutorialStudy.pdf:</strong> script used in the experiment with and without Newton to be consistent with all participants.</p> </li> <li> <p><strong>Woz_Script.pdf:</strong> script wizard used to maintain consistent Newton responses among the participants.</p> </li> <li><strong>Dockerfile</strong>: docker definition of newton-docker.tar.gz.</li> <li> <p><strong>newton-docker.tar.gz:</strong> docker image that contains both the tool and the analysis files.</p> </li> <li><strong>LICENSE: </strong>license file describing the license of data and code files.</li> </ul> <p> </p> <p><strong>1. Forms.zip</strong></p> <p>The forms zip contains the following files:</p> <ul> <li> <p><strong>Demographics.pdf: </strong>a PDF form used to collect demographic information from participants before the experiments</p> </li> <li> <p><strong>Post-Task Control (without the tool).pdf:</strong> a PDF form used to collect data from participants about challenges and interactions when performing the task without Newton </p> </li> <li> <p><strong>Post-Task Newton (with the tool).pdf:</strong> a PDF form used to collect data from participants after the task with Newton.</p> </li> <li> <p><strong>Post-Study Questionnaire.pdf:</strong> a PDF form used to collect data from the participant after the experiment.</p> </li> </ul> <p> </p> <p><strong>2. Experiments.zip</strong></p> <p>The experiments zip contains two types of folders:</p> <ul> <li> <p><strong>exp[participant’s number]-c[number of dataset used for control task]e[number of dataset used for experimental task]</strong>. Example: exp1-c2e1 (experiment participant 1 - control used dataset 2, experimental used dataset 1)</p> </li> <li> <p><strong>sandboxing[sandboxer’s number].</strong> Example: sandboxing1 (experiment with sandboxer 1)</p> </li> </ul> <p> </p> <p>Every experiment subfolder contains:</p> <ul> <li> <p><strong>warmup.json: </strong>a JSON file with the results of Newton-Participant interactions in the chat for the warmup task.</p> </li> <li> <p><strong>warmup.ipynb: </strong>a Jupyter notebook file with the participant’s results from the code provided by Newton in the warmup task.</p> </li> <li> <p><strong>sample1.csv: </strong>Death Event dataset.</p> </li> <li> <p><strong>sample2.csv: </strong>Heart Disease dataset.</p> </li> <li> <p><strong>tool.ipynb: </strong>a Jupyter notebook file with the participant’s results from the code provided by Newton in the experimental task.</p> </li> <li> <p><strong>python.ipynb: </strong>a Jupyter notebook file with the participant’s results from the code they tried during the control task.</p> </li> <li> <p><strong>results.json:</strong> a JSON file with the results of Newton-Participant interactions in the chat for the task with Newton.</p> </li> </ul> <p> </p> <p>To load an experiment chat log into Newton, add the following code to the notebook:</p> <pre><code>import anachat import json with open("result.json", "r") as f: anachat.comm.COMM.history = json.load(f) </code></pre> <p>Then, click on the notebook name inside Newton chat</p> <p>Note 1: the subfolder for P6 is exp6-e2c1-serverdied because the experiment server died before we were able to save the logs. We reconstructed them using the notebook newton_remake.ipynb based on the video recording.</p> <p>Note 2: The sandboxing occurred during the development of Newton. We did not collect all the files, and the format of JSON files is different than the one supported by the attached version of Newton.</p> <p> </p> <p><strong>3. Responses.zip</strong></p> <p>The responses zip contains the following files:</p> <ul> <li> <p><strong>demographics.csv: </strong>a CSV file containing the responses collected from participants using the demographics form</p> </li> <li> <p><strong>task_newton.csv: </strong>a CSV file containing the responses collected from participants using the post-task newton form.</p> </li> <li> <p><strong>task_control.csv:</strong> a CSV file containing the responses collected from participants using the post-task control form.</p> </li> <li> <p><strong>post_study.csv:</strong> a CSV file containing the responses collected from participants using the post-study control form.</p> </li> </ul> <p> </p> <p><strong>4. Analysis.zip</strong></p> <p>The analysis zip contains the following files:</p> <ul> <li> <p><strong>1.Challenge.ipynb:</strong> a Jupyter notebook file that performs the statistical tests and creates the perceptions of challenges figure.</p> </li> <li> <p><strong>2.Interactions.py: </strong>a Python file that creates the participants’ JSON files.</p> </li> <li> <p><strong>3.Interactions.Graph.ipynb: </strong>a Jupyter notebook file that creates the participant’s interaction figure.</p> </li> <li> <p><strong>4.Interactions.Count.ipynb:</strong> a Jupyter notebook file that counts participants’ interaction with each figure.</p> </li> <li> <p><strong>config_interactions.py:</strong> this file contains the definitions of interaction colors and grouping</p> </li> <li> <p><strong>interactions.json: </strong>a JSON file with the interactions during the Newton task of each participant based on the categorization.</p> </li> <li> <p><strong>requirements.txt: </strong>dependencies required to run the code to generate the graphs and json analysis.</p> </li> </ul> <p> </p> <p>To run the analyses, please follow the steps:</p> <p>1- Extract Analysis.zip and cd into the directory</p> <p>2- Install Python 3.10, and then the analysis dependencies with the following command:</p> <pre><code>pip install -r requirements.txt</code></pre> <p>3- Run Jupyter Notebook/Lab and execute all cells of <strong>1.Challenge.ipynb</strong>. It will generate the challenges figure.</p> <p>4- Run <strong>2.Interactions.py</strong> using the following command:</p> <pre><code>python 2.Interactions.py</code></pre> <p>This file was created manually by individually categorizing each interaction of the participants. The execution will generate the file interactions.json with the graph definitions of the interactions.</p> <p>5- Run Jupyter Notebook/Lab and execute all cells of <strong>3.Interactions.Graph.ipynb</strong>. It will create the interactions graph visualization.</p> <p>6- Run Jupyter Notebook/Lab and execute all cells of <strong>4.Interactions.Count.ipynb</strong>. It will create the interactions table.</p> <p> </p> <p><strong>5. newton.zip</strong></p> <p>The newton zip contains the <strong>source code of the Jupyter Lab extension</strong> we used in the experiments. Read the <strong>README.md</strong> file inside it for instructions on how to install and run it.</p> <p> </p> <p><strong>6. newton-docker.tar.gz</strong></p> <p>This file contains the Docker image with the replication package in a configured environment for both the tool and the dat analyses.</p> <p>To import the image, run:</p> <pre><code><span>docker load <</span> <span>newton-docker.tar.gz</span></code></pre> <p>Then, start the container by running:</p> <pre><code><span>docker run -p 8888:8888 -it newton</span></code></pre> <p>Finally, start Jupyter Lab:</p> <pre><code><span>jupyter lab --collaborative --ip="*" --port=8888 --allow-root</span></code></pre> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.