Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,943

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,943 results for “machine learning”

Learn how ShareScore rates datasets ↗
zenodo40/100

Towards Generalization of Machine Learning Models: An Arabic Sentiment Analysis Dataset

<p>This data set consists of approximately 1.64 Million Arabic tweets (shared by their IDs) posted from 2009 to 2020, and their corresponding sentiment using a three-point classification system of Positive, Negative and Neutral/Mixed.&nbsp;No specific locations and/or keywords were specified throughout the data collection to obtain variation in the dialects and topics represented within the dataset. It is important to note that any biases in the proposed dataset in relation to the dialects and/or topics discussed were unintentional.</p> <p><strong>Please&nbsp;use the following citation if you use this data in a paper:</strong></p> <blockquote> <p>Abdaljalil, S., Hassanein, S., Mubarak, H., &amp; Abdelali, A. (2023). Towards Generalization of Machine Learning Models: A Case Study of Arabic Sentiment Analysis.&nbsp;<strong>Proceedings of the International AAAI Conference on Web and Social Media,&nbsp;17(1), 971-980.</strong></p> </blockquote> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Apr 2023View details →
zenodo40/100

DIETxPOSOME - Selection of potentially useful papers from literature mining and machine learning protocols

<p>DIETxPOSOME database concerning literature selection of potentially useful papers retrieved from PubMed search API, concerning contaminants quantification in food items of worldwide highest supply and using FoodMine code (text matching filter) and machine learning (ML) protocols. A list of 2,442 papers that potentially contained relevant information, covering the period between 2000 and 2022, was compiled from an initial number of 1,932,345 papers.</p>

opencc-by-4.0Apr 2023View details →
zenodo40/100

DIETxPOSOME - Summary statistics from papers obtained from literature mining and machine learning protocols

<p>DIETxPOSOME database concerning literature selection of potentially useful papers retrieved from PubMed search API, concerning contaminants quantification in food items of worldwide highest supply and using FoodMine code (text matching filter) and machine learning (ML) protocols.&nbsp;11,723 data points were collected from 254 papers from the last two decades&nbsp;in 72 foods to obtain relevant information on 96 contaminants, including heavy metals, polychlorinated biphenyls, dioxins, furans, polycyclic aromatic hydrocarbons (PAHs), pesticides, mycotoxins, and heterocyclic aromatic amines (HAAs).</p> <p>&nbsp;</p>

opencc-by-4.0May 2023View details →
zenodo40/100

Machine Learning the Tip of the Red Giant Branch

<p># Reproduction package for the Paper &quot;Machine Learning the Tip of the Red Giant Branch&quot;<br> ## Authors<br> Mitchell T. Dennis (mtde226@hawaii.edu)<br> Jeremy Sakstein (sakstein@hawaii.edu)<br> ## Software<br> MESA version 15140 (http://mesa.sourceforge.net/) &nbsp;<br> MESASDK version 20210401 (http://www.astro.wisc.edu/~townsend/static.php?ref=mesasdk) &nbsp;<br> GFORTRAN GCC version 9.2.0<br> ## Citation Policy<br> If you use any part of this reproduction package for independent work, we recommend you cite the following papers:</p> <p>* This paper<br> * Astrophys. J. Suppl. 192, 3 (2011)<br> * Astrophys. J. Suppl. 208, 4 (2013)<br> * Astrophys. J. Suppl. 234, 34 (2018)<br> * Astrophys. J. Suppl. 243, 10 (2019)</p> <p>&nbsp;</p> <p>For more information, see the Readme.md</p>

opencc-by-4.0Oct 2022View details →
zenodo40/100

Exploring the utility of regulatory network-based machine learning for gene expression prediction in maize

<p>Relevant Data and Code for&nbsp;<em>Exploring the utility of regulatory network-based &nbsp;machine learning for gene expression prediction in maize&nbsp;</em>by Taylor Ferebee and Edward Buckler.</p> <p><strong>Input&nbsp;Data</strong></p> <p>The inputs&nbsp;of the models are enclosed in&nbsp;<em>Input_data-2022-001.zip</em></p> <p><strong>Output Data</strong></p> <p>The outputs of the models are enclosed in&nbsp;<em>Output_Results-2022-001.zip</em></p> <p><strong>Relevant Code&nbsp;</strong></p> <p>The code for all analyses is enclosed in <em>Code_Archive.zip&nbsp;</em></p>

opencc-by-4.0May 2023View details →
zenodo40/100

BRCA1-specific machine learning model predicts variant pathogenicity with high accuracy - Supplementary material

<p>Figure S1: Distribution of the reviewed 141&nbsp;<em>BRCA1</em>&nbsp;missense variants; Figure S2:&nbsp;The Shapely values for the&nbsp;<em>BRCA1</em>&nbsp;XGBoost models; Figure S3: The Shapely values of the&nbsp;<em>BRCA1</em>&nbsp;XGBoost model used to predict the functional assays&rsquo; results for variants of uncertain significance;&nbsp;Table S1: The receiver operating characteristic (ROC) curve analysis for the different in silico predictions; Table S2: Cross validation of the BRCA1 model in 5 different random training and test samples; Table S3: Pathogenicity prediction and prioritization of the 31,058 unreviewed BRCA1 variants from the BRCA Exchange database.</p>

opencc-by-4.0Apr 2023View details →
zenodo40/100

Large-scale Docking Datasets for Machine Learning

<p><strong>Large-scale virtual screening has become a valuable tool for early-phase drug discovery. Recent expansions of commercial chemical space have made it computationally intractable to evaluate all compounds in the libraries. Machine learning is one of the methods that aim to prioritize specific subsets of these vast libraries. In order to put these methods to the test, access to large-scale datasets is beneficial. To help the community benchmark their work, we share the docking scores of several ultralarge virtual screening campaigns.</strong></p> <p><strong>The datasets we provide contain canonical SMILES, compound identifiers, and docking scores. We docked two different chemical libraries against eight different biological targets with therapeutic relevance. The first dataset contained approximately 15.5 million molecules adhering to the &quot;Rule-of-Four&quot;, whereas the second datasets consists of approximately 235 million &quot;lead-like&quot; molecules. The biological targets represent different classes of proteins and binding sites.</strong></p> <p><strong>More details on the datasets and our methods can be found on (<a href="http://github.com/carlssonlab/conformalpredictor">https://github.com/carlssonlab/conformalpredictor</a>) and our pre-print (<a href="https://doi.org/10.26434/chemrxiv-2023-w3x36">https://doi.org/10.26434/chemrxiv-2023-w3x36</a>). </strong></p> <p><strong>Please feel free to download and use these datasets for your own research purposes. We only ask that you cite our pre-print and datasets appropriately if you use it in your work. Thank you for your interest in our research!</strong></p>

opencc-by-4.0May 2023View details →
zenodo40/100

Dataset for "Brain-machine interface learning is facilitated by specific patterning of distributed cortical feedback"

<p><strong>Dataset for the manuscript entitled: Brain-machine interface learning is facilitated by specific patterning of distributed cortical feedback.</strong></p> <p><strong><em>DOI of the manuscript:</em>&nbsp;<a href="https://doi.org/10.1126/sciadv.adh1328">10.1126/sciadv.adh1328</a></strong></p> <p><em><strong>Abstract of the manuscript:</strong></em></p> <p>Neuroprosthetics offer great hope for motor-impaired patients. One obstacle is that fine motor control requires near-instantaneous, rich somatosensory feedback. Such distributed feedback may be recreated in a brain-machine interface using distributed artificial stimulation across the cortical surface. Here, we hypothesized that neuronal stimulation must be contiguous in its spatiotemporal dynamics in order to be efficiently integrated by sensorimotor circuits. Using a closed-loop brain-machine interface, we trained head-fixed mice to control a virtual cursor by modulating the activity of motor cortex neurons. We provided artificial feedback in real time with distributed optogenetic stimulation patterns in the primary somatosensory cortex. Mice developed a specific motor strategy and succeeded to learn the task only when the optogenetic feedback pattern was spatially and temporally contiguous while it moved across the topography of the somatosensory cortex. These results reveal new properties of sensorimotor cortical integration and set new constraints on the design of neuroprosthetics.</p> <p><strong><em>Description of the variables in the data storage dictionary:</em></strong></p> <ol> <li>cursor_positions: sequence of the virtual cursor position over a session</li> <li>range_of_rewardable_cursor_position: range of cursor positions that can be rewarded. Upper threshold excluded. Lower threshold included. &nbsp;</li> <li>cursor_times: timing of the cursor positions provided by the cursor_positions data, in the same clock as spike and lick times.&nbsp;</li> <li>lick_times: timing of all recorded licks. &nbsp;</li> <li>reward_times:timing of the opening of the valve that releases the water reward.&nbsp;</li> <li>master_spike_times: time of the spikes of the master neurons.</li> <li>master_spike_shape: spike shape of each spike stored in master_spike_times. The shipe shapes are shown for 3s (30kHz sampling rate), for each 4 electrode of the corresponding tetrode.</li> <li>neighbor_spike_times': same as master_spike_times for neighbor neurons.</li> <li>neighbor_spike_shape': same as master_spike_shape for neighbor neurons. &nbsp;&nbsp;</li> </ol> <p><em><strong>General structure of the data set:</strong></em></p> <p>Each variable is a hierarchical tree of lists: [<em>Protocol</em>][<em>Mouse</em>][<em>Session</em>].&nbsp;</p> <p><em>Protocol</em> takes one of the following values:&nbsp;0: Bar feedback | 1: Full shuffle |&nbsp;2: No Feedback |&nbsp;5: Playback Structured Feedback |&nbsp;7: Spontaneous activity |&nbsp;8: Barrel Shuffle | 9: Frame shuffle.</p> <p><em>Mouse </em>and&nbsp;<em>Session</em> iterate through respectively the mice that were involved in the protocole, and the sessions, generally 5 in total.</p> <p><em><strong>Python code to load the hdf5 File</strong></em></p> <p>The code below relies on libraries available on a standard, mac anaconda install of the jupyter notebook system on the 25/03/2024. It provides with several nested lists of Numpy arrays.</p> <p>For instance, to access the series of cursor positions of Mouse M during session number S of Protocole P, index as follows:&nbsp;</p> <p>CP = cursor_positions_DATA[P][M][S]</p> <p># 1 - Load the libraries</p> <p>&nbsp; &nbsp; import h5py<br>&nbsp; &nbsp; import numpy as np<br>&nbsp; &nbsp; import pylab as pl</p> <p>&nbsp; &nbsp; filename = "./data.h5"<br>&nbsp; &nbsp; f = h5py.File(filename, "r")</p> <p># 2 - Extraction of the cursor position, time, lick and reward time, as well as the activity of the master and neighbor neurons.&nbsp;</p> <p>&nbsp; &nbsp; cursor_positions = f['cursor_positions']</p> <p>&nbsp; &nbsp; cursor_positions_DATA = []<br>&nbsp; &nbsp; for Protocole in range(9):<br>&nbsp; &nbsp; &nbsp; &nbsp; DATA_protocole = []<br>&nbsp; &nbsp; &nbsp; &nbsp; P = f[cursor_positions[Protocole][0]]<br>&nbsp; &nbsp; &nbsp; &nbsp; for Mouse in range(16): &nbsp; &nbsp;<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; # Loop through the mice, and collect&nbsp;<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; M = f[P[Mouse][0]]<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; if not(list(M) == [0,1]) and (len(M) == 5):<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; DATA_mouse = []<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; for Session in range(5):<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; S = f[M[Session][0]]<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; bfr = np.array(S)<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; DATA_mouse.append(bfr)&nbsp;<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; DATA_protocole.append(DATA_mouse)<br>&nbsp; &nbsp; &nbsp; &nbsp; cursor_positions_DATA.append(DATA_protocole)</p> <p>&nbsp;</p> <p>&nbsp; &nbsp; cursor_times = f['cursor_times']</p> <p>&nbsp; &nbsp; cursor_times_DATA = []<br>&nbsp; &nbsp; for Protocole in range(9):<br>&nbsp; &nbsp; &nbsp; &nbsp; DATA_protocole = []<br>&nbsp; &nbsp; &nbsp; &nbsp; P = f[cursor_times[Protocole][0]]<br>&nbsp; &nbsp; &nbsp; &nbsp; for Mouse in range(16): &nbsp; &nbsp;<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; # Loop through the mice, and collect&nbsp;<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; M = f[P[Mouse][0]]<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; if not(list(M) == [0,1]) and (len(M) == 5):<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; DATA_mouse = []<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; for Session in range(5):<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; S = f[M[Session][0]]<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; bfr = np.array(S)<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; DATA_mouse.append(bfr)&nbsp;<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; DATA_protocole.append(DATA_mouse)<br>&nbsp; &nbsp; &nbsp; &nbsp; cursor_times_DATA.append(DATA_protocole)</p> <p>&nbsp;</p> <p>&nbsp; &nbsp; lick_times = f['lick_times']</p> <p>&nbsp; &nbsp; lick_times_DATA = []<br>&nbsp; &nbsp; for Protocole in range(9):<br>&nbsp; &nbsp; &nbsp; &nbsp; DATA_protocole = []<br>&nbsp; &nbsp; &nbsp; &nbsp; P = f[lick_times[Protocole][0]]<br>&nbsp; &nbsp; &nbsp; &nbsp; for Mouse in range(16): &nbsp; &nbsp;<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; # Loop through the mice, and collect&nbsp;<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; M = f[P[Mouse][0]]<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; if not(list(M) == [0,1]) and (len(M) == 5):<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; DATA_mouse = []<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; for Session in range(5):<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; S = f[M[Session][0]]<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; bfr = np.array(S)<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; DATA_mouse.append(bfr)&nbsp;<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; DATA_protocole.append(DATA_mouse)<br>&nbsp; &nbsp; &nbsp; &nbsp; lick_times_DATA.append(DATA_protocole)</p> <p>&nbsp;</p> <p>&nbsp; &nbsp; reward_times = f['reward_times']</p> <p>&nbsp; &nbsp; reward_times_DATA = []<br>&nbsp; &nbsp; for Protocole in range(9):<br>&nbsp; &nbsp; &nbsp; &nbsp; DATA_protocole = []<br>&nbsp; &nbsp; &nbsp; &nbsp; P = f[reward_times[Protocole][0]]<br>&nbsp; &nbsp; &nbsp; &nbsp; for Mouse in range(16): &nbsp; &nbsp;<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; # Loop through the mice, and collect&nbsp;<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; M = f[P[Mouse][0]]<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; if not(list(M) == [0,1]) and (len(M) == 5):<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; DATA_mouse = []<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; for Session in range(5):<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; S = f[M[Session][0]]<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; bfr = np.array(S)<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; DATA_mouse.append(bfr)&nbsp;<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; DATA_protocole.append(DATA_mouse)<br>&nbsp; &nbsp; &nbsp; &nbsp; reward_times_DATA.append(DATA_protocole)</p> <p>&nbsp;</p> <p>&nbsp; &nbsp; master_spike_times = f['master_spike_times']</p> <p>&nbsp; &nbsp; master_spike_times_DATA = []<br>&nbsp; &nbsp; for Protocole in range(9):<br>&nbsp; &nbsp; &nbsp; &nbsp; DATA_protocole = []<br>&nbsp; &nbsp; &nbsp; &nbsp; P = f[master_spike_times[Protocole][0]]<br>&nbsp; &nbsp; &nbsp; &nbsp; for Mouse in range(16): &nbsp; &nbsp;<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; # Loop through the mice, and collect&nbsp;<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; M = f[P[Mouse][0]]<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; if not(list(M) == [0,1]) and (len(M) == 5):<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; DATA_mouse = []<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; for Session in range(5):<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; S = f[M[Session][0]]<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; DATA_unit = []<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; for Unit in range(len(S)):<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; if not(list(S) == [0,1]):<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; N = f[S[Unit][0]]<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; bfr = np.array(N)<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; DATA_unit.append(bfr)&nbsp;<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; DATA_mouse.append(DATA_unit)<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; DATA_protocole.append(DATA_mouse)<br>&nbsp; &nbsp; &nbsp; &nbsp; master_spike_times_DATA.append(DATA_protocole)</p> <p>&nbsp;</p> <p>&nbsp; &nbsp; neighbor_spike_times = f['neighbor_spike_times']</p> <p>&nbsp; &nbsp; neighbor_spike_times_DATA = []<br>&nbsp; &nbsp; for Protocole in range(9):<br>&nbsp; &nbsp; &nbsp; &nbsp; DATA_protocole = []<br>&nbsp; &nbsp; &nbsp; &nbsp; P = f[neighbor_spike_times[Protocole][0]]<br>&nbsp; &nbsp; &nbsp; &nbsp; for Mouse in range(16): &nbsp; &nbsp;<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; # Loop through the mice, and collect&nbsp;<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; M = f[P[Mouse][0]]<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; if not(list(M) == [0,1]) and (len(M) == 5):<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; DATA_mouse = []<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; for Session in range(5):<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; S = f[M[Session][0]]<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; DATA_unit = []<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; for Unit in range(len(S)):<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; if not(list(S) == [0,1]):<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; N = f[S[Unit][0]]<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; bfr = np.array(N)<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; DATA_unit.append(bfr)&nbsp;<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; DATA_mouse.append(DATA_unit)<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; DATA_protocole.append(DATA_mouse)<br>&nbsp; &nbsp; &nbsp; &nbsp; neighbor_spike_times_DATA.append(DATA_protocole)</p> <p>&nbsp;</p>

opencc-by-4.0May 2023View details →
zenodo40/100

A Tagged Traffic Accident Dataset for Machine Learning

<p>This dataset contains tagged accident data and is provided for reproducibility for our journal paper&nbsp;</p> <p><strong>Pablo Moriano, Andy Berres, Haowen Xu, Jibonananda Sanyal. &ldquo;Spatiotemporal Features of Traffic Help Reduce Automatic Accident Detection Time.&rdquo; <em>Expert Systems with Applications</em> 244 (2024): 122813. <a title="Persistent link using digital object identifier" href="https://doi.org/10.1016/j.eswa.2023.122813" target="_blank" rel="noopener">https://doi.org/10.1016/j.eswa.2023.122813</a></strong></p> <p>The accompanying Data in Brief publication discusses the methodology behind the creation of these data.</p> <p><strong>Berres, Andy, Pablo Moriano, Haowen Xu, Sarah Tennille, Lee Smith, Jonathan Storey, and Jibonananda Sanyal. "A Traffic Accident Dataset for Chattanooga, Tennessee."&nbsp;<em>Data in Brief</em> (2024): 110675.</strong></p> <p>&nbsp;</p> <p>The zip folder&nbsp;<strong><em>annotatedData.zip</em></strong> contains two subfolders: <strong><em>allData</em></strong> and <strong><em>bestData</em></strong>. The <em>bestData</em> folder contains all data for which a full neighborhood of five sensors upstream and five sensors downstream is available, whereas <em>allData</em> includes everything from <em>bestData</em> as well as data with a smaller number of neighboring sensors. Each folder contains one subfolder called <strong><em>accidents</em></strong> and one subfolder called <strong><em>non-accidents</em></strong>. The <em>accidents</em> folder contains one file per accident. The <em>non-accidents</em> folder contains files for the same location, day of the week and time as a corresponding accident, for each week during which there was no accident impact on the traffic.</p> <p>The file names in both folders are formatted as follows: <strong>yyyy-mm-dd-hhmm-rrrrrXaaa.a.csv</strong>, consisting of date (yyyy-mm-dd), time (hhmm in 24-hour format), and sensor name (rrrrrXaaa.a), which consists of road name (rrrrr; 5 alphanumerical characters), heading (X), and mile marker (aaa.a). For example, the file <em>2020-11-03-1611-00I24W182.8.csv </em>&nbsp;contains data for an accident which occurred at 4:11 p.m. on November 3, 2020 on I-24 Westbound near the radar sensor at mile marker 182.8.</p> <p>The content of each CSV file is a timeseries of radar data beginning 15 minutes prior to the reported incident and ending 15 minutes after the reported incident. It also contains metadata, such as the accident type, etc. Each CSV file contains the following columns:</p> <ul> <li><strong>incident at sensor(i)</strong>: 1 for yes (<em>accidents</em> folder), 0 for no (<em>non-accidents</em> folder)</li> <li><strong>road</strong>: road name with heading, e.g. 00I24E</li> <li><strong>mile</strong>: mile marker of nearest radar sensor, e.g. 182.8</li> <li><strong>type</strong>: accident type, e.g. &ldquo;Prop Damage (over)&rdquo; for property damage exceeding a certain threshold. For non-accidents, the type is given as &ldquo;None&rdquo;.</li> <li><strong>date</strong>: date of the data sample. For accidents, this is the date on which the accident occurred. For non-accidents, this is the date for which the non-accident data sample is collected.</li> <li><strong>incident_time</strong>: time the reference accident was reported in hh:mm. This is the time which is provided in E-TRIMS as the time the 911 call was made.</li> <li><strong>incident_hour</strong>: just the hour from the incident_time, in integer format.</li> <li><strong>data_time</strong>: timestamp for the timeseries contained in the file in hh:mm:ss format. The timeseries consists of 30 second timesteps.</li> <li><strong>weather</strong>: weather during <em>data_time</em>, based on data collected from NASA POWER. We used dry bulb temperature (&deg;C), precipitation (mm/h), and wind speed (m/s) from the raw NASA POWER data to produce the classifications of <em>rain</em> (at least 1mm precipitation and temperatures above 2&deg;C), <em>snow </em>(at least 1mm precipitation and temperatures at or below 2&deg;C), and <em>wind</em> (wind speeds over 30 mph or 13.5 m/s). If there were no inclement weather conditions, we set the category to <em>&ldquo;--"</em>.</li> <li><strong>light</strong>: light conditions during data_time. To produce this field, we collected sunrise, sunset, civil twilight start and civil twilight end times from <a href="https://sunrise-sunset.org">https://sunrise-sunset.org</a>, and derived the categories dawn, daylight, dusk, and dark using these start and end times.</li> <li>The last 33 columns contain radar data for the 11 sensors surrounding the accident or non-accident. For each sensor, we collected <em>speed</em> (mean over 30-second interval in miles per hour, or empty if no vehicles passed), <em>volume</em> (count of all vehicles passing during 30-second interval), and <em>occupancy</em> (mean % of occupancy over 30-second interval).&nbsp; These three variables are grouped in triples, of <strong>speed (k), volume (k), occupancy (k)</strong>, where <em>k</em> indicates the sensor number relative to the closest sensor <em>i</em> to the incident, <em>k&lt;i</em> indicate upstream sensors and <em>k&gt;i</em> indicate downstream sensors. For example, <strong>speed (i-5)</strong> refers to the mean speed at the sensor which is 5 hops upstream from the accident, and <strong>volume(i+1) </strong>refers to the number of vehicles at the sensor immediately downstream from the accident.</li> </ul> <p>The folder <strong><em>metaData.zip</em></strong> contains the following files:</p> <ul> <li><strong>Accidents.csv</strong>: cleaned-up accidents file with all accidents which happened on Chattanooga area highways between November 1, 2020 and April 29, 2021. We have removed accidents which happened on non-highway roads, and we have corrected the timestamps (which were in 12-hour format but missing a.m./p.m. markers) by cross-referencing light and weather conditions.</li> <li><strong>WeatherDict</strong><strong>.json:</strong> a dictionary containing the weather data synthesized from NASA POWER.</li> <li><strong>LightDict.json</strong>: a dictionary containing the light data synthesized from Sunrise-and-Sunset.</li> <li><strong>SensorTopology.csv</strong>: neighborhood information for each radar sensor in the Chattanooga area.</li> <li><strong>SensorZones.geojson</strong>: polygons used to determine the nearest radar sensor for each accident location. Each polygon is tagged with the corresponding radar sensor&rsquo;s name.</li> </ul>

opencc-by-4.0May 2023View details →
zenodo40/100

Exploring the Diagnostic Markers of Essential Tremor: A study based on Machine Learning Algorithms

<p>supplementary&nbsp;information files for the manuscript&nbsp;Exploring the Diagnostic Markers of Essential Tremor: A study based on Machine Learning Algorithms</p>

opencc-by-4.0May 2023View details →
zenodo40/100

Understanding cirrus clouds using explainable machine learning

<p>This repository contains the data for the paper:</p> <p>Authors:&nbsp;Kai Jeggle&nbsp;, David Neubauer&nbsp;, Gustau Camps-Valls&nbsp;and Ulrike Lohmann<br> Titel:&nbsp;Understanding cirrus clouds using explainable machine learning<br> Date: 2023</p> <p>Note that the scripts can be found in the accompanying package (https://github.com/tabularaza27/explaining_cirrus)</p>

opencc-by-4.0Jan 2023View details →
zenodo40/100

Supplementary Material for Nine tips for ecologists using machine learning

<p>R code and data to reproduce the analysis in &quot;Nine tips for ecologists using machine learning&quot; (https://arxiv.org/abs/2305.10472)</p>

opencc-by-4.0May 2023View details →
zenodo40/100

Prediction of Diblock Copolymer Lamellar Morphologies in 3D and Complex Geometry via Machine Learning

<p>3d morphology evolution of diblock copolymer via 3D UNet.</p> <p>Manuscript:&nbsp;Prediction of Diblock Copolymer Lamellar Morphologies in 3D and Complex Geometry via Machine Learning</p>

opencc-by-4.0May 2023View details →
zenodo40/100

Discrimination of Quartz Genesis Based on Explainable Machine Learning

<p>The investigation of trace elements in quartz samples was conducted via ambient laser ablation inductively coupled plasma mass spectrometry (LA-ICP-MS). The sampled deposits encompassed diverse geological formations, namely porphyry, pegmatite, granite, skarn, epithermal, carlin, greisen, and orogenic deposits. The objective of the study was to employ eight specific trace elements (Al, Ti, Li, Ge, Fe, Na, K, and P) for the characterization of the training samples.</p>

opencc-by-4.0Jun 2023View details →
zenodo40/100

Hainan gibbon (Nomascus hainanus) bioacoustics dataset for machine learning

<p>Data accompanying the paper:&nbsp;&quot;<strong>Empowering Deep Learning Acoustic Classifiers with Human-like Ability to Utilize Contextual Information for Wildlife Monitoring</strong>&quot;</p> <p>We provide the audio data (.wav) used to&nbsp;test our neural network classifier along with the corresponding labelled text files (.svl). The audio and labelled files can easily be viewed in Sonic Visualiser. Drag and drop the audio file. Create the spectrogram layer. Drag and drop the corresponding .svl file.</p> <p>The dataset provided here is a subset of the full dataset provided here:&nbsp;10.5281/zenodo.3991714. This dataset has additional files that were manually annotated, which were not manually verified in the original version (10.5281/zenodo.3991714).</p> <p><strong>Files provided</strong></p> <ul> <li><strong>Audio_x.zip</strong> -- we provide x number of .zip files containing audio files, numbered 1 to 4. These were created in batches to simplify downloads.</li> <li><strong>Annotations.zip -</strong>- .svl files which contain the manually verified labels. These files can be read in Sonic Visualiser, or as .XML files in a programming language. We took care to annotate the start and stop time of each gibbon call. The height of each bounding box is not important as the frequency range of Hainan gibbons is already known.</li> <li><strong>model_weights_tensorflow.hdf5&nbsp;</strong>-- the Tensorflow model. Load the model using: model = tf.keras.models.load_model(model_filepath) note that the model expects a three channel input as explained in the research article.</li> </ul>

opencc-by-nc-sa-4.0Jun 2023View details →
zenodo40/100

Challenges and Limitations in the Design and Implementation of Fair and Equitable Machine Learning Algorithms in Healthcare

<p>We provide the programs in Python, a CSV file with references, images used in the article and a text corpus generated with Python.</p>

opencc-by-4.0Jun 2023View details →
zenodo40/100

Upscaling soil organic carbon measurements at the continental scale using multivariate clustering analysis and machine learning

<p><strong>Data Description</strong>:</p> <p>To improve SOC estimation in the United States, we upscaled site-based SOC measurements to the continental scale using&nbsp;multivariate geographic clustering (MGC)&nbsp;approach coupled with machine learning models. First, we used the&nbsp;MGC approach&nbsp;to segment the United States at 30 arc second resolution based on principal component information from environmental covariates (gNATSGO soil properties, WorldClim bioclimatic variables, MODIS biological&nbsp;variables, and physiographic variables) to&nbsp;20 SOC regions. We then trained separate random forest model ensembles for each of the SOC regions identified using environmental covariates and soil profile measurements from the International Soil Carbon Network (ISCN)&nbsp;and an Alaska soil profile data. We estimated United States SOC for 0-30 cm and 0-100 cm depths were 52.6&nbsp;+&nbsp;3.2 and 108.3&nbsp;+&nbsp;8.2 Pg C, respectively.</p> <p>Files in collection (32):</p> <p>Collection contains 22 soil properties geospatial rasters,&nbsp;4 soil SOC geospatial rasters,&nbsp;2 ISCN site&nbsp;SOC observations&nbsp;csv files, and 4 R scripts</p> <p>gNATSGO&nbsp;TIF files:</p> <p>├── available_water_storage_30arc_30cm_us.tif&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;[30 cm depth soil&nbsp;available&nbsp;water storage]<br> ├── available_water_storage_30arc_100cm_us.tif &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; [100 cm depth soil&nbsp;available&nbsp;water storage]<br> ├── caco3_30arc_30cm_us.tif &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp;[30 cm depth soil CaCO3 content]<br> ├── caco3_30arc_100cm_us.tif &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; [100 cm depth soil CaCO3 content]<br> ├── cec_30arc_30cm_us.tif &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; [30 cm depth soil cation exchange capacity]<br> ├── cec_30arc_100cm_us.tif &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; [100 cm depth soil cation exchange capacity]<br> ├── clay_30arc_30cm_us.tif&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;[30 cm depth soil clay content]<br> ├── clay_30arc_100cm_us.tif&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;[100 cm depth soil clay content]<br> ├── depthWT_30arc_us.tif&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; [depth to water table]<br> ├── kfactor_30arc_30cm_us.tif&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;[30 cm depth soil erosion factor]<br> ├── kfactor_30arc_100cm_us.tif&nbsp;&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; [100 cm depth soil erosion factor]<br> ├── ph_30arc_100cm_us.tif &nbsp;&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; [100 cm depth soil pH]<br> ├── ph_30arc_100cm_us.tif &nbsp;&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; [30 cm depth soil pH]<br> ├── pondingFre_30arc_us.tif &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; [ponding frequency]<br> ├── sand_30arc_30cm_us.tif&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; [30 cm depth soil sand content]<br> ├── sand_30arc_100cm_us.tif&nbsp;&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;[100 cm depth soil sand content]<br> ├── silt_30arc_30cm_us.tif&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; [30 cm depth soil silt content]<br> ├── silt_30arc_100cm_us.tif &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp; [100 cm depth soil silt content]<br> ├── water_content_30arc_30cm_us.tif&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp;[30 cm depth soil water content]<br> └── water_content_30arc_100cm_us.tif&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;[100 cm depth soil water content]</p> <p>SOC TIF&nbsp;files:</p> <p>├──30cm SOC mean.tif&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;[30 cm depth soil SOC]<br> ├──100cm SOC mean.tif&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;[100 cm depth soil SOC]<br> ├──30cm SOC CV.tif&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;[30 cm depth soil SOC coefficient of variation]<br> └──100cm SOC CV.tif&nbsp;&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp;&nbsp;[100 cm depth soil SOC&nbsp;coefficient of variation]</p> <p>site&nbsp;observations csv files:</p> <p>ISCN_rmNRCS_addNCSS_30cm.csv&nbsp; &nbsp; &nbsp; &nbsp;30cm ISCN sites SOC replaced NRCS sites with NCSS centroid removed data</p> <p>ISCN_rmNRCS_addNCSS_100cm.csv&nbsp; &nbsp; &nbsp; &nbsp;100cm ISCN sites SOC replaced NRCS sites with NCSS centroid removed data</p> <p><br> <strong>Data format</strong>:</p> <p>Geospatial files are provided in Geotiff format in Lat/Lon WGS84 EPSG: 4326 projection at 30 arc second resolution.</p> <p><strong>Geospatial projection</strong>:&nbsp;</p> <pre><code>GEOGCS["GCS_WGS_1984", DATUM["D_WGS_1984", SPHEROID["WGS_1984",6378137,298.257223563]], PRIMEM["Greenwich",0], UNIT["Degree",0.017453292519943295]] (base) [jbk@theseus ltar_regionalization]$ g.proj -w GEOGCS["wgs84", DATUM["WGS_1984", SPHEROID["WGS_1984",6378137,298.257223563]], PRIMEM["Greenwich",0], UNIT["degree",0.0174532925199433]] </code></pre> <p>&nbsp;</p>

opencc-by-4.0Jan 2023View details →
zenodo40/100

Climate-Invariant Machine Learning

<p>The &quot;Climate-Invariant Machine Learning&quot;&nbsp; manuscript&#39;s accompanying data is organized into two folders:</p> <ul> <li>&quot;CIML_Fig_Data_v2.zip&quot; contains the data necessary to reproduce all the manuscript&#39;s figures by running the Jupyter notebook <a href="https://github.com/tbeucler/CBRAIN-CAM/blob/master/notebooks/tbeucler_devlog/090_Climate_Invariant_Paper_Figures_v2.ipynb">at this link</a> and to train climate-invariant models by running the Jupyter notebook&nbsp;<a href="https://colab.research.google.com/github/tbeucler/CBRAIN-CAM/blob/master/Climate_Invariant_Guide.ipynb">at this link</a>.</li> <li>&quot;CIML_SPCAM5_Initialization&quot; contains the data necessary to intialize and re-run the three SPCAM5, Earth-like simulations used in the manuscript.</li> </ul> <p>See SI A of the manuscript and the notebooks for more details.</p> <p>This is a pre-release: The release will be final if the manuscript if accepted for publication after peer-review.</p>

opencc-by-4.0Dec 2021View details →
zenodo40/100

Molecular simulations and machine learning potentials for graphene on liquid copper

<p>Dataset for the paper:</p> <p>Gao, H. et al. Graphene at Liquid Copper Catalysts: Atomic-Scale Agreement of Experimental and First-Principles Adsorption Height. Advanced Science 9, 2204684 (2022). (DOI: 10.1002/advs.202204684)</p> <p>&nbsp;</p> <p>zenodo/dataset/: Training and test sets for MTP</p> <p>zenodo/md/: Initial atomic models of the Gr-Cu interface, input files and resulting trajectories of MD simulations</p> <p>zenodo/potentials/: Trained MTP potential files</p> <p>&nbsp;</p>

opencc-by-4.0Aug 2022View details →
dryad40/100

Machine learning analysis of wing venation patterns accurately identifies Sarcophagidae, Calliphoridae and Muscidae fly species

<p>In medical, veterinary, and forensic entomology, the ease and affordability of image data acquisition have resulted in whole-image analysis becoming an invaluable approach for species identification. Krawtchouk moment invariants are a classical mathematical transformation that can extract local features from an image, thus allowing subtle species-specific biological variations to be accentuated for subsequent analyses. We extracted Krawtchouk moment invariant features from binarised wing images of 759 male fly specimens from the Calliphoridae, Sarcophagidae, and Muscidae families (13 species and a species variant). Subsequently, we trained the Generalized, Unbiased, Interaction Detection and Estimation (GUIDE) random forests classifier using linear discriminants derived from these features and inferred the species identity of specimens from the test samples. Five-fold cross validation results show a 98.56 ± 0.38% (standard error) mean identification accuracy at the family level, and a 91.04 ± 1.33% mean identification accuracy at the species level. The mean F1-score of 0.89 ± 0.02 reflects good balance of precision and recall properties of the model. The present study consolidates findings from previous small pilot studies of the usefulness of wing venation patterns for inferring species identities. Thus, the stage is set for the development of a mature data analytic ecosystem for routine computer image-based identification of fly species that are of medical, veterinary, and forensic importance.</p>

opencc-zeroJul 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record