Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,782

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

1,782 results for “algorithms”

Learn how ShareScore rates datasets ↗
zenodo44/100

The COUGHVID crowdsourcing dataset: A corpus for the study of large-scale cough analysis algorithms

<p><strong>Overview</strong></p> <p>Cough audio signal classification has been successfully used to diagnose a variety of respiratory conditions, and there has been significant interest in leveraging Machine Learning (ML) to provide widespread COVID-19 screening. The COUGHVID dataset provides over 30,000 crowdsourced cough recordings representing a wide range of subject ages, genders, geographic locations, and COVID-19 statuses. Furthermore, experienced pulmonologists labeled more than 2,000 recordings to diagnose medical abnormalities present in the coughs, thereby contributing one of the largest expert-labeled cough datasets in existence that can be used for a plethora of cough audio classification tasks.&nbsp;As a result, the COUGHVID dataset contributes a wealth of cough recordings for training ML models to address the world&rsquo;s most urgent health crises.</p> <p><strong>Private Set and Testing Protocol</strong></p> <p>Researchers interested in testing their models on the private test dataset should contact us at coughvid@epfl.ch, briefly explaining the type of validation they wish&nbsp;to make, and their obtained results obtained through&nbsp;cross-validation with the public data. Then, access to the unlabeled recordings will be provided, and&nbsp;the researchers should&nbsp;send the predictions of their models on these recordings. Finally,&nbsp;the&nbsp;performance metrics of the predictions will be sent to the researchers. The private testing data is not included in any file within our Zenodo record, and it can only be accessed by contacting the COUGHVID team at the aforementioned e-mail address.</p> <p><strong>New Semi-Supervised Labeling</strong></p> <p>The third version of the COUGHVID dataset contains thousands of additional recordings obtained through October 2021. Additionally, the recordings containing coughs were re-labeled according to a semi-supervised learning algorithm that combined the user labels with those of the expert physicians, which were&nbsp;modeled using ML and expanded on the previously unlabeled data. These labels can be found in the &quot;status_SSL&quot; column of the &quot;metadata_compiled.csv&quot; file.</p>

opencc-by-4.0Sep 2020View details →
zenodo44/100

Data set for the paper: Intercomparison of ocean colour algorithms for picophytoplankton carbon in the ocean

<p>This dataset contains the phytoplankton carbon,Cphy, obtained from in situ counts of phytoplankton cells using ow cytometry presented in the paper [13]. The location and time of the samples have been matched with the satelllite data in the Ocean&nbsp; Colour Climate Change Initiative (OCCCI) dataset. This dataset is the match between the in situ Cphy and the products from using the OCCCI inputs (i.e. chlorophyll concentration, backscattering coecient, phytoplankton absorption) with 6 different algorithms. This document describes the dataset details: data sources, computation of Cphy, selected data.</p>

opencc-by-4.0Nov 2017View details →
zenodo44/100

Experiment on the performance of different machine learning algorithms for classification - Results

<h2>Results of a short performance study of machine learning algorithms</h2> <h3>Context and methodology</h3> <ul> <li>This data was produced while performing a university project to examine the performance of various machine learning algorithms on different prediction datasets</li> <li>The data serves the purpose of comparing the metrics of performing the different tasks</li> <li>The dataset contains a number of matrices for every classifier and every dataset</li> <li>The data was produced with python scripts provided further down and with the usage of the external datasets: <ul> <li>Membership Woes Dataset (OpenML): <a href="https://api.openml.org/d/44224">https://api.openml.org/d/44224</a></li> <li>Zoo dataset (UCI): <a href="https://doi.org/10.24432/C5R59V">https://doi.org/10.24432/C5R59V</a></li> <li>Breast Cancer Dataset: <a href="https://github.com/moritx/performance-experiment-machine-learning/tree/main/data">https://github.com/moritx/performance-experiment-machine-learning/tree/main/data</a></li> <li>Loan Dataset: <a href="https://github.com/moritx/performance-experiment-machine-learning/tree/main/data">https://github.com/moritx/performance-experiment-machine-learning/tree/main/data</a></li> </ul> </li> </ul> <h3>Technical details</h3> <ul> <li>The data consists of one JSON file</li> <li>The source code for producing this data is available at&nbsp;<a href="https://doi.org/10.5281/zenodo.11085222">https://doi.org/10.5281/zenodo.11085222</a></li> </ul> <h3>Structure of the data</h3> <p>[ {"classifier": ...,<br>"dataset": ...,<br>"hyper_parameters": ...,<br>"cross_validation_results": {<br>&nbsp; &nbsp; "fit_time": {} ,<br>&nbsp; &nbsp; "score_time": ...,<br>&nbsp; &nbsp; "metrics": {}<br>},&nbsp;<br>"holdout_test_results": ...}, ]</p>

opencc-by-4.0Apr 2024View details →
zenodo44/100

Datasets for "Advancing Drug-Target Interactions Prediction: Leveraging a Large-Scale Dataset with a Rapid and Robust Chemogenomic Algorithm"

<p>All datasets required to reproduce the results of publication "Drug-Target Interactions Prediction at Scale: the Komet Algorithm with the LCIdb Dataset"</p>

opencc-by-4.0Feb 2024View details →
zenodo44/100

RAIN: Reinforcement Algorithms for Improving Numerical Weather and Climate Models

<p>This Zenodo repository contains the runs data for the paper <strong>RAIN: Reinforcement Algorithms for Improving Numerical Weather and Climate Models </strong>[<a href="https://arxiv.org/abs/2408.16118" target="_blank" rel="noopener">https://arxiv.org/abs/2408.16118</a>] presented at the NeurIPS 2024 workshop on Tackling Climate Change with Machine Learning, as well as the Master of Research (MRes) report <strong>Towards improving weather and climate models using reinforcement learning</strong> at the University of Cambridge<strong>.</strong> For questions, please contact Pritthijit Nath, <a href="mailto:pn341@cam.ac.uk" target="_blank" rel="noopener">pn341@cam.ac.uk</a>. Full documentation is available in the README.md file of the associated&nbsp;<a href="https://github.com/nathzi1505/climate-rl" target="_blank" rel="noopener">GitHub repo</a>.</p>

opencc-by-4.0Jun 2024View details →
zenodo44/100

Web-based Editor for Entity-Relationship Modeling with SQL Transformation Algorithm

<p>With the daily growth of data in all kinds of sectors, such as <span>Information Technolo</span><span>gies (IT)</span>, healthcare, education, commerce or telecommunication, it becomes important to use a high-performance system to manage all this data in the best possible way. Indeed, in the absence of good data management, it can be difficult for these sectors to prevent data loss, ensure safe maintenance or guarantee data security. For this reason, an effective data management system is the Entity-Relationship model. Indeed, thanks to this model, all kinds of sectors have the possibility of designing and organizing their data in relational databases, thus improving data security and better internal communication. Furthermore, it is interesting to modernize the classic approach of the Entity-Relationship model and its visual representations of data in the present day. The technologies of Augmented and Virtual Reality respond precisely to this challenge of innovation in the Entity-Relationship model. Therefore, the first objective of this thesis is to implement the Entity-Relationship model in a Meta-Modeling Platform for Augmented and Virtual Reality. The second objective is to implement an algorithm that transforms an Entity-Relationship model into SQL statements to create database tables. The Entity-Relationship model is used to design the logic of a database, and the implementation of the algorithm is used to create the database.</p>

opencc-by-4.0Jun 2024View details →
zenodo44/100

Dataset and Source Code for the Paper: A Framework for Developing Strategic Cyber Threat Intelligence from Advanced Persistent Threat Analysis Reports Using Graph-Based Algorithms

<p>Here are the data set and source code related to the paper: "A Framework for Developing Strategic Cyber Threat Intelligence from Advanced Persistent Threat Analysis Reports Using Graph-Based Algorithms"</p> <p>1- aptnotes-downloader.zip : contains source code that downloads all APT reports listed in https://github.com/aptnotes/data and https://github.com/CyberMonitor/APT_CyberCriminal_Campagin_Collections</p> <p>2- apt-groups.zip : contains all APT group names gathered from https://docs.google.com/spreadsheets/d/1H9_xaxQHpWaa4O_Son4Gx0YOIzlcBWMsdvePFX68EKU/edit?gid=1864660085#gid=1864660085 and https://malpedia.caad.fkie.fraunhofer.de/actors&nbsp;and https://malpedia.caad.fkie.fraunhofer.de/actors</p> <p>3- apt-reports.zip : contains all deduplicated APT reports gathered from https://github.com/aptnotes/data and https://github.com/CyberMonitor/APT_CyberCriminal_Campagin_Collections</p> <p>4- countries.zip : contains country name list.</p> <p>5- ttps.zip : contains all MITRE techniques gathered from https://attack.mitre.org/resources/attack-data-and-tools/</p> <p>6- malware-families.zip : contains all malware family names gathered from https://malpedia.caad.fkie.fraunhofer.de/families</p> <p>7- ioc-searcher-app.zip : contains source code that extracts IoCs from APT reports. Extracted IoC files are provided in report-analyser.zip. Original code repo can be found at https://github.com/malicialab/iocsearcher</p> <p>8- extracted-iocs.zip : contains extracted IoCs by ioc-searcher-app.zip</p> <p>9- report-analyser.zip : contains source code that searchs APT reports, malware families, countries and TTPs. I case of a match, it updates files in extracted-iocs.zip.</p> <p>10- cti-transformation-app.zip : contains source code that transforms files in extracted-iocs.zip to CTI triples and saves into Neo4j graph database.</p> <p>11- graph-db-backup.zip : contains volume folder of Neo4j Docker container. When it is mounted to a Docker container, all CTI database becomes reachable from Neo4j web interface. Here is how to run a Neo4j Docker container that mounts folder in the zip:</p> <p>docker run -d --publish=7474:7474 --publish=7687:7687 --volume={PATH_TO_VOLUME}/DEVIL_NEO4J_VOLUME/neo4j/data:/data --volume={PATH_TO_VOLUME}/DEVIL_NEO4J_VOLUME/neo4j/plugins:/plugins --volume={PATH_TO_VOLUME}/DEVIL_NEO4J_VOLUME/neo4j/logs:/logs --volume={PATH_TO_VOLUME}/DEVIL_NEO4J_VOLUME/neo4j/conf:/conf --env 'NEO4J_PLUGINS=["apoc","graph-data-science"]' --env NEO4J_apoc_export_file_enabled=true --env NEO4J_apoc_import_file_enabled=true --env NEO4J_apoc_import_file_use__neo4j__config=true --env=NEO4J_AUTH=none neo4j:5.13.0</p> <h4><strong>web interface: http://localhost:7474</strong></h4> <h4><strong>username: neo4j</strong></h4> <h4><strong>password: neo4j</strong></h4> <p>&nbsp;</p>

opencc-by-4.0Dec 2023View details →
zenodo44/100

Optimisation of business processes tenant distribution in the Cloud with a genetic algorithm

<p>Used data and obtained results for the paper Optimisation of business processes tenant distribution in the Cloud with a genetic algorithm.</p> <p>The reader can find the following files :</p> <ul> <li>configuration_types.csv contains the cloud resource types (the name is the EC2 instance for database, and for the BPM engine separated by an underscore), their price and their capacity</li> <li>tenants_uni.csv contains the customers and their minimum and maximum BPM task throughput</li> <li>results_[number of tenants]_seg.csv files contain the results for the previous heuristic (segmentation only)</li> <li>results_<em>[number of tenants]</em>_ga_<em>[duration]</em>.csv files contain the results for the genetic algorithm coupled to the iterative heuristic tests</li> <li>solver_<em>[number of tenants]</em>_ga_<em>[duration]</em>.csv files contain the results for the genetic algorithm coupled to the restricted model solved tests</li> </ul>

opencc-by-4.0Feb 2018View details →
zenodo44/100

Ground truth recordings for validation of spike sorting algorithms

<p><strong>Ground-truth recordings for validation of spike sorting algorithms</strong><br> &nbsp;</p> <p>This datasets is composed of simultaneous loose patch recordings of Ganglion Cells in mice retina, combined with dense extra-cellular recordings (252 channels). The details of the dataset can be found here <a href="https://elifesciences.org/articles/34518">https://elifesciences.org/articles/34518</a></p> <p><strong>Probe layout</strong></p> <p>The probe layout can be found as mea_256.prb. This is a 16x16 Multi Electrode Array with 30um spacing. Only 252 channels are extra-cellular signals, and the 4 corners are devoted to triggers/sync/juxta.</p> <p><strong>Struture of the data</strong></p> <p>In this dataset, you will find several individual recordings, at max 5min long each (but please do not hesitate to contact us if interested by longer recordings).&nbsp;The extra-cellular data are saved as 16bits unsigned integer, with a variable offset at the beginning of the file. The value of this offset is given, for every datafile, in the additional text file (padding value (see following for more details)).&nbsp;The files have already been filtered with a Butterworth filter of order 3 with a cut-off frequency at 100Hz</p> <p><strong>Structure of a given dataset</strong></p> <p>Please read carefully the following to understand how to load and perform spike sorting with the data. In every .tar.gz file, you will find:</p> <ul> <li>&nbsp;a jpg image, displaying a small chunk of the juxta-cellular signal (top left), with detected peaks and threshold. The extra-cellular spike triggered waveform, across all channels, for the juxta-spike times (top right). In the bottom, you can see the juxta-cellular spikes, for all the detected triggers (left), and on the right the voltage on the channel where the Spike Triggered Average of the extra-cellular waveform is peaking the most.</li> <li>a file .juxta.raw, as float32, with the juxta-cellular trace at 20kHz, no data offset</li> <li>a file .raw, as uint16, with the extra-cellular signals recorded for 256 channels at a sampling rate of 20kHZ. In fact, only 252 channels are extra-cellular signals, the 4 corners of the arrays are devoted to juxta-cellular and sync signals (see probe layout mea_256.prb)</li> <li>a file .triggers.npy containing the spike times of the juxta-cellular spikes, detected using a threshold of k.MAD. The exact value of k can vary on a per dataset basis, and is written in the .txt file (threshold)</li> <li>a .txt file describing some information for a given dataset, such as the threshold value used to detect the spikes, the channel in the raw file where the juxta-cellular signal is located, the minimal value of the peak for the STA (and on which channel it is located), and the header size to read the raw data</li> <li>a .params file, if you want to analyze the data with SpyKING CIRCUS</li> </ul> <p><strong>How to load the raw data in numpy</strong></p> <pre><code class="language-python">#Using the offset value from the txt file, we can load the data with memmap arrays data=numpy.memmap('mydata.raw', dtype='uint16', offset=offset, mode='r') data=data.reshape(len(data)//256, 256) #Then for example, to display the first second of channel 0 one_channel = data[:20000, 0].astype('float32') #If we want to center data around 0 one_channel -= 2**15 - 1 #And if we want to display data in micro volt, we must use the gain factor of 0.1042 provided in the header one_channel *= 0.1042</code></pre> <p>&nbsp;</p>

opencc-by-4.0Mar 2018View details →
zenodo44/100

Datasets for "BB correction" algorithm evaluation and analysis

<p>The datasets contain the raw and processed data relevant to the &quot;BB (black body) correction&quot; method, an algorithm to correct for scattering and systematic bias in quantitative neutron imaging from experimental measurements. Five CT datasets imaging various hydrogenous and non hydrogenous&nbsp; materials are uploaded. The raw data folders for each dataset contains the sample projections over the tomographic range as well as the image used for referencing: the dark current, the open beam files and the additional &quot;BB images&quot;, obtained with an interposed grid of neutron absorbers (BBs) from which scattering and bias components are estimated. The processed folders contain the CT reconstruction with and without the BB correction. All processing was done with the MuhRec open source software (https://zenodo.org/record/1311698#.W30exxixVtg)</p>

opencc-by-4.0Oct 2018View details →
zenodo44/100

Use of EO data to test machine learning algorithms

<p>This data set is composed of one Sentinel-2 image of Darwin City, Australia. The objective of this work is to use EO data to compare the performance of different machine learning algorithms.</p>

opencc-by-4.0Nov 2018View details →
zenodo44/100

CVD2014 - A database for evaluating no-reference video quality assessment algorithms

<p>The CVD video database is developed to provide an useful tool for researchers in the validation and developing processes of no-reference (NR) objective video quality assessment (VQA) algorithms. It consists of 234 videos from five different scenes captured by 78 different cameras (mobile phones, compact camera, video camera, SLR). The subjective experiments are conducted following the Single-Stimulus (SS) procedure to collect ratings of video quality.</p> <p><strong>Setup</strong></p> <p>We implement our experiments according to the Single Stimulus methodology using VQone MATLAB toolboxon high quality monitors (Eizo ColorEdge CG241W) with 1920x1200 pixel resolution in a dark room (ambient light &lt; 20 lux). Video stimuli were displayed at their original size of VGA (640 x 480) or HD (1280 x 720). The subjects viewing distance (80 cm) was controlled by a string hanging from a ceiling and they were instructed to keep their head steady next to it. The monitors were calibrated to according to sRGB (target values were: 6500 K, 80 lux, and gamma 2.2) using EyeOne Pro calibrator (X-rite co.). The laboratory setup is showed in the figure below.</p> <p><strong>Subjects</strong></p> <p>Subjects (n = 30, 30, 28, 33, 30, 32 and 27 for Tests 1 - 7 respectively) were na&iuml;ve in a sense that they did not study or work with image quality or related fields. They were recruited through student mailing lists consisting mainly humanities and behavioral science students. Subjects&rsquo; vision was controlled for the near visual acuity, near contrast vision (near F.A.C.T.) and color vision (Farnsworth D15) before the participation. They received movie tickets as a reward.</p> <p><strong>Procedure</strong></p> <p>Subjects evaluated one video sample at a time and all video samples of one scene were presented in a row. The order of video samples and scenes was randomized. Subjects had the option to view video samples again as many times as they wanted.</p> <p><strong>Data</strong></p> <p>The results are processed and reported in the form of Mean Opinion Score (MOS) for the tested video samples. In addition, we provide the whole raw data from the subjective experiments instead of just pre-calculated mean opinion scores from each video sample. This allows further analyses to be made by those who wish to use this database and gives them better opportunity to utilize the data to its full potential.</p> <p>Realignment study (test 7) contains the data from the additional study in which the mappings from the test and scene specific quality scales (test 1-6) to the global quality scale were formed. The global scale is valuable when studying and developing VQA algorithms. With the global scale, all of the samples (234 video samples in the case of the CVD2014) are in the same scale, and the performance analysis for algorithms can be conducted with a high number of samples.</p> <p><strong>If you use this database in your research, we kindly ask that you follow The Copyright notice below and cite the following paper:</strong></p> <p>&nbsp;</p> <p>M. Nuutinen, T. Virtanen, M. Vaahteranoksa, T. Vuori, P. Oittinen and J. H&auml;kkinen, &quot;CVD2014&mdash;A Database for Evaluating No-Reference Video Quality Assessment Algorithms,&quot; in <em>IEEE Transactions on Image Processing</em>, vol. 25, no. 7, pp. 3073-3086, July 2016. doi: 10.1109/TIP.2016.2562513</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>-----------COPYRIGHT NOTICE STARTS WITH THIS LINE------------</p> <p>Copyright (c) 2014 The University of Helsinki<br> All rights reserved.</p> <p>Permission is hereby granted, without written agreement and without license or royalty fees, to use, copy, modify, and distribute this database (the videos, the images, the results and the source files) and its documentation for any purpose, provided that the copyright notice in its entirely appear in all copies of this database, and the original source of this database,Visual Cognition research group (www.helsinki.fi/psychology/groups/visualcognition/index.htm) and the Institute of Behavioral Science (www.helsinki.fi/ibs/index.html) at the University of Helsinki (www.helsinki.fi/university/), is acknowledged in any publication that reports research using this database. Individual videos and images may not be used outside the scope of this database (e.g. in marketing purposes) without prior permission.</p> <p>The database and our paper are to be cited in the bibliography as: M. Nuutinen, T. Virtanen, M. Vaahteranoksa, T. Vuori, P. Oittinen and J. H&auml;kkinen, &quot;CVD2014&mdash;A Database for Evaluating No-Reference Video Quality Assessment Algorithms,&quot; in <em>IEEE Transactions on Image Processing</em>, vol. 25, no. 7, pp. 3073-3086, July 2016.<br> doi: 10.1109/TIP.2016.2562513</p> <p>-----------------------------------------------------------------------------</p> <p>LIMITATION OF LIABILITY</p> <p>UNIVERSITY OF HELSINKI SHALL IN NO CASE BE LIABLE IN CONTRACT, TORT OR OTHERWISE FOR ANY LOSS OF REVENUE, PROFIT, BUSINESS OR GOODWILL OR ANY DIRECT, INDIRECT, SPECIAL, CONSEQUENTIAL, INCIDENTAL OR PUNITIVE COST, DAMAGES OR EXPENSE OF ANY KIND HOWEVER CAUSED OR HOWEVER ARISING UNDER OR IN CONNECTION WITH THE USE OF THIS DATABASE.</p> <p>THE UNIVERSITY OF HELSINKI SPECIFICALLY DISCLAIMS ANY WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE. THE DATABASE PROVIDED HEREUNDER IS ON AN &quot;AS IS&quot; BASIS, AND THE UNIVERSITY OF HELSINKI HAS NO OBLIGATION TO PROVIDE MAINTENANCE, SUPPORT, UPDATES, ENHANCEMENTS, OR MODIFICATIONS.</p> <p>THIS AGREEMENT SHALL BE CONSTRUED AND INTERPRETED IN ACCORDANCE WITH THE LAWS OF FINLAND, EXCLUDING ITS RULES FOR CHOICE OF LAW.</p> <p>-----------COPYRIGHT NOTICE ENDS WITH THIS LINE------------</p>

opencc-by-4.0Jun 2016View details →
zenodo44/100

Position Weighted Backpressure Control algorithms Codes and Vissim Simulations

<p>The attached file&nbsp;contains the simulation implementation of the Position Weighted Back Pressure Control algorithm, tested for various traffic demand scenarios. The comparison simulations with Fixed time, Back-pressure and Capacity aware back-pressure controls are also supplied. For more information about the testing procedures and analysis, see the article:</p> <p>Li, L. and Jabari, S.E., 2018. Position weighted backpressure intersection control for connected urban networks.&nbsp;<em>arXiv preprint arXiv:1810.11406</em>.</p> <p>The usage and modification of the attached files are subject to the citation of the above article or the dataset DOI:&nbsp;10.5281/zenodo.3236757.</p> <p>&nbsp;</p>

opencc-by-4.0May 2019View details →
zenodo44/100

Time series of turbidity in Northern Patagonia using the Nechad algorithms (v2009 and v2016) at 665 nm. Time series 2016-2020.

<p>Time series of turbidity in Northern Patagonia using the Nechad algorithms (v2009 and v2016) at 665 nm. Time series 2016-2020.</p> <p>Our study aimed to evaluate the spatio-temporal variability of turbidity from Sentinel-2 (S2) images in the Reloncav&iacute; sound and fjord, in Northern Patagonia, Chile, a coastal ecosystem that is intensively used by finfish and shellfish aquaculture. To this end, we downloaded 123 S2 images and assembled a five-year time series (2016-2020) covering five study sites (R1 to R5) located along the axis of the fjord and seaward into the sound. We used Acolite to perform the atmospheric correction and estimate turbidity with two algorithms proposed by Nechad et al. (2009, 2016 Nv09 and Nv16, respectively).</p> <p>Columns (R) represent the spatial distribution of study sites (see Figure 2).</p> <p>Link: https://doi.org/10.1016/j.ecoinf.2024.102814</p> <p>For more information see materials and methods.</p> <p>Nv2009 or Nv09 are the results obtained for the Nechad algorithm version 2009. Similar to Nv2016 or Nv16 are the results obtained for the Nechad algorithm version 2016.</p>

opencc-by-4.0Sep 2024View details →
zenodo44/100

Data set for risk management in the allocation of vehicles to tasks in transport companies using a heuristic algorithm

<p>The purpose of this dataset is to enable the replication of the research results presented in the article: Izdebski, M. (2023). Risk management in the allocation of vehicles to tasks in transport companies using a heuristic algorithm. Archives of Transport, 67(3), 139-153. https://doi.org/10.5604/01.3001.0053.7463 - published online: 2023-09-30, which discusses the allocation problem of vehicles to tasks, taking into account risk issues.</p> <p>Dataset contains:</p> <ul> <li>Readme.txt: description of the dataset</li> <li>InputData.xlsx: Contains the input data used in the model</li> <li>DistributionFit.xlsx: Compliance testing and distribution parameters for road accidents of any type and collision-type</li> <li>OutputAssignment.xlsx: Results of assignment and alghoritm tests</li> </ul> <p>The dataset was created as part of the E-Laas project (Energy optimal urban logistics As A Service).<br>Project implemented as part of the call ERA-NET Cofund Urban Accessibility and Connectivity (ENUAC China Call) organized by JPI Urban Europe and the National Natural Science Foundation of China (NSFC). This project has received funding from the European Union&rsquo;s Horizon 2020 research and innovation programme under grant agreement No 875022.<br>&nbsp;E-Laas project is carried out in an international consortium. Project coordinator in Europe: Chalmers University of Technology (Sweden), project coordinator in China: Shanghai University (China), consortium members: Tsinghua University (China), Warsaw University of Technology (Poland), cooperation partners: Stockholms stad, Trafikkontoret (Sweden), ParkUnload (Spain), Metropolis GZM (Poland), Shanghai Urban-Rural Construction and Transportation Department (China), Volvo Group Trucks Technology and Operations (Sweden).<br>- The Chinese part of the project is funded by National Natural Science Foundation of China.<br>- The Swedish part of the project is funded by Swedish Energy Agency.<br>- The Polish part of the project is funded by the National Science Centre, Poland (project no. 2022/04/Y/ST8/00134). The value of the co-financing is PLN 878,107.00. Project duration 27/04/2023 - 26/04/2026 (36 months).</p>

opencc-zeroSep 2024View details →
zenodo44/100

Fuτure - dataset for studies, development, and training of algorithms for reconstructing and identifying hadronically decaying tau leptons

<h1>&nbsp;Data description</h1> <h2>MC Simulation</h2> <p><br>The <strong>Fu&tau;ure</strong> dataset is intended for studies, development, and training of algorithms for reconstructing and identifying hadronically decaying tau leptons. The dataset is generated with Pythia 8, with the full detector simulation being performed by Geant4 with the CLIC-like detector setup CLICdet (CLIC_o3_v14) setup. Events are reconstructed using the Marlin reconstruction framework and interfaced with Key4HEP. Particle candidates in the reconstructed events are reconstructed using the PandoraPF algorithm.</p> <p>In this version of the dataset no &gamma;&gamma; -&gt; hadrons background is included.</p> <h2>Samples</h2> <p><br>This dataset contains e+e- samples with Z-&gt;&tau;&tau;, ZH,H-&gt;&tau;&tau; and Z-&gt;qq events, with approximately 2 million events simulated in each category.</p> <p>The following processes e+e- were simulated with Pythia 8 at sqrt(s) = 380 GeV:</p> <ul> <li>p8_ee_qq_ecm380 [Z -&gt; qq events]</li> <li>p8_ee_ZH_Htautau [ZH -&gt; Ztautau]</li> <li>p8_ee_Z_Ztautau_ecm380 [ZH -&gt; Ztautau]</li> </ul> <p>The .root files from the MC simulation chain are eventually processed by the software found in&nbsp;<a href="https://github.com/HEP-KBFI/ml-tau-en-reg">Github</a> in order to create flat ntuples as the final product.</p> <h2><br>Features</h2> <p><br>The basis of the ntuples are the particle flow (PF) candidates from PandoraPF. Each PF candidate has four momenta, charge and particle label (electron / muon / photon / charged hadron / neutral hadron). The PF candidates in a given event are clustered into jets using generalized kt algorithm for ee collisions, with parameters p=-1 and R=0.4. The minimum pT is set to be 0 GeV for both generator level jets and reconstructed jets. The dataset contains the four momenta of the jets, with the PF candidates in the jets with the above listed properties.</p> <p>Additionally, a set of variables describing the tau lifetime are calculated using the software in <a href="https://github.com/HEP-KBFI/ml-tau-en-reg">Github</a>. As tau lifetime is very short, these variables are sensitive to true tau decays.&nbsp;In the calculation of these lifetime variables, we use a linear approximation.</p> <p>In summary, the features found in the flat ntuples are:</p> <p>&nbsp;</p> <table> <tbody> <tr> <td><strong>Name</strong></td> <td><strong>Description</strong></td> </tr> <tr> <td>reco_cand_p4s</td> <td>4-momenta per particle in the reco jet.</td> </tr> <tr> <td>reco_cand_charge</td> <td>Charge per particle in the jet.</td> </tr> <tr> <td>reco_cand_pdg</td> <td>PDGid per particle in the jet.</td> </tr> <tr> <td>reco_jet_p4s</td> <td>RecoJet 4-momenta.</td> </tr> <tr> <td>reco_cand_dz</td> <td>Longitudinal impact parameter per particle in the jet. For future steps. Fill value used for neutral particles as no track parameters can be calculated.</td> </tr> <tr> <td>reco_cand_dz_err</td> <td>Uncertainty of the longitudinal impact parameter per particle in the jet. For future steps. Fill value used for neutral particles as no track parameters can be calculated.</td> </tr> <tr> <td>reco_cand_dxy</td> <td>Transverse impact parameter per particle in the jet. For future steps. Fill value used for neutral particles as no track parameters can be calculated.</td> </tr> <tr> <td>reco_cand_dxy_err</td> <td>Uncertainty of the transverse impact parameter per particle in the jet. For future steps. Fill value used for neutral particles as no track parameters can be calculated.</td> </tr> <tr> <td>gen_jet_p4s</td> <td>GenJet 4-momenta. Matched with RecoJet within a cone of radius dR &lt; 0.3.</td> </tr> <tr> <td>gen_jet_tau_decaymode</td> <td>Decay mode of the associated genTau. Jets that have associated leptonically decaying taus are removed, so there are no DM=16 jets. If no GenTau can be matched to GenJet within dR &lt; 0.4, a fill value is used.</td> </tr> <tr> <td>gen_jet_tau_p4s</td> <td>Visible 4-momenta of the genTau. If no GenTau can be matched to GenJet within dR&lt;0.4, a fill value is used.</td> </tr> </tbody> </table> <p>The ground truth is based on stable particles at the generator level, before detector simulation. These particles are clustered into generator-level jets and are matched to generator-level &tau; leptons as well as reconstructed jets. In order for a generator-level jet to be matched to generator-level &tau; lepton, the &tau; lepton needs to be inside a cone of dR = 0.4. The same applies for the reconstructed jet, with the requirement on dR being set to dR = 0.3. For each reconstructed jet, we define three target values related to &tau; lepton reconstruction:</p> <ul> <li>&nbsp;a binary flag <strong>isTau</strong> if it was matched to a generator-level hadronically decaying &tau; lepton. <strong>gen_jet_tau_decaymode</strong> of value -1 indicates no match to generator-level hadronically decaying &tau;.</li> <li>&nbsp;the categorical decay mode of the &tau; <strong>gen_jet_tau_decaymode</strong> in terms of the number of generator level charged and neutral hadrons. Possible <strong>gen_jet_tau_decaymode</strong> are {0, 1, . . . , 15}.</li> <li>&nbsp;if matched, the visible (neglecting neutrinos), reconstructable pT of the &tau; lepton. This is inferred from the <strong>gen_jet_tau_p4s</strong></li> </ul> <h2>Contents:</h2> <ul> <li>qq_test.parquet</li> <li>qq_train.parquet</li> <li>zh_test.parquet</li> <li>zh_train.parquet</li> <li>z_test.parquet</li> <li>&nbsp;z_train.parquet</li> <li>data_intro.ipynb</li> </ul> <h2>Dataset characteristics</h2> <p>&nbsp;</p> <table> <tbody> <tr> <td><strong>File</strong></td> <td><strong># Jets</strong></td> <td><strong>Size</strong></td> </tr> <tr> <td>z_test.parquet</td> <td> <pre>870 843</pre> </td> <td>171 MB</td> </tr> <tr> <td>z_train.parquet</td> <td> <pre>3 483 369</pre> </td> <td>681 MB</td> </tr> <tr> <td>zh_test.parquet</td> <td> <pre>1 068 606</pre> </td> <td>213 MB</td> </tr> <tr> <td>zh_train.parquet</td> <td> <pre>4 274 423</pre> </td> <td>851 MB</td> </tr> <tr> <td>qq_test.parquet</td> <td> <pre>6 366 715</pre> </td> <td>1.4 GB</td> </tr> <tr> <td>qq_train.parquet</td> <td> <pre>25 466 858</pre> </td> <td>5.6 GB</td> </tr> </tbody> </table> <p>The dataset consists of 6 files of 8.9 GB in total.</p> <h2>How can you use these data?</h2> <p>The .parquet files can be directly loaded with the Awkward Array Python library.<br>An example how one might use the dataset and the features is given in&nbsp;<strong>data_intro.ipynb</strong></p>

opencc-by-4.0Jul 2024View details →
zenodo44/100

Overview: Simulation Module + Optimisation Algorithm + LCA Support Tool

<p>Overview of the Model2Bio elements: Simulation Module + Optimisation Algorithm + LCA Support Tool</p> <p>Platforms showing all the processes of the Model2Bio project. From&nbsp;the beginning of the Simulation Module,&nbsp;the&nbsp;variables for the production line models (agriculture and food production industries) to bio-products created from the residues.</p>

opencc-by-4.0Jun 2021View details →
zenodo44/100

Dataset for A Fast Frozen Phonon Algorithm Using Mixed Static Potentials

<p>Multislice simulation dataset for &quot;A Fast Frozen Phonon Algorithm Using Mixed Static Potentials&quot;</p>

opencc-by-4.0May 2021View details →
zenodo44/100

Dataset for Algorithms and Complexity for Counting Configurations in Steiner Triple Systems

<p>This dataset contains the classification of full n-line configurations (for all n &lt;= 13, filename: &quot;full_line_config_&lt;n&gt;.txt.gz&quot;) and w_3 configurations (for all w &lt;= 16, filename: &quot;w_3_config_&lt;w&gt;.txt.gz&quot;) together with the sizes of minimum generating sets. Each file lists &quot;m&lt;s&gt;&quot; so that s is the size of the minimum generating set of the subsequent configuration, which is denoted by writing the points of each of its lines row-wise. For example</p> <p>m3<br> 0 1 4<br> 0 2 6<br> 0 3 5<br> 1 2 5<br> 1 3 6<br> 2 3 4<br> 4 5 6</p> <p>is the Fano plane and its minimum generating set has size 3.</p> <p>Additionally, the file &quot;fulllineconjecture.txt.gz&quot; contains the 623 Steiner triple systems of order 25 (i.e., all rows which contain curly brackets) used in Theorem 7 in the paper below, followed by a row starting with 1 and then describing the number of occurrences of all 179 full n-line configurations for n &lt;= 8, i.e., first the number of occurrences of Pasch configurations, then mitre configurations, then the 5 full 6-line configurations, the 19 full 7-line configurations, and finally the 153 full 8-line configurations contained in the STS(25) in the preceding row. The ordering follows the ordering within the files &quot;full_line_config_&lt;n&gt;.txt.gz&quot;. This file is built so that omitting all lines with curly brackets is a valid gap code and results in a prove of said theorem (i.e., zgrep -v &quot;{&quot; fulllineconjecture.txt.gz | gap yields 180).</p> <p><br> Further details can be found in the corresponding publication</p> <p>&quot;Algorithms and Complexity for Counting Configurations in Steiner Triple Systems&quot;</p> <p>by Daniel Heinlein and Patric R. J. &Ouml;sterg&aring;rd.</p> <p>All files are compressed with gzip.</p>

opencc-by-4.0Sep 2021View details →
zenodo44/100

Direct sun retrievals of nitrogen dioxide (NO2) total columns from Brewer #067, Rome, Italy (reprocessed with algorithm BNALG2)

<p>Cloud-screened and quality-filtered direct sun retrievals of nitrogen dioxide (NO2) vertical column densities (VCDs) derived from MkIV Brewer #067 measurements in Rome (wavelengths 425.02, 431.40, 437.35, 442.83, 448.08, and 453.20 nm) and processed using the Brewer Nitrogen Dioxide Algoritm BNALG2. Calibration is carried out with Bootstrap Estimation techniques. The values represent averages of 5 samples.</p> <p>In the latest version, days with obviously erroneous data (NO2 VCD &gt; 99.9% percentile) have been removed.</p> <p>A detailed description of the method has been accepted as a research article by the ESSD journal (H. Di&eacute;moz et al., Advanced NO2 retrieval technique for the Brewer spectrophotometer applied to the 20-year record in Rome, Italy, Earth Syst. Sci. Data, 2021).</p>

opencc-by-4.0Apr 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record