Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
311
datasets available to search
ShareScore release 0.9.0
Dataset results
311 results for “new dataset”
New Challenges in Point Cloud Visual Quality Assessment: A Systematic Review (Dataset)
<p>This dataset is a collection of annotated information on the scientific papers screened and analyzed for the systematic review of the literature in Point Cloud Visual Quality Assessment. </p> <p>The data is structured as follows:</p> <ul> <li>General information <ul> <li>Document title</li> <li>Authors</li> <li>Year of publication</li> <li>Venue (Conference or Journal title)</li> <li>Citations (number)</li> <li>URL/DOI</li> </ul> </li> </ul> <ul> <li>About the content <br> <ul> <li>Content Type: Point clouds (PC), Colored Point clouds (CPC), Meshes, Dynamic Point Clouds (DPC)</li> <li>Content source: Source of the content used in a subjective QA test or the evaluation of one or more QA metrics</li> </ul> </li> </ul> <ul> <li>About metric benchmarks <ul> <li>Subjective Ground-truth Data: Dataset(s) Source of the subjective scores used as ground-truth in a QA metric benchmark</li> <li>Assessed Metrics: Types of metrics assessed in a benchmark (JPEG standards, IQM, NR, State-of-the-art, others)</li> <li>Performance Measures: PLCC, SROCC, KRCC, RMSE, OR, others</li> </ul> </li> </ul> <ul> <li>About Objective QA metrics <ul> <li>Metric: Name given to the metric introduced in this paper</li> <li>Base: 3D-based or Projection-based</li> <li>Categories: Categories that characterize the approach of the proposed metric (Feature-based, Learning-Based, Perceptual-based, IQM, others) </li> <li>Reference: Full-Reference (FR), Reduced-Reference (RR) or No-Reference (NR)</li> </ul> </li> </ul> <ul> <li>About Subjective QA experiments <ul> <li>Display: Type of display (2D, 3D, AR, MR, VR) and interaction approach (passive, interactive, 3DoF, 6DoF) used in the described experiment.</li> <li>Rendering: Type of rendering used to display the stimuli (Points, Squares, Cubes, Surface)</li> <li>Lab/Remote: The experiment was run in one or more lab environments, or remotely (Lab, Cross-Lab, Remote)</li> <li>Rating: Subjective rating methodology used in the experiment (ACR, DSIS, PWC, others)</li> <li>Dataset: Name of the new subjective dataset if the experiment's results were published.</li> <li>Observers: Number of observers </li> <li>Distortion type: Types of distortions applied to the stimuli and assessed in the experiment</li> </ul> </li> </ul>
New datasets obtained from experimental installations with centralized control
<p>The dataset contains the data that local controller 1 (LC1) received from local controller 2 (LC2) during normal system operation. The system was created and it is located in the Laboratory for Manufacturing Automation at the Faculty of Mechanical Engineering, University of Belgrade. The system is based on a smart sensor (electromagnetic linear encoder with local controller) and a smart actuator (rodless pneumatic cylinder with electro-pneumatic pressure regulator and local controller), and the main goal is to achieve the desired position of the piston on the pneumatic cylinder. The dataset includes 7 signals that represent a combination of different piston trajectories. Each signal was recorded during the 200 minutes of piston movement along a defined trajectory, where the length of each signal is 400,000 samples. Table 2 shows the list of collected signals, whereas a detailed description of the system can be found in [1].</p> <p>This dataset was developed with the support of the Science Fund of the Republic of Serbia, Grant No. 6523109, AI - MISSION4.0, 2020-2022.</p>
THÖR-Magni (Demo Subset): a new multi-modal context-rich dataset of human-robot motion
<p>The Magni Human Motion Dataset provides high-quality tracking information from motion capture, eye-gaze trackers, and on-board robot sensors in a semantically rich environment. To induce natural behavior of recorded participants, we utilized loosely scripted task assignment, which induced participants to navigate through a dynamic laboratory environment in a natural and purposeful way. The dataset sets a high-quality standard as realistic and accurate data is enhanced with semantic information, enabling development of new algorithms that rely not only on tracking information but also on contextual cues of moving agents, static and dynamic environments.</p> <p> </p> <p>Link to dashboard that uses the data: https://magni-dash.streamlit.app/</p> <p><br> Here we publish a subset of the final dataset, to accompany the presentation at the 2023 IEEE International Conference on Robotics and Automation (ICRA)</p>
Dataset for : A New Era in Software Security: Towards Self-Healing Software via Large Language Models and Formal Verification
<p>We present a novel solution combining Large Language Model (LLM) capabilities with Formal Verification strategies to falsify and automatically repair software vulnerabilities. Initially, we employ Bounded Model Checking (BMC) to locate the software vulnerability and derive a counterexample. Relying on mathematical proofs, counterexamples provide evidence that the system behaves incorrectly or contains a vulnerability, thereby preventing the generation of false positive alerts. The counterexample that has been detected, along with the source code, are provided to the LLM engine. Our approach involves establishing a specialized prompt language for conducting code debugging and generation to understand the vulnerability's root cause and repair the code. Finally, we use BMC to verify the corrected version of the code generated by the LLM. As a proof of concept, we create \esbmcai based on the Efficient SMT-based Context-Bounded Model Checker (ESBMC) and a pre-trained Transformer model, specifically gpt-3.5-turbo, to detect and fix errors in C programs. We generated a dataset comprising $1{,}000$ C code samples, each consisting of $20$ to $50$ lines of C code. Experimental results show that our proposed method achieved an impressive success rate of up to $80$\% in repairing vulnerable code, encompassing buffer overflow, arithmetic overflow, and pointer dereference failures. To our knowledge, \esbmcai represents the first proposal for a pioneering initiative to integrate a Large Language Model (LLM) with software model checking. We advocate that this automated approach has the potential to incorporate into the software development lifecycle's continuous integration and deployment (CI/CD) process. </p> <p> </p> <p>The uploaded dataset contains 1000 codes, each comprising 20 to 50 lines of C code generated with gpt-3.5-turbo. The material also consists of a version of ESBMC statically compiled with all dependencies, a classifier script, and the output file.</p> <p> </p> <p> </p>
[Luminescence Dataset] The loess sequence of Dolní Věstonice, Czech Republic: A new OSL‐based chronology of the last climatic cycle
<p><strong>Dolní Věstonice - Luminescence Dataset</strong></p> <table> <tbody> <tr> <td> <p><strong>Applicable Licence</strong></p> </td> <td> <p>CC-BY-NC</p> </td> </tr> <tr> <td> <p><strong>Data curators:</strong></p> </td> <td> <p>Markus Fuchs, Sebastian Kreutzer</p> </td> </tr> <tr> <td> <p><strong>Reference original work</strong></p> </td> <td> <p>Fuchs, M., Kreutzer, S., Rousseau, D. D., Antoine, P., Hatté, C., Lagroix, F., Moine, O., Gauthier, C., Svoboda, J., and Lisá, L.: The loess sequence of Dolní Věstonice, Czech Republic: A new OSL‐based chronology of the last climatic cycle, Boreas, 42, 664–677, https://doi.org/10.1111/j.1502-3885.2012.00299.x, 2013.</p> </td> </tr> <tr> <td> <p><strong>How to refer/cite this dataset</strong></p> </td> <td> <p>Please the information auto-generated by Zendo</p> </td> </tr> </tbody> </table> <p><strong>Scope</strong></p> <p>This dataset contains the original luminescence data used to create the chronostratigraphy of the Dolní Věstonice loess profile. By publishing this dataset, we aim to support studies on the general characteristics of luminescence behaviours of natural minerals and support the FAIR guidelines for data sharing. This data is primary <strong>unmodified measurement data with all the possible errors and typos!</strong></p> <p><em>Please note: The dataset was compiled with the greatest care. However, the dataset comes without any guarantee. Human errors are always possible. If you feel that, after reading the original study and this document, the metadata describing the dataset is insufficient, please contact the data curators so that this document can be updated accordingly. Contrary, the primary cannot be updated/modified!</em></p> <p><strong>Dataset structure</strong></p> <p>The dataset consists of sequence files (<code>SEQ</code>) and measurement data (<code>BIN</code>). To each <code>.bin</code> file, there should be one corresponding sequence file. Measurement data are tagged with sample names in the <code>.bin</code> file, e.g., unless, in case of mistakes, samples can be distinguished in the file. </p> <ul> <li> <p><code>...BIN/</code></p> <ul> <li> <p><code>BIN/A-VALUE/:</code>measurement data with <em>a</em>-value measurements; the alpha irradiation was partly done on an external source</p> </li> <li> <p><code>BIN/MAIN/</code> measurement data with data used to estimate the equivalent dose of each sample</p> </li> <li> <p><code>BIN/PREHEAT_PLATEAU</code> files with combined preheat and dose recovery test results</p> </li> <li> <p><code>BIN/MISC</code> various additional measurements as indicated by the name</p> </li> </ul> </li> <li> <p><code>...SEQ/</code></p> <ul> <li> <p><code>SEQ/A-VALUE/</code></p> </li> <li> <p><code>SEQ/MAIN</code></p> </li> <li> <p><code>SEQ/PREHEAT_PLATEAU</code></p> </li> <li> <p><code>SEQ/MISC</code></p> </li> </ul> </li> <li> <p><code>...DE_CSV_EXPORT</code> The extracted equivalent dose values from the measurements with uncertainties in s. </p> </li> </ul> <p><em>Please note that samples were partly measured on different machines. All files are selections; in the course of the study, we carried out various additional measurements, in particular, preliminary tests; those data are not included in the dataset to keep the dataset comprehensible. </em></p> <p><strong>Used abbreviations</strong></p> <p>Abbreviations as used in the measurement and sequence file names. For technical details, we refer to the original study and the reference therein.</p> <table> <tbody> <tr> <td> <p>ABBREVIATION</p> </td> <td> <p>TERM</p> </td> <td> <p>DESCRIPTION</p> </td> </tr> <tr> <td> <p><code>CG</code></p> </td> <td> <p>coarse grain</p> </td> <td> <p>refers to the used grain size fraction, here: 90-200 µm</p> </td> </tr> <tr> <td> <p><code>DRT</code></p> </td> <td> <p>dose-recovery test</p> </td> <td> <p>measurement data with dose recovery test results; in this study, such measurements were carried out in combination with different preheat settings</p> </td> </tr> <tr> <td> <p><code>FG</code></p> </td> <td> <p>fine grain</p> </td> <td> <p>refers to the used grain size fraction, here: 4-11 µm</p> </td> </tr> <tr> <td> <p><code>IRSLT</code></p> </td> <td> <p>Infrared stimulated test</p> </td> <td> <p>test for feldspar contamination</p> </td> </tr> <tr> <td> <p><code>LM</code></p> </td> <td> <p>linear modulation</p> </td> <td> <p>linearly modulated optically stimulated luminescence measurements</p> </td> </tr> <tr> <td> <p><code>MAIN</code></p> </td> <td> <p>main measurement</p> </td> <td> <p>usually, the measurement used to estimate the equivalent dose</p> </td> </tr> <tr> <td> <p><code>MG</code></p> </td> <td> <p>medium grain</p> </td> <td> <p>refers to the used grain size fraction, here: 38-63 µm</p> </td> </tr> <tr> <td> <p><code>mz</code></p> </td> <td> <p>Moritz</p> </td> <td> <p>Name of the used Risø OSL/TL DA-15 reader</p> </td> </tr> <tr> <td> <p><code>mx</code></p> </td> <td> <p>Max</p> </td> <td> <p>Name of the used Risø OSL/TL DA-15 reader</p> </td> </tr> <tr> <td> <p><code>NEU</code></p> </td> <td> <p>neu</p> </td> <td> <p>a German word translating to 'new'</p> </td> </tr> <tr> <td> <p><code>oDRT</code></p> </td> <td> <p>oDRT</p> </td> <td> <p>a typo for just <code>DRT</code></p> </td> </tr> <tr> <td> <p><code>PreaHeat</code></p> </td> <td> <p>preheat test</p> </td> <td> <p>test against different preheat temperatures</p> </td> </tr> <tr> <td> <p><code>Q</code></p> </td> <td> <p>quartz</p> </td> <td> <p>the measured mineral composition, often in combination with <code>CG</code> , <code>MG</code> , or F<code>G</code> . Example: <code>FGQ</code> : fine grain quartz</p> </td> </tr> </tbody> </table>
A Benchmark Dataset for Semi-Automatic Seismic Interpretation Based on a New Zealand's Seismic Survey
<p>Open access to curated datasets positively impacts on scientific research of machine learning and deep learning techniques. It is a fact that benchmarks and public datasets prepared for data science assist researchers interested in evaluating, testing, and building new data-driven methodologies for specific domain areas.</p> <p>In geosciences, there has been a remarkable growth of public datasets arranged to address machine learning challenges related to the oil and gas industry, particularly for reserves exploration and data interpretation. </p> <p>For these reasons, we present the Taranaki dataset, which is a collection of seismic horizons interpreted for a seismic stratigraphic interpretation study in the Taranaki Basin, offshore New Zealand. This data comprises fourteen seismic horizons that mark stratigraphic discordances in the Tui-3D seismic dataset. We annotated five seismic horizons on 33 inline sections and nine horizons on 19 crossline sections.</p> <p>Besides, we present the results of a series of experiments that compare a method of interpolation and a method of deep learning for seismic segmentation. The deep learning experiments evaluated the result of different image tile sizes to train the model, which is presented separately in this dataset. </p> <p>Finally, we evaluated both methodologies to interpret the horizons of this dataset in selected seismic sections. Also, we assessed the absolute error of each method with the ground truth interpretations proposed in this dataset.</p>
Geochemical sediment fingerprinting dataset from Oroua river catchment, New Zealand
<p>Geochemical dataset collected to determine key source contributions to overbank sediment deposition for specific particle size fractions as described in "Vale, S., Smith, H., Matthews, A., & Boyte, S. (2020). Determining sediment source contributions to overbank deposits within stopbanks in the Oroua River, New Zealand, using sediment fingerprinting. <i>Journal of Hydrology (New Zealand)</i>, <i>59</i>(2), 147-172." </p>
Toxic Content Detection in online social networks: a new dataset from Brazilian Reddit Communities
<p>This is new dataset of 2,500 manually annotated examples of comments extracted from the top 10 largest Brazilian subreddits on Reddit. The dataset has been annotated by crowd-sourcing efforts with contributions from the departments of computer science (DCC) and the linguistic group @ UFMG. As part of our contribution to the toxicity automatic detection and moderation of online social networks, we're making the dataset public for research.</p> <h3>Dataset</h3> <p>The dataset contains 2,500 manually annotated comments from the most popular brazilian communities on Reddit. The data sampling proccess was a stratified sampling by the number of generated publications by subreddit and the month of publication. The list of communities collected is presented below. The collected data period ranges from January 2022 to December 2022.</p> <p> </p> <table> <tbody> <tr> <td><strong>Subreddit</strong></td> <td><strong>Posts</strong></td> <td><strong>Comments</strong></td> </tr> <tr> <td>r/brasil</td> <td>110,829 </td> <td>2,136,866</td> </tr> <tr> <td>r/desabafos</td> <td>115,876</td> <td>1,211,643</td> </tr> <tr> <td>r/futebol</td> <td>35,826</td> <td>1,214,412</td> </tr> <tr> <td>r/saopaulo</td> <td>7,308</td> <td>81,969</td> </tr> <tr> <td>r/eu_nvr</td> <td>12,631</td> <td>188,620</td> </tr> <tr> <td>r/botecodoreddit</td> <td>7,059</td> <td>57,298</td> </tr> <tr> <td>r/conversas</td> <td>21,967</td> <td>326,061</td> </tr> <tr> <td>r/investimentos</td> <td>9,756</td> <td>141,823</td> </tr> <tr> <td>r/tiodopave</td> <td>2,371</td> <td>11,584</td> </tr> <tr> <td>r/brasilivre</td> <td>67,301</td> <td>1,219265</td> </tr> <tr> <td>Total</td> <td>390,924</td> <td>6,589,541</td> </tr> </tbody> </table> <p> </p> <h3><strong>Annotation proccess</strong></h3> <p>The annotators were divided into groups of raters and each group was assigned a batch of comments to label. The raters were then asked to label a comment as <strong>Toxic</strong>, <strong>Non-toxic</strong>, <strong>I do not know</strong> and <strong>Missing info</strong>. During the annotation process, the raters were encouraged to assign one of the uncertain labels when they're not sure about the toxicity of a comment or the context is missing. </p> <h3>Available data</h3> <p>The dataset is available as csv file and the label was assigned as a majority vote among the raters. The available data are the original collected comment id and body. The label was created from the original classification from the annotators. No data processing has been done on this version of the dataset. The overall schema of the dataset if presented below.</p> <p>- <strong>id</strong>: The unique identifier of the comment on the Reddit platform<br>- <strong>body</strong>: The original comment text publication<br>- <strong>is_toxic</strong>: The final label of a given comment. The label is <strong>0</strong> for non-toxic comments, <strong>1</strong> for toxic comments and <strong>-1</strong> for comments where the raters disagreed about the toxicity.</p>
Melting glass fibres recovered from wind turbine blades into new glass fibres for wind turbine blades: Dataset
<p><strong><span>Melting glass fibres recovered from wind turbine blades into new glass fibres for wind turbine blades: Dataset.</span></strong></p> <p><span>In the study titled “Melting glass fibres recovered from wind turbine blades into new glass fibres for wind turbine blades”, four different types of glass fibres were manufactured with varying fractions of recycled fibre powder, 0 wt%, 1.64 wt%, 1.90 wt% or 1.96 wt%. Furthermore, these glass fibre types were used to manufacture composite specimens and characterised by static tension tests in fibre and transverse directions. </span></p> <p><span>This dataset is a collection of 13 Excel files.</span></p> <p><span>The “Glass fibre properties and Weibull analysis” Excel file summarise the single fibre tensile testing of glass fibre types and strength analysis using unimodal 2-parameter Weibull theory. There are four sheets in the Excel files for 0 wt%, 1.64 wt%, 1.90 wt% or 1.96 wt% glass fibres. </span></p> <p><span>The Excel files “Single glass fibres-Stress-strain curves-0 %, 1.64 %, 1.90 %, 1.96 %” contains the raw data of individual fibres obtained from the single fibre tensile testing experiments.</span></p> <p><span>There are eight Excel files containing the raw data of static tensile tests in the fibre and transverse direction of the composites made with glass fibres were manufactured with varying fractions of recycled fibre powder, 0 wt%, 1.64 wt%, 1.90 wt% or 1.96 wt%.</span></p>
CLDF dataset derived from Carroll et al. "Yamfinder: The Southern New Guinea Lexical Database"
<p>Cite the source of the dataset as:</p> <blockquote> <p>Carroll, Matthew J., Barth, Wolfgang, Nicholas Evans, I Wayan Arka, Christian Döhler, Eri Kashima, Volker Gast, Tina Gregor, Kate L. Lindsey, Julia Miller, Emil Mittag, Bruno Olsson, Dineke Schokkin, Jeff Siegel, Charlotte van Tongeren, Kyla Quinn. [DATE ACCESSED]. Yamfinder: Southern New Guinea Lexical Database. Available online at: http://www.yamfinder.com</p> </blockquote>
Towards new demography proxies and regional chronologies: Radiocarbon dates from archaeological contexts located in the Czech Republic covering the period between 10,000 BC and AD 1250 (dataset)
<p>The dataset was created within the project “<em>Land use, social transformations and woodland in Central European Prehistory. Modelling approaches to human-environment interactions</em>” funded by the Czech Science Foundation (19-20970Y). This dataset represents the largest and the most comprehensive collection of archaeological radiocarbon dates from the Czech Republic to date. The dataset offers 1579 samples from 347 archaeological sites dating from Early Mesolithic (10 000 BC) to Medieval Period (AD 1250). Published in a simple spreadsheet format, the database offers researchers a quick tool for further analyses. It is important to highlight that dates we collected originated only from archaeological contexts, which means that we have excluded some radiocarbon dates produced through palaeoecological research without a direct relationship to past human activities, such as pollen records or samples from fossilized trees in river beds. The dataset is intended to be used for demographic modelling of population numbers during periods without written records, i.e. prehistory.</p>
Dataset for "IRIS analyser assessment reveals sub-hourly variability of isotope ratios in carbon dioxide at Baring Head, New Zealand's atmospheric observatory in the Southern Ocean"
<p>Dataset for</p> <p>Sperlich, P., Brailsford, G. W., Moss, R. C., McGregor, J., Martin, R. J., Nichol, S., Mikaloff-Fletcher, S., Bukosa, B., Mandic, M., Schipper, I., Krummel, P. and Griffiths, A. D.: IRIS analyser assessment reveals sub-hourly variability of isotope ratios in carbon dioxide at Baring Head, New Zealand's atmospheric observatory in the Southern Ocean, Atmos. Meas. Tech., https://doi.org/10.5194/amt-15-1-2022, 2022.</p>
Habitat Protection Indexes - new monitoring measures for the conservation of threatened marine habitats - Datasets and supporting files
<p>The supporting datasets, scripts, and supplementary information for the manuscript, "Habitat Protection Indexes - new monitoring measures for the conservation of threatened marine habitats," are available within this repository.</p> <p>We conduct an analysis on the coverage of protected areas that cover six threatened marine and coastal and developed two indexes, the Local Proportion of Habitat Protected Index and the Global Proportion of Habitat Protected Index, describing the protection of these habitats locally and globally. The habitats considered are the following: cold corals, warm water corals, knolls and seamounts, mangroves, saltmarshes, and seagrasses.</p> <p>The index scores of each jurisdiction are made available for download in the dataset: <em>habitat_protection_indexes_average.csv</em></p> <p>The habitat specific index scores for each jurisdiction are made available for download in the dataset: <em>habitat_protection_indexes.csv. </em></p> <p>Column name descriptions are available in the text file: <em>Column_name_descriptions_20220301</em></p> <p>The scripts used to run the workflow to calculate the indexes, create figures, and calculate statistics for the manuscript are also included. The script <em>01_Workflow sources</em> the first 9 scripts in the <em>scripts</em> folder to calculate the indexes which relies on the functions script within the functions folder. The rest of the scripts in the folder create the figures and calculate the statistics for the manuscript.</p> <p>A readme pdf file is included here to ease with reproducing the workflow, but we strongly suggest to please visit our github (<a href="https://github.com/jkumagai96/Marine_Habitat_protection">https://github.com/jkumagai96/Marine_Habitat_protection</a>) to reproduce the entire calculation where we provide detailed information on how to run the workflow and package management.</p>
Where2Test Saxony-Czechia COVID-19 new cases dataset
<p>Data in the repository were used in the study "Fine-scale variation in the effect of national border on COVID-19 spread: A case study of the Saxon-Czech border region", published in <a href="https://www.sciencedirect.com/journal/spatial-and-spatio-temporal-epidemiology">Spatial and Spatio-temporal Epidemiology</a>.</p> <p>This repository consists of two files:</p> <p><strong>saxony-westczechia_cases7</strong></p> <p>Weekly numbers of new COVID-19 cases in all municipalities in Saxony and Northwestern Czechia (Liberec, Ústí nad Labem, and Karlovy Vary regions) in the first half of 2021. Data are extracted from the websites <a href="https://www.coronavirus.sachsen.de">coronavirus.sachsen</a> and <a href="https://onemocneni-aktualne.mzcr.cz/covid-19">onemocneni-aktualne.mzcr.cz/covid-19</a>. The missing values were interpolated, and daily values were recalculated to weekly values.</p> <p><strong>municipalities</strong></p> <p>The second file consists of a list of all municipalities with their names, geometries, and population values. For Germany, we used the dataset <a href="https://hub.arcgis.com/datasets/esri-de-content::gemeindegrenzen-2018-mit-einwohnerzahl/about">"Gemeindegrenzen 2018 mit Einwohnerzahl"</a> (© GeoBasis-DE / BKG, Statistisches Bundesamt (Destatis) (2020), <a href="http://www.govdata.de/dl-de/by-2-0">dl-de/by-2-0</a>) as a source of geometries and population sizes of the municipalities (“<em>Gemeinde”</em>) in Saxony. Czech population numbers on the municipality level ("obec") were taken from the <a href="https://www.czso.cz/csu/czso/population-of-municipalities-1-january-2021">Czech Statistical Office</a>, while the geometries were obtained from <a href="https://www.cuzk.cz/ruian/RUIAN.aspx">RÚIAN</a> (@<a href="http://geoportal.cuzk.cz">Czech Office for Surveying, Mapping and Cadastre</a>, 2021). To keep the same geometry detail on both sides of the borders, we applied the Douglas-Peucker simplification algorithm implemented in the Python library <a href="https://github.com/mattijn/topojson">TopoJSON</a>.</p>
Dataset: A New Infrared Criterion for Selecting Active Galactic Nuclei to Lower Luminosities
<p>The dataset accompanying the paper "A New Infrared Criterion for Selecting Active Galactic Nuclei to Lower Luminosities" (Hviding et al. in prep)</p>
New dataset obtained from 2D positioning system with distributed control
<p>The dataset contains the obtained trajectory (encoder measurements) of the 2D positioning system, as well as the reference (commanded) trajectory. The control task is distributed to the low-level controllers for x and y axes synchronized using IEEE 1588 Precision Time Protocol (PTP), where the movement of the axes is realized based on the data that the low-level controllers receive from the high-level controller. The dataset includes 60 signals (30 measurements for x and y axis) that represent obtained trajectories and 2 signals (x and y axes) that represent commanded trajectory, where the length of each signal is 61,000 samples. The system was designed and built by the Cyber-Physical Systems Lab at the Pratt School of Engineering, Duke University, where it is located. Table 1 shows the list of collected signals, whereas a detailed description of the system can be found in [1]. For more information, see [2, 3].</p>
New hourly HCHO total columns dataset
<p>The new hourly HCHO dataset is generated by ground-based high-resolution FTIR, GEOS-Chem model and machine learning approach over Hefei, China.</p>
amel-github/sars-ani: Releasing new data fields in SARS-ANI dataset
<p>2022-06-20 - Release v1.1</p> <p>The original SARS-ANI dataset displayed common and scientific names of the animal host as found in the information source and/or inferred from the literature or expert knowledge.<br> Misspelled animal names and errors in taxonomy can lead to incorrect scientific conclusions and poor policy design. Moreover, harmonized host names can aid integrating other datasets (e.g. data on host biological traits, geographic distribution, or association with other pathogens).<br> Therefore, for each event, we programmatically performed taxonomic validation of the animal host name, using the R package taxize (Chamberlain et al. 2013). For more information on our validation process, see the R script <strong>sars_ani_validation.R.</strong></p> <p>Version 1.1. contains seven fields related to the identification of the animal host:</p> <ul> <li> <p>host_com_orig: Most specific designation of the animal host provided by the source(s), in English.</p> </li> <li> <p>host_sci_orig: Scientific name of the animal host as mentioned in the source(s) (scientific names are harmonized so that only the first letter of the genus is capitalized).</p> </li> <li> <p>host_com_res: Common name of the animal host, harmonized against the National Center for Biotechnology Information (<a href="https://www.ncbi.nlm.nih.gov/">NCBI</a>) taxonomic backbone.</p> </li> <li> <p>host_sci_res: Scientific name of the animal host (resolved to species or subspecies level), harmonized against the National Center for Biotechnology Information (<a href="https://www.ncbi.nlm.nih.gov/">NCBI</a>) taxonomic backbone.</p> </li> <li> <p>host_colloq: The colloquial name of the host, i.e. the name commonly used to identify the animal in non-specialist language (e.g. "tiger" for "Sumatran tiger").</p> </li> <li> <p>host_sci_spec_res: The scientific name of the host resolved to the species level.</p> </li> <li> <p>family: Animal family of the animal host.</p> </li> </ul>
Dataset of multi-objective optimization results for a new latent energy storage approach in buildings based on several phase change materials with different melting temperatures
<p>This dataset comprises the multi-objective optimization results obtained for a new latent energy storage approach based on several phase change materials (PCMs) with different melting temperatures in buildings. The results were obtained for a small office building in eight climate-representative locations according to the ASHRAE 169-2020 climate classification and within the WMO Region VI (Europe).</p> <p>The dataset contains:</p> <p>- The EnergyPlus baseline models employed as a case study for each climate.</p> <p>- The Pareto fronts obtained after the multi-objective optimization in each climate.</p> <p>- The EnergyPlus models for the best designs achieved on the Pareto fronts in terms of annual total load reductions.</p>
Large-Eddy Simulation of Wind Turbine Flows: A New Evaluation of Actuator Disk Models - Dataset
<p>Main data used in the following paper: Revaz, T.; Porté-Agel, F. Large-Eddy Simulation of Wind Turbine Flows: A New Evaluation of Actuator Disk Models. <em>Energies</em> <strong>2021</strong>, <em>14</em>, 3745. https://doi.org/10.3390/en14133745</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.