Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
9,300
datasets available to search
ShareScore release 0.7.1
Dataset results
9,300 results for “detection”
DCASE 2024 Task 5: Few-shot Bioacoustic Event Detection Development Set
<p><strong>General Description:</strong></p> <p>The development set for task 5 of DCASE 2024 "Few-shot Bioacoustic Event Detection" consists of 217 audio files acquired from different bioacoustic sources. The dataset is split into training and validation sets. </p> <p>Multi-class annotations are provided for the training set with positive (POS), negative (NEG) and unkwown (UNK) values for each class. UNK indicates uncertainty about a class. </p> <p>Single-class (class of interest) annotations are provided for the validation set, with events marked as positive (POS) or unkwown (UNK) provided for the class of interest. </p> <p><strong>Folder Structure:</strong></p> <p><em>Development_set.zip</em></p> <p>|_Development_Set/</p> <p> |__Training_Set/</p> <p> |___JD/</p> <p> |____*.wav</p> <p> |____*.csv</p> <p> |___HT/</p> <p> |____*.wav</p> <p> |____*.csv</p> <p> |___BV/</p> <p> |____*.wav</p> <p> |____*.csv</p> <p> |___MT/</p> <p> |____*.wav</p> <p> |____*.csv</p> <p> |___WMW/</p> <p> |____*.wav</p> <p> |____*.csv</p> <p> </p> <p> |__Validation_Set/</p> <p> |___HB/</p> <p> |____*.wav</p> <p> |____*.csv</p> <p> |___PB/</p> <p> |____*.wav</p> <p> |____*.csv</p> <p> |___ME/</p> <p> |____*.wav</p> <p> |____*.csv</p> <p> |___PB24/</p> <p> |____*.wav</p> <p> |____*.csv</p> <p> |___RD/</p> <p> |____*.wav</p> <p> |____*.csv</p> <p> |___PW/</p> <p> |____*.wav</p> <p> |____*.csv</p> <p> </p> <p><em>Development_set_annotations.zip</em> has the same structure but contains only the *.csv files</p> <p> </p> <p><strong>Dataset statistics</strong></p> <p>Some statistics on this dataset are as follows, split between training and validation set and their sub-folders:</p> <p>-----------------------------------------------------<br>TRAINING SET<br>-----------------------------------------------------<br>Number of audio recordings | 174<br>Total duration | 21 hours<br>Total classes | 47<br>Total events | 14229<br>-----------------------------------------------------<br>TRAINING SET/BV<br>-----------------------------------------------------<br>Number of audio recordings | 5<br>Total duration | 10 hours<br>Total classes | 11<br>Total events | 9026<br>Sampling rate | 24000 Hz<br>-----------------------------------------------------<br>TRAINING SET/HT<br>-----------------------------------------------------<br>Number of audio recordings | 5<br>Total duration | 5 hours<br>Total classes | 5<br>Total events | 611<br>Sampling rate | 6000 Hz<br>-----------------------------------------------------<br>TRAINING SET/JD<br>-----------------------------------------------------<br>Number of audio recordings | 1<br>Total duration | 10 mins<br>Total classes | 1<br>Total events | 357<br>Sampling rate | 22050 Hz<br>-----------------------------------------------------<br>TRAINING SET/MT<br>-----------------------------------------------------<br>Number of audio recordings | 2<br>Total duration | 1 hour and 10 mins<br>Total classes | 4<br>Total events | 1294<br>Sampling rate | 8000 Hz<br>-----------------------------------------------------<br>TRAINING SET/WMW<br>-----------------------------------------------------<br>Number of audio recordings | 161<br>Total duration | 4 hours and 40 mins<br>Total classes | 26<br>Total events | 2941<br>Sampling rate | various sampling rates<br>-----------------------------------------------------</p> <p>-----------------------------------------------------<br>VALIDATION SET<br>-----------------------------------------------------<br>Number of audio recordings | 43<br>Total duration | 49 hours and 57 minutes<br>Total classes | 7<br>Total events | 3504<br>-----------------------------------------------------<br>VALIDATION SET/HB<br>-----------------------------------------------------<br>Number of audio recordings | 10<br>Total duration | 2 hours and 38 minutes<br>Total classes | 1<br>Total events | 712<br>Sampling rate | 44100 Hz<br>-----------------------------------------------------<br>VALIDATION SET/PB<br>-----------------------------------------------------<br>Number of audio recordings | 6<br>Total duration | 3 hours<br>Total classes | 2<br>Total events | 292<br>Sampling rate | 44100 Hz<br>-----------------------------------------------------<br>VALIDATION SET/ME<br>-----------------------------------------------------<br>Number of audio recordings | 2<br>Total duration | 20 minutes<br>Total classes | 2<br>Total events | 73<br>Sampling rate | 44100 Hz<br>-----------------------------------------------------<br>VALIDATION SET/PB24<br>-----------------------------------------------------<br>Number of audio recordings | 4<br>Total duration | 2 hours<br>Total classes | 2<br>Total events | 350<br>Sampling rate | 44100 Hz<br>-----------------------------------------------------<br>VALIDATION SET/RD<br>-----------------------------------------------------<br>Number of audio recordings | 6<br>Total duration | 18 hours<br>Total classes | 1<br>Total events | 1372<br>Sampling rate | 48000 Hz<br>-----------------------------------------------------<br>VALIDATION SET/PW<br>-----------------------------------------------------<br>Number of audio recordings | 15<br>Total duration | 24 hours<br>Total classes | 1<br>Total events | 705<br>Sampling rate | 96000 Hz<br>-----------------------------------------------------</p> <p><strong>Annotation structure</strong></p> <p>Each line of the annotation csv represents an event in the audio file. The column descriptions are as follows:</p> <p>TRAINING SET<br>---------------------<br>Audiofilename, Starttime, Endtime, CLASS_1, CLASS_2, ...CLASS_N</p> <p>VALIDATION SET<br>---------------------<br>Audiofilename, Starttime, Endtime, Q</p> <p> </p> <p><strong>Classes</strong></p> <p>DCASE2024_task5_training_set_classes.csv and DCASE2024_task5_validation_set_classes.csv provide a table with class code correspondence to class name for all classes in the Development set. Additionally, DCASE2024_task5_validation_set_classes.csv also provides a recording names column.</p> <p>DCASE2024_task5_training_set_classes.csv<br>---------------------<br>dataset, class_code, class_name</p> <p>DCASE2024_task5_validation_set_classes.csv<br>---------------------<br>dataset, recording, class_code, class_name</p> <p> </p> <p><strong>Evaluation Set</strong></p> <p>The Evaluation set for this task will be released on the 1 June 2024</p> <p><strong>Open Access:</strong></p> <p>This dataset is available under a Creative Commons Attribution 4.0 International (CC BY 4.0) license.<br> </p> <p><strong>Contact info:</strong></p> <p>Please send any feedback or questions to:</p> <p>Burooj Ghani - burooj.ghani@naturalis.nl | Ines Nolasco - i.dealmeidanolasco@qmul.ac.uk</p> <p>Alternately, join us on Slack: <a href="https://join.slack.com/t/dcase/shared_invite/zt-12zfa5kw0-dD41gVaPU3EZTCAw1mHTCA">task-fewshot-bio-sed</a></p> <p> </p>
Data from: The effect of probe density coverage on the detection of oenological tannins in quartz crystal microbalance with dissipation monitoring (QCM-D) experiments
<p>Polyphenols, crucial compounds in grapes, musts, and wines, influence grape ripening, must fermentation, and final wine quality. Current detection methods for polyphenols are expensive, time-consuming, and reliant on specialized laboratories and personnel. This study proposes the use of a functionalized acoustic sensor to address these limitations and efficiently detect oenological polyphenols.</p> <p>The method employs a quartz crystal microbalance with dissipation monitoring (QCM-D) combined with a gelatin-based probe layer to detect the target analyte. The sensor is functionalized by optimizing probe coverage density, accomplished through the use of 12-mercaptododecanoic acid (12-MCA) for probe immobilization onto the gold sensor surface, along with dithiothreitol (DTT) as a reducing and competitive binding agent. Varying concentrations of 12-MCA and DTT allow for control over probe density, with QCM-D measurements demonstrating effective adjustment, ranging from 0.2 × 10^13 to 2 × 10^13 molecules cm^−2. The study also explores the interaction between the probe and tannins, confirming the ability of the sensor to detect them. Notably, lower probe coverage yields higher detection signals when normalized to probe immobilization signals. Additionally, significant alterations in the mechanical properties of the functionalization layer occur after interaction with samples.</p> <p>Combining QCM-D with gelatin functionalization presents promising applications in the wine industry. This approach enables real-time monitoring, requires minimal sample preparation, and offers high sensitivity for quality control purposes.</p>
Change detection technique comparison in long-term wetland monitoring: datasets and maps of the Poitevin Marsh (France)
<h3>For a full description of the methodology and results, please see the following article:</h3> <div> <div>Demarquet, Q., Rapinel, S., Gore, O., Dufour, S., Hubert-Moy, L., 2024. Continuous change detection outperforms traditional post-classification change detection for long term monitoring of wetlands. <em>International Journal of Applied Earth Observation and Geoinformation </em>133, 104142. <a href="https://doi.org/10.1016/j.jag.2024.104142">https://doi.org/10.1016/j.jag.2024.104142</a></div> <div> </div> <div>--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------</div> </div> <h3># Datasets</h3> <p>Points datasets are projected in WGS84 (EPSG:4326), and are provided in the open source GeoPackage format.</p> <p>The first dataset (<strong>Dataset_1.gpkg</strong>) contains training and validation points for random forest classification of EUNIS habitats in the Poitevin Marsh. This dataset consists of 3360 training and 840 validation points (total: 4200).<br>Fields description:</p> <ul> <li>"<em>ID</em>": unique identifier</li> <li>"<em>CLASS</em>": EUNIS first level habitat type, classified as following:<br> <ul> <li>1: EUNIS habitat A</li> <li>2: EUNIS habitat B</li> <li>3: EUNIS habitat C1J5</li> <li>4: EUNIS habitat C3</li> <li>5: EUNIS habitat E</li> <li>6: EUNIS habitat G</li> <li>7: EUNIS habitat I</li> <li>8: EUNIS habitat J</li> </ul> </li> <li>"<em>DATE</em>": Date associated with EUNIS habitat sample</li> <li>"<em>LON</em>": Point longitude in decimal degrees</li> <li>"<em>LAT</em>": Point latitude in decimal degrees</li> <li>"<em>TYPE</em>": Either training ("<em>train</em>") or validation ("<em>test</em>") sample</li> </ul> <p>The second dataset (<strong>Dataset_2.gpkg</strong>) contains points for the Olofsson correction method. This dataset consists of 326 points where the change classes are classified as following: -10 (wetland loss), 10 (wetland gain), 100 (stable existing wetland), and 200 (stable damaged wetland).<br>Fields description:</p> <ul> <li>"<em>ID</em>": unique identifier</li> <li>"<em>LON</em>": Point longitude in decimal degrees</li> <li>"<em>LAT</em>": Point latitude in decimal degrees</li> <li>"<em>REFERENCE</em>": Change class reference</li> <li>"<em>CCDC</em>": Change class obtained from the Continuous Change Detection and Classification approach</li> <li>"<em>PCCD</em>": Change class obtained from the Post-Classification Change Detection approach</li> </ul> <p>Supplementary layout files (<strong>Dataset_1.qml</strong> and <strong>Dataset_2.qml</strong>) support formatting of the points in QGIS software.</p> <p>--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------</p> <h3># EUNIS habitat</h3> <p>Maps are projected in WGS84 (EPSG:4326), and are provided in the GeoTiff format at 30m of spatial resolution. </p> <p>Habitat maps are given for the two approaches in years 1984 and 2022:</p> <ul> <li>CCDC: Continuous Change Detection and Classification (<strong>CCDC_HABITAT_1984.tif</strong> and <strong>CCDC_HABITAT_2022.tif</strong>)</li> <li>PCCD: Traditional post-classification approach (<strong>PCCD_HABITAT_1984.tif </strong>and <strong>PCCD_HABITAT_2022.tif</strong>)</li> </ul> <p>Supplementary layout files (<strong>CCDC_HABITAT_1984.qml, CCDC_HABITAT_2022.qml, PCCD_HABITAT_1984.qml, PCCD_HABITAT_2022.qml</strong>) support formatting of raster layers in QGIS software.</p> <p>--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------</p> <h3># Change detection during the 1984-2022 period</h3> <p>Maps are projected in WGS84 (EPSG:4326), and are provided in the GeoTiff format at 30m of spatial resolution. Raster values follow the classification scheme used in Dataset_2.</p> <p>Change detection maps are given for the two approaches:</p> <ul> <li>CCDC: Continuous Change Detection and Classification (<strong>CCDC_CHANGE_1984_2022.tif</strong>)</li> <li>PCCD: Traditional post-classification approach (<strong>PCCD_CHANGE_1984_2022.tif</strong>)</li> </ul> <p>Supplementary layer files (<strong>CCDC_CHANGE_1984_2022.qml</strong> and<strong> PCCD_CHANGE_1984_2022.qml</strong>) support formatting of the raster layers in QGIS software.</p> <p>--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------</p> <h3># GEE repository</h3> <p>To get direct access to GEE scripts and assets, please follow those two links:</p> <p>https://code.earthengine.google.com/?accept_repo=users/demarquetquentin/CCDC_Poitevin</p> <p>https://code.earthengine.google.com/?asset=projects/ee-quen-dem/assets/CCDC_Poitevin</p>
Dataset for publication: An inter-laboratory study characterizes the impact of bioinformatic approaches on genome-based cluster detection for foodborne bacterial pathogens
<p>This dataset is part of a dry-lab interlaboratory study conducted across Germany, regarding bacterial outbreak detection based on NGS data, with a focus on bioinformatic analysis of four species to identify potential variability caused by different data analysis approaches and human interpretation. Participants were asked to follow their usual in-house protocols while adhering to the general guidelines. A quality assessment (with sample exclusion) was followed by 7-gene Multilocus-Sequence Typing (MLST), core genome Multilocus Sequencing Typing (cgMLST), and SNP calling. The participants were then asked to identify clusters. The study was not intended to resemble a standard proficiency test with a passing/failing grade, but rather to investigate and quantify obvious variability in the results and, where possible, the reasons for it. For this purpose, the datasets included borderline cases in terms of quality.</p>
Data for publication 'Detection of Artificial Seed-like Objects from UAV Imagery'
<p>This resource contains the datasets supporting the model development as published in the article 'Detection of Artificial Seed-like Objects from UAV Imagery' (https://doi.org/10.3390/rs15061637).</p> <p>In the last two decades, unmanned aerial vehicle (UAV) technology has been widely utilized as an aerial survey method. Recently, a unique system of self-deployable and biodegradable microrobots akin to winged achene seeds was introduced to monitor environmental parameters in the air above the soil interface, which requires geo-localization. This research focuses on detecting these artificial seed-like objects from UAV RGB images in real-time scenarios, employing the object detection algorithm YOLO (You Only Look Once). Three environmental parameters, namely, daylight condition, background type, and flying altitude, were investigated to encompass varying data acquisition situations and their influence on detection accuracy. Artificial seeds were detected using four variants of the YOLO version 5 (YOLOv5) algorithm, which were compared in terms of accuracy and speed. The most accurate model variant was used in combination with slice-aided hyper inference (SAHI) on full resolution images to evaluate the model’s performance. It was found that the YOLOv5n variant had the highest accuracy and fastest inference speed. After model training, the best conditions for detecting artificial seed-like objects were found at a flight altitude of 4 m, on an overcast day, and against a concrete background, obtaining accuracies of 0.91, 0.90, and 0.99, respectively. YOLOv5n outperformed the other models by achieving a mAP0.5 score of 84.6% on the validation set and 83.2% on the test set. This study can be used as a baseline for detecting seed-like objects under the tested conditions in future studies.</p>
AWOFRO : Annotated tweet corpus of mixed Wolof-French for detecting obnoxious messages
<p>These data are tweets of mixed Wolof-French codes annotated by three(3) annotators. <br>They were extracted during the period from 1 January 2021 to 31 May 2023.</p> <p>Content description :</p> <p><strong>Corpora.rar</strong> : The dataset contains 3510 annotated tweets</p>
Ultrasensitive ctDNA detection for preoperative disease stratification in early-stage lung adenocarcinoma
<p>Code and data for the MS <strong>"Ultrasensitive ctDNA detection for preoperative disease stratification in early-stage lung adenocarcinoma"</strong></p>
A subset of the EMARS dataset in MY24 and MY26 converted from the sigma-p hybrid coordinate to the pressure coordinate and a list of local dust storms detected during the MYs in western Arcadia Planitia
<p>This dataset includes a subset of EMARS' background mean data (Greybush et al., 2019) converted from the sigma-p hybrid coordinate to the pressure coordinate. Only MY24 and MY26 were used to generate the figures shown in Ogohara (submitted to JGR Planets). <br>Updates from the original EMARS are:</p> <ul> <li>The vertical coordinate has been converted from the sigma-p hybrid coordinate to the pressure coordinate.</li> <li>The variables expressing the Earth date (e.g., year, month, day, etc.) have been combined into one variable, earth_date.</li> <li>A new variable, emars_date, has been created from emars_sol and mars_hour.</li> </ul> <p>In addition, this dataset provides two lists of local dust storms events during MY24 and MY26 which were detected in western Arcadia Planitia using a deep learning-based method proposed by Ogohara and Gichu (2022). The lists are:</p> <ul> <li>[Data Set S1] List of global image swath files examined. Only file names of MGS/MOC red band images are listed. The list consists of 5 columns indicating image ID, observation date, orbit number, solar longitude, and filter name (RED).</li> <li>[Data Set S2] List of global image swath files containing identified dust storms, as well as some attributes of the detected dust storms. Only file names of red band images are listed. The list consists of 7 columns indicating image ID, observation date, orbit number, solar longitude, center longitude and latitude, and area (km2.)</li> </ul>
Dataset for marine vessel detection from Sentinel 2 images in the Finnish coast
<p>This dataset contains annotated marine vessels from 15 different Sentinel-2 product, used for training object detection models for marine vessel detection. The vessels are annotated as bounding boxes, covering also some amount of the wake, if present.</p> <h2>Source data</h2> <div> <div>Individual products used to generate annotations are shown in the following table:</div> </div> <div> </div> <div> <table style="width: 58.034%; height: 411.47px;"> <tbody> <tr style="height: 19.5938px;"> <td style="width: 16.0296%; height: 19.5938px;"><strong>Location</strong></td> <td style="width: 79.3617%; height: 19.5938px;"><strong>Product name</strong></td> </tr> <tr style="height: 39.1875px;"> <td style="width: 16.0296%; height: 39.1875px;">Archipelago sea</td> <td style="width: 79.3617%; height: 39.1875px;">S2A_MSIL1C_20220515T100031_N0510_R122_T34VEM_20240617T162344.SAFE</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 16.0296%; height: 19.5938px;"> </td> <td style="width: 79.3617%; height: 19.5938px;">S2B_MSIL1C_20220619T100029_N0510_R122_T34VEM_20240627T204751.SAFE</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 16.0296%; height: 19.5938px;"> </td> <td style="width: 79.3617%; height: 19.5938px;">S2A_MSIL1C_20220721T095041_N0510_R079_T34VEM_20240712T224506.SAFE</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 16.0296%; height: 19.5938px;"> </td> <td style="width: 79.3617%; height: 19.5938px;">S2A_MSIL1C_20220813T095601_N0510_R122_T34VEM_20240717T115958.SAFE</td> </tr> <tr style="height: 39.1875px;"> <td style="width: 16.0296%; height: 39.1875px;">Gulf of Finland</td> <td style="width: 79.3617%; height: 39.1875px;">S2B_MSIL1C_20220606T095029_N0510_R079_T35VLG_20240619T111429.SAFE</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 16.0296%; height: 19.5938px;"> </td> <td style="width: 79.3617%; height: 19.5938px;">S2B_MSIL1C_20220626T095039_N0510_R079_T35VLG_20240620T013500.SAFE</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 16.0296%; height: 19.5938px;"> </td> <td style="width: 79.3617%; height: 19.5938px;">S2B_MSIL1C_20220703T094039_N0510_R036_T35VLG_20240702T075354.SAFE</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 16.0296%; height: 19.5938px;"> </td> <td style="width: 79.3617%; height: 19.5938px;">S2A_MSIL1C_20220721T095041_N0510_R079_T35VLG_20240712T224506.SAFE</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 16.0296%; height: 19.5938px;">Bothnian Bay</td> <td style="width: 79.3617%; height: 19.5938px;">S2A_MSIL1C_20220627T100611_N0510_R022_T34WFT_20240628T041908.SAFE</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 16.0296%; height: 19.5938px;"> </td> <td style="width: 79.3617%; height: 19.5938px;">S2B_MSIL1C_20220712T100559_N0510_R022_T34WFT_20240718T033027.SAFE</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 16.0296%; height: 19.5938px;"> </td> <td style="width: 79.3617%; height: 19.5938px;">S2B_MSIL1C_20220828T095549_N0510_R122_T34WFT_20240708T035231.SAFE</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 16.0296%; height: 19.5938px;">Bothnian Sea</td> <td style="width: 79.3617%; height: 19.5938px;">S2B_MSIL1C_20210714T100029_N0500_R122_T34VEN_20230224T120043.SAFE</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 16.0296%; height: 19.5938px;"> </td> <td style="width: 79.3617%; height: 19.5938px;">S2B_MSIL1C_20220619T100029_N0510_R122_T34VEN_20240627T204751.SAFE</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 16.0296%; height: 19.5938px;"> </td> <td style="width: 79.3617%; height: 19.5938px;">S2A_MSIL1C_20220624T100041_N0510_R122_T34VEN_20240714T110124.SAFE</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 16.0296%; height: 19.5938px;"> </td> <td style="width: 79.3617%; height: 19.5938px;">S2A_MSIL1C_20220813T095601_N0510_R122_T34VEN_20240717T115958.SAFE</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 16.0296%; height: 19.5938px;">Kvarken</td> <td style="width: 79.3617%; height: 19.5938px;">S2A_MSIL1C_20220617T100611_N0510_R022_T34VER_20240627T094433.SAFE</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 16.0296%; height: 19.5938px;"> </td> <td style="width: 79.3617%; height: 19.5938px;">S2B_MSIL1C_20220712T100559_N0510_R022_T34VER_20240718T033027.SAFE</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 16.0296%; height: 19.5938px;"> </td> <td style="width: 79.3617%; height: 19.5938px;">S2A_MSIL1C_20220826T100611_N0510_R022_T34VER_20240705T062429.SAFE</td> </tr> </tbody> </table> </div> <div> <div> </div> <div>Even though the reference data IDs are for L1C products, L2A products from the same acquisition dates can be used along with the annotations. However, Sen2Cor has been known to produce incorrect reflectance values for water bodies.</div> <div> </div> <div>The corresponding L2A product identifiers are:</div> </div> <div> </div> <div> <table style="width: 58.034%; height: 411.47px;"> <tbody> <tr style="height: 19.5938px;"> <td style="width: 16.0296%; height: 19.5938px;"><strong>Location</strong></td> <td style="width: 79.3617%; height: 19.5938px;"><strong>Product name</strong></td> </tr> <tr style="height: 39.1875px;"> <td style="width: 16.0296%; height: 39.1875px;">Archipelago sea</td> <td style="width: 79.3617%; height: 39.1875px;">S2A_MSIL2A_20220515T100031_N0400_R122_T34VEM_20220515T141508.SAFE</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 16.0296%; height: 19.5938px;"> </td> <td style="width: 79.3617%; height: 19.5938px;">S2B_MSIL2A_20220619T100029_N0510_R122_T34VEM_20240628T011619.SAFE</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 16.0296%; height: 19.5938px;"> </td> <td style="width: 79.3617%; height: 19.5938px;">S2A_MSIL2A_20220721T095041_N0510_R079_T34VEM_20240713T035445.SAFE</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 16.0296%; height: 19.5938px;"> </td> <td style="width: 79.3617%; height: 19.5938px;">S2A_MSIL2A_20220813T095601_N0510_R122_T34VEM_20240717T165127.SAFE</td> </tr> <tr style="height: 39.1875px;"> <td style="width: 16.0296%; height: 39.1875px;">Gulf of Finland</td> <td style="width: 79.3617%; height: 39.1875px;">S2B_MSIL2A_20220606T095029_N0510_R079_T35VLG_20240619T162121.SAFE</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 16.0296%; height: 19.5938px;"> </td> <td style="width: 79.3617%; height: 19.5938px;">S2B_MSIL2A_20220626T095039_N0510_R079_T35VLG_20240620T063951.SAFE</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 16.0296%; height: 19.5938px;"> </td> <td style="width: 79.3617%; height: 19.5938px;">S2B_MSIL2A_20220703T094039_N0510_R036_T35VLG_20240702T130032.SAFE</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 16.0296%; height: 19.5938px;"> </td> <td style="width: 79.3617%; height: 19.5938px;">S2A_MSIL2A_20220721T095041_N0510_R079_T35VLG_20240713T035445.SAFE</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 16.0296%; height: 19.5938px;">Bothnian Bay</td> <td style="width: 79.3617%; height: 19.5938px;">S2A_MSIL2A_20220627T100611_N0510_R022_T34WFT_20240628T095704.SAFE</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 16.0296%; height: 19.5938px;"> </td> <td style="width: 79.3617%; height: 19.5938px;">S2B_MSIL2A_20220712T100559_N0510_R022_T34WFT_20240718T063657.SAFE</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 16.0296%; height: 19.5938px;"> </td> <td style="width: 79.3617%; height: 19.5938px;">S2B_MSIL2A_20220828T095549_N0510_R122_T34WFT_20240708T091048.SAFE</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 16.0296%; height: 19.5938px;">Bothnian Sea</td> <td style="width: 79.3617%; height: 19.5938px;">S2B_MSIL2A_20210714T100029_N0500_R122_T34VEN_20230224T182455.SAFE</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 16.0296%; height: 19.5938px;"> </td> <td style="width: 79.3617%; height: 19.5938px;">S2B_MSIL2A_20220619T100029_N0510_R122_T34VEN_20240628T011619.SAFE</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 16.0296%; height: 19.5938px;"> </td> <td style="width: 79.3617%; height: 19.5938px;">S2A_MSIL2A_20220624T100041_N0510_R122_T34VEN_20240714T162313.SAFE</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 16.0296%; height: 19.5938px;"> </td> <td style="width: 79.3617%; height: 19.5938px;">S2A_MSIL2A_20220813T095601_N0510_R122_T34VEN_20240717T165127.SAFE</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 16.0296%; height: 19.5938px;">Kvarken</td> <td style="width: 79.3617%; height: 19.5938px;">S2A_MSIL2A_20220617T100611_N0510_R022_T34VER_20240627T130404.SAFE</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 16.0296%; height: 19.5938px;"> </td> <td style="width: 79.3617%; height: 19.5938px;">S2B_MSIL2A_20220712T100559_N0510_R022_T34VER_20240718T063657.SAFE</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 16.0296%; height: 19.5938px;"> </td> <td style="width: 79.3617%; height: 19.5938px;">S2A_MSIL2A_20220826T100611_N0510_R022_T34VER_20240705T120522.SAFE</td> </tr> </tbody> </table> </div> <div><br> <div>The raw products can be acquired from <a href="https://dataspace.copernicus.eu" target="_blank" rel="noopener">Copernicus Data Space Ecosystem.</a> The products listed above can be unavailable due to e.g. processing level updates and old versions being deleted. In those cases, try searching with the tile identifier and acquisition date in order to get the correct product ID.</div> <br> <h2>Annotations</h2> <br> <div>The annotations are bounding boxes drawn around marine vessels so that some amount of their wakes, if present, are also contained within the boxes. The data are distributed as geopackage files, so that one geopackage corresponds to a single Sentinel-2 tile, and each package has separate layers for individual products as shown below:</div> <br> <blockquote> <div>T34VEM</div> <div>|-20220515</div> <div>|-20220619</div> <div>|-20220721</div> <div>|-20220813</div> </blockquote> <br> <div>All layers have a column <strong>id</strong>, which has the value <strong>b</strong><strong>oat</strong> for all annotations.</div> <br> <div>CRS is EPSG:32634 for all products except for the Gulf of Finland (35VLG), which is in EPSG:32635. This is done in order to have the bounding boxes to be aligned with the pixels in the imagery.</div> <br> <div>As tiles 34VEM and 34VEN have an overlap of 9.5x100 km, 34VEN is not annotated from the overlapping part to prevent data leakage between splits.</div> <br> <h3>Annotation process</h3> The minimum size for an object to be considered as a potential marine vessel was set to 2x2 pixels. Three separate acquisitions for each location were used to detect smallest objects, so that if an object was located at the same place in all images, then it was left unannotated. The data were annotated by two experts. <div> </div> <table style="width: 63.327%; height: 391.876px;"> <tbody> <tr style="height: 39.1875px;"> <td style="width: 72.7285%; height: 39.1875px;"><strong>Product name</strong></td> <td style="width: 23.0224%; height: 39.1875px;"><strong>Number of annotations</strong></td> </tr> <tr style="height: 19.5938px;"> <td style="width: 72.7285%; height: 19.5938px;"> <div> <div>S2A_MSIL1C_20220515T100031_N0510_R122_T34VEM_20240617T162344.SAFE</div> </div> </td> <td style="width: 23.0224%; height: 19.5938px;">183</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 72.7285%; height: 19.5938px;"> <div> <div>S2B_MSIL1C_20220619T100029_N0510_R122_T34VEM_20240627T204751.SAFE</div> </div> </td> <td style="width: 23.0224%; height: 19.5938px;">519</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 72.7285%; height: 19.5938px;"> <div> <div>S2A_MSIL1C_20220721T095041_N0510_R079_T34VEM_20240712T224506.SAFE</div> </div> </td> <td style="width: 23.0224%; height: 19.5938px;">1518</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 72.7285%; height: 19.5938px;"> <div> <div>S2A_MSIL1C_20220813T095601_N0510_R122_T34VEM_20240717T115958.SAFE</div> </div> </td> <td style="width: 23.0224%; height: 19.5938px;">1371</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 72.7285%; height: 19.5938px;"> <div> <div>S2B_MSIL1C_20220606T095029_N0510_R079_T35VLG_20240619T111429.SAFE</div> </div> </td> <td style="width: 23.0224%; height: 19.5938px;">277</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 72.7285%; height: 19.5938px;"> <div> <div>S2B_MSIL1C_20220626T095039_N0510_R079_T35VLG_20240620T013500.SAFE</div> </div> </td> <td style="width: 23.0224%; height: 19.5938px;">1205</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 72.7285%; height: 19.5938px;"> <div> <div>S2B_MSIL1C_20220703T094039_N0510_R036_T35VLG_20240702T075354.SAFE</div> </div> </td> <td style="width: 23.0224%; height: 19.5938px;">746</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 72.7285%; height: 19.5938px;"> <div> <div>S2A_MSIL1C_20220721T095041_N0510_R079_T35VLG_20240712T224506.SAFE</div> </div> </td> <td style="width: 23.0224%; height: 19.5938px;">971</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 72.7285%; height: 19.5938px;"> <div> <div>S2A_MSIL1C_20220627T100611_N0510_R022_T34WFT_20240628T041908.SAFE</div> </div> </td> <td style="width: 23.0224%; height: 19.5938px;">122</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 72.7285%; height: 19.5938px;"> <div> <div>S2B_MSIL1C_20220712T100559_N0510_R022_T34WFT_20240718T033027.SAFE</div> </div> </td> <td style="width: 23.0224%; height: 19.5938px;">162</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 72.7285%; height: 19.5938px;"> <div> <div>S2B_MSIL1C_20220828T095549_N0510_R122_T34WFT_20240708T035231.SAFE</div> </div> </td> <td style="width: 23.0224%; height: 19.5938px;">98</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 72.7285%; height: 19.5938px;"> <div> <div>S2B_MSIL1C_20210714T100029_N0500_R122_T34VEN_20230224T120043.SAFE</div> </div> </td> <td style="width: 23.0224%; height: 19.5938px;">450</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 72.7285%; height: 19.5938px;"> <div> <div>S2B_MSIL1C_20220619T100029_N0510_R122_T34VEN_20240627T204751.SAFE</div> </div> </td> <td style="width: 23.0224%; height: 19.5938px;">66</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 72.7285%; height: 19.5938px;"> <div> <div>S2A_MSIL1C_20220624T100041_N0510_R122_T34VEN_20240714T110124.SAFE</div> </div> </td> <td style="width: 23.0224%; height: 19.5938px;">424</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 72.7285%; height: 19.5938px;"> <div> <div>S2A_MSIL1C_20220813T095601_N0510_R122_T34VEN_20240717T115958.SAFE</div> </div> </td> <td style="width: 23.0224%; height: 19.5938px;">399</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 72.7285%; height: 19.5938px;"> <div> <div>S2A_MSIL1C_20220617T100611_N0510_R022_T34VER_20240627T094433.SAFE</div> </div> </td> <td style="width: 23.0224%; height: 19.5938px;">83</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 72.7285%; height: 19.5938px;"> <div> <div>S2B_MSIL1C_20220712T100559_N0510_R022_T34VER_20240718T033027.SAFE</div> </div> </td> <td style="width: 23.0224%; height: 19.5938px;">184</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 72.7285%; height: 19.5938px;">S2A_MSIL1C_20220826T100611_N0510_R022_T34VER_20240705T062429.SAFE</td> <td style="width: 23.0224%; height: 19.5938px;">88</td> </tr> </tbody> </table> <br><br> <h3>Annotation statistics</h3> <br>Sentinel-2 images have spatial resolution of 10 m, so below statistics can be converted to pixel sizes by dividing them by 10 (diameter) or 100 (area).</div> <div> <table> <tbody> <tr> <td> </td> <td><strong>mean</strong></td> <td><strong>min</strong></td> <td><strong>25%</strong></td> <td><strong>50%</strong></td> <td><strong>75%</strong></td> <td><strong>max</strong></td> </tr> <tr> <td><strong>Area (m²)</strong></td> <td>5305.7</td> <td>567.9</td> <td>1629.9</td> <td>2328.2</td> <td>5176.3</td> <td>414795.7</td> </tr> <tr> <td><strong>Diameter (m)</strong></td> <td>92.5</td> <td>33.9</td> <td>57.9</td> <td>69.4</td> <td>108.3</td> <td>913.9</td> </tr> </tbody> </table> <br><br> <div>As most of the annotations cover also most of the wake of the marine vessel, the bounding boxes are significantly larger than a typical boat. There are a few annotations larger than 100 000 m², which are either cruise or cargo ships that are travelling along ordinal directions instead of cardinal directions, instead of e.g. smaller leisure boats.</div> <br> <div>Annotations typically have diameter less than 100 meters, and the largest diameters correspond to similar instances than the largest bounding box areas.</div> <br> <h3>Train-test-split</h3> <br> <div>We used tiles 34VEN and 34VER as the test dataset. For validation, we split the other three tile areas into 5x5 equal sized grid, and used 20 % of the area (i.e 5 cells) for the validation. The same split also makes it possible to do cross-validation.</div> <div> </div> <div> </div> <div> </div> </div> <div> <h3>Post-processing</h3> </div> <div><br> <div>Before evaluating, the predictions for the test set are cleaned using the following steps:</div> <br> <div>1. All prediction whose centroid points are not located on water are discarded. The water mask used contains layers `jarvi` (Lakes), `meri` (Sea) and `virtavesialue` (Rivers as polygon geometry) from the Topographical database by the National Land Survey of Finland. Unfortunately this also discards all points not within the Finnish borders.</div> <div>2. All predictions whose centroid points are located on water rock areas are discarded. The mask is the layer `vesikivikko` (Water rock areas) from the Topographical database.</div> <div>3. All predictions that contain an above water rock within the bounding box are discarded. The mask contains classes `38511`, `38512`, `38513` from the layer `vesikivi` in the Topographical database.</div> <div>4. All predictions that contain a lighthouse or a sector light within the bounding box are discarded. Lighthouses and sector lights come from Väylävirasto data, `ty_njr` class ids are 1, 2, 3, 4, 5, 8</div> <div>5. All predictions that are wind turbines, found in Topographical database layer `tuulivoimalat`</div> <div>6. All predictions that are obviously too large are discarded. The prediction is defined to be "too large" if either of its edges is longer than 750 meters.</div> </div> <div> </div> <div>Model checkpoint for the best performing model is available on Hugging Face platform: <a href="https://huggingface.co/mayrajeo/marine-vessel-detection-yolov8">https://huggingface.co/mayrajeo/marine-vessel-detection-yolo</a><br> <h2>Usage</h2> The simplest way to chip the rasters into suitable format and convert the data to COCO or YOLO formats is to use <a href="https://github.com/mayrajeo/geo2ml">geo2ml</a>. First download the raw mosaics and convert them into GeoTiff files and then use the following to generate the datasets. <div> </div> To generate COCO format dataset run</div> <div> </div> <div> <pre><code>from geo2ml.scripts.data import create_coco_dataset raster_path = '<path_to_raster>' outpath = '<path_to_save_the_dataset>' poly_path = '<path_to_gpkg>' layer = '<date_of_raster>' create_coco_dataset(raster_path=raster_path, polygon_path=poly_path, target_column='id', gpkg_layer=layer, outpath=outpath, save_grid=False, dataset_name='<name_of_dataset>', gridsize_x=320, gridsize_y=320, ann_format='box', min_bbox_area=0)</code></pre> </div> <div><br> <div>To generate YOLO format dataset run</div> <div> <pre><code>from geo2ml.scripts.data import create_yolo_dataset raster_path = '<path_to_raster>' outpath = '<path_to_save_the_dataset>' poly_path = '<path_to_gpkg>' layer = '<date_of_raster>' create_yolo_dataset(raster_path=raster_path, polygon_path=poly_path, target_column='id', gpkg_layer=layer, outpath=outpath, save_grid=False, gridsize_x=320, gridsize_y=320, ann_format='box', min_bbox_area=0)</code></pre> </div> </div>
Data for "Detection of metabolite-protein interactions in complex biological samples by high-resolution relaxometry: towards interactomics by NMR"
<p>Raw NMR data for relaxometry experiments, divided by donor sample. For every donor sample 2 or 3 different samples were used in order to record data at 19 different magnetic fields.</p> <p>Data from fast field-cycling relaxometry. All the data is in one xlsx file, divided by donor sample.</p> <p>Relaxometry results for alanine, lactate, creatinine and glutamine, obtained from the fitting of their relaxation decays recorded at 19 different fields, divided by donor sample.</p>
WaveFake: A data set to facilitate audio DeepFake detection
<p>The main purpose of this data set is to facilitate research into audio DeepFakes. We hope that this work helps in finding new detection methods to prevent such attempts. These generated media files have been increasingly used to commit <a href="https://www.vice.com/en/article/pkyqvb/deepfake-audio-impersonating-ceo-fraud-attempt">impersonation attempts</a> or <a href="https://www.wired.com/story/telegram-still-hasnt-removed-an-ai-bot-thats-abusing-women/">online harassment</a>. You can find the accompanying code repository on <a href="https://github.com/RUB-SysSec/WaveFake">GitHub</a>.</p> <p>The data set consists of 104,885 generated audio clips (16-bit PCM wav). We examine multiple networks trained on two reference data sets. First, the <a href="https://keithito.com/LJ-Speech-Dataset/">LJSpeech</a> data set consisting of 13,100 short audio clips (on average 6 seconds each; roughly 24 hours total) read by a female speaker. It features passages from 7 non-fiction books and the audio was recorded on a MacBook Pro microphone. Second, we include samples based on the <a href="https://sites.google.com/site/shinnosuketakamichi/publication/jsut">JSUT</a> data set, specifically, basic5000 corpus. This corpus consists of 5,000 sentences covering all basic kanji of the Japanese language (4.8 seconds on average; roughly 6.7 hours total). The recordings were performed by a female native Japanese speaker in an anechoic room. Finally, we include samples from a full text-to-speech pipeline (16,283 phrases; 3.8s on average; roughly 17.5 hours total). Thus, our data set consists of approximately 175 hours of generated audio files in total. Note that we do not redistribute the reference data.</p> <p>We included a range of architectures in our data set:</p> <ul> <li><a href="https://arxiv.org/abs/1910.06711">MelGAN</a></li> <li><a href="https://arxiv.org/abs/1910.11480">Parallel WaveGAN</a></li> <li><a href="https://arxiv.org/abs/2005.05106">Multi-Band MelGAN</a></li> <li><a href="http://arxiv.org/abs/2005.05106">Full-Band MelGAN</a></li> <li><a href="https://arxiv.org/abs/2010.05646">HiFi-GAN</a></li> <li><a href="https://arxiv.org/abs/1811.00002">WaveGlow</a></li> </ul> <p>Additionally, we examined a bigger version of MelGAN and include samples from a full TTS-pipeline consisting of a conformer and parallel WaveGAN model.</p> <p><strong>Collection Process</strong></p> <p>For WaveGlow, we utilize the <a href="https://github.com/NVIDIA/waveglow">official implementation</a> (commit 8afb643) in conjunction with the official pre-trained network on <a href="https://pytorch.org/hub/nvidia_deeplearningexamples_waveglow/">PyTorch Hub</a>. We use a popular implementation available on <a href="https://github.com/kan-bayashi/ParallelWaveGAN">GitHub</a> (commit 12c677e) for the remaining networks. The repository also offers pre-trained models. We used the pre-trained networks to generate samples that are similar to their respective training distributions, <a href="https://keithito.com/LJ-Speech-Dataset/">LJ Speech</a> and <a href="https://sites.google.com/site/shinnosuketakamichi/publication/jsut">JSUT</a>. When sampling the data set, we first extract Mel spectrograms from the original audio files, using the pre-processing scripts of the corresponding repositories. We then feed these Mel spectrograms to the respective models to obtain the data set. For sampling the full TTS results, we use the <a href="https://github.com/espnet/espnet">ESPnet</a> project. To make sure the generated phrases do not overlap with the training set, we downloaded the <a href="https://commonvoice.mozilla.org/en/datasets">common voices data set</a> and extracted 16.285 phrases from it.</p> <p>This data set is licensed with a CC-BY-SA 4.0 license.</p> <p>This work was supported by the Deutsche Forschungsgemeinschaft (DFG, German Research Foundation) under Germany's Excellence Strategy -- EXC-2092 CaSa -- 390781972.</p>
Reproduction package for paper "How far are we from reproducible research on code smell detection? A systematic literature review"
<p>Checklist and data extracted from publications analyzed for "How far are we from reproducible research on code smell detection? A systematic literature review" paper, together with processing scripts and calculations of Cohen's Kappa.</p> <p>Paper that describes details of the data is available here: https://doi.org/10.1016/j.infsof.2021.106783</p>
In-Plane and Out-of-Plane MEMS Piezoresistive Cantilever Sensors for Nanoparticle Mass Detection (Data)
<p>Origin projects, figures and LabVIEW software used for the article "In-Plane and Out-of-Plane MEMS Piezoresistive Cantilever Sensors for Nanoparticle Mass Detection", published in <em>Sensors </em>on 22 Jan 2020.</p>
Reference dataset for comparison of cloud detection algorithms for Sentinel-2 imagery
<p>Sentinel-2 cloud mask reference dataset generated and analyzed as part of Tarrio, K., Tang, X., Masek, J.G., Claverie, M., Ju, J., Qiu, S., Zhu, Z. and Woodcock, C.E., 2020. Comparison of cloud detection algorithms for Sentinel-2 imagery. Science of Remote Sensing, 2, p.100010. [https://www.sciencedirect.com/science/article/pii/S2666017220300092](https://www.sciencedirect.com/science/article/pii/S2666017220300092)</p> <p><strong>1. Reference masks</strong></p> <p><strong>Algorithms:</strong></p> <ul> <li>Fmask 1.x</li> <li>Fmask 2.x</li> <li>Fmask 4.x</li> <li>Tmask</li> <li>Sen2Cor</li> <li>MAJA</li> <li>LaSRC</li> </ul> <p><strong>Locations:</strong></p> <ul> <li>South Africa (35JPM)</li> <li>Senegal (28PDC)</li> <li>Switzerland (32TLT)</li> <li>France (31TCJ, 31TFJ)</li> <li>Morocco (29RNQ)</li> </ul> <p><strong>Standardized legend:</strong></p> <p>Original algorithm outputs were standardized to the same categorical legend.</p> <ul> <li>0 = clear land</li> <li>1 = clear water</li> <li>2 = cloud shadow</li> <li>3 = snow/ice</li> <li>4 = cloud</li> </ul> <p>All reference masks processed to both 10m and 30m resolution, with the exception of Tmask, which is available only at a 30m resolution.</p> <p><strong>Mask naming convention:</strong></p> <p>All processed masks are named according to the following convention:<br> M<*resolution*><*S2 MGRS tile ID*><*YYYY*><*DOY*><*algorithm*><br> e.g. **M30T28PDC2016351TMASK**</p> <p><br> <strong>2. Interpreted sample points</strong></p> <p>Sample points were selected based on agreement among different map products. This record includes a shapefile with the final interpretations for each of the sampled sites. (See publication for additional information.)</p>
The detection of radio emission from known X-ray flaring star EXO 040830−7134.7
<p>This is the radio light curve of known X-ray flaring star EXO 040830−7134.7 observed by MeerKAT as part of ThunderKAT. These data are part of a publication in the Monthly Notice of the Royal Astronomical Society (Driessen et al., Accepted 2021 November 25. Received 2021 November 25; in original form 2021 August 25).</p> <p>The light curve is from the full-time-integration, full-frequency-integration images of VW Hyi, as processed by the LOFAR Transients Pipeline (<a href="https://tkp.readthedocs.io/en/latest/introduction.html">TraP</a>).</p> <p>The columns in the file are:</p> <ul> <li>mjd: the modified Julian Date (MJD) of the observation. The MJD is given by MJD=JD-2400000.5 where JD is the Julian Date</li> <li>f_int_Jy: the integrated flux density of the source in Jansky (Jy) determined by the LOFAR TraP</li> <li>f_int_err_Jy: the uncertainty on f_int_Jy in Jansky determined by the LOFAR TraP</li> <li>freq_eff_Hz: the effect frequency in Hertz (Hz) as determined by the LOFAR TraP</li> <li>taustart_ts: the ISO 8601 time of the observation in Coordinated Universal Time (UTC)</li> </ul> <p>The files were made using the Pandas package, so we recommend Python users load them using</p> <pre><code>import pandas as pd pd.read_csv(filename, comment='#')</code></pre> <p>If you use the data shared here please ensure that you cite the MNRAS paper (Driessen at al. 2021) and the Zenodo DOI: 10.5281/zenodo.5084298.</p> <p>The MeerKAT telescope is operated by the South African Radio Astronomy Observatory, which is a facility of the National Research Foundation, an agency of the Department of Science and Innovation.<br> LND acknowledges support from the European Research Council (ERC) under the European Union's Horizon 2020 research and innovation programme (grant agreement No 694745).</p>
Supramolecular Self-Healing Sensor Fiber Composites for Damage Detection in Piezoresistive Electronic Skin for Soft Robots
<p>Self-healing materials can prolong the lifetime of structures and products by enabling the repairing of damage. However, detecting the damage and the progress of the healing process remains an important issue. In this study, self-healing, piezoresistive strain sensor fibers (ShSFs) are used for detecting strain deformation and damage in a self-healing elastomeric matrix. The ShSFs were embedded in the self-healing matrix for the development of self-healing sensor fiber composites (ShSFC) with elongation at break values of up to 100%. A quadruple hydrogen-bonded supramolecular elastomer was used as a matrix material. The ShSFCs exhibited a reproducible and monotonic response. The ShSFCs were investigated for use as sensorized electronic skin on 3D-printed soft robotic modules, such as bending actuators. Depending on the bending actuator module, the electronic skin was loaded under either compression (pneumatic-based module) or tension (tendon-based module). In both configurations, the ShSFs could be successfully used as deformation sensors, and in addition, detect the presence of damage based on the sensor signal drift. The sensor under tension showed better recovery of the signal after healing, and smaller signal relaxation. Even with the complete severing of the fiber, the piezoresistive properties returned after the healing, but in that case, thermal heat treatment was required. With their resilient response and self-healing properties, the supramolecular fiber composites can be used for the next generation of soft robotic modules</p>
BirdVox-296h: a large-scale dataset for detection and classification of flight calls
<p>BirdVox 296 hours dataset (BirdVox-296h)<br> ====================================</p> <p>Version 2.1, May 2022.</p> <p><br> Created By<br> ----------</p> <p>Andrew Farnsworth (1), Benjamin Mark Van Doren (1), Steve Kelling (1), Vincent Lostanlen (2), Justin Salamon (3), Aurora Cramer (4), Juan Pablo Bello (4)</p> <p>(1): Cornell Lab of Ornithology (CLO)<br> (2): Laboratoire des Sciences du Numérique de Nantes (LS2N), CNRS<br> (3): Adobe Research<br> (4): New York University</p> <p>https://wp.nyu.edu/birdvox<br> <br> </p> <p>Description<br> ---------------</p> <p>The BirdVox-296h dataset contains 148 audio recordings, each two hours in duration. These recordings come from ROBIN autonomous recording units, placed near Ithaca, NY, USA during the fall 2015. They were captured by nine different sensors, originally numbered 1, 2, 3, 4, 5, 6, 7, 8, and 10.<br> <br> Ornithologist Andrew Farnsworth used the Raven software to pinpoint and label every avian flight call in time and frequency. He found 26138 sound events, of which 21546 are flight calls from Passeriformes. Of those, 13385 are identifiable in terms of family, and 8669 are identifiable in terms of both family and species. The annotation process took over 600 hours.</p> <p>The dataset can be used, among other things, for the research, development and testing of machine listening models for bird migration monitoring.</p> <p> </p> <p>Data Files<br> ------------</p> <p>The BirdVox-296h_wav folder contains 148 recordings as WAV files, sampled at 24 kHz, with a single channel (mono). Each recording lasts exactly two hours and is named according to the following format:</p> <p>YYYY-MM-DD_hh-mm-ss_unitUU.wav</p> <p>Where Y means Year, M means Month, D means Day, h means hour, m means minute, and s means second. This date format corresponds to the start time of the recording file, expressed in Coordinated Universal Time (UTC).</p> <p>The field UU contains two digits corresponding to the identifier of the autonomous recording unit (i.e., bioacoustic sensor). UU is either equal to 01, 02, 03, 04, 05, 06, 07, 08, or 10. Note that 09 is absent from the list because sensor 09 failed during the acquisition campaign.</p> <p> </p> <p>Metadata Files<br> -------------------</p> <p>The BirdVox-296h_csv-annotations folder contains CSV files, one for each audio file. The columns of each CSV file are:</p> <p>ID,Time (s),Frequency (Hz),Taxonomy Code,Fine Label,Medium Label,Coarse Label</p> <p><br> "Taxonomy Code" is compliant with the BirdVoxClassify software: github.com/BirdVox/BirdVoxClassify</p> <p>"Fine Label", "Medium Label", and "Coarse Label" most often correspond to species, family and order respectively.</p> <p> </p> <p>The BirdVox-296h_gps-coordinates.csv file contains the approximate GPS coordinates of the sensors (latitudes and longitudes rounded to 2 decimal points) of all nine sensors.</p> <p> </p> <p> </p> <p>Conditions of Use<br> -----------------</p> <p>Dataset created by Andrew Farnsworth, Steve Kelling, Vincent Lostanlen, Justin Salamon, Aurora Cramer, and Juan Pablo Bello.</p> <p>The BirdVox-full-night dataset is offered free of charge under the terms of the Creative Commons Attribution 4.0 International (CC BY 4.0) license:<br> https://creativecommons.org/licenses/by/4.0/</p> <p>The dataset and its contents are made available on an "as is" basis and without warranties of any kind, including without limitation satisfactory quality and conformity, merchantability, fitness for a particular purpose, accuracy or completeness, or absence of errors. Subject to any liability that may not be excluded or limited by law, Cornell Lab of Ornithology is not liable for, and expressly excludes all liability for, loss or damage however and whenever caused to anyone by any use of the BirdVox-full-night dataset or any part of it.</p> <p> </p> <p>Feedback<br> -------------</p> <p>Please help us improve BirdVox-296h by sending your feedback to:<br> vincent.lostanlen@ls2n.fr and af27@cornell.edu</p> <p>In case of a problem, please include as many details as possible.</p> <p> </p> <p>Acknowledgements<br> --------------------------</p> <p>Jessie Barry, Ian Davies, Tom Fredericks, Jeff Gerbracht, Sara Keen, Holger Klinck, Anne Klingensmith, Ray Mack, Peter Marchetto, Ed Moore, Matt Robbins, Ken Rosenberg, and Chris Tessaglia-Hymes.</p> <p>We acknowledge that the land on which the data was collected is the unceded territory of the Cayuga nation, which is part of the Haudenosaunee (Iroquois) confederacy.</p>
Dataset for "CVD growth of self-assembled 2D and 1D WS2 nanomaterials for the ultrasensitive detection of NO2"
<p>This file contains the raw data used in the paper entitled CVD growth of self-assembled 2D and 1D WS2 nanomaterials for the ultrasensitive detection of NO2 published in Sensors and Actuators: B. Chemical 326 (2021) 128813</p> <p>DOI: <a href="https://doi.org/10.1016/j.snb.2020.128813">10.1016/j.snb.2020.128813</a></p>
Outputs of the Jupyter Notebook - Detecting floating objects using Deep Learning and Sentinel-2 imagery
<p>The dataset contains the outputs of the notebook "Detecting floating objects using Deep Learning and Sentinel-2 imagery" published in the ocean modelling section of The Environmental Data Science Book.</p> <p><strong>Contributions</strong></p> <p><em>Notebook</em></p> <ul> <li> <p>Jamila Mifdal (author), European Space Agency Φ-lab, <a href="https://github.com/jmifdal">@jmifdal</a></p> </li> <li> <p>Raquel Carmo (author), European Space Agency Φ-lab, <a href="https://github.com/raquelcarmo">@raquelcarmo</a></p> </li> <li> <p>Alejandro Coca-Castro (reviewer), The Alan Turing Institute, <a href="https://github.com/acocac">@acocac</a></p> </li> </ul> <p><em>Modelling codebase</em></p> <ul> <li> <p>Jamila Mifdal (author), European Space Agency Φ-lab, <a href="https://github.com/jmifdal">@jmifdal</a></p> </li> <li> <p>Raquel Carmo (author), European Space Agency Φ-lab, <a href="https://github.com/raquelcarmo">@raquelcarmo</a></p> </li> <li> <p>Marc Rußwurm (author), EPFL-ECEO, <a href="https://github.com/MarcCoru">@marccoru</a></p> </li> </ul>
Dataset for "LoRa Sensor Network Development for Air Quality Monitoring or Detecting Gas Leakage Events; DOI: 10.3390/s20216225"
<p>This excel file contains the raw data used in the paper " LoRa Sensor Network Development for Air Quality Monitoring or Detecting Gas Leakage Events; DOI: 10.3390/s20216225 " In particular it comprises sensor measurements and pollutant data from the automated air quality monitoring stations in the Tarragona area.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.