Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,487
datasets available to search
ShareScore release 0.9.0
Dataset results
1,487 results for “tags”
S1000 corpus, large-scale tagging results and other supplementary files
<p>Data associated with the S1000 corpus</p><p>The tagger software for which the dictionary files in <a href="https://zenodo.org/api/files/b8a0e221-3cc3-4db5-a2e9-f19a1bd2e5cb/tagger-organisms-dictionary-S1000.tar.gz">tagger-organisms-dictionary-S1000.tar.gz </a>can be used with can be found here: <a href="https://github.com/larsjuhljensen/tagger">https://github.com/larsjuhljensen/tagger</a></p><p>The online version of the annotation documentation can be found here: <a href="https://katnastou.github.io/s1000-corpus-annotation-guidelines/">https://katnastou.github.io/s1000-corpus-annotation-guidelines/</a></p><p>The S1000 corpus split in training, development and test sets in BRAT format can be found in <a href="https://zenodo.org/api/records/10285825/files/S1000-corpus.tar.gz">S1000-corpus.tar.gz</a><a href="https://zenodo.org/api/files/b8a0e221-3cc3-4db5-a2e9-f19a1bd2e5cb/S1000-corpus.tar.gz?versionId=ac7ce430-c265-49bb-8c8f-9b5f8e271cbe"> </a>and in CoNLL format here: <a href="https://zenodo.org/api/files/b8a0e221-3cc3-4db5-a2e9-f19a1bd2e5cb/s1000-conll.tar.gz">s1000-conll.tar.gz</a></p><p>The tagging results of Jensenlab tagger for the S1000 test set are here: <a href="https://zenodo.org/api/files/b8a0e221-3cc3-4db5-a2e9-f19a1bd2e5cb/S1000-jensenlab-tagger.tar.gz?versionId=d8d9c9f5-ee3b-4738-aefa-a4a95475d25d">S1000-jensenlab-tagger.tar.gz</a></p><p>The result from the large scale run in entire PubMed and PMC Open Access articles for Jensenlab tagger is provided here: <a href="https://zenodo.org/api/files/b8a0e221-3cc3-4db5-a2e9-f19a1bd2e5cb/Jensenlab_tagger_large_scale_matches_with_rank.tsv.gz?versionId=48825928-9fc9-423c-8a4c-4f8994e95805">Jensenlab_tagger_large_scale_matches_with_rank.tsv.gz</a></p><p>The model used for the large scale run of the transformer-based method is here: <a href="https://zenodo.org/api/files/b8a0e221-3cc3-4db5-a2e9-f19a1bd2e5cb/S1000_Transformer_based_tagger_large_scale_model.tar.gz?versionId=8e974f64-9abc-4449-a377-e3f97e91d612">S1000_Transformer_based_tagger_large_scale_model.tar.gz</a> and the results from the large scale tagging here: <a href="https://zenodo.org/api/files/b8a0e221-3cc3-4db5-a2e9-f19a1bd2e5cb/Transformer_based_tagger_large_scale_matches_with_rank.tsv.zip?versionId=dc21a6ba-9763-4130-9f02-0341a885c692">Transformer_based_tagger_large_scale_matches_with_rank.tsv.zip</a></p>
Arctic Grayling length, weight and tag data from Arctic LTER Streams project, Toolik Filed Station Alaska, 1985 to 2018
Since 1983, the Streams Project at the Toolik Field Station has monitored physical, chemical, and biological parameters in a 5-km, fourth-order reach of the Kuparuk River near its intersection with the Dalton Highway and the Trans-Alaska Pipeline. In 1989, similar studies were begun on a 3.5-km, third-order reach of a second stream, Oksrukuyik Creek. Fish were collected on each river. Station locations, representing kilomter values certain distances from original phosphorus dripper (see method) were noted. 1985 to 2012 long-term tagging file for Arctic Grayling (Thymallus arcticus) on the Kuparuk River. All grayling adults and juveniles captured during the field season are measured, weighed, tagged and released. Grayling were tagged originally with a colored tag with a number. In 1993, researchers started pit tagging the grayling. These pit tags can be read with an antenna to track the migration of the grayling throughout the Kuparuk River system. Arctic grayling young-of-the-year (YOY) were caught multiple times during each summer and measured and weighed as well. This file combines the data from the following data sets: Dataset ID Short name 10325 1985-2012_Kuparuk_Grayling_Tags 10327 1986-2012_Kuparuk_YOY 10329 1989-2011_Oksrukuyik_Grayling_Tags 10330 1989-2012_Oksrukuyik_YOY
Fish tag data remotely detected using whole stream antennas or hand held tag readers in the Kuparuk, Itkilik, and Sagavanirktok drainages near Toolik Field Station, Alaska, from 2010 to 2017
From 2009 to 2017, the FISHSCAPE Project (grant numbers 1719267, 1417754, and 0902153), based at Toolik Field Station, has monitored physical, chemical, and biological parameters within three watersheds: The Kuparuk (including Toolik Lake and Toolik outlet stream); The Sagavanirktok (primarily Oksrukuyik Creek, but also including sections of the Ailish and Atigun Rivers and the Galbraith Lakes); and The Itkillik (primarily the I-Minus outlet stream, a tributary that that feeds into the Itkilik River). Target species were primarily Arctic grayling and Lake trout, although Arctic char, Burbot, Dolly varden, round whitefish, and slimey sculpin were also captured. This file contains the detectioned fish tags using whole stream or hand-held antennas in the three watersheds. We had no field season in 2014 and thus did not deploy antennaes. Fish were tagged with Passive Integrated Transponder (PIT) tags which can be read with a whole stream antenna to track the migration of the fish, predominately Arctic grayling, throughout the systems. Fish tags detected with a handheld readers are designated in Site ID as "XXX_capture". For "capture" fish time is arbitraily set at '7:00:00'' of the day of capture and tagging because actual time was not recorded. The individual fish data (date, tag number, length, weight, species) associated with the tag can be found in the 2009-2017_FISHSCAPE_fish_tagging file.
Fish tagging data (length, weight, tag number) from the Kuparuk, the Sagavanirktok (primarily Oksrukuyik Creek) and the Itkillik (primarily the I-Minus outlet stream) watersheds, 2009 - 2017
Since 2009, the FISHSCAPE Project (grant number 1719267, 1417754, and 0902153), based at Toolik Field Station, has monitored physical, chemical, and biological parameters within three watersheds: The Kuparuk (including Toolik Lake and Toolik outlet stream); The Sagavanirktok (primarily Oksrukuyik Creek, but also including sections of the Ailish and Atigun Rivers and the Galbraith Lakes); and The Itkillik (primarily the I-Minus outlet stream, a tributary that that feeds into the Itkilik River). Target species were primarily Arctic grayling and Lake trout, although Arctic char, Burbot, Dolly varden, round whitefish, and slimey sculpin were also captured. Fish were collected on each river/lake. Coordinates and/or specific station locations were noted. All fish captured during the field season are measured, weighed, tagged (if large enough) and released. If fish were not previously tagged, they were tagged with Passive Integrated Transponder (PIT) tags which can be read with a whole stream antenna to track the migration of the fish, predminately Arctic grayling, throughout the systems.
Plexcitonic Nanorattles as Highly Efficient SERS-Encoded Tags
<p>Recent publication: </p> <h1>Plexcitonic Nanorattles as Highly Efficient SERS-Encoded Tags</h1> <div> </div> <div> <div> <div> <div><span><a href="https://onlinelibrary.wiley.com/authored-by/Est%C3%A9vez%E2%80%90Varela/Carla">Carla Estévez-Varela</a><span>, </span></span><span><a href="https://onlinelibrary.wiley.com/authored-by/N%C3%BA%C3%B1ez%E2%80%90S%C3%A1nchez/Sara">Sara Núñez-Sánchez</a><span>, </span></span><span><a href="https://onlinelibrary.wiley.com/authored-by/Pi%C3%B1eiro%E2%80%90Varela/Paula">Paula Piñeiro-Varela</a><span>, </span></span><span><a href="https://onlinelibrary.wiley.com/authored-by/Aberasturi/Dorleta+Jim%C3%A9nez">Dorleta Jiménez de Aberasturi</a><span>, </span></span><span><a href="https://onlinelibrary.wiley.com/authored-by/Liz%E2%80%90Marz%C3%A1n/Luis+M.">Luis M. Liz-Marzán</a><span>, </span></span><span><a href="https://onlinelibrary.wiley.com/authored-by/P%C3%A9rez%E2%80%90Juste/Jorge">Jorge Pérez-Juste</a><span>, </span></span><span><a href="https://onlinelibrary.wiley.com/authored-by/Pastoriza%E2%80%90Santos/Isabel">Isabel Pastoriza-Santos</a></span></div> </div> </div> </div> <div> <div><span>First published: </span><span>27 November 2023</span></div> <br> <div><a href="https://doi.org/10.1002/smll.202306045">https://doi.org/10.1002/smll.202306045</a></div> </div> <p> </p> <p>Abstract: Plexcitonic nanoparticles exhibit strong light-matter interactions, mediated by localized surface plasmon resonances, and thereby promise potential applications in fields such as photonics, solar cells, and sensing, among others. Herein, these light-matter interactions are investigated by UV-visible and surface-enhanced Raman scattering (SERS) spectroscopies, supported by finite-difference time-domain (FDTD) calculations. Our results reveal the importance of combining plasmonic nanomaterials and J-aggregates with near-zero-refractive index. As plexcitonic nanostructures nanorattles are employed, based on J-aggregates of the cyanine dye 5,5,6,6-tetrachloro-1,1-diethyl-3,3-bis(4-sulfobutyl)benzimidazolocarbocyanine (TDBC) and plasmonic silver-coated gold nanorods, confined within mesoporous silica shells, which facilitate the adsorption of the J-aggregates onto the metallic nanorod surface, while providing high colloidal stability. Electromagnetic simulations show that the electromagnetic field is strongly confined inside the J-aggregate layer, at wavelengths near the upper plexcitonic mode, but it is damped toward the J-aggregate/water interface at the lower plexcitonic mode. This behavior is ascribed to the sharp variation of dielectric properties of the J-aggregate shell close to the plasmon resonance, which leads to a high opposite refractive index contrast between water and the TDBC shell, at the upper and the lower plexcitonic modes. This behavior is responsible for the high SERS efficiency of the plexcitonic nanorattles under both 633 nm and 532 nm laser illumination. SERS analysis showed a detection sensitivity down to the single-nanoparticle level and, therefore, an exceptionally high average SERS intensity per particle. These findings may open new opportunities for ultrasensitive biosensing and bioimaging, as superbright and highly stable optical labels based on the strong coupling effect.</p>
Audio tagging of avian dawn chorus recordings in California, Oregon, and Washington
<p><strong>General Summary</strong></p> <p>This acoustic data collection includes 1,575 5-minute soundscape recordings randomly selected from passive acoustic recordings made at 525 sites during 2022 on federally managed lands in western California, Oregon, and Washington, USA. We fully labeled 141 recordings (11.75 hrs) with 39,717 annotations for 118 sound types, including 58 avian species, two mammalian species, six aggregated biotic sounds, and eight non-biotic sound types. An additional 215 recordings were partially annotated with 1,466 annotations. The remaining unlabeled recordings have been included to facilitate novel research applications and methodological evaluations. Beyond the labeled soundscape recordings, we have included township and range identifications and 38 environmental covariates for each recording location.</p> <p><strong>Data Collection</strong></p> <p>Lesmeister et al. (2021) collected passive acoustic recordings during 2022 in support of long-term monitoring of federally threatened northern spotted owl (<em>Strix occidentalis caurina) </em>populations under the Northwest Forest Plan Effective Monitoring Program (U. S. Fish and Wildlife Service 1990, U. S. Department of Agriculture and U. S. Department of the Interior 1994). These data were collected at 643 hexagons that were randomly selected from a tessellation of 5 km2 hexagons covering the entire range of the northern spotted owl (Northern California, Oregon, Washington) under a selective constraint that hexagons contain ≥ 50 % forest-capable lands (<em>def.</em> forested lands or lands capable of developing closed-canopy forests) and be ≥ 25% federal ownership (Davis et al., 2011).</p> <p>Each hexagon was sampled by four Song Meter 4 (SM4) acoustic recording units (Wildlife Acoustics, Maynard, MA) deployed in a standardized spatial arrangement, such that recorders on a site were placed ≥ 500 m apart and were ≥ 200 m from the edge of the sampling hexagon boundary. Recorders were mounted to small trees (15 – 20 cm diameter at breast height) approximately 1.5 m above the ground and were placed on mid-to-upper slopes and ≥ 50 m from roads, trails, and streams. The SM4 devices each have two built-in omnidirectional microphones with a signal-to-noise ratio of 80 dB, typical at 1 kHz, and a recording bandwidth of 20 Hz – 48 kHz. Each device recorded ~11 hours of audio daily for six weeks from March to August at a sampling rate of 32 kHz. The daily recording schedule included a 4-hour window from two hours before sunrise to two hours after sunrise, a 4-hour window from one hour before sunset to 3 hours after sunset, and 10-minute recordings outside the two longer recording blocks at the start of every hour.</p> <p><strong>Data Sampling</strong></p> <p>The goal of this project was to develop a tagged audio dataset (hereafter project dataset) focused on the avian dawn chorus, which is an ecologically important period for the study of avian behavior (McNamara et al. 1987, Staicer et al. 1996, Zhang et al. 2015) and monitoring avian biodiversity (Bibby et al. 2000), but remains a challenging problem for acoustic classification systems (Duan et al. 2013, Stowell 2022). Passive acoustic monitoring on our sites occurs throughout the day. We filtered the full dataset to recordings collected between May and August during the hour immediately after sunrise. From the recordings meeting our filtering criteria, we randomly selected three 5-minute files from each site, which were assigned ordinal labels 'A, 'B,' or 'C.' The final project dataset comprised 131.25 hours of acoustic data.</p> <p><strong>Annotation Protocol</strong></p> <p>We randomly selected 141 sites from the project dataset and fully annotated each recording at a 2-second resolution. We applied labels to each 2-second window of the selected recordings following a predefined sound phonology library (available in the 'metadata.tsv' file), which concatenated the 2021 eBird taxonomy codes (Clements list; Clements et al. 2022) with standardized sonotype codes that incremented depending on the species repertoire (i.e., 'call_1,' 'song_1,' 'drum_1'). For example, 'herthr_song_1' is the label for Hermit Thrush, song_1. Unknown signals were labeled 'unknown,' and clips with no biotic signals (or noise classes of interest documented in metadata.tsv) were labeled 'empty.' Windows were labeled 'complete' and considered fully annotated when every signal was assigned an annotation. Files were deemed fully annotated when every 2-second window contained the 'complete' label.</p> <p><strong>Environmental Covariates</strong></p> <p>Sampling locations will not be published to afford protections for Federally Threatened or Endangered species which may occur on our sites. However, we provide the State, Township, and Range for each sampling location along with the site-specific values for 38 forest structure, topographic, and climatic environmental covariates developed by the Landscape Ecology, Modeling, Mapping, and Analysis group in the Pacific Northwest (<a href="https://lemma.forestry.oregonstate.edu/data">https://lemma.forestry.oregonstate.edu/data</a>; Ohmann and Gregory 2002). State, Township, and Range values are sufficient to explore geographic variation in species- or community-specific call and song phenology and the extracted environmental covariates may provide useful contextual information for novel machine-learning developments (Liu et al. 2018). </p> <p><strong>Description of Data Format</strong></p> <p>The fully annotated audio files can be accessed by downloading and extracting "annotated_recordings.zip." Partially annotated and non-annotated audio files can be accessed by downloading and extracting "additional_recordings_part_1.zip" or "additional_recordings_part_2.zip." Acoustic file names contain site and replicate indicators, such that file "Site_001_Rep_A.wav' was recorded on site 1 and is the A replicate random draw from the available set of dawn chorus recordings. The site and replicate numbers link to additional recording information in "files.tsv," annotations in "annotations.tsv" and "partial_annotations.tsv," as well as site and replicate specific environmental characteristics in "environmental_characteristics.tsv."</p> <p>Metadata describing sound classes and environmental characteristics can be found in "metadata.tsv," and "environmental_characteristics_metadata.tsv."</p> <p><strong>Acknowledgments</strong></p> <p>Acoustic data collection was funded and collected by the US Forest Service and the US Bureau of Land Management. Annotation work was funded by Google. We would also like to thank the many biologists that collected and processed the data compiled here. The use of trade or firm names in this publication is for reader information and does not imply endorsement by the U.S. Government of any product or service.</p>
Corpus of Occitan Written Traditional Folktales Annotated with Part-Of-Speech (OWT-Tag)
<p>This resource contains 5 extracts of texts in Occitan which were manually annotated with lemmas and parts-of-speech, following the Grace standard. It was produced during the ExpressioNarration project, funded by a Marie Curie Individual Fellowship, in order to evaluate the performance of an Occitan Part-Of-Speech tagger, Talismane, to the specifities of the corpus of the project called Oral Occitan (OcOr), also available on https://zenodo.org/record/1451753#.W78FJWOYSpo.<br> Each extract contains around 1500 words. They are extracted from 'Contes et proverbes populaires recueillis en armagnac et Contes populaires recueillis en agenais' de J.-F. Bladé, 'Coundes biarnés, couéilhuts aüs parsàas miéytadès dou péys dé Biarn' de J.-V. Lalanne, 'Contes populaires du Languedoc' de L. Lambert and 'Contes populaires recueillis dans la Grande-Lande' de F. Arnaudin.<br> The annotation process is described in the following article available on https://www.openscience.fr/IMG/pdf/iste_modocv1n1_2.pdf.</p>
White spruce trees tagged measured for total height and girth at 10 centimeter height, and leader length, Coldfoot, Alaska 2015, 2016
White spruce seedlings have colonized the site of the Coldfoot transplant garden (CF, 67°15′32″N, 150°10′12″W) since the original garden was established in 1982. Some trees are 2-3 meter tall. All seedlings and trees within the current (2014) garden were tagged, located with a Global Positioning System (GPS) receiver, and measured in 2015 and 2016 for total height and girth at 10 centimeter height and leader length.
Frequency Tagging of Syntactic Structure or Lexical Properties
Open the record for dataset details and reuse information.
Data: Methods for tagging an ectoparasite, the salmon louse Lepeophtheirus salmonis
<p>Monitoring individuals within populations is a cornerstone in evolutionary ecology, yet<span> </span>individual tracking of invertebrates and particularly parasitic organisms remains rare. To address this gap, we describe here a method for attaching radio frequency identification<span> </span>(RFID) tags to individual adult females of a marine ectoparasite, the salmon louse<span> </span><em><span>Lepeophtheirus salmonis</span></em>. Comparing two alternative types of glue, we found that one of them<span> </span>(2-octyl cyanoacrylate, <em><span>2oc</span></em>) gave a significantly higher tag retention rate than the other (ethyl<span> </span>2-cyanoacrylate, <em><span>e2c</span></em>). This glue comparison test also resulted in a higher loss rate of adult ectoparasites from the population where tagging was done using <em><span>2oc</span></em>, but this included males<span> </span>not tagged and thus could also suggest a mere tank effect. Corroborating this, a more extensive analysis using data collected over two years showed no significant difference in<span> </span>mortality after repeated exposure to the <em><span>2oc </span></em>glue, nor did it show any significant effect of the<span> </span>tagging procedure on the reproduction of female salmon lice. The proportion of RFID-tagged<span> </span>individuals followed a negative exponential decline, with tag retention among the living<span> </span>female population generally high. The projected retention was found to be about 88% after<span> </span>30 days or 80% after 60 days, although one of the four batches of glue used, purchased from<span> </span>a different supplier, appeared to give significantly lower tag retention and with greater initial<span> </span>loss (74% and 60% respectively). Overall, we find that RFID tagging is a simple and effective technology that enables documenting individual life histories for invertebrates of a suitable size, including marine and parasitic species, and that it can be used over long periods of study.</p>
Dataset for "ZnO decorated Graphene-based NFC tag for personal NO2 exposure monitoring during a workday"
<p>Dataset with all measurements performed and related to the publication "ZnO decorated Graphene-based NFC tag for personal NO2 exposure<br>monitoring during a workday" Published in Sensors MDPI 2024 by A. Santos and co-workers.</p>
Bio-logger Ethogram Benchmark: A benchmark for computational analysis of animal behavior, using animal-borne tags
<p>This repository contains the datasets and experiment results presented in our <a href="https://arxiv.org/abs/2305.10740">arxiv paper</a>:</p> <blockquote> <p>B. Hoffman, M. Cusimano, V. Baglione, D. Canestrari, D. Chevallier, D. DeSantis, L. Jeantet, M. Ladds, T. Maekawa, V. Mata-Silva, V. Moreno-González, A. Pagano, E. Trapote, O. Vainio, A. Vehkaoja, K. Yoda, K. Zacarian, A. Friedlaender, "A benchmark for computational analysis of animal behavior, using animal-borne tags," 2023.</p> </blockquote> <p>Standardized code to implement, train, and evaluate models can be found at <a href="https://github.com/earthspecies/BEBE/">https://github.com/earthspecies/BEBE/</a>. </p> <p>Please note the licenses in each dataset folder.</p> <p><strong>Zip folders beginning with "formatted":</strong> These are the datasets we used to run the experiments reported in the benchmark paper. </p> <p><strong>Zip folders beginning with "raw": </strong>These are the unprocessed datasets used in BEBE. Code to process these raw datasets into the formatted ones used by BEBE can be found at <a href="https://github.com/earthspecies/BEBE-datasets/">https://github.com/earthspecies/BEBE-datasets/</a>.</p> <p><strong>Zip folders beginning with "experiments": </strong>Results of the cross-validation experiments reported in the paper, as well as hyperparameter optimization. Confusion matrices for all experiments can also be found here. Note that dt, rf, and svm refer to the feature set from Nathan et al., 2012.</p> <p><em>Results used in Fig. 4 of <a href="https://arxiv.org/abs/2305.10740">arxiv paper</a> (deep neural networks vs. classical models)</em><br>{dataset}_ harnet_nogyr<br>{dataset}_CRNN<br>{dataset}_CNN<br>{dataset}_dt<br>{dataset}_rf<br>{dataset}_svm<br>{dataset}_wavelet_dt<br>{dataset}_wavelet_rf<br>{dataset}_wavelet_svm</p> <p><em>Results used in Fig. 5D of <a href="https://arxiv.org/abs/2305.10740">arxiv paper</a> (full data setting)<br></em>If dataset contains gyroscope (HAR, jeantet_turtles, vehkaoja_dogs):<br>{dataset}_harnet_nogyr<br>{dataset}_harnet_random_nogyr<br>{dataset}_harnet_unfrozen_nogyr<br>{dataset}_RNN_nogyr<br>{dataset}_CRNN_nogyr<br>{dataset}_rf_nogyr<br><br>Otherwise:<br>{dataset}_harnet_nogyr<br>{dataset}_harnet_unfrozen_nogyr<br>{dataset}_harnet_random_nogyr<br>{dataset}_RNN_nogyr<br>{dataset}_CRNN<br>{dataset}_rf</p> <p><em>Results used in Fig. 5E of <a href="https://arxiv.org/abs/2305.10740">arxiv paper</a> (reduced data setting)<br></em>If dataset contains gyroscope (HAR, jeantet_turtles, vehkaoja_dogs):<br>{dataset}_harnet_low_data_nogyr<br>{dataset}_harnet_random_low_data_nogyr<br>{dataset}_harnet_unfrozen_low_data_nogyr<br>{dataset}_RNN_low_data_nogyr<br>{dataset}_wavelet_RNN_low_data_nogyr<br>{dataset}_CRNN_low_data_nogyr<br>{dataset}_rf_low_data_nogyr</p> <p>Otherwise:<br>{dataset}_harnet_low_data_nogyr<br>{dataset}_harnet_random_low_data_nogyr<br>{dataset}_harnet_unfrozen_low_data_nogyr<br>{dataset}_RNN_low_data_nogyr<br>{dataset}_wavelet_RNN_low_data_nogyr<br>{dataset}_CRNN_low_data<br>{dataset}_rf_low_data<br><br></p> <p><strong>CSV files</strong>: we also include summaries of the experimental results in experiments_summary.csv, experiments_by_fold_individual.csv, experiments_by_fold_behavior.csv. </p> <p><em>experiments_summary.csv - results averaged over individuals and behavior classes<br></em>dataset (str): name of dataset<br>experiment (str): name of model with experiment setting <br>fig4 (bool): True if dataset+experiment was used in figure 4 of <a href="https://arxiv.org/abs/2305.10740">arxiv paper</a><br>fig5d (bool): True if dataset+experiment was used in figure 5d of <a href="https://arxiv.org/abs/2305.10740">arxiv paper</a><br>fig5e (bool): True if dataset+experiment was used in figure 5e of <a href="https://arxiv.org/abs/2305.10740">arxiv paper</a><br>f1_mean (float): mean of macro-averaged F1 score, averaged over individuals in test folds<br>f1_std (float): standard deviation of macro-averaged F1 score, computed over individuals in test folds<br>prec_mean, prec_std (float): analogous for precision<br>rec_mean, rec_std (float): analogous for recall<em><br><br>experiments_by_fold_individual.csv - results per individual in the test folds<br></em>dataset (str): name of dataset<br>experiment (str): name of model with experiment setting <br>fig4 (bool): True if dataset+experiment was used in figure 4 of <a href="https://arxiv.org/abs/2305.10740">arxiv paper</a><br>fig5d (bool): True if dataset+experiment was used in figure 5d of <a href="https://arxiv.org/abs/2305.10740">arxiv paper</a><br>fig5e (bool): True if dataset+experiment was used in figure 5e of <a href="https://arxiv.org/abs/2305.10740">arxiv paper</a><br>fold (int): test fold index<br>individual (int): individuals are numbered zero-indexed, starting from fold 1<br>f1 (float): macro-averaged f1 score for this individual<br>precision (float): macro-averaged precision for this individual<br>recall (float): macro-averaged recall for this individual<em><br></em></p> <p><em>experiments_by_fold_behavior.csv - results per behavior class, for each test fold<br></em>dataset (str): name of dataset<br>experiment (str): name of model with experiment setting <br>fig4 (bool): True if dataset+experiment was used in figure 4 of <a href="https://arxiv.org/abs/2305.10740">arxiv paper</a><br>fig5d (bool): True if dataset+experiment was used in figure 5d of <a href="https://arxiv.org/abs/2305.10740">arxiv paper</a><br>fig5e (bool): True if dataset+experiment was used in figure 5e of <a href="https://arxiv.org/abs/2305.10740">arxiv paper</a><br>fold (int): test fold index<br>behavior_class (str): name of behavior class<br>f1 (float): f1 score for this behavior, averaged over individuals in the test fold<br>precision (float): precision for this behavior, averaged over individuals in the test fold<br>recall (float): recall for this behavior, averaged over individuals in the test fold<br>train_ground_truth_label_counts (int): number of timepoints labeled with this behavior class, in the training set<em><br></em></p>
Gene tagging and gene deletion resources for Leishmania mexicana MNYC/BZ/62/M379 Cas9/T7 strain
<p><em>primers_barcodes.csv</em>: List of primer sequences necessary for N and C terminus gene tagging, as well as gene deletion in the Leishmania mexicana MNYC/BZ/62/M379 Cas9/T7 strain. Each row contains the gene name and the DF (downstream forward), DR (downstream reverse), DSG (downstream guide sRNA), UF (upstream forward), UFB (upstream forward including a gene-unique 17nt barcode sequence), UR (upstream reverse), USG (upstream guide sRNA), VF (verification forward) and VR (verification reverse) primer sequences. Empty cells indicate that it was not possible to design this primer for this gene. Guide sRNA perfect match and off-target counts are included as well. The primer sequences were designed using LeishGEdit (http://www.leishgedit.net). Barcode sequences and assigned IDs for the unique identification of knock-out or tagged strains are included as separate columns. For recommended methods for endogenous tagging or gene deletion see Beneke <em>et al., </em>R. Soc. Open Sci.4170095 (2017), for generating barcoded deletion mutants see Beneke and Gluenz, Mol. Biochem. Parasitol. 239 (2020).</p> <p><em>genome.gff</em>: Annotated genome of the <em>L. mexicana</em> MNYC/BZ/62/M379 strain, genetically modified to express T7 RNA polymerase and Cas9. The annotation is provided in a combined GFF3 / FASTA format that also includes the sequences of the chromosomes and small contigs. The annotation also specifies polyadenylation sites (PAS features) and splice leader acceptor sites (SLAS features) which were used to refine the boundaries of protein-coding sequences as well as 3' and 5' untranslated regions over the reference genome of <em>L. mexicana</em> MNYC/BZ/62/M379<em>.</em></p> <p><em>c9t7_sequences.fasta</em>: Raw chromosome and contig sequences in FASTA format.</p> <p><em>c9t7_transcripts.fasta</em>: mRNA transcript sequences in FASTA format (includes 5' and 3' UTRs).</p> <p><em>c9t7_transcript_CDSs.fasta</em>: Coding sequences in FASTA format.</p> <p><em>c9t7_predicted_protein_sequences.fasta</em>: Predicted protein amino acid sequences in FASTA format.</p> <p>Note: This version provides an update for <em>c9t7_transcript_CDSs.fasta, c9t7_predicted_protein_sequences.fasta</em> and <em>genome.gff</em>, correcting an off-by-one sequence coordinate in 48 of the the protein-coding genes.</p>
Resources for The Fundamental Limit of Jet Tagging
<p>Resources related to The Fundamental Limit of Jet Tagging (arxiv:2411.02628)</p> <p>Includes:</p> <p>-In total 11,200,0000 qcd and 11,200,0000 top jets generated from corresponding transformer-based models (Tmodels) trained using the JetClass Dataset.</p> <p>-The trained Tmodels.</p> <p>-The LLR predictions from the Tmodels, for different number of constituents.</p> <p>-The predictions corresponding to the classifier-based jet taggers.</p>
Data for Mellado et al. The impacts of marking on bats: mark-recapture models for assessing injury rates and tag loss. Journal of Mammalogy. 103:100-110. DOI:10.1093/jmammal/gyab153
<p>Data sets used in Mellado et al. The impacts of marking on bats: mark-recapture models for assessing injury rates and tag loss. Journal of Mammalogy. 103:100-110. (https://doi.org/10.1093/jmammal/gyab153)</p> <p>File Descriptions:</p> <p>CapHistTagLoss.txt - Capture histories for <em>Carollia perspicillata</em> identifying if individual was captured with both tags (B), arm bands (A), collar (C), not captured (0) or not monitored (dot). Covariates included are Sex, Forearm Length and Scaled Mass Index.<br> CaptHistTagInj.txt - Capture histories for <em>Carollia perspicillata</em> identifying if individual was captured with no lesions from arm band (A), minor injury (I), major injury (M), not captured (0) or not monitored (dot). Covariates included are Sex, Forearm Length and Scaled Mass Index.<br> LesionOccurrence.txt - Censored time-to-event data for survival analysis. Recorded events were the occurrence of lesions of any type due to arm bands.<br> RingCondition.txt - Censored time-to-event data for survival analysis. Recorded events were the occurrence of damage to arm bands.<br> SMI.txt - Longitudinal data for individual <em>Carollia perspicillata</em> Scaled Mass Index, identifying individual records, the occurrence of lesions, sex, month, year</p> <p> </p>
Dataset for "A 21 m Operation Range RFID Tag for "Pick to Light" Applications with a Photovoltaic Harvester"
<p>In the paper, a novel Radio-Frequency Identification (RFID) tag for “pick to light” applications is presented. The proposed tag architecture shows the implementation of a novel voltage limiter and a supply voltage (VDD) monitoring circuit to guarantee a correct operation between the tag and the reader for the “pick to light” application. The feasibility to power the tag with different photovoltaic cells is also analyzed, showing the influence of the illuminance level (lx), type of source light (fluorescent, LED or halogen) and type of photovoltaic cell (photodiode or solar cell) on the amount of harvested energy. Measurements show that the photodiodes present a power per unit package area for low illuminance levels (500 lx) of around 0.08 μW/mm<sup>2</sup>, which is slightly higher than the measured one for a solar cell of 0.06 μW/mm<sup>2</sup>. However, solar cells present a more compact design for the same absolute harvested power due to the large number of required photodiodes in parallel. Finally, an RFID tag prototype for “pick to light” applications is implemented, showing an operation range of 3.7 m in fully passive mode. This operation range can be significantly increased to 21 m when the tag is powered by a solar cell with an illuminance level as low as 100 lx and a halogen bulb as source light.</p> <p>This dataset contains some of the data gathered during the experimental work developed and used in the paper.</p>
USM Dataset - A Dataset for Polyphonic Sound Event Tagging in Urban Sound Monitoring Scenarios
<p>This dataset includes 24,000 5-seconds-long polyphonic stereo soundscapes composed of sounds taken from the FSD50k dataset:</p> <p>- Eduardo Fonseca, Xavier Favory, Jordi Pons, Frederic Font, Xavier Serra. FSD50K: an Open Dataset of Human-Labeled Sound Events (<a href="https://arxiv.org/abs/2010.00475">https://arxiv.org/abs/2010.00475</a>)</p> <p>FSD50k samples used in the USM dataset were selected to allow for commercial usage.</p> <p>Find more details about the USM dataset at <a href="https://github.com/jakobabesser/USM">https://github.com/jakobabesser/USM</a></p>
Pop-up satellite archival tagging data of Atlantic bluefin tuna in the Gulf of Lions, Northwestern Mediterranean Sea
<p>24 Atlantic bluefin tuna (Thunnus thynnus) individuals (117–158 cm fork length) were tagged with pop-up archival tags in the Gulf of Lion, NW-Mediterranean Sea between 2015 and 2016.</p> <p><strong>Tag programming and data</strong></p> <p>The tags applied (miniPATs by Wildlife Computers, https://wildlifecomputers.com) can record depth and temperature time series (denoted hereafter as DepthTS and TempTS, respectively) at a temporal resolution of 3–5 s (depending on the predefined deployment duration) and a vertical resolution of 0.5 m. Based on these data, the tag calculates and stores additional data products such as PAT-style Depth–Temperature profiles (PDT), time at depth data, and time at temperature. After pop-up, the tags transmit user-defined data products and subsets from the recorded data sets. All our tags were configured to transmit the following data products: daily light curves, DepthTS, and PDT. In order to maximize data coverage of the transmitted datasets, we decreased the temporal resolution of the DepthTS and PDT data after the first tagging campaign in 2015 from 150 to 600 s and 6 to 24 h, respectively. For both years, deployment durations were set to 150 and 90 d during spring (April–May) and summer (August–September), respectively. A description of the electronic tagging procedure can be found in <a href="https://doi.org/10.1093/icesjms/fsaa083">Bauer et al. (2020)</a>.</p> <p>Seven tags were physically recovered, providing the complete archived time series data at a resolution of 3–5 s. Nineteen tags provided more than 7 d of complete DepthTS data (i.e. without transmission gaps). Three tags from 2016 had deployment durations of <1 week (#15P0983, #15P0985, and #15P0986) because of hardware failure.</p> <p>Provided files contain raw tag data (transmitted and recovered datasets) from the Wildlife Computers Data Portal as well as related GPE3 model runs (geolocation estimates).</p> <p>We thank the crews of the Cyngali and Roussillon Fishing recreational fishing vessels for their cooperation during the tagging cruises. This tagging study was part of the BLUEMED project and funded by the French National Research Agency (ANR; Project-ID ANR-14-ACHN-0002).</p> <p> </p>
SETDB1 knock-out and reexpression of a WT, a catalytic-dead variant (CA) or a NLS-tagged SETDB1 protein
<p>In order to generate the Setdb1 cKO allele, exons 15 and 16, which encode the core amino acids of the catalytic domain, have been flanked with two lox-P sites recognized by Cre recombinase enzyme. Cre-oestrogen receptor fusion gene Mer-Cre-Mer has been introduced to induce acute Setdb1 KO after Tamoxifen treatment. Setdb1 cKO mESCs where Setdb1 expression is stably rescued by wild type 3xFlag-Setdb1 (WT) or by the catalytic dead mutant 3xFlag-Setdb1 (CA) were generously given by Pr Yoichi Shinkai. The catalytic dead mutant was obtained with a single lysine mutation. In the lab, Setdb1 cKO mESCs, where Setdb1 expression is stably rescued by NLS-3xFlag-Setdb1 which localizes only in the nucleus, have been established (3 NLS sequences have been added in order to retain Setdb1 in the nucleus).</p> <p>The quality of the RNA was determined on the Agilent 2100 Bioanalyzer (Agilent Technologies, Palo Alto, CA, USA), the RNA integrity number was above 8 for all the samples. To construct libraries, 1 mg of high-quality total RNA sample was processed using Truseq® stranded total RNA kit (Illumina®). After the removal of ribosomal RNAs (using Ribo-zero® rRNA), confirmed by QC control on pico chipTM on the Agilent 2100 Bioanalyzer (Agilent Technologies, Palo Alto, CA, USA), total RNA molecules are fragmented and reverse-transcribed using random primers. Replacement of dTTP by dUTP during the second strand synthesis will permit to achieve the strand specificity. Addition of a single A base to the cDNA is followed by ligation of adapters. Libraries were quantified by qPCR using the KAPA Library Quantification Kit for Illumina Libraries (KapaBiosystems) and library profiles were assessed using the DNA High SensitivityTMHS kit on an Agilent Bioanalyzer 2100. Libraries were sequenced on an Illumina® Nextseq 500 instrument using 75 base-lengths read V2 chemistry in a paired-end mode.</p>
Top Quark Tagging Reference Dataset
<p>A set of MC simulated training/testing events for the evaluation of top quark tagging architectures.</p> <p>In total 1.2M training events, 400k validation events and 400k test events. Use “train” for training, “val” for validation during the training and “test” for final testing and reporting results.</p> <p><strong>Description</strong></p> <ul> <li> <p>14 TeV, hadronic tops for signal, qcd diets background, Delphes ATLAS detector card with Pythia8</p> </li> <li> <p>No MPI/pile-up included</p> </li> <li> <p>Clustering of particle-flow entries (produced by Delphes E-flow) into anti-kT 0.8 jets in the pT range [550,650] GeV</p> </li> <li> <p>All top jets are matched to a parton-level top within ∆R = 0.8, and to all top decay partons within 0.8</p> </li> <li> <p>Jets are required to have |eta| < 2</p> </li> <li> <p>The leading 200 jet constituent four-momenta are stored, with zero-padding for jets with fewer than 200</p> </li> <li> <p>Constituents are sorted by pT, with the highest pT one first</p> </li> <li> <p>The truth top four-momentum is stored as truth_px etc.</p> </li> <li> <p>A flag (1 for top, 0 for QCD) is kept for each jet. It is called is_signal_new</p> </li> <li> <p>The variable "ttv" (= test/train/validation) is kept for each jet. It indicates to which dataset the jet belongs. It is redundant as the different sets are already distributed as different files.</p> </li> </ul>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.