Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
7,523
datasets available to search
ShareScore release 0.7.1
Dataset results
7,523 results for “Annotation”
An annotated high-content fluorescence microscopy dataset with Hoechst 33342-stained nuclei and manually labelled outlines
<p>Here we present a benchmarking dataset of fluorescence microscopy images with Hoechst 33342-stained nuclei together with annotations of nuclei, nuclear fragments and micronuclei. Images were randomly selected from an RNA interference screen with a modified U2OS osteosarcoma cell line, acquired on a Thermo Fischer CX7 high-content imaging system at 20x magnification. Labelling was performed by a single annotator and reviewed by a biomedical expert.</p> <p>The dataset contains 50 images showing over 2000 labelled nuclear objects in total, which is sufficiently large to train well-performing neural networks for instance or semantic segmentation. It is pre-split into training, development and test set, each in a zip file. The dataset should be referred to as Aitslab_bioimaging1. A brief article describing the dataset is also available (Arvidsson M, Kazemi Rashed S, Aits S. <a href="https://doi.org/10.1016/j.dib.2022.108769">10.1016/j.dib.2022.108769</a> )</p> <p><strong>Dataset description:</strong></p> <p>Fluorescence microscopy images: original .C01 files and files converted to 8-bit .png format (Grayscale)</p> <p>Annotations: 24-bit .png format (RGB)</p> <p>Script used to convert C01 to png images: C01_to_png.py file with python code and readme.md file with instructions to run it</p>
3D Nuclei annotations and StarDist 3D model(s) (rat brain)
<p><strong>Name</strong>: 3D Nuclei annotations and StarDist3D model(s) (rat brain)</p> <p><strong><em>Images: </em></strong>From a large tiling acquisition ( https://doi.org/10.5281/zenodo.6646128 ) individual Tile (xyz : 1024x1024x62) were downsampled and cropped (128x128x62). Four crops, from different tiles (./annotations_BIOP/images/) were manually annotated with ITK-SNAP (./annotations_BIOP/masks/)</p> <p>These four images, and their corresponding masks, were cropped into four quadrants (./crops_BIOP_v1/) in order to get 16 different images (64x64x62).</p> <p><strong><em>Conda environment</em></strong><em>: </em>A conda environment was created using the yml file <em>stardist0.8_TF1.15.yml</em></p> <p><strong><em>Training : </em></strong>Training was performed using the jupyter notebook <em>1-Training_notebook.ipynb</em>.<br> Three different trainings (with the same random seed, same anisotropy, patch size and grid) were performed and produced three different models (./models/)</p> <p>Validation images (from the random seed used) were exported to ease the visual inspection of the results(./val_rdm42/).</p> <p><strong><em>Validation: </em></strong>To save metrics in a csv file and compare predictions to the annotations the jupyter notebook <em>2-QC_notebook.ipynb </em>can be used on the validation folder.</p> <p><strong>Large images</strong>: To test the model on larger images one can use Whole_ds441.tif (or Crop_ds441.tif )<br> These images were obtained using the plugin <a href="https://imagej.net/plugins/bigstitcher/">BigSticher </a>on the raw data ( https://doi.org/10.5281/zenodo.6646128 ), resaved as h5 and exported the downsample by 4 version.</p> <p> </p> <p> </p>
CafeteriaSA corpus: Scientific abstracts annotated across different food semantic resources
<p>In the last decades, a great amount of work has been done in predictive modeling of issues related to human and environmental health. Resolution of issues related to healthcare is made possible by the existence of several biomedical vocabularies and standards, which play a crucial role in understanding health information, together with a large amount of health-related data. However, despite the large number of available resources and work done in the health and environmental domains, there is a lack of semantic resources that can be utilized in the food and nutrition domain, as well as their interconnections. For this purpose, in an European Food Safety Authority-funded project CAFETERIA, we have developed the first annotated corpus of 500 scientific abstracts that consists of 6,407 annotated food entities with regard to Hansard taxonomy, 4,299 for FoodOn, and 3,623 for SNOMED-CT. The CafeteriaSA corpus will enable further development of natural language processing methods for food information extraction from textual data that will allow extracting of food information from scientific textual data.</p>
Annotating Cognates in Phylogenetic Studies of South-East Asian Languages [Supplement]
<p>Source code and data accompanying the study "<strong>Annotating Cognates in Phylogenetic Studies of South-East Asian Languages" by M.-S. Wu and J.-M. List.</strong></p>
HISTORIAN: a large-scale HISTORIcal film dataset with cinematographic ANnotation
<p>Developing automated tools for sustainable film preservation of extensive historical film collections assumes an understanding of fundamental cinematographic settings. In order to be able to investigate new approaches to detect and classify cinematographic settings, this paper proposes a novel large-scale historical film dataset with cinematographic annotations (HISTORIAN), i.e., shot boundaries, shot types, camera movements. The dataset consists of 98 digitized original analog film reels related to the Second World War and 10593 film shots manually annotated by human film experts. Moreover, annotations for overscan areas such as sprocket holes are included. A baseline film analysis pipeline is introduced and evaluated. To the best of our knowledge, HISTORIAN is the first dataset that covers the challenges and characteristics of historical film documentaries and provides novel possibilities for exploring automatic film analysis tools.</p> <p>This repository presents a tiny set including a few examples for demonstration.</p> <p>A link to the Github repository (including helper scripts and readme) can be found <a href="https://github.com/dahe-cvl/historian_dataset">here</a></p> <p> </p> <p> </p> <p> </p> <p> </p>
CafeteriaFCD corpus: Food consumption data annotated with regard to different food semantic resources
<p>The FoodBase curated version which contains 1,000 manually evaluated recipes, annotated with the appropriate semantic tags from the Hansard Taxonomy, FoodON and SNOMED-CT.</p>
Mataws annotated Web service collection
<p><strong>Description. </strong>The Mataws annotated Web service collection is a set of WS descriptions under the WSDL and OWL-S formats. It contains 816 descriptions, which were originally only syntactically described, and were annotated using our tool Mataws. Consequently, each description appears twice (once in a syntactical version, and once in a semantic version). The descriptions are also classified thematically.</p> <p>Our collection is based primarily on the FullDataset collection of the Assam project (<a href="http://www.andreas-hess.info/projects/annotator/">http://www.andreas-hess.info/projects/annotator/</a>), which we extended using WS descriptions found on the web. These individual files were classified thematically with the rest of the WSDL files, and used to assess the quality of annotation of Mataws.</p> <p><strong>Source code. </strong>The source code of our tool Mataws is available online: <a href="https://github.com/CompNet/mataws">https://github.com/CompNet/mataws</a></p> <p><strong>License. </strong>The annotated descriptions are shared under a Creative Commons 0 license. The original descriptions belong to their authors.</p> <p><strong>Citation. </strong>If you use our dataset, please cite the following article:</p> <ul> <li>Aksoy, C., Labatut, V., Cherifi, C. & Santucci, J.-F (2011). MATAWS: A Multimodal Approach for Automatic WS Semantic Annotation. In International Conference on Networked Digital Technologies. Macau, CN : Springer. ⟨<a href="https://hal.archives-ouvertes.fr/hal-00620566">hal-00620566</a>⟩ - DOI: <a href="https://doi.org/10.1007/978-3-642-22185-9_27">10.1007/978-3-642-22185-9_27</a></li> </ul> <p><br><code>@InProceedings{Aksoy2011,</code><br><code> author = {Aksoy, Cihan and Labatut, Vincent and Cherifi, Chantal and Santucci, Jean-François},</code><br><code> title = {{MATAWS}: A Multimodal Approach for Automatic WS Semantic Annotation},</code><br><code> booktitle = {3\textsuperscript{rd} International Conference on Networked Digital Technologies},</code><br><code> year = {2011},</code><br><code> volume = {136},</code><br><code> series = {Communications in Computer and Information Science},</code><br><code> pages = {319-333},</code><br><code> address = {Macau, CN},</code><br><code> publisher = {Springer},</code><br><code> doi = {10.1007/978-3-642-22185-9_27},</code><br><code>}</code></p>
Gene annotation files for Fraxinus excelsior (European ash) genome assembly BATG-0.5
<p>Gene annotation files for <em>Fraxinus excelsior</em> genome assembly v. BATG-0.5, published in Nature (doi:10.1038/nature20786). These files were previously hosted on the Ash Tree Genomes website (http://www.ashgenome.org/transcriptomes) and first made available for download via that site on 2016-02-08.</p> <p>The following annotation files are available:</p> <p>### GFF file of all gene models (all isoforms)<br> Fraxinus_excelsior_38873_TGAC_v2.gff3</p> <p>### FASTA file of all cDNA sequences (all isoforms)<br> Fraxinus_excelsior_38873_TGAC_v2.gff3.cdna.fa</p> <p>### FASTA file of all CDS DNA sequences (all isoforms)<br> Fraxinus_excelsior_38873_TGAC_v2.gff3.cds.fa</p> <p>### FASTA file of all peptide sequences (all isoforms)<br> Fraxinus_excelsior_38873_TGAC_v2.gff3.pep.fa</p> <p>### Functional annotation for each gene model (all isoforms)<br> Fraxinus_excelsior_38873_TGAC_v2.gff3.functional_annotation.tsv</p> <p>### GFF file of all gene models (longest isoform only)<br> Fraxinus_excelsior_38873_TGAC_v2.longestCDStranscript.gff3</p> <p>### FASTA file of all cDNA sequences (longest isoform only)<br> Fraxinus_excelsior_38873_TGAC_v2.longestCDStranscript.gff3.cdna.fa</p> <p>### FASTA file of all CDS DNA sequences (longest isoform only)<br> Fraxinus_excelsior_38873_TGAC_v2.longestCDStranscript.gff3.cds.fa</p> <p>### FASTA file of all peptide sequences (longest isoform only)<br> Fraxinus_excelsior_38873_TGAC_v2.longestCDStranscript.gff3.pep.fa</p> <p>### GFF file for gene models identified as probable transposable element related sequences (excluded from the other files)<br> Fraxinus_excelsior_38873_TGAC_v2.transposable_elements.gff3</p> <p><br> NB: The annotation files include preliminary annotations for genes within the organellar scaffolds (gene models FRAEX38873_v2_000400370-FRAEX38873_v2_000401330), which were not reported in the publication of the BATG0.5 assembly (doi:10.1038/nature20786).</p>
Annotation and Orthofinder results for three Mytilus species genomes.
<p>Annotations and Orthofinder results for three Mytilus species genomes: accessions JAKGDF000000000 (MgalMED), JAKGDG000000000 (MeduEUS), and JAKGDH000000000 (MeduEUN).</p>
A collection of fully-annotated soundscape recordings from the Western United States
<p>This collection contains 33 hour-long soundscape recordings, which have been annotated with 20,147 bounding box labels for 56 different bird species from the Western United States. The data were recorded in 2018 in the Sierra Nevada, California, USA. This collection has partially been featured as test data in the 2021 BirdCLEF competition and can primarily be used for training and evaluation of machine learning algorithms.</p> <p><strong>Data collection</strong></p> <p>Measuring the effects of forest management activities in the Sierra Nevada, California, USA can reveal a potential correlation with avian population density and diversity. For this dataset, passive acoustic surveys were conducted in the Lassen and Plumas National Forests in May-August 2018. Survey grid cells (4 km<sup>2</sup>) were randomly selected from a 6,000-km<sup>2</sup> area, and SWIFT recording units were deployed at locations conducive to sound propagation (e.g., ridges rather than gullies) within those cells. The sensitivity of the used microphones was -44 (+/-3) dB re 1 V/Pa. The microphone's frequency response was not measured, but is assumed to be flat (+/- 2 dB) in the frequency range 100 Hz to 7.5 kHz. The analog signal was amplified by 38 dB and digitized (16-bit resolution) using an analog-to-digital converter (ADC) with a clipping level of -/+ 0.9 V. Recording units recorded continuously 17:00 - 23:59, 0:00 - 10:00, one-hour files were stored as uncompressed WAVE sampled at 32 kHz and later converted to FLAC. Parts of this dataset have previously been used in the 2021 BirdCLEF competition.</p> <p><strong>Sampling and annotation protocol</strong></p> <p>We subsampled data for this collection by selecting locations that spanned the full elevational and latitudinal gradients of our study area (~840 – 1700 m asl and 39.41 – 40.71°N), and thus represent a broad range of plant communities. A single annotator boxed every bird call he could recognize, ignoring those that are too faint or unidentifiable. Raven Pro software was used to annotate the data. Provided labels contain full bird calls that are boxed in time and frequency. The annotator was allowed to combine multiple consecutive calls of one species into one bounding box label if pauses between calls were shorter than five seconds. We use eBird species codes as labels, following the 2021 eBird taxonomy (Clements list).</p> <p><strong>Files in this collection</strong></p> <p>Audio recordings can be accessed by downloading and extracting the “soundscape_data.zip” file. Soundscape recording filenames contain a sequential file ID, recording date and timestamp in PDT. As an example, the file “SNE_001_20180509_050002.flac” has sequential ID 001 and was recorded on May 9th 2018 at 05:00:02 PDT. Ground truth annotations are listed in “annotations.csv” where each line specifies the corresponding filename, start and end time in seconds, low and high frequency in Hertz and an eBird species code. These species codes can be assigned to scientific and common name of a species with the “species.csv” file. The approximate recording location with longitude and latitude can be found in the “recording_location.txt” file.</p> <p><strong>Acknowledgements </strong></p> <p>The collection and annotation of this dataset was funded by the U.S. Forest Service Region 5 and the California Department of Fish and Wildlife.</p>
A collection of fully-annotated soundscape recordings from the Island of Hawai'i
<p>This collection contains 635 soundscape recordings with a total duration of almost 51 hours, which have been annotated by expert ornithologists who provided 59,583 bounding box labels for 27 different bird species from the Hawaiian Islands, including 6 threatened or endangered native birds. The data were recorded between 2016 and 2022 at four sites across Hawai‘i Island. This collection has partially been featured as test data in the 2022 BirdCLEF competition and can primarily be used for training and evaluation of machine learning algorithms.</p> <p><strong>Data collection</strong></p> <p>Soundscapes for this collection were recorded for various research projects by the Listening Observatory for Hawaiian Ecosystems (LOHE) at the University of Hawai‘i at Hilo. The recordings were collected using Wildlife Acoustics Inc. Song Meters (models 2, 4, or Mini), as 16-bit wav files at a sampling rate of 44.1 kHz, using the default gain settings of each model. Further specifics for each recording, such as recording location and habitat type, can be found in the metadata provided. Soundscapes in this collection vary in length, ranging from just under a minute to 9 minutes in duration. All audio was unified, converted to FLAC, and resampled to 32 kHz for this collection. Parts of this dataset have previously been used in the 2022 BirdCLEF competition.</p> <p><strong>Sampling and annotation protocol</strong></p> <p>This collection is a subset of the files recorded over the course of the LOHE lab’s respective studies. The data were subsampled for annotation by aurally scanning the recordings and visually scanning spectrograms generated using Raven Pro software for target species of interest to the individual research project for which each recording was collected. Recordings that did not contain vocalizations of the species of interest were excluded from full annotation and thus this collection. </p> <p>Using Raven Pro, annotators were asked to create a selection box around every bird call they could recognize, ignoring those that were too faint or unidentifiable at a spectrogram window size of 700 points. Provided labels contain full bird calls that are boxed in time and frequency. Annotators were allowed to combine multiple consecutive calls of the same species into one bounding box label if pauses between calls were shorter than 0.5 seconds. We converted labels to eBird species codes, following the 2021 eBird taxonomy (Clements list).</p> <p><strong>Files in this collection</strong></p> <p>Audio recordings can be accessed by downloading and extracting the “soundscape_data.zip” file. Soundscape recording filenames contain a sequential file ID, site ID, recording date, and timestamp in HST. As an example, the file “UHH_001_S01_20161121_150000.flac” has sequential ID 001 and was recorded at site S01 on Nov 21st, 2016 at 15:00:00 HST. Ground truth annotations are listed in “annotations.csv” where each line specifies the corresponding filename, start and end time in seconds, low and high frequency in Hertz, and an eBird species code. These species codes can be assigned to the scientific and common name of a species with the “species.csv” file. The approximate recording location with Universal Transverse Mercator (UTM) coordinates and other metadata can be found in the “recording_location.csv” file.</p> <p><strong>Acknowledgements </strong></p> <p>Compiling this extensive dataset was a major undertaking, and we are very thankful to the domain experts who helped to collect and manually annotate the data for this collection. Specifically, we want to thank Charlotte Forbes-Perry with the Pacific Cooperative Studies Unit, University of Hawai'i at Hawai‘i Volcanoes National Park as well as the following current and past members of the LOHE lab (in alphabetical order): Keith Burnett, Saxony Charlot, Noah Hunt, Caleb Kow, Elizabeth Lough, and Bret Mossman.</p> <p>Access and permits to record soundscapes were provided by (in alphabetical order): Hakalau Forest National Wildlife Refuge, the State of Hawai‘i Department of Land and Natural Resources Division of Forestry and Wildlife, and the U.S. Fish and Wildlife Service.</p> <p>We would also like to acknowledge our funding sources (in alphabetical order): The National Park Service Inventory and Monitoring Division, the National Science Foundation, and the U.S. Army Engineer Research and Development Center.</p>
A collection of fully-annotated soundscape recordings from the Northeastern United States
<p>This collection contains 285 hour-long soundscape recordings, which have been annotated by expert ornithologists who provided 50,760 bounding box labels for 81 different bird species from the Northeastern USA. The data were recorded in 2017 in the Sapsucker Woods bird sanctuary in Ithaca, NY, USA. This collection has (partially) been featured as test data in the 2019, 2020 and 2021 BirdCLEF competition and can primarily be used for training and evaluation of machine learning algorithms.</p> <p><strong>Data collection</strong></p> <p>As part of the Sapsucker Woods Acoustic Monitoring Project (SWAMP), the K. Lisa Yang Center for Conservation Bioacoustics at the Cornell Lab of Ornithology deployed 30 first-generation SWIFT recorders in the surrounding bird sanctuary area in Ithaca, NY, USA. The sensitivity of the used microphones was -44 (+/-3) dB re 1 V/Pa. The microphone's frequency response was not measured, but is assumed to be flat (+/- 2 dB) in the frequency range 100 Hz to 7.5 kHz. The analog signal was amplified by 33 dB and digitized (16-bit resolution) using an analog-to-digital converter (ADC) with a clipping level of -/+ 0.9 V. This ongoing study aims to investigate the vocal activity patterns and seasonally changing diversity of local bird species. The data are also used to assess the impact of noise pollution on the behavior of birds. Recordings were recorded 24 h/day in 1-hour uncompressed WAVE files at 48 kHz, converted to FLAC and resampled to 32 kHz for this collection. Parts of this dataset have previously been used in the 2019, 2020 and 2021 BirdCLEF competition.</p> <p><strong>Sampling and annotation protocol</strong></p> <p>We subsampled data for this collection by randomly selecting one 1-hour file from one of the 30 different recording units for each hour of one day per week between Feb and Aug 2017. For this collection, we excluded recordings that were shorter than one hour or did not contain a bird vocalization. Annotators were asked to box every bird call they could recognize, ignoring those that are too faint. Raven Pro software was used to annotate the data. Provided labels contain full bird calls that are boxed in time and frequency. Annotators were allowed to combine multiple consecutive calls of one species into one bounding box label if pauses between calls were shorter than five seconds. We use eBird species codes as labels, following the 2021 eBird taxonomy (Clements list).</p> <p><strong>Files in this collection</strong></p> <p>Audio recordings can be accessed by downloading and extracting the “soundscape_data.zip” file. Soundscape recording filenames contain a sequential file ID, recording date, and timestamp in UTC. As an example, the file “SSW_001_20170225_010000Z.flac” has sequential ID 001 and was recorded on Feb 25th, 2017 at 01:00:00 UTC. Ground truth annotations are listed in “annotations.csv” where each line specifies the corresponding filename, start and end time in seconds, low and high frequency in Hertz, and an eBird species code. These species codes can be assigned to scientific and common name of a species with the “species.csv” file. Unidentifiable calls have been marked with “????” and are included in the ground truth annotations. The approximate recording location with longitude and latitude can be found in the “recording_location.txt” file.</p> <p><strong>Acknowledgements </strong></p> <p>Compiling this extensive dataset was a major undertaking, and we are very thankful to the domain experts who helped to collect and manually annotate the data for this collection (individual contributors in alphabetic order): Jessie Barry, Sarah Dzielski, Cullen Hanks, W. Alexander Hopping, Robert Koch, Jim Lowe, Jay McGowan, Ashik Rahaman, Yu Shiu, Laurel Symes, and Matt Young. </p> <p><strong>Version history</strong></p> <p>Version 2: Unidentifiable calls have been marked with “????” and added as bounding box labels to the ground truth annotations.<br> Version 1: Initial release.</p>
D3 annotation with CSO Classifier
<p>The <a href="https://zenodo.org/record/7069915">DBLP Discovery Dataset </a>(D3) is a newly created dataset of research papers in the field of Computer Science which can support several tasks like identifying trends in research activity, productivity, focus, bias, accessibility, and impact. This dataset stems from DBLP and integrates additional information from the full-texts. We argue that papers classified with their research topics can improve the identification of research trends. To this end, we used the <a href="https://github.com/angelosalatino/cso-classifier">CSO Classifier</a> to annotate all the papers within D3 and we made such extension available for research purposes.</p> <p> </p> <p>More info: <a href="https://www.salatino.org/wp/annotating-d3-dataset-with-the-cso-classifier/">https://www.salatino.org/wp/annotating-d3-dataset-with-the-cso-classifier/</a></p> <p>More info pdf: <a href="https://www.salatino.org/wp/wp-content/uploads/2022/09/Annotating-D3-dataset-with-the-CSO-Classifier.pdf">https://www.salatino.org/wp/wp-content/uploads/2022/09/Annotating-D3-dataset-with-the-CSO-Classifier.pdf</a></p>
TARA Pacific CTAX colony morphological annotations release version 1_1
<p>PHOTO dataset (<em>in situ</em> photos) and the colony morphometric analysis using <em>in situ </em>photographs of two scleractinian corals and one hydrozoan coral taken during the TARA Pacific Expedition: <em>Pocillopra </em>spp., <em>Porites </em>spp. and <em>Millepora </em>spp. respectively.</p>
Spectral Libraries for Metabolome Annotation Workflow (MAW)
<p>MassBank saved at 2022-09-12 10:28:52 with release version 2022.06 as mbankNIST.rda (MsBackendMsp)<br> GNPS saved at 2022-09-12 13:37:42 as gnps.rda (MsBackendMsp)<br> HMDB saved with the release version 4 as hmdb.rda (MsBackendHmdb)</p> <p>All .rda files can be reloaded into R session using the respective Backends. These databases were created for MAW version 1.</p> <p>hmdb_dframe_str.csv is downloaded from HMDB Downloads for structural information on HMDB IDs present in the HMDB version 4 spectral data.</p>
Joseph Haydn - String Quartets Op.20 - Harmonic Analysis Annotations Dataset
<p>This dataset accompanies the Master Thesis from the same author. It is a manually-annotated corpus of harmonic analysis in **harm syntax.</p> <p>The dataset contains the following scores:<br> Haydn, Joseph<br> 1. E-flat major, op. 20 no. 1, Hob. III-31<br> I. Allegro moderato<br> II. Menuetto. Allegretto<br> III. Affettuoso e sostenuto<br> IV. Finale. Presto<br> 2. C major, op. 20 no. 2, Hob. III-32 <br> I. Moderato<br> II. Capriccio. Adagio<br> III. Menuetto. Allegretto<br> IV. Fuga a 4 soggetti<br> 3. G minor, op. 20 no. 3, Hob. III-33<br> I. Allegro con spirito<br> II. Menuetto. Allegretto<br> III. Poco adagio<br> IV. Finale. Allegro molto<br> 4. D major, op. 20 no. 4, Hob. III-34<br> I. Allegro di molto<br> II. Un poco adagio e affettuoso<br> III. Menuet alla Zingarese & Trio<br> IV. Presto e scherzando<br> 5. F minor, op. 20 no. 5, Hob. III-35<br> I. Allegro moderato<br> II. Menuetto<br> III. Adagio<br> IV. Finale. Fuga a due soggetti<br> 6. A major, op. 20 no. 6, Hob. III-36<br> I. Allegro di molto e scherzando<br> II. Adagio. Cantabile<br> III. Menuetto. Allegretto<br> IV. Fuga a 3 soggetti. Allegro</p>
Data for common data models to streamline metabolomics processing and annotation, and implementation in a Python pipeline
<p>This upload contains the HZV029 Plasma and HZV029 Two-Phase dataset for reviewers of the "Data for common data models to streamline metabolomics processing and annotation, and implementation in a Python pipeline" submission. </p> <p>Both datasets will be uploaded to metabolomics workbench and the upload completed before final publication of the manuscript. For the he HZV029 Plasma datasets only the final run is included for any sample (i.e., failed injections or other samples with data quality issues that were reran during acquisition were omitted).</p> <p>Also included in the upload is the source code for the MetDataModel and the pcpfm at the time of manuscript re-submission and the pcpfm itself. If you find this upload in the future, please check out the github repos for more updated versions:</p> <p>https://github.com/shuzhao-li-lab/PythonCentricPipelineForMetabolomics</p> <p>https://github.com/shuzhao-li-lab/metDataModel</p> <p>The github repo does not store the input the data for space reasons, they only have the notebooks. However, the .zip here has both the notebooks by themselves in the notebook subdirectory and a separate directory with the notebooks and the data used to generate all the figures and results in the manuscript.</p> <p><strong>Some information that is needed to rerun this analysis:</strong></p> <p>Sequence files are critical to the functioning of the pipeline. The sequence files for all analyses are provided under sequence_files.zip. These can be used to recapitulate the analysis by eitehr changing the filepath to each acquisition to where you put it on your sytem or by placing the sequence file in the same directory as the mzml or raw. In the latter case, the pipeline will search for filenames matching the sample names. The sequence files also store some sample metadata such as the type of sample a given acquisition is (unknown, pooled, qc, etc...)</p> <p>.raw to .mzML conversion works well on MacOS but may not work well on other systems. You will need to use the ability to specify your own conversion command or convert files outside of the pipeline. </p> <p>To replicate the results, you do need to have the annotation sources downloaded which can be done using the pipeline. MS2 annotation requires the files in the AcquireX directory which is MS2 acquisitions on pooled HZV029 plasma samples.</p> <p>For the comparison between MetaboAnalystR and the pcpfm, subsets of the datasets were used. These subsets and the sequence files are in Subsets_for_performance_testing.zip. The sequences are also in the sequence_files directory as well</p> <p>The notebooks reference data in the analysis folders. Copies of these files are located with the notebooks to ease reproduction of the exact results in the paper; however, to do so, you will need to change paths to this data in the notebook. This lets the notebooks be ran during a rerun without copying intermediates back and forth and it keeps the github repo clean.</p> <p><strong>Version History:</strong></p> <p>This version is after reviewer comments and is for resubmission.</p> <p> </p> <p><strong>Contributions:</strong></p> <p>Joshua M Mitchell implemented the pipeline and was first author on the manuscript. Shuzhao Li is the corresponding author on the manuscript. </p> <p>Maheshwor Thapa performed the experiments to collect the HZV029 data. Yuanye Chi helped with testing and documenting the pipeline. </p> <p>Jiangou (Jeff) Xia and Zhiqiang Pang provided the R portion of the analysis. </p>
SunspotsYoloDataset: annotated solar images captured with smart telescopes (January 2023 - May 2024)
<p><strong>SunspotsYoloDataset</strong> is a set of 1690+380+128 high-resolution RGB astronomical images captured with smart telescopes with specific solar filters and annotated with the positions of sunspots that are effectively in the images. Two instruments were used for several months from Luxembourg and France between January 2023 and May 2024: a Stellina smart telescope (<a href="https://vaonis.com/stellina">https://vaonis.com/stellina</a>) and a Vespera smart telescope (<a href="https://vaonis.com/vespera">https://vaonis.com/vespera</a>).</p> <p><strong>SunspotsYoloDataset</strong> can be used to train YOLO detection models on solar images, enabling the prediction of unexpected events such as Borealis Aurora with astronomical equipment accessible to the public.</p> <p><strong>SunspotsYoloDataset</strong> is formatted with the YOLO standard, i.e., with separated files for images and annotations, usable by state-of-the-art training tools and graphical software like MakeSense (<a href="https://www.makesense.ai">https://www.makesense.ai</a>). More precisely, there is a ZIP file containing RGB images in JPEG format (minimal compression), and text files containing the positions of sunspots. Each RGB image has a resolution of 640 × 640 pixels.</p> <p>For more details about the dataset, please contact the author: olivier.parisot@list.lu .</p> <p>For more information about Luxembourg of Science and Technology (LIST), please consult: <a href="https://www.list.lu">https://www.list.lu</a> .</p> <p> </p>
Breast Masses Dataset with Precisely Annotated Sequential Mammograms
<p><strong>BREAST MASSES DATASET WITH PRECISELY ANNOTATED SEQUENTIAL MAMMOGRAMS</strong></p> <p><strong>Citing the Dataset</strong></p> <p>The dataset is released under a Creative Commons Attribution license, so it is mandatory to cite the dataset if you use it in your work in any form. Published academic papers should use the academic paper citation of our paper. Personal works, such as projects or blog posts, should provide a URL to this Zenodo page, though a reference to our paper would also be appreciated.</p> <p><em>Academic paper citation</em></p> <p>TBA</p> <p><em>Personal use citation</em></p> <p>Include a link to this Zenodo page - 10.5281/zenodo.11446259</p> <p><strong>ACKNOWLEDGMENT</strong></p> <p>The publication of this paper is supported by the European Union’s Horizon 2020 research and innovation programme under grant agreement No 739551 (KIOS CoE) and the Government of the Republic of Cyprus through the Cyprus Deputy Ministry of Research, Innovation and Digital Policy.</p> <p><strong>Contact Information</strong></p> <p>If you would like further information about the dataset, or if you experience any issues downloading files, please contact us at cloizi01@ucy.ac.cy.</p> <p><strong>General Information</strong></p> <p>This dataset consists of 100 pairs of mammograms, from two temporally sequential rounds. Specifically, this dataset includes the prior and recent mammograms with two mammographic views for each patient. This is a complete dataset for the detection and classification of breast masses, using sequential mammograms. It contains normal (BI-RADS 1), benign (BI-RADS 2), and biopsy-confirmed malignant cases (BI-RADS 6). For each mammogram, an image with precise annotation of each individual mass, by two expert radiologists, is provided.</p> <p><strong>More details are available in the README.txt</strong></p>
Annotated Data in Spanish for Toxicity and Insults in Digital Social Networks
<p>This repository contains data sets and materials for a gold standard elaboration on toxicity and incivility in the digital sphere based on human coding to benchmark algorithmic classification tasks with transformers and LLMs. <strong>The labelling progress is 62%</strong>.</p> <p>We are labelling two samples of novel datasets of political digital interactions on Twitter (rebranded as X). The first set comprises almost 5 million data points from three Latin American protest events: (a) protests against the coronavirus and judicial reform measures in Argentina during August 2020; (b) protests against education budget cuts in Brazil in May 2019; and (c) the social outburst in Chile stemming from protests against the underground fare hike in October 2019. We are focusing on interactions in Spanish to elaborate a gold standard for digital interactions in this language, therefore, we prioritise Argentinian and Chilean data. The second set contains more than 31 million messages and more than 9 million interactions between 2010 and 2022, covering the election of members of the first Constitutional Convention in Chile, the drafting process and the referendum in which the proposal was rejected.</p> <p>This project is generously funded by the <strong>OpenAI Academic Programme</strong>, <strong>2024 FAE-UDP Research Grant</strong>, and partially by the <strong>St Hilda's College Muriel Wise Fund at the University of Oxford</strong>. The <a href="https://training-datalab.com/"><strong>Training Data Lab</strong></a> research group also logistically supports this project.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.