Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

4,612

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

4,612 results for “labeling”

Learn how ShareScore rates datasets ↗
zenodo48/100

Cortical slice labelled with anti GFP and VAMP2 antibodies - sample image for software testing of "Contacting synapse" protocol

<p><strong>Image 1.tif is a Brain slice</strong>. This 16 bits confocal stack of pictures ((801x711 pixels x33 z slices - pixel size 78.17 nm) of a brain slice has been taken at 93x (LeicaHC PL APO CS2 93x/1.30 GLYC) in sequential mode with two channels : one dedicated to the GFP detection, and the other one to synpatic boutons labelled with VAMP2 protein. VAMP2 protein are expressed at glutamatergic presynaptic sites and is usually found apposed to Post Synaptic Density. This is a good sample to test &quot;contacting synapse&quot; software. Here GFP cells were electroporated with various plasmid. The aim of the software is to identify if expression of those plasmid within the GFP labelled cell, influence the density of synapse contacting this GFP cells. Here presynaptic contact are identified through the use of antibodies to VAMP2 proteins.</p>

opencc-by-4.0Jun 2022View details →
zenodo48/100

An annotated high-content fluorescence microscopy dataset with EGFP-Galectin-3-stained cells and manually labelled outlines

<p>Here we present a benchmarking dataset of fluorescence microscopy images with EGFP-Galectin-3-stained cells together with annotations of their outlines. Images were randomly selected from an RNA interference screen with a modified U2OS osteosarcoma cell line, acquired on a Thermo Fischer CX7 high-content imaging system at 20x magnification.&nbsp;</p> <p>The dataset contains 60 images showing over 2000 labelled nuclear objects in total, which is sufficiently large to train well-performing neural networks for instance or semantic segmentation. It is pre-split into training, development and test set, each in a zip file. The dataset should be referred to as Aitslab_bioimaging2.</p> <p>For most of the images, nuclear staining and annotations have been published previously in the dataset Aitslab_bioimaging1 (https://doi.org/10.5281/zenodo.6657260). The conversion script to produce the png images from the C01 images was published together with this dataset.</p>

opencc-by-4.0Jun 2024View details →
zenodo48/100

Topic Labels of "Dynamic Topic Modelling for Exploring the Scientific Literature on Coronavirus: An Unsupervised Labelling Technique"

<p>These are the labels generated with the method proposed in the article <em>"Dynamic Topic Modelling for Exploring the Scientific Literature on Coronavirus: An Unsupervised Labelling Technique".</em> These labels are for the 100 and 200 DTM topic models, trained both with the whole corpus and with only the COVID-19 period data&nbsp;</p> <p>&nbsp;</p> <p>For the generation of these labels you can go to the original published work or to the linked Zenodo resource.</p>

opencc-by-4.0Jul 2024View details →
zenodo48/100

Comparison and practical review of segmentation approaches for label-free microscopy

<p>This dataset contains microscopic images of PNT1A cell line captured by multiple microcopic without use of any labeling and a manually annotated ground truth for subsequent use in segmentation algorithms. Dataset also includes images reconstructed according to the methods described below in order to ease further segmentation.&nbsp;</p> <p>See&nbsp;Vicar et al. Cell segmentation methods for label-free contrast microscopy: review and comprehensive comparison. BMC Bioinformatics (2019) 20:360. DOI&nbsp;<a href="https://doi.org/10.1186/s12859-019-2880-8">10.1186/s12859-019-2880-8</a></p> <p>Code using this dataset is available at&nbsp;<a href="https://github.com/tomasvicar/Cell-segmentation-methods-comparison">https://github.com/tomasvicar/Cell-segmentation-methods-comparison</a></p> <p><strong>Materials and methods&nbsp;</strong></p> <p>Cells were cultured in RPMI-1640 medium supplemented with antibiotics (penicillin 100 U/ml and streptomycin 0.1 mg/ml) with 10%&nbsp;fetal bovine serum. Prior microscopy acquisition, cells were maintained at 37 cenigrade in a humidified incubator with 5% CO2. Intentionally, high passage number of cells was used (&gt;30) in order to describe distinct morphological heterogeneity of cells (rounded and spindle-shaped, relatively small to large polyploid cells). For acquisition purposes, cells were cultivated in Flow chambers &micro;-Slide I Luer Family (Ibidi, Martinsried, Germany).</p> <p>Quantitative phase imaging (QPI)&nbsp;microscopy was performed on Tescan Q-PHASE (Tescan, Brno, Czech republic), with objective Nikon CFI Plan Fluor 10x/0.30 captured by Ximea MR4021MC (Ximea, M&uuml;nster, Germany). Imaging is based on the original concept of coherence-controlled holographic microscope \cite{Kolman:10,Slaby:13}, images are shown in grayscale with units of pg/&micro;m2.</p> <p>DIC microscopy was performed on microscope Nikon A1R (Nikon, Tokyo, Japan), with objective Nikon CFI Plan Apo VC 20x/0.75 captured by CCD camera Jenoptik ProgRes MF (Jenoptik, Jena, Germany).&nbsp;</p> <p>HMC microscopy was performed on microscope Olympus IX71 (Olympus, Tokyo, Japan), with objective Olympus CplanFL N 10x/0.3 RC1 captured by CCD camera Hamamatsu Photonics ORCA-R2 (Hamamatsu Photonics K.K., Hamamatsu, Japan).</p> <p>PC microscopy was performed on a Nikon Eclipse TS100-F microscope, with a Nikon CFI Achro ADL 10x/0.25 objective captured by CCD camera Jenoptik ProgRes MF.</p> <p><strong>Folder structure and file and filename description</strong><br> <br> <em>folder &quot;source data+groundtruth&quot;</em><br> - includes raw microscopic data&nbsp;<br> &nbsp; (uncompressed 16-bit for DIC, HMC and PC,&nbsp;32-bit for QPI)<br> - includes manualy annotated groundtruth&nbsp;(zip file - imageJ ROI file, 1bit png mask)</p> <p>e.g.&nbsp;<br> DIC_01_raw.tif<br> DIC_01_groundtruth_imagejROI.zip<br> DIC_01_groundtruth_mask.png</p> <p><br> <em>folder &quot;reconstructions&quot;</em></p> <p>includes reconstructed images using reconstructions with highest dice coefficient achieved.&nbsp;</p> <p>for DIC and HMC: rDIC-Koos, rDIC-Yin, and rWeka<br> for PC: rPC-Top-Hat, rDIC-Yin, and rWeka<br> for QPI: rWeka</p> <p>note that for rWeka images numbered 01 for DIC, HMC and PC and 01-03 for QPI were used for learning.</p> <p><strong>Abbreviations</strong><br> DIC, differential image contrast<br> HMC, Hoffman modulation contrast<br> PC, phase contrast<br> QPI, quantitative phase imaging<br> rDIC-Koos, DIC/HMC image reconstruction according to Koos et al, Sci Rep. 2016;6:30420<br> rDIC-Yin, DIC/HMC image reconstruction according to Yin et al, Inf Process Med Imaging. 2011;22:384-97.<br> rPC-Yin, PC image reconstruction according to Yin et al, &nbsp;Med Im Anal. 2012; 16(5):1047<br> rPC-Top-Hat, Top-Hat filter according to Dewan et al, IEEE Transactions on Biomedical Circuits and<br> Systems.2014;8(5):716-728<br> rWeka, probability map using Trainable Weka segmentation according to Arganda-Carreras et al. Bioinformatics. 2017</p>

opencc-by-4.0May 2018View details →
zenodo48/100

Extracellular recordings and juxtacellular labelling with glass electrodes in the mouse medial septum and hippocampus

<p>This repository contains MAT files consisting&nbsp;of&nbsp;simultaneously recorded mouse medial septal and hippocampal&nbsp;local field potentials&nbsp;(20 kHz sampling rates) and spikes from&nbsp;single medial septal&nbsp;cells. Data were recorded with glass electrodes&nbsp;during spontaneous&nbsp;movement and rest periods, followed by juxtacellular labelling of the medial septal cell. Text files of the&nbsp;spike times and detected hippocampal CA1 theta (5-12 Hz) oscillation trough times are associated with each MAT file.</p> <p>The files are organised by cell (neuron) name. For further details,&nbsp;see the CSV file included with the dataset. These recorded and labelled single cells were originally reported in Joshi et al 2017, Viney et al 2018, and Salib et al 2019.</p> <p>Each MAT file contains the following channels, exported from the original Spike2 (smr) recording files:</p> <p>(1) Details of the recording</p> <p>(2) Detected spikes (in seconds) from the single medial septal cell</p> <p>(3) Movement detection (eg. accelerometer or rotary encoder)</p> <p>(4) Local field potential (medial septum), in mV</p> <p>(5) Local field potential (hippocampal CA1), in mV;&nbsp;see CSV file for precise location (e.g. within stratum pyramidale)</p> <p>This dataset is made available under a Creative Commons Attribution 4.0 International&nbsp;(CC BY 4.0) license: If you share or adapt these data you must give appropriate credit, provide a link to the license, and indicate if changes were made.</p>

opencc-by-4.0Jul 2023View details →
zenodo48/100

European consumers ́ preference and willingness to pay for food products labelled as obtained by a circular farming system -in relation to environmental attitudes and consumption behaviours

<p>Data was collected with questionnaire-based research carried out in Belgium, Croatia, Hungary, Italy, Poland, and Spain as part of a European project. The survey questions were designed to obtain the Willingness to pay using 2 different methodologies the discrete choice experiment and the open-end choice experiment. The survey also included questions about consumers environmental attitudes, and consumption behavior (purchase, use and recycling), to identify if them have influence on preferences towards more sustainable food products. The 3 analyzed food products were pork meat, milk and bread, all of them obtained through different agricultural production systems (circular, conventional, and organic agriculture). The sample was stratified in terms of gender and age to be representative to the average population in each country. Furthermore, respondents included in this study were those that are mainly, or in part responsible for the household food shopping. The questionnaire was translated to the languages of the countries involved in the data collection and pre-launched using a pilot sample of 50 consumers in each case study country. &nbsp;Finally, a total of 5,362 validated questionnaires were obtained. Data was collected online using the Qualtrics market research company, and Net panel market company for Hungary from June 2021 to January 2022.</p>

opencc-by-4.0Sep 2023View details →
edi48/100

MCR LTER: Coral Reef: Computer Vision: Moorea Labeled Corals

The Moorea Labeled Corals dataset is a subset of the MCR LTER packaged for computer vision research. It contains 2055 images from three habitats IDs: fringing reef outer 10m and outer 17m, from 2008, 2009 and 2010. It also contains random point annotation (row, col, label) for the nine most abundant labels, four non coral labels: (1) Crustose Coralline Algae (CCA), (2) Turf algae, (3) Macroalgae and (4) Sand, and five coral genera: (5) Acropora, (6) Pavona, (7) Montipora, (8) Pocillopora, and (9) Porites. These nine classes account for 96% of the annotations and total to almost 400,000 points. These nine classes are the ones analyzed in (Beijbom, 2012); less-abundant genera not treated in the automation are also present in the dataset. These data were published in Beijbom O., Edmunds P.J., Kline D.I., Mitchell G.B., Kriegman D., 'Automated Annotation of Coral Reef Survey Images', IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Providence, Rhode Island, 2012. [BibTex] [pdf] These data are a subset of the raw data from which knb-lter-mcr.4 is derived. This material is based upon work supported by the U.S. National Science Foundation under Grant No. OCE 16-37396 (and earlier awards) as well as a generous gift from the Gordon and Betty Moore Foundation. Research was completed under permits issued by the French Polynesian Government (Délégation à la Recherche) and the Haut-commissariat de la République en Polynésie Francaise (DTRT) (Protocole d'Accueil 2005-2018). This work represents a contribution of the Moorea Coral Reef (MCR) LTER Site.

openCC (other)May 2012View details →
zenodo44/100

Envargs - Labelling Environmental Belgian Case Law

<p>This dataset contains cases from the Belgian administrative law, the topic of law they belong to, as well as an additional binary label that determines if the case can be considered as an &#39;environmental law&#39; case.&nbsp;</p> <p><strong>Files description:&nbsp;</strong><br> <strong>- topics_env.csv :</strong> the columns contains&nbsp;the list of specific topics that are considered to be &#39;environmental law&#39; topics. Under each topic is the list of case_numbers that are related to that topic of law.&nbsp;<br> <br> <strong>- topics_non_env.csv :&nbsp;</strong>the columns contains&nbsp;the list of specific topics that are considered <strong>not</strong> to be &#39;environmental law&#39; topics. Under each topic is the list of case_numbers that are related to that topic of law.&nbsp;<br> <br> <strong>- labelled_case_dataset.csv :</strong> this file contains all the cases together with their text and their label (0 if non-environmental, 1 if environmental). The columns are the following: [&#39;case_number&#39;, &#39;topic&#39;, &#39;year&#39;, &#39;case_text&#39;, &#39;label&#39;]. The &#39;case_text&#39; filed contain the raw text of the case (no pre-processing applied).&nbsp;<br> <br> <strong>- labelled_case_dataset_lemmatized.csv:&nbsp;</strong>this file contains all the cases together with their text and their label (0 if non-environmental, 1 if environmental). The columns are the following: [&#39;case_number&#39;, &#39;topic&#39;, &#39;year&#39;, &#39;case_text&#39;, &#39;label&#39;]. The &#39;case_text&#39; filed contain the lemmatised version of the cases, stopwords and punctuation signs&nbsp;have been removed too.&nbsp;<br> <br> <strong>- labelled_case_dataset_no_stopwords.csv </strong>:&nbsp;this file contains all the cases together with their text and their label (0 if non-environmental, 1 if environmental). The columns are the following: [&#39;case_number&#39;, &#39;topic&#39;, &#39;year&#39;, &#39;case_text&#39;, &#39;label&#39;]. The &#39;case_text&#39; filed contain the text of the cases, after having removed stopwords and punctuation signs.&nbsp;</p>

opencc-by-4.0Jun 2020View details →
zenodo44/100

Fluorescently-labelled zebrafish pronephroi + ground truth classes (normal/cystic) + trained CNN model

<p>This upload contains :</p> <p>- <strong>images.zip:&nbsp;</strong> microscope images of fluorescently-labelled pronephroi in larvae of the <em>Tg(wt1b:EGFP)</em> transgenic zebrafish line showing 2 morphologies (normal vs cystic) upon injection with Co-Mo or ift172-MO, respectively. Images were obtained using &nbsp;an ACQUIFER Imaging Machine widefield high content screening microscope.</p> <p>Reference:&nbsp;</p> <p>Pandey, G., Westhoff, J., Schaefer, F. and Gehrig, J. (2019). <strong>A Smart Imaging Workflow for Organ-Specific Screening in a Cystic Kidney Zebrafish Disease Model</strong>. International Journal of Molecular Sciences <em>20</em>, 1290, doi:<a href="https://doi.org/10.3390/ijms20061290">10.3390/ijms20061290</a>.</p> <p>&nbsp;</p> <p>- <strong>Annotations-***.csv : </strong>Tables containing ground-truth category classes (normal vs cystic) for the images in the zip file.</p> <p>The tables contain&nbsp;columns with the image filename, folder and category.</p> <p>Note&nbsp;:&nbsp;<strong>the Folder column should be updated with the root folder directory once downloaded on your machine.</strong></p> <p>These&nbsp;files were&nbsp;generated with the Fiji plugin <em>single-class (button)</em>&nbsp;from the <em>Qualitative-Annotations</em> update site.</p> <p>The 2 files contain&nbsp;the same information, they only differ in the formatting&nbsp;of the category, the <em>singleColumn </em>file has a single category column while the <em>multiColumn</em> has 2 columns (normal/cystic) with 0/1 encoding.</p> <p>The choice of category encoding solely depends on how the table is used, i.e. in which training workflow, home-made script or software.</p> <p>- <strong>trainedModel.zip :&nbsp;</strong>This archive contains 2 files:&nbsp;<strong>(1) </strong>a h5 file corresponding to a trained deep-learning model to classify the images of the dataset in the 2 categories (normal vs cystic), and <strong>(2)</strong>&nbsp;a text file containing the class names. Both files are necessary to predict the category of new images similar to the one in the dataset, for instance using the published KNIME workflows.</p>

opencc-by-4.0Aug 2020View details →
zenodo44/100

Node2Vec model - Czech Wikidata (knowledge graph / labels / l80 / rw40)

<p>Node2Vec&nbsp;embedding model trained on Czech wikidata (from October 2020) labels using gensim implementation of Word2Vec with the following parameters for random walks:</p> <ul> <li>length of walk = 80</li> <li>number of random walks = 40</li> </ul>

opencc-by-4.0Jan 2021View details →
zenodo44/100

Node2Vec model - Czech Wikidata (knowledge graph / labels / l160 / rw40)

<p>Node2Vec&nbsp;embedding model trained on Czech wikidata (from October 2020) labels using gensim implementation of Word2Vec with the following parameters for random walks:</p> <ul> <li>length of walk = 160</li> <li>number of random walks = 40</li> </ul>

opencc-by-4.0Jan 2021View details →
zenodo44/100

Cu dataset – A copper ore labeled images dataset for segmentation training and testing

<p>This dataset is composed of 121 pairs of correlated images. Each pair contains one image of a copper ore sample acquired through reflected light microscopy (RGB, 24-bit), and the corresponding binary reference image (8-bit), in which the pixels are labeled as belonging to one of two classes: ore (0) or embedding resin (255).</p> <p>The sample came from a copper ore from Yauri Cusco (Peru) with a complex mineralogy, mainly composed of sulfides, oxides, silicates, and native copper. It was classified by size. The fraction +74-100 &mu;m was cold mounted with epoxy resin and subsequently ground and polished.</p> <p>Correlative microscopy was employed for image acquisition. Thus, 121 fields were imaged on a reflected light microscope with a 20&times; (NA 0.40) objective lens and on a scanning electron microscope (SEM). In sequence, they were registered, resulting in images of 1017&times;753 pixels with a resolution of 0.53 &micro;m/pixel. As matter of fact, some images (the images No. 2, 3, 24, 25, 46, 47, 69, 91, and 113) have slightly smaller sizes because they were cropped during the registration procedure to correct co-localization errors of the order of a few pixels. Finally, the images from SEM were thresholded to generate the reference images.</p> <p>Further description of this sample and its imaging procedure can be found in the work by Gomes and Paciornik (2012).</p> <p>This dataset was created for developing and testing deep learning models on semantic segmentation tasks. The paper of Filippo et al. (2021) presented a variant of the DeepLabv3+ model (Chen et al., 2018) that reached mean values of 90.56% and 92.12% for overall accuracy and F1 score, respectively, for 5 rounds of experiments (training and testing), each with a different, random initialization of network weights.</p> <p>For further questions and suggestions, please do not hesitate to contact us.</p> <p>&nbsp;</p> <p><strong>Contact email</strong>: ogomes@gmail.com</p> <p>&nbsp;</p> <p>If you use this dataset in your own work, please cite this DOI: 10.5281/zenodo.5020566</p> <p>&nbsp;</p> <p>Please also cite this paper, which provides additional details about the dataset:</p> <p>Michel Pedro Filippo, Ot&aacute;vio da Fonseca Martins Gomes, Gilson Alexandre Ostwald Pedro da Costa, Guilherme Lucio Abelha Mota. <em>Deep learning semantic segmentation of opaque and non-opaque minerals from epoxy resin in reflected light microscopy images</em>. <strong>Minerals Engineering</strong>, Volume 170, 2021, 107007, https://doi.org/10.1016/j.mineng.2021.107007.</p>

opencc-by-4.0Jun 2021View details →
zenodo44/100

Quantitative Content Analysis Data for Hand Labeling Road Surface Conditions in New York State Department of Transportation Camera Images

<p><strong>Foundational Codebook and Data:&nbsp;</strong></p> <p>Traffic camera images from the New York State Department of Transportation (511ny.org) are used to create a hand-labeled dataset of images classified into to one of six road surface conditions: 1) severe snow, 2) snow, 3) wet, 4) dry, 5) poor visibility, or 6) obstructed. Six labelers (authors Sutter, Wirz, Przybylo, Cains, Radford, and Evans) went through a series of four labeling trials where reliability across all six labelers were assessed using the Krippendorff&rsquo;s alpha (KA) metric (Krippendorff, 2007). The online tool by Dr. Freelon (Freelon, 2013; Freelon, 2010) was used to calculate reliability metrics after each trial, and the group achieved inter-coder reliability with KA of 0.888 on the 4th trial. This process is known as quantitative content analysis, and three pieces of data used in this process are shared, including: 1) a PDF of the codebook which serves as a set of rules for labeling images, 2) images from each of the four labeling trials, including the use of New York State Mesonet weather observation data (Brotzge et al., 2020), and 3) an Excel spreadsheet including the calculated inter-coder reliability (ICR) metrics and other summaries used to asses reliability after each trial. The data are included in NYSDOT_quantitative_content_analysis.zip.</p> <p>The broader purpose of this work is that the six human labelers, after achieving inter-coder reliability,&nbsp;can then label large sets of images independently, each contributing to the creation of larger labeled dataset&nbsp;used for&nbsp;training supervised machine learning models to predict road surface conditions from camera images. The xCITE lab&nbsp;(xCITE, 2023) is used to store&nbsp;camera images from 511ny.org, and the lab provides computing resources for training machine learning models.</p> <p><strong>Obstructed Class Variation: </strong></p> <p>There are many applications for labeling roadside camera images, and as a variation of the foundational codebook, an addendum codebook provides another version of labeling the obstructed class. Specifically, this variation prioritizes labeling an image as &ldquo;obstructed&rdquo; only in extreme circumstances where there is a camera- or image- specific problem that prevents the assessment of any road surfaces. For labelers who want to use this version of the obstructed class (in this document) and also the other five weather-related classes (in the foundational codebook), the guidance is to use both documents in tandem, making sure to use the obstructed rules/definitions in this document while disregarding the obstructed rules/definitions in the foundational codebook. Alternatively, this codebook may be used alone in applications where the goal is to solely classify obstructed vs not obstructed.&nbsp;To ensure reliability and quality of this variation, quantitative content analysis was conducted on this addendum codebook, just as it was for the foundational codebook. Two labelers were tested with a sample of 30 images and achieved inter-coder reliability with Krippendorff's Alpha of 0.934 after one trial. The data, including the addendum codebook and labeling trial data (images and results) are included in ObstructedVariation_quantitative_content_analysis.zip.</p> <p>This material is based upon work supported by the U.S. National Science Foundation under Grant No. RISE-2019758.</p>

opencc-by-4.0Sep 2023View details →
zenodo44/100

Today's cat is tomorrow's dog: accounting for time-based changes in the labels of ML vulnerability detection approaches (Replication Package Part 3: OpenSSL dataset)

<h1><strong>The Replication Package of</strong></h1> <h1><strong>"Today's cat is tomorrow's dog: accounting for time-based changes in the labels of ML vulnerability detection approaches"</strong></h1> <h3><strong>Part 3 (OPENSSL Dataset)</strong></h3> <div> <div>This repository includes:</div> <ol> <li><em><strong>Code.zip</strong></em> that contains the codes to replicate some parts of this study:<br>a.&nbsp;<em>1_generate_datasets</em> implements our methodology to generate the datasets.<br>b.&nbsp;<em>2_run_models</em> runs the ML models during the evaluation.<br>c.&nbsp;<em>3_result_replication </em>generates charts presented in the paper from the ML evaluation results.</li> <li><em><strong>Datasets.zip</strong></em> that contain 2 folders:<br>a.&nbsp;<em>original</em> datasets: 1 from <a href="https://github.com/CGCL-codes/VulDeePecker" target="_blank" rel="noopener">NVD Vuldeepecker</a> and 3 extracted from&nbsp;<a href="https://github.com/ZeoVan/MSR_20_Code_vulnerability_CSV_Dataset" target="_blank" rel="noopener">BigVul</a>.<br> <div> <div>b. <em>OPENSSL</em> datasets: train, validation, test sets for each time of observation extracted using our methodology from <a href="https://github.com/ZeoVan/MSR_20_Code_vulnerability_CSV_Dataset" target="_blank" rel="noopener">BigVul</a>&nbsp;dataset for project <em>openssl</em>.</div> </div> </li> <li><em><strong>Pretrained-models.zip</strong></em>&nbsp;that we generated during our evaluation (3 test results for each time point in the timeline [2013-2019]).</li> <li><em><strong>Results.zip</strong></em> of our evaluation, the folder <em>ALL</em> contains the overall results and other folders are results by model.</li> </ol> <p><strong>UPDATED version 5<br></strong>- added a GLOBAL_README.md which contains the 3 stages and how they are connected to each other<br>- updated LineVul.ipynb: import AdamW from torch.optim instead of transformers<br>- updated README.md in Code2Vec with the prerequisites of Java to run gradlew for astmine</p> <p><strong>UPDATED version 6<br></strong>- updated CodeBert.ipynb: import AdamW from torch.optim instead of transformers</p> <p>Documentations</p> <ol> <li><em><strong>INSTALL.pdf&nbsp;</strong></em>: how to install the codes</li> <li><em><strong>README.pdf</strong></em>: readme file</li> <li><em><strong>REQUIREMENTS.pdf</strong></em>: hardware and software requirements</li> <li><em><strong>STATUS.pdf</strong></em>&nbsp;: status for artifact submission</li> <li><em><strong>LICENSE.pdf</strong></em>: the license of this artifact</li> <li><em><strong>PAPER.pdf</strong></em>: the camera-ready version of the paper</li> </ol> </div> <div> <div>Please refer to the following repositories for the other datasets and pre-trained models:</div> <div>- Part 1 NVD Vuldeeepecker :&nbsp;<a href="https://doi.org/10.5281/zenodo.8207883" target="_blank" rel="noopener">https://doi.org/10.5281/zenodo.8207883</a></div> - Part 2 LINUX :&nbsp;<a href="https://doi.org/10.5281/zenodo.10960662" target="_blank" rel="noopener">https://doi.org/10.5281/zenodo.10960662</a><br> <div>- Part 4 POPPLER : <a href="https://doi.org/10.5281/zenodo.14713143">https://doi.org/10.5281/zenodo.14713143</a></div> <div>&nbsp;</div> <div>This work was partly funded by the EU under the H2020 Program AssureMOSS (Grant n. 952647) and the Horizon Europe Program Sec4AI4Sec (Grant n. 101120393), by the Italian Ministry of University and Research (MUR) under the P.N.R.R. &ndash; NextGenerationEU grant n.\ PE00000014 (SERICS subproject COVERT), and by the Dutch Research Council (NWO) under the grant NWA.1215.18.006 (Theseus) and grant KIC1.VE01.20.004 (HEWSTI).&nbsp;</div> </div>

opencc-by-4.0Apr 2024View details →
zenodo44/100

A German Language Labeled Dataset of Tweets

<p>Our dataset contains 8,048 German language tweets related to Jewish life from a four-year timespan.&nbsp;</p><p>The dataset consists of 18 samples of tweets with the keyword "Juden" or "Israel." The samples are representative samples of all live tweets (at the time of sampling) with these keywords respectively over the indicated time period. Each sample was annotated by two expert annotators using an Annotation Portal that visualizes the live tweets in context. We provide the annotation results based on the agreement of two annotators, after discussing discrepancies (Jikeli et al. 2022: 3-6).&nbsp;</p><p>&nbsp;Overall, 335 tweets (4%) were labelled as antisemitic following the IHRA Working Definition of Antisemitism. 1345 tweets (17 %) come from 2019, 1364 tweets (17 %) from 2020, 2639 tweets (33 %) from 2021 and 2700 tweets (34 %) from 2022.&nbsp;</p><p>About half of the tweets, a total of 4,493 tweets (56 %) come from queries with the keyword "Juden," which is representative of a continuous time period from January 2019 to December 2022: 864 tweets (19 %) come from 2019, 891 tweets (20 %) from 2020, 1364 tweets (30 %) from 2021 and 1374 (31 %). 148 out of the 4493 tweets, so 3% from the query with "Juden" are antisemitic.&nbsp;</p><p>The other part of the tweets, a total of 3,555 (44 %)&nbsp; results of queries with the keyword "Israel". 481 tweets (14 %) of the keywords containing Israel stem from 2019, 473 (13 %) come from 2020, 1275 tweets (36 %) from 2021 and 1326 tweets (37 %) are from 2022. Out of all tweets from the "Israel" query, 187 (5 %)&nbsp; are antisemitic.&nbsp;</p><p>The csv file contains diacritics and special characters of the German language (e.g., "ä", "ü", "ö", "ß"), which should be taken into account when opening it with anything other than a text editor.&nbsp;</p><p><strong>Acknowledgements</strong>&nbsp;</p><p>This work used Jetstream2 at Indiana University through allocation HUM200003 from the Advanced Cyberinfrastructure Coordination Ecosystem: Services &amp; Support (ACCESS) program, which is supported by National Science Foundation grants #2138259, #2138286, #2138307, #2137603, and #2138296.&nbsp;&nbsp;</p><p>We are grateful for the support of Indiana University's Observatory on Social Media (OSoMe) (Davis et al. 2016) and the contributions and annotations of all team members in our Social Media &amp; Hate Research Lab at Indiana University's Institute for the Study of Contemporary Antisemitism, especially Grace Bland, Elisha S. Breton, Kathryn Cooper, Robin Forstenhäusler, Sophie von Máriássy, Mabel Poindexter, Jenna Solomon, Clara Schilling, Emma Shriberg and Victor Tschiskale.&nbsp;</p>

opencc-by-4.0Oct 2023View details →
zenodo44/100

2 million histological images of breast cancer tumors with her2 labels

<p><strong>Data Description</strong><br> This is a 2 million set of non-overlapping image patches from hematoxylin &amp; eosin (H&amp;E) stained histological images of human breast cancer tumor tissue.</p> <p>The anonymized dataset comes from a cohort of BC patients from the A. C. Camargo Cancer Center (ACCCC, N = 504). All patients were treated for breast cancer at the ACCCC between 2019 and 2021. As part of their diagnosis, in HER2 IHC score 2+ cases, patients&#39; HER2 status was determined following the ASCO guidelines updated in 2018, with visual evaluation of IHC assay and either a FISH or DDISH test. All cases with metastasis or neoadjuvant treatment were excluded.</p> <p>A total of 426 H&amp;E stained high resolution images (40x magnification) were scanned from biopsy and resection tissue samples with a Leica Aperio AT2 scanner. Ethical approval of the ACCCC study was given by the ethics committee of the Funda&ccedil;&atilde;o Ant&ocirc;nio Prudente. We divided the cases into the following 3 groups according to the results of the IHC and ISH tests: HER2-negative, HER2-low and HER2-high.</p> <p>The slides were divided into 256 px x 256 px tiles at 0.5 um/pixel magnification. Then, we used a custom trained ConvNext-tiny neural network to only include tiles from the tumor region and its environment, generating a total of 2051877 image patches.</p> <p>A sample is considered her2-negative with an IHC score of 0; her2-low with an IHC score of 1+ or an IHC score of 2+ with a negative ISH-based test result, and her2-high with an IHC score of 2+ with a positive ISH-based test or an IHC score of 3+.</p> <p>The accompanying code used for training&nbsp;the models is available at https://github.com/tojallab/wsi-mil</p>

opencc-by-4.0Oct 2023View details →
zenodo44/100

Improving the accuracy of automated labeling of specimen images datasets via a confidence-based process - Datasets

<p>This dataset contains supporting data for a research project aimed at analysing herbarium samples from the New England area at a large scale with deep learning techniques. Details on the methodology are shared in the acompanying paper (to be published).</p> <p>Content:</p> <ul> <li>dataset600k_withAI.csv : A dataset of over 600.000 herbarium samples with its record metadata and a corresponding AI phenological annotations with matching confidence scores. The entirety of the record headers are provided, extracted directly from the NEVP portal. In addition, the AI labels are defined by the following headers. These 8 columns represent 4 binary classifiers with the Presence/Absence of each 4 traits and corresponding confidence (as a percentage - presence/absence percentages sum to 1).<br> <ul> <li> <table> <tbody> <tr> <td>Flowering</td> <td>Not Flowering</td> <td>Budding</td> <td>Not Budding</td> <td>Fruiting</td> <td>Not Fruiting</td> <td>Reproductive</td> <td>Not Reproductive</td> </tr> </tbody> </table> </li> </ul> </li> </ul> <ul> <li>data_species_with_statuses.csv: A processed dataset summarizing flowering period shift at a species level. Two types of headers are provided. <ul> <li>First metadata concerning the flowering shift and the data used to compute that value:&nbsp; <ul> <li> <table> <tbody> <tr> <td>genus</td> <td>genus_species</td> <td>slope</td> <td>nb_specimens</td> <td>p_value_significance</td> <td>trend_category</td> </tr> <tr> <td>Genus of the species</td> <td>Binomial name of the species</td> <td>Regression slope defining the flowering shift as a slope</td> <td>Number of herbarium specimens used to compute the shift</td> <td>P-value significance of the slope being non-zero. ('Non Significant'/'Significant')</td> <td>Summary of the shift as a binary characteristic ('Earlier'/'Later')</td> </tr> </tbody> </table> </li> </ul> </li> <li>Second, metadata summarizing various traits associated to each species: <ul> <li> <table> <tbody> <tr> <td>lifeform_status</td> <td>native_introduced_status</td> <td>wetland_status</td> <td>seasonality_average</td> <td>seasonality_spread</td> </tr> <tr> <td>Growth form from the USDA PLANTS Database. 'Forb_Herb', 'Shrub_Tree' or 'Vine'</td> <td>'Native'/'Introduced' status from the USDA PLANTS Database.</td> <td> <p>National Wetland Plant List (NWPL) Wetland Indicator Status within the Northcentral and Northeast Region</p> <p>'OBL'/'FACW'/'FAC'/'FACU'/'UPL'</p> </td> <td>A characteristic of the flowering season of the species based on the mean Day of Year of the analysed specimens: if &lt;=180: 'Early', else 'Late'</td> <td>A characteristic of the flowering season of the species based on the spread of the flowering season. Less than 28 days: 'Narrow', larger: 'Large'.</td> </tr> </tbody> </table> <p>&nbsp;</p> </li> </ul> </li> </ul> </li> <li>phylogenetic_tree.tre: The raw data used to generate the visualization of the flowering seasonality character and the detected flowering shift foreach species on a phylogenetic tree.</li> <li>phylogenetic_processed_dataset.csv: The processed dataset resuting from the&nbsp;phylogenetic signal analysis. For each trait, an associated significance binary value is provided.</li> </ul>

opencc-by-4.0Nov 2024View details →
zenodo44/100

19th Century United States Newspaper images predicted as Photographs with labels for "human", "animal", "human-structure" and "landscape"

<p>The Dataset contains images derived from the Newspaper Navigator (news-navigator.labs.loc.gov/), a dataset of images drawn from the Library of Congress Chronicling America collection (<a href="https://chroniclingamerica.loc.gov/">chroniclingamerica.loc.gov/</a>).&nbsp;</p> <blockquote> <p>[The Newspaper Navigator dataset] consists of extracted visual content for 16,358,041 historic newspaper pages in&nbsp;<em>Chronicling America</em>. The visual content was identified using an object detection model trained on annotations of World War 1-era Chronicling America pages, including annotations made by volunteers as part of the&nbsp;<a href="https://labs.loc.gov/work/experiments/beyond-words/">Beyond Words</a>&nbsp;crowdsourcing project.</p> <p>source:<a href="https://news-navigator.labs.loc.gov/"> https://news-navigator.labs.loc.gov/</a></p> </blockquote> <p>One of these categories is &#39;photographs&#39;. This dataset contains a sample of these images with additional labels indicating if the photograph has one or more of the following labels: &quot;human&quot;, &quot;animal&quot;, &quot;human-structure&quot; and &quot;landscape&quot;</p> <p>The data is organised as follows:</p> <ul> <li>The images themselves can be found in `images.zip`</li> <li>`newspaper-navigator-sample-metadata.csv` contains metadata about each image drawn from the Newspaper Navigator Dataset.</li> <li>`multi_label.csv` contains the labels for the images as a CSV file</li> <li>`annotations.csv` conains the labels for the images with additional metadata</li> </ul> <p>This dataset was created for use in an under-review Programming Historian tutorial (<a href="http://programminghistorian.github.io/ph-submissions/lessons/computer-vision-deep-learning-pt2">http://programminghistorian.github.io/ph-submissions/lessons/computer-vision-deep-learning-pt2</a>) The primary aim of the data was to provide a realistic example dataset for teaching computer vision for working with digitised heritage material. The data is shared here since it may be useful for others. <strong>This data documentation is a work in progress and will be updated when the Programming Historian tutorial is released publicly. </strong></p> <p>The metadata CSV file contains the following columns:</p> <p>- filepath<br> - pub_date<br> - page_seq_num<br> - edition_seq_num<br> - batch<br> - lccn<br> - box<br> - score<br> - ocr<br> - place_of_publication<br> - geographic_coverage<br> - name<br> - publisher<br> - url<br> - page_url<br> - month<br> - year<br> - iiif_url</p>

openother-openJan 2022View details →
zenodo44/100

Raw EEG Data for: Learning from Label Proportions in Brain-Computer Interfaces

<p>If you prefer to use the preprocessed and epoched data, please refer to: https://zenodo.org/record/192684</p> <p>Note that this repository ontains only the visual paradigm with the N=13 subjects recorded at 31 EEG channels, as described in the above link. We copied the relevant section of the description below:</p> <blockquote> <p>This data repository contains raw EEG of an EEG experiment utilizing visual event-related potentials (ERPs) with N=13 healthy subjects.</p> <p>The dataset is used and described in the following journal article:</p> <p><em>H&uuml;bner, D., Verhoeven, T., Schmid, K., M&uuml;ller, K. R., Tangermann, M., &amp; Kindermans, P. J. (2017). Learning from label proportions in brain-computer interfaces: online unsupervised learning with guarantees. PloS one, 12(4), e0175856.</em></p> <p><strong>Please cite the above article when using the data.</strong></p> <p>The data set with N=13 subjects is different to ordinary ERP datasets in the sense that the train of stimuli to spell one character (68) is divided into repetitions of two interleaved sequences with length 8 and 18, respectively. We added &#39;#&#39; symbols to the spelling matrix which should never be attended by the subject and hence, are non-targets by definition. The first, shorter sequence, now highlights only ordinary characters, while the second sequence also highlights &#39;#&#39; -- visual blank symbols. By construction, sequence 1 has a higher target ratio than sequence 2. These known, but different target and non-target proportions are then used to reconstruct the target and non-target class means. This approach which does not need explicit class labels is termed Learning from Label Proportions (LLP). It can be used to decode brain signals without prior calibration session. More details can be found in the article.</p> <p>In another study, the above data set was used to simulate a new unsupervised mixture approach which combines the mean estimation of the unsupervised expectation-maximization algorithm by Kindermans et al. (2012, PLoS One) with the means obtained with the LLP approach. This leads to an unsupervised solution for which the performance is as good as in the supervised scenario. Please find more details in the following article:</p> <p><em>Verhoeven, T., H&uuml;bner, D., Tangermann, M., M&uuml;ller, K. R., Dambre, J., &amp; Kindermans, P. J. (2017). Improving zero-training brain-computer interfaces by mixing model estimators. Journal of neural engineering, 14(3), 036021.</em></p> </blockquote> <p>The data was recorded with BrainVision recorder. A new file was recorded for every group of 7 characters. The .eeg file contains the RAW EEG data in the format as described in the .vhdr file. Events / stimuli markers are provided in the .vmrk files. Note that there is a wrapper available to use this data in MOABB here: TODO INSERT LINK</p> <p>The subjects had the task to spell a specific sentence with 63 letters. In the online experiment, this was repeated 3 times and each time the online unsupervised classifier was reset at the start of the sentence.</p>

opencc-by-4.0Jan 2022View details →
zenodo44/100

A Weakly-Labeled Stance Dataset during the 2019 South American Protests

<p>Research across different disciplines has documented the expanding polarization in social media. However, much of it focused on the US political system or its culturally controversial topics. In this work, we explore polarization on Twitter in a different context, namely the protest that paralyzed several countries in the South American region in 2019. By leveraging users&rsquo; endorsement of politicians&#39; tweets and hashtag campaigns with defined stances towards the government of each country (for or against), we construct a weakly labeled stance dataset with hundreds of thousands of users. Moreover, through the synergistic usage of network-focused methods applied on news sharing patterns and language-focused methods, we validate our labeling methodology by showing that these stances partition the users into meaningful communities. That is, we show that polarization in users&#39; news sharing patterns was consistent with their stances towards the government and that polarization in their language mainly manifested along ideological, political, or protest-related lines.</p>

opencc-by-4.0Apr 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record