Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,399

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,399 results for “Manual”

Learn how ShareScore rates datasets ↗
edi56/100

Manually-collected discharge data for multiple inflow and outflow tributaries at Falling Creek Reservoir, Beaverdam Reservoir, and Carvins Cove Reservoir, Virginia, USA from 2019-2025

Discharge rates at multiple inflow streams into Falling Creek Reservoir (Vinton, Virginia, USA), Beaverdam Reservoir (Vinton, Virginia, USA), and Carvins Cove Reservoir (Roanoke, Virginia, USA), and one outflow at Falling Creek Reservoir were measured manually using multiple methods from 2019-2025. Falling Creek Reservoir, Beaverdam Reservoir, and Carvins Cove Reservoir are owned and operated by the Western Virginia Water Authority as drinking water sources for Roanoke, Virginia. The dataset consists of discharge rates calculated using one of four methods: handheld flowmeter, salt injection, velocity float, or bucket method. Data were collected weekly to monthly from February through October 2019 at Falling Creek and Beaverdam Reservoir, and approximately weekly to seasonally at Falling Creek and Carvins Cove from 2020 to 2025. The dataset is accompanied by a maintenance log and quality assurance/quality control analysis scripts.

openCC (other)Jan 2026View details →
zenodo52/100

BioASQ-QA: A manually curated corpus for Biomedical Question Answering

<p>The BioASQ question answering (QA) benchmark dataset contains questions in English, along with golden standard (reference) answers and related material. The dataset has been designed to reflect real information needs of biomedical experts and is therefore more realistic and challenging than most existing datasets. Furthermore, unlike most previous QA benchmarks that contain only exact answers, the BioASQ-QA dataset also includes ideal answers (in effect summaries), which are particularly useful for research on multi-document summarization. The dataset combines structured and unstructured data. The material linked with each question comprise documents and snippets, which are useful for Information Retrieval and Passage Retrieval experiments, as well as concepts that are useful in concept-to-text Natural Language Generation. Researchers working on paraphrasing and textual entailment can also measure the degree to which their methods improve the performance of biomedical QA systems. Last but not least, the dataset is continuously extended, as the BioASQ challenge is running and new data are generated.</p>

opencc-by-2.5Dec 2022View details →
zenodo48/100

An annotated high-content fluorescence microscopy dataset with EGFP-Galectin-3-stained cells and manually labelled outlines

<p>Here we present a benchmarking dataset of fluorescence microscopy images with EGFP-Galectin-3-stained cells together with annotations of their outlines. Images were randomly selected from an RNA interference screen with a modified U2OS osteosarcoma cell line, acquired on a Thermo Fischer CX7 high-content imaging system at 20x magnification.&nbsp;</p> <p>The dataset contains 60 images showing over 2000 labelled nuclear objects in total, which is sufficiently large to train well-performing neural networks for instance or semantic segmentation. It is pre-split into training, development and test set, each in a zip file. The dataset should be referred to as Aitslab_bioimaging2.</p> <p>For most of the images, nuclear staining and annotations have been published previously in the dataset Aitslab_bioimaging1 (https://doi.org/10.5281/zenodo.6657260). The conversion script to produce the png images from the C01 images was published together with this dataset.</p>

opencc-by-4.0Jun 2024View details →
zenodo48/100

Oráculo manual. Schopenhauer's marginalia (TEI-XML)

<p>XML-TEI file used in the project&nbsp;<em>Schopenhauer&#39;s Library.&nbsp;Annotations and marks in his Spanish books</em>&nbsp;<a href="http://schopenhauer.uni.wroc.pl">http://schopenhauer.uni.wroc.pl</a></p> <p>&nbsp;</p>

opencc-by-4.0Mar 2018View details →
zenodo48/100

Multi-Dimensional Data Viewer (MDV) user manual for data exploration: "Systematic analysis of YFP traps reveals common discordance between mRNA and protein across the nervous system"

<table> <tbody> <tr> <td> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp;Please also see the latest version of the repository:<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp;<a href="https://doi.org/10.5281/zenodo.6374011">https://doi.org/10.5281/zenodo.6374011</a> and<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp;our website: <a href="https://ilandavis.com/jcb2023-yfp">https://ilandavis.com/jcb2023-yfp</a></p> </td> </tr> </tbody> </table> <p>&nbsp;</p> <p>The explosion in the volume of biological imaging data challenges the available technologies for data interrogation and its intersection with related published bioinformatics data sets. Moreover, intersection of highly rich and complex datasets from different sources provided as flat csv files requires advanced informatics skills, which is time consuming and not accessible to all. &nbsp;Here, we provide a &ldquo;user manual&rdquo; to our new paradigm for systematically filtering and analysing a dataset with more than 1300 microscopy data figures using Multi-Dimensional Viewer (MDV) -<a href="https://mdv.molbiol.ox.ac.uk/projects/mdv_project/7012?view=RNA+%2F+Protein+Distribution">link</a>, a solution for interactive multimodal data visualisation and exploration. The primary data we use are derived from our published systematic analysis of 200 YFP traps reveals common discordance between mRNA and protein across the nervous system (<a href="https://doi.org/10.1083/jcb.202205129">eprint link</a>). This manual provides the raw image data together with the expert annotations of the mRNA and protein distribution as well as associated bioinformatics data. We provide an explanation, with specific examples, of how to use MDV to make the multiple data types interoperable and explore them together. We also provide the open-source python code <a href="https://github.com/ilandavislab/Annotate.OMERO.Fig">(github link)</a> used to annotate the figures, which could be adapted to any other kind of data annotation task.</p>

opencc-by-4.0Mar 2023View details →
zenodo48/100

A pangenome-guided manually curated library of transposable elements for Zymoseptoria tritici

<p>A manually-curated TE consensus library generated using a panel of 19 reference genomes for&nbsp;<em>Zymoseptoria tritici</em><sup>1-3</sup>&nbsp;along with reference genome assemblies for the sister species&nbsp;<em>Z. ardabiliae</em>,&nbsp;<em>Z. brevis</em>,&nbsp;<em>Z. pseudotritici</em>, and&nbsp;<em>Z. passerinii<sup>4</sup></em>.&nbsp;</p> <p>&nbsp;</p> <p><strong>Methods</strong></p> <p>Putative TE consensus sequences were first obtained by annotating all 23 genome assemblies<sup>1&ndash;4</sup>&nbsp;with Earl Grey with default settings (v3.0;&nbsp;<a href="https://github.com/TobyBaril/EarlGrey">https://github.com/TobyBaril/EarlGrey</a>)<sup>5,6</sup>. Consensus sequences generated from each reference genome were clustered using CD-Hit-Est (v4.8.1)<sup>7,8</sup>&nbsp;to group sequences with 90% similarity across 80% of the longer sequence length (<em>-n 8 -d 0 -aL 0.8 -c 0.90 -G 0 -g 1 -b 500 -r 1</em>)&nbsp;to reduce redundancy whilst preventing the collapsing of chimeric sequences. Consensus sequences &lt;100bp were removed, as these are unlikely to represent true TE sequences. Each consensus sequence was then subject to manual curation as described by Goubert et al. (2022)<sup>9</sup>. Briefly, genomic copies of each TE were obtained using a &ldquo;BLAST, Extract, Extend&rdquo; process to recover genomic copies from each of the 23 reference genome assemblies with 1,000 flanking bases at either end<sup>9,10</sup>. For families with &gt;100 BLASTN hits, the 25 longest hits were selected, along with 75 random hits. Multiple alignments were generated for each putative TE family using MAFFT (v7.505) with the --auto flag<sup>11</sup>. Columns composed of &gt;=80% gaps were removed with T-COFFEE (v13.45.0.4846264)<sup>12</sup>. Subsequently, all sequence alignments were manually curated to define TE boundaries and remove regions of low conservation and rare insertions. Following manual curation, new majority-rule consensus sequences were generated with EMBOSS (v6.6.0.0) cons<sup>13</sup>. TE-Aid (<a href="https://github.com/clemgoub/TE-Aid/">https://github.com/clemgoub/TE-Aid/</a>) was used to aid visual inspection and to identify diagnostic features for classification of extended consensus sequences. Following this, TIRs were recorded if present, and nhmmscan (HMMER v3.3.2)<sup>14</sup>&nbsp;was used to identify homology to known curated elements in Dfam (v3.7). Combining this information, each TE consensus sequence was manually classified using available information following the naming convention &lsquo;&gt;ZymTri_2023_family_[n]#[Classification]/[Family]&rsquo;. Consensus sequences classified with low confidence have a &lsquo;?&rsquo; added to the name, as well as the string &lsquo;_LowConf&rsquo;. To reduce redundancy in the final TE library, sequences were clustered to the family-level using the 80-80-80 rule implemented in CD-hit-est<sup>9,15 </sup>(<em>-d 0 -aS 0.8 -c 0.8 -G 0 -g 1 -b 500 -r 1</em>). The representative sequence for each cluster was manually selected to select the sequence with the highest classification confidence, also defined as the &lsquo;most intact consensus&rsquo;. Chimeric sequences erroneously clustered were manually separated to retain sequences for the chimeric TE and the individual elements that generated the chimer.</p> <p>&nbsp;</p> <p><strong>References</strong></p> <p>1.&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;Badet, T., Oggenfuss, U., Abraham, L., McDonald, B. A. &amp; Croll, D. A 19-isolate reference-quality global pangenome for the fungal wheat pathogen Zymoseptoria tritici.&nbsp;<em>BMC Biol.</em>&nbsp;<strong>18</strong>, 12 (2020).</p> <p>2.&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;Goodwin, S. B.&nbsp;<em>et al.</em>&nbsp;Finished genome of the fungal wheat pathogen Mycosphaerella graminicola reveals dispensome structure, chromosome plasticity, and stealth pathogenesis.&nbsp;<em>PLoS Genet.</em>&nbsp;<strong>7</strong>, e1002070 (2011).</p> <p>3.&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;Plissonneau, C., Hartmann, F. E. &amp; Croll, D. Pangenome analyses of the wheat pathogen Zymoseptoria tritici reveal the structural basis of a highly plastic eukaryotic genome.&nbsp;<em>BMC Biol.</em>&nbsp;<strong>16</strong>, 5 (2018).</p> <p>4.&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;Feurtey, A.&nbsp;<em>et al.</em>&nbsp;Genome compartmentalization predates species divergence in the plant pathogen genus Zymoseptoria.&nbsp;<em>BMC Genomics</em>&nbsp;<strong>21</strong>, 588 (2020).</p> <p>5.&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;Baril, T., Imrie, R. M. &amp; Hayward, A. Earl Grey: a fully automated user-friendly transposable element annotation and analysis pipeline. (2022) doi:10.21203/rs.3.rs-1812599/v1.</p> <p>6.&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;Baril, T., Galbraith, J. &amp; Hayward, A.&nbsp;<em>Earl Grey</em>. (Zenodo, 2023). doi:10.5281/ZENODO.8116025.</p> <p>7.&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;Li, W. &amp; Godzik, A. Cd-hit: a fast program for clustering and comparing large sets of protein or nucleotide sequences.&nbsp;<em>Bioinformatics</em>&nbsp;<strong>22</strong>, 1658&ndash;1659 (2006).</p> <p>8.&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;Fu, L., Niu, B., Zhu, Z., Wu, S. &amp; Li, W. CD-HIT: accelerated for clustering the next-generation sequencing data.&nbsp;<em>Bioinformatics</em>&nbsp;<strong>28</strong>, 3150&ndash;3152 (2012).</p> <p>9.&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;Goubert, C.&nbsp;<em>et al.</em>&nbsp;A beginner&rsquo;s guide to manual curation of transposable elements.&nbsp;<em>Mob. DNA</em>&nbsp;<strong>13</strong>, 7 (2022).</p> <p>10.&nbsp;&nbsp;&nbsp;Camacho, C.&nbsp;<em>et al.</em>&nbsp;BLAST+: Architecture and applications.&nbsp;<em>BMC Bioinformatics</em>&nbsp;<strong>10</strong>, 1&ndash;9 (2009).</p> <p>11.&nbsp;&nbsp;&nbsp;Katoh, K. &amp; Standley, D. M. MAFFT multiple sequence alignment software version 7: Improvements in performance and usability.&nbsp;<em>Mol. Biol. Evol.</em>&nbsp;<strong>30</strong>, 772&ndash;780 (2013).</p> <p>12.&nbsp;&nbsp;&nbsp;Notredame, C., Higgins, D. G. &amp; Heringa, J. T-coffee: a novel method for fast and accurate multiple sequence alignment.&nbsp;<em>J. Mol. Biol.</em>&nbsp;<strong>302</strong>, 205&ndash;217 (2000).</p> <p>13.&nbsp;&nbsp;&nbsp;Rice, P., Longden, L. &amp; Bleasby, A. EMBOSS: The European Molecular Biology Open Software Suite.&nbsp;<em>Trends Genet.</em>&nbsp;<strong>16</strong>, 276&ndash;277 (2000).</p> <p>14.&nbsp;&nbsp;&nbsp;Wheeler, T. J. &amp; Eddy, S. R. nhmmer: DNA homology search with profile HMMs.&nbsp;<em>Bioinformatics</em>&nbsp;<strong>29</strong>, 2487&ndash;2489 (2013).</p> <p>15.&nbsp;&nbsp;&nbsp;Wicker, T.&nbsp;<em>et al.</em>&nbsp;A unified classification system for eukaryotic transposable elements.&nbsp;<em>Nat. Rev. Genet.</em>&nbsp;<strong>8</strong>, 973&ndash;982 (2007).</p>

opencc-by-4.0Sep 2023View details →
edi48/100

Eight Mile Lake Research Watershed, Carbon in Permafrost Experimental Heating Research (CiPEHR): CiPEHR snow depth manual data 2009-2025

The Carbon in Permafrost Experimental Heating Research (CiPEHR) project addresses the following questions: 1) Does ecosystem warming cause a net release of C from the ecosystem to the atmosphere?, 2) Does the decomposition of old C, that comprises the bulk of the soil C pool, influence ecosystem C loss?, and 3) How do winter and summer warming alone, and in combination, affect ecosystem C exchange? We are answering these questions using a combination of field and laboratory experiments to measure ecosystem carbon balance and radiocarbon isotope ratios at a warming experiment located in an upland tundra field site near Healy, Alaska in the foothills of the Alaska Range. This data set includes manual measurements of snow depth collected in early spring on winter warming and control treatment plots.

openOpenOct 2025View details →
zenodo44/100

A Data Set of 255,000 Randomly Selected and Manually Classified Extracted Ion Chromatograms for Evaluation of Peak Detection Methods

<p>Non-targeted mass spectrometry (MS) has become an important method over the last years in the fields of metabolomics and environmental research. While more and more algorithms and workflows become available to process a large number of data sets nontargeted, there still exist few manually evaluated universal test data sets for refining and evaluating these methods. The first step of non-targeted screening, peak detection (and refinement of it) is arguably the most important step for non-targeted screening. However, the absence of a model data set makes it harder for researchers to evaluate peak detection methods. In this Data Descriptor, we provide a manually checked data set consisting of 255,000 EICs (5000 peaks randomly sampled from across 51 samples) for the evaluation on peak detection and gap filling algorithms. The data set was created from a previous real-world study, of which a subset was used to extract and manually classify ion chromatograms by three mass spectrometry experts. The data set consists of:</p> <ul> <li>51 converted mass spectral files in mzML format</li> <li>An .RData-file containing the extracted ion chromtograms (EICs)</li> <li>The randomly selected subset and the original output table of MZmine in .csv-format</li> <li>Example .xlsx files for the classification</li> <li>2 central classification tables</li> <li>Several tables with additional information about the sampling, chemical analysis and expert jugdement on EICs</li> </ul> <p>For a full description of the experiment and the data set, please read the related Data Descriptor with the title &quot;A data set of 255000 randomly selected and manually classified extracted ion chromatograms for evaluation of peak detection methods&quot; in Metabolites (https://www.mdpi.com/journal/metabolites; DOI: https://doi.org/10.3390/metabo10040162).</p>

opencc-by-4.0Feb 2020View details →
zenodo44/100

Manually Annotated Instances of Ich ('I') from the German KoLas Corpus

<p>Dataset used in Andresen/Knorr (2020). The dataset comprises 360 instances of <em>ich</em> (&#39;I&#39;) taken from the German learner corpus KoLaS (Andresen/Knorr 2017, see <a href="http://hdl.handle.net/11022/0000-0001-B732-8">http://hdl.handle.net/11022/0000-0001-B732-8</a> for full corpus access) and manually annotated with categories taken from Steinhoff (2007).</p> <p>Column descriptions:</p> <ul> <li>document: name of the document by which it can be found in the KoLaS corpus</li> <li>code_annotator1 - code_annotator4: Annotations by four annotators. Possible values: Verfasser-<em>Ich</em> (author <em>I</em>), Forscher-<em>Ich</em> (researcher <em>I</em>), Erz&auml;hler-<em>Ich</em> (narrator <em>I</em>)</li> <li>max_agreement_freq: Highest number of anntators that agreed on one label</li> <li>max_agreement_label: Label on which the highest number of annotators agreed</li> <li>context_before: 150 characters of context before the match</li> <li>match: the match itself (either <em>ich</em> or <em>Ich</em>)</li> <li>context_after: 150 characters of context after the match</li> </ul> <p><strong>References</strong></p> <p>Andresen M, Knorr D. KoLaS &ndash; Ein Lernendenkorpus in der Schreibberatungsausbildung einsetzen. <em>Zeitschrift Schreiben</em>. Published online July 5, 2017:10-17.</p> <p>Andresen M, Knorr D. Exploring the Use of the Pronoun I in German Academic Texts with Machine Learning. In: Burghardt M, M&uuml;ller-Birn C, eds. <em>Methoden und Anwendungen der Computational Humanities</em>. Lecture Notes in Informatics (LNI). Gesellschaft f&uuml;r Informatik; 2020.</p> <p>Steinhoff T. Zum ich-Gebrauch in Wissenschaftstexten. <em>Zeitschrift f&uuml;r germanistische Linguistik</em>. 2007;35(1-2):1&ndash;26.</p>

opencc-by-4.0Aug 2020View details →
zenodo44/100

African Swine Fever Worldwide Epidemiology Data - OIE Webscrape example - Geocoded using Google API and Manual

<p>Example African Swine Fever dataset generated by programs described in following publication&nbsp;</p> <p>Title: Web-scraping programmatic techniques in aggregating difficult to access OIE WAHIS animal disease outbreak information; using African Swine Fever in Europe as an example.</p> <p>Short running title: Methods for web-scraping OIE WAHIS data.</p> <p>Abstract: This study describes and makes available new methods for acquiring difficult to access, publicly available, disease surveillance data. It uses World Organisation for Animal Heath (OIE) data on African Swine Fever (ASF) outbreaks in Belarus and its neighbouring European countries to showcase the importance of adequate disease surveillance data to inform decision-making. The data acquired from these methods allow for large-scale, geospatial outbreak mapping and summary statistics of any terrestrial disease listed on the OIE World Animal Health Information System (WAHIS) database. These techniques will make important epidemiological data more accessible to the scientific community and aid in gaining further insight into the occurrence and spread of OIE listed diseases in a timely manner, fulfilling an important function of disease surveillance.</p>

opencc-by-4.0Nov 2020View details →
zenodo44/100

Manually Annotated Drone Imagery (RGB) Dataset for automatic coastline delineation of Southern Baltic Sea, Poland with polyline annotations (0.1.1)

<p><strong>Overview:</strong></p> <p>The &nbsp;Manually Annotated Drone Imagery Dataset (MADRID) consists of hand annotated high resolution RGB images taken in two different types of coasts in Poland, Miedzyzdroje - cliff coast and in Mrzezyno - dune coast in 2022-2023. All images were converted into a uniform format of 1440x2560 pixels, polyline annotated and set into file structure format suited for semantic segmentation tasks (See "Usage" notes below for more details).</p> <p>The raw images of our dataset were captured Zenmuse L1 Sensor (RGB) mounted on a DJI Matrice 300 RTK Drone. Total of 4895 images were captured, however the dataset contains 3876 images with each image annotated with coastline. The dataset only include images with coastlines that are visually identifiable with the human eye. For the annotations of the images, CVAT v2.13 open-source software was utilized.</p> <p><strong>Usage:</strong></p> <p>The compressed RAR file contains two folders train and test. Each folder contains the file that represents the date at which the image was captured in the format of (year, month, day), number of the image and the name of the drone utilized to capture the image. For example, DJI_20220111140051_0051_Zenmuse-L1-mission and DJI_20220111140105_0053_Zenmuse-L1-mission. Additionally, the test folder contains annotations (one per image) which are extracted from the original XML annotation file provided in the CVAT 1.1 image format.</p> <p>Archives were compressed using RAR compression. They can be decompressed in a terminal by opening and extracting Madrid_v0.1_Data.zip.</p> <p>The subset of the data with the name Madrid_subset_data.zip has been added which contains a small portion of train and test images for purpose of inspecting the dataset without downloading the entire dataset.</p> <p>The training images for both training data and testing data are structured as follows.</p> <pre><code>Train/ └── images/ └── DJI_20220111140051_0051_Zenmuse-L1-mission.JPG └── DJI_20220111140105_0053_Zenmuse-L1-mission.JPG └── ...<br>└── masks/ └── DJI_20220111140051_0051_Zenmuse-L1-mission.PNG └── DJI_20220111140105_0053_Zenmuse-L1-mission.PNG └── ...<br><br>Test/ └── images/ └── DJI_20220111140051_0051_Zenmuse-L1-mission.JPG └── DJI_20220111140105_0053_Zenmuse-L1-mission.JPG └── ...<br>└── masks/ └── DJI_20220111140051_0051_Zenmuse-L1-mission.PNG └── DJI_20220111140105_0053_Zenmuse-L1-mission.PNG └── ...</code></pre> <p>&nbsp;</p>

opencc-by-4.0May 2024View details →
zenodo44/100

Manually curated transcriptomics data collection for toxicogenomic assessment of engineered nanomaterials

<p>Toxicogenomics (TGx) approaches are increasingly applied to gain insight into the possible toxicity mechanisms of engineered nanomaterials (ENMs). Omics data can be valuable to elucidate the mechanism of action of chemicals and develop predictive models in toxicology. While vast amounts of transcriptomics data from ENM exposures have already been accumulated, a unified, easily accessible and reusable collection of transcriptomics data for ENMs is currently lacking. In an attempt to improve the FAIRness of already existing transcriptomics data for nanomaterials, we curated a collection of homogenized transcriptomics data from human, mouse and rat ENM exposures <em>in vitro</em> and <em>in vivo</em>.</p>

opencc-by-4.0Jul 2020View details →
zenodo44/100

Towards a systematic approach to manual annotation of code smells - C# Dataset of Long Method and Large Class code smells

<p>This dataset includes open-source projects written in C# programing language, annotated for the presence of Long Method and God Class code smells. Each instance was manually annotated by at least two annotators.&nbsp;We explain our motivation and methodology for creating this dataset in our <a href="https://www.techrxiv.org/articles/preprint/Towards_a_systematic_approach_to_manual_annotation_of_code_smells/14159183/1">preprint</a>:</p> <p>Luburić, N., Prokić, S., Grujić, K.G., Slivka, J., Kovačević, A., Sladić, G. and Vidaković, D., 2021. Towards a systematic approach to manual annotation of code smells.&nbsp;</p> <p>The dataset contains two excel datasheets:</p> <ul> <li><em>DataSet_Large Class.xlsx</em> &ndash; C# classes annotated for the Large Class code smell severity.</li> <li><em>DataSet_Long Method.xlsx</em> &ndash; C# methods annotated for the Long method code smell severity.</li> </ul> <p>&nbsp;The columns in the datasheet represent:</p> <ul> <li><em>Code Snippet ID</em> &ndash; the full name of the code snippet.&nbsp; <ul> <li>For classes, this is the package/namespace name followed by the class name. The full name of inner classes also contains the names of any outer classes (e.g., <em>namespace.subnamespace.outerclass.innerclass</em>).</li> <li>For methods, this is the full name of the class and the methods&rsquo;s signature (e.g., <em>namespace.class.method(param1Type, param2Type)</em> ).</li> </ul> </li> <li><em>Link </em>&ndash; The GitHub link to the code snippet, including the commit and the start and end LOC.</li> <li><em>Code Smell </em>&ndash; code smell for which the code snippet is examined (Large Class or Long Method).</li> <li><em>Project Link </em>&ndash; the link to the version of the code repository that was annotated.</li> <li><em>Metrics </em>&ndash; a list of metrics for the code snippet, calculated by our <a href="https://github.com/Clean-CaDET/platform#readme">platform</a>. Our dataset provides 25 class-level metrics for Large Class detection and 18 method-level metrics for Long Method detection The list of metrics and their definitions is available <a href="https://github.com/Clean-CaDET/platform/blob/c4acff95ec00ff6c25fa62dde4818c1f40e39d39/CodeModel/CaDETModel/CodeItems/CaDETMetrics.cs">here</a>.</li> <li><em>Final annotation </em>&ndash; a single severity score calculated by a majority vote.&nbsp;</li> <li><em>Annotators </em>&ndash; each annotator&#39;s (1, 2, or 3) assigned severity score.</li> </ul> <p>To help guide their reasoning for evaluating the presence and the severity of a code smell, three annotators independently annotated whether the considered heuristics apply to an evaluated code snippet. We provide these results in two separate excel datasheets:</p> <ul> <li><em>LargeClass_Heuristics.xlsx </em>- C# classes annotated for the presence of heuristics relevant for the Large Class code smell.</li> <li><em>LongMethod_Heuristics.xlsx </em>- C# classes annotated for the presence of heuristics relevant for the Large Class code smell.</li> </ul> <p>The columns of these two datasheets are:</p> <ul> <li><em>Code Snippet ID </em>- the full name of the code snippet (matching the IDs from <em>DataSet_Large Class.xlsx </em>and <em>DataSet_Long Method.xlsx</em>)</li> <li><em>Annotators</em> &ndash; heuristics labelled by each of the annotators (1, 2, or 3).</li> <li><em>Heuristics </em>&ndash; whether the heuristic is applicable to the examined code snippet or not (Section 1.2.4 lists heuristics relevant for the Large Class detection, and Section 1.2.5 lists the heuristics relevant for the Long Method detection).</li> </ul>

opencc-by-4.0May 2022View details →
zenodo44/100

An annotated high-content fluorescence microscopy dataset with Hoechst 33342-stained nuclei and manually labelled outlines

<p>Here we present a benchmarking dataset of fluorescence microscopy images with Hoechst 33342-stained nuclei together with annotations of nuclei, nuclear fragments and micronuclei. Images were randomly selected from an RNA interference screen with a modified U2OS osteosarcoma cell line, acquired on a Thermo Fischer CX7 high-content imaging system at 20x magnification. Labelling was performed by a single annotator and reviewed by a biomedical expert.</p> <p>The dataset contains 50 images showing over 2000 labelled nuclear objects in total, which is sufficiently large to train well-performing neural networks for instance or semantic segmentation. It is pre-split into training, development and test set, each in a zip file. The dataset should be referred to as Aitslab_bioimaging1. A brief article describing the dataset is also available (Arvidsson M, Kazemi Rashed S, Aits S. <a href="https://doi.org/10.1016/j.dib.2022.108769">10.1016/j.dib.2022.108769</a> )</p> <p><strong>Dataset description:</strong></p> <p>Fluorescence microscopy images: original .C01 files and files converted to 8-bit .png format (Grayscale)</p> <p>Annotations: 24-bit .png format (RGB)</p> <p>Script used to convert C01 to png images:&nbsp;C01_to_png.py file with python code and readme.md file with instructions to run it</p>

opencc-by-4.0Jun 2022View details →
zenodo44/100

EukRibo: a manually curated eukaryotic 18S rDNA reference database

<p>EukRibo is a manually curated database of reference small-subunit ribosomal RNA gene (18S rDNA) sequences of eukaryotes, specifically aimed at taxonomic annotation of high-throughput metabarcoding datasets. Unlike other reference databases of ribosomal genes, it is not meant to exhaustively capture all publicly available 18S rDNA sequences from the INSDC repositories, but to represent a subset of highly trustable sequences covering the whole known diversity of eukaryotes, with a focus on protists, manually verified taxonomic identifications, and relatively low genetic redundancy.</p> <p>EukRibo is part of a suite of public resources generated by the UniEuk project (www.unieuk.org), which are all designed to follow a common taxonomic framework for maximal interoperability. The high level of taxonomic accuracy of EukRibo, together with a newly designed, phylogenetically-informed annotation approach, allow high confidence in the taxonomic annotation of environmental metabarcodes, as well as identification of new eukaryotic diversity at various taxonomic levels using a connected components approach.</p> <p>*&nbsp;&nbsp; *&nbsp;&nbsp; *</p> <p>Accompanying preprint available at <a href="https://doi.org/10.1101/2022.11.03.515105">https://doi.org/10.1101/2022.11.03.515105</a>.</p> <p>*&nbsp;&nbsp; *&nbsp;&nbsp; *</p> <p><strong>EukRibo ReadMe file, versions 1 and 2</strong></p> <p>Each EukRibo release consists of <strong>4 files</strong>:<br> - a <strong>tsv table </strong>containing the taxonomic and other information about the 18S rDNA sequences included in the release<br> - a <strong>fasta file </strong>containing the <strong>full sequences </strong>as retrieved from the INSDC repositories (NCBI, EMBL-EBI/ENA, DDBJ)<br> - a <strong>fasta file </strong>containing the <strong>variable region V4 </strong>extracted from all these sequences (based on the fragment amplified with the Tara-Oceans V4 primers)<br> - a <strong>fasta file </strong>containing the <strong>variable region V9 </strong>extracted from the subset of sequences where it is present (based on the fragment amplified with the Tara-Oceans V9 primers)</p> <p>The primary goal of EukRibo was to be used to annotate the EukBank meta-dataset of available V4 metabarcoding datasets, and therefore all sequences included in EukRibo contain the variable region V4.<br> Only a subset of these sequences (about 75%) also contain the variable region V9; this is because many 18S rDNA sequences in the INSDC repositories stop before the V9 fragment.</p> <p>Sequences with slightly incomplete V4 or V9 fragments were kept if phylogenetically useful - i.e. if they are the only available representatives of a certain taxonomic lineage.<br> <strong>V4&nbsp;&nbsp; &nbsp;</strong>We allowed up to 50 missing positions in the relatively conserved area at the 5&#39; end of the V4 fragment (for an average fragment length of about 380 bp); no sequence incomplete at the 3&#39; end of the V4 fragment is included.<br> <strong>V9&nbsp;&nbsp; &nbsp;</strong>We allowed up to 30 missing positions in the relatively conserved area at the 3&#39; end of the V9 fragment (for an average length of about 135 bp); no sequence incomplete at the 5&#39; end of the V9 fragment is included.<br> We allowed a higher proportion of missing positions for the V9 region because being more conservative would imply losing too many sequences, including entire taxonomic lineages.</p> <p><strong>Version 1 of EukRibo</strong><br> This is the starting version of EukRibo that was used for the taxonomic annotation of the EukBank dataset, with taxonomy strings that were fixed as of October 2020.<br> - Contains 46,345 sequences with a sufficiently complete V4 region; 46,299 with the actual complete V4 region and 46 (about 0.1%) with missing positions at the 5&#39; end.<br> - Of these, 34,438 also include a sufficiently complete V9 region; 23,226 with the actual complete V9 region and 11,206 (about 33%) with missing positions at the 3&#39; end.</p> <p><strong>Version 2 of EukRibo</strong><br> This is a version of EukRibo that was made taxonomically compatible with version 3 of the EukProt database (<a href="https://doi.org/10.1101/2020.06.30.180687">https://doi.org/10.1101/2020.06.30.180687</a>), with taxonomic revisions as of July 2022 as well as additional information on the included selection of sequences that was not provided in the tsv file of version 1.<br> - Contains the exact same selection of sequences as in version 1, with the addition of genus <em>Meteora</em>, the last remaining known supergroup-level eukaryotic lineage for which an 18S rDNA was not previously available. (The <em>Meteora </em>sequence contains the full V4 fragment but does not include a sufficiently complete V9 fragment.)<br> - Only 34,432 sequences with a sufficiently complete V9 region are now retained because of 6 previously unrecognised chimeric sequences where the V9 fragment does not originate from the same organism as the V4 fragment.</p> <p><strong>Files in EukRibo version 1</strong>:<br> 46345_EukRibo.tsv.gz<br> 46345_EukRibo_full_seqs.fas.gz<br> 46345_EukRibo_V4.fas.gz<br> 34438_EukRibo_V9.fas.gz</p> <p>The tsv file contains 6 columns:<br> <strong>gb_accession </strong>- INSDC accession number of the sequence<br> <strong>supergroup</strong>, <strong>taxogroup1</strong>, <strong>taxogroup2 </strong>- binning of the taxa into strictly monophyletic clades of evolutionary and/or ecological significance<br> <strong>UniEuk_taxonomy_string </strong>- full UniEuk-compatible taxonomic annotation of the sequence<br> - an unlimited number of levels is allowed (going down to strain for isolated organisms or to clone for environmental sequences)<br> - informal names are used for phylogenetically supported clades without formal name<br> <strong>V9 </strong>- presence (&#39;Y&#39;) or absence (&#39;N&#39;) of a sufficiently complete V9 fragment in the sequence</p> <p><strong>Files in EukRibo version 2</strong>:<br> 46346_EukRibo-02.tsv.gz<br> 46346_EukRibo-02_full_seqs.fas.gz<br> 46346_EukRibo-02_V4.fas.gz<br> 34432_EukRibo-02_V9.fas.gz</p> <p>The tsv file now contains 12 columns:<br> <strong>gb_accession</strong>, <strong>supergroup</strong>, <strong>taxogroup1</strong>, <strong>taxogroup2</strong>, <strong>UniEuk_taxonomy_string</strong><br> &nbsp;&nbsp; &nbsp;- same columns as in version 1<br> <strong>alternative_strain_names </strong>(new) - provides alternative strain/isolate names when known to help cross-linking genetic data coming from the same organism<br> <strong>V4 </strong>(new) - indicates whether the V4 fragment is complete (&#39;yes - complete&#39;) or missing positions at the 5&#39; end (&#39;yes - partial&#39;)<br> <strong>V9 </strong>(emended content) - now contains more precise information than in version 1 about whether it is complete (&#39;yes - complete&#39;), missing positions at the 3&#39; end (&#39;yes - partial&#39;), or was excluded, and the 6 possible reasons why (&#39;no - missing&#39;, &#39;no - too incomplete&#39;, &#39;no - chimera&#39;, &#39;no - bad quality&#39;, &#39;no - deletion in V9&#39;, &#39;no - Ns in V9&#39;)<br> <strong>EukProt_ID_same_strain </strong>(new) - accession of EukProt datasets from the same isolate<br> <strong>EukProt_ID_different_strain </strong>(new) - accession of EukProt datasets from a different isolate of the same species<br> <strong>columns_modified_since_previous_version </strong>(new) - lists all of the 6 pre-existing columns that have a modified content compared to version 1<br> <strong>remarks </strong>(new) - additional information such as presence of an intron in the V9 fragment, taxonomic identity of the two parts of chimeric sequences, or the presence of Ns or a deletion in the V4 or the V9 fragment (but insufficient to warrant exclusion)</p>

opencc-by-4.0Mar 2022View details →
zenodo44/100

Manual in-situ measurements of snow depth and snow water equivalent at the Polish Polar Station Hornsund - winter seasons 2021/2022and 2022/2023

<p>The dataset presents manual measurements of snow depth and snow water equivalent collected at the Polish Polar Station Hornsund in Svalbard during the winter seasons of 2021/2022 and 2022/2023.</p> <p>Snow depth measurements have been conducted at the same location by the Station's overwintering personnel since August 1982. Snow depth is calculated from a mean of three snow stakes to avoid the effects of the drifting snow. Measurements are taken manualy, on a daily basis.&nbsp;</p> <p>Snow water equivalent measurements have also been carried out at the same points by the Station's overwintering crew since October 1982. These measurements are performed every five days using a VS-43 snow tube. However, measurements are not taken when the snow depth is less than 5 cm.</p>

opencc-by-4.0Sep 2024View details →
zenodo44/100

RDF version of the data from Saarimaki et al. Manually curated transcriptomics data collection for toxicogenomic assessment of engineered nanomaterials (Version 1.0.0) [Zenodo Dataset] (2020)

<p>This is an RDFied version of the dataset published by&nbsp;Saarimaki et al. Manually curated transcriptomics data collection for toxicogenomic assessment of engineered nanomaterials (Version 1.0.0) [Zebodo Dataset] (2020)</p> <p>The original dataset publication DOI:&nbsp;<a href="http://doi.org/10.5281/zenodo.4146981">http://doi.org/10.5281/zenodo.4146981</a></p> <p>The Original publication authors:&nbsp;Saarimaki, Laura Aliisa, Federico, Antonio, Lynch, Iseult, Papadiamantis, Anastasios G., Tsoumanis, Andreas, Melagraki, Georgia, Afantitis, Antreas, Serra, Angela, &amp; Greco, Dario</p>

opencc-by-4.0Nov 2021View details →
zenodo44/100

AddressGB Manual Evaluation Sample

<p>Random sample of 7,200 manually geo-coded <a href="https://doi.org/10.5255/UKDA-SN-7481-2">I-CeM</a> addresses. The sample comprises 1,000 addresses from each England and Wales census (1851, 1861, 1881, 1891, 1901, and 1911) and 200 addresses from each Scottish census (1851, 1861, 1871, 1881, 1891, and 1901).</p> <p>Addresses have been linked to two sources of geo-coded address data: GB1900 and OS Open Roads.</p> <p>GB1900 contains transcriptions of text labels from the Second Edition County Series six-inch-to-one-mile maps covering the whole of Great Britain, published by the Ordnance Survey between 1888 and 1914. To obtain GB1900 dataset, visit the <a href="http://www.visionofbritain.org.uk/data/#tabgb1900">GB1900 website</a>.</p> <p>OS Open Roads is the Ordnance Survey's Open access modern road vector data. To obtain the OS Open Roads dataset, visit the <a href="https://www.ordnancesurvey.co.uk/business-government/products/open-map-roads">OS Website</a>.</p>

opencc-by-4.0Nov 2024View details →
zenodo44/100

TF-Marker: A comprehensive manually curated database for transcription factors and related markers in specific cell and tissue types in human.

<p>Here, we developed the TF-Marker database (TF-Marker, http://bio.liclab.net/TF-Marker/) which is committed to a comprehensive manual curation of TFs and related markers with experimental evidence in specific cell and tissue types in human. Currently, through reviewing <strong>2,091</strong> published literature, we have manually classified TFs and related markers into five types according to their functions: 1) <strong>TF</strong>: TFs, which regulate the expression of markers; 2) <strong>T Marker</strong>: markers, which are regulated by TFs (TF and T Marker pairs can identify cell types more specifically); 3) <strong>I Marker</strong>: markers, which influence the activity of TFs (I Markers can also influence the development of specific cells and tissues); 4) <strong>TFMarker</strong>: TFs, which play roles as markers (TFMarkers are cell/tissue-specific TFs used as cell markers in biology experiments); and 5) <strong>TF Pmarker</strong>: TFs, which play roles as potential markers. By curating thousands of published literature, <strong>5,905</strong> entries including <strong>1,316</strong> TFs, <strong>1,092</strong> T Markers, <strong>473</strong> I Markers, <strong>1,600</strong> TFMarkers and <strong>1,424</strong> TF Pmarkers, were annotated in <strong>383</strong> cell types and <strong>95</strong> tissue types in human. Moreover, TF-Marker divided markers into disease markers and tissue/cell-specific markers. TF-Marker is an elaborate database, which provides TFs and related markers supported by experimental evidence. We believe TF-Marker will provide strong support for research into cell/tissue-specific TFs and related markers.</p>

opencc-by-4.0Oct 2021View details →
zenodo44/100

Semi-automatic and manual shallow landslide inventories of two extreme rainfall events.

<p>This dataset contains the polygons of automatic ( PL) and manually (ML)&nbsp;&nbsp;based shallow landslides related to two extreme rainfall events. In KML format, the dataset can be visualized on GIS software or&nbsp;&nbsp;Google Earth.</p><p>With more details, it is possible to find:</p><ul><li>AOI_2016: The study area of the extreme rainfall of November 2016,&nbsp; Tanerello and Arroscia Valleys NW Italy.</li><li>The&nbsp; 2016_PL:&nbsp; The inventory of potential shallow landslides semi-automatically&nbsp;&nbsp;mapped on the base of Sentinel-2 images&nbsp;&nbsp;related to extreme rainfall events that hit NW Italy in November 2016</li><li>The&nbsp; 2016_ML:&nbsp; The inventory of shallow landslides manually mapped on high-resolution images of Google Earth, related to extreme rainfall events that hit NW Italy in November 2016</li><li>AOI_2019_large: The study area of the extreme rainfall of October&nbsp;2019&nbsp;&nbsp;NW Italy.</li><li>AOI_2019: The testing&nbsp;area of the extreme rainfall of October&nbsp;2019,&nbsp;Gavi Area&nbsp;NW Italy.</li><li>The&nbsp; 2019_PL_all: The inventory of potential shallow landslides semi-automatically&nbsp;&nbsp;mapped on the base of Sentinel-2 images&nbsp;&nbsp;related to extreme rainfall events that hit NW Italy in October 2019 (whole Study&nbsp;area)</li><li>The&nbsp; 2019_PL:&nbsp; The inventory of potential shallow landslides semi-automatically&nbsp;&nbsp;mapped on the base of Sentinel-2 images&nbsp;&nbsp;related to extreme rainfall events that hit NW Italy in October 2019 (Gavi test area)</li><li>The&nbsp; 2019_ML: The inventory of shallow landslides manually mapped on high-resolution images of Google Earth, related to extreme rainfall events that hit NW Italy in October 2019</li></ul><p>GEE_Script: A list of codes used in Google Earth Engine to produce NDVI time series or averaged NDVI on some sample studied areas are reported in the attached PDF.&nbsp; The code may be pasted and copied to the Google Earth Engine console.&nbsp;</p><p>The codes (if an account on &nbsp;Google Earth Engine is active) may be reached directly from the following URLs:&nbsp;</p><p><strong>Script 1. </strong>NDVI time series of some sampled areas to select the best pair of images for the PL creation (Tanarello and Arroscia Valley and GAVI AOIs; Fig. 16 of the paper). Link to GEE: <a href="https://code.earthengine.google.com/998af951fcb74519589bf8e722bb30b0?noload=true">https://code.earthengine.google.com/998af951fcb74519589bf8e722bb30b0?noload=true</a></p><p><strong>Script 2.</strong> sampled NDVI time series from different intersection cases for the Tanarello and Arroscia Valley study area (2016&nbsp; Event). Link to&nbsp; GEE: <a href="https://code.earthengine.google.com/b622cb64f90771ced78ef73bad9cc50f?noload=true">https://code.earthengine.google.com/b622cb64f90771ced78ef73bad9cc50f?noload=true</a></p><p><strong>Script 3.&nbsp;</strong>Sampled NDVI time series from different land-use cases for the Gavi study area (2019&nbsp; Event). Link to&nbsp; GEE: <a href="https://code.earthengine.google.com/f686c60b78a3dee0b2a2c94a259ccff2?noload=true">https://code.earthengine.google.com/f686c60b78a3dee0b2a2c94a259ccff2?noload=true</a></p><p><strong>Script 4.</strong> Multi-temporal-averaged NDVIvar &nbsp;&nbsp; Link to GEE Script: &nbsp;<a href="https://code.earthengine.google.com/bfc2e570bb675372c4c482eef682be4a?noload=true">https://code.earthengine.google.com/bfc2e570bb675372c4c482eef682be4a?noload=true</a>&nbsp;for the whole Gavi study area (2019 flood) and&nbsp; &nbsp;<a href="https://code.earthengine.google.com/89e1c0a1361860cd407b7e6ab8bb95de?noload=true">https://code.earthengine.google.com/a3390b262cef1b5f42837c88d8791b5b?noload=true</a>&nbsp;for the entire Arroscia-Tanarello study area</p><p>The full description of the methodology can be found in the paper of&nbsp; Notti et al., 2023</p><p>Notti, D., Cignetti, M., Godone, D., and Giordan, D.: Semi-automatic mapping of shallow landslides using free Sentinel-2 images and Google Earth Engine, Nat. Hazards Earth Syst. Sci., 23, 2625–2648, <a href="https://doi.org/10.5194/nhess-23-2625-2023">https://doi.org/10.5194/nhess-23-2625-2023</a>, 2023</p>

opencc-by-4.0Jun 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record