Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

2,129

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

2,129 results for “scores”

Learn how ShareScore rates datasets ↗
zenodo44/100

Bird plumage brightness scores and blood parasite prevalence values of North American passerine species

<p>Dataset with bird plumage brightness scores and blood parasite prevalence values for 114&nbsp;North American passerine host species. One file contains the data table. One file contains a table with descriptions of the columns in the data table.</p> <p>Note: These data were reconstructed from files used in Read &amp; Harvey 1989 (<a href="https://doi.org/10.1038/339618a0">https://doi.org/10.1038/339618a0</a>) with column headings inferred with the help of&nbsp;Read 1991 (<a href="https://doi.org/10.1086/285225">https://doi.org/10.1086/285225</a>).</p>

opencc-by-4.0Jul 2023View details →
zenodo44/100

Audit Opinion Score Dataset

<p>Please read the Code Book.</p> <p>Author: Tomohiro Hosoi</p> <p>&nbsp;</p> <p>1. About the Dataset</p> <p>This dataset contains information on the Auditor General South Africa&rsquo;s audit opinions based on the Public Finance Management Act (PFMA) and Municipal Finance Management Act (MFMA). The dataset aims to understand the trend of audit opinion.</p> <p>&nbsp;</p> <p>2. Procedure</p> <ul> <li>The Auditor General gives five audit opinions. Namely, (i) unqualified with no findings; (ii) unqualified with findings; (iii) qualified; (iv) adverse; (v) disclaimed; (vi) outstanding.</li> <li>Sheets named &ldquo;PFMA_Original&rdquo; and &ldquo;MFMA_Original&rdquo; are records of these audit opinions.</li> <li>The coder gives each opinion score based on the following rule; (i) unqualified with no findings is 5 points; (ii) unqualified with findings is 4 points; (iii) qualified is 3 points; (iv) adverse is 2 points; (v) disclaimed is 1 point; (vi) outstanding is 0 point.</li> <li>Sheets named &ldquo;PFMA_Score&rdquo; and &ldquo;MFMA_Score&rdquo; were records of these points.</li> </ul>

opencc-by-4.0Feb 2023View details →
zenodo44/100

Dataset of CT scans, slice photographs, and visual browning scores of 120 'Kanzi' apples

<p><strong>Summary</strong></p><p>This dataset is a collection of CT scans, slice photographs, and visual browning scores of 120 'Kanzi' apples.</p><p><br><strong>Description</strong></p><p><i>Sample information</i></p><p>In 2022, 120 'Kanzi' apples that had been stored under CA conditions (4 °C, 1 kPa O2, 1.5 kPa CO2) for 8 months were obtained from FruitMasters, The Netherlands. The fruit was grown in orchards surrounding Geldermalsen, the Netherlands, and harvested at physiological maturity in 2021.</p><p><i>CT acquisition</i></p><p>The dataset is acquired in the FleX-ray Laboratory, developed by TESCAN-XRE, located at CWI in Amsterdam. The CT scanner consists of a cone-beam microfocus polychromatic X-ray point source, and a 1944x1536 pixel, 14-bit, flat detector panel (Dexela1512NDT). Full details can be found in [Coban 2020].&nbsp; A cone beam geometry with a circular trajectory was used to acquire 1440 projection images at an exposure time of 100ms, a tube peak voltage of 90kV, a current of 550uA, and 2 times binning, halving the detector resolution. Volumes were reconstructed with the FDK algorithm and a voxel size of 129.3um. Beam hardening correction was used from the FleXbox package [Kostenko 2020]. To make sure that the grey values could be compared between scans the spectral sensitivity of the scanner was first estimated for each scan individually and the average of these estimates was used for beam hardening correction on all CT scans. All apples were scanned with the stem side on top. Moreover, a line was drawn on all apples from the stem to the calyx. The apples were put in the CT scanner so that the line was facing the X-ray source.</p><p>The CT volumes are saved as .tiff stacks. All volumes have been cropped to remove the background.</p><p><i>Slicing and photograph acquisition</i></p><p>One day after CT scanning, the apples were sliced using a modified meat-slicing machine (CaterChef, house brand of EMGA, Mijdrecht, The Netherlands), which is illustrated in the file slicing_machine_labels.png. The sliding surface of the meat-slicing machine was replaced by a transparent acrylic sheet, and a camera was placed behind the slicing surface. While in the machine, each apple was kept in place by a suction cup so that it could not rotate during the slicing. All apples were sliced from the stem end to the calyx end, with a slice thickness of roughly 4mm. Every time before slicing, a picture was taken of the remaining part of the apple through the transparent sliding surface. To ensure that all apples were roughly aligned to the CT scans, the apples were oriented so that the line drawn earlier was on top.</p><p>The slice photographs are saved as .png files. All photographs have been cropped to remove the background and to center the apple in the image.</p><p><i>Visual browning scores</i></p><p>After each apple was sliced it was also visually inspected, and a score from one to ten was given to describe the amount of browning in the apple.</p><p><strong>Related paper</strong></p><p>When using this dataset please consider citing the following paper. It explains how the dataset was collected and used for the first time:</p><p>Dirk Elias Schut, Rachael Maree Wood, Anna Katharina Trull, Rob Schouten, Robert van Liere, Tristan van Leeuwen, Kees Joost Batenburg, "Detecting internal disorders in fruit by CT. Part 1: Joint 2D to 3D image registration workflow for comparing multiple slice photographs and CT scans of apple fruit", 2023, <a href="https://arxiv.org/abs/2310.01987">arXiv preprint arXiv:2310.01987</a></p><p><br><strong>Research group</strong><br>This dataset was produced in a collaboration between the Computational Imaging group at Centrum Wiskunde &amp; Informatica (CWI), and GREEFA.</p><p><a href="https://www.cwi.nl/research/groups/computational-imaging">https://www.cwi.nl/research/groups/computational-imaging</a><br><a href="https://www.greefa.com/nl/">https://www.greefa.com/nl/</a></p><p><strong>Contact details</strong><br>dirk [dot] schut [at] cwi [dot] nl</p><p><strong>Acknowledgments</strong><br>This work was funded by the Dutch Research Council (NWO) through the UTOPIA project (ENWSS.2018.003). The authors also acknowledge TESCAN-XRE NV for their collaboration and support of the FleX-ray laboratory.</p>

opencc-by-4.0Jul 2023View details →
zenodo44/100

Metagenomes: metadata, taxonomic abundances, PlasX scores, circularity, and contig sequences

<p>This repository contains files that describe the ~36 million contigs that were assembled from 1,782 metagenomes</p> <p>metagenomes_contigs.tar.gz</p> <ul> <li>Fasta files of contig sequences. One file per metagenome</li> </ul> <p>metagenomes_metadata.txt</p> <ul> <li>Information about the 1,782 metagenomes</li> </ul> <p>metagenomes_taxonomic_abundances.txt.gz</p> <ul> <li>Taxonomic abundances&nbsp;in metagenomes, inferred by kraken2 and bracken</li> </ul> <p>metagenomes_contigs_summary.txt.gz</p> <ul> <li>PlasX scores</li> <li>Metagenomic coverage and detection</li> <li>Circularity</li> </ul>

opencc-by-4.0Nov 2021View details →
zenodo44/100

Watkins natural accessions yellow rust disease resistant scores

<p>The file contains&nbsp;the phenotypic information from field trails conducted in Kenya at the Kenya Agriculture and Livestock&nbsp;Research Organisation (KALRO) and the Ethiopian Institute of Agricultural Research (EIAR).&nbsp;&nbsp;Data are separated into the three rusts (yellow rust (Yr), stem rust (Sr) and leaf rust (Lr)), although data is not complete at all locations. When possible both seedling and adult plant data is provided. Adult scores include several observation across the growing season.&nbsp;For scoring rust severity, the modified Cobb scale (Peterson et al. 1948) was used to determine the percentage of tissue infected (0-100%) with rust and infection response (S, MS, MR and R, corresponding to susceptible, moderately susceptible, moderately resistant and resistant).&nbsp;</p> <p>Accession codes relate to Watkins landraces and their country of origin and accession names are indicated.&nbsp;Locations, dates and disease scores are indicated. Missing data is indicated as &quot;-&quot;.&nbsp;</p> <p>For more detailed passport data and access to germplasm visit the John Innes Centre Germplasm Resources Unit (<a href="https://www.seedstor.ac.uk/search-browseaccessions.php?idCollection=39">SeedStor</a>). Additional germplasm resources and populations developed from the Watkins accessions can be found here:&nbsp;<a href="https://wisplandracepillar.jic.ac.uk/">https://wisplandracepillar.jic.ac.uk/</a>&nbsp;&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Sep 2023View details →
zenodo44/100

Marine Stewardship Council (MSC) Fisheries Standard v2.0 assessment scores

<p>MSC fishery assessment scores manually collated by Marine Stewardship Council (MSC). See about file for complete details.</p> <p>This data has been manually collated by MSC from reports prepared by third party assessors. This dataset shows, for each Unit of Assessment (UoA), scoring data from the Public Certification Report of the UoA&#39;s most recent initial assessment or re-assessment conducted against the Default Tree v2.0 (including Default Tree v2.01) of the MSC Fisheries Standard as of 26 February 2021. A Unit of Assessment is a unique combination of target stock, fishing method or gear, geographic area, and operators or vessels being assessed which may eventually carry the MSC certification status. This excludes UoA which failed their assessment because failing scores are not consistently reported (they are often simply considered a &lsquo;fail&rsquo;).</p> <p>The following constraints and limitations apply to this dataset:<br> 1. This data has been manually collated by MSC from fishery reports prepared by third party assessors. MSC carries out data assurance to a standard that is fit for the purpose the information is used for, including being complete, accurate and as up to date as possible. &nbsp;If accuracy is paramount to a finite resolution, receivers are asked to validate data against assessment reports which can all be found published on fisheries.msc.org. The MSC is not responsible for any issues arising to any parties as a result of any information provided therein.<br> 2. This dataset shows, for each UoA, scoring data from the certification report of the UoA&#39;s most recent initial assessment or re-assessment conducted against v2.01 of the Standard as of 26 February 2021. This excludes UoA which failed their assessment.<br> 3. UoA details (uoa_details) can be joined to scoring data (scores) by the Event_unitbk. This will join the details of the scored UoA with the score that was awarded to the UoA.</p> <p>This dataset consists of two tables:</p> <p>1. uoa_details<br> This table contains details describing UoA included in the dataset. UoA descriptions include the name of fishery as published on Track a Fishery (fisheries.msc.org), species, gear, and other description of the UoA that was assessed.<br> <br> 2. scores<br> Scoring results to the lowest level (scoring guidepost of the scoring issue) for each UoA, as presented in the scoring tables of the report. One row represents a single guidepost result of a single scoring issue. However, the performance indicator (PI) &lsquo;Score&rsquo; will be duplicated for all scoring issues in the same PI. For example, if a UoA received a score of 80 for a performance indicator, and in this performance indicator there are 2 scoring issues each scored at 3 guideposts, there will be 6 total rows showing the results for each guidepost of each scoring issue, but all rows will show 80 for the overall &#39;Score&#39; of the PI. One row also exists for each Principle score of each UoA, in which the principle will be noted in the &#39;PI&#39; column, and the Principle score in &#39;Score&#39;. Data has been copied from the scoring tables in the Public Certification Report.</p>

opencc-by-4.0Sep 2023View details →
zenodo40/100

Paralog variant classification and scoring

<p><em>Para_zscore </em>data</p> <p>Input data, annotation of all hg19 missense variants, score for every gene having a paralog in the human genes. This dataset is a supplement for the publication Lal. et al.</p> <p>Information on the files, scripts to generate and use the <em>para_zscore</em> are available under</p> <p>https://git-r3lab.uni.lu/genomeanalysis/paralogs.</p> <p>Version 3582386 updates:</p> <p>- Annovar annotation file for hg38 added</p> <p>&nbsp;</p>

opencc-by-4.0Jun 2017View details →
zenodo40/100

Gini index mean score grain size estimates for Murray formation rocks (Gale crater, Mars) from ChemCam LIBS data (sols 766-1804)

<p>This repository includes main research results from the Journal of Geophysical Research: Planets manuscript entitled <em>&quot;Grain Size Variations in the Murray Formation: Stratigraphic Evidence for Changing Depositional Environments in Gale Crater, Mars&quot;</em>.&nbsp; These datasets were also included as Supplementary tables with the manuscript. Please find the plain language summary of the manuscript and the captions for these datasets below:</p> <p><strong><em>Plain language summary:</em></strong>&nbsp;The lowest exposed rocks of the Murray formation in Gale crater, Mars are interpreted as ancient lake deposits based on <em>Curiosity </em>rover data. However, the duration and temporal variability of this ancient lake is still an open question. Here we characterize the vertical distribution of deposits within the entire Murray formation using new grain size information. Characterizing grain size in rocks provides information about the speed of past fluid flows, which is crucial for interpreting depositional environments. However, measuring grain size in images is rarely possible for martian rocks. Thus, we estimate grain sizes with the Gini Index Mean Score (GIMS), a grain-size proxy that uses ChemCam Laser Induced Breakdown Spectroscopy data. GIMS results indicate that the Murray formation is dominated by rocks with mud-sized grains (i.e., mudstones), suggesting mud-sized grain settled in a low energy lake environment.&nbsp; Mud cracks occur in some of the mudstones, indicating drying periods in a lake. Rocks with sand-sized grains (i.e., sandstones) and cross bedding occur at specific intervals, suggesting episodes with stream channels and wind-blown sand dunes. The dominance of lake deposits interspersed with stream deposits suggests that liquid water was present in Gale crater for tens of thousands to millions of years.</p> <p><strong>Table captions:</strong></p> <p><em>Table S3. </em>All Murray formation targets used in the Gini mean index score (GIMS) analysis&nbsp;(sols 766-1804) with summary information, general grain size estimates from MAHLI and RMI images if known, and G<sub>MEAN </sub>values with associated standard deviation errors. For targets with N/A grain sizes, grains could not be resolved in any of the images, or images were not available. Protrusions in the rocks that are not grains are likely diagenetic nodules or concretions. Targets with G<sub>MEAN</sub>=0.07 have transitional GSRs, indicated by GSR1/GSR2. Targets names are merged in the same cell for those analyses that were taken on the same rock exposure. Next to the target names, the symbol * denotes that the ChemCam target was imaged by the MAHLI, the symbol ** denotes that the dust removal tool was used before the MAHLI image was taken, and ~ signifies that a location close to the ChemCam target was imaged by the MAHLI.</p> <p><em>Table S4</em>. The mean, median, minimum and maximum G<sub>MEAN</sub> and the minimum and maximum grain size regime (GSR) for each locality in the Murray formation.</p>

opencc-by-4.0Jan 2020View details →
zenodo40/100

A protocol to assess the risk of dementia among patients with coronary artery diseases using CAIDE score-Extended Data

<p>The contents of this extended data file are-<br> 01. Consent form (English &amp; Bengali Version)<br> 02. Interview Questionnaire (English &amp; Bengali Version)<br> These contents will help to address the objectives of the study that attempted to identify the risk of long-term dementia among coronary artery disease patients in Bangladesh.</p>

opencc-bySep 2020View details →
zenodo40/100

All vs all matrix of tanimoto scores for GNPS AllPositive dataset

<p>All vs all matrix of tanimoto scores of the &quot;AllPositive&quot; dataset, created from GNPS positive ionmode mass spectra (95320 spectra after filtering). To simplify this matrix, all spectra are grouped by the first 14 characters of their respective&nbsp;inchikeys, which leaves 12,846 inchikeys.</p> <p>In version 2&nbsp;a&nbsp;lookup table, the metadata file,&nbsp;is provided that records which row/column of the matrix corresponds to which inchikey.</p>

opencc-by-4.0Nov 2020View details →
zenodo40/100

High-quality large curated dataset of protein sequences (1.83 million) and their corresponding Position Specific Scoring Matrices

<p>As part of his&nbsp;master thesis at the Rostlab, which is located at the Technical University of Munich (TUM),&nbsp;Mr. Issar Arab&nbsp;developed&nbsp;the first language model that encodes evolutionary information of proteins explicitly. The pre-training involved the creation of a novel high-quality dataset of protein sequences (around 1.83&nbsp;million proteins, or ~0.8 Billion amino acids) with their corresponding Position Specific Scoring Matrices (PSSMs).&nbsp; Those matrices reflect the relative frequency of each amino acid at each position in a protein and is derived from evolutionarily related proteins.</p> <p>Mr. Arab makes this work publicly available to help other researchers speed up their work to leverage AI to learn the representation of protein evolutionary information more explicitly. The set of sequences was derived by extracting all PSSMs from the&nbsp;<a href="https://predictprotein.org/">PredictProtein</a>&nbsp;(PP) cache, which&nbsp;were also part o the UniProt&nbsp;Reference Cluster with 50% sequence identity (uniref50 2019_12). The overlap between PP and uniref50 was further filtered to only include high-quality samples, e.g. only multiple sequence alignments with a certain number of aligned sequences were considered. The processing led to&nbsp;a training set of 1.83 Million sequences, a validation set of 879 instances, and a test set of 879 entries.&nbsp;The training data of proteins is reduced to 40% sequence identity, with respect to the validation/test sets, and contains sequences ranging between 18 and 9858 residues in length.</p> <p>Refer to the Jupyter notebook for a detailed description of the files'&nbsp;structure and a Python code snippet to correctly manipulate&nbsp;this data.</p> <p>To access the full original&nbsp;work, please visit the following link:&nbsp; <a href="https://mediatum.ub.tum.de/node?id=1579236">Manuscript</a>&nbsp;<br><br><strong>Note:</strong> The dataset was recently used to fine tune a protein sequence language model (<a href="https://github.com/issararab/PEvoLM">PEvoLM</a>). The work was presented at the CIBCB'23 conference. If you use PEvoLM or this dataset in your work, please cite the following publication:</p> <p>- Issar Arab, <strong>PEvoLM: Protein Sequence Evolutionary Information Language Model</strong>, <em>IEEE Conference on Computational Intelligence in Bioinformatics and Computational Biology (CIBCB), Eindhoven, Netherlands</em>, (2023), pp. 1-8, doi:<a href="https://ieeexplore.ieee.org/document/10264890">10.1109/CIBCB56990.2023.10264890</a></p>

opencc-by-4.0Sep 2020View details →
zenodo40/100

Anopheles gambiae (AgamP4) genome conservation score

<p>The conservation score storage&nbsp;is a result of&nbsp;a bioinformatics pipeline that integrates a systematic analysis of the data on genetic variation in more than 1,000 wild-caught Anopheles gambiae individuals and conserved syntenic regions of 19 Anopheles species and 3 phylogenetically more distant species of dipterans.</p> <p>The results of this analysis are gathered in the HDF5&nbsp;data storage system that allows for flexible extraction and bioinformatic manipulation at each genomic position in AgamP4 reference genome.</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Aug 2020View details →
zenodo40/100

PIR data and EEG scoring for Wellcome Open Research methods paper (Brown et al 2016)

<p>PIR data and EEG-scored sleep in the Wellcome Open Research article:</p> <p>'COMPASS: Continuous Open Mouse Phenotyping of Activity and Sleep Status'</p> <p> </p> <p>1sensorPIRvsEEGdata.csv  -  PIR based actigraphy for mice to compare to EEG-scored sleep</p> <p>EEG_4mice10sec.csv  -  Manually scored sleep from EEG files (.edf) from 10.5281/zenodo.160118</p> <p>blandAltLandD.csv  -  paired estimates of sleep by PIR and EEG methods (sum of 4 mice over 1 day in 30min bins)</p> <p><br> 1monthPIRsleep.csv  - 1 month of activity for for figure 4</p> <p><br> 24mice_activity_LD1week.csv  - activity and sleep for 24 wt mice (for hierarchical clustering in figure 4)<br> 24mice_sleep_LD1week.csv </p> <p>     </p> <p> </p>

opencc-zeroOct 2016View details →
zenodo40/100

OTMM Score Structure Experiments for FMA 2016

<p>otmm-score-structure-experiments</p> <p>Structure Analysis Experiments on Ottoman-Turkish Makam Music Scores</p> <p>This repository contains the experiments to find the optimal melodic and lyrics similarity threshold, conducted in the paper:</p> <p>Şentürk, S., &amp; Serra X. (2016). A method for structural analysis of Ottoman-Turkish makam music scores. In Proceedings of 6th International Workshop on Folk Music Analysis, (pp. 39-46)., Dublin, Ireland.</p> <p>For the details of the experiments, please refer to the paper. Please cite the publication above in any work using these experiments.</p> <p>The submodule turkish_makam_section_dataset stores the test scores in the SymbTr-txt format. The experiments folder have the experimental results and evaluation. Each folder in this folder stores the sections extracted from each score for the given threshold, e.g. folder "0_6" has the results obtained using a similarity threshold of 0.6 for both melodic and lyrical relationship computation. The extracted sections for each score are stored in a csv file, which has the same name as the SymbTr-name (makam--form--usul--name--composer) of the analyzed score. The fields are:</p> <p>start_note: The starting note index in the SymbTr-txt score end_note: The ending note index in the SymbTr-txt score name: The basic semantic name of the section. Right now, it is the name (TESLİM, ARANAĞME...) annotated in the score for instumental sections or "VOCAL_SECTION" for vocal sections. melodic_structure: Melodic semiotic label lyric_structure: Lyrical semiotic label lyrics: Lyrics of the section slug: The processed version of "name" field with the Turkish characters and special characters handled.</p> <p>results.json stores the evaluation and the statistics of the experiments.</p> <p>To run the experiments you have to install Jupyter notebook and the requirements.</p> <p>For additional information please contact the authors.</p>

openother-openNov 2016View details →
zenodo40/100

Concept detection scores for the IACC.3 dataset (TRECVID AVS Task)

<p>We provide concept detection scores for the IACC.3 dataset (600 hr internet archive videos), which is used in the TRECVID Ad-hoc Video Search (AVS) task [1]. Concept detection scores for 1345 concepts (1000 ImageNet concepts provided for the ILSVRC challenge [2] and 345 TRECVID SIN concepts [3]) have been generated as follows:<br> 1) To generate scores for the ImageNet concepts, 5 pre-trained ImageNet networks were applied on the IACC.3 dataset and their output was fused in terms of arithmetic mean.<br> 2) To generate scores for the TRECVID SIN concepts, two pre-trained ImageNet networks were fine-tuned on these concepts using a combination of our methods presented in the following papers: [4], [5]. We provide two different sets of concept scores for the TRECVID SIN concepts: a) The output of the two fine-tuned networks was fused in terms of arithmetic mean in order to return a single score for each concept. b) The last fully-connected layer was used as feature to train SVM classifiers separately for each fine-tuned network and each concept. Then, the SVM classifiers were applied on the IACC.3 dataset and the prediction scores of the SVMs for the same concept were fused in terms of arithmetic mean in order to return a single score for each concept. We evaluated the two different sets of concepts in terms of MXInfAP on a subset of 38 TRECVID SIN concepts for which ground-truth annotation exists, and the MXInfAP of each set of concept scores is: a) 30.04% for the networks' direct output, b) 35.81% for the SVM classifiers.</p> <p>Three different files of concept detection scores can be downloaded (after unpacking the compressed file):<br> 1) scores_ImageNet.txt<br> 2a) scores_SIN_direct.txt<br> 2b) scores_SIN_svm.txt<br> In total there are 335944 rows in each file; 1002 columns in the first file and 347 columns in each of the other two. Each row in any of these files corresponds to a different video shot; the video shot IDs appear in the first two columns. (Note: the shot IDs are the ones from the mp7 files in the TRECVID AVS master shot reference, with the format shotFILENUMBER_SHOTNUMBER). Then, each column (except for the fist two) corresponds to a different concept, with all concept scores being in [0,1] range. The higher the score the more likely that the corresponding concept appears in the video shot. Files “concept_names_ImageNet.txt” and “concept_names_SIN.txt” indicate the order of the concepts that is used in the concept score files. </p> <p>[1] G. Awad, J. Fiscus, M. Michel et al. 2016. TRECVID 2016: Evaluating Video Search, Video Event Detection, Localization, and Hyperlinking. In TRECVID 2016 Workshop. NIST, USA.<br> [2] O. Russakovsky, J. Deng, H. Su et al. 2015. ImageNet Large Scale Visual Recognition Challenge. Int. Journal of Computer Vision (IJCV) 115, 211–252.<br> [3] G. Awad, C. Snoek, A. Smeaton, and G. Quénot. 2016. TRECVid semantic indexing of video: a 6-year retrospective. ITE Transactions on Media Technology and Applications, 4 (3). pp. 187-208.<br> [4] N. Pittaras, F. Markatopoulou, V. Mezaris, I. Patras. 2017. Comparison of Fine-tuning and Extension Strategies for Deep Convolutional Neural Networks, Proc. 23rd Int. Conf. on MultiMedia Modeling (MMM'17), Reykjavik, Iceland, Springer LNCS vol. 10132, pp. 102-114, Jan. 2017.<br> [5] F. Markatopoulou, V. Mezaris, and I. Patras. 2016. Deep Multi-task Learning with Label Correlation Constraint for Video Concept Detection, Proc. ACM Multimedia 2016, Amsterdam, Oct. 2016.</p>

opencc-by-4.0Feb 2017View details →
zenodo40/100

Concept detection scores for the MED16train dataset (TRECVID MED task)

<p>We provide concept detection scores for the MED16train dataset which is used at the TRECVID Multimedia Event Detection (MED) task [1]. First, each video is decoded into a set of keyframes at fixed temporal intervals (2 keyframes per second). Then, we calculated concept detection scores for the two following concept sets: i) 487 sport-related concepts from YouTube Sports-1M Dataset[1] and ii) 345 TRECVID SIN concepts [3]. The scores have been generated as follows:<br> 1) For the 487 concepts for the Sports-1M Dataset, a Googlenet network [4] originally trained on 5055 ImageNet concepts was fine-tuned, following the extension strategy of [2] with one extension layer of dimension 128.<br> 2) For the 345 TRECVID SIN concepts, a pre-trained Googlenet network [4] on 5055 ImageNet concepts was fine-tuned on these concepts, again following the extension strategy of [2] with one extension layer of dimension 1024. </p> <p>After unpacking the compressed file two different folders can be found, namely "Prob_sports_MED16train" and "Prob_SIN_MED16train", one for each concept set. We provide one file for every video of the MED16train dataset for each concept set. Each file consists of N columns (where N = 345 for TRECVID SIN and N = 487 for Sports-1M Dataset) and M rows (where M is the number of extracted keyframes for the corresponding video). Each column corresponds to a different concept, with all concept scores being in the range [0,1]. The higher the score the more likely that the corresponding concept appears in the keyframe. Two additional files are provided; files "sports_487_Classes.txt" and "SIN_345_Classes.txt" indicate the order of the concepts that is used in the concept score files.</p> <p>[1] A. Karpathy, G. Toderici, S. Shetty, T. Leung, R. Sukthankar and L. Fei-Fei, "Large-scale video classification with convolutional neural networks", In Proceedings of the IEEE conference on Computer Vision and Pattern Recognition, pp. 1725-1732, 2014.<br> [2] N. Pittaras, F. Markatopoulou, V. Mezaris and I. Patras, "Comparison of Fine-tuning and Extension Strategies for Deep Convolutional Neural Networks", Proc. 23rd Int. Conf. on MultiMedia Modeling (MMM'17), Reykjavik, Iceland, Springer LNCS vol. 10132, pp. 102-114, Jan. 2017.<br> [3] G. Awad, C. Snoek, A. Smeaton, and G. Quénot, "TRECVid semantic indexing of video: a 6-year retrospective", ITE Transactions on Media Technology and Applications, 4 (3). pp. 187-208, 2016.<br> [4] C. Szegedy, Wei Liu, Yangqing Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke and A. Rabinovich, "Going deeper with convolutions", In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 1-9, 2015.</p>

opencc-by-4.0Mar 2017View details →
zenodo40/100

Integrating AlphaFold pLDDT Scores into CABS-flex for Enhanced Protein Flexibility Simulations

<div>This dataset accompanies the publication "Integrating AlphaFold pLDDT Scores into CABS-flex for enhanced protein flexibility simulations".</div> <div>This project was funded by the OPUS grant from the National Science Centre, Poland [2020/39/B/NZ2/01301].</div> <div>&nbsp;</div> <div>Training_set_protein_chains.txt and Whole_set_protein_chains.txt have lists of all PDB ID and chain used.</div> <div>Description_of_runs.csv has a list of every run tested. Run number correponds to csv file in Results_run.tar.gz.</div> <div>Every csv file has following columns:&nbsp;</div> <div> <ul> <li>PDB - PDB ID and chain&nbsp;</li> <li>Total_residues</li> <li>%_C - Percent of secondary structure assigned as coil by DSSP</li> <li>%_H - Percent of secondary structure assigned as helix by DSSP</li> <li>%_E - Percent of secondary structure assigned as sheet by DSSP</li> <li>%_T - Percent of secondary structure assigned as turn by DSSP</li> <li>pLDDT_mean - Average pLDDT score across all residues</li> <li>pLDDT_std - Standard deviation of pLDDT scores across all residues</li> <li>Unique_restraints - Number of unique restraints created by CABS-flex</li> <li>RMSF_CABS_R1_corr - RMSF correlation between CABS-flex and first MD simulation from ATLAS</li> <li>RMSF_CABS_R2_corr - RMSF correlation between CABS-flex and second MD simulation from ATLAS</li> <li>RMSF_CABS_R3_corr - RMSF correlation between CABS-flex and third MD simulation from ATLAS</li> <li>Highest_RMSF_corr - Highest RMSF correlation out of three</li> </ul> </div> <div>&nbsp;</div>

opencc-by-4.0Oct 2024View details →
zenodo40/100

Handwritten ASAP Short Answer Scoring

<p>This dataset is based on the Short Answer Scoring (SAS) dataset of the&nbsp;<a href="https://www.kaggle.com/c/asap-sas">Automated Student Assessment Prize (ASAP)</a>.&nbsp;<br>Although the original dataset was conducted on handwritten content, the scans are not available.<br>To analyze the full pipeline from Handwritten Answers to Automated Scoring, we let students rewrite some answers.<br>The texts used from SAS are from the test set and from the training set.&nbsp;</p>

opencc-by-4.0Jun 2020View details →
zenodo40/100

Data for "A learned score function improves the power of mass spectrometry database search"

<div> <h1>DATA for "A learned score function improves the power of mass spectrometry database search"</h1> <br> <div>These data files are associated with the following publication:</div> <br> <div> <ul> <li>Varun Ananth, Justin Sanders, Melih Yilmaz, Sewoong Oh and William Stafford Noble. "<a title="biorXiv Preprint Link" href="https://www.biorxiv.org/content/10.1101/2024.01.26.577425v2" target="_blank" rel="noopener">A learned score function improves the power of mass spectrometry database search</a>". Bioinformatics (Proceedings of the ISMB). &nbsp;2024.</li> </ul> </div> <br> <div>For the benchmarking data, we used a dataset that is publicly available on ProteomeXchange (PXD028735). The paper that introduced this dataset is:</div> <br> <div> <ul> <li>Van Puyvelde, B., Daled, S., Willems, S., Gabriels, R., Gonzalez de Peredo, A., Chaoui, K., Mouton-Barbosa, E., Bouyssi&eacute;, D., Boonen, K., Hughes, C. J., Gethings, L. A., Perez-Riverol, Y., Bloomfield, N., Tate, S., Schiltz, O., Martens, L., Deforce, D., &amp; Dhaenens, M. (2022). A comprehensive LFQ benchmark dataset on modern day acquisition strategies in proteomics. In Scientific Data (Vol. 9, Issue 1). Springer Science and Business Media LLC. https://doi.org/10.1038/s41597-022-01216-6</li> </ul> </div> <br> <div>More specifically, the following `.raw` files were downloaded:</div> <br> <ul> <li><code>LFQ_Orbitrap_DDA_Ecoli_01.raw</code></li> <li><code>LFQ_Orbitrap_DDA_Human_01.raw</code></li> <li><code>LFQ_Orbitrap_DDA_Yeast_01.raw</code></li> </ul> <br> <div>Those files can be accessed via FTP&nbsp;<a title="Link to ProteomeXchange: PXD028735" href="https://ftp.pride.ebi.ac.uk/pride/data/archive/2022/02/PXD028735/" target="_blank" rel="noopener">here</a>.</div> <br> <div>We upload here the annotated <code>.mgf</code> files created from these <code>.raw</code> files, as described in our paper.</div> <br> <div>The human, yeast, and E. coli .fasta files used in all database searches were downloaded from UniProt on 11/6/23, 4:30 PM.</div> <br> <div> <ul> <li>Bateman, A., Martin, M.-J., Orchard, S., Magrane, M., Ahmad, S., Alpi, E., Bowler-Barnett, E. H., Britto, R., Bye-A-Jee, H., Cukura, A., Denny, P., Dogan, T., Ebenezer, T., Fan, J., Garmiri, P., da Costa Gonzales, L. J., Hatton-Ellis, E., Hussein, A., &hellip; Zhang, J. (2022). UniProt: the Universal Protein Knowledgebase in 2023. In Nucleic Acids Research (Vol. 51, Issue D1, pp. D523&ndash;D531). Oxford University Press (OUP). https://doi.org/10.1093/nar/gkac1052</li> </ul> </div> <br> <div>We include these files here, with only minor modifications to replace `U` amino acids with `X` so that all amino acids fall into Casanovo-DB's vocabulary.</div> </div>

opencc-by-4.0Mar 2024View details →
zenodo40/100

Homologous Missense Constraint scores

<p>For all missense variants with HMC scores:</p> <p>1. If one missense variant could be mapped to the multiple Pfam domains, we report&nbsp;the worst HMC scores (smallest, more likely to be deleterious) since it's&nbsp;the worst scenario.&nbsp;&nbsp;&nbsp;</p> <p>2. The genomic positions are provided with genome build in both hg19 and hg38.&nbsp;</p>

opencc-by-4.0Mar 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record