Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
14,447
datasets available to search
ShareScore release 0.7.1
Dataset results
14,447 results for “Identification”
TriGraphSlant - benchmark set for writer identification - writers were asked to write in unnatural slant
<p> <br> Disclaimer and terms of use:<br> ============================<br> <br> /*****************************************************************************\<br> * *<br> * *<br> * This is the TrigraphSlant (Img version) Distribution, release 18/3/2011 * *<br> * *<br> * This distribution contains 188 images of scanned handwritten text, *<br> * scanned at resolution 300dpi Canon LiDE 25, grey scale, *<br> * by 47 Dutch writers, four pages per writer, from four *<br> * writing conditions, one condition per page. The conditions are: *<br> * 1. [AN] Copy text A in your natural handwriting. *<br> * 2. [BN] Copy text B in your natural handwriting. *<br> * 3. [BL] Copy text B and slant your handwriting to the *<br> * left as much as possible. *<br> * 4. [BR] Copy text B and slant your handwriting to the *<br> * right as much as possible. *<br> * The codes AN, BN, BL and BR refer to subsets into which the collected *<br> * pages of the writers were subdivided. AN represents a collection of *<br> * authentic documents; BN, BL and BR can be seen as collections of *<br> * questioned documents. To avoid structural effects of fatigue, the order *<br> * of item 3 and 4 was randomized at each collection: half of the subjects *<br> * wrote the BR page before the BL page. The data were collected at three *<br> * sites, in three cities: The Hague: NFI (N...), Donders Institute for *<br> * Brain, Cognition and Behaviour, Radboud University Nijmegen (D...) *<br> * and the Artificial Intelligence Dept. of University of Groningen (R...) *<br> * *<br> * Copyright The International Unipen Foundation, 2010, All rights reserved *<br> *******************************************************************************<br> * *<br> * *<br> * DISCLAIMER AND COPYRIGHT NOTICE FOR ALL DATA CONTAINED ON THIS CARRIER: *<br> * *<br> * *<br> * 1) PERMISSION IS HEREBY GRANTED TO USE THE DATA FOR RESEARCH *<br> * PURPOSES. IT IS NOT ALLOWED TO DISTRIBUTE THIS DATA FOR COMMERCIAL *<br> * PURPOSES. *<br> * *<br> * *<br> * 2) PROVIDER GIVES NO EXPRESS OR IMPLIED WARRANTY OF ANY KIND AND ANY *<br> * IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR PURPOSE ARE *<br> * DISCLAIMED. *<br> * *<br> * 3) PROVIDER SHALL NOT BE LIABLE FOR ANY DIRECT, INDIRECT, SPECIAL, *<br> * INCIDENTAL OR CONSEQUENTIAL DAMAGES ARISING OUT OF ANY USE OF THIS *<br> * DATA. *<br> * *<br> * 4) THE USER SHOULD REFER TO THE FOLLOWING ARTICLE ON THIS DATA SET: *<br> * *<br> * A.A. Brink, R.M.J. Niels, R.A. van Batenburg, C.E. van den Heuvel, *<br> * L.R.B. Schomaker, Towards robust writer verification by correcting *<br> * unnatural slant, Pattern Recognition Letters, Volume 32, Issue 3, *<br> * 1 February 2011, Pages 449-457, ISSN 0167-8655, *<br> * DOI: 10.1016/j.patrec.2010.10.010. *<br> * *<br> * 5) THE RECIPIENT SHOULD REFRAIN FROM PROLIFERATING THE DATA SET TO THIRD *<br> * PARTIES EXTERNAL TO HIS/HER LOCAL RESEARCH GROUP. PLEASE REFER INTERESTED *<br> * RESEARCHERS TO HTTP://UNIPEN.ORG FOR OBTAINING THEIR OWN COPY. *<br> \*****************************************************************************/<br> <br> Abstract<br> <br> Towards robust writer verification by correcting unnatural slant<br> <br> A.A. Brink, , R.M.J. Niels, R.A. van Batenburg, C.E. van den Heuvel, <br> and L.R.B. Schomaker, <br> <br> a Institute of Artificial Intelligence and Cognitive Engineering (ALICE), <br> University of Groningen, P.O. Box 407, 9700 AK Groningen, The Netherlands<br> <br> b Donders Institute for Brain, Cognition and Behaviour, Radboud University Nijmegen, <br> P.O. Box 9104, 6500 HE Nijmegen, The Netherlands<br> <br> c Netherlands Forensic Institute, P.O. Box 24044, 2490 AA Den Haag, The Netherlands<br> <br> Received 11 September 2009. Available online 30 October 2010.<br> <br> Slant is a salient feature of Western handwriting and it is considered to be an<br> important writer-specific feature. In disguised handwriting however, slant is<br> often modified. It was tested whether slant is indeed an important factor and it<br> was tested whether the distorting effect of deliberate slant change can be<br> countered by a simple shear transform. This was done in two off-line writer<br> verification experiments in image processing conditions of slant elimination and<br> slant correction. The experiments were performed using three features based on<br> statistical pattern recognition, including the state-of-the-art features<br> Fraglets and Hinge. A new public dataset was created and used, containing<br> natural and slanted handwriting by 47 writers. A striking result is that the<br> average natural slant value is much less important for biometric systems than is<br> usually assumed: eliminating slant yields just a 1-5% performance loss. A<br> second result is that the effects of deliberate slant change cannot be fully<br> countered by a simple shear transform: it raises performance on the distorted<br> handwriting from 53-68% to 64-90%, but this is still lower than normal<br> operation on natural handwriting: 97-100%.<br> <br> Research highlights<br> - The value of slant as a writer identification feature has been overrated. <br> - Deliberate slant change can be partly countered by the shear transform. <br> - Deliberate slant change introduces non-affine distortions to the handwriting. <br> - A new dataset of deliberately slanted handwriting was introduced.<br> <br> Keywords: Handwriting biometrics; Writer verification; Slant; Disguise; Statistical<br> pattern recognition</p> <p> </p>
Identification and Functional Characterization of an Alternative Cancer-derived PD-L1 Isoform (supplemental data)
<p>The enclosed files contain all of the supplemental data from: Identification and Functional Characterization of an Alternative Cancer-derived PD-L1 Isoform. The files include the complete tables in CSV-formatted files.</p>
FitHiChIP: Identification of significant chromatin contacts from HiChIP data
<p>FitHiChIP is a computational method for identifying chromatin contacts among regulatory regions such as enhancers and promoters from HiChIP/PLAC-seq data.</p> <p><strong>Functionalities</strong> of FitHiChIP include:</p> <p>1) Calling significant interactions / loops / contacts from a HiChIP / PLAC-seq data </p> <p>2) Identifying peaks (enriched segments) from a HiChIP data (i.e. HiChIP peak caller)</p> <p>3) Finding differential loops among non-differential loci between two different categories of HiChIP samples, each with one or more replicates.</p> <p><strong>GitHub page</strong>: <a href="https://github.com/ay-lab/FitHiChIP">github.com/ay-lab/FitHiChIP</a></p> <p><strong>Documentation</strong>: <a href="https://ay-lab.github.io/FitHiChIP/">https://ay-lab.github.io/FitHiChIP/</a></p> <p><strong>Citation</strong>: Please check the above documentation regarding citation of FitHiChIP</p> <p>About this repository: All the data and results provided here correspond to the published manuscript. The file <strong>Data_Summary.xlsx</strong> summarizes for each figure, corresponding tables storing the related datasets.</p>
Electronic Identification survey on small ruminants
<p>The results from the survey made in the 7 countries on farmers about the current use of Electronic Identification and main barriers and motivations in October 2018.</p>
Supplementary material: Efficient in vivo screening method for the identification of C4 photosynthesis inhibitors based on cell suspensions of the single-cell C4 plant Bienertia sinuspersici
<p>Data described in Minges et al. (2019) Efficient <em>in vivo</em> screening method for the identification of C<sub>4</sub> photosynthesis inhibitors based on cell suspensions of the single-cell C<sub>4</sub> plant <em>Bienertia sinuspersici</em>. doi: <a href="https://doi.org/10.3389/fpls.2019.01350">10.3389/fpls.2019.01350</a></p> <p> </p>
Data From: Harnessing Deep Belief Networks for Selective HDAC6 Inhibitors Identification
<p>This dataset contains the results from virtual screening of SPECS library after predicting the selectivity of molecules against HDAC6 over using a Deep Belief Network model that was trained and tested by the authors.</p> <p>The docking poses of 10 molecules with best docking scores were given along with their MMGBSA scores. </p> <p>The MD trajectory files along with the analysis were also provided for the selected three molecules as well as the reference molecule (co-crystalized ligand).</p>
Genome-wide tool for rapid de novo identification and visualisation of interspersed and tandem
<p><span>Genomic repeats are functionally ubiquitous structural units found in all genomes. Studying these repeats of different origins is essential for the evolution and adaptation of a given organism. These repeating patterns have manifold signatures and structures with varying degrees of homology, making their identification challenging. To address this challenge, we developed a new algorithm and software that can rapidly and accurately detect any repeated sequences <em>de novo</em> with varying degrees of homology in genomic sequences in interspersed or clustered repeats. Numerous forms of repeated sequences and complex patterns can be identified, even for complex sequence variants and implicit or mixed types of repeat blocks. Direct and inverted-repeat elements, perfect and imperfect microsatellite repeats, and any short- or long-tandem repeat belonging to a wide range of higher-order repeat structures of telomers or large satellite sequences can be detected. By combining precision and versatility, our tool contributes significantly to elucidating the intricate landscape of genomic repeats.</span></p>
vPro-MS peptide spectral library for the identification of human-pathogenic viruses by untargeted proteomics
<p>The viral proteomics workflow (vPro-MS) enables identification of human-pathogenic viruses from patient samples by untargeted proteomics. vPro-MS is based on an in-silico derived peptide library covering the human virome in <a href="https://www.uniprot.org/" rel="nofollow">UniProtKB</a> (331 viruses, 20,386 genomes, 121,977 peptides). vPro-MS is intended to identify human-pathogenic viruses from DiaNN (<a href="https://github.com/vdemichev/DiaNN">https://github.com/vdemichev/DiaNN</a>) outputs of either DIA or diaPASEF data. A scoring algorithm (vProID) assesses the confidence of virus identification and the results are finally summarized in a report table. </p> <p>The vPro Peptide Library folder contains 3 peptide FASTA files (Contaminants.fasta, Human.fasta, vPro.Virus.fasta), which were used to predict the spectral library (vPro-lib.predicted.speclib). Please note, that the additional commands “--cut” and “--duplicate-proteins” are needed to reprocess the prediction in DiaNN. This spectral library should be used to identify peptide sequences from samples of human origin using DiaNN. Furthermore, the folder contains the metadata file of the viral peptide sequences (vPro.Peptide.Library.txt) and a summary file of the virus taxonomy covered by the library (Taxonomy.Summary.txt). The metadata file is used by the vPro script to identify viruses from the DiaNN main report.</p>
Alphafold2_ab_initio iterative predictions for folding intermediate identification
<p>PDB ids starts from 1 and 8, rmsds, plddts, t-sne embeddings.</p> <p>Check related biorxiv preprint: AlphaFold2 knows some protein folding principles; DOI: https://doi.org/10.1101/2024.08.25.609581.</p>
Review of Polydora species from Brazil, with identification key and description of two new species (Annelida: Spionidae)
<p>Supplementary Material for the article published by the <strong><em>Ocean and Coastal Research</em> Journal</strong></p> <p>Complete information on the material examined during this study and records by other authors is given in Supplementary Tables S1−11. A list of the museums and other collections (and their acronyms) holding the samples which are reported in this study is given in Table ESM12.</p>
Identification of Genes Regulating Dexamethasone Resistance and Prognostic Model Development in Acute Lymphoblastic Leukemia
<p>This study investigates the mechanisms of dexamethasone resistance in acute lymphoblastic leukemia (ALL) and presents a prognostic model to predict patient outcomes and immunotherapy responses. By analyzing gene expression data, we identified autophagy-related genes associated with dexamethasone resistance, particularly focusing on STK38L’s role in modulating autophagy via ULK1. Our results reveal that high STK38L expression enhances dexamethasone resistance by promoting autophagy markers LC3II/LC3I and beclin-1. This study provides valuable insights into the molecular basis of dexamethasone resistance and highlights STK38L as a potential biomarker and therapeutic target for improving ALL treatment strategies.</p>
PhasAGE Training School 1 - Linear motifs identification and prediction - PRACTICAL
<p>The Training School 1 <strong>“Computational Methods to Study Protein Phase Separation”</strong> is the first edition of a series of PhasAGE training activities.</p> <p>The goal of this course is to provide participants with the basic knowledge to understand the phenomenon of <strong>Phase Separation</strong>, its role in biological processes and diseases. In addition, the course will provide <strong>an overview of the available computational resources</strong> to navigate this knowledge. Participants will have <strong>hands-on training</strong> in tools and resources available for life sciences, to collect information from the literature on biomolecular phase transitions, identify features triggering phase transitions, mutations associated with diseases, known or predicted PTMs and molecular interaction sites.</p>
Image-based Many-language Programming Language Identification - Replication Package
<p>This dataset contains the data, software, and instructions needed to replicate the findings of the paper:</p> <p>Francesca Del Bonifro, Maurizio Gabbrielli, Antonio Lategano, and Stefano Zacchiroli. Image-based Many-language<br> Programming Language Identification. <a href="https://peerj.com/computer-science/"><em>PeerJ Computer Science</em></a>, 2021 (to appear). DOI: <a href="https://dx.doi.org/10.7717/peerj-cs.631">10.7717/peerj-cs.631</a></p> <p>After retrieving the full dataset, extract the replication-package.zip archive and follow the instructions described in the README.md file.</p>
xPore: Identification of differential RNA modifications from nanopore direct RNA sequencing - SGNEx data
<p>xPore is a Python package for identification and quantification of differential RNA modifications from direct RNA sequencing.</p> <p>The detailed usage is documented at <a href="https://xpore.readthedocs.io/en/latest/">https://xpore.readthedocs.io/en/latest</a>, while all scripts and source code are available at <a href="https://github.com/GoekeLab/xpore">https://github.com/GoekeLab/xpore</a>.</p> <p>All the preprocessed datasets used in the paper are provided here. </p> <p>Please cite our paper below when using these data.<br> Ploy N. Pratanwanich et al. "Detection of differential RNA modifications from direct RNA sequencing of human cell lines." bioRxiv (2020).</p>
Research Data Identification - RDAlliance
<p>A diagram that aims to align with an elaboration of the Research Data Alliance Dynamic Data Citation Working Group recommendations, as documented in:</p> <p>Andreas Rauber, Ari Asmi, Dieter van Uytvanck, and Stefan Pröll. (2016). Identification of reproducible subsets for data citation, sharing and re-use. https://www.force11.org/sites/default/files/d7/project/81/ieee-tcdl-dc-2016_paper_1.pdf</p> <p>(Which later lead to: Rauber, Andreas, Asmi, Ari, van Uytvanck, Dieter, & Proell, Stefan. (2015). Data Citation of Evolving Data: Recommendations of the Working Group on Data Citation (WGDC). https://doi.org/10.15497/RDA00016)</p>
Data for: Unmasking the Effects of Orthography, Semantics, and Phonology on 2AFC Visual Word Perceptual Identification
<p>This data was used in analyses for "Unmasking the Effects of Orthography, Semantics, and Phonology on 2AFC Visual Word Perceptual Identification".</p>
Vectra Polatis image of human colorectal cancer (CRC1) from: A SIMPLI (Single-cell Identification from MultiPLexed Images) approach for spatially resolved tissue phenotyping at single-cell resolution.
<p>Two 4 µm thick serial sections were cut from CRC1 FFPE block using a microtome. The first slide was dewaxed and rehydrated before carrying out HIER with Antigen Retrieval Reagent-Basic (R&D Systems). The tissue was then blocked and incubated with the anti-CD3 antibody (Dako, Supplementary Table 2) followed by horseradish peroxidase (HRP) conjugated anti-rabbit antibody (Dako) and stained with 3,3' diaminobenzidine (DAB) substrate (Abcam) and haematoxylin. Areas with CD3<sup>+</sup> infiltration in the proximity of the tumour invasive margin were identified by a clinical pathologist (M. R-J.)</p> <p>The second slide was stained with a panel of six antibodies (CD8, PD1, Ki67, PDL1, CD68, GzB, Supplementary Table 2), Opal fluorophores and 4’,6-diamidino-2-phenylindole (DAPI) on a Ventana Discovery Ultra automated staining platform (Roche). Expected expression and cellular localisation of each marker as well as fluorophore brightness were used to minimise fluorescence spillage upon antibody-Opal pairing. Following a one-hour incubation at a 60°C, the slide was subjected to an automated staining protocol on an autostainer. The protocol involved deparaffinisation (EZ-Prep solution, Roche), HIER (DISC. CC1 solution, Roche) and seven sequential rounds of: one hour incubation with the primary antibody, 12 minutes incubation with the HRP-conjugated secondary antibody (DISC. Omnimap anti-Ms HRP RUO or DISC. Omnimap anti-Rb HRP RUO, Roche) and 16 minute incubation with the Opal reactive fluorophore (Akoya Biosciences). For the last round of staining, the slide was incubated with Opal TSA-DIG reagent (Akoya Biosciences) for 12 minutes followed by Opal 780 reactive fluorophore for our hour (Akoya Biosciences). A denaturation step (100°C for 8 minutes) was introduced between each staining round in order to remove the primary and secondary antibodies from the previous cycle without disrupting the fluorescent signal. The slide was counterstained with DAPI (Akoya Biosciences) and coverslipped using ProLong Gold antifade mounting media (Thermo Fisher Scientific). The Vectra Polaris automated quantitative pathology imaging system (Akoya Biosciences) was used to scan the labelled slide. Six fields of view, within the area selected by the pathologist, were scanned at 20x and 40x magnification using appropriate exposure times and loaded into inForm{Kramer, 2018 #23} for spectral unmixing and autofluorescence isolation using the spectral libraries. After spectral unmixing and merging of six 20x fields of view for a total of >5mm<sup>2</sup> ROI (Table 2), one single-tiff image was extracted for each marker and its intensity was rescaled from 0 to 1 with custom R scripts.</p>
Dataset used in the publication entitled "Decomposition Problem in Process of Selective Identification and Localization of Voltage Fluctuation Sources in Power Grids" presented at 2022 20th International Conference on Harmonics and Quality of Power (ICHQP)
<p>Dataset obtained from experimental research carried out in a real power grid. Based on the dataset, the problem of decomposition in identification of sources of voltage fluctuations has been presented in the publication: Kuwałek P., Decomposition Problem in Process of Selective Identification and Localization of Voltage Fluctuation Sources in Power Grids, <em>Proceedings of the 20th International Conference on Harmonics and Quality of Power</em>, IEEE , art. no. 43, 2022, Italy, Naples. The description of the power grid model is presented in this publication. The research results are part of the work under the project entitled "Voltage fluctuation diagnostic focused on identification and localization disturbing loads in power grids" funded by the National Science Centre, Poland - 2021/41/N/ST7/00397.</p>
Processed data for "Model identification of neural encoding (MINE)" publication
<p>This dataset contains mouse and zebrafish data processed by MINE. These datafiles were used to generate the publication figures for the mouse cortical dataset [m<em>usall.hdf5</em>] (Figure 5) and the zebrafish whole-brain [<em>main_analysis.hdf5</em>] (Figures 6 and 7) and reticulospinal datasets [r<em>s_analysis.hdf5</em>] (Figure 6).</p> <p> </p> <p><em>Musall.hdf5 </em>contains reordered data from "Musall, S., Kaufman, M.T., Juavinett, A.L. <em>et al.</em> Single-trial neural dynamics are dominated by richly varied movements. <em>Nat Neurosci</em> <strong>22</strong>, 1677–1686 (2019)."</p> <p>The contents of each dataset are described in <em>DataContent_xxx.pdf</em></p>
Pinterest dataset for age and gender identification in author profiling
<p>This dataset was used for the experiments presented in the article "Reconstructive Classification for Age and Gender Identification in Social Networks" - IEEE Transactions on Computational Social Systems</p> <p>The dataset contains text data from 548,761 pins corresponding to 264 users of Pinterest.</p> <p>There are 7 files.</p> <p>The first 5 files correspond to the extracted textual features from the pins that are aggregated per user: ats, emojis/emoticons, hashtags, links, and words.</p> <p>There are 264 lines in each file (one per user), as the concatenation of the extracted features from all the pins corresponding to each user.</p> <p>The last 2 files are the labels for the age and gender of the users. There are also 264 lines (one per user).</p> <p>For age, there are 4 possible labels: 18-24, 25-34, 35-46, and 50+</p> <p>For gender, there are 2 possible labels: F and M</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.