Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,047

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

1,047 results for “interpretability”

Learn how ShareScore rates datasets ↗
zenodo44/100

A Bioconductor workflow for processing, evaluating and interpreting expression proteomics data

<p>Files for users of the workflow "A Bioconductor workflow for processing, evaluating and interpreting expression proteomics data". Files include Proteome Discoverer (v2.5) processing and consensus workflows for both TMT and LFQ expression proteomics data. Also provided are the output .txt files of a corresponding Proteome Discoverer identification search, as required for users to follow the workflow themselves. For raw data please refer to PRIDE. Appendix is provided as a PDF.</p>

opencc-by-4.0Apr 2023View details →
zenodo44/100

Speech Corpus of Interpreted Premier Press Conferences (SCIPPC)

<p>SCIPPC v1.0 is a parallel corpus of consecutive interpreting between Mandarin Chinese and English and vice versa in two Chinese premiers&rsquo; press conferences in March 2003&ndash;2007 and 2013&ndash;2017. The conferences were held after sessions of the National People&rsquo;s Congress and the Chinese People&rsquo;s Political Consultative Conference. They were moderated by spokespersons of the Congress and the Chinese Ministry of Foreign Affairs and attended by journalists, who asked the premiers questions.</p> <p>SCIPPC v1.0 includes source speeches by approximately 170 speakers and interpretations by six different staff interpreters of the Chinese Ministry of Foreign Affairs, who worked into their B language. It contains 192,209 tokens (source: 108,296, target: 83,913; Chinese: 112,528, English: 79,681) and 19 h 43 min 5 s of video recordings. It is fully transcribed and aligned at the recording&ndash;transcript and source&ndash;target transcript levels.</p>

opencc-by-sa-4.0Sep 2024View details →
zenodo44/100

Improvement of regulations interpretation and formalisation for information need definition - Municipality of Ascoli Piceno, Italy

<p>CHEK Digital Building Permit Maturity Model (CDBPMM) as developed within the HORIZON EUROPE project 'Change toolkit for Digital Building Permit'.</p> <p>(CHEK)&nbsp;https://chekdbp.eu&nbsp;</p> <p>It is described in the CHEK project deliverable D2.1.</p> <p>This project has received funding from the European Union's Horizon Europe program under Grant Agreement No.101058559.</p> <p>The aim of CHEK is to remove barriers preventing municipalities from adopting digital building permit processes by developing, connecting, and aligning scalable solutions in the regulatory and policy context, in open standards and interoperability (geospatial and BIM), in closing knowledge gaps through education, in renewing municipal processes, and in deploying technology.&nbsp;</p>

opencc-by-4.0Sep 2024View details →
zenodo44/100

A foundation model enpowered by a multi-modal prompt engine for universal seismic geobody interpretation across surveys

<p>A multi-type geobody dataset for training SAG model, including channel, paloekarst, salt body, and so on.</p> <p>A foundation model enpowered by a multi-modal prompt engine for universal seismic geobody interpretation across surveys (<a href="https://arxiv.org/abs/2409.04962">[2409.04962] A foundation model enpowered by a multi-modal prompt engine for universal seismic geobody interpretation across surveys (arxiv.org)</a>)</p> <p>&nbsp;</p>

opencc-by-4.0Oct 2024View details →
zenodo44/100

Interpretable prediction for anticancer sensitivity of glycoside amides

<p><a href="https://anti-cancer.eu/" target="_blank" rel="noopener">Development Timeline: Selective Anticancer Logic of Glycoside Amides</a></p> <p>/ Second supplemented edition /</p> <h2>📊 Interpretable Prediction Dataset:</h2> <h3>Transparent Modeling of Anticancer Sensitivity to Glycoside Amides</h3> <p>The current dataset presents the <strong>exact results</strong> of our interpretable prediction model for anticancer sensitivity to glycoside amides. It is provided in *.xlsx format and includes:</p> <ul> <li> <p>✅ Analytical data</p> </li> <li> <p>✅ Theoretical framework</p> </li> <li> <p>✅ Authorial conclusions</p> </li> <li> <p>✅ Full filtered dataset</p> </li> <li> <p>✅ Complete raw data</p> </li> </ul> <p>All information is organized in a <strong>user-friendly structure</strong>, fully compatible with standard data export formats and ready for integration into clinical modeling, pharmaceutical analysis, or transcriptomic mapping.</p> <div>&nbsp;</div> <h3>🔍 Transparency and Scientific Integrity</h3> <p>This dataset is not a closed interpretation. It reflects <strong>our original research findings</strong>, derived from a specific theoretical and biochemical framework. We fully acknowledge that the data may be interpreted differently depending on the analytical model, clinical context, or pharmacological assumptions.</p> <p>That is precisely why we have chosen to publish the <strong>exact numerical calculations</strong>&mdash;not just summaries or visualizations. This decision underscores our commitment to <strong>transparency</strong>, <strong>scientific reproducibility</strong>, and <strong>open dialogue</strong> with the broader research community.</p> <div>&nbsp;</div> <h3>🧠 A Platform for Collaboration</h3> <p>We invite clinicians, researchers, and data scientists to explore the dataset, challenge its assumptions, and build upon its structure. Whether used for comparative modeling, transcriptomic validation, or therapeutic design, the data is intended to serve as a <strong>foundation for further inquiry</strong>, not a final verdict.</p>

opencc-by-4.0Jan 2023View details →
zenodo44/100

Practitioners Interpretation of conditional Requirements

<p>This dataset contains (1) the survey protocol and (2) the survey responses of the study on how practitioners interpret conditional requirements. A readme provides guidance on how to comprehend the data.</p>

opencc-by-4.0Jun 2021View details →
zenodo44/100

Multi-decade land use and land cover samples for Brazil based in a stratified sampling design and visual interpretation of Landsat data (1985 — 2018)

<p>This dataset is composed&nbsp;by 85,152 random points throughout the Brazilian territory selected according to a stratified sampling design, based in&nbsp;127 regular&nbsp;regions&nbsp;and six&nbsp;slope classes&nbsp;(<a href="https://www.usgs.gov/centers/eros/science/usgs-eros-archive-digital-elevation-shuttle-radar-topography-mission-srtm-1-arc?qt-science_center_objects=0#qt-science_center_objects">SRTM</a>). Each sample was visually inspected by three independent&nbsp;interpreters, which associated all the land use and land cover (LULC)&nbsp;changes between 1985 and 2018, on a <strong>yearly basis</strong>,&nbsp;using as reference two <strong>Landsat</strong> images per year, a <strong>MODIS</strong> NDVI time series and&nbsp;high resolution images from <strong>Google Earth</strong>.&nbsp;</p> <p>This&nbsp;process was guided by a <a href="https://www.lapig.iesa.ufg.br/chave/">reference labeling protocol</a> which established the follow LULC classes:</p> <ul> <li><strong>Annual crop:</strong> Areas occupied with short to medium-term crops, usually with a vegetative cycle of less than one year, which after harvest needs to be re-planted.&nbsp;</li> <li><strong>Aquaculture:</strong> Artificial lakes, where aquaculture and/or salt production activities predominate</li> <li><strong>Beach and dune (Other):</strong> Sandy areas, with bright white color, where there is no vegetation predominance of any kind.</li> <li><strong>Forest formation:</strong> Vegetation types with predominance of tree species, with continuous canopy formation</li> <li><strong>Grassland formation:</strong> Grassland formations with predominance of herbaceous stratum</li> <li><strong>Mangrove (Other):</strong> Dense and Evergreen Forest formations, often flooded by tide and associated with the mangrove coastal ecosystem.</li> <li><strong>Mining (Other):</strong> Areas where clear signs of extensive mineral extractions are present, shows clear exposure of the soil by the action of heavy machinery. Only regions surrounding the AhkBrasilien (AHK) and the CPRM digital reference data were considered.</li> <li><strong>Not observed:</strong> Areas blocked by clouds or atmospheric noise, or with absence of ground observation masked out from analysis.</li> <li><strong>Other non-forest natural formations:</strong> Marshes (with fluvio-marine influence).</li> <li><strong>Other non-vegetated area (Other):</strong> Non-permeable surface areas (infrastructure, urban expansion or mining) not mapped into their classes</li> <li><strong>Pasture:</strong> Pasture areas, natural or planted, related with farming activity. In particular in the Pampa and Pantanal biomes part of the area classified as Grassland Formation also includes pasture areas.</li> <li><strong>Perennial crop:</strong> Areas occupied with crops with a long cycle (more than one year), which allow successive harvests without the need for new crop.&nbsp;</li> <li><strong>Rocky outcrop (Other)</strong>: Naturally exposed rocks without soil cover, often with the partial presence of rupicolous vegetation and high slope.&nbsp;</li> <li><strong>Salt flat (Other):</strong> &quot;Apicuns&quot; or Salt flats are formations often without tree vegetation, associated to a higher, hypersaline and less flooded area in the mangrove, generally in the transition between this area and the continent.</li> <li><strong>Savanna formation:</strong> Savanna formations with defined tree and shrub-herbaceous stratum</li> <li><strong>Semi-perennial crop:</strong> Cultivated areas with sugar cane</li> <li><strong>Tree plantation:</strong> Planted tree species for commercial use (e.g. Eucalyptus, Pinus and Araucaria)</li> <li><strong>Urban infrastructure:</strong> Urban areas with predominance of non-vegetated surfaces, including roads, highways and constructions.</li> <li><strong>Water:</strong> Rivers, lakes, dams, reservoir and other water bodies</li> <li><strong>Wetland:</strong> Wetlands with fluvial influence or swampy areas</li> </ul> <p>To enable a proper area estimation and accuracy assessment (<a href="https://www.tandfonline.com/doi/abs/10.1080/01431161.2014.930207">Stehman, 2014</a>) the dataset is provided with the&nbsp;<strong>sampling probability</strong> for each sample (<em>brazil_lulc_samples_1985_2018</em> and <em>brazil_lulc_samples_1985_2018_row_wise</em>)&nbsp;and the <strong>sampling weight</strong> (<em>brazil_lulc_samples_1985_2018_row_wise</em>), which was adjusted to disregard the &quot;<strong>Not observed&quot; </strong>class. The number of votes for the associated LULC class (visual interpretation agreement) and an indication if the sample is between two different LULC<strong> </strong>classes (<strong>border flag</strong>) are also provided.</p> <p>The samples were used to produce&nbsp;several&nbsp;<strong><a href="https://github.com/lapig-ufg/tvi-analysis">area estimation analyses</a></strong>, including&nbsp;land use and land cover dynamics, historical deforestation and agricultural expansion of Brazil. A publication describing in detail the methodology and the analysis&nbsp;is under preparation.</p>

opencc-by-4.0Jul 2021View details →
zenodo44/100

Extracting interpretable rules with Bayesian Networks. A case study of intrinsic human hazardous properties of silver nanoforms for the Safety Dimension of Safe and Sustainable by design paradigm.

<p>Three different datasets: toxicological attributes in i) lung and ii) intestinal cell line along with system dependent features and iii) system independent pchem properties) were merged. Each row represents one set of experimental testing conditions and related system dependent nanodescriptors based on the exposure dose and NFs pre-treatment (for intestinal assessments). The system independent inputs are NF specific and independent of experimental conditions. Data is captured via FAIR principles where the reader can find the origin (institution) of each data, the responsible data creators (experimentalists), the raw measurements, the protocols followed and the instrumentations used for each experiment. .</p>

opencc-by-4.0Nov 2022View details →
zenodo44/100

Song Interpretation Dataset

<p>The Song Interpretation Dataset combines data from two sources: (1) music and metadata from the Music4All Dataset and (2) lyrics and user interpretations from SongMeanings.com. We design a music metadata-based matching algorithm that aligns matching items in the two datasets with each other. In the end, we successfully match 25.47% of the tracks in the Music4All Dataset.</p> <p>The dataset contains audio excerpts from 27,834 songs (30 seconds each, recorded at 44.1 kHz), the corresponding music metadata, about 490,000 user interpretations of the lyric text, and the number of votes given for each of these user interpretations. The average length of the interpretations is 97 words. Music in the dataset covers various genres, of which the top 5 are: Rock (11,626), Pop (6,071), Metal (2,516), Electronic (2,213) and Folk (1,760).&nbsp;</p> <p>For more details, please refer to our paper &quot;Interpreting Song Lyrics with an Audio-Informed Pre-trained Language Model&quot;.</p>

opencc-by-4.0May 2022View details →
zenodo44/100

MammoTab 22: a giant and comprehensive dataset for Semantic Table Interpretation

<p>MammoTab is a dataset designed to evaluate semantic table annotation approaches.</p> <p>It includes two types of annotation:</p> <ol> <li>cell/mentions to Knowledge Graph (KG) entity matching (CEA task) and;</li> <li>column to KG&nbsp;class matching (CTA task).</li> </ol> <p>It is composed of 980254 tables extracted from 21149260 Wikipedia pages and annotated through Wikidata v. 20220708. The dataset is compliant with the data format used in&nbsp;<a href="https://www.cs.ox.ac.uk/isg/challenges/sem-tab/2019/index.html">SemTab2019</a>.</p>

opencc-by-4.0Aug 2022View details →
zenodo44/100

FIND: A Function Interpretation Dataset and Benchmark for Evaluating Interpretability Methods

<p><strong>FIND</strong> is an&nbsp;interactive dataset&nbsp;for evaluating AI interpretability methods on black box functions.&nbsp;</p> <p>This dataset contains all function files for the <strong>FIND</strong> benchmark and JSON files with associated metadata. The utilities provided in the associated&nbsp;<strong>FIND</strong>&nbsp;<a href="https://github.com/multimodal-interpretability/FIND">GitHub Repository</a>&nbsp;support running and evaluating&nbsp;interpretation of the functions with user-defined&nbsp;interpreters.</p>

openmit-licenseJun 2023View details →
zenodo44/100

Report on Transformers interpretability for Natural Language Processing: A case study on Technical Debt classification

<p>Transformer models have significantly advanced the field of natural language processing (NLP), achieving exceptional results in various tasks. However, these models are often seen as &quot;black boxes&quot;, providing limited insight into the factors influencing their predictions. It has become crucial to develop and utilise methods for interpreting and explaining these models to uncover their complex inner workings. This report discusses the latest techniques and tools that aid in a more profound understanding of transformer models within NLP. Additionally, it explores a vital industrial use case: Technical Debt (TD) classification. In this context, the report leverages transformer model interpretability tools and Retrieval Augmented Generation (RAG) to analyse and understand the characteristics of text in Github issues, distinguishing between TD and non-TD.</p> <p>This report thoroughly outlines an approach to improve the transparency and reproducibility of machine learning models, with a special emphasis on TD classification. It integrates the RAG approach and exploits feature attribution techniques, presenting a route to create AI systems that are not only high-performing but also demonstrably trustworthy and comprehensible. Through a detailed examination of word patterns in TD classification and the innovative use of the RAG approach, the research highlights a strong dedication to promoting transparency and responsibility in AI systems, potentially ushering in a new phase in machine learning research that focuses on clarity and dependability.</p>

opencc-by-4.0Sep 2023View details →
zenodo44/100

Resolving the Interpretation of Magnetic Coercivity Components from Backfield Isothermal Remanence Curves Using Unmixing of Non-linear Preisach Maps: Application to Loess-Paleosol Sequences

<p>The data set includes:</p> <p>1. Non-linear Preisach measurements for Lunca and Costinești loess-paleosol sections</p> <p>2. IRM coercivity distributions from Lunca and Costinesti interpolated on a common sequence of fields</p> <p>3. Lunca granulometry data</p> <p>4. Median Grain size for Costineși section.</p> <p>5. Magnetic susceptibility data measured at Lunca section (Constantin et al., 2015)</p> <p>6. IRM acquisition curves derived from backfield IRM data through rescaling for Costinesti section</p> <p>7. Costinesti rock magnetic data (Necula et al., 2015)</p>

opencc-by-4.0May 2023View details →
edi44/100

Interpreting the smells of predation: How alarm cues and kairomones induce different prey defenses.

1. For phenotypically plastic organisms to produce phenotypes that are well matched to their environment, they must acquire information about their environment. For inducible defences, cues from damaged prey and cues from predators both have the potential to provide important information, yet we know little about the relative importance of these separate sources of information for behavioural and morphological defences. We also do not know the point during a predation event at which kairomones are produced, i.e. whether they are produced constitutively, during prey attack or during prey digestion. 2. We exposed leopard frog tadpoles (Rana pipiens) to nine predator cue treatments involving several combinations of cues from damaged conspecifics or heterospecifics, starved predators, predators only chewing prey, predators only digesting prey or predators chewing and digesting prey. 3. We quantified two behavioural defences. Tadpole hiding behaviour was induced only by cues from crushed tadpoles. Reduced tadpole activity was induced only by cues from predators digesting tadpoles or predators chewing + digesting tadpoles. 4. We also quantified tadpole mass and two size-adjusted morphological traits that are known to be phenotypically plastic. Mass was unaffected by the cue treatments. Relative body length was affected (i.e. there were differences among some treatments), but none of the treatments significantly differed from the no-predator control. Relative tail depth was affected by the treatments and deeper tails were induced only when tadpoles were exposed to cues from predators digesting tadpoles or cues from predators chewing + digesting tadpoles. 5. These results demonstrate that some prey species can discriminate among a diverse set of potential cues from heterospecific prey, conspecific prey and predators. Moreover, the results illustrate that the cues responsible for the full suite of behavioural and morphological defences are not induced by tadpole crushing nor

openCC (other)Jun 2024View details →
zenodo40/100

Archäologische Chronologie und historische Interpretation: Die Merowingerzeit in Süddeutschland (Correspondence Analysis Data Set)

<p>This data set is a supplement to the book &quot;Arch&auml;ologische Chronologie und historische Interpretation: Die Merowingerzeit in S&uuml;ddeutschland&quot; (De Gruyter, 2016) and comprises the archaeological data and the results of the correspondence analysis of Merovingian-period graves from southern Germany and their chronological classification. The data sets for female and male burials can be downloaded as PDF, EXCEL and CSV files.</p>

opencc-by-nc-4.0Aug 2016View details →
zenodo40/100

cigKast: A data of 3D synthetic seismic volumes with labeled paleokarsts for deep-learning-based paleokarst interpretation

<p>cigKarst is a dataset created by the <a href="http://cig.ustc.edu.cn/">Computational Interpretation Group (CIG)</a> for the deep-learning-based peleokarst interpretation in 3D seismic images, <a href="http://cig.ustc.edu.cn/xinming/list.htm" target="_blank" rel="noopener">Xinming Wu</a> is the main contributor to the dataset.</p> <p>This dataset contains 120 pairs of synthetic 3D seismic images and the corresponding label images with the ground truth of the paleokarst systems simulated in the seismic images. More detail of building this dataset is discussed in the paper published at the journal of JGR Solid Earth:</p> <p><strong>Wu, X.</strong>, S. Yan, J. Qi, and H. Zeng, 2020, Deep learning for characterizing paleokarst collapse features in 3D seismic images.&nbsp;<strong>JGR, Solid Earth</strong>, Vol. 125(9), 1-23, e2020JB019685.&nbsp;<a href="http://cig.ustc.edu.cn/_upload/tpl/05/cd/1485/template1485/papers/wu2020karst.pdf">[PDF]</a>. doi: 10.1029/2020JB019685</p> <p>Below are some brief description of the dataset:</p> <p>1) The "seismic.zip" contains 120 3D seismic images, each image is with the dimension of 256X256X256;</p> <p>&nbsp;2) The "karst.zip" contains 120 3D label images of the karsts. Each label image is with the same dimension of 256X256X256. The values in a label image are set with ones in the karst areas while zeros elsewhere, which is why the compressed label images in the karst.zip is much smaller than the&nbsp;seismic images compressed in the seismic.zip</p>

opencc-by-4.0Nov 2020View details →
zenodo40/100

Dataset and Jupyter worksheet interpreting the (results from) small- and wide-angle scattering data from a series of boehmite/epoxy nanocomposites. Accompanies the publication "Competition of nanoparticle-induced mobilization and immobilization effects on segmental dynamics of an epoxy-based nanocomposite"

<p>Dataset and Jupyter worksheet interpreting the (results from) small- and wide-angle scattering data from a series of boehmite/epoxy nanocomposites. Accompanies the publication &quot;Competition of nanoparticle-induced mobilization and immobilization effects on segmental dynamics of an epoxy-based nanocomposite&quot;, by Paulina Szymoniak, Brian R. Pauw, Xintong Qu, and Andreas Sch&ouml;nhals.</p> <p>Datasets are in three-column ascii (processed and azimuthally averaged data) from a Xenocs NanoInXider SW&nbsp;instrument. Monte-Carlo analyses were performed using McSAS 1.3.1, other analyses are in the Python 3.7&nbsp;worksheet. Graphics and result tables are output by the worksheet.&nbsp;</p>

opencc-by-4.0Dec 2020View details →
zenodo40/100

FIGURES 10 – 15 in Mature larva of Stenichnus godarti (Latreille) (Coleoptera: Staphylinidae, Scydmaeninae): redescription, hypothesis of displaced epicranial suture and alternative interpretation of homology between chaetotaxic structures

FIGURES 10 – 15. Larva of Stenichnus godarti. Head in dorsal (10) and ventral (11) views; left (12) and right (13) antenna in dorsal view; right (14) and left (15) maxilla in ventral view. Abbreviations: Ag, antennal gland; An 1 – 3, antennomere I – III; Cd, cardo; Da, dorsoanterior seta; De, dorsoepicranial seta; Df, dorsofrontal seta; Dl, dorsolateral seta; Dp, dorsoposterior seta; Es, frontal arm of epicranial suture; Est, epicranial stem; L, lateral seta; l, lentiform structure; La, labral anterior seta; Ma, mala; Md. mandible; MdS, mandibular seta; Mn, mentum; Mxp 1 – 3, maxillary palpomere II – III; Ptp, posterior tentorial pit; SA, sensory appendage; Smn, submentum; sol, solenidion; St, stemma; Stp, stipes; V, ventral seta; Va, ventroanterior seta; Vl, ventrolateral seta.

opencc-zeroDec 2016View details →
zenodo40/100

Fig. 7 in A New Interpretation of the Oldest Fossil Bee (Hymenoptera: Apidae)

Fig. 7. Left lateral habitus illustration of holotype worker of Cretotrigona prisca (illustration courtesy of D. A. Grimaldi, AMNH).

opencc-by-4.0Apr 2000View details →
zenodo40/100

Fig. 6 in A New Interpretation of the Oldest Fossil Bee (Hymenoptera: Apidae)

Fig. 6. Artist's reconstruction of Cretotrigona prisca as it might have appeared in flight (painting courtesy of Michael Rothman).

opencc-by-4.0Apr 2000View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record