Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
189
datasets available to search
ShareScore release 0.9.0
Dataset results
189 results for “feature model”
Supplementary table, figures and DNA sequences of sorghum gene models SbiRTx430.01G455400 and SbiRTx.02G006600 that feature primers, gRNAs and indels created
Open the record for dataset details and reuse information.
Predictor complexity and feature selection affect Maxent model transferability: evidence from global freshwater invasive species
Open the record for dataset details and reuse information.
Data from: Structural and functional features of medium spiny neurons in the BACHDΔN17 mouse model of Huntington’s disease
Open the record for dataset details and reuse information.
Hyperbrain features of team mental models within a juggling paradigm: a proof of concept
<p>To capture the neural schemas underlying the notion of shared and complementary mental models, we examined the functional connectivity patterns and hyperbrain features of a juggling dyad involved in cooperative motor tasks of increasing difficulty. Jugglers' cortical activity was measured using two synchronized 32-channel EEG systems during dyadic juggling performed with 3, 4, 5 and 6 balls. Individual and hyperbrain functional connections were quantified through coherence maps calculated across all electrode pairs in the theta and alpha bands (4-8 Hz and 8-12 Hz). Graph metrics were used to typify the topology and efficiency of the functional networks.</p> <p>The datasets uploaded are the two dataset of both jugglers used for this study.</p>
Best learned models : Land Cover Classification with Gaussian Processes using spatio-spectro-temporal features
<p>Best learned models (Gaussian Processes, Random Forest, Multilayer Perceptron and Lightweight Temporal Self-Attention models) for each region based on the classification data set DS-A with seed 0 (see description <a href="https://zenodo.org/deposit/7099785">here</a>). </p><p>Models: GP non spatial, GP spatial (sum), GP spatial (product),RF non spatial, RF spatial, MLP non spatial, MLP spatial, LTAE non spatial, LTAE spatial</p><p>For further details see section VI-C of the pre-print article "Land Cover Classification with Gaussian Processes using spatio-spectro-temporal features ". This article is available <a href="https://hal.archives-ouvertes.fr/hal-03781332">here</a>.</p><p>The implementation of the models is available in the <a href="https://gitlab.cesbio.omp.eu/belletv/land_cover_southfrance_gp">open source repository</a>.</p>
SAMPLER representations of FFPE TCGA-lung WSIs using tile-level features of the MMIL-Transformer model
<p>Here we provide single-scale SAMPLER representations of the FFPE TCGA-lung (LUAD and LUSC) WSIs using tile-level features provided in https://github.com/hustvl/MMIL-Transformer. To learn more about SAMPLER please visit https://github.com/TheJacksonLaboratory/SAMPLER.</p><p>The SAMPLER representations are provided as a single python pickle file. This pickle file contains a dictionary where each key is a WSI ID and each entry is the SAMPLER representation of the WSI.</p>
Evaluating Feature Attribution Methods in the Image Domain: Benchmark results and model parameters
<p>This dataset contains the experimental results, adversarial patches and model parameters used in the paper Evaluating Feature Attribution Methods in the Image Domain.</p>
Data for DualNetGO: A Dual Network Model for Protein Function Prediction via Effective Feature Selection
<p>Data used in the paper, including annotation files, graph embeddings from TransformerAE, and protein attributes for both human and mouse, and for cafa3 data. Extract and place them in the <em>data </em>folder.</p>
Original features and code of clinical prediction model
Open the record for dataset details and reuse information.
How Configurable is the Linux Kernel? Analyzing Two Decades of Feature-Model History
<p>Reproduction package for the TOSEM'25 paper "How Configurable is the Linux Kernel? Analyzing Two Decades of Feature-Model History"</p>
Predictive modeling for clinical features associated with Neurofibromatosis Type 1
<p>Objective: Perform a longitudinal analysis of clinical features associated with Neurofibromatosis Type 1 (NF1) based on demographic and clinical characteristics, and to apply a machine learning strategy to determine feasibility of developing exploratory predictive models of optic pathway glioma (OPG) and attention-deficit/hyperactivity disorder (ADHD) in a pediatric NF1 cohort.</p> <p>Methods: Using NF1 as a model system, we perform retrospective data analyses utilizing a manually-curated NF1 clinical registry and electronic health record (EHR) information, and develop machine-learning models. Data for 798 individuals were available, with 578 comprising the pediatric cohort used for analysis.</p> <p>Results: Males and females were evenly represented in the cohort. White children were more likely to develop OPG (OR: 2.11, 95%CI: 1.11-4.00, p=0.02) relative to their non-white peers. Median age at diagnosis of OPG was 6.5 years (1.7-17.0), irrespective of sex. Males were more likely than females to have a diagnosis of ADHD (OR: 1.90, 95%CI: 1.33-2.70, p<0.001), and earlier diagnosis in males relative to females was observed. The gradient boosting classification model predicted diagnosis of ADHD with an AUROC of 0.74, and predicted diagnosis of OPG with an AUROC of 0.82.</p> <p>Conclusions: Using readily available clinical and EHR data, we successfully recapitulated several important and clinically-relevant patterns in NF1 semiology specifically based on demographic and clinical characteristics. Naïve machine learning techniques can be potentially used to develop and validate predictive phenotype complexes applicable to risk stratification and disease management in NF1.</p>
Blinded Predictions and Post-hoc Analysis of the Second Solubility Challenge Data: Exploring Training Data and Feature Set Selection for Machine and Deep Learning Models
<p>Training and test datasets and scripts for training models.</p>
Artifact for "Towards Deterministic Compilation of Binary Decision Diagrams From Feature Models"
<p>This artifact supplies the tooling and original data used in the evaluation of the corresponding submission #52 `Towards Deterministic Compilation of Binary Decision Diagrams From Feature Models'' accepted at SPLC'24. The artifact allows reproducing all the figures and tables of the submission. In addition, the artifact allows for easy replicating of the results for additional input models.</p>
Trained Random Forest model and scaler parameters on new physical and tsfel features from seismic data of 150s length.
Open the record for dataset details and reuse information.
Trained random forest models on 5000 traces per class based on updated features
Open the record for dataset details and reuse information.
The features of Tissues and Patches for "Predicting microsatellite instabilitiy from histology images with a three-level hierarchical graph fusion model"
<p>This repository contains features and corresponding coordinates of patches and tissues extracted from 430 and 326 histologic images from patients with colorectal and gastric cancers from the TCGA cohort (original whole section SVS images are freely available at https://portal.gdc.cancer.gov/). All images in this library are from formalin-fixed paraffin-embedded (FFPE) diagnostic sections (“DX” on the GDC Data Portal). This blog explains this in detail: http://www.andrewjanowczyk.com/download-tcga-digital-pathology-images-ffpe/</p> <p><strong>Preprocessing.</strong></p> <p>All SVS slices were pre-processed as follows.</p> <p>According to “Deep learning can predict microsatellite instability directly from histology in gastrointestinal cancer” these histology images were categorized into The histology images were classified as “MSS” (microsatellite stable) or “MSIMUT” (microsatellite unstable or highly mutated) according to “Deep learning can predict microsatellite instability directly from histology in gastrointestinal cancer”, which corresponds to the division of the training and test sets in the article.<br><br></p> <p>Patches were extracted at 40x objective magnification and 20x objective magnification, respectively, and the corresponding features were extracted by pre-training resnet48, respectively</p> <p>The features of Tissues are thumbnails obtained at 2.5x objective magnification and further extracted by MedSAM after extracting the masks of the tissues.</p>
Cell features for "Datasets for "Predicting microsatellite instabilitiy from histology images with a three-level hierarchical graph fusion model""
<p>This repository contains features and corresponding coordinates of cells extracted from 430 and 326 histologic images from patients with colorectal and gastric cancers from the TCGA cohort (original whole section SVS images are freely available at https://portal.gdc.cancer.gov/). All images in this library are from formalin-fixed paraffin-embedded (FFPE) diagnostic sections (“DX” on the GDC Data Portal). This blog explains this in detail: http://www.andrewjanowczyk.com/download-tcga-digital-pathology-images-ffpe/</p> <p><strong>Preprocessing.</strong></p> <p>All SVS slices were pre-processed as follows.</p> <p>According to “Deep learning can predict microsatellite instability directly from histology in gastrointestinal cancer” these histology images were categorized into “MSS” (microsatellite stable) or “MSIMUT” (microsatellite unstable or highly mutated) and corresponded to the article dividing the training and test sets.<br><br></p> <p>The features of all cells were extracted by Hovernet and Transnuseg at 40x objective magnification for extraction masking and further feature extraction</p>
Dataset for ASE'24 Efficient Slicing of Feature Models via Projected d-DNNF Compilation
<p>This dataset includes various feature models as dimacs. Each dimacs includes a header indicating variables to be projected.</p> <p>The dataset was used for evaluating pd4 within the work Efficient Slicing of Feature Models via Projected d-DNNF Compilation at ASE'24.</p>
Essential Features of Escalated Force and Negotiated Management Models
<p><strong><span>Skrzypek, Maciej, Essential Features of Escalated Force and Negotiated Management Models</span></strong></p> <p><strong><span>This dataset was elaborated for the research project <em>Civil Disorder in Pandemic-ridden European Union</em>. The latter was financially supported by the National Science Centre, Poland [grant number 2021/43/B/HS5/00290].</span></strong></p>
Haploid, diploid, and pooled exome capture recapitulate features of biology and paralogy in two non-model tree species
<p>Despite their suitability for studying evolution, many conifer species have large and repetitive giga-genomes (16-31Gbp) that create hurdles to producing high coverage SNP datasets that capture diversity from across the entirety of the genome. Due in part to multiple ancient whole genome duplication events, gene family expansion and subsequent evolution within <i>Pinaceae</i>, false diversity from the misalignment of paralog copies creates further challenges in accurately and reproducibly inferring evolutionary history from sequence data. Here, we leverage the cost-saving benefits of pool-seq and exome-capture to discover SNPs in two conifer species, Douglas-fir (<i>Pseudotsuga menziesii</i> var. <i>menziesii </i>(Mirb.) Franco, <i>Pinaceae</i>) and jack pine (<i>Pinus banksiana</i> Lamb., <i>Pinaceae</i>). We show, using minimal baseline filtering, that allele frequencies estimated from pooled individuals show a strong positive correlation with those estimated by sequencing the same population as individuals (r > 0.948), on par with such comparisons made in model organisms. Further, we highlight the utility of haploid megagametophyte tissue for identifying sites that are likely due to misaligned paralogs. Together with additional minor filtering, we show that it is possible to remove many of the loci with large frequency estimate discrepancies between individual and pooled sequencing approaches, improving the correlation further (r > 0.973). Our work addresses bioinformatic challenges in non-model organisms with large and complex genomes, highlights the use of megagametophyte tissue for the identification of paralog sites, and suggests the combination of pool-seq and exome capture to be robust for further evolutionary hypothesis testing in these systems.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.