Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
69
datasets available to search
ShareScore release 0.7.1
Dataset results
69 results for “feature selection”
Summary of the most important features for selected ABs
<p>These data summarizes the relevant findings and the identified limitations (in terms of "Category", "Technology", "Properties", "Limitation", and "Applicability to railway"), coming from the overview of different Alternative Bearers (ABs), carried out in deliverable D21 (AB4Rail project, www.ab4rail.eu).<br> The results have provided an overview of several technologies, each of them showing specific characteristics. The heterogeneous nature of different ABs allows to provide a plethora of available communication technologies to be potentially used by the Adaptable Communication System (ACS) for different railway scenarios. All the selected ABs provide the IP interconnection feature since they are Integrated within OSI reference model.<br> In this way, it collects the planned objectives of deliverable D2.1, expressed as a technological overview of selected ABs, as possible candidates coexisting with Traditional Bearers (TBs) for supporting railway applications.</p>
Feature selection on microbial profiles of CRC samples with chopin2 (powered by hdlib)
<p>This Zenodo entry contains the result of the feature selection algorithm implemented through a backward variable elimination strategy in <a href="https://github.com/cumbof/chopin2" target="_blank" rel="noopener">chopin2</a> (powered by <a href="https://github.com/cumbof/hdlib" target="_blank" rel="noopener">hdlib</a>) applied on <a href="https://github.com/biobakery/MetaPhlAn" target="_blank" rel="noopener">MetaPhlAn3</a> microbial profiles of a public dataset of metagenomic stool samples collected from patients affected by the colorectal cancer (CRC) as well as from healthy individuals.</p> <p>Microbial profiles have been extracted through the <a href="https://bioconductor.org/packages/release/data/experiment/html/curatedMetagenomicData.html" target="_blank" rel="noopener">curatedMetagenomicData</a> package for R under the IDs <em>ThomasAM_2018a</em>, <em>ThomasAM_2018b</em>, and <em>ThomasAM_2019_a</em>.</p> <p>The feature selection algorithm is implemented as a backward variable elimination method, and it makes use of the vector-symbolic architecture described in <a href="https://doi.org/10.3390/a13090233" target="_blank" rel="noopener">Cumbo F 2020</a>.</p> <p>Deposited data is described below:</p> <ul> <li><em>datasets.tar.gz</em>: it contains the datasets used as input of <em>chopin2</em> as the result of merging the three datasets with relative abundances mentioned above, also stratified by age and sex (with prefix RA). The same datasets have been also binarized (with prefix BIN);</li> <li><em>hd-models.tar.gz</em>: it contains the output of the feature selection performed with <em>chopin2</em> (powered by <em>hdlib</em>) on the datasets with both relative abundance and binary profiles (RA and BIN);</li> <li><em>ml-models.tar.gz</em>: it contains the result of the feature selection produced with classical wrapper-based techniques (i.e., Random Forest, Decision Tree, Support Vector Machine, Logistic Regression, and Extreme Gradient Boosting) in addition to a Python 3.8 script to reproduce the results.</li> </ul> <p>Please note that the datasets <em>RA__ThomasAM__species.csv</em> and <em>BIN__ThomasAM__species.csv</em> are also included into the <em>datasets.tar.gz</em> archive.</p>
Raw and processed hydro-meteorological variables of Jucar river basin for feature selection
<p>The dataset Processed data – input WQEISS.csv was employed for the input variable selection step in Zaniolo et al., 2018. It includes monthly values of 28 hydro-meteorological variables and indexes of Jucar river basin, Spain, for the period 1986-2000, namely:</p> <ul> <li>2 temporal features: day and month of the year;</li> <li>12 inputs to the Jucar State Index: average monthly storage and groundwater levels, average three months river runoff, and cumulated areal precipitation over 12 months;</li> <li>8 additional observed variables in the basin: three months average outflows from, and inflows to, the main reservoirs, and mean monthly areal temperatures;</li> <li>6 traditional drought indicators: Standardized Precipitation Index (SPI) and Standardized Precipitation and Evaporation Index (SPEI). SPI and SPEI indicators are computed on mean monthly data over the entire basin for 3, 6, and 12 months time aggregations.</li> </ul> <p>The last column of the dataset reports the target variable, i.e., the monthly nominal shortage of water conveyed to the irrigation districts simulated via AQUATOOL model. For further details on the dataset please consult Zaniolo et al., 2018, or the dedicated website <a href="http://www.nrm.deib.polimi.it/?page_id=2438">http://www.nrm.deib.polimi.it/?page_id=2438</a></p> <p>The unprocessed data used to compute indices and temporal cumulations in Processed data – input WQEISS.csv are reported in table Raw Data.csv. Public observations of rainfall, streamflows and storage levels come from the SAIH (Hydrological Automatic Information System) of the CHJ (Jucar Hydrological Confederation). Users can directly download data for the last 12 months on the dedicated webpage <a href="http://saih.chj.es/chj/saih/?f">http://saih.chj.es/chj/saih/?f</a> while previous data records are provided for free by CHJ upon request. Observations from piezometers are downloadable from the Piezometric Network Information section section of the CHJ <a href="https://www.chj.es/es-es/medioambiente/redescontrol/Paginas/Piezometr%C3%ADa.aspx">https://www.chj.es/es-es/medioambiente/redescontrol/Paginas/Piezometr%C3%ADa.aspx</a>.</p>
The features of the selected papers in the field of air quality prediction
<p>The table is a part of a submitted manuscript (Iskandaryan, D., Ramos, F., & Trilles, S. The Role of Datasets in Air Quality Prediction. Submitted to Atmosphere.) and includes the following features extracted from the selected papers: <em>Year, Case Study, Prediction Target, Dataset Type, Data Rate, Period (Days), Open Data, Algorithm, Time Granularity and Evaluation Metric</em>. The relevant papers were selected from a systematic review in <em>Air Quality Prediction Using Machine Learning Technologies. </em>The works were queried in Association for Computing Machinery, IEEE Xplore, Scopus and Web of Science databases using the following query: ("machine learning") AND ("prediction"OR "forecast") AND ("air quality" OR "air pollution"), which was being applied to title, abstract and keywords. After filtering the results guided by the Preferred Reporting Items for Systematic Reviews and Meta-Analyses, ninety-three papers were selected. The goal of this review is to understand which features are used in the field, in particular to answer the following questions: 1) What types of datasets are used to improve air quality predictions?; and 2) What characteristics of the dataset are important for efficient and effective air quality forecasting? <br> Twenty-six datasets were used by the authors as supplemental air quality data in order to predict air quality more accurately. Those datasets are: "MET"- meteorological data; "Spatial"- topographical characteristics, the locations of the stations; "Temporal"-includes the day of the month, day of the week, the hour of the day; "AOD"- aerosol optical depth; "Social Media"- microblog data; "Traffic"; "PBL Height"- planetary boundary layer height; "Land Use"; "BEV"- Built Environment Variables; "UV Index"; "SP"- Sound Pressure; "PD"-Population Density; "Human Movements"- floating population and estimated traffic volume; "Altitude"; "OMI-SO2"-Satellite-retrieved SO2 from Ozone Monitoring Instrument-SO2; "PPS"- Pollution Point Source; "TS"-Transportation Source; "WFD’"- weather forecast data; "POI Distribution"; "FAPE"- factory air pollution emission; "RND"- Road Network Distribution; "Elevation"; "AEI"- Anthropogenic Emission Inventory; "NDVI"; "Chemical"- chemical component forecast data (organic carbon, black carbon, sea salt, etc.); "Emission".</p>
Benchmarking eliminative radiomic feature selection for head and neck lymph node classification - Supplemental data
<p>Supplementary files for the publication "Benchmarking eliminative radiomic feature selection for head and neck lymph node classification"</p>
Dataset: Location- and feature-based selection histories make independent, qualitatively distinct contributions to urgent visuomotor performance
<p>This dataset (packaged as the zip file history_share.zip) accompanies the article titled "Location- and feature-based selection histories make independent, qualitatively distinct contributions to urgent visuomotor performance" by EE Oor, E Salinas, and TR Stanford which is available as a preprint in bioRxiv. The experimental results in the article are based on behavioral data collected from 2 monkey subjects during performance of a visuomotor task (the compelled oddball task), as described in the text. This dataset contains the trial-by-trial behavioral results collected for each subject and upon which all subsequent analyses were based.</p> <p>In addition to the trial-wise data arrays (stored in the files dataC.csv, dataN.csv, and dataCN.csv), the package includes Matlab functions and scripts (*.m files) used to analyze the data and recreate the results and figures in the article. Instructions and specifics are detailed in the README file. </p>
CLIP Features and Selected Relevance Judgments Subset for TRECVID Ad-hoc Search (2019-2023)
<div> <div> <div> <div> <div> </div> </div> </div> </div> </div> <div> <div> <div> <div> <div> <div> <p>This repository contains CLIP features and annotations for a subset of V3C images, based on their relevance to selected queries from the TREC Video Retrieval Evaluation (TRECVID) Ad-hoc Video Search (AVS) task. The data includes annotations for AVS queries and judgments conducted in TRECVID from 2019 to 2023 [1], using the V3C1 and V3C2 collections [2]. Specifically, the TRECVID-AVS collection covers 89 queries, with video shots manually labeled as relevant (1), non-relevant (0), or not annotated (-1).</p> <p>We used approximately 2.6 million keyframes extracted from these video shots, mapping the annotations to the corresponding keyframes (note that there may not be a one-to-one correspondence between TRECVID shotID since multiple frames might be extracted from a single shot). Image representations are based on CLIP ViT-H/14 - LAION-2B features [3]. The timestamps of the keyframes and their CLIP features are sourced from the VISIONE repository [4].</p> <p>Given the incomplete nature of the TRECVID ground truth (where only a subset of video segments were judged per query), we focused on queries with at least 200 positive and 1400 negative annotations. This resulted in 80 datasets, each containing 1500 images—10% labeled as relevant and 90% as non-relevant.</p> <h3>Contents of the Repository:</h3> <ol> <li> <p><strong>Query-specific CSV Files:</strong> For each of the 80 selected AVS query (e.g., <code>1591</code>), the corresponding CSV file (e.g., <code>1591.csv</code>) contains a column for each image, where:</p> <ul> <li><strong>VISIONE image ID</strong> is in the first row.</li> <li><strong>CLIP features</strong> are in the subsequent rows.</li> <li><strong>Relevance annotations</strong> are in the last row: <code>1</code> for relevant, <code>0</code> for non-relevant.</li> </ul> </li> <li> <p><strong>Post-processed Datasets:</strong></p> <ul> <li><code>dataset_normalized.zip</code>: L2-normalized CLIP features.</li> <li><code>dataset_softmax.zip</code>: CLIP features converted into probabilities using a softmax function.</li> <li><code>dataset_logistic.zip</code>: CLIP features converted into probabilities using a logistic function followed by L1 normalization.</li> </ul> </li> <li> <p><strong>Text Feature Data:</strong> <code>clip_laion_text_features.csv</code> contains additional details for each query, including the query ID, query text, and L2 normalized CLIP features extracted from the query text.</p> </li> </ol> <h3>Citation and Usage:</h3> <p>This data was used in the experiments described in:</p> <p>Lucia Vadicamo, Francesca Scotti, Alan Dearle, Richard Connor, <em>Comparative Analysis of Relevance Feedback Techniques for Image Retrieval</em>, in Proceedings of the 31st International Conference on Multimedia Modeling (MMM 2025).</p> <p>The data is released under a Creative Commons Attribution license. If you use it in your research, please cite the above work. </p> <h3>References:</h3> <p>[1]TRECVID Data: <a href="https://www-nlpir.nist.gov/projects/trecvid/trecvid.data.html">https://www-nlpir.nist.gov/projects/trecvid/trecvid.data.html</a><br>[2] Rossetto, L., Schuldt, H., Awad, G., Butt, A.A.: V3C - A research video collection. <em>In: International Conference on Multimedia Modeling</em>, pp. 349–360. Springer (2019).<br>[3] https://huggingface.co/laion/CLIP-ViT-H-14-laion2B-s32B-b79K<br>[4] VISIONE Repository: <a href="https://zenodo.org/records/8188570">https://zenodo.org/records/8188570</a></p> </div> </div> </div> </div> </div> </div> <p> </p> <p> </p> <p> </p>
FIG. 27. Astrapotherium magnum, selected basicranial features. A in Cranial Morphology And Phylogenetic Relationships Of Trigonostylops Wortmani, An Eocene South American Native Ungulate
FIG. 27. Astrapotherium magnum, selected basicranial features. A, AMNH VP-9278, right basicranium in ventral aspect. Prominent impression labelled "?vasc sulc" may have conducted extracranial venous structures similar to basicranial plexuses seen in extant Equus (fig. 6; see also similarly positioned sulcus in Tapirus, fig. 38: feature 5). Basicapsular fenestra is hidden in this perspective, except for extreme rostral and caudal ends. B, Digitally reconstructed left petrosal of A . magnum MACN A 3208, reversed and rotated to permit com-
Physiology data for: Biomechanical origins of proprioceptor feature selectivity and topographic maps in the Drosophila leg
<p>Our ability to sense and move our bodies relies on proprioceptors, sensory neurons that detect mechanical forces within the body. Because they are located within complex and dynamic peripheral tissues, the underlying mechanisms of proprioceptor feature selectivity remain poorly understood. Using single-nucleus RNA sequencing, we found that proprioceptor subtypes in the <em>Drosophila</em> leg express similar complements of mechanosensory and other ion channels. However, anatomical reconstruction of the proprioceptive organ and connected tendons revealed major biomechanical differences between proprioceptor subtypes. We constructed a computational model that identified a biomechanical mechanism for joint angle selectivity and predicted the existence of a goniotopic map of joint angle among position-tuned proprioceptors, which we confirmed using calcium imaging. Our findings suggest that biomechanical specialization is a key determinant of proprioceptor feature selectivity in <em>Drosophila</em>. The discovery of proprioceptive maps in the fly leg reveals common organizational principles between proprioception and other topographically organized sensory systems.</p>
Synaptic basis of feature selectivity in hippocampal neurons
Open the record for dataset details and reuse information.
Physiology data for: Biomechanical origins of proprioceptor feature selectivity and topographic maps in the Drosophila leg
Open the record for dataset details and reuse information.
Predictor complexity and feature selection affect Maxent model transferability: evidence from global freshwater invasive species
<p>This dataset contains the following:</p> <ol> <li>Occurrence datasets of five global freshwater invasive species (African sharptooth catfish <i>Clarias gariepinus</i>, Mozambique tilapia <i>Oreochromis mossambicus</i>, American bullfrog <i>Lithobates catesbeianus</i>, red swamp crayfish <i>Procambarus clarkii</i>, and Australian redclaw crayfish <i>Cherax quadricarinatus</i>)</li> <li>Background points for presence-only ecological niche modelling (e.g., Maxent)</li> <li>Example R script (with annotations inline) to conduct model tuning and transferability assessments using Maxent</li> </ol>
Data set from 'Sequential Feature Selection for Power System Event Classification Utilizing Wide-Area PMU Data'
<p>The increasing penetration of intermittent, nonsynchronous<br> generation has led to a reduction in total power<br> system inertia. Low inertia systems are more sensitive to sudden<br> changes, and more susceptible to secondary issues that can result<br> in large scale events. Due to the short time frames involved,<br> automatic methods for power system event detection and diagnosis<br> are required. Wide-area monitoring systems can provide<br> the data required to detect and diagnose events; however due to<br> the increasing quantity of data it is next to impossible for power<br> system operators to manually process raw data. The important<br> information is required to be extracted and presented to system<br> operators for real/near-time decision making and control. This<br> paper demonstrates an approach for the wide-area classification<br> of a number of power system events. A mixture of sequential<br> feature selection and linear discriminant analysis is adopted<br> to reduce the dimensionality of PMU data. Successful event<br> classification is obtained by employing quadratic discriminant<br> analysis on wide-area synchronized frequency, phase angle and<br> voltage measurements. The reliability of the proposed method is<br> evaluated using simulated case studies and benchmarked against<br> other classification methods.</p>
Optimizing the detection of biological signals through a semi-automated feature selection tool
<p>This repository contains the .xml files generated by the MZmine (Schmid et al., 2023) software version 2.53 during the preprocessing of the raw data (.raw) according to the recommendations of the 2007 Metabolomics Standards Initiative (Summer et al., 2007).</p> <p> </p> <p>Schmid R et al. Integrative analysis of multimodal mass spectrometry data in MZmine 3. Nat Biotechnol. 2023 Apr;41(4):447-449.</p> <p>Sumner et al. Proposed minimum reporting standards for chemical analysis Chemical Analysis Working Group (CAWG) Metabolomics Standards Initiative (MSI). Metabolomics. 2007 Sep;3(3):211-221. </p>
StreetScouting dataset: A Street-Level Image dataset for finetuning and applying custom object detectors for urban feature selection
<p>The dataset consists of two .zip files.</p> <p>The first .zip file named "annotated dataset" contains a folder named “annotated dataset" with annotated street images. It consists of the folder “images” that has 763 image files. The image format is PNG. 432 images have dimensions of 1080 x 2160 and 331 have dimensions of 866 x 2400. The filenames are random uuids. The “annotated dataset” folder also contains the annotations in the file “coco_annotations.json”. Annotations are provided in COCO format. Table 1 shows the total number of annotated objects per class.</p> <table align="center" summary="Total number of annotated objects per class"> <caption><em>Table 1. Total number of annotated objects per class</em></caption> <thead> <tr> <th scope="col"><strong><em>Class</em></strong></th> <th scope="col"><strong><em>Annotated Objects</em></strong></th> </tr> </thead> <tbody> <tr> <td><em>Tree</em></td> <td>1922</td> </tr> <tr> <td><em>Waste Bin</em></td> <td><em>223</em></td> </tr> <tr> <td><em>Recycling Bin</em></td> <td>181</td> </tr> <tr> <td><em>Lighting Pole</em></td> <td><em>716</em></td> </tr> <tr> <td><em>Shop Storefront</em></td> <td><em>628</em></td> </tr> </tbody> </table> <p>The second .zip file is named "routes" and contains a folder named “routes” with consecutive frames of four different driving routes in the city of Thessaloniki and their corresponding GPS signal. So the folder “routes” contains 4 folders in the following format “VID_<YYYYMMDD>_<HHmmSS>” where Y denotes digits for year, M denotes digits for month, D denotes digits for day, H denotes digits for hour, m denotes digits for minutes and S denotes digits for seconds. Not the filename represents the start of the collection sequence. All street data was collected in 2022. Each route folder has the “images” folder which contains the consecutive street image data. Image data in this folder is in JPEG format. Each filename in ‘images’ has the frame_<id>.jpg format where id denotes the order of the frame. Table 2 shows more details regarding the number of frames and frame dimension of the driving routes.</p> <table align="center" summary="Total number of annotated objects per class"> <caption><em>Table 2. Total number frames and frame dimensions for each of the routes</em></caption> <thead> <tr> <th scope="col"><strong><em>Route Name</em></strong></th> <th scope="col"><strong><em>Frames Number</em></strong></th> <th scope="col"><strong><em>Frame Dimension</em></strong></th> <th scope="col"><strong><em>Route duration</em></strong></th> </tr> </thead> <tbody> <tr> <td><em>VID_20220617_111456</em></td> <td><em>41,650</em></td> <td><em>1080 x 2160</em></td> <td>1h, 9m, 26s</td> </tr> <tr> <td><em>VID_20220210_112926</em></td> <td><em>23.035</em></td> <td><em>866 x 2400</em></td> <td>38m, 26s</td> </tr> <tr> <td><em>VID_20220209_114831</em></td> <td><em>18.000</em></td> <td><em>1080 x 2160</em></td> <td>30m, 3s</td> </tr> <tr> <td><em>VID_20220209_123323</em></td> <td><em>18.273</em></td> <td><em>1080 x 2160</em></td> <td>30m, 30s</td> </tr> </tbody> </table> <p>Each route folder contains a “gps.json” file which contains latitude and longitude information for each frame. This file is essentially a JSON list of objects that each object contains the “frame_name” attribute and the corresponding “coordinates” object which contains the “latitude” and "longitude" attributes.</p>
Feature selective adaptation of numerosity perception
Open the record for dataset details and reuse information.
Landscape features and seasonal habitat predicts lek-site selection and lek size of a <em>Tympanuchus</em> grouse
Open the record for dataset details and reuse information.
Dual-feature selectivity enables bidirectional coding in visual cortical neurons
Open the record for dataset details and reuse information.
Predictor complexity and feature selection affect Maxent model transferability: evidence from global freshwater invasive species
Open the record for dataset details and reuse information.
FIGURE 27. Selected morphometric features for 34 in Systematic revision of the flatfish genus Peltorhamphus Günther, 1862 (Teleostei: Pleuronectiformes: Rhombosoleidae), including description of a new species from Southeastern New Zealand, with biological and ecological summaries for the species
FIGURE 27. Selected morphometric features for 34 specimens of Peltorhamphus kryptostomus n. sp. 33.2–145.1 mm SL. A-D. Body depth (BD), Head length (HL), Head width (HW), and Ocular-side pectoral fin (OSP) expressed as percent of SL versus SL (in mm), respectively. E–J. Dorsal head width (DHW), Snout length (SNL), Eye diameter (ED), Interorbital width (IO), Upper jaw length (UJL), and Eye to upper mouth distance (EUM) expressed as percent of HL versus HL (in mm), respectively.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.