Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

3,709

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

3,709 results for “feature”

Learn how ShareScore rates datasets ↗
zenodo48/100

Extended datasets from MM-IMDB and Ads-Parallelity dataset with the features from Google Cloud Vision API

<p>This is extended datasets from&nbsp;MM-IMDB [<a href="https://openreview.net/forum?id=S12_nquOe">Arevalo+ ICLRW&#39;17</a>], Ads-Parallelity [<a href="https://arxiv.org/abs/1807.08205">Zhang+ BMVC&#39;18</a>]&nbsp;dataset with the features from Google Cloud Vision API. These datasets are stored in jsonl (JSON Lines) format.</p> <p><strong>Abstract (from our paper):</strong></p> <p>There is increasing interest in the use of multimodal data in various web applications, such as digital advertising and e-commerce.&nbsp;Typical methods for extracting important information from multimodal data rely on a mid-fusion architecture that combines the feature representations from multiple encoders.&nbsp;However, as the number of modalities increases, several potential problems with the mid-fusion model structure arise, such as an increase in the dimensionality of the concatenated multimodal features and missing modalities.&nbsp;To address these problems, we propose a new concept that considers multimodal inputs as a set of sequences, namely, deep multimodal sequence sets (DM<sup>2</sup>S<sup>2</sup>).&nbsp;Our set-aware concept consists of three components that capture the relationships among multiple modalities: (a) a BERT-based encoder to handle the inter- and intra-order of elements in the sequences, (b) intra-modality residual attention (IntraMRA) to capture the importance of the elements in a modality, and (c) inter-modality residual attention (InterMRA) to enhance the importance of elements with modality-level granularity further.&nbsp;Our concept exhibits performance that is comparable to or better than the previous set-aware models.&nbsp;Furthermore, we demonstrate that the visualization of the learned InterMRA and IntraMRA weights can provide an interpretation of the prediction results.</p> <p><strong>Dataset (MM-IMDB and Ads-Parallelity):</strong></p> <p>We extended&nbsp;two multimodal datasets, namely, MM-IMDB [<a href="https://openreview.net/forum?id=S12_nquOe">Arevalo+ ICLRW&#39;17</a>], Ads-Parallelity [<a href="https://arxiv.org/abs/1807.08205">Zhang+ BMVC&#39;18</a>] for the empirical experiments. The MM-IMDB dataset contains 25,925 movies with multiple labels (genres). We used the original split provided in the dataset and reported the F1 scores (micro, macro, and samples) of the test set. The Ads-Parallelity dataset contains 670 images and slogans from persuasive advertisements to understand the implicit relationship (parallel and non-parallel) between these two modalities. A binary classification task is used to predict whether the text and image in the same ad convey the same message.</p> <p>We transformed the following multimodal information (i.e., visual, textual, and categorical data) into textual tokens and fed these into our proposed model. We used the <a href="https://cloud.google.com/vision">Google Cloud Vision API</a>&nbsp;for the visual features to obtain the following four pieces of information as tokens: (1) text from the OCR, (2) category labels from the label detection, (3) object tags from the object detection, and (4) the number of faces from the facial detection. We input the labels and object detection results as a sequence in order of confidence, as obtained from the API. We describe the visual, textual, and categorical features of each dataset below.</p> <p><em><strong>MM-IMDB</strong></em>:&nbsp;We used the title and plot of movies as the textual features, and the aforementioned API results based on poster images as visual features.</p> <p><em><strong>Ads-Parallelity</strong></em>: We used the same API-based visual features as in MM-IMDB. Furthermore, we used textual and categorical features consisting of textual inputs of transcriptions and messages, and categorical inputs of natural and text concrete images.</p>

opencc-by-4.0Sep 2022View details →
zenodo48/100

Training Datasets for Epilepsy Analysis: Preprocessing and Feature Extraction from EEG Time Series

<h2>The files include the 20 training datasets, in csv format, from 20 epileptic patients. Each set of data is described by 1080 features extracted using the sliding window technique.</h2>

opencc-by-4.0Mar 2024View details →
zenodo48/100

Feature selection on microbial profiles of CRC samples with chopin2 (powered by hdlib)

<p>This Zenodo entry contains the result of the feature selection algorithm implemented through a backward variable elimination strategy in&nbsp;<a href="https://github.com/cumbof/chopin2" target="_blank" rel="noopener">chopin2</a> (powered by <a href="https://github.com/cumbof/hdlib" target="_blank" rel="noopener">hdlib</a>) applied on&nbsp;<a href="https://github.com/biobakery/MetaPhlAn" target="_blank" rel="noopener">MetaPhlAn3</a> microbial profiles of a public dataset of metagenomic stool samples collected from patients affected by the colorectal cancer (CRC) as well as&nbsp;from healthy individuals.</p> <p>Microbial profiles have been extracted through the <a href="https://bioconductor.org/packages/release/data/experiment/html/curatedMetagenomicData.html" target="_blank" rel="noopener">curatedMetagenomicData</a> package for R under the IDs&nbsp;<em>ThomasAM_2018a</em>, <em>ThomasAM_2018b</em>, and <em>ThomasAM_2019_a</em>.</p> <p>The feature selection algorithm is implemented as a backward variable elimination method, and it makes use of the vector-symbolic architecture described in&nbsp;<a href="https://doi.org/10.3390/a13090233" target="_blank" rel="noopener">Cumbo F 2020</a>.</p> <p>Deposited data is described below:</p> <ul> <li><em>datasets.tar.gz</em>: it contains the datasets used as input of <em>chopin2</em>&nbsp;as the result of merging the three datasets with relative abundances mentioned above, also stratified by age and sex (with prefix RA). The same datasets have been also binarized (with prefix BIN);</li> <li><em>hd-models.tar.gz</em>: it contains the output of the feature selection performed with&nbsp;<em>chopin2</em> (powered by <em>hdlib</em>) on the datasets with both relative abundance and binary profiles (RA and BIN);</li> <li><em>ml-models.tar.gz</em>: it contains the result of the feature selection produced with classical wrapper-based techniques (i.e., Random Forest, Decision Tree, Support Vector Machine, Logistic Regression, and Extreme Gradient Boosting) in addition to a Python 3.8 script to reproduce the results.</li> </ul> <p>Please note that the datasets <em>RA__ThomasAM__species.csv</em> and&nbsp;<em>BIN__ThomasAM__species.csv</em>&nbsp;are also included into the&nbsp;<em>datasets.tar.gz</em> archive.</p>

opencc-zeroMay 2024View details →
zenodo48/100

Raw and processed hydro-meteorological variables of Jucar river basin for feature selection

<p>The dataset Processed data &ndash; input WQEISS.csv was employed for the input variable selection step in Zaniolo et al., 2018. It includes monthly values of 28 hydro-meteorological variables and indexes of Jucar river basin, Spain, for the period 1986-2000, namely:</p> <ul> <li>2 temporal features: day and month of the year;</li> <li>12 inputs to the Jucar State Index: average monthly storage and groundwater levels, average three months river runoff, and cumulated areal precipitation over 12 months;</li> <li>8 additional observed variables in the basin: three months average outflows from, and inflows to, the main reservoirs, and mean monthly areal temperatures;</li> <li>6 traditional drought indicators: Standardized Precipitation Index (SPI) and Standardized Precipitation and Evaporation Index (SPEI). SPI and SPEI indicators are computed on mean monthly data over the entire basin for 3, 6, and 12 months time aggregations.</li> </ul> <p>The last column of the dataset reports the target variable, i.e., the monthly nominal shortage of water conveyed to the irrigation districts simulated via AQUATOOL model. For further details on the dataset please consult Zaniolo et al., 2018, or the dedicated website <a href="http://www.nrm.deib.polimi.it/?page_id=2438">http://www.nrm.deib.polimi.it/?page_id=2438</a></p> <p>The unprocessed data used to compute indices and temporal cumulations in Processed data &ndash; input WQEISS.csv are reported in table Raw Data.csv. Public observations of rainfall, streamflows and storage levels come from the SAIH (Hydrological Automatic Information System) of the CHJ (Jucar Hydrological Confederation). Users can directly download data for the last 12 months on the dedicated webpage <a href="http://saih.chj.es/chj/saih/?f">http://saih.chj.es/chj/saih/?f</a> while previous data records are provided for free by CHJ upon request. Observations from piezometers are downloadable from the Piezometric Network Information section section of the CHJ&nbsp; <a href="https://www.chj.es/es-es/medioambiente/redescontrol/Paginas/Piezometr%C3%ADa.aspx">https://www.chj.es/es-es/medioambiente/redescontrol/Paginas/Piezometr%C3%ADa.aspx</a>.</p>

opencc-by-4.0Feb 2018View details →
zenodo48/100

Processed features in support of Liebeskind et al (2018)

<p>Processed feature matrices used in Liebeskind et al. (2018). Supporting code: https://github.com/marcottelab/plum</p> <p>Datasets 1 - 4 correspond to those used in Figure 4:</p> <p>Dataset 1: No AP-MS, yeast CF-MS, training species: Human</p> <p>Dataset 2: AP-MS, yeast CF-MS, training species: Human</p> <p>Dataset 3: AP-MS, yeast CF-MS, training species: Human, Yeast</p> <p>Dataset 4: AP-MS, no yeast CF-MS, training species: Human, Yeast</p> <p>&quot;.train_labeled.missing_annotated.csv&quot; files are those used for training the model and include only orthogroups for which interactions are known in the training species. These known interactions come either from gold-standard test sets, such as CORUM or EMBL&#39;s training portal, or from the fact that at least of the orthogroup pairs is missing in the focal taxon.</p> <p>&quot;.missing_annotated.csv&quot; files were used for prediction, and include the entire feature matrices, plus known missing pairs. Note that there is no dataset 3 file. This is because data sets 2 and 3 differ only in the training species used, so dataset 3 predictions used dataset2_07302018.missing_annotated.csv as a feature matrix.</p> <p>dataset4_prediction_07302018.csv contains the predictions for all pairs on data set 4, the best performing data set that was used for all downstream analyses.</p>

opencc-by-4.0Aug 2018View details →
zenodo48/100

Impact Areas and Dynamical Features associated with Mediterranean Cyclones (1980-2019)

<p>The dataset includes NetCDF files of Impact Areas and Dynamical Features associated with Mediterranean Cyclones (henceforth MedCyclones).</p> <p>MedCyclone tracks correspond to confidence-level 5 tracks from Flaounas et al. (2023), <a href="https://doi.org/10.5194/wcd-4-639-2023">https://doi.org/10.5194/wcd-4-639-2023</a>.</p> <p>Temporal frequency: 6h (00, 06, 12, 18 UTC)<br>Years: 1980 &ndash; 2019<br>Spatial resolution: 0.5 deg<br>Grid extension: &nbsp;0-70N, 40W-65E</p> <p>&nbsp;</p> <h2>Dynamical Features</h2> <p>Files "dynfeats_bool_rmax2000_YYYY.nc" include the following list of variables, describing connected boolean objects:</p> <ul> <li><strong>r_500</strong>, <strong>r_1000</strong>: central areas of fixed 500 or 1000 km radius around MedCyclone centres;</li> <li><strong>WCB</strong>: warm conveyor belts related to MedCyclones (i.e., overlapping with r_500 in at least one grid point). Each <strong>WCB</strong> is eventually separated into inflow (<strong>WCBin</strong>, up to 800 hPa) and ascent (<strong>WCBout</strong>, between 800 and 400 hPa) regions. Ref. at <a href="https://doi.org/10.1175/JCLI-D-12-00720.1">https://doi.org/10.1175/JCLI-D-12-00720.1</a>, <a href="https://doi.org/10.5194/wcd-5-537-2024">https://doi.org/10.5194/wcd-5-537-2024</a>;</li> <li><strong>fronts</strong>: cold fronts related to MedCyclones (i.e., overlapping with r_500 in at least one grid point). Ref. at <a href="https://doi.org/10.5194/gmd-17-6137-2024">https://doi.org/10.5194/gmd-17-6137-2024</a>;</li> <li><strong>DI</strong>: dry instrusions related to MedCyclones (i.e., overlapping with r_1000 in at least one grid point). Ref. at <a href="https://doi.org/10.1175/JCLI-D-16-0782.1">https://doi.org/10.1175/JCLI-D-16-0782.1</a>;</li> <li><strong>r_1000_Nodynfeat</strong>: the central 1000 km area excluding regions of MedCyclone WCB, fronts and DI objects.</li> </ul> <p>The criteria for the identification of WCB, fronts and DI objects are described in Section 2.3 of Portal et al. (2024), <a href="https://doi.org/10.5194/wcd-5-1043-2024">https://doi.org/10.5194/wcd-5-1043-2024</a>.</p> <p>Additionally, we note that :<br>i. a weaker overlap constraint was used to associate DI objects to MedCyclones (r_1000 compared to r_500 for WCB and fronts objects) because of the relatively large distance of the DI airstream from the cyclone centre;<br>ii. in this dataset, all connected objects related to MedCyclones are cropped within a 2000 km area circle from the cyclone centre for two reasons. Firstly, the dynamical-feature related surface impacts usually weaken with the distance from the cyclone centre. Secondly, to cut connected objects composed by multiple overlapping features of the same kind - this often happens for fronts in summer because of their high detection density. Far from the cyclone centre, these objects are usually unrelated with the MedCyclone circulation.</p> <p>&nbsp;</p> <h2>Impact Areas</h2> <p>Files "IAs_bool_rmax2000_YYYY.nc" include boolean impact areas, combining a central area (r_1000 or r_500) and cyclone-related WCB, CF and DI objects. The three types of impact area are described in the following :</p> <ol> <li><strong>IA01</strong> is composed by a 1000 km radius circle around the cyclone centre (r_1000) extended by cyclone-related WCB, fronts and DI;</li> <li><strong>IA02</strong> is composed by a 500 km radius circle around the cyclone centre (r_500) extended by cyclone-related WCB, fronts and DI</li> <li><strong>IA03</strong> is composed by a 500 km radius circle around the cyclone centre (r_500) extended by cyclone-related WCB and fronts (DI is neglected).</li> </ol> <p>As discussed in Section 3.1 and Appendix A of Portal et al. (2024) (<a href="https://doi.org/10.5194/wcd-5-1043-2024">https://doi.org/10.5194/wcd-5-1043-2024</a>), IA01, composed by a central area of 1000 km, is adequate for intercepting long-range wind impacts associated with MedCyclones. IA02 and IA03, on the contrary, are better devised for detecting impacts expected at shorter distances from the cyclone centre, such as rainfall, thunderstorm and storm surges. In particular, IA03 neglects the DI region, which is normally of little interest for cyclone-related moist processes, involved in producing precipitation. Noetheless, DI remains relevant for the identification of strong cyclone-related winds.</p> <p>&nbsp;</p> <h3>Case Studies</h3> <p>A pdf file providing the visualisation and description of impact areas and dynamical features of all MedCyclones occurring in 1980 is available at the link <a href="https://boris.unibe.ch/192315/">https://boris.unibe.ch/192315/</a>. Note that in the examples the dynamical features are not cropped at 2000 km from the cyclone centres, as for the present dataset.&nbsp;</p> <p>Note that many of the "Annotations and Limitations" listed below derive from the attentive analysis of these study cases.</p> <h3>Annotations and limitations</h3> <ul> <li>In the case of more than one MedCyclone centre per timestep, the dasaset does not distinguish the impact areas / dynamical features associated with each centre.</li> <li>Because of the automated criteria for associating WCB, fronts and DI objects to MedCyclones, at times objects close to the centre but unrelated to the MedCyclone's circulation, are considered to be cyclone-related and included in the impact area.</li> <li>Elaborating on the point above, at times fronts responsible for Mediterranean cyclogenesis (and not produced by the cyclonic circulation itself) are included in the MedCyclone impact area.</li> <li>When computing statistics over a long time interval (e.g., climatology), the effects of erroneous associations of dynamical features to MedCyclone impact areas are attenuated by the aggregation of large quantity of data.&nbsp;</li> <li>Over a long time interval (e.g., climatology) the choice of a 1000 km fixed-radius impact area provides similar statistics to IA01, although in the first case it is not possible to isolate the role played by the different features composing the MedCyclones.</li> </ul>

opencc-by-4.0Sep 2024View details →
zenodo48/100

Feature Template Angular Power Spectra

<p>This data was used in the machine learning analysis of the Cosmic Microwave Background data in: https://github.com/IndiraOcampo/CMB_ML_based_model_selection.git and https://dx.doi.org/10.1088/1475-7516/2025/02/004</p> <p>The objective is to train a neural network architecture on the different polarization modes (TT, TE, EE and joint) to perform model selection between the standard cosmological model, &Lambda;CDM and a model that introduces a Feature Template (FT) in the primordial power spectrum - related to the early Universe physics.</p> <p>The first row corresponds to the multipole moment "\ell" and the remaining ones correspond to the different components of the Cl's angular power spectrum, for the different values of A_lin (the feature oscilation parameter). While A_0 = 10^-2 is a reasonable value that still agrees with observations, A_0 = 0 corresponds to the &Lambda;CDM model.</p> <p>Finally, our aim is to apply SHAP to perform feature importance (interpretability) in our results.</p>

openmit-licenseSep 2024View details →
zenodo48/100

Tectonic Map and Compressional Feature Map of Mare Tranquillitatis, Moon

<p>ArcGIS shapefiles of the tectonic (Compressional_Tectonism and&nbsp;Extensional_Tectonism) and compressional feature (Feature_Map_Classes) maps of Mare Tranquillitatis. This is complementary data for &quot;Timing and Origin of Compressional Tectonism in Mare Tranquillitatis&quot;&nbsp;published on JGR: Planets by Frueh et al. (Available on&nbsp;<a href="https://doi.org/10.1029/2022JE007533">https://doi.org/10.1029/2022JE007533</a>).&nbsp;</p> <ul> <li>Compressional_Tectonism: Polylines of wrinkle ridges, lobate scarps, and unidentified features, as well as their geodesic length, coordinates, and bearing.</li> <li>Extensional_Tectonism:&nbsp;Polylines of large graben and normal faults, as well as their geodesic length, coordinates, and bearing.</li> <li>Feature_Map_Classes: Compressional tectonic feature map, including their assigned erosional states.</li> </ul> <p>For a&nbsp;detailed description of the mapping process, features, and erosional states, we currently refer to&nbsp;our publication (Frueh et al., 2023).</p> <p>&nbsp;</p> <p>Frueh, T.,&nbsp;Hiesinger, H.,&nbsp;van der Bogert, C. H.,&nbsp;Clark, J. D.,&nbsp;Watters, T. R., &amp;&nbsp;Schmedemann, N.&nbsp;(2023).&nbsp;Timing and origin of compressional tectonism in Mare Tranquillitatis.&nbsp;<em>Journal of Geophysical Research: Planets</em>,&nbsp;128, e2022JE007533.&nbsp;<a href="https://doi.org/10.1029/2022JE007533">https://doi.org/10.1029/2022JE007533</a></p>

opencc-by-4.0Jan 2023View details →
zenodo48/100

Transfer learning for galaxy feature detection: Finding Giant Star-forming Clumps in low redshift galaxies using Faster R-CNN

<p>This repository contains the data released in the paper 'Transfer learning for galaxy feature detection: Finding Giant Star-forming Clumps in low redshift galaxies using Faster R-CNN'&nbsp;<em>(DOI: <a href="https://doi.org/10.1093/rasti/rzae013">10.1093/rasti/rzae013</a>).</em></p> <p>We release a detailed catalogue of Giant Star-forming Clumps (GSFCs), detected for the full set of Galaxy Zoo: Clump Scout&nbsp;galaxies observed by SDSS using the Faster R-CNN architecture with the Zoobot classification-CNN as a feature extraction backbone.</p> <p>The final models and code are made publicly available via Github:&nbsp;<a href="https://github.com/ou-astrophysics/Faster-R-CNN-for-Galaxy-Zoo-Clump-Scout">https://github.com/ou-astrophysics/Faster-R-CNN-for-Galaxy-Zoo-Clump-Scout</a>.</p> <p>We will release updates if needed via Zenodo versioning. We recommend using the latest version of this repository. You can check the version you are currently viewing on the right-hand sidebar.</p> <p>Please cite the paper (DOI: <a href="https://doi.org/10.1093/rasti/rzae013">10.1093/rasti/rzae013</a>) when using the data in this repository.</p> <p>The csv-file <em>FRCNN_Zoobot_SDSS_GZCS_detections.csv</em>&nbsp;has the following columns. Alternatively, the file <em>FRCNN_Zoobot_SDSS_GZCS_detections.gzip</em> contains the same data but stored as a parquet-file.</p> <table> <tbody><tr> <th>Column name</th> <th>Description</th> </tr> </tbody><tbody> <tr> <td>specobjid</td> <td>SDSS spec object ID</td> </tr> <tr> <td>dr7objid</td> <td>SDSS DR7 object ID</td> </tr> <tr> <td>clump_id</td> <td>Clump index</td> </tr> <tr> <td>clump_label_id</td> <td>Clump label ID (1 or 2)</td> </tr> <tr> <td>clump_label_name</td> <td>Clump label name</td> </tr> <tr> <td>clump_score</td> <td>Detection score for the clump</td> </tr> <tr> <td>clump_centre_ra</td> <td>Clump centroid RA in degrees</td> </tr> <tr> <td>clump_centre_dec</td> <td>Clump centroid dec in degrees</td> </tr> <tr> <td>clump_flux_u</td> <td>Clump u-band flux in Jy</td> </tr> <tr> <td>clump_flux_g</td> <td>Clump g-band flux in Jy</td> </tr> <tr> <td>clump_flux_r</td> <td>Clump r-band flux in Jy</td> </tr> <tr> <td>clump_flux_i</td> <td>Clump i-band flux in Jy</td> </tr> <tr> <td>clump_flux_z</td> <td>Clump z-band flux in Jy</td> </tr> <tr> <td>clump_flux_err_u</td> <td>Clump u-band flux error in Jy</td> </tr> <tr> <td>clump_flux_err_g</td> <td>Clump g-band flux error in Jy</td> </tr> <tr> <td>clump_flux_err_r</td> <td>Clump r-band flux error in Jy</td> </tr> <tr> <td>clump_flux_err_i</td> <td>Clump i-band flux error in Jy</td> </tr> <tr> <td>clump_flux_err_z</td> <td>Clump z-band flux error in Jy</td> </tr> <tr> <td>clump_mag_u</td> <td>Clump u-band magnitude (AB-mag)</td> </tr> <tr> <td>clump_mag_g</td> <td>Clump g-band magnitude (AB-mag)</td> </tr> <tr> <td>clump_mag_r</td> <td>Clump r-band magnitude (AB-mag)</td> </tr> <tr> <td>clump_mag_i</td> <td>Clump i-band magnitude (AB-mag)</td> </tr> <tr> <td>clump_mag_z</td> <td>Clump z-band magnitude (AB-mag)</td> </tr> <tr> <td>clump_ext_mag_u</td> <td>Clump u-band extinction (E(B-V), AB-mag)</td> </tr> <tr> <td>clump_ext_mag_g</td> <td>Clump g-band extinction (E(B-V), AB-mag)</td> </tr> <tr> <td>clump_ext_mag_r</td> <td>Clump r-band extinction (E(B-V), AB-mag)</td> </tr> <tr> <td>clump_ext_mag_i</td> <td>Clump i-band extinction (E(B-V), AB-mag)</td> </tr> <tr> <td>clump_ext_mag_z</td> <td>Clump z-band extinction (E(B-V), AB-mag)</td> </tr> <tr> <td>clump_mag_corr_u</td> <td>Clump corrected u-band magnitude (AB-mag)</td> </tr> <tr> <td>clump_mag_corr_g</td> <td>Clump corrected g-band magnitude (AB-mag)</td> </tr> <tr> <td>clump_mag_corr_r</td> <td>Clump corrected r-band magnitude (AB-mag)</td> </tr> <tr> <td>clump_mag_corr_i</td> <td>Clump corrected i-band magnitude (AB-mag)</td> </tr> <tr> <td>clump_mag_corr_z</td> <td>Clump corrected z-band magnitude (AB-mag)</td> </tr> <tr> <td>clump_mag_corr_u_g</td> <td>Clump colour (u-g)</td> </tr> <tr> <td>clump_mag_corr_g_r</td> <td>Clump colour (g-r)</td> </tr> <tr> <td>clump_mag_corr_r_i</td> <td>Clump colour (r-i)</td> </tr> <tr> <td>clump_mag_corr_i_z</td> <td>Clump colour (i-z)</td> </tr> <tr> <td>clump_flux_ratio</td> <td>Est. clump/galaxy near-UV flux ratio (u-band)</td> </tr> <tr> <td>is_clump_3pct</td> <td>Flag (True/False) if clump/galaxy flux ratio is &gt;3%</td> </tr> <tr> <td>is_clump_8pct</td> <td>Flag (True/False) if clump/galaxy flux ratio is &gt;8%</td> </tr> <tr> <td>galaxy_ra</td> <td>Host galaxy RA in degrees</td> </tr> <tr> <td>galaxy_dec</td> <td>Host galaxy dec in degrees</td> </tr> <tr> <td>galaxy_z</td> <td>Host galaxy redshift</td> </tr> <tr> <td>galaxy_mag_u</td> <td>Host galaxy u-band magnitude (AB-mag)</td> </tr> <tr> <td>galaxy_mag_g</td> <td>Host galaxy g-band magnitude (AB-mag)</td> </tr> <tr> <td>galaxy_mag_r</td> <td>Host galaxy r-band magnitude (AB-mag)</td> </tr> <tr> <td>galaxy_mag_i</td> <td>Host galaxy i-band magnitude (AB-mag)</td> </tr> <tr> <td>galaxy_mag_z</td> <td>Host galaxy z-band magnitude (AB-mag)</td> </tr> <tr> <td>galaxy_mag_err_u</td> <td>Host galaxy u-band magnitude error (AB-mag)</td> </tr> <tr> <td>galaxy_mag_err_g</td> <td>Host galaxy g-band magnitude error (AB-mag)</td> </tr> <tr> <td>galaxy_mag_err_r</td> <td>Host galaxy r-band magnitude error (AB-mag)</td> </tr> <tr> <td>galaxy_mag_err_i</td> <td>Host galaxy i-band magnitude error (AB-mag)</td> </tr> <tr> <td>galaxy_mag_err_z</td> <td>Host galaxy z-band magnitude error (AB-mag)</td> </tr> <tr> <td>galaxy_flux_u</td> <td>Host galaxy u-band flux in Jy</td> </tr> <tr> <td>galaxy_flux_g</td> <td>Host galaxy g-band flux in Jy</td> </tr> <tr> <td>galaxy_flux_r</td> <td>Host galaxy r-band flux in Jy</td> </tr> <tr> <td>galaxy_flux_i</td> <td>Host galaxy i-band flux in Jy</td> </tr> <tr> <td>galaxy_flux_z</td> <td>Host galaxy z-band flux in Jy</td> </tr> <tr> <td>galaxy_expAB_r</td> <td>Host galaxy axis ratio from SDSS</td> </tr> <tr> <td>galaxy_expRad_r</td> <td>Host galaxy exponential fit scale radius from SDSS</td> </tr> <tr> <td>galaxy_lmass</td> <td>Host galaxy log mass in MSun</td> </tr> <tr> <td>galaxy_lssfr</td> <td>Host galaxy log specific SFR</td> </tr> <tr> <td>galaxy_mag_corr_u</td> <td>Host galaxy corrected u-band magnitude (AB-mag)</td> </tr> <tr> <td>galaxy_mag_corr_g</td> <td>Host galaxy corrected g-band magnitude (AB-mag)</td> </tr> <tr> <td>galaxy_mag_corr_r</td> <td>Host galaxy corrected r-band magnitude (AB-mag)</td> </tr> <tr> <td>galaxy_mag_corr_i</td> <td>Host galaxy corrected i-band magnitude (AB-mag)</td> </tr> <tr> <td>galaxy_mag_corr_z</td> <td>Host galaxy corrected z-band magnitude (AB-mag)</td> </tr> </tbody> </table> <p>&nbsp;</p>

opencc-by-4.0Aug 2023View details →
zenodo48/100

PITS Apparent Depth Profiles for Mars Global Cave Candidate Catalog (MGC3) Features

<p>Apparent depth profiles calculated by the Pit Topography from Shadows (PITS) tool for the majority of the features in the Mars Global Cave Candidate Catalog (MGC<sup>3</sup>). PITS is a Python framework for automatically calculating apparent depth profiles for Martian and Lunar pits from just a single cropped satellite image. These images can also be single- or multi-band, such as in the case of the Mars Reconnaissance Orbiter (MRO) HiRISE camera.&nbsp;You can learn more about PITS by reading its <a href="https://academic.oup.com/rasti/article/2/1/492/7241547">journal article</a> in RAS Techniques and Instruments, going to its <a href="https://github.com/dlecorre387/Pit-Topography-from-Shadows/">GitHub repository</a> or reading the following <a href="https://www.danlecorre.com/post/first-paper-published">post</a>.</p> <p>Since not all catalogued cave candidates on Mars will be pits, PITS has so far&nbsp;been applied to the following MGC<sup>3</sup> subcategories:</p> <ul> <li>Atypical Pit Craters (APCs).</li> </ul> <p>With plans to extend this to:</p> <ul> <li>Lava tube skylights,</li> <li>small rimless pits,</li> <li>generic, amorphous pits,</li> <li>and polar pits.</li> </ul> <p>This totals 123&nbsp;apparent depth profiles&nbsp;in CSV format, which have been derived automatically by PITS for 88 APCs. Therefore, these profiles can be plotted as the user prefers, and/or used in combination with other data to reveal more about this particular APC on the surface of Mars.</p> <p>Each depth profile&#39;s CSV file is named according to the HiRISE Reduced Data Record Version 1.1. (RDRV11) that it was calculated upon (e.g. ESP_011386_2065_RED_profile.csv for the red-band version of the HiRISE image ESP_011386_2065). Where there are multiple MGC3 APCs contained within a single image, the file names are numbered generally from the most northern&nbsp;to southernmost, or most westerly to easterly. ESRI shapefiles for the location of all&nbsp;APCs in each HiRISE image have been provided in polygon (containing the extents used to crop the larger HiRISE product) and point format in order to give context in these intances.&nbsp;</p> <p>As the headers suggest, the first four columns represent the shadow length (<em><span class="math-tex">\(L\)</span></em>), apparent depth (<em><span class="math-tex">\(h\)</span></em>), and the upper/lower bounds of <span class="math-tex">\(\Delta h\)</span>, respectively, before they have been corrected for non-zero emission angles (<span class="math-tex">\(\varepsilon\)</span>) at the time of image acquisition. Whereas the latter four columns represent the same quantities after <span class="math-tex">\(\varepsilon\)</span>-correction. How this correction is derived and applied is explained in the PITS journal article linked above.</p>

opencc-by-4.0Aug 2023View details →
OpenNeuro44/100

Adaptive memory distortions are predicted by feature representations in parietal cortex

Open the record for dataset details and reuse information.

openCC0Jan 2021View details →
zenodo44/100

P4KxSpotify: A Dataset of Pitchfork Music Reviews and Spotify Musical Features

<p>18,403 music reviews scraped from Pitchfork, including relevant metadata such as author, review date, record release year, score, and genre, along with those album&#39;s audio features pulled from Spotify&#39;s API.</p>

opencc-by-4.0Jan 2020View details →
zenodo44/100

Exploiting Statistical and Structural Features for the Detection of Domain Generation Algorithms

<p>This repository contains a&nbsp;dataset for the research of domain generation algorithms (DGAs) and machine learning. More precisely, it targets dictionary-based DGAs.</p> <p><em>Constantinos Patsakis, Fran Casino: &quot;Exploiting Statistical and Structural Features for the Detection of Domain Generation Algorithms&quot;,&nbsp;Journal of Information Security and Applications, 2021.</em></p> <p>Features ordered as in the shared dataset:</p> <ul> <li>Family: DGA that the domain belongs to</li> <li>SLD: SLD of the Domain</li> <li>L-LEN: The length of Domain</li> <li>L-DIG: The number of digits in Domain</li> <li>L-CON-MAX: The maximum number of consecutive consonants Domain</li> <li>R-CON-VOW: Number of consonants divided by L-LEN&nbsp;</li> <li>L-SYM: The number of special characters</li> <li>R-SYM-LEN: L-SYM divided by L-LEN</li> <li>R-Dom-3G: Ratio of benign grams in Dom-3G</li> <li>R-Dom-4G: Ratio of benign grams in Dom-4G</li> <li>R-Dom-5G: Ratio of benign grams in Dom-5G</li> <li>L-W2: Number of words with more than 2 characters in Domain</li> <li>L-W3: Number of words with more than 3 characters in Domain</li> <li>R-WS-LEN: Dom-WS divided by L-LEN</li> <li>R-WDS-LEN: Dom-WDS divided by L-LEN</li> <li>R-W2-LEN: Dom-W2 divided by L-LEN</li> <li>R-W3-LEN: Dom-W3 divided by L-LEN</li> <li>M2-Dom-Ws: 2-Chain Markov English grams applied to Dom-WS</li> <li>M2-Dom-WDS: 2-Chain Markov English grams applied Dom-WDS</li> <li>E-Dom-WS: Entropy of Dom-WS&nbsp;</li> <li>E-Dom-WDS: Entropy of Dom-WDS</li> <li>E-Dom-W2: Entropy of Dom-W2</li> <li>E-Dom-W3: Entropy of Dom-W3</li> </ul>

opencc-by-4.0Aug 2020View details →
zenodo44/100

Associated Data: RASPD+: Fast protein-ligand binding free energy prediction using simplified physicochemical features

<p>Additional digital data to &quot;RASPD+: Fast protein-ligand binding free energy prediction using simplified physicochemical features&quot; (ChemRxiv preprint:<a href="https://doi.org/10.26434/chemrxiv.12636704.v1">https://doi.org/10.26434/chemrxiv.12636704</a>).</p> <p>Associated code can be found at:&nbsp;<a href="https://github.com/HITS-MCM/RASPDplus">https://github.com/HITS-MCM/RASPDplus</a></p> <p>Files:</p> <ul> <li>weights.tar.gz: contains the model weights of one random dataset split and its associated crossvalidation folds. Used for standard RASPD+ evaluation.</li> <li>additional_model_replicates.tar.gz: contains the remaining models trained on the full set of descriptors.</li> <li>external_test_sets.tar.gz: contains the descriptor tables for all external test sets used</li> <li>dude.tar.gz: contains the descriptor tables for and several identifier lists for evaluation on the Directory of Useful Decoys - Enhanced (DUD-E)</li> <li>run_outputs.tar.gz: Performance metric data and predicted values created during the model training and evaluation runs. Basis for the figures and metrics in the manuscript.</li> </ul> <p>&nbsp;</p>

opencc-by-4.0Jul 2020View details →
zenodo44/100

1QIsaa data collection (binarized images, feature files, and plotting scripts) for writer identification test using artificial intelligence and image-based pattern recognition techniques

<p><strong>The Great Isaiah Scroll (1QIsa<sup>a</sup>) data set for writer identification</strong></p> <p>This data set is collected for the ERC project:<br> The Hands that Wrote the Bible: Digital Palaeography and Scribal Culture of the Dead Sea Scrolls<br> PI: Mladen Popović<br> Grant agreement ID: 640497</p> <p>Project website: <a href="https://cordis.europa.eu/project/id/640497">https://cordis.europa.eu/project/id/640497</a><br> <br> <strong>Copyright (c) </strong>&nbsp;&nbsp; &nbsp;University of Groningen, 2021. All rights reserved.<br> <strong>Disclaimer and copyright notice for all data contained on this .tar.gz file:</strong></p> <p><strong>1)</strong> permission is hereby granted to use the data for research purposes. It is not allowed to distribute this data for commercial purposes.</p> <p><strong>2) </strong>provider gives no express or implied warranty of any kind, and any implied warranties of merchantability and fitness for purpose are disclaimed.</p> <p><strong>3) </strong>provider shall not be liable for any direct, indirect, special, incidental, or consequential damages arising out of any use of this data.</p> <p><strong>4) </strong>the user should refer to the first public article on this data set:<br> <br> <em>Popović, M., Dhali, M. A., &amp; Schomaker, L. (2020). Artificial intelligence-based writer identification generates new evidence for the unknown scribes of the Dead Sea Scrolls exemplified by the Great Isaiah Scroll (1QIsa<sup>a</sup>). arXiv preprint arXiv:2010.14476.</em><br> <br> BibTeX:</p> <pre>@article{popovic2020artificial, title={Artificial intelligence based writer identification generates new evidence for the unknown scribes of the Dead Sea Scrolls exemplified by the Great Isaiah Scroll (1QIsaa)}, author={Popovi{\&#39;c}, Mladen and Dhali, Maruf A and Schomaker, Lambert}, journal={arXiv preprint arXiv:2010.14476}, year={2020} }</pre> <p><strong>5) </strong>the recipient should refrain from proliferating the data set to third parties external to his/her local research group. Please refer interested researchers to this site for obtaining their own copy.</p> <p><strong>Organisation of the data:</strong></p> <p>The .tar.gz file contains three directories: images, features, and plots. The included &#39;README&#39; file contains all the instructions.</p> <p>The &#39;images&#39; directory contains NetPBM images of the columns of 1QIsa<sup>a</sup>. The NetPBM format is chosen because of its simplicity. Additionally, there is no doubt about lossy compression in the processing chain. There are two images for each of the Great Isaiah Scroll columns: one is the direct binarized output from the BiNet (<em>arxiv.org/abs/1911.07930</em>) system, and the other one is the manually cleaned version of the binarized output. &nbsp; The file names for the direct binarized output are of the format &#39;1QIsaa_col&lt;columnnr&gt;.pbm&#39;, for example, &#39;1QIsaa_col15.pbm&#39;. And, for the cleaned version, the format is &#39;1QIsaa_col&lt;columnnr&gt;_cleaned.pbm&#39;, for example, &#39;1QIsaa_col15_cleaned.pbm&#39;. Note: the image files are not in a separate directory; they will be extracted in the same place. However, due to the unique naming, there is no problem extracting them in one single directory.</p> <p>The &#39;features&#39; directory contains feature files computed for each of the column images. There are two types of feature files: Hinge and Adjoined. They are distinguishable by their extension, for example, &#39;1QIsaa_col15_cleaned.hinge&#39; and &#39;1QIsaa_col15_cleaned.adjoined&#39;. They are also arranged in separate directories for ease of use.</p> <p>The &#39;plots&#39; directory contains a simple python script to perform PCA on the feature files and then visualize them in a 3D plot. The file takes the location of feature files as an input. The &#39;README_plot&#39; file contains examples of how-to-run in the terminal.</p> <p><strong>Brief description:</strong><br> According to ImageMagick&#39;s&#39; identify&#39; tool, the original images are in grayscale (.jpg) from Brill collection, in &#39;8-bit Gray 256c&#39;. &nbsp;These images pass through multiple preprocessing measures to become suitable for pattern recognition-based techniques. The first step in preprocessing is the image-binarization technique. In order to prevent any classification of the text-column images based on irrelevant background patterns, a specific binarization technique (BiNet) was applied, keeping the original ink traces intact. After performing the binarization, the images were cleaned further by removing the adjacent columns that partially appear on the target columns&#39; images. Finally, few minor affine transformations and stretching corrections were performed in a restrictive manner. These corrections are also targeted for aligning the texts where the text lines get twisted due to the leather writing surface&#39;s degradation. Hence, the clean images are there in the directory along with the direct binarized images. No effort has been made to obtain a balanced set in any way.</p> <p><strong>Tools:</strong><br> <strong>Binarization:</strong><br> The BiNet tool is available for scientific use upon request (m.a.dhal(at)rug.nl)</p> <p><strong>Image Morphing:</strong><br> In the original article, data augmentation was performed using image morphing. The tool is available on GitHub:<br> https://github.com/GrHound/imagemorph.c</p> <p><strong>Features for writer identification:</strong><br> Lambert Schomaker<br> http://www.ai.rug.nl/~lambert/allographic-fraglet-codebooks/allographic-fraglet-codebooks.html<br> http://www.ai.rug.nl/~lambert/hinge/hinge-transform.html<br> <em><strong>1.&nbsp;</strong>L. Schomaker &amp; M. Bulacu (2004). Automatic writer identification using connected-component contours and edge-based features of upper-case Western script. IEEE Transactions on Pattern Analysis and Machine Intelligence, Vol 26(6), June 2004, pp. 787 - 798.<br> <strong>2. </strong>Bulacu, M. &amp; Schomaker, L.R.B. (2007). Text-independent Writer Identification and Verification Using Textural and Allographic Features, &nbsp;IEEE Trans. on Pattern Analysis and Machine Intelligence (PAMI), Special Issue - Biometrics: Progress and Directions, April, 29(4), p. 701-717.</em><br> &nbsp;<br> The features (hinge, fraglets) have been combined in a single MS Windows application, GIWIS, which is available for scientific use upon request (l.r.b.schomaker(at)rug.nl)</p> <p><strong>If you have any question, please contact us:</strong><br> Maruf A. Dhali &lt;m.a.dhali(at)rug.nl&gt;<br> Lambert Schomaker &lt;l.r.b.schomaker(at)rug.nl&gt;<br> Mladen Popović &lt;m.popovic(at)rug.nl&gt;</p> <p><strong>Please cite our papers if you use this data set:</strong><br> <em><strong>1.</strong> Popović, M., Dhali, M. A., &amp; Schomaker, L. (2020). Artificial intelligence based writer identification generates new evidence for the unknown scribes of the Dead Sea Scrolls exemplified by the Great Isaiah Scroll (1QIsa<sup>a</sup>). arXiv preprint arXiv:2010.14476.<br> <strong>2. </strong>Dhali, M. A., de Wit, J. W., &amp; Schomaker, L. (2019). Binet: Degraded-manuscript binarization in diverse document textures and layouts using deep encoder-decoder networks. arXiv preprint arXiv:1911.07930.</em></p>

opencc-by-4.0Jan 2021View details →
zenodo44/100

Porto Santo landscape features and endemic lichens occurrence data

<p>Landscape features of Porto Santo island of and observation data of endemic lichens belonging to Sparrius et al. 2017, Bryologist.</p>

openmit-licenseJul 2017View details →
zenodo44/100

MelanoDB: Dataset files of clinical and molecular features of advanced melanoma patients treated with MAPK inhibitors

<p>MAPK inhibitors have significantly improved overall survival in patients with metastatic melanoma disease but their efficacy is still limited by primary or acquired resistance. Several studies have attempted to predict response to MAPK inhibitor therapy, however the lack of a consistent cohort prevents better definition of associations between treatment efficacy and clinical and/or molecular features. Here, we present MelanoDB, a collection of patients with metastatic melanoma treated with MAPK inhibitors. We formatted data from 8 different studies for a total of 417 cases to gather common clinical and molecular features. Whole or partial exome sequencing is available for 191 cases and gene expression for 132 cases. We provide a web application to explore the integrated data and its distribution among the collected studies, and we share this dataset to the scientific community according to FAIR principles</p> <p>These data are available under the licence CC-BY-SA.</p> <p>Here we provide a web application viewer of the database content: http://melanodb-ircm.montp.inserm.fr/</p> <p>You are requested to cite this repository in the case of using these data in a publication.</p>

opencc-by-4.0Jul 2024View details →
zenodo44/100

CLDF dataset derived from Lee's "Phonological Features of Caijia" from 2023

<p>Cite the source of the dataset as:</p> <blockquote> <p>Lee, Man Hei (2023): Phonological features of Caijia that are notable from a diachronic perspective. Journal of Historical Linguistics. DOI: https://doi.org/10.1075/jhl.21025.lee</p> </blockquote>

opencc-by-4.0Aug 2024View details →
zenodo44/100

Dataset: Features of animal babbling in the vocal ontogeny of the gray mouse lemur

<p>Dataset used in the unsupervised cluster analysis of the publication "Features of animal babbling in the vocal ontogeny of the gray mouse lemur (<em>Microcebus murinus</em>)"</p> <p><strong>Abstract</strong></p> <p>In human infants babbling is an important developmental stage of vocal plasticity to acquire maternal language. To investigate parallels in the vocal development of human infants and non-human mammals, seven key features of human babbling were defined, which are up to date only shown in bats and marmosets. This study will explore whether these features can also be found in gray mouse lemurs by investigating how infant vocal streams gradually resemble the structure of the adult trill call, which is not present at birth. Using unsupervised clustering, we distinguished six syllable types, whose sequential order gradually reflected the adult trill. A subset of adult syllable types was produced by several infants, with the syllable production being rhythmic, repetitive, and independent of the social context. The temporal structure of the calling bouts and the tempo-spectral features of syllable types became adult-like at the age of weaning. The age-dependent changes in the acoustic parameters differed between syllable types, suggesting that they cannot solely be explained by physical maturation of the vocal apparatus. Since gray mouse lemurs exhibit five features of animal babbling, they show parallels to the vocal development of human infants, bats, and marmosets.</p> <p>&nbsp;</p> <p>For details concerning the recording of the calling bouts confer to the publication at doi:10.1038/s41598-023-47919-7</p>

opencc-by-sa-4.0Dec 2023View details →
zenodo44/100

MRI Neonatal Lung Segmentation and 3D Morphologic Features

<p>We developed an ensemble of deep convolutional neural networks (2D-UNets) to perform automated neonatal lung segmentation from MRI sequences. A three-dimensional reconstruction is used to calculate MRI features for lung volume, shape, pixel intensity, and surface.</p> <p>In addition, ML Models for severity prediction of Bronchopulmonary Dysplasia (BPD) are implemented as an applied example of the use of MRI lung volumetric features for disease prognosis.</p> <p>This dataset comprises:</p> <ul> <li>Three pretrained 2D-UNet Models for Neonatal MRI Lung Segmentation.</li> <li>Resulting performances and features per MRI-sequence.</li> </ul> <p>See Publication:</p> <p>Automated MRI Lung Segmentation and 3D Morphologic Features for Quantification of Neonatal Lung Disease (2023)</p> <p><a href="https://doi.org/10.1148/ryai.220239">https://doi.org/10.1148/ryai.220239</a></p>

opencc-by-4.0Oct 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record