Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
58
datasets available to search
ShareScore release 0.7.1
Dataset results
58 results for “Feature Extraction”
CoAID dataset with multiple extracted features (both sparse and dense) and degraded by OCR
<p>This is the same datasets as:</p> <p>Guillaume Bernard. (2022). CoAID dataset with multiple extracted features (both sparse and dense) (1.0) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.6630405</p> <p>But with texts degraded by OCR as described in:</p> <p>Guillaume Bernard. (2022). CoAID dataset texts with OCR degradations (1.0) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.6630710</p>
Fibvid dataset with multiple extracted features (both sparse and dense)
<p>This is a publication of the FibVid dataset originaly dedicated to fake news detection. We changed here the purpose of this dataset in order to use it in the context of event tracking in press documents.</p> <p>Kim, Jisu, Jihwan Aum, SangEun Lee, Yeonju Jang, Eunil Park, et Daejin Choi. 2021. « FibVID: Comprehensive Fake News Diffusion Dataset during the COVID-19 Period ». <em>Telematics and Informatics</em> 64 (novembre): 101688. <a href="https://doi.org/10.1016/j.tele.2021.101688">https://doi.org/10.1016/j.tele.2021.101688</a>.</p> <p>In this dataset, we provide multiple features extracted from the text itself. <strong>Please note the text is missing from the dataset published in the CSV format for copyright reasons. You can download the original datasets and manually add the missing texts from the original publications.</strong></p> <p>Features are extracted using:</p> <p>- A corpus of reference articles in multiple languages languages for TF-IDF weighting. (<em>features_news</em>) [1]</p> <p>- A corpus of tweets reporting news for TF-IDF weighting. (<em>features_tweets)</em> [1]</p> <p>- A S-BERT model [2] that uses <em>distiluse-base-multilingual-cased-v1 </em>(called <em>features_use</em>) [3]</p> <p>- A S-BERT model [2] that uses <em>paraphrase-multilingual-mpnet-base-v2 </em>(called <em>features_mpnet</em>) [4]</p> <p><strong>References:</strong></p> <p>[1]: Guillaume Bernard. (2022). Resources to compute TF-IDF weightings on press articles and tweets (1.0) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.6610406</p> <p>[2]: Reimers, Nils, et Iryna Gurevych. 2019. « Sentence-BERT: Sentence Embeddings Using Siamese BERT-Networks ». In <em>Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)</em>, 3982‑92. Hong Kong, China: Association for Computational Linguistics. <a href="https://doi.org/10.18653/v1/D19-1410">https://doi.org/10.18653/v1/D19-1410</a>.</p> <p>[3]: https://huggingface.co/sentence-transformers/distiluse-base-multilingual-cased-v1</p> <p>[4]: https://huggingface.co/sentence-transformers/paraphrase-multilingual-mpnet-base-v2</p>
Event Registry dataset with multiple extracted features (both sparse and dense)
<p>This is a republication of the Event Registry dataset originaly published by:</p> <p>Rupnik, Jan, Andrej Muhic, Gregor Leban, Primoz Skraba, Blaz Fortuna, et Marko Grobelnik. 2016. « News Across Languages - Cross-Lingual Document Similarity and Event Tracking ». <em>Journal of Artificial Intelligence Research</em> 55 (janvier): 283‑316. <a href="https://doi.org/10.1613/jair.4780">https://doi.org/10.1613/jair.4780</a>.</p> <p>And reorganised for document tracking by:</p> <p>Miranda, Sebastião, Artūrs Znotiņš, Shay B. Cohen, et Guntis Barzdins. 2018. « Multilingual Clustering of Streaming News ». In <em>2018 Conference on Empirical Methods in Natural Language Processing</em>, 4535‑44. Brussels, Belgium: Association for Computational Linguistics. <a href="https://www.aclweb.org/anthology/D18-1483/">https://www.aclweb.org/anthology/D18-1483/</a>.</p> <p>In this dataset, we provide multiple features extracted from the text itself. <strong>Please note the text is missing from the dataset published in the CSV format for copyright reasons. You can download the original datasets and manually add the missing texts from the original publications.</strong></p> <p>Features are extracted using:</p> <p>- A corpus of reference articles in multiple languages languages for TF-IDF weighting. (<em>features_news</em>) [1]</p> <p>- A corpus of tweets reporting news for TF-IDF weighting. (<em>features_tweets)</em> [1]</p> <p>- A S-BERT model [2] that uses <em>distiluse-base-multilingual-cased-v1 </em>(called <em>features_use</em>) [3]</p> <p>- A S-BERT model [2] that uses <em>paraphrase-multilingual-mpnet-base-v2 </em>(called <em>features_mpnet</em>) [4]</p> <p><strong>References:</strong></p> <p>[1]: Guillaume Bernard. (2022). Resources to compute TF-IDF weightings on press articles and tweets (1.0) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.6610406</p> <p>[2]: Reimers, Nils, et Iryna Gurevych. 2019. « Sentence-BERT: Sentence Embeddings Using Siamese BERT-Networks ». In <em>Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)</em>, 3982‑92. Hong Kong, China: Association for Computational Linguistics. <a href="https://doi.org/10.18653/v1/D19-1410">https://doi.org/10.18653/v1/D19-1410</a>.</p> <p>[3]: https://huggingface.co/sentence-transformers/distiluse-base-multilingual-cased-v1</p> <p>[4]: https://huggingface.co/sentence-transformers/paraphrase-multilingual-mpnet-base-v2</p>
Training Datasets for Epilepsy Analysis: Preprocessing and Feature Extraction from EEG Time Series
<h2>The files include the 20 training datasets, in csv format, from 20 epileptic patients. Each set of data is described by 1080 features extracted using the sliding window technique.</h2>
Event Registry dataset with multiple extracted features (both sparse and dense) and degraded by OCR
<p>This is the same dataset as:</p> <p>Guillaume Bernard. (2022). Event Registry dataset with multiple extracted features (both sparse and dense) (1.0) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.6630367</p> <p>But with texts degraded by OCR as described in:</p> <p>Guillaume Bernard. (2022). Event Registry dataset texts with OCR degradations and synthesised segmentation (1.0) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.6631305</p>
Event Registry titles dataset with multiple extracted features (both sparse and dense) and degraded by OCR
<p>This is the same dataset as:</p> <p>Guillaume Bernard. (2022). Event Registry titles only dataset with multiple extracted features (both sparse and dense) (1.0) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.6630447</p> <p>But with texts degraded by OCR as described in:</p> <p>Guillaume Bernard. (2022). Event Registry titles dataset texts with OCR degradations (1.0) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.6630828</p>
FibVid dataset with multiple extracted features (both sparse and dense) and degraded by OCR
<p>This is the same dataset as:</p> <p>Guillaume Bernard. (2022). Fibvid dataset with multiple extracted features (both sparse and dense) (1.0) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.6630409</p> <p>But with texts degraded by OCR as described in:</p> <p>Guillaume Bernard. (2022). FibVid dataset texts with OCR degradations (1.0) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.6630758</p>
Graph topological features extracted from expression profiles of neuroblastoma patients
<p><strong>Introduction</strong></p> <p>This dataset contains the data described in the paper titled "A deep neural network approach to predicting clinical outcomes of neuroblastoma patients." by Tranchevent, Azuaje and Rajapakse. More precisely, this dataset contains the topological features extracted from graphs built from publicly available expression data (see details below). This dataset does not contain the original expression data, which are available elsewhere. We thank the scientists who did generate and share these data (please see below the relevant links and publications).</p> <p> </p> <p><strong>Content</strong></p> <p>File names start with the name of the publicly available dataset they are built on (among "Fischer", "Maris" and "Versteeg"). This name is followed by a tag representing whether they contain raw data ("raw", which means, in this case, the raw topological features) or TF formatted data ("TF", which stands for TensorFlow). This tag is then followed by a unique identifier representing a unique configuration. The configuration file "Global_configuration.tsv" contains details about these configurations such as which topological features are present and which clinical outcome is considered.</p> <p>The code associated to the same manuscript that uses these data is at <a href="https://gitlab.com/biomodlih/SingalunDeep">https://gitlab.com/biomodlih/SingalunDeep</a>. The procedure by which the raw data are transformed into the TensorFlow ready data is described in the paper.</p> <p> </p> <p><strong>File format</strong></p> <p>All files are TSV files that correspond to matrices with samples as rows and features as columns (or clinical data as columns for clinical data files). The data files contain various sets of topological features that were extracted from the sample graphs (or Patient Similarity Networks - PSN). The clinical files contain relevant clinical outcomes.</p> <p>The raw data files only contain the topological data. For instance, the file "Fischer_raw_2d0000_data_tsv" contains 24 values for each sample corresponding to the 12 centralities computed for both the microarray (<em>Fischer-M</em>) and RNA-seq (<em>Fischer-R</em>) datasets. The TensorFlow ready files do not contain the sample identifiers in the first column. However, they contain two extra columns at the end. The first extra column is the sample weights (for the classifiers and because we very often have a dominant class). The second extra column is the class labels (binary), based on the clinical outcome of interest.</p> <p> </p> <p><strong>Dataset details</strong></p> <p>The <em>Fischer</em> dataset is used to train, evaluate and validate the models, so the dataset is split into train / eval / valid files, which contains respectively 249, 125 and 124 rows (samples) of the original 498 samples. In contrast, the other two datasets (<em>Maris</em> and <em>Versteeg</em>) are smaller and are only used for validation (and therefore have no training or evaluation file).</p> <p>The <em>Fischer</em> dataset also has more data files because various configurations were tested (see manuscript). In contrast, the validation, using the <em>Maris</em> and <em>Versteeg</em> datasets is only done for a single configuration and there are therefore less files.</p> <p>For <em>Fischer</em>, a few configurations are listed in the global configuration file but there is no corresponding raw data. This is because these items are derived from concatenations of the original raw data (see global configuration file and manuscript for details).</p> <p> </p> <p><strong>References</strong></p> <p>This dataset is associated with Tranchevent L., Azuaje F.. Rajapakse J.C., A deep neural network approach to predicting clinical outcomes of neuroblastoma patients.</p> <p>If you use these data in your research, please do not forget to also cite the researchers who have generated the original expression datasets.</p> <p><em>Fischer</em> dataset:</p> <ul> <li>Zhang W. et al., Comparison of RNA-seq and microarray-based models for clinical endpoint prediction. Genome Biology 16(1) (2015). doi:10.1186/s13059-015-0694-1</li> <li>Wang C. et al., The concordance between RNA-seq and microarray data depends on chemical treatment and transcript abundance. Nat. Biotechnol. 32(9), 926–932. doi:10.1038/nbt.3001</li> </ul> <p><em>Versteeg</em> dataset:</p> <ul> <li>Molenaar J.J. et al., Sequencing of neuroblastoma identifies chromothripsis and defects in neuritogenesis genes. Nature 483(7391), 589–593. doi:10.1038/nature10910</li> </ul> <p><em>Maris</em> dataset:</p> <ul> <li>Wang Q. et al., Integrative genomics identifies distinct molecular classes of neuroblastoma and shows that multiple genes are targeted by regional alterations in DNA copy number. Cancer Res. 66(12), 6050–6062. doi:10.1158/0008-5472.CAN-05-4618</li> </ul>
Custom script for feature extraction from Genbank files
<p><span>Microbes are thought to be distributed and circulated around the world, but the connection between marine and terrestrial microbiomes is largely unknown. We use <em>Plantibacter</em>, a representative plant-associated genus, as our research model to show the global distribution and adaptation of plant-related bacteria in plant-free environments, especially in the remote Southern Ocean and the deep Atlantic Ocean. The marine isolates and their plant-associated relatives shared over 98% whole-genome average nucleotide identity (ANI), indicating recent divergence and ongoing speciation from plant-related niches to marine environments. Comparative genomics revealed that the marine strains acquired new genes via horizontal gene transfer from non-<em>Plantibacter</em> species and refined existing genes through positive selection to improve adaptation to new habitats. Meanwhile, marine strains retained the ability to interact with plants, such as modifying root system architecture and promoting germination. <em>Plantibacter </em>species were further found to be widely distributed in marine environments, revealing an unrecognized phenomenon that plant-associated microbiomes have colonized the ocean, which could serve as a reservoir for plant growth-promoting microbes. This study demonstrates the presence of an active reservoir of terrestrial plant growth-promoting bacteria in remote marine systems and advances our understanding of the microbial connections between plant-associated and plant-free environments at the genome level.</span></p>
Collection of literature and extracted features in the field of under-frequency load shedding for the period 1954-2020
<p>The number of publications in the field of UFLS research is rapidly increasing. This confirms that the topic is timely, but due to the amount of publications, it is easy to lose track of the prevailing concepts driving the latest technologies. To overcome this problem, researchers resort to subjective review articles in the literature in which individual authors attempt to systematically categorize UFLS algorithms according to their own understanding of a variety of approaches. Since this is not exactly a trivial task and is undoubtedly subject to personal interpretation, it is better to use specialized mathematical techniques for this purpose. Recently, it has been shown that the use of clustering techniques and graph theory can be a useful tool for systematic literature reviews.</p> <p>If one chooses to search for similarities between UFLS algorithms using such techniques, one must first collect the relevant literature in the field and extract important information. Therefore, this collection provides an Excel spreadsheet (.xlsx) of 381 publications in the field of UFLS protection for the period from 1954 to 2020, with the following information for each publication:</p> <ol> <li><em>authors</em>,</li> <li><em>title</em>,</li> <li><em>year</em> of publication (journal, conference proceedings, doctoral dissertation, master's thesis, bachelor's thesis, other),</li> <li>the <em>source</em> from which we obtained the publication,</li> <li><em>DOI</em> (Digital Object Identifier),</li> <li>extracted <em>general features</em>, and</li> <li>extracted <em>specific features</em>.</li> </ol> <p>Relevant literature was obtained from various open access journals, impact factor journals, repositories of various universities, and archives of various electric utilities and transmission system operators.</p> <p>If the publication in the row can be assigned to the feature specified in column I3-BJ3, the corresponding cell contains the value "1". Otherwise, i.e., if the feature specified in column I3-BJ3 cannot be assigned to the publication, the cell is empty. Detailed descriptions of the individual feature can be found in "Description_of_features.docx". Namely, two identical documents are available, one in English and one in Slovene.</p> <p>Updated versions with more recent data will be uploaded with a differing version number and doi.</p>
Feature Extraction Using Hidden Markov Model for a Phonetic Process
<p>Speech is one of the primary forms of communication among humans. In real life, a dictionary is used to seek the pronunciation of a complex word; but, for computers, this look-up table is called a phonetic dictionary. A speech recognition process tags a word-utterance to its phoneme structure, thereby returning the grapheme representation. However, the speech recognition process is challenging because of the contextual relationship between words and sentences, dependent on speakers’ intentions. Further, factors influencing time, accents, noisy environment, and data security impose accuracy threats. The present research study proposes a new hybrid speech recognition model by considering three significant aspects: sound generation through phonetic representation, sound acoustics for transmission, and sound reception on how the sound is received. These steps are achieved through a speech-to-text model divided into various stages such as noise removal, speech-pause detection, feature extraction through framing, and windowing by adopting Hidden Markov Model (HMM). The implementation is performed on a phonetic tool, Praat. The robustness of the model is estimated using evaluation metrics such as f-measure and accuracy, resulting in 98% and 99% scores, respectively. Thus, the proposed approach efficiently transforms the spoken words into their corresponding text.</p>
Data from: How many specimens make a sufficient training set for automated three dimensional feature extraction?
<p>Deep learning has emerged as a robust tool for automating feature extraction from 3D images, offering an efficient alternative to labour-intensive and potentially biased manual image segmentation methods. However, there has been limited exploration into the optimal training set sizes, including assessing whether artificial expansion by data augmentation can achieve consistent results in less time and how consistent these benefits are across different types of traits. In this study, we manually segmented 50 planktonic foraminifera specimens from the genus Menardella to determine the minimum number of training images required to produce accurate volumetric and shape data from internal and external structures. The results reveal unsurprisingly that deep learning models improve with a larger number of training images with eight specimens being required to achieve 95% accuracy. Furthermore, data augmentation can enhance network accuracy by up to 8.0%. Notably, predicting both volumetric and shape measurements for the internal structure poses a greater challenge compared to the external structure, due to low contrast differences between different materials and increased geometric complexity. These results provide novel insight into optimal training set sizes for precise image segmentation of diverse traits and highlight the potential of data augmentation for enhancing multivariate feature extraction from 3D images. </p>
Extracting Ridge and Valley Lines in Mountainous Areas from Airborne Lidar Data by Utilizing Line Feature Strength
<p><strong><span>Background</span></strong><strong><span>:</span></strong><span> </span><span>DEMs (digital elevation models) are very important in many fields, such as in Geomatics and in water conservation of mountainous areas etc. Geomorphic feature lines are necessary data for the topography interpolation and computation from DEMs.</span></p> <p><strong><span>Methods</span></strong><strong><span>:</span></strong><span> </span><span>Instead of the parameter space, we propose a novel automatic extraction of Geomorphic feature lines in the feature space from discrete airborne LiDAR (Light detection and ranging) data by TVM (tensor voting method) developed originally for image data in this article. A tensor field for discrete airborne LiDAR points is first established and then utilizing the TVM, a new geometric feature metric of data, the line feature strength, was captured. A practical line growing method based on the local maximum line feature strength is proposed in the article.</span></p> <p><strong><span>Results</span></strong><strong><span>:</span></strong><span> </span><span>Compared with the general line growing that is based on a certain threshold, our line growing method is quite effective, in particular for the extraction of primary and minor ridge and valley lines in mountainous areas.</span></p> <p><strong><span>Conclusions</span><span>:</span></strong><span> </span><span>The method presented in this paper is fast and automated and can furnish operators with a wealth of detailed information about minor line features. This will enable the extraction of ridge and valley lines tailored to specific requirements. It is no doubt that the method developed here can be generalized to a large amount of Lidar data.</span></p>
BRAIN Journal-Motor Imagery signal Classification for BCI System Using Empirical Mode Décomposition and Bandpower Feature Extraction-Figure 5. PSD (dB/Hz) vs freaquency (Hz) of each IMF showen in fig 4 in channel C4 (a) and in C3 (b)
<p> In Fig 5, we noted that ocular artifact frequency is generally low around 5Hz with high amplitude. This artifact appears mainly in IMF3 and IMF4. Finally, band power was applied for the new signal. As a last step, the logarithm of the BP is calculated in order to transform the distribution of this feature to a more Gaussian like shape, because the classifiers we used, such as HMMs and SVM assume normally distributed features.</p>
BRAIN Journal-Motor Imagery signal Classification for BCI System Using Empirical Mode Décomposition and Bandpower Feature Extraction-Figure 4. The EMD decomposition results for subject 2 when he imagines left hand movement
<p>Fig. 4 shows the EMD decomposition result of one-trial (left hand movement imagination) for subject 2 in the channels C3 and C4 respectively (the pre-filtered EEG signal used for this illustration is not corrupted by blinking artifact.). Each channel is decomposed into ten IMFs and one residue</p>
BRAIN Journal-Motor Imagery signal Classification for BCI System Using Empirical Mode Décomposition and Bandpower Feature Extraction-Figure 5b. PSD (dB/Hz) vs freaquency (Hz) of each IMF showen in fig 4 in channel C4 (a) and in C3 (b)
<p>Therefore, the new signal is reconstructed by keeping only the two first IMFs. EMD also allows eliminating the artifacts in the EEG during the recording sessions like eye blinks and eyeball movements. In Fig 5, we noted that ocular artifact frequency is generally low around 5Hz with high amplitude. This artifact appears mainly in IMF3 and IMF4. Finally, band power was applied for the new signal. As a last step, the logarithm of the BP is calculated in order to transform the distribution of this feature to a more Gaussian like shape, because the classifiers we used, such as HMMs and SVM assume normally distributed features.</p>
BRAIN Journal-Motor Imagery signal Classification for BCI System Using Empirical Mode Décomposition and Bandpower Feature Extraction-Figure 3. Hybrid EMD-BP approach for one trail feature extraction
<p>In this work, we propose a direct nonlinear approach to extract the more relevant IMFs corresponding to the different frequency components in the and bands and then obtain the BP in order to use them as features for mental task classification (see Fig. 3). The feature vector p used for the demonstration in this paper is composed, for each sample I, 1 < i < 2048, in a given trial (among a total of 160 trials) of four bandpower, calculated of the rhythms and in positions C3 and C4 Trad et al., 2011).</p>
BRAIN Journal-Motor Imagery signal Classification for BCI System Using Empirical Mode Décomposition and Bandpower Feature Extraction-Figure 4b. The EMD decomposition results for subject 2 when he imagines left hand movement
<p>d et al., 2011). Fig. 4 shows the EMD decomposition result of one-trial (left hand movement imagination) for subject 2 in the channels C3 and C4 respectively (the pre-filtered EEG signal used for this illustration is not corrupted by blinking artifact.). Each channel is decomposed into ten IMFs and one residue.</p>
BRAIN Journal-Motor Imagery signal Classification for BCI System Using Empirical Mode Décomposition and Bandpower Feature Extraction-Figure 6. The general conception of our asynchronous system BCI (offline - online) for reinforcement of a joystick movement
<p>Once the motor imagery is identified, a command may be associated to this mental task in order to control a machine (Prataksita et al., (2014)) (Guger et al., 1999). In this work, we constructed a new Simuhnk/MathWork model to translate on-line the EEG signals into low-level commands. Fig. 6 shows our experimental EEG-based BCI System </p>
BRAIN Journal-Motor Imagery signal Classification for BCI System Using Empirical Mode Décomposition and Bandpower Feature Extraction-Figure 1. General architecture of an online (BCI)
<p>One major challenge of our BCI system is to describe the signals EEG by a few relevant values called features i.e. step 3 in Fig (1). The success of the mental imagery classification depends on the choice of features used to characterize the raw EEG signals. These features can then be used in step 4 in order to classify the user’s mental state. Several approaches for feature extraction have been proposed in literature. </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.