Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
668
datasets available to search
ShareScore release 0.9.0
Dataset results
668 results for “embedding”
Dataset of "Perovskite QDs embedded in polymer as a wavelength-shifting layer for UV-sensitized silicon sensors"
<p>Detection of UV radiation is becoming increasingly important for many applications. Here we present novel UV sensor construction on the basis of standard Si detector modification. Wavelength shifting mechanism is achieved by the luminescence effect of perovskite quantum dots embedded in polymer layers. We comprehensively characterize these composite materials, various sensor modification routes and the optical properties of UV-enhanced visible sensors. Modified S1227-16 BG silicon photodiodes and S13360-1375 CS MPPC photodetectors with the enhanced UV response are successively manufactured.micrographs; cross sections of model simulation or prediction (MSP).</p>
Word Embedding of Amazon Product Review Corpus
<p>A word embedding of the <a href="https://www.cs.uic.edu/~liub/FBS/sentiment-analysis.html#datasets">Amazon Product Review Corpus</a> (<a href="https://www.doi.org/10.1145/1341531.1341560">Jindal and Liu, 2008</a>).</p> <p>Created using <a href="https://code.google.com/archive/p/word2vec/">Word2Vec</a> in CBOW mode, 500 dimensions and window size 5.</p> <p>Words have been lemmatised and particle verbs have been merged into a single token (e.g. <code>calm_down</code>).</p> <ul> </ul> <p> </p> <p><strong>Attribution</strong></p> <p>This dataset was created as part of the following publication:</p> <p>Marc Schulder, Michael Wiegand, Josef Ruppenhofer and Benjamin Roth (2017). <strong>"Towards Bootstrapping a Polarity Shifter Lexicon using Linguistic Features"</strong>. Proceedings of the 8th International Joint Conference on Natural Language Processing (IJCNLP). Taipei, Taiwan, November 27 - December 3, 2017. <a href="https://doi.org/10.5281/zenodo.3365609">DOI: 10.5281/zenodo.3365609</a>.</p> <p>If you use the data in your research or work, please cite the publication.</p>
Experimental data of dissipative embedded column base connections tested under cyclic lateral loading
<p>This experimental dataset is comprised of the following items:</p> <p>(a) the deduced experimental data of conventional/dissipative embedded column base connection specimens, which contains base moment, column drift ratio, and axial shortening responses (TestData.xlsx);</p> <p>(b) photos of each specimen taken during cyclic loading (C-N-0_Test_Photos.7z, D-M1-1_Test_Photos.7z, D-M1-3_Test_Photos.7z, D-M1-5_Test_Photos.7z, D-M2-2_Test_Photos.7z);</p> <p>(c) characteristic videos for each specimen that demonstrate the cyclic behavior (Test_Video.7z);</p> <p>(d) Digital image correlation (DIC) images taken during cyclic loading to obtain strain fields near the steel column/reinforced concrete foundation interface (C-N-0_DIC_Photos.7z, D-M1-1_DIC_Photos.7z, D-M1-3_DIC_Photos.7z, D-M1-5_DIC_Photos.7z, D-M2-2_DIC_Photos.7z); </p> <p>(e) Videos that demonstrate strain fields of column flanges of both conventional and dissipative embedded column base connection specimens (DIC_Video.7z) </p> <p>Please read the "README" file contained in each folder for more detailed information regarding each data.</p> <p> </p>
Diachronic word embeddings from 19th-century newspapers digitised by the British Library (1800-1919)
<p>Word vectors related to the paper <em>Machines in the media: semantic change in the lexicon of mechanization in 19th-century British newspapers </em>by Nilo Pedrazzini and Barbara McGillivray (2022).</p> <p>The embeddings were trained on a 4.2-billion-word corpus of 19th-century British newspapers using Word2Vec and the following parameters:</p> <pre><code>sg = True min_count = 1 window = 3 vector_size = 200 epochs = 5</code></pre> <p>The embeddings are divided into periods of ten years each, with the vectors from each decade aligned to the ones from the most recent decade (1910s) using Orthogonal Procrustes.</p> <p>See related GitHub repository for the full documentation: <a href="https://github.com/Living-with-machines/DiachronicEmb-BigHistData">https://github.com/Living-with-machines/DiachronicEmb-BigHistData</a></p> <p>Project webpage (Living with Machines): <a href="https://livingwithmachines.ac.uk/">https://livingwithmachines.ac.uk/</a></p>
Predictive simulations of ionization energies of solvated halide ions with relativistic embedded Equation of Motion Coupled-Cluster Theory: Figures
<p>This entry contains the sources for the figures included in the body of the paper titled "Predictive simulations of ionization energies of solvated halide ions with relativistic embedded Equation of Motion Coupled-Cluster Theory", by Yassine Bouchafra, Avijit Shee, Florent Réal, Valérie Vallet and André Severo Pereira Gomes, as well as those found in the supplementary information.</p> <p>It accompanies the dataset found at the DOI: 10.5281/zenodo.1477004</p> <p> </p> <p> </p>
3D model of antenna system embedded into building envelope for improved cellular signal transmission through load-bearing walls
<p>The purpose of this dataset is to supplement the data presented in our journal publication "Electromagnetic–Thermal Analyses of Distributed Antennas Embedded Into a Load-Bearing Wall" (see <a href="https://ieeexplore.ieee.org/document/10151683">https://ieeexplore.ieee.org/document/10151683</a>).</p> <p>This dataset contains the 3-D discretized model, without the internal numerical mesh, of the unit cell of the spiral antenna system embedded in a load bearing wall. The 3D model is in .STP format (see ISO 10303-21:2016), which can be imported into most commercial computer-aided design (CAD) software. The wall's dielectric properties are calculated using the model described in ITU-R P.2040-2 (<a href="https://www.itu.int/rec/R-REC-P.2040/en">https://www.itu.int/rec/R-REC-P.2040/en</a>, material parameter and calculation model are on pages 22-23). Materials used in the antenna system and their electrical and thermal parameters are given in the file materials.txt</p>
Pre-training Audio Embeddings
<p>Pre-trained audio embeddings (VGGish, OpenL3, YAMNet) of <a href="https://www.upf.edu/web/mtg/irmas">IRMAS</a> and <a href="https://zenodo.org/record/1432913">OpenMIC-2018</a> datasets, released in the following paper:</p> <p>Changhong Wang, Brian McFee, and Gaël Richard. "<strong>Transfer Learning and Bias Correction with Pre-trained Audio Embeddings</strong>". <em>Proceedings of the <a href="https://ismir2023.ismir.net/">International Society for Music Information Retrieval (ISMIR) Conference</a></em>, 2023.</p>
Research data supporting for "The embedded research librarian: a project partner"
<p>This dataset contains the data that supports the following paper: Féret, R. and Cros, M., 2019. The embedded research librarian: a project partner. <em>LIBER Quarterly</em>, 29(1), pp.1–20. DOI: <a href="https://dx.doi.org/10.18352/lq.10304">10.18352/lq.10304</a></p> <p>The dataset contains 3 files related to the bibliographic metadata of the publications of the 7 H2020 projects supported by the University Library of Lille and a general file providing the data for the table, figure 2 and 3 and for the data on H2020 projects coordinators:</p> <ul> <li>figures: this file contains the information related to the projects supported by the Library, including the data presented in the figure 2 (tab 1), the figure 3 (tab 2), the table 1 (tab 3) and the data on 2020 project coordinators (tab 4).</li> <li>wos_publications : the data extracted from the Web of Science for 106 publications (.txt, UTF-8, Windows), searched on the base of the 7 H2020 projects Cordis number.</li> <li>refined_wos_publications : the same data after having been transformed into a .xlsx format in the tool OpenRefine.</li> <li>processed_publications : contains the main bibliographic data (authors, article title, source title, DOI, date of publication) and their open status.</li> </ul> <p><strong>Abstract of the paper</strong><br> This paper presents new services developed by the Lille University Library for European and National research project coordinators. This is a specific audience that libraries are not used to target, with a widely recognised institutional status and academic background. Supporting them in their coordination activities is an opportunity to gain a new role for libraries, which starts from the design of research at the submission stage and lasts several years after, during the project lifetime. These services help coordinators to meet their funders’ expectations on open access and research data management. It is also a way to develop new collaborations with research units and some university services, such as the Grant Office. The Lille University Library has already supported the writing of forty grant proposals since 2017, including about thirty since early 2019. The Library currently follows twelve projects on open access, research data management or both. This second figure is likely to increase in 2020 due to the number of projects supported at submission stage since the beginning of 2019. The paper describes our set of services and the lessons we learned from our approach.</p>
Latin embeddings
<p>Lemma embeddings for Latin can be downloaded <a href="https://embeddings.lila-erc.eu/samples/download/">here</a>.</p> <p>Embeddings have been evaluated on a novel benchmark for the synonym selection task available in the file <em>syn-selection-benchmark-Latin.tsv</em>.</p> <p>You can visually explore the embeddings online: <a href="https://embeddings.lila-erc.eu/">https://embeddings.lila-erc.eu/</a>.</p> <p>These resources are licensed under a <a href="http://creativecommons.org/licenses/by-nc-sa/4.0/">Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License</a>.</p>
Word embeddings learnt on MEDLINE abstracts
<p>Accompanying a preprint manuscript and code repository, this folder contains both raw text data and learnt word embeddings. The data source is the set of MEDLINE articles published on or after 2000. Preprocessing consists of extraction of each article's title and abstract and some minor text processing. The result is a corpus of 10.5 million documents in a single 14 GB file. </p> <p>word2vec and fastText are used to learn word embeddings on this corpus and three sets of word embeddings are shared here: 1) word2vec skip-gram, 2) word2vec CBOW, and 3) fastText skip-gram. All three sets use the default parameters of the software (e.g. context=5) with the exception of hierarchical softmax optimization and dimension=200.</p> <p>Preprint manuscript: https://arxiv.org/abs/1705.06262<br> GitHub repository: https://github.com/vincentmajor/ctsa_prediction</p>
Jacdac: Service-based Prototyping of Embedded Systems (PLDI 2024 Artifact Evaluation)
<p>This artifact allows others to reproduce and explore the results seen in "Jacdac: Service-based Prototyping of Embedded Systems". The artifact contains a prebuilt docker image and the Dockerfile source used to produce the prebuilt docker image. Evaluators should follow the README contained in this artifact for complete instruction.</p>
scGraph2Vec: a deep generative model for gene embedding augmented by Graph Neural Network and single-cell omics data
<p>This repository contains the training data and source code to reproduce the results of our paper:<br>scGraph2Vec: a deep generative model for gene embedding augmented by Graph Neural Network and single-cell omics data</p> <p>More description can be also found in GitHub (https://github.com/LPH-BIG/scGraph2Vec).</p>
Embedding Evaluation Data for South African Languages
<p><strong>WordSim and Simlex Data for South African Languages</strong></p> <ul> <li>Setswana</li> <li>Sepedi</li> </ul> <p><strong>Embedding Evaluation Data for South African Languages</strong></p> <p><strong>Dataset Information\</strong></p> <p>The datasets(Simlex and WordSim) contain pairs of Setswana and Sepedi words that have been assigned similarity ratings by humans to measure semantic relatedness. The word-pairs(Simlex and WordSim) are manually translated from English to Setswana and Sepedi. The evaluation task aims to find the degree of correlation between the scores provided by the model and the human rating, the score of the model is collected by computing the cosine similarity of corresponding vectors for word pairs.</p> <p>Online Repository link</p> <ul> <li><a href="https://zenodo.org/record/5673974">Zenodo Data Repository</a> - Link to the data repository.</li> </ul> <p>Authors</p> <ul> <li><strong>Vukosi Marivate</strong> - <a href="https://twitter.com/vukosi">@vukosi</a></li> <li><strong>Valencia Wagner</strong></li> <li><strong>Mack Makgatho</strong></li> <li><strong>Tshephisho Sefara</strong></li> </ul> <p>See also the list of <a href="https://github.com/dsfsi/embedding-eval-data//contributors">contributors</a> who participated in this project.</p> <p>Citing the dataset</p> <p>To appear in conference proceedings</p> <blockquote> <p>@article{Makgatho_Marivate_Sefara_Wagner_2022, title={Training Cross-Lingual embeddings for Setswana and Sepedi}, <br> volume={3}, <br> url={https://upjournals.up.ac.za/index.php/dhasa/article/view/3822}, <br> DOI={10.55492/dhasa.v3i03.3822}, <br> number={03},<br> journal={Journal of the Digital Humanities Association of Southern Africa },<br> author={Makgatho, Mack and Marivate, Vukosi and Sefara, Tshephisho and Wagner, Valencia}, <br> year={2022}, <br> month={Feb.}}</p> </blockquote>
Dataset: Environment effects on X-ray absorption spectra with quantum embedded real-time Time-dependent density functional theory approaches
<p>This dataset collects the outputs from real-time TDDFT simulation of X-ray absorption of halides in model systems, using the frozen density embedding (FDE) and block-orthogonalized Manby-Miller embedding (BOMME), as well as processing tools and scripts used to carry out the calculations.</p>
DBpedia RDF2Vec Graph Embeddings
<p>DBpedia graph embeddings using RDF2Vec. RDF2Vec embedding generation code can be found <a href="https://github.com/dwslab/jRDF2Vec">here</a> and is based on a publication by Portisch et al. [1].</p> <p>The embeddings dataset consists of 200-dimensional vectors of DBpedia entities (from 1/9/2021).</p> <p>Figure of cosine similarities between a selected set of DBpedia entities are provided in the dataset <a href="https://zenodo.org/record/6384728/files/heatmap.pdf?download=1">here</a>.</p> <p> </p> <p><strong>Generating Embeddings</strong></p> <p>The code for generating these embeddings can be found <a href="https://github.com/EDAO-Project/DBpediaEmbedding">here</a>.</p> <p>Run the run.sh script that wraps all the necessary commmands to generate embeddings</p> <pre><code class="language-bash">bash run.sh</code></pre> <p>The script downloads a set of DBpedia files, which are listed in <code>dbpedia_files.txt</code>. It then builds a Docker image and runs a container of that image that generates the embeddings for the DBpedia graph defined by the DBpedia files.</p> <p>A folder <code>files</code> is created containing all the downloaded DBpedia files, and a folder <code>embeddings/dbpedia</code> is created containing the embeddings in <code>vectors.txt</code> along a set of random walk files.</p> <p> </p> <p><strong>Run Time of Embeddings Generation</strong></p> <p>Generating embeddings can take more than a day, but it depends on the number of DBpedia files chosen to be downloaded. Following are some basic run time statistics when embeddings are generated on a 64 GB RAM, 8 cores (AMD EPYC), 1 TB SSD, 1996.221 MHz machine.</p> <ul> <li><strong>Total</strong>: 1 day, 8 hours, 52 minutes, 41 seconds</li> <li><strong>Walk generation</strong>: 0 days, 7 minutes, 24 minutes, 36 seconds</li> <li><strong>Training</strong>: 1 day, 1 hour, 28 minutes, 5 seconds</li> </ul> <p> </p> <p><strong>Parameters Used</strong></p> <p>Here is listed the parameters used to generate the embeddings provided here:</p> <ul> <li><strong>Number of walks per entity</strong>: 100</li> <li><strong>Depth (hops) per walk</strong>: 4</li> <li><strong>Walk generation mode</strong>: RANDOM_WALKS_DUPLICATE_FREE</li> <li><strong>Threads</strong>: # of processors / 2</li> <li><strong>Training mode</strong>: sg</li> <li><strong>Embeddings vector dimension</strong>: 200</li> <li><strong>Minimum word2vec word count</strong>: 1</li> <li><strong>Sample rate</strong>: 0.0</li> <li><strong>Training window size</strong>: 5</li> <li><strong>Training epochs</strong>: 5</li> </ul>
Classical Tibetan Word Embeddings
<p>Classical Tibetan word embeddings trained with FastText based on the 2018 version of the BDRC corpus, a segmented version of which is available on Zenodo:</p> <p>Meelen, Marieke, & Roux, Élie. (2020). The Annotated Corpus of Classical Tibetan (ACTib) - Version 2.0 (Segmented & POS-tagged) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.3951503</p> <p>This is the first version trained with default FastText settings (100D) for a pilot study on Chinese-Tibetan crosslinguistic Semantic Textual Similarity:</p> <p>Felbur, Rafal, Marieke Meelen & Paul Vierthaler (2022), 'Crosslinguistic Semantic Textual Similarity of Buddhist Chinese and Classical Tibetan' in <em>Journal of Open Humanities Data</em>.</p> <p>This research was done with generous funding from the Open Philology project. This project (running 2018–2022) is funded by the European Research Council (ERC) under the Horizon 2020 program (Advanced Grant agreement No 741884). It is based at the Leiden University Institute for Area Studies.</p>
Deep Reference Mining from Scholarly Literature in the Arts and Humanities - Pre-trained word embeddings
<p>Pre-trained word vectors of dimensionality 100 and 300 for the publication: Deep Reference Mining from Scholarly Literature in the Arts and Humanities, submitted to Frontiers in Digital Humanities.</p> <p>The corpus of scholarly publications from which these vectors were trained is under copyright, therefore we publish these vectors for reproducibility. Please refer to the publication's repository for further details: <a href="https://github.com/dhlab-epfl/LinkedBooksDeepReferenceParsing">https://github.com/dhlab-epfl/LinkedBooksDeepReferenceParsing</a>.</p> <p>These vectors were trained using Gensim 3.1.0. The corpus was preprocessed as follows:</p> <ol> <li>word tokenization with NLTK word_punct tokenizer.</li> <li>digits were converted into the $NUM$ token</li> <li>words less frequent than 5 times, for every document, were converted to the $UNK$ token</li> <li>vectors were trained using the function: Word2Vec(window=5, min_count=5, sg=1)</li> </ol>
Embeddings from protein language models predict conservation and variant effects
<p>For this work, we used protein language model representations (embeddings) to predict sequence conservation without multiple sequence alignments (MSAs). Embeddings alone predicted residue conservation almost as accurately from single sequences as ConSeq using MSAs (two-state Matthew Correlation Coefficient – MCC - for ProtT5 embeddings of 0.596±0.006 vs. 0.608±0.006 for ConSeq).</p> <p><strong><em>ConSurf10k</em>- Dataset for the development of ProtT5cons:</strong> The method (ProtT5cons) predicting residue conservation used <em>ConSurf-DB </em>(Ben Chorin et al. 2020). This resource provided sequences and conservation for 89,673 proteins. For all, experimental high-resolution three-dimensional (3D) structures were available in the Protein Data Bank (PDB) (Berman et al. 2000). As standard-of-truth for the conservation prediction, we used the values from ConSurf-DB generated using HMMER (Mistry et al. 2013), CD-HIT (Fu et al. 2012), and MAFFT-LINSi (Katoh and Standley 2013) to align proteins in the PDB (Burley et al. 2019). For proteins from families with over 50 proteins in the resulting MSA, an evolutionary rate at each residue position is computed and used along with the MSA to reconstruct a phylogenetic tree. The ConSurf-DB conservation scores ranged from 1 (most variable) to 9 (most conserved). The PISCES server (Wang and Dunbrack 2003) was used to redundancy reduce the data set such that no pair of proteins had more than 25% pairwise sequence identity. We removed proteins with resolutions >2.5Å, those shorter than 40 residues, and those longer than 10,000 residues. The resulting data set (ConSurf10k) with 10,507 proteins (or domains) was randomly partitioned into training (9,392 sequences), cross-training/validation (555) and test (519) sets.</p> <p>Uploaded data:</p> <ul> <li>ConSuf10k_PDBid_seq_cons.fasta: fasta file with PDBid, sequence and conservation annotation</li> <li>consurf10k_test_ids.txt: txt file with id's of test set</li> <li>consurf10k_train_ids.txt: txt file with id's of train set</li> <li>consurf10k_val_ids.txt: txt file with id's of cross-validation set</li> </ul>
Datasets for 'Mandrake: visualising microbial population structure by embedding millions of genomes into a low-dimensional representation'
<p>Datasets for the paper '<strong>Mandrake: visualising microbial population structure by embedding millions of genomes into a low-dimensional representation</strong>'</p> <p>Files:</p> <ul> <li>616k* - Files for the analysis of 661k bacterial genomes from the SRA (note typo 616-661k). Includes mandrake output and input files (.npz)</li> <li>gps_acc - Files for the analysis of 20k S. pneumoniae accessory genomes from the GPS project. Original accessory matrix is gps_gene_presence_absence.Rtab</li> <li>sc2million_v1* - Files for the analysis of ~1M SARS-CoV-2 genomes. sc2million_v3.npz are the input distances.</li> <li>sce<commit hash>.qdrep - Nvidia systems profile of code at that commit hash</li> <li>sce<commit hash>.ncu-rep - Nvidia kernel profile of code at that commit hash</li> </ul>
German CBOW FastText embeddings with min count 250
<p>FastText embeddings built from Common Crawl german dataset</p> <table> <caption>Parameters</caption> <thead> <tr> <th scope="col">Parameters</th> <th scope="col">Value(s)</th> </tr> </thead> <tbody> <tr> <td>Dimensions</td> <td>256 and 384</td> </tr> <tr> <td>Context window</td> <td>5</td> </tr> <tr> <td>Negative sampled</td> <td>10</td> </tr> <tr> <td>Epochs</td> <td>1</td> </tr> <tr> <td>Number of buckets</td> <td>131072 or 262144</td> </tr> <tr> <td>Min n</td> <td>3</td> </tr> <tr> <td>Max n</td> <td>6</td> </tr> </tbody> </table>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.