Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
668
datasets available to search
ShareScore release 0.9.0
Dataset results
668 results for “embedding”
Universal Knowledge Graph Embeddings
<p>The dataset provides embeddings for entities and relations in DBpedia (English) and Wikidata. The two knowledge graphs are first merged using a novel approach that we developed by leveraging the sameAs links between them. Then, we used the state-of-the-art embedding model ConEx to compute embeddings of the merge. Our embeddings are called universal knowledge graph embeddings.</p>
Experimental 3D convex hull tiling of a wild moss specimen ( Rhytidiadelphus or Eurhynchium ?) embedded in 1% agarose
<p>Experimental 3D convex hull tiling of a wild moss specimen ( Rhytidiadelphus or Eurhynchium ?) embedded in 1% agarose</p> <p>Raw data consists of 18 tiles of varying number of Z stacks, depending on where the 3D convex hull points were determined.</p> <p>Images acquired on a Zeiss Lightsheet Z1 using a W Plan-Apochromat 20x/1.0 UV-VIS objective simultaneous dual camera acquisition of autofluorescence under 488nm and 561nm excitation.</p> <p>Original voxel sizes are (x,y,z) = (0.326, 0.326, 2.0) microns</p> <p>Data provided here are tiles fused using BigStitcher, downsampled to the following voxel sizes</p> <p>- Moss_Fused_Downsampled.tif: (4, 4, 24) microns<br> - Moss_Fused_Downsampled_Anisotropy_2.tif: (12, 12, 24) microns</p> <p>Sample acquisition, sample preparation, imaging and stitching performed by Olivier Burri, at the Ecole Polytechnique Fédérale de Lausanne BioImaging and Optics Platform (EPFL - SV - PTECH - PTBIOP)</p>
Diachronic and diatopic word embeddings from newspapers digitised by the British Library (1830-1889): North and South England
<p>Diachronic word embeddings (decade-level) trained with Word2Vec (via Gensim) on different geographic subcorpora of the Heritage Made Digital British and the Living with Machines historical newspaper collections:</p> <p>- North England (north.zip)</p> <p>- South England (south.zip)</p> <p>At the moment, for each subcorpus, Word2Vec models are available for each decade in the period 1830-1889. More models are on the way for the following:</p> <p>- each decade in the periods 1780-1829 and 1890-1920 for both North and South England.</p> <p>- diachronic models for the following regions: Scotland, Wales, and Midlands.</p> <p>The models were trained using the following parameters:</p> <pre><code>sg = True min_count = 1 window = 5 vector_size = 200 epochs = 5</code></pre> <p>Like the embeddings in <a href="https://doi.org/10.5281/zenodo.7181682">this repository</a>, the model for each decade was aligned to the most recent one with Orthogonal Procrustes.</p> <p>See related GitHub repository for the full documentation: <a href="https://github.com/Living-with-machines/DiachronicEmb-BigHistData">https://github.com/Living-with-machines/DiachronicEmb-BigHistData</a>.</p> <p>Project website (Living with Machines): <a href="https://livingwithmachines.ac.uk/">https://livingwithmachines.ac.uk/</a></p> <p>Data related to: Nilo Pedrazzini & Barbara McGillivray, <em>Diachronic and diatopic word embeddings from British historical newspapers</em>, presented at AIUCD (Convegno dell’Associazione per l’Informatica Umanistica e la Cultura Digitale) in Siena (Italy), June 2023.</p>
Joint embedding of vertebrate brain single-cell RNA-Seq using sequence or structure
<p>Embeddings of single-cell RNA-Seq data from three adult vertebrate brain datasets into Orthogroup feature space or Structural cluster feature space. Orthogroups were generated using OrthoFinder v5.5.0; Structural clusters were assigned by using FoldSeek to cluster AlphaFold-v4 structural predictions.<br> <br> The three datasets used as the basis for these embeddings were:</p> <ul> <li>sample <a href="https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSM3768152">"Brain8"</a> from the <a href="https://www.frontiersin.org/articles/10.3389/fcell.2021.743421/full">Jiang et al. 2021</a> zebrafish cell atlas (files beginning with GSM3768152)</li> <li>sample <a href="https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSM2906405">"Brain1"</a> from the <a href="https://www.sciencedirect.com/science/article/pii/S0092867418301168#sec4">Han et al. 2018</a> mouse cell atlas (files beginning with GSM2906405)</li> <li>sample <a href="https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSM6214268">"Xenopus_brain_COL65"</a> from the <a href="https://www.nature.com/articles/s41467-022-31949-2">Liao et al. 2022</a> Xenopus laevis adult cell atlas (files beginning with GSM6214268)</li> </ul> <p>For each dataset, we also generated a standardized cell type annotation file based on the author's originally provided cell type annotation data. The first column is the cell barcode for that species and the second column is the original study's cell type annotation for that cell.</p> <p>For the Xenopus brain data, we removed around ~18k cells that were not annotated in the original data to simplify data analyses - these are reflected in the files with the "subsampled" suffix. Subsampled versions of the data are also available for the joint embedding space (prefixed with "DrerMmusXlae").</p> <p>For the final datasets used in our analyses, we also provide features x cell matrices as .h5ad files for smaller file sizes and faster loading using Scanpy. </p> <p>For visualizing our UMAP plots of our top200 embedding space, we provide ".tsv" files with a variety of metrics and the x and y positions of each cell in the UMAP. See "DrerMmusXlae_adultbrain_FoldSeek_plotlydata.tsv" and "DrerMmusXlae_adultbrain_OrthoFinder_plotlydata.tsv"</p> <p>These data are part of the Arcadia Science Pub titled <a href="https://doi.org/10.57844/arcadia-vw5e-2670">"Comparing gene expression across species based on protein structure instead of sequence"</a>.</p>
Deconvolved STED nanoscopy images of the nuclear phosphatidylinositol 4,5-bisphosphate and nuclear speckle marker SON together with deconvolved confocal images of DAPI stained nuclei in human formalin-fixed paraffin-embedded skin warts sections
<p>The collection and analysis of formalin-fixed paraffin-embedded (FFPE) human skin sections was approved by the local ethics-committee at the Department of Pathology, University of Cologne, Germany. Written informed consentwas obtained from all patients in accordance with the Declaration of Helsinki. For biopsy materials from archival paraffin blocks of human skin, an informed consent was obtained from all the subjects and ethical approval obtained from the Ethics Committee at the University of Cologne. Surgically removed human FFPE skin biopsies were sectioned into 4 µm sections. Sections were dewaxed, and indirectly immunofluorescently labeled against nuclear phosphatidylinositol 4,5-bisphosphate (nPI(4,5)P2) using 5 µg/mL rabbit primary polyclonal antibody (Echelon Biosciences Inc. Z-A045, clone 2C11). The primary antibody against nPI(4,5)P2 was recognized by the goat secondary antibody conjugated with Abberrior Star 635P (Abberior 2-0002-007-5). Sections were indirectly immunofluorescently labeled against nuclear speckle marker SON using 1 µg/mL rabbit primary polyclonal antibody (Abcam ab121759). The primary antibody against SON was recognized by the goat secondary antibody conjugated with Abberrior Star 580 (Abberrior ST580-1002). Sections were co-stained by DAPI 1:1000 in PBS for 5 min.</p> <p>Imaging of nPI(4,5)P2-635P channel was performed on Leica TCS SP8 STED 3x inverted DMi8 microscope with pulsed white light laser 470-640 nm 1.5 mW and 775 nm pulse STED laser >1.5 W controlled by Leica Application Suite X software and equipped with HC PL APO CS2 100x/1.40 OIL objective used with Leica Type F immersion oil n=1.518. Unidirectional xyz scanning speed was 400 Hz, line accumulation 8. Pixel size was 20 nm in X and Y. Channel settings: 7% 633 nm laser; 775 Notch filter; 50% 775 nm STED laser; 30% 3D STED; HyD 639-698 nm, photon-counting mode, gain 100, gating 0.3-10 ns. Imaging of SON-580 channel was performed on Leica TCS SP8 STED 3x inverted DMi8 microscope with pulsed white light laser 470-640 nm 1.5 mW and 775 nm pulse STED laser >1.5 W controlled by Leica Application Suite X software and equipped with HC PL APO CS2 100x/1.40 OIL objective used with Leica Type F immersion oil n=1.518. Unidirectional xyz scanning speed was 400 Hz, line accumulation 8. Pixel size was 20 nm in X and Y. Channel settings: 10% 585 nm laser; 775 Notch filter; 80% 775 nm STED laser, 30% 3D STED; Hybrid detector (HyD) 589-616 nm, photon-counting mode, gain 100, gating 0.4-10 ns.</p> <p>Z-stacks of STED images were deconvolved using Huygens Professional 22.10 software (Scientific Imaging B.V.). Data sets were processed using Workflow Processor. The workflow consisted of selecting images, setting up the microscopy and deconvolution parameters and saving deconvolved images as 8-bit TIFF single files for individual channels (which were later used for the quantitative analyses; see below). Microscopy parameters were optimized and set as follows. Sampling intervals were ≤20 nm in X and Y and ≤20 nm in Z. Numerical aperture was 1.4; refractive indexes of the lens immersion oil was 1.518 and of the embedding media 1.458; objective quality was good, coverslip position was 0 µm and imaging direction was downward. For nPI(4,5)P2-635P STED channel the backprojected pinhole was 216 nm; excitation (ex.) and emission (em.) wavelengths (λ) were 633 and 651 nm, resp., ex. fill factor 2. STED depletion mode was pulsed, saturation factor 25, STED λ = 775, STED immunity factor 10 and STED 3X was 30%. Classic MLE algorithm with stabilization of Z-slices was used and signal-to-noise ratio was 5.1. For SON-580 STED channel the backprojected pinhole was 195 nm; excitation (ex.) and emission (em.) wavelengths (λ) were 585 and 602 nm, resp., ex. fill factor 2. STED depletion mode was pulsed, saturation factor 20, STED λ = 775, STED immunity factor 10 and STED 3X was 30%. Classic MLE algorithm with stabilization of Z-slices was used and signal-to-noise ratio was 4.</p>
Datasets for the article "Embedded Digital Phase Noise Analyzer for Optical Frequency Metrology"
<p>These files contain the datasets related to the figures presented in the article Donadello et al. (2023) "Embedded Digital Phase Noise Analyzer for Optical Frequency Metrology", IEEE Transactions on Instrumentation and Measurement, 72, pp. 1–12, Article Sequence Number: 2005412. Available at: https://doi.org/10.1109/TIM.2023.3288255.<br> Each dataset is related to the respective figure number in the article, as indicated in the filename (e.g. "data_fig3a.csv" contains the data used to produce Fig. 3, subplot a).</p> <p>All the files are formatted in the comma-separated format.</p> <p>Description:<br> "data_fig3a.csv": coefficients for IQ demodulation and filtering, with f_int=200kHz.<br> "data_fig3b.csv": frequency response of demodulation filtering, with f_int=200kHz and f_int=10kHz.<br> "data_fig4a.csv": power spectral density (PSD) of frequency signals, at different signal amplitudes.<br> "data_fig4b.csv": overlapping Allan deviation of frequency signals, with either OCXO and maser references.<br> "data_fig4c.csv": measured signal frequency and amplitude as a function of nominal frequency deviation.<br> "data_fig4d.csv": measured signal amplitude as a function of nominal amplitude.<br> "data_fig5.csv": example signals related to the synchronized acquisition of RF inputs.<br> "data_fig7a.csv": synchronized time series of frequency deviation acquired over a fiber link in the self-heterodyne interference scheme.<br> "data_fig7b.csv": PSD of the frequency signals acquired over a fiber link in the self-heterodyne interference scheme.<br> "data_fig8a-without-corr.csv": synchronized time series of frequency deviation acquired over a fiber link in the heterodyne interference scheme without frequency drift correction.<br> "data_fig8a-with-corr.csv": synchronized time series of frequency deviation acquired over a fiber link in the heterodyne interference scheme with frequency drift correction.<br> "data_fig8b.csv": PSD of the frequency signals acquired over a fiber link in the heterodyne interference scheme.</p>
Quantized versions of fastText embeddings for 158 languages
<p>These are compressed (quantized) versions of word embeddings for 158 languages originally available from https://fasttext.cc/docs/en/crawl-vectors.html . Following steps were performed to reduce file size:</p> <ol> <li>the output matrix is discarded</li> <li>the quantization is performed with parameters "-qnorm -dsub 1"</li> </ol> <p>The average cosine similarity between original and quantized vectors for frequent words is 0.99. The file size is 4-6 times smaller.</p>
FastText Spanish Medical Embeddings
<p>[Plan TL/medicine/word embeddings] Word embeddings generated from Spanish corpora that include: (a) the full-text in Spanish available in SciELO.org (until December/2018), (b) all articles from the following Wikipedia categories: Pharmacology, Pharmacy, Medicine and Biology (during December/2018) and (c) the concatenation of the previous two corpora.</p> <p>We used fastText to train the word embeddings.</p> <p>For more information, we refer to the corresponding article: <a href="https://www.aclweb.org/anthology/W19-1916/">https://www.aclweb.org/anthology/W19-1916/</a></p> <p>Copyright (c) 2021 Secretaría de Estado de Digitalización e Inteligencia Artificial</p>
Luxembourgish word embedding (User comments from RTL.lu)
<p>This dataset is a word embedding model trained on Luxembourgish user comments from the media platform RTL.lu. It contains data from roughly 544k Luxembourgish texts published between December 2008 and December 2018. See the documentation file for detailed info.</p>
Reversible H2 oxidation and evolution by hydrogenase embedded in a redox polymer film
<p>This data set contains the source data of the nature catalysis manuscript NATCATAL-20033655A.</p> <p>All files are named according to the respective figure numbers of the manuscript. </p>
Reversible H2 oxidation and evolution by hydrogenase embedded in a redox polymer film
<p>This Upload contains the source data sets of all measurements shown in the Nature Catalysis manuscript NATCATAL-20033655A.</p> <p>All folder and file names correspond to the respective figure numbers of the manuscript.</p>
Bilingual English-German word embedding models for scientific text
<p>This data set contains three word embedding models, constructed from the same training corpus of English and German parallel scientific texts (abstracts and research project descriptions). All text was pre-processed by language-specific stemming with the Porter stemming algorithm, removing numbers, and lower-casing.</p> <p>The first model is a 1000-dimensional Latent Semantic Analysis model, constructed from concatenating the English and German texts. The input data was a m×n (297,852×923,864) document-term matrix of tf-idf weights. This was processed with truncated SVD. There are two files, the word vectors in file lsa_1000_Vmat.csv (the V* term by latent factors matrix of right singular values) and the dimension weights in lsa_1000_d_weights.csv (the 1000 values of the diagonal of the <span class="math-tex">\(\Sigma\)</span> matrix.</p> <p>lsa_1000_Vmat.csv has two fields, the term and its vector representation in LSA space, separated by a "|" character. The structure looks like this:</p> <p>tarifplural|{5.00599733151825e-08,-1.43071379136936e-08,8.32862290483082e-08,-6.08010721687266e-08,1.15831140150142e-07,-2.46470313387358e-08,3.43215595753282e-07,6.24301666802575e-07,-2.62907158945831e-07,-1.04120313981517e-07,4.5864574355164e-07,-2.31799632277312e-07,8.37354377858843e-07,8.22507467711628e-07,4.07585381069368e-07,-4.26358988941922e-08,-8.38652991154651e-07,1.98091851171759e-07,-3.94768548759816e-08,-4.28802181962385e-07, ...}</p> <p>The other two models are a basic Random Indexing and a Reflective Random Indexing model, contained in same file, RI_training.csv. Both models have 1000 dimensions. The data structure is as follows.</p> <ul> <li>language: either "en" (English) or "de" (German), the language of the term</li> <li>term: the term as a character string</li> <li>term_collection_count: integer, number of times the term occurred in the training data</li> <li>c_vector: vector of 1000 reals, RI context vector of the term. formatted like this: "{0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0.12309149,0,0,-0.12309149,0,0,0,0,0,0,0,0,0,0,0,0,0,0, ...}"</li> <li>n_docs: integer, number of different documents which contained the term</li> <li>c_vector_o2: vector of 1000 reals, RRI context vector of the term, formatted like c_vector above</li> </ul> <p>1,034,860 rows.</p> <p>All files are aggressively compressed with GNU gzip and will require much more disk space when uncompressed. Note the special formatting of the vector numeric variables, which are different for the two models.</p>
The development of the Leuven Embedded Figures Test (L-EFT)
<p>The Embedded Figures task has a long history as an important clinical and psychological test. There is however some ambiguity as to what exactly the test measures. This ambiguity is brought into clear focus by the fact that some researchers use the test as a measure of a local or global perceptual bias while others regard the test as a good measure of a much broader cognitive capacity related to intelligence or executive function. Given the importance of this test, particularly in clinical domains such as Autism, we have set out to develop a new version of the embedded figures test that more systematically manipulates the perceptual factors that contribute to the effective embedding of a target in a complex context. The result from two experiments will be presented, in which a range of factors, including continuity, complexity, closure and symmetry are revealed as potentially important. Based on these two experiments, a new set of stimuli will be presented which will form the basis of our new version of the Embedded Figures Test, which we plan to launch as an online test using the format of the Leuven Perceptual Organization Screening Test (L-POST). By more systematically manipulating the perceptual factors that contribute to effective embedding, we hope to offer a much more sensitive test, and a test that is better able to differentiate between genuine perceptual, as appose to executive, contributions to performance on this test.</p>
Supporting materials for manuscript entitled "Investigation of guided wave propagation in pipes fully- and partially-embedded in concrete"
<p>This set contains data in support of some of the figures appearing in an open access manuscript entitled 'Investigation of guided wave propagation in pipes fully- and partially-embedded in concrete,' published at the Journal of the Acoustical Society of America (DOI:10.1121/1.4972118) by the authors.</p> <p>The data set contains the numerical output from finite element (FE) modelling, Semi-analytical FE (SAFE) modelling, simulations using the Disperse software, and experimental measurements of guided wave transmission loss in full-scale laboratory tests. The set contains three separate data files. Details on the specific data is provided within an extra file ('readme' file).</p>
Research data supporting "Enzyme Prodrug Therapy Engineered into Electrospun Fibers with Embedded Liposomes for Controlled, Localized Synthesis of Therapeutics"
<p>Research data supporting the publication: Chandrawati R. et al., 2017, Enzyme Prodrug Therapy Engineered into Electrospun Fibers with Embedded Liposomes for Controlled, Localized Synthesis of Therapeutics, Advanced Healthcare Materials. DOI: 10.1002/adhm.201700385</p>
Reference model and embedding for human kidney endothelial cell mapping
<p>Reference model and embedding for human kidney endothelial cell mapping</p><p>The reference model serves as a basis for the mapping of new data to the HLCA using scArches (Lotfollahi et al., https://doi.org/10.1038/s41587-021-01001-7). </p>
Thai Word Embeddings (word2vec) Trained on Oscar Corpus
<p>A large Thai word2vec model trained on Oscar corpus and tokenized and normalized with PyThaiNLP. The model can be loaded using gensim, it is saved in binary format. </p>
Aquaporins Embedded in a Cell Membrane with O2 Diffusion
<p>This is an illustration of an aquaporin embedded in a cell membrane with O2 diffusion as well. </p>
Early Irish Analogy Dataset for Word Embedding Evaluation
<p>An embedding evaluation dataset for Early Irish described in the paper "<a href="https://aclanthology.org/2023.insights-1.10.pdf">Do not Trust the Experts: How the Lack of Standard Complicates <span>NLP</span> for Historical <span>I</span>rish</a>".</p> <p>Traditionally, analogy datasets are based on pairwise semantic proportion, and therefore every question has a single correct answer. Given the high level of variation in historical languages, such a strict definition of a correct answer seems unjustified. Therefore, Early Irish Analogy Dataset follows the <a href="https://vecto.space/projects/BATS/">Bigger Analogy Test Set (BATS)</a> and provides several correct answers to each analogy question. </p> <p>Morphological and spelling variation data are extracted from the <a href="https://dil.ie/">eDIL</a>, a historical dictionary of medieval Irish. Unlike BATS, no distinction is made between inflection types due to eDIL's structure. The raw data amounted to 2,370 spelling variation and 9,690 morphological variation questions, from which 150 examples were randomly selected for each of the subsets to be comparable in size with the synonym and antonym subsets. The synonym and antonym subsets are translations of the correspondent BATS parts obtained by reverse-searching the eDIL and proofread by four expert evaluators. The dataset includes 98 entries in the synonym subset and 109 entries in the antonym subset, upon which three or more experts agreed.</p>
Gene embeddings used in GenePT: A Simple But Hard-to-Beat Foundation Model for Genes and Cells Built From ChatGPT
<p>These are the pulled NCBI (and UniProt, when applicable) summaries of genes, as well as the corresponding OpenAI text embeddings (text-embedding-ada-002 and text-embedding-3-large) computed on the summaries. See methods details in Chen and Zou (2024+).</p> <p>The unzipped folder contains four different files: </p> <ol> <li>NCBI_summary_of_genes.json (NCBI gene card summary of human genes)</li> <li>NCBI_UniProt_summary_of_genes.json (NCBI gene card and UniProt protein (when applicable) summary of human genes)</li> <li>GenePT_gene_embedding_ada_text.pickle (a dictionary of numpy array where gene names (upper case) are keys and text-embedding-ada-002 embeddings of the summary in 1. are the values)</li> <li>GenePT_gene_protein_embedding_model_3_text.pickle (a dictionary of numpy array where gene names (upper case) are keys and text-embedding-3-large embeddings of the summary in 1. are the values)</li> </ol> <p>Reference:</p> <p>Chen YT, Zou J. (2024+) GenePT: A Simple But Effective Foundation Model for Genes and Cells Built From ChatGPT. bioRxiv preprint: <a href="https://www.biorxiv.org/content/10.1101/2023.10.16.562533v1">https://www.biorxiv.org/content/10.1101/2023.10.16.562533v1</a>.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.