Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
5,526
datasets available to search
ShareScore release 0.7.1
Dataset results
5,526 results for “information”
Revised database of the Soil Information System of Latin America and the Caribbean, SISLAC
<p>This dataset contains the revised version of the SISLAC database in three formats: comma-separated values (.csv), microsoft access (.mdb) and PostGIS database (.backup). This database was reviewed and the inconsistencies found in the profiles and in the description of their horizons were corrected. Consists of two tables, one for the description of the profiles and the other with the description of the horizons and their properties. The key field between both tables is the profile identifier, column <strong><em>profile_id</em></strong>.</p>
Supporting data for "Estimating animal density for a community of species using information obtained only from camera-traps"
<p>Data underlying a paper published in Methods in Ecology and Evolution (<a href="https://doi.org/10.1111/2041-210X.13930">https://doi.org/10.1111/2041-210X.13930</a>).</p> <p>These data are suitable for estimating animal density using the Random Encounter Model and include: i) detection counts for 35 species across 510 camera-trap locations; ii) movement speeds (estimated by tracking animal movements in camera-trap image sequences), iii) activity times (filtered so that records of the same species at the same location are > 60 minutes apart), and iv) measurements of the angular and radial distance from camera-traps for animals that were detected.</p>
Dataset for "Too much information: CDCL solvers need to forget and perform restarts"
<p>This repository contains all generated data and evaluations of the paper "<strong>Too much information: CDCL solvers need to forget and perform restarts</strong>" by Tom Krüger, Jan-Hendrik Lorenz, and Florian Wörz.</p> <p>In particular, this collection contains the scripts for obtaining the sets <span class="math-tex">\(\mathbb{L}\)</span> (<em>cores</em>) and reconstructing our sampled sets <span class="math-tex">\(L\)</span> (<em>ext_bitstrings</em>). Furthermore, all data obtained by calling <span class="math-tex">\(\mathrm{CDCLSolver}(\mathscr{F} \cup L)\)</span> can be found. Additionally, we included visual and statistical evaluations used in this paper.</p>
Project Information Model resulting from the on-site survey
<p>The On-site Analysis and Verification Service (ODAVS) is a service developed within the Encore project. It allows users to check any constructability issues regarding renovation projects of residential buildings by means of surveys facilitated by a mixed reality tool. </p> <p>This dataset includes two IFC files of the renovation studies developed by Univpm and JEA (both partners of Encore project) at JEA experimental building in Caceres, and assessed on-site on 2021 December 14th and 15th. The dataset also includes an XML file containing the list of 33 URLS pointing to audio files previously published on a different Zenodo dataset [1]. Note that the access to the Zenodo dataset [1] is restricted. The interested users must ask the dataset authors for permission to access the dataset itself. In the XML file, for each comment, the GUID of the IFC object referred by the audio comment itself is given.</p> <p>References</p> <p>[1] "Pictures and Videos Collected During ODAVS Activity", by Carbonari A. and Vaccarini M., DOI 10.5281/zenodo.6531860, URL: https://doi.org/10.5281/zenodo.6531860</p> <p> </p>
Alignment used in "A phylogenomically informed five-order system for the closest relatives of land plants"
<p>Alignment that served as the basis for the phylogenomic analyses presented in "A phylogenomically informed five-order system for the closest relatives of land plants" — preprint on bioRxiv doi: https://doi.org/10.1101/2022.07.06.499032</p>
Supporting Information for Disclosing Spin-Polarized Bonds on Isolable Molecules
<p>The file corresponds to the Bachelor Thesis of Ms. Elena Paulus. It contains the xyz coordinates of all optimized structures and their corresponding electronic energy in Hartree.</p>
Deep Deconvolution of Object Information Modulated by a Refractive Lens Using Lucy-Richardson-Rosen Algorithm
<p>A refractive lens is one of the simplest, cost-effective and easily available imaging elements. With a spatially incoherent illumination, a refractive lens can faithfully map every object point to an image point in the sensor plane, when the object and image distances satisfy the imaging conditions. However, static imaging is limited to the depth of focus, beyond which the point-to-point mapping can be only obtained by changing either the location of the lens or the imaging sensor. In this study, the depth of focus of a refractive lens in static mode has been expanded using a recently developed computational reconstruction method, Lucy-Richardson-Rosen algorithm (LRRA). The technique consists of three steps. In this first step, the point spread functions (PSFs) were recorded along different depths and stored in the computer as PSF library. In the next step, the object intensity distribution was recorded. The LRRA was then applied to deconvolve the object information from the recorded intensity distributions in the final step. The results of LRRA were compared against two well-known reconstruction methods namely Lucy-Richardson algorithm and non-linear reconstruction. The data corresponding to experimental analysis is given in the manuscript. (Preprints Link:). The theoretical simulation data is given here.</p>
Electronic Supporting Information for Catalytic Ammonia Oxidation to Dinitrogen by a Nickel Complex
<p>The dataset provides electronic supporting information in the format of XYZ molecular files, formatted Gaussian checkpoint files, and cube files for atomic spin density distributions for selected complexes obtained while investigating the catalytic mechanism of ammonia oxidation to dinitrogen using a N-heterocyclic carbene containing nickelocene complex.</p> <p>The level of theory used for all calculations is omega-B97xD with def2TZVP basis set. All calculations were performed using the Gaussian16 suite of programmes.</p> <p><strong>Model Set 1</strong> contains the metal free compounds and were used to calculate the overall thermodynamics of the ammonia oxidation reaction.</p> <p><strong>Model Set 2</strong> corresponds to the most truncated, in vacuo optimized structures.</p> <p><strong>Model Set 3</strong> comprises from non-truncated, realistic structures embedded in polarizable continuum model of benzene.</p> <p> </p>
Italian Soil Information System
<p>The application offers both soil and climatic information, related to 1:500,000 scale geography. The soil information systems of Italy is made up of a hierarchy of geographical layers, which includes soil regions, aimed at correlating the soils of Italy with the other European countries, soil systems, for the correlation of soil at the national level, and soil sub-systems, for the regional level. The databases provides an inventory of Italian soilscapes at two reference scale, Soil region (1:5,000,000) and Soil systems (1:500,000). Relation between entities (soil-typological-unit, soil derived profile, benchmark soil profile) and geography (soil mapping units), was also stored. Geographical layers (administrative boundaries, regionand province main towns, soil regions, and rasters of climatic variables) and correlation entities was also stored. Soil regions are a regionally restricted part of the soil cover characterized by a typical climate and parent-material association. Soil systems illustrate main Italian soilscapes and are composed of homogeneous areas as for physiography, lithology, river drainage network, and land cover. Each cartographic unit was described as combination of: i) a major landform, established according to main morphological process, slope, hypsometry, kind and degree of drainage, ii) two lithological types, iii) three land cover attributes. Maximum seven land components were recognized in each land system. A land component was a specific combination of morphology, lithology and land cover that was not delineated. 1412 soil observations have been stored on the Ms Access database mostly complete with 4284 analyzed soil horizons for the most common analytical parameters (pH in water; Carbon (C) - organic; Carbonate (CO3--) - Total; Clay, Sand, and Silt fraction; Available water capacity - estimated volumetric) and 2039 photos. Several climatic variables, relevant for soil evaluation and management, have been collected and stored. Climatical maps have been produced by spatialization of longterm statistics related to meteorological stations. The map of precipitation (AnnualRainfall) was obtained by ordinary kriging of 1,613 stations completed of long termannual average values. Mean annual air temperature was obtained by ordinary kriging of 944 stations of long-term average annual air temperature. The humidity index was obtained with a simple calculation (Annual Rainfall/Mean Annual Air Temperature). Average temperature of the soil to 50 cm was calculated on the basis of the average air temperature of the long term and field capacity of the soil at the same depth, accordingto the equation: Tm = soil to 0.5 m tm air + (field capacity to 0.5 m - 20.7) / 7.9 (Constantini et al., 1999). The mean annual soil temperatures map was obtained byordinary kriging of 6,660 data of average soil temperature at 50 cm.</p>
The institutional perspective on informal housing
<p><strong>Coordinador del Seminario:</strong> Carlos A. Navarrete Ulloa.<br> <strong>Expositor</strong>: Luis Adolfo Ortega Granados</p> <p><strong>Comité Ejecutivo PRONACE-Vivienda</strong><br> Fernando Córdova Canela, Centro Universitario de Arte, Arquitectura y Diseño, Universidad de Guadalajara (UdeG). Francisco Javier Porras Sánchez, Instituto de Investigaciones Dr. José María Luis Mora.<br> Gabriel Castañeda Nolasco, Universidad Autónoma de Chiapas (UNACH). Carlos A. Navarrete Ulloa, Centro Universitario de Tonalá, (UdeG).</p> <p>Exposición realizada en el marco del PRONACE Vivienda en el cual se comenta la lectura:</p> <p>Dekel. T. (2020) The Institutional Perspective on Informal Housing. Habitat International, 106, 102287 https://doi.org/10.1016/j.habitatint.2020.102287</p>
El sector informal en la Ciudad de México Caso de estudio de la Delegación Iztapalapa
<p><strong>Coordinador del Seminario:</strong> Carlos A. Navarrete Ulloa.<br> <strong>Expositor</strong>: Luis Adolfo Ortega Granados</p> <p><strong>Comité Ejecutivo PRONACE-Vivienda</strong><br> Fernando Córdova Canela, Centro Universitario de Arte, Arquitectura y Diseño, Universidad de Guadalajara (UdeG).<br> Francisco Javier Porras Sánchez, Instituto de Investigaciones Dr. José María Luis Mora.<br> Gabriel Castañeda Nolasco, Universidad Autónoma de Chiapas (UNACH).<br> Carlos A. Navarrete Ulloa, Centro Universitario de Tonalá, (UdeG).</p> <p>Exposición realizada en el marco del PRONACE Vivienda en el cual se comenta la lectura:<br> Martínez-Luis, D., Pérez-Fernández, A., Pat-Fernández, L. A., Caamal-Cauich, I., Franco-Gutiérrez, M. J., & García-Cabrera, L. G. (2019). El sector informal en la Ciudad de México. Caso de estudio de la Delegación Iztapalapa. Estudios sociales. Revista de alimentación contemporánea y desarrollo regional, 29(53). https://www.ciad.mx/estudiosociales/index.php/es/article/view/725</p>
Applying Sensor Fusion to Augment Hyperspectral Data with Depth Information
<p>The research data for the paper "Applying Sensor Fusion to Augment Hyperspectral Data with Depth Information"<br> <br> Data in the archive "hyperdepth.tar.gz" includes:</p> <p><br> <strong>calibration_images/</strong><br> includes preprocessed images for calibrating both cameras</p> <p><strong>pointclouds/</strong><br> Includes individual hyperspectral point clouds for each view (front, rightmost, right, leftmost, left with postfixes correspondingly: edesta, oikea, oikea2, vasen, vasen2)<br> <br> <strong>raw_images/</strong><br> Two directories "day5" and "day6" which include the raw hyperspectral images and kinect images<br> <br> Some extra images are included which were not used in the research paper.</p> <p> </p> <p><strong>2022-03-11_112336_stereocalibration.json</strong> includes calibration results (mainly the intrinsic camera matrix and extrinsic parameters) for the setup.</p>
GERDAT010 Dataset for literature search linked to publication "Information needs of older patients newly diagnosed with cancer"
<p>Dataset of the literature search belonging to the publication "Information needs of older patients newly diagnosed with cancer"</p>
Shool drop-out ut in Brazil: rates per city and informations about schools
<p>The dataset presented here is a combination of three databases created by INEP (Brazil), and referes to the years of 2014/2015: </p> <p>- Drop-out rates by city,</p> <p>- Questionnaires to principals about their schools,</p> <p>- Questionnaires about school structure.</p> <p>The original databases and dictionaires are avalilable here:</p> <p>http://portal.inep.gov.br/web/guest/indicadores-educacionais</p> <p>http://portal.inep.gov.br/artigo/-/asset_publisher/B4AQV9zFY7Bv/content/divulgados-os-microdados-do-sistema-nacional-de-avaliacao-da-educacao-basica/21206</p> <p> </p> <p> </p>
Oregon Wolfe Barley (Hordeum vulgare) Informative & Spectacular Subset (ISS) vegetative stage growth data
<p>Oregon Wolfe Barley Informative & Spectacular Subset (ISS) was raised at the Ag Alumni Seed Phenotyping Facility (AAPF) at Purdue University (West Lafayette, Indiana, USA) for 42 days. There were 18 genotypes, with two replicates for each genotype (total plants: 36). AAPF is a controlled environment high-throughput phenotyping facility with automated imaging and irrigation systems. A virtual tour of AAPF can be found at <a href="https://ag.purdue.edu/aapf/virtual-tour.html">https://ag.purdue.edu/aapf/virtual-tour.html</a>.</p> <p>Seeds were sown in a 6 L pot with 2.8 L of Profile Porous Ceramic Greens Grade and Berger BM6 each with 10g of Osmocote. Five hundred ml of Turface was laid on top of each pot to avoid effect of algae for RGB data derivation. The growth temperature in the chamber was 72/68 degrees Fahrenheit day/night. Relative humidity was set at 60%. Lighting was 16 h day/8 h night.</p> <p>Plants were imaged with RGB camera from one top and 12 side views three times a week, ranging between 10 days from planting (equivalent to sowing, Dfp) to 42 Dfp. Ground reference data of plant height and tiller count were measured twice a week. </p> <p> </p> <p>RGB imaging data were stored in “OWB_RGB.xlsx”. Datasheet “Information” describes the variables in datasheets for top view, side average view and every side view.</p> <p> </p> <p>Ground reference data for plant height and tiller count were stored in “OWB_ground_reference.xlsx”. Datasheet “Information” describes the variables in datasheet “Data”.</p>
Structural Inheritance in the Eastern Cordillera, NW Argentina: Low‐Temperature Thermochronology of the Cianzo Basin - Supporting Information
<p>Supporting information accompanying the publication "Structural Inheritance in the Eastern Cordillera, NW Argentina: Low‐Temperature Thermochronology of the Cianzo Basin" published in Tectonics. The dataset contains (U-Th-Sm)/He and apatite fission track data from the Cianzo Basin, Jujuy, Argentina, and accompanying figures.</p> <p>Table S1 contains full single-grain results from apatite (AHe) and zircon (ZHe) (U-Th-Sm)/He analyses. Table S2 and S3 contain AFT results including full counting data from apatite fission track (AFT) analyses.</p> <p>Figure S1 supports AHe and ZHe data with plots showing relationships between cooling ages, eU, Ft and ESR. Figures S2–S4 support AFT data with radial plots.</p>
Supporting Information for "New 3D velocity model (mTAB3D) for absolute hypocenter location in southern Iberia and the westernmost Mediterranean"
<p>These files comprise supplementary information for the paper entitled "New 3D velocity model (mTAB3D) for absolute hypocenter location in southern Iberia and the westernmost Mediterranean" (Sánchez-Roldán et al., 2024a)</p> <p>These results were obtained after performing a relocation using the 3D P-wave velocity model mTAB3D (Sánchez-Roldán et al. 2024b).</p> <p>In "Files.zip", we provide the eight files with the absolute locations and the uncertainty parameters (extracted from the 68% confidence ellipse of the PDF’s) obtained after performing the relocation using mIGN1D and mTAB3D. The absolute location files follow this format:</p> <p>origin_time(YYYY-mm-ddTHH:MM:SS) longitude(º) latitude(º) depth(km) magnitude(mbLg)</p> <p>• origin_time: Hypocenter’s origin time after the relocation.</p> <p>• longitude: Hypocenter’s longitude in decimal degrees after the relocation.</p> <p>• latitude: Hypocenter’s latitude in decimal degrees after the relocation.</p> <p>• depth: Hypocenter’s depth in kilometers.</p> <p>• magnitude: Hypocenter’s magnitude (mbLg) computed by the Spanish Seismic Network.</p> <p>The files with the uncertainty values:</p> <p>horizontal_uncertainty(km) vertical_uncertainty(km) rms(s) no_arrivals</p> <p>• horizontal_uncertainty: Obtained after computing the geometrical mean between the horizontal semi-minor and semi-major axes of the 68% confidence ellipse in kilometers.</p> <p>• vertical_uncertainty: Vertical semi-axis of the 68% confidence ellipse.</p> <p>• rms: root-mean-square of residuals at maximum likelihood or expectation hypocenter.</p> <p>• no_arrivals: number of readings used for the absolute location.</p> <p><br>File S1. File_S1.dat: Eastern Betics Shear Zone catalog’s absolute locations with mIGN1D.</p> <p>File S2. File_S2.dat: Eastern Betics Shear Zone catalog’s statistics with mIGN1D.</p> <p>File S3. File_S3.dat: Eastern Betics Shear Zone catalog’s absolute locations with mTAB3D.</p> <p>File S4. File_S4.dat: Eastern Betics Shear Zone catalog’s statistics with mTAB3D.</p> <p>File S5. File_S5.dat: Al Hoceima 2016 catalog’s absolute locations with mIGN1D.</p> <p>File S6. File_S6.dat: Al Hoceima 2016 catalog’s statistics with mIGN1D.</p> <p>File S7. File_S7.dat: Al Hoceima 2016 catalog’s absolute locations with mTAB3D.</p> <p>File S8. File_S8.dat: Al Hoceima 2016 catalog’s statistics with mTAB3D.</p> <p>Additionally, we provide two figures showing the location of those hypocenters (alboran.jpg and ebsz.jpg), which are included as Figures 3 and 5, respectively, in Sánchez-Roldán et al. (2024a).</p> <p>References:</p> <p><span>Sánchez-Roldán, J. L.</span>, <span>Álvarez-Gómez, J. A.</span>, <span>Martínez-Díaz, J. J.</span>, <span>Herrero-Barbero, P.</span>, <span>Perea, H.</span>, <span>Cantavella, J. V.</span>, & <span>Lozano, L.</span> (<span>2024a</span>). <span>New 3D velocity model (mTAB3D) for absolute hypocenter location in southern Iberia and the westernmost mediterranean</span>. <em>Earth and Space Science</em>, <span>11</span>, e2023EA00299. <a href="https://doi.org/10.1029/2023EA002993">https://doi.org/10.1029/2023EA002993</a></p> <p>Sánchez-Roldán, J. L., Álvarez-Gómez, J. A., Martínez-Díaz, J. J., Herrero-Barbero, P., Perea, H., Lozano, L., & Cantavella, J. V. (2024b). MTAB3D: a 3-D velocity model for absolute hypocenter location in southern Iberia and westernmost Mediterranean. (v1.0) [Data set]. Zenodo. <a href="https://doi.org/10.5281/zenodo.7766525" target="_blank" rel="noopener">https://doi.org/10.5281/zenodo.7766525</a></p> <div> </div> <p> </p> <p> </p>
Supplementary Information and Data for "Unveiling the 3D Morphology of Epitaxial GaAs/AlGaAs Quantum Dots"
<p>Raw and processed TEM and AFM data for the article <strong><em>Unveiling the 3D Morphology of Epitaxial GaAs/AlGaAs Quantum Dots</em></strong>.</p> <p>Paper: <a href="https://doi.org/10.1021/acs.nanolett.4c02182" target="_blank" rel="noopener">https://doi.org/10.1021/acs.nanolett.4c02182</a></p> <p>Preprint: <a href="https://arxiv.org/abs/2405.16073" target="_blank" rel="noopener">https://arxiv.org/abs/2405.16073</a></p> <p>The TEM data has a PDF information file included with description of the file types and how to open them.</p> <p>The AFM Nanosurf .nid files can be opened, e.g., with <a href="http://gwyddion.net/" target="_blank" rel="noopener">Gwyddion</a>.</p>
Examining LGBTQ+-related Concepts in the Semantic Web: Link Discovery, Concept Drift, Ambiguity, and Multilingual Information Reuse
<div> <h1>Examining LGBTQ+-related Concepts in the Semantic Web</h1> </div> <div> <h2>Introduction</h2> </div> <p>Welcome to the project. We study the links between LGBTQ+ ontologies and structured vocabularies. More specifically, we focus on GSSO, Homosaurus, QLIT, and Wikidata. The code is free for use with the license GPL 3,0. You can resue/extend the code for free as long as you give credits to us in your publication/data. Citation information will be added after the corresponding paper gets accepted. The paper is under submission and will be included soon. </p> <p>If you would like to extend this work, you may want to contact the experts in the acknowledgement before releasing your data/code about legal and ethical issues. The DOI for this version is 10.5281/zenodo.12684870. The latest code can be found at https://github.com/Multilingual-LGBTQIA-Vocabularies/Examing_LGBTQ_Concepts. </p> <p>To reproduce the results or extend our work, you need to take the following steps.</p> <div> <h2>Step 1: Preparing the data</h2> </div> <p>In this project, the following datasets were used:</p> <ul> <li>QLIT: version 1.0</li> <li>Homosaurus: version 3.5 and version 2.3</li> <li>Wikidata: retrieved from the SPARQL Endpoint (<a href="https://query.wikidata.org/sparql" rel="nofollow">https://query.wikidata.org/sparql</a>) and processed between 5th May and 8th May, 2024.</li> <li>GSSO: we used gsso.owl (version 2.0.10) obtained from its Github (<a href="https://github.com/Superraptor/GSSO">https://github.com/Superraptor/GSSO</a>).</li> <li>LCSH was obtained from the official website: <a href="https://id.loc.gov/authorities/subjects.html" rel="nofollow">https://id.loc.gov/authorities/subjects.html</a> on 9th May, 2024. The LCSH data was converted to its HDT format.</li> </ul> <p>Please put the corresponding files in the following folders (and change its names where necessary) to make sure that the Python scripts can find your code.</p> <ul> <li>./data/GSSO/gsso.owl</li> <li>./data/Homosaurus/v2.ttl and ./data/Homosaurus/v3.ttl</li> <li>./data/LCSH/lcsh.hdt (we used its HDT format for fast query and analysis). The original file is also attached: subjects.skosrdf.nt.</li> <li>./data/QLIT/Qlit-v1.ttl</li> </ul> <p>The case of Wikidata is more complicated. The following scripts were used for the retrival of data. These scripts are all in the folder ./data/wikidata/</p> <ul> <li>We used the Wikidata SPARQL endpoint: <a href="https://query.wikidata.org/" rel="nofollow">https://query.wikidata.org/</a></li> </ul> <p>The following relations from Wikidata were used while extracting triples.</p> <ul> <li>Wikidata - GSSO: <a href="http://www.wikidata.org/prop/direct/P9827" rel="nofollow">http://www.wikidata.org/prop/direct/P9827</a></li> <li>Wikidata - Homosaurus 2: <a href="http://www.wikidata.org/prop/direct/P6417" rel="nofollow">http://www.wikidata.org/prop/direct/P6417</a></li> <li>Wikidata - Homosaurus 3: <a href="http://www.wikidata.org/prop/direct/P10192" rel="nofollow">http://www.wikidata.org/prop/direct/P10192</a></li> <li>Wikidata - LCSH: <a href="http://www.wikidata.org/prop/direct/P244" rel="nofollow">http://www.wikidata.org/prop/direct/P244</a></li> </ul> <p>The generated files are:</p> <ul> <li>'wikidata-homosaurus-v2-links.nt'</li> <li>'wikidata-homosaurus-v3-links.nt'</li> <li>'wikidata-gsso-links.nt'</li> <li>'wikidata-qlit-links.nt'</li> <li>'wikidata-lcsh-links-all.nt'</li> </ul> <p>Please note that the case of Wikdiata-LCSH is more complicated: there are so many links that are nothing to do with the entities in our scope. We restrict it to only entities in the scope of this paper. See below for more details.</p> <p>You can find all the scripts in the corresponding folder in the data folder.</p> <p>All the SPARQL queries used can be found in the folder ./SPARQL/</p> <p>Note! For GSSO, the following two mistakes were corrected while preprocessing:</p> <ul> <li><a href="https://www.wikidata.org/wiki/Q1823134" rel="nofollow">https://www.wikidata.org/wiki/Q1823134</a> should not be used as a relation. We have replaced it with <a href="http://www.wikidata.org/prop/direct/P244" rel="nofollow">http://www.wikidata.org/prop/direct/P244</a>.</li> <li>Instead of referring to the page, we refer to the entity. We use <a href="http://www.wikidata.org/entity/" rel="nofollow">http://www.wikidata.org/entity/</a>* instead of <a href="https://www.wikidata.org/wiki/" rel="nofollow">https://www.wikidata.org/wiki/</a>*</li> </ul> <p>The redirection test was conducted on 30th April, 2024, between 6PM and 8PM. The files can be found in the folder of ./data/Homosaurus/redirect/.</p> <div> <h2>Integrating the data</h2> </div> <p>In the folder ./integrated_data/, you can find all the scripts related to the integrated data. Unfortunately, due to the CC-BY-NC-ND license of GSSO and Homosaurus, the integrated data will not be made available. But you can generate it with the instructions above and by using the following scripts.</p> <p>The script ./integrated_data/integrate.py takes advantage of the data generated. It first integrates a list of files of links. Then we go through the links between Wikidata and LCSH. Only those that are in the scope of the study are included.</p> <ul> <li>If your steps are correct and using the same version as we did, you should be able to get four files:</li> <li>a) the integrated file as integrated.nt</li> <li>b) the links that are relevant for this study: wikidata-lcsh-links-selected.nt.</li> <li>c) a plot of the distribution of the size of WCCs</li> <li>d) a mapping of entities and their corresponding ID of WCCs.</li> </ul> <div> <h2>Weakly Connected Components</h2> </div> <p>The weakly connected components (WCCs) were computed for the following three purposes:</p> <p>a) Discovering missing links. See the section below for details.</p> <p>b) The WCCs can be used for manual examination. These are entities that form clusters about related concepts. The intuition is that the larger they are, the more likely there is concept drift/change, ambiguity, and mistakes.</p> <p>c) Multilingual information reuse. Smaller WCCs with exactly one entity from each dataset (e.g. Homosaurus and Wikidata) can then be used to suggest labels for the one with fewer labels for some given languages. See below for more details.</p> <p>As mentioned above, the distribution has been plotted. You can find this plot here: ./integrated_data/frequency.png</p> <p>In the folder ./integrated_data/weakly_connected_components/, you can find all the WCCs and their links.</p> <p>Two examples were given in the folder. The largest WCC about sex, gender, fucking, etc. The other is about BDSM and fetish.</p> <div> <h2>Discovering missing and outdated links</h2> </div> <p>Taking advantage of WCCs, we can further find missing and outdated links. The scripts are in the folder ./discover_missing_links.</p> <p>Three examples were given. The first two is about discovering missing links. The last one is about finding outdated links.</p> <ul> <li> <p>The script ./discover_missing_links/discover_H3_LCSH.py and ./discover_missing_links/discover_QLIT_LCSH.py are scripts that outputs links that could be missing in Homosaurus and QLIT respectively. This was computed by looking at the WCCs. If two entities are both involved in the same WCC, there could be a link between them. The csv files in the same folder are the corresponding links found.</p> </li> <li> <p>The script ./discover_missing_links/find_qlit_outdated_links/ is used to discover the outdated links between QLIT and Homosaurus v3. There was only one link found.</p> </li> <li> <p>The 105 potentially missing links were taken for further review by Swedish-speaking experts from the QLIT team, which showed that 78 (72.38%) suggested links should be included: 38 (36.19%) can be included using skos:exactMatch and another 38 (36.19%) using skos:closeMatch. 28 (26.67%) suggested links are incorrect. The manual annotation are included in the file ./discover_missing_links/Annotated_found_new_links_qlit-lcsh.xlsx.</p> </li> </ul> <div> <h2>Multilingual Information Reuse</h2> </div> <p>You can find two attempts in the folders about the use of GSSO and Wikidata for Homosaurus respectively.</p> <ul> <li>./WCC-based-gsso-multilingual_info_reuse/</li> <li>./WCC-based-wikidata-multilingual_info_reuse/</li> </ul> <p>Additionally, we provide also some code for the reuse of Wikidata multilingual info for QLIT. It's in the folder</p> <ul> <li>./WCC-based-QLIT-info-reuse-from-Wikidata/</li> </ul> <p>They follow very similar steps:</p> <ol> <li> <p>Compute the one-to-one mapping using the WCCs. The script is named compute-one-to-one-mapping.py</p> </li> <li> <p>Extract the multilingual labels from sources. The corresponding file is extract_multilingual_labels_from_one_to_one_mappings.py</p> </li> <li> <p>Provide the extracted multilingual as suggestions for targeting entities. The name of the corresponding files are like "*suggesting-labels.py", where the * is replaced by the actual source/target.</p> </li> </ol> <p>For GSSO, we use the following relations:</p> <ul> <li><a href="http://www.w3.org/2000/01/rdf-schema#label" rel="nofollow">http://www.w3.org/2000/01/rdf-schema#label</a></li> <li><a href="http://www.geneontology.org/formats/oboInOwl#hasRelatedSynonym" rel="nofollow">http://www.geneontology.org/formats/oboInOwl#hasRelatedSynonym</a></li> <li><a href="http://www.geneontology.org/formats/oboInOwl#hasSynonym" rel="nofollow">http://www.geneontology.org/formats/oboInOwl#hasSynonym</a></li> <li><a href="http://www.geneontology.org/formats/oboInOwl#hasExactSynonym" rel="nofollow">http://www.geneontology.org/formats/oboInOwl#hasExactSynonym</a></li> <li><a href="http://purl.org/dc/terms/replaces" rel="nofollow">http://purl.org/dc/terms/replaces</a></li> <li><a href="https://www.wikidata.org/wiki/Property:P5191" rel="nofollow">https://www.wikidata.org/wiki/Property:P5191</a></li> <li><a href="https://www.wikidata.org/wiki/Property:P1813" rel="nofollow">https://www.wikidata.org/wiki/Property:P1813</a></li> <li><a href="https://schema.org/alternateName" rel="nofollow">https://schema.org/alternateName</a></li> <li><a href="http://www.w3.org/2002/07/owl#annotatedTarget" rel="nofollow">http://www.w3.org/2002/07/owl#annotatedTarget</a></li> </ul> <p>Additioinally, we found the relation to be studied in the future: <a href="http://www.geneontology.org/formats/oboInOwl#hasNarrowSynonym" rel="nofollow">http://www.geneontology.org/formats/oboInOwl#hasNarrowSynonym</a></p> <p>For Wikidata, there are only two:</p> <ul> <li><a href="http://www.w3.org/2000/01/rdf-schema#label" rel="nofollow">http://www.w3.org/2000/01/rdf-schema#label</a></li> <li><a href="http://www.w3.org/2004/02/skos/core#altLabel" rel="nofollow">http://www.w3.org/2004/02/skos/core#altLabel</a></li> </ul> <div> <h2>Additional analysis</h2> </div> <p>Additionally, we perform an analysis using only redirection and replacement for GSSO and Homosaurus. The scripts are in the folder ./additional_test_gsso_multilingual_info_reuse. We consider also Homosaurus v2. This additional analysis shows the following:</p> <ul> <li> <p>For the Turkish language, in total there are 103 triples about labels about 23 entities. The average suggested labels per entity is 3.0.</p> </li> <li> <p>For the Spanish language, in total there are 205 triples about labels about 43 entities. The average suggested labels per entity is 2.12.</p> </li> <li> <p>For the French language, in total there are 277 triples about labels about 47 entities. The average suggested labels per entity is 2.19.</p> </li> <li> <p>For the Danish language, in total there are 115 triples about labels about 47 entities. The average suggested labels per entity is 2.70.</p> </li> </ul> <p>Some analysis about the replacement relations of Homosaurus is in the folder ./data/Homosaurus/replace_relations_homosaurus/.</p> <p>Finally, some additional analysis is included in the folder ./analysis_integrated_graph. Currently, there is only one that is about outdated entities in Homosaurus v3. Some more analysis will be added in the future.</p> <div> <h2>Acknowledgement</h2> </div> <p>The authors appreciate the help of the following researchers:</p> <ul> <li>Siska Humlesjö, QLIT, Göteborgs Universitet (<a href="mailto:siska.humlesjo@lir.gu.se">siska.humlesjo@lir.gu.se</a>)</li> <li>Olov Kriström, former member of QLIT</li> <li>Jack van der Wel, IHLIA (<a href="mailto:jack@ihlia.nl">jack@ihlia.nl</a>)</li> <li>Clair Kronk, GSSO (<a href="mailto:clair.kronk@mountsinai.org">clair.kronk@mountsinai.org</a>)</li> </ul> <div> <p>If you would like to extend this work, you may want to contact them before releasing your data/code about legal and ethical issues.</p> <h2>Contact</h2> </div> <ul> <li>Shuai Wang, Vrije Universiteit Amsterdam (<a href="mailto:shuai.wang@vu.nl">shuai.wang@vu.nl</a>)</li> <li>Maria Adamidou, Vrije Universiteit Amsterdam (<a href="mailto:m.adamidou@student.vu.nl">m.adamidou@student.vu.nl</a>)</li> </ul> <p> </p> <p>Thank you very much for your interest in our project!</p>
Open database on distributional information on European pollinators
<p>(abstract) This dataset was produced in the framework of the work package 1 (task 1) of the Horizon EU project Safeguard. We aimed to mobilise EU experts and data to compile and make available distributional data for bees, butterflies, moths and hoverflies. This will allow us to assess the magnitude, scale and extent of status and trends in pollinator distributions, diversity, abundance, communities and plant-pollinator networks.</p> <p>(method) Regarding distribution data for bees, UMons have been in contact with 23 bee taxonomists, 52 national champions and 5 museums. To date, we collected 52 bio-geographical databases of European bees from both restricted (i.e. databases shared under ad hoc agreement) and public (i.e. openly accessible databases) sources. Regarding distributional data for hoverflies, the starting point was the recently published in the IUCN Red List of hoverflies. To expand the number of species with precise distributional data on syrphid flies, UNSPMF further contacted taxonomists working with this species group : Gunilla Stahls from Finland; Jeroen van Steenis, Wouter van Steenis and Gerard Pennards from Netherlands; Grigory Popov from Ukraine; Santos Rojo from Spain; Axel Ssymank from Germany; Libor Mazanek from Czech Republic; Daniele Sommaggio from Italy. They provided additional data and conducted validation of the existing data, but also engaged additional experts who provided the data. For the butterflies and the moth, the data was collected by UFZ and come from an original initiative of the scientific expert on those two groups. As the publication of the row data of some databases (e.g. bees from The Netherlands) required the clustering of the spatial records to geographic grid squares (e.g. 10x10 km²), we simplified all the records in the present dataset.</p> <p>(dataset) We consider as a data, a record that includes the following information: the name of the species, the coordinates where the species was collected. Additional information were collected (e.g. collector, determinator, number of the individuals collected, sex, data owner and reference code) but were not displayed in the present dataset. The aggregation of bee databases include 4,837,731 row data for bees, 680,641 row data for hoverflies, 1,209,320 row data for butterflies and 6,862,835 row data for moths.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.