Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
167
datasets available to search
ShareScore release 0.7.1
Dataset results
167 results for “Identifier mapping”
Mapping the Atlantic Ocean i.e. the Gulf of Maine to identify suitable cultivation sites for kelp species
<p>Input source:</p> <ul> <li>Temperature data</li> <li>Depth data</li> <li>Wave data</li> <li>Nutrients data</li> <li>Current data</li> <li>Marine use data</li> </ul> <p><strong>All from other available sources outside the project</strong></p> <p> </p> <p>DATA SET GENERATED:</p> <ul> <li>Environmental data</li> <li>Training/validation data</li> <li>The socioeconomic datasets</li> </ul> <ul> <li>Map of suitable sites</li> <li>Model using GIS</li> </ul>
Raw data for the article "Games on Climate Change: Identifying Development Potentials through Advanced Classification and Game Characteristics Mapping"
<p>Raw data used for the article "Gerber, Andreas, Markus Ulrich, Flurin X. Wäger, Marta Roca-Puigròs, João S.V. Gonçalves, and Patrick Wäger. 2021. "Games on Climate Change: Identifying Development Potentials through Advanced Classification and Game Characteristics Mapping" <em>Sustainability</em> 13, no. 4: 1997. <a href="https://doi.org/10.3390/su13041997">https://doi.org/10.3390/su13041997</a>"</p> <p>The documents include the raw data (both as .csv and .xlsx files with the same content), as well as the publication (.pdf file). The data collection process and the data itself are described in the publication. The data is published as "supplementary material" on the publisher's homepage.</p>
Mapping between zbMATH Open identifiers, DOIs, ORCIDs and arXiv identifiers
<p>The second version of the mapping between zbMATH Open identifiers for <a href="https://www.wikidata.org/w/index.php?title=Property:P1556&oldid=1755821772">authors</a> and <a href="https://www.wikidata.org/w/index.php?title=Property:P894&oldid=1766254659">documents</a> and <a href="https://www.wikidata.org/wiki/Property:P356">DOIs</a> and <a href="https://www.wikidata.org/wiki/Property:P496">ORCIDs</a> in CSV format.</p> <ul> <li>The file authors.csv contains the mapping between zbMATH Open author id and ORCIDs for 38 159 authors.</li> <li>The file documents.csv contains the mapping between zbMATH Open document id and DOI for 2 813 563 documents.</li> </ul> <p>Beginning from this version, we also provide the mapping between <a href="https://www.wikidata.org/w/index.php?title=Property:P894&oldid=1766254659">documents</a> and <a href="https://www.wikidata.org/w/index.php?title=Property:P818&oldid=2151126552">arXiv</a> in CSV format</p> <ul> <li>The file arxiv.csv contains the mapping between zbMATH Open document id and arXiv identifiers for 528 640 documents.</li> </ul> <p>See https://zbmath.org/about/ (section Full Text Links) for a live version of this dataset. That version is more current but less reproducible. Moreover, the dataset here is restricted to documents with a permanent zbMATH Open identifier in the form <code>Zbl d+.d+</code>.</p> <p> </p> <p> </p>
EOL full taxon identifier map
<p>A mapping of taxon identifiers from EOL resources, of the form: node_id, resource_pk, resource_id, page_id, preferred_canonical_for_page</p> <ul> <li>node_id: internal to EOL; useful for some API calls</li> <li>resource_pk: identifier according to the classification provider </li> <li>resource_id: identifies the classification provider (see below) </li> <li>page_id: EOL taxon concept identifier; official, for sharing </li> <li>preferred_canonical_for_page: canonical name preferred by EOL for this taxon concept </li> </ul> <p>commas within entries are "escaped, by, quoting", and quotes-within-quotes are ""double quoted""</p> <p>resource_ids, their names and descriptions, are available at: <a href="https://eol.org/resources.json">https://eol.org/resources.json </a></p> <p>To view one at a time: https://eol.org/resources/[enter ID here] </p> <p>There is also a summary file with resources ids, links, and resource names here: <a href="https://github.com/KatjaSchulz/eolResources">https://github.com/KatjaSchulz/eolResources</a></p> <p>And there is a smaller mapping file just for the major EOL classification sources: <a href="../doi/10.5281/zenodo.13769681">EOL taxon identifier map</a></p>
DoubleChEC program to identify transcription factor binding sites from mapped ChEC-seq data
<p>ChIP-seq (chromatin immunoprecipitation followed by sequencing) is commonly used to identify genome-wide protein-DNA interactions. However, ChIP-seq often gives a low yield, which is not ideal for quantitative outcomes. An alternative method to ChIP-seq is ChEC-seq (Chromatin endogenous cleavage with high-throughput sequencing). In this method, the endogenous TF (transcription factor) of interest is fused with MNase (micrococcal nuclease) that non-specifically cleaves DNA near binding sites. Compared to the <a href="https://www.nature.com/articles/ncomms9733" rel="nofollow">original ChEC-seq method</a>, the <a href="https://sites.northwestern.edu/bricknerlab/" rel="nofollow">modified version</a> requires far less amplification. Since <a href="https://github.com/macs3-project/MACS/tree/master#introduction">MACS3</a> failed to identify peaks in data generated from the modified ChEC-seq method, a new peak finder has been developed specifically for it.</p> <p>There are three functions in the <em><code>peak_finder/</code></em>. <code>callpeaks()</code> is used to identify peaks from BAM files. <code>goanalysis()</code> is used to make GO (Gene Ontology) term plots from peaks. <code>bedtomeme()</code> is a wrapper function to perform <a href="https://meme-suite.org/meme/tools/meme" rel="nofollow">MEME analysis</a> in R <strong>after <a href="https://meme-suite.org/meme/doc/download.html" rel="nofollow">MEME Suite</a> is installed locally</strong>.</p>
Maps of ecosystem multifunctionality and ecological connectivity for identifying Green Infrastructure networks in the European Alps
<p>High resolution raster datasets (20 meters) containing the results of an ecological connectivity and an ecosystem multifunctionality assessment for identifying Green Infrastructure networks in 10 pilot regions of the European Alps, modelled as part of the LUIGI Interreg Alpine Space project. Pilot regions include: department of Isère (FR), departments of Savoie and Haute-Savoie (FR), Munich Metropolitan Region (DE), Central Area of Salzburg (AT), South Burgenland (AT), Goriška region (SI), South Tyrol (IT), canton of Grisons (CH), Metropolitan City of Milan (IT), and Metropolitan City of Turin (IT). For a preview of the data and the results available for each pilot region <a href="https://www.alpine-space.org/projects/luigi/en/project-results/d.t1.2.1-pilot-regions-policy-briefs">click here</a></p> <p>Further information on the LUIGI project is available at: <a href="https://www.alpine-space.org/projects/luigi/en/home">https://www.alpine-space.org/projects/luigi/en/home</a></p> <p><a href="https://webassets.eurac.edu/31538/1661510408-luigi-wp1-technical-annex-mapping-a-green-infrastructure-network-in-the-alpine-space.pdf">https://webassets.eurac.edu/31538/1661510408-luigi-wp1-technical-annex-mapping-a-green-infrastructure-network-in-the-alpine-space.pdf </a></p> <p>The datasets include:</p> <ul> <li>a map for ecosystem service-based multifunctionality calculated out of the average of 11 standardized ecosystem service indicators: water provision, crop potential, timber production, fodder provision, pollination potential, carbon sequestration, nitrogen retention, natural hazard mitigation, runoff retention, outdoor recreation, and landscape aesthetics.</li> <li>a map of the modelled Ecological Network composed of core areas and ecological corridors. Corridors are modelled for medium-large forest mammal species and represent least-cost pathways connecting core areas. Different classes indicate areas with different levels of current ecological connectivity starting from core areas to areas in cities or anthropized land with no connectivity. Modeled corridors are presented in two classes to mirror different levels of prioritization and management actions.</li> <li>a map of the resistance of the landscape to the movement of forest mammal species. The landscape resistance raster has been developed by reclassifying and aggregating a high resolution (5m) land use and land cover map. Resistance values have been determined in relation to the naturalness of different land use and land cover classes. In this context, land use or landscape resistance is intended as the opposite of habitat suitability.</li> </ul>
Derby database for mapping secondary to primary HMDB identifiers
<p>The data (hmdb_metabolites, released on 17/11/2021) used to create this ID mapping database was downloaded from HMDB (<em>Human Metabolome Database, </em>website URL: https://hmdb.ca/). </p> <p>This database was used for the <a href="https://github.com/tabbassidaloii/BridgeDbDemoBioSB2022">BridgeDb demo at BioSB 2022</a> conference.</p> <p>The scripts used to create this database based on HGNC: https://github.com/tabbassidaloii/create-bridgedb-secondary2primary</p> <p>This work was funded by the <a href="https://fairplus-project.eu/">FAIRplus project</a> (grant agreement no 802750) and <a href="https://www.nwo.nl/en/researchprogrammes/open-science/open-science-fund/open-science-fund-2021-awarded-grants">NWO Open Science Fund</a> (grant no <a href="https://www.nwo.nl/en/projects/203001121">203.001.121</a>).</p>
Association mapping identified novel candidate loci affecting wood formation in Norway spruce
<p>Data sets associated with the study for the Association mapping and identification of novel candidate loci affecting wood formation in Norway spruce</p>
EOL taxon identifier map
<p>A mapping of taxon identifiers from major classification sources to EOL, of the form:</p> <p>node_id, resource_pk, resource_id, page_id, preferred_canonical_for_page</p> <ul> <li>node_id: internal to EOL; useful for some API calls</li> <li>resource_pk: identifier according to the classification provider</li> <li>resource_id: identifies the classification provider (see below)</li> <li>page_id: EOL taxon concept identifier; official, for sharing</li> <li>preferred_canonical_for_page: canonical name preferred by EOL for this taxon concept</li> </ul> <p>resource_ids, their names and descriptions, are available at: <a href="https://eol.org/resources.json" target="_blank" rel="nofollow noopener">https://eol.org/resources.json</a> </p> <div> <p>There is also a summary file with resources ids, links, and resource names here: <a href="https://github.com/KatjaSchulz/eolResources">https://github.com/KatjaSchulz/eolResources</a></p> </div> <p>To view one at a time: <a href="https://eol.org/resources/" target="_blank" rel="nofollow noopener">https://eol.org/resources/</a>[enter ID here]</p> <p>Included classification providers:</p> <p>1-> EOL Dynamic Hierarchy version 2.1, <a href="https://eol.org/resources/1">https://eol.org/resources/1</a><br>5 -> IUCN, <a href="https://eol.org/resources/5">https://eol.org/resources/5</a> <br>459 -> World Register of Marine Species (WoRMS), <a href="https://eol.org/resources/459">https://eol.org/resources/459 </a><br>676 -> National Center for Biotechnology Information (NCBI), <a href="https://eol.org/resources/676">https://eol.org/resources/676</a><br>695 -> Integrated Taxonomic Information System (ITIS), <a href="https://eol.org/resources/695">https://eol.org/resources/695</a> <br>724 -> EOL Dynamic Hierarchy version 1.1, <a href="https://eol.org/resources/724">https://eol.org/resources/724</a> <br>767 -> GBIF classification, <a href="https://eol.org/resources/767">https://eol.org/resources/767</a> <br>1072 -> Wikidata hierarchy, <a href="https://eol.org/resources/1072">https://eol.org/resources/1072</a> <br>1162 -> Catalogue of Life (COL), <a href="https://eol.org/resources/1162">https://eol.org/resources/1162</a> <br>1174 -> Dynamic Hierarchy Version 2.2.3 - Test, <a href="https://eol.org/resources/1174">https://eol.org/resources/1174</a></p> <p>commas within entries are "escaped, by, quoting", and quotes-within-quotes are ""double quoted"" </p> <p>There is also a much larger mapping file for all EOL resources: <a href="../doi/10.5281/zenodo.13253932">EOL full taxon identifier map</a></p>
DoubleChEC program to identify transcription factor binding sites from mapped ChEC-seq data
Open the record for dataset details and reuse information.
Genome-wide association mapping to identify genetic loci for cold tolerance and cold recovery during germination in rice
<p>To investigate the genetic architecture underlying cold tolerance during germination in rice (<i>Oryza sativa</i>), we conducted a genome-wide association study (GWAS) using a novel diversity panel of 257 rice accessions from around the world and 5,185 SNP markers from a 7K SNP marker array. Genotyping was performed using a 7K Illumina iSelect custom-designed array by following the Infinium HD Array Ultra Protocol. The 7K array, called the C7AIR, was designed by Dr. Susan McCouch's Lab at Cornell University and consists of 7,098 SNPs (Morales et al. 2020, under review). After genotyping 257 rice accessions with the 7K array (C7AIR), poor-performing SNP markers (SNPs of call rate <90%; minor allele frequency <5%; or heterozygosity >20%) were removed from the dataset. For our study, a subset of 5,185 high-quality SNP markers obtained after filtering was used to perform the genome-wide association analysis. The dataset representing the genotype data of 5,185 SNP markers by 257 rice accessions is presented here.</p>
Improved multi-ancestry fine-mapping identifies cis-regulatory variants underlying molecular traits and disease risk
<p>sushie.molqtl.weights.tar.gz contains ancestry-specific eQTL and pQTL weights trained on mRNA and protein levels measured in American European, American African, and American Hispanic ancestries from TOPMed-MESA and GENOA studies. Column “a1” is the counting allele.</p> <p>mesa.*.fusion.tar.gz contains the weights in FUSION format.</p> <p>sushie_real_data_results.tar.gz contains all the real data analyzed in the sushie project.</p> <p>sushie_sim_data_results.tar.gz contains all the sim data analyzed in the sushie project.</p> <p>sushie_analysis_codes.tar.gz contains all the codes and scripts to generate and analyze these data.</p>
Data from: mapping endemic freshwater fish richness to identify high priority areas for conservation: an ecoregion approach
<p>Freshwater ecosystems are experiencing accelerating global biodiversity loss. Thus, knowing where these unique ecosystems' species richness reaches a peak can facilitate their conservation planning. By hosting more than 290 freshwater fishes, Iran is a major freshwater fish hotspot in the Middle East. Considering the accelerating rate of biodiversity loss, there is an urgent need to identify species rich areas and understanding of the mechanisms driving biodiversity distribution. In this study, we gathered distribution records of all endemic freshwater fishes of Iran (85 species) to develop their richness map and determine the most critical drivers of their richness patterns from an ecoregion approach. We performed a generalized linear model (GLM) with quasi-Poisson distribution to identify contemporary and historical determinants of endemic freshwater fish richness. We also quantified endemic fish similarity among the 15 freshwater ecoregions of Iran. Results showed that endemic freshwater fish richness is highest in the Zagros Mountains while moderate level of richness was observed between Zagros and Alborz Mountains. High, moderate and low richness of endemic freshwater fish match with Upper Tigris & Euphrates, Namak, and Kavir & Lut Deserts ecoregions respectively. Kura - South Caspian Drainages and Caspian Highlands were the most similar ecoregions and Orumiyeh was the most unique ecoregion according to endemic fish presence. Precipitation and precipitation change velocity since the Last Glacial Maximum were the most important predictors of endemic freshwater fish richness. Areas identified to have the highest species richness have high priority for the conservation of freshwater fish in Iran, therefore, should be considered in future protected areas development.</p>
Data from: Male mouse recombination maps for each autosome identified by chromosome painting
<p>Linkage maps constructed from genetic analysis of gene order and crossover frequency provide few clues to the basis of the genomewide distribution of meiotic recombination, such as chromosome structure, that influences meiotic recombination. To bridge this gap, we have generated the first cytological recombination map that identifies individual autosomes in the male mouse. We prepared meiotic chromosome (synaptonemal complex [SC]) spreads from 110 mouse spermatocytes, identified each autosome by multicolor fluorescence in situ hybridization of chromosome- specific DNA libraries, and mapped 12,000 sites of recombination along individual autosomes, using immunolocalization of MLH1, a mismatch repair protein that marks crossover sites. We show that SC length is strongly correlated with crossover frequency and distribution. Although the length of most SCs corresponds to that predicted from their mitotic chromosome length rank, several SCs are longer or shorter than expected, with corresponding increases and decreases in MLH1 frequency. Although all bivalents share certain general recombination features, such as few crossovers near the centromeres and a high rate of distal recombination, individual bivalents have unique patterns of crossover distribution along their length. In addition to SC length, other, as-yet-unidentified, factors influence crossover distribution leading to hot regions on individual chromosomes, with recombination frequencies as much as six times higher than average, as well as cold spots with no recombination. By reprobing the SC spreads with genetically mapped BACs, we demonstrate a robust strategy for integrating genetic linkage and physical contig maps with mitotic and meiotic chromosome structure.</p>
Spatial mapping of the hepatocellular carcinoma landscape identifies unique intratumoural perivascular-immune neighbourhoods
<p>The uploaded data includes results from imaging mass cytometry (IMC) data collected from hepatocellular carcinoma patients. The associated publication can be found <a href="https://journals.lww.com/hepcomm/fulltext/2024/11010/spatial_mapping_of_the_hcc_landscape_identifies.11.aspx">here</a> (Marsh-Wakefield <em>et al.</em>, 2024, <em>Hepatology Communications</em>).</p> <p>The CSV file contains segmented cells from IMC data. This includes the Patient, ROI, and Group each cell is assigned. Marker signal intensities underwent arcsine transformation and were rescaled. The “simprof_cluster” column contains the final iteration of clustering following initial X-shift clustering.</p> <p>Notes on additional columns:</p> <ul> <li>“Sample” is barcoded such that the first three digits are the ablation number, followed by the region on the TMA, the group, and the patient. I.e., “[ablation.number]_[TMA.location]_[group]_[patient]”.</li> <li>“x” and “y” refer to the coordinates of samples.</li> <li>“Group” refers to the tissue type. Included non-tumour (NT), invasive margin (IM), and tumour (T) regions.</li> <li>The area for each ROI has been calculated (µm^2 and mm^2).</li> <li>“Batch” refers to TMA.</li> <li>In most cases each area from each patient has three ablation sites. Three samples have an extra ablation site due to technical difficulties during the ablation, and hence have a “split” sample.</li> </ul> <p>The PDF file contains patient information associated with the IMC data.</p> <p>DOI of dataset:</p> <p>10.5281/zenodo.10622397</p> <p>Any further questions can be addressed to Felix Marsh-Wakefield felix.marsh-wakefield@sydney.edu.au</p>
Spatial Mapping of Mobile Genetic Elements and their Cognate Hosts in Complex Microbiomes - Identifying the host taxon of a previously undescribed plasmid
<p>We investigated the taxonomic association of an unknown plasmid within a plaque biofilm of a patient diagnosed with stage 3 periodontitis. We combined long- and short- read sequencing to identify a complete plasmid with minimal homology to any sequence in the RefSeq database. The plasmid carried several predicted genes for mobilization and toxin-antitoxin systems. We designed MGE-FISH probes for the plasmid and combined this MGE-FISH stain with an 18-genera HiPR-FISH panel.</p> <p>Images are labeled by collection time such that the laser order for a given field of view (fov) is: 488nm Lambda, 514nm Lambda, 561nm Lambda, 633nm Airyscan, 405nm Lambda. We used Flye (https://github.com/fenderglass/Flye) to assemble the plasmid using long read Nanopore sequencing only and we used OPERA-MS (https://github.com/CSB5/OPERA-MS) to do hybrid assembly with Illumina short reads and Nanopore long reads. The assemblies are in the fasta files and the reads that map to the assemblies are in the fastq files. </p>
Data for "Identifying and Fitting Eclipse Maps of Exoplanets with Cross-Validation" (Hammond et al. 2024)
<p>This archive contains the data and scripts needed to reproduce the analysis in "Identifying and Fitting Eclipse Maps of Exoplanets with Cross-Validation" (Hammond et al. 2024).</p> <p>Contents</p> <p>data/: Input data and posterior distributions of different model fits</p> <p>datasets/: Observational datasets</p> <p>figures/: Folder to save figures in</p> <p>fluxes/: Saved lightcurves for auxiliary plotting purposes</p> <p>archive_paper_plotter.ipynb: Example script to plot fitted eclipse maps</p> <p>eclipse_pixel_sampling.py: Script to fit eclipse map and test k-fold CV score</p> <p>paper_eclipse_suite.py: Script to use simulated or observational data to fit an eclipse map</p>
Association mapping identified novel candidate loci affecting wood formation in Norway spruce
<p>Genotypic data set for the association mapping in Norway spruce for wood formation and tracheid traits.</p>
Inputs & results for "Identifying circular DNA using short-read mapping"
<div>* `inputs`: Contains genome assemblies, annotations, and sample sheets for the Nextflow pipeline</div> <div> * `parasite_species`: Contains the sample sheet and input data for the parasite/related species dataset</div> <div> * `parasitoid_wasps`: Contains the sample sheet and input data for the parasitoid wasp dataset</div> <div> </div> <div>* `results`: Contains filtered BAM and coverage files, figures, and `GenomeInfo`-filtered example files for the parasitoid wasp and parasite/related species datasets</div> <div> * `parasite_species`:</div> <div> * `bam_files`: Contains BAM files with mapped distances >= 1 kb</div> <div> * `coverage_files`: Contains coverage depth files filtered by the BAM files</div> <div> * `insert_filtered_results`: Contains examples of filtered file outputs from the `GenomeInfo` class. The complete set of outputs can be found on Zenodo.</div> <div> * `fig`: Figures generated for the pub.</div> <div> * `parasitoid_wasps`:</div> <div> * `bam_files`: Contains BAM files with mapped distances >= 1 kb</div> <div> * `coverage_files`: Contains coverage depth files filtered by the BAM files</div> <div> * `fig`: Figures generated for the pub.</div> <div> * `hyposoter_didymator_blastx_results`: Contains the manual BLASTx results from the _Hyposoter didymator_ search.</div>
Supplementary tables for the paper: "Comprehensive Mapping of the AOP-Wiki Database: Identifying Biological and Disease Gaps"
<p><strong>Supplementary tables for the paper: "Comprehensive Mapping of the AOP-Wiki Database: Identifying Biological and Disease Gaps"</strong></p> <p><em>Original Research Article</em><br><strong>Frontiers in Toxicology</strong>, March 8, 2024<br>Section: Regulatory Toxicology<br><strong>Volume 6 - 2024</strong> | <a href="https://doi.org/10.3389/ftox.2024.1285768" target="_new" rel="noopener">https://doi.org/10.3389/ftox.2024.1285768</a></p> <p><strong>Authors</strong>:<br>Thomas Jaylet, Thibaut Coustillet, Nicola M. Smith, Barbara Viviani, Birgitte Lindeman, Lucia Vergauwen, Oddvar Myhre, Nurettin Yarar, Johanna M. Gostner, Pablo Monfort-Lanzas, Florence Jornod, Henrik Holbech, Xavier Coumoul, Dimosthenis A. Sarigiannis, Philipp Antczak, Anna Bal-Price, Ellen Fritsche, Eliska Kuchovska, Antonios K. Stratidakis, Robert Barouki, Min Ji Kim, Olivier Taboureau, Marcin W. Wojewodzic, Dries Knapen, Karine Audouze</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.