Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

6,766

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

6,766 results for “project”

Learn how ShareScore rates datasets ↗
zenodo44/100

Model results and configuration files for "Large modeling uncertainty in projecting decadal surface ozone changes over urban and industrial regions of China"

<p>This repository includes files as described below:</p> <p><strong>1. namelist_CBMZ09_example.input, namelist_MOZART202_example.input:</strong></p> <p>Two WRF-chem namelist files for CBMZ and MOZART simulation.</p> <p>They are modified according to the namelist from <a href="https://github.com/wrfchem-leeds/WRFotron">https://github.com/wrfchem-leeds/WRFotron</a>.</p> <p><strong>2. wps_namelist_example.wps:</strong></p> <p>namelist for WRF Preprocessing System (WPS)</p> <p><strong>3. temporal_hourly_scale_factor_emission.csv:</strong></p> <p>Hourly scale factors for emissions.</p> <p>Hourly allocation is applied to all emission data (i.e., emissions for 2017, 2030 and perturbated emissions of NOx, VOCs).</p> <p><strong>4. vertical_emission_ratio.csv</strong></p> <p>Vertical shares (ratios) of emissions.</p> <p>Emissions from sectors of power and industry are vertically allocated based on this file. Vertical allocation is conducted for all emission data.</p> <p>These shares are suggested by MICS-ASIA III intercomparison framework.</p> <p><strong>5. 01_2030_2017_simulations.zip: </strong></p> <p>Simulated MDA8 ozone under future (2030) and 2017 emission scenarios by the two chemical mechanisms (i.e., CBMZ, MOZART).</p> <p><strong>6. 02_perturbations_of_NOxVOCs.zip:</strong></p> <p>Simulated MDA8 ozone given perturbations of NOx and VOCs emissions by the two chemical mechanisms.</p> <p><strong>7. 03_hourly_diff_O3_NOx_OH_HNO3.zip: </strong></p> <p>Differences of hourly simulated concentrations of O3, NOx, OH and HNO3 during July in the Base-2017 scenario between CBMZ and MOZART (CBMZ - MOZART).</p>

opencc-by-4.0Feb 2023View details →
zenodo44/100

Interviews with experts on digitalisation in agriculture, forestry, and rural areas (H2020 DESIRA project, WP1)

<p>Interviews with experts on digitalisation in agriculture, forestry, and rural areas (H2020 DESIRA project, WP1).</p> <p>The scripts and the answers are provided. Two groups of experts have been interviewed: the first group with expertise in ICT, and the second group with expertise in socio-economic aspects.&nbsp;&nbsp;</p>

opencc-by-4.0Feb 2023View details →
zenodo44/100

Results of the crowd-mapping action within the project TeRRIFICA [Dataset No. 1 dated 2022-09-19]

<p>The dataset includes the results of the crowd-mapping action within the project &quot;Territorial RRI fostering innovative climate action&quot; - TeRRIFICA (Horizon 2020 under GA 824489) dated 2022-09-19. The data are points added to the map by the users (volunteers) and represent locations where climate change-related issues occur regarding air temperature,&nbsp;air quality,&nbsp;water,&nbsp;soil,&nbsp;and wind (SPOTS). The second part of the dataset is related to the crowd-mapping users and their anonymized characteristics (USERS). More details are available at&nbsp;https://terrifica.eu/.</p>

opencc-by-4.0Sep 2022View details →
zenodo44/100

ASSIST-IOT Open Call Project RAZOR DATASETS (INSIGHIO)

<p>Example datasets for road anomaly detection produced in the context of RAZOR Open Call ASSIST-IoT Project, carried out by INSIGHIO.</p>

opencc-by-4.0Feb 2023View details →
zenodo44/100

Future Projections for Posidonia oceanica and Zostera marina in Europe

<p>Future projections of seagrass biomass for <em>Posidonia oceanica</em> and <em>Zostera marina </em>in Europe.</p> <p>Supporting data for T4.1 of Horizon 2020 project FutureMARES.</p> <p>The file naming convention is {species}_{cmip6_model}_{ssp}, where <em>species</em> identifies whether the model run is for <em>P. oceanica</em> or <em>Z. marina</em>, <em>cmip6_model</em> names the CMIP6 model uses to drive the projections, and <em>ssp</em> identifies which of SSP126, SSP245 and SSP585 were used.</p> <p>Projections cover the year 1995-2099.</p> <p>&nbsp;</p>

opencc-by-4.0Feb 2023View details →
zenodo44/100

GEOLAB - Transnational Access project QC-CEM - Mapping quick clay with geophysical methods

<p>Quick clay is characterised by complete collapse and liquid-like mobility when overloaded. Quick clay is found primarily in Norway and Sweden, but also exists in Finland, Russia, Canada and Alaska. Quick clay landslides, with their retrogression characteristics and extreme mobility, pose significant risk to human lives, infrastructure, property and surrounding ecosystems. Hence, the proper characterization of quick clay sites is essential for ensuring the safety and resilience of infrastructure in Norway and elsewhere in Europe.<br> The current practice for mapping quick clay in Norway relies heavily on borehole data with either rotary sounding or total sounding and core samples tested in the laboratory. The only method for identifying quick clay with certainty is physical testing in the laboratory, but it is time-consuming, expensive and gives limited information, i.e., only at the depths and locations where the samples are taken. In Norway, rotary sounding and total soundings are frequently used in mapping of quick clay. There is increasing interest in using geophysical methods such as Electrical Resistivity Tomography (ERT) to supplement the results from soundings, particularly in early stage of ground investigation for mapping of quick clay. ERT is a near surface geophysical method that uses direct current to measure the earth&#39;s electrical resistivity. The current is injected into the subsurface through steel electrodes installed 10-20 cm into the ground, and the apparent resistivity distribution along a profile or area is measured. Using data processing and inverse modelling a 2D or 3D resistivity model of the subsurface can be derived.<br> Geophysical methods such as ERT show capability to identify not quick clay such as sand, silt, dry crust, moraine and bed rock reasonably accurate, but the identification of quick clay is still generally limited. The detection of leached clay (thus potentially quick clay) is however possible.<br> Transnational Access project QC-CEM is funded through the 1st call for proposal for the GEOLAB project. This project aims at testing various geophysical methods for their capability for soil characterisation, particularly for detecting quick clay.</p> <p>The objectives of the QC-CEM project are:<br> (i) to test different configurations of Electrical Resistivity Tomography survey for detection of quick clay<br> (ii) to test innovative and efficient electromagnetic based methods for mapping of quick clay. Results from this investigation is not available to share at this stage.<br> (iii) to investigation the effectiveness of cross-interpretation using different geophysical methods for soil characterisation. The results from this activity will be published in open publication after they are processed.</p>

opencc-by-4.0Mar 2023View details →
zenodo44/100

Hesperomys Project v23.3.0

<p>An export of data from the <a href="https://hesperomys.com">Hesperomys Project</a> version 23.3.0. The Hesperomys Project is a database of taxonomy and nomenclature, focused on mammals but also covering some other groups, principally other fossil tetrapods. The database contains information such as:</p> <ul> <li>Taxonomic classification for all mammals, living and extinct</li> <li>References to original citations for the vast majority of names</li> <li>Type specimens and type localities for numerous names</li> </ul> <p>The full database is available online at hesperomys.com. This export contains:</p> <ul> <li>name.csv: Data on names, including taxonomic context, authority, citation, type locality, type specimen, and classification of the etymology.</li> <li>taxon.csv: Data on taxa, including classification and authority</li> <li>collection.csv: Data on collections that contain type specimens, including name, location, and number of type specimens in the database</li> </ul> <p>The code used to generate the exports is on <a href="https://github.com/JelleZijlstra/taxonomy/blob/9a68bec908264ddfcb1d623b4c427ebfb689396f/taxonomy/db/export.py">GitHub</a>.</p> <p>Release notes for version 23.3.0:</p> <ul> <li>Database <ul> <li>Incorporate some data from the African Chiroptera Database (thanks to Victor Van Cakenberghe).</li> <li>Incorporate many missed names from the Mammal Diversity Database (with more to come)</li> <li>Add SMF mammalian type specimens from a type catalog I located.</li> <li>Various new data, including some additional original citations and a few new species.</li> </ul> </li> <li>Backend <ul> <li>Add the ability to associate type catalogs and collection databases with Collection objects.</li> <li>Add the ability to link to collection database entries for type specimens.</li> <li>Support name aliases, in order to provide more familiar citation forms for some personal names. This feature is not yet widely used.</li> </ul> </li> <li>Frontend <ul> <li>Add bibliographic notes on <em>Zoology of the Erebus and Terror</em> and <em>Histoire naturelle des Mammif&egrave;res</em>.</li> <li>Add new pages on data sources and scores.</li> </ul> </li> </ul>

opencc-by-4.0Mar 2023View details →
zenodo44/100

The Effect of Soundscape Composition on Bird Vocalization Classification in a Citizen Science Biodiversity Monitoring Project

<p>This archive includes sound clips (.wav files) and associated mel-scale spectrograms of bird vocalizations for 54 species in Sonoma County, California, USA. These data were used for training and validating convolutional neural network (CNN) models for bird species detection. We also include xeno-canto training and validation mel spectrograms&nbsp;used to pretrain CNNs. Details on these data are explained in the paper by Clark et al. (2023) titled &quot;The effect of soundscape composition on bird vocalization classification in a citizen science biodiversity monitoring project&quot;. These data are available for use without restrictions, with no warranty on data quality or utility for a given application. We request that any work that does use these data cite the Clark et al. (2023) paper.<br> <br> Clark, M.L., Salas, L., Baligar, S., Quinn, C., Snyder, R.L., Leland, D., Schackwitz, W., Goetz, S.J., Newsam, S. (2023). The effect of soundscape composition on bird vocalization classification in a citizen science biodiversity monitoring project. <em>Ecological Informatics</em>.&nbsp;<a href="https://doi.org/10.1016/j.ecoinf.2023.102065">https://doi.org/10.1016/j.ecoinf.2023.102065</a></p> <p>Associated code for training CNN models,&nbsp;performing inference, and applying post-classification corrections can be found in the GitHub archive&nbsp;<a href="https://github.com/pointblue/Soundscapes2Landscapes/tree/master/CNN_Bird_Species">https://github.com/pointblue/Soundscapes2Landscapes/tree/master/CNN_Bird_Species</a></p> <p>Raw sound data from the Soundscapes to Landscapes project are available upon request: Dr. Matthew Clark, matthew.clark@sonoma.edu</p> <p>These data were collected as part of the&nbsp;Soundscapes to Landscapes project (<a href="https://soundscapes2landscapes.org/">soundscapes2landscapes.org</a>),&nbsp;funded by NASA&rsquo;s Citizen Science for Earth Systems Program (CSESP) 16-CSESP 2016-0009 under cooperative agreement 80NSSC18M0107.<br> <br> ----------------------------<br> This depository&nbsp;includes the following archives:</p> <ul> <li> <p>mel_specs.zip: contains 2-sec mel spectrograms split into training (&ldquo;tr&rdquo;), validation (&ldquo;val&rdquo;), testing (&ldquo;test&rdquo;) data for each target bird species (n = 54) used to fine-tune the CNNs. Select spectrogram files are appended with &ldquo;aug&rdquo; if they are augmented versions for the training data.</p> </li> <li> <p>wav.zip: contains the associated wav-format sound recordings used to generate the training, validation, testing mel spectrograms found in mel_specs.zip.</p> </li> <li> <p>Xeno-canto_pretrain.tar: contains 2-sec mel spectrograms split into training and validation data for 40 bird species used for CNN pre-training that were generated using a warbleR segmentation methodology described in the paper. The sound files used to generate these mel spectrograms came from the Kaggle competition,&nbsp;<a href="https://www.kaggle.com/datasets/imoore/xenocanto-bird-recordings-dataset">https://www.kaggle.com/datasets/imoore/xenocanto-bird-recordings-dataset</a><br> Mel spectrogram naming reflects the XC number used for cataloging on Xeno-canto in the format XC123456_2.png. The six numbers following the XC characters can be used to search for unique recordings on Xeno-canto (<a href="https://xeno-canto.org/">https://xeno-canto.org/</a>) using the search query &ldquo;nr:123456&rdquo; in the search tool or queried using the Xeno-canto API (<a href="https://xeno-canto.org/explore/api">https://xeno-canto.org/explore/api</a>). Unique recording names can be extracted from the mel spectrogram filenames.</p> </li> <li> <p>soundscape_test_wavs.zip: the wav-format&nbsp;sound recordings&nbsp;used to perform soundscape testing.</p> </li> </ul>

opencc-by-4.0Mar 2023View details →
zenodo44/100

AIRSEAL Project - Data sample

<p>The attached files contain test data (main parameters measured during rotating labyrinth seals testing).<br> The .csv data is structured as follows:</p> <p>&nbsp;&nbsp;&nbsp;&nbsp; column1 =&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; Time [s]<br> &nbsp;&nbsp;&nbsp;&nbsp; column2 =&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; Tut&nbsp; [deg C]<br> &nbsp;&nbsp;&nbsp;&nbsp; column3 =&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; PR&nbsp;&nbsp; [non dimentional]<br> &nbsp;&nbsp;&nbsp;&nbsp; column4 =&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; N&nbsp;&nbsp;&nbsp; [rpm]<br> &nbsp;&nbsp;&nbsp;&nbsp; column5 =&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; CLR&nbsp; [mm]<br> &nbsp;&nbsp;&nbsp;&nbsp; column6 =&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; Qma&nbsp; [kg/s]<br> &nbsp;&nbsp;&nbsp;&nbsp; column7 =&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; Swirler_angle [deg]<br> &nbsp;<br> The &quot;11&quot; in file names is related to the first configuration (which has been tested in the project): METCO casing and axial flow (0 deg swirl).<br> The &quot;1&quot; to &quot;30&quot; file indexes are related to test number (stabilized regime: pressure ratio, rotor speed and temperature).</p>

opencc-by-4.0Mar 2023View details →
zenodo44/100

IceAq Project - Ground Penetrating Profiles Data

<p>Ground Penetrating Radar profiles on 47 sections.&nbsp;Equipment used: Cobra Plug-In by Radarteam</p> <p>45 .sgy files containing the GPR data,&nbsp;45 text files containing the corresponding GPS data,&nbsp;1 spreadsheet giving the correlation table between the .sgy and GPS text files.&nbsp;Total 424 Mb.</p> <p>Origin: MSCA Project Number: 885891 Project Acronym: IceAq</p> <p>Project title: Proglacial and subglacial aquifers: their evolution under climate change and the potential impacts in terms of resources and natural hazards, through the case of eastern Iceland</p> <p>Period of data collection: June 2022, precise periods specified for each file.</p> <p>Area of data collection: Iceland, south-east of Vatnaj&ouml;kull, Su&eth;ursveit and M&yacute;rar.</p> <p>Accessibility: The data can be read using e.g. the following open source softwares: LibreOffice (https://www.libreoffice.org), QGIS (<a href="https://qgis.org/">https://qgis.org</a>), and Geopsy (https://www.geopsy.org/).</p>

opencc-by-4.0Mar 2023View details →
zenodo44/100

Project "Public services management system to improve the quality and accessibility of services" (01.2.2-LMT-K-718-03-0019) interviews

<p>The dataset of depersonalized qualitative semi-structured interviews with the representatives of public sector organisations providing public services in Lithuania. The data were collected in December-November, 2022 as a part of the project &quot;Public services management system to improve the quality and accessibility of services&quot; (&quot;Vie&scaron;ųjų paslaugų vadybos sistema paslaugų kokybei ir prieinamumui gerinti&quot;), grant no. 01.2.2-LMT-K-718-03-0019, funded by the Lithuanian research council. The interviews are in the Lithuanian language.</p>

opencc-by-4.0Mar 2023View details →
zenodo44/100

Assessment of current and future invasive plants in protected dune habitats of the Atlantic coastal region for the LIFE DUNIAS project (LIFE20 NAT/BE/001442)

<p>This .csv file contains the raw data from the risk screening supplementing the LIFE DUNIAS horizon scan for (invasive) alien species in protected habitats of Atlantic coastal dune ecosystems (<a href="https://doi.org/10.21436/inbor.86703335">Adriaens et al. 2022</a>). We gladly refer to the annexes and methods section in this report for more explanation about the fields and their contained values.</p> <p>The file contains the following fields:</p> <p><em>TaxonName</em>: original taxonomic name of the considered alien species</p> <p><em>WorkName</em>:&nbsp;taxonomic name of the considered alien species after lumping of subspecies, closely related species of a complex, functionally similar species of the same genus (see chapter 3.1)</p> <p><em>hab_xxxx</em> (1110,&nbsp;1130,&nbsp;1140,&nbsp;1210,&nbsp;1230, 1310,&nbsp;1320,&nbsp;1330,&nbsp;2110,&nbsp;2120,&nbsp;2130,&nbsp;2140,&nbsp;21A0,&nbsp;2150,&nbsp;2190,&nbsp;2160,&nbsp;2170,&nbsp;2180): susceptibility of habitat for the alien species (4-digit code refering to the Annex I habitat under the Habitats Directive)&nbsp;</p> <p><em>occ_XX</em> (BE,&nbsp;FR,&nbsp;IE,&nbsp;NL,&nbsp;ES,&nbsp;UK, DK,&nbsp;DE,&nbsp;PT,&nbsp;ALL): occupancy of the alien species in different countries of the Atlantic European region (as the number of 10km<sup>2</sup> squares per country). Country codes: BE = Belgium, FR = France, IE = Ireland, NL = Netherlands, ES = Spain, UK = United Kingdom, DK = Denmark, DE = Germany, PT = Portugal, ALL = total for all countries.</p> <p><em>scor_XXX_xxxx</em>: score of the assessment per criterium (INT = introduction, EST = establishment, SPR = spread, IMP = ecological impact, ALL = overall score) and per habitat group (salt = salties, sand = sandies,&nbsp;shru = shrubbies)&nbsp;conf_<em>XXX_xxxx</em>: confidence on the scores&nbsp;of the assessment per criterium (INT = introduction, EST = establishment, SPR = spread, IMP = ecological impact, ALL = overall score) and per habitat group (salt = salties, sand = sandies,&nbsp;shru = shrubbies)</p> <p><em>scor_ALL_MAX</em>: maximum ecological impact score of the alien taxon across all habitats</p>

opencc-zeroApr 2023View details →
zenodo44/100

The OHEJP BeONE Project – Escherichia coli genome assembly dataset

<p><strong>Dataset</strong></p> <p>This dataset comprises the genome assemblies of 308 <em>Escherichia coli</em> samples collected by the BeONE Consortium on behalf of the One Health European Joint Programme &ldquo;BeONE: Building Integrative Tools for One Health Surveillance&rdquo; (<a href="https://onehealthejp.eu/jrp-beone/">https://onehealthejp.eu/jrp-beone/</a>). Additionally, a complementary dataset is also made available (<a href="https://zenodo.org/record/7120057">https://zenodo.org/record/7120057</a>), comprising genome assemblies of 1,999&nbsp;<em>E. coli</em> samples selected among the Whole-Genome Sequencing (WGS) data publicly available in the European Nucleotide Archive (ENA) or in the National Center for Biotechnology Information (NCBI) Sequence Read Archive (SRA).</p> <p>File &ldquo;<strong>BeONE_Ec_metadata.xlsx</strong>&rdquo; contains the genome assembly statistics for each isolate, including European Nucleotide Archive accession numbers, in-silico Multi Locus Sequence Type and Serotype, and information regarding year of sampling, country and source.</p> <p>The archive &ldquo;<strong>BeONE_Ec_assemblies.zip</strong>&rdquo; contains all the genome assemblies (.fasta format) of each isolate presented in the metadata file.</p> <p>&nbsp;</p> <p><strong>Dataset selection and curation</strong></p> <p>This anonymized dataset of <em>E. coli</em>&nbsp;genome assemblies was generated using Next Generation Sequencing data collected within the BeONE Consortium available at the European Nucleotide Archive under BioProject Accession Number&nbsp;<a href="https://www.ebi.ac.uk/ena/browser/view/PRJEB57098">PRJEB57098</a>. Read quality control, trimming and assembly were performed with Aquamis v1.3.9 (<a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC8145556/">Deneke et al. 2021</a>) using default parameters. Assembly quality control (QC), including contamination assessment, as well as MLST ST determination were performed with the same pipeline. All genome assemblies passing the QC were included in the final dataset. Among the others, we noticed that a considerable proportion of assemblies was flagged as &ldquo;QC fail&rdquo; exclusively due to the &ldquo;NumContamSNVs&rdquo; parameter, suggesting that this setting might have been too strict. After manual inspection of a random subset, assemblies for which the percentage of reads corresponding to the correct species was &gt;98% were recovered and integrated in the final dataset (those samples are labeled in the Metadata file). In total, 308 isolates passed the dataset curation step and were included in the final dataset.&nbsp;In-silico serotyping was performed with&nbsp;<a href="https://github.com/B-UMMI/seq_typing">seq_typing</a>&nbsp;v2.2.</p> <p>&nbsp;</p> <p><strong>Funding</strong></p> <p>This work was supported by funding from the European Union&rsquo;s Horizon 2020 Research and Innovation programme under grant agreement No 773830: One Health European Joint Programme.&nbsp;</p> <p>&nbsp;</p> <p><strong>Acknowledgements</strong></p> <p>We thank the National Distributed Computing Infrastructure of Portugal (INCD) for providing the necessary resources to run the genome assemblies. INCD was funded by FCT and FEDER under the project 22153-01/SAICT/2016.</p>

opencc-by-4.0Oct 2022View details →
zenodo44/100

The OHEJP BeONE Project – Campylobacter jejuni genome assembly dataset

<p><strong>Dataset</strong></p> <p>This dataset comprises the genome assemblies of 610 <em>Campylobacter jejuni&nbsp;</em>samples collected by the BeONE Consortium on behalf of the One Health European Joint Programme &ldquo;BeONE: Building Integrative Tools for One Health Surveillance&rdquo; (<a href="https://onehealthejp.eu/jrp-beone/">https://onehealthejp.eu/jrp-beone/</a>). Additionally, a complementary dataset is also made available (<a href="https://zenodo.org/record/7120166">https://zenodo.org/record/7120166</a>), comprising genome assemblies of&nbsp;3,076&nbsp;<em>C. jejuni</em>&nbsp;samples selected among the Whole-Genome Sequencing (WGS) data publicly available in the European Nucleotide Archive (ENA) or in the National Center for Biotechnology Information (NCBI) Sequence Read Archive (SRA).</p> <p>File &ldquo;<strong>BeONE_Cj_metadata.xlsx</strong>&rdquo; contains the genome assembly statistics for each isolate, including European Nucleotide Archive accession numbers and&nbsp;<em>in-silico</em> Multi Locus Sequence Type,&nbsp;and information regarding year of sampling, country and source.</p> <p>The archive &ldquo;<strong>BeONE_Cj_assemblies.zip</strong>&rdquo; contains all the genome assemblies (.fasta format) of each isolate presented in the metadata file.</p> <p>&nbsp;</p> <p><strong>Dataset selection and curation</strong></p> <p>This anonymized dataset of <em>C. jejuni</em>&nbsp;genome assemblies was generated using Next Generation Sequencing data collected within the BeONE Consortium available at the European Nucleotide Archive under BioProject Accession Number <a href="https://www.ebi.ac.uk/ena/browser/view/PRJEB57119">PRJEB57119</a>. Read quality control, trimming and assembly were performed with Aquamis v1.3.9 (<a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC8145556/">Deneke et al. 2021</a>) using default parameters. Assembly quality control (QC), including contamination assessment, as well as MLST ST determination were performed with the same pipeline. All genome assemblies passing the QC were included in the final dataset. Among the others, we noticed that a considerable proportion of assemblies was flagged as &ldquo;QC fail&rdquo; exclusively due to the &ldquo;NumContamSNVs&rdquo; parameter, suggesting that this setting might have been too strict. After manual inspection of a random subset, assemblies for which the percentage of reads corresponding to the correct species was &gt;98% were recovered and integrated in the final dataset (those samples are labeled in the Metadata file). In total, 610 isolates passed the dataset curation step and were included in the final dataset.</p> <p>&nbsp;</p> <p><strong>Funding</strong></p> <p>This work was supported by funding from the European Union&rsquo;s Horizon 2020 Research and Innovation programme under grant agreement No 773830: One Health European Joint Programme.&nbsp;</p> <p>&nbsp;</p> <p><strong>Acknowledgements</strong></p> <p>We thank the National Distributed Computing Infrastructure of Portugal (INCD) for providing the necessary resources to run the genome assemblies. INCD was funded by FCT and FEDER under the project 22153-01/SAICT/2016.</p>

opencc-by-4.0Oct 2022View details →
zenodo44/100

The OHEJP BeONE Project – Listeria monocytogenes genome assembly dataset

<p><strong>Dataset</strong></p> <p>This dataset comprises the genome assemblies of 1,426 <em>Listeria monocytogenes</em> samples collected by the BeONE Consortium on behalf of the One Health European Joint Programme &ldquo;BeONE: Building Integrative Tools for One Health Surveillance&rdquo; (<a href="https://onehealthejp.eu/jrp-beone/">https://onehealthejp.eu/jrp-beone/</a>). Additionally, a complementary dataset is also made available (<a href="https://zenodo.org/record/7116878">https://zenodo.org/record/7116878</a>), comprising genome assemblies of 1,874 <em>L. monocytogenes</em> samples selected among the Whole-Genome Sequencing (WGS) data publicly available in the European Nucleotide Archive (ENA) or in the National Center for Biotechnology Information (NCBI) Sequence Read Archive (SRA).</p> <p>File &ldquo;<strong>BeONE_Lm_metadata.xlsx</strong>&rdquo; contains the genome assembly statistics for each isolate, including European Nucleotide Archive accession numbers and <em>in-silico</em> Multi Locus Sequence Type, and information regarding year of sampling, country and source.</p> <p>The archive &ldquo;<strong>BeONE_Lm_assemblies.zip</strong>&rdquo; contains all the genome assemblies (.fasta format) of each isolate presented in the metadata file.</p> <p>&nbsp;</p> <p><strong>Dataset selection and curation</strong></p> <p>This anonymized dataset of <em>L. monocytogenes </em>genome assemblies was generated using Next Generation Sequencing data collected within the BeONE Consortium available at the European Nucleotide Archive under BioProject Accession Number <a href="https://www.ebi.ac.uk/ena/browser/view/PRJEB57166">PRJEB57166</a>. Read quality control, trimming and assembly were performed with Aquamis v1.3.9 (<a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC8145556/">Deneke et al. 2021</a>) using default parameters. Assembly quality control (QC), including contamination assessment, as well as MLST ST determination were performed with the same pipeline. All genome assemblies passing the QC were included in the final dataset. Among the others, we noticed that a considerable proportion of assemblies was flagged as &ldquo;QC fail&rdquo; exclusively due to the &ldquo;NumContamSNVs&rdquo; parameter, suggesting that this setting might have been too strict. After manual inspection of a random subset, assemblies for which the percentage of reads corresponding to the correct species was &gt;98% were recovered and integrated in the final dataset (those samples are labeled in the Metadata file). In total, 1,426 isolates passed the dataset curation step and were included in the final dataset.</p> <p>&nbsp;</p> <p><strong>Funding</strong></p> <p>This work was supported by funding from the European Union&rsquo;s Horizon 2020 Research and Innovation programme under grant agreement No 773830: One Health European Joint Programme.</p> <p>&nbsp;</p> <p><strong>Acknowledgements</strong></p> <p>We thank the National Distributed Computing Infrastructure of Portugal (INCD) for providing the necessary resources to run the genome assemblies. INCD was funded by FCT and FEDER under the project 22153-01/SAICT/2016.</p> <p>&nbsp;</p>

opencc-by-4.0Oct 2022View details →
zenodo44/100

The OHEJP BeONE Project – Salmonella enterica genome assembly dataset

<p><strong>Dataset</strong></p> <p>This dataset comprises the genome assemblies of 1,540&nbsp;<em>Salmonella enterica</em> samples collected by the BeONE Consortium on behalf of the One Health European Joint Programme &ldquo;BeONE: Building Integrative Tools for One Health Surveillance&rdquo; (<a href="https://onehealthejp.eu/jrp-beone/">https://onehealthejp.eu/jrp-beone/</a>). Additionally, a complementary dataset is also made available (<a href="https://zenodo.org/record/7119735">https://zenodo.org/record/7119735</a>), comprising genome assemblies of 1,434&nbsp;<em>S. enterica</em>&nbsp;samples selected among the Whole-Genome Sequencing (WGS) data publicly available in the European Nucleotide Archive (ENA) or in the National Center for Biotechnology Information (NCBI) Sequence Read Archive (SRA).</p> <p>File &ldquo;<strong>BeONE_Se_metadata.xls</strong>x&rdquo; contains the genome assembly statistics for each isolate, including European Nucleotide Archive accession numbers, in-silico Multi Locus Sequence Type and Serotype, and information regarding year of sampling, country and source.</p> <p>The archive &ldquo;<strong>BeONE_Se_assemblies.zi</strong>p&rdquo; contains all the genome assemblies (.fasta format) of each isolate presented in the metadata file.</p> <p>&nbsp;</p> <p><strong>Dataset selection and curation</strong></p> <p>This anonymized dataset of <em>S. enterica</em>&nbsp;genome assemblies was generated using Next Generation Sequencing data collected within the BeONE Consortium available at the European Nucleotide Archive under BioProject Accession Number <a href="http://www.ebi.ac.uk/ena/browser/view/PRJEB57179">PRJEB57179</a>. Read quality control, trimming and assembly were performed with Aquamis v1.3.9 (<a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC8145556/">Deneke et al. 2021</a>) using default parameters. Assembly quality control (QC), including contamination assessment, as well as MLST ST determination were performed with the same pipeline. All genome assemblies passing the QC were included in the final dataset. Among the others, we noticed that a considerable proportion of assemblies was flagged as &ldquo;QC fail&rdquo; exclusively due to the &ldquo;NumContamSNVs&rdquo; parameter, suggesting that this setting might have been too strict. After manual inspection of a random subset, assemblies for which the percentage of reads corresponding to the correct species was &gt;98% were recovered and integrated in the final dataset (those samples are labeled in the Metadata file). In total, 1,540&nbsp;isolates passed the dataset curation step and were included in the final dataset.&nbsp;In-silico serotyping was performed with SeqSero2 v1.2.1 (<a href="https://pubmed.ncbi.nlm.nih.gov/31540993/">Zhang et al. 2019</a>).</p> <p>&nbsp;</p> <p><strong>Funding</strong></p> <p>This work was supported by funding from the European Union&rsquo;s Horizon 2020 Research and Innovation programme under grant agreement No 773830: One Health European Joint Programme.&nbsp;</p> <p>&nbsp;</p> <p><strong>Acknowledgements</strong></p> <p>We thank the National Distributed Computing Infrastructure of Portugal (INCD) for providing the necessary resources to run the genome assemblies. INCD was funded by FCT and FEDER under the project 22153-01/SAICT/2016.</p>

opencc-by-4.0Oct 2022View details →
zenodo44/100

Zonal Statistics of Weather Indicators for Brazilian Municipalities from the TerraClimate Project

<p>This dataset contains 14 parquet-format files with monthly data.</p> <table align="center"> <tbody> <tr> <td>File</td> <td>Indicator</td> <td>Unit</td> </tr> <tr> <td>aet.parquet</td> <td>Actual Evapotranspiration</td> <td>mm</td> </tr> <tr> <td>def.parquet</td> <td>Climate Water Deficit</td> <td>mm</td> </tr> <tr> <td>pdsi.parquet</td> <td>Palmer Drought Severity Index (PDSI)</td> <td>unitless</td> </tr> <tr> <td>pet.parquet</td> <td>Precipitation</td> <td>mm</td> </tr> <tr> <td>ppt.parquet</td> <td>Potential evapotranspiration</td> <td>mm</td> </tr> <tr> <td>q.parquet</td> <td>Runoff</td> <td>mm</td> </tr> <tr> <td>soil.parquet</td> <td>Soil Moisture</td> <td>mm</td> </tr> <tr> <td>srad.parquet</td> <td>Downward surface shortwave radiation</td> <td>W/m2</td> </tr> <tr> <td>swe.parquet</td> <td>Snow water equivalent</td> <td>mm</td> </tr> <tr> <td>tmax.parquet</td> <td>Maximun Temperature</td> <td>&deg;C</td> </tr> <tr> <td>tmin.parquet</td> <td>Minimum Temperature</td> <td>&deg;C</td> </tr> <tr> <td>vap.parquet</td> <td>Vapor pressure</td> <td>kPa</td> </tr> <tr> <td>vpd.parquet</td> <td>Vapor Pressure Deficit</td> <td>kpq</td> </tr> <tr> <td>ws.parquet</td> <td>Wind speed</td> <td>m/s</td> </tr> </tbody> </table> <p>&nbsp;</p>

opencc-by-4.0Apr 2023View details →
zenodo44/100

Validation data set on land cover changes for RapidAI4EO project

<p>This is a reference data set collected for validation of the monthly land cover maps at a 3m and at a 10m resolution produced in the WP5. The reference data set has been collected by using Geo-Wiki toolbox for visual interpretation of very high-resolution images, including Planet data and Google maps. The data set has been collected over 3 AOIs. Each reference sample site corresponds to a 30m-by-30m box and includes information about monthly land cover type over the period 2018-2020. Land cover legend is the same as in ESA WorldCover map at a 10m resolution (https://worldcover2021.esa.int/).</p> <p>Fields:</p> <p>&quot;rowid&quot; &ndash; unique row identifier;</p> <p>&quot;sampleid&quot; &ndash; unique sample site identifier in the Geo-Wiki database;</p> <p>&quot;samplegroupid&quot; &ndash; group id with values 257(Portugal), 258 (Belgium), 259(Sicily);</p> <p>&quot;x_min&quot;,&quot;x_max&quot;,&quot;y_min&quot;,&quot;y_max&quot; &ndash; bounding box coordinates of each sample site (30m x 30m), in WGS84</p> <p>&quot;X2018_1&quot;,&quot;X2018_2&quot;,&hellip;, &quot;X2020_12&quot; &ndash; dominant land cover class in each sample site in each month from January 2018 to December 2020;</p> <p>Land cover codes:</p> <p>10 &ndash; Tree cover</p> <p>20 - Shrubland</p> <p>30 - Grassland</p> <p>40 - Cropland</p> <p>50 &ndash; Urban/built-up</p> <p>60 - Bare/Sparse vegetation</p> <p>80 - Water</p> <p>90 - Wetland</p> <p>110 - Burnt</p> <p>120 &ndash; Not sure</p>

opencc-by-4.0Apr 2023View details →
zenodo44/100

Food production in the context of UrbanGreenUP project for Valladolid City

<p>Production of food in urban orchards (agriculture, eggs, etc.). Measurement of the amount of food produced. The production of food will be measured by tones/Ha per year.</p> <p>Measurement of the amount of food produced. If it cannot be measured, an estimate of the amount generated will be made.</p> <p>Users will be asked directly using surveys.</p>

opencc-by-4.0Mar 2023View details →
zenodo44/100

1000 Genomes Project Transposable Element database

<p>Multi-sample VCF with transposable elements across individuals in the 1KGP dataset. Transposable elements were called using RetroSeq&nbsp;</p>

opencc-by-4.0Apr 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record