Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

28,952

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

28,952 results for “Distributed”

Learn how ShareScore rates datasets ↗
zenodo44/100

Background data 'Effect of biotic dependencies in species distribution models: The future distribution of Thymallus thymallus under consideration of Allogamus auricollis'

<p>Background data of the paper 'Effect of biotic dependencies in species distribution models: The future distribution of Thymallus thymallus under consideration of Allogamus auricollis'</p>

opencc-by-nd-4.0May 2017View details →
zenodo44/100

Biolinks, datasets and algorithms supporting semantic-based distribution and similarity for scientific publications

<p><strong>Background: </strong>Finding articles related to a publication of interest remains a challenge in the Life Sciences domain as the number of scientific publications grows day by day. Publication repositories such as PubMed and Elsevier provides a list of similar articles. There, similarity is commonly calculated based on title, abstract and some keywords assigned to articles. Here we present the datasets and algorithms used in Biolinks. Biolinks uses ontological concepts extracted from publication and makes it possible to calculate a distribution score according to semantic groups as well as a semantic similarity based on either all identified annotations or narrowed to one or more particular semantic groups. Biolinks supports both title and abstract only as well as full-text.</p> <p><strong>Materials: </strong>In a previous work [1], 4,240 articles from the TREC-05 collection [2] were selected. The title-and-abstract for those 4,240 articles were annotated with Unified Medical Language System (UMLS) concepts, such annotations are refer to as our TA-dataset and correspond to the JSON files under the pubmed folder in the JSON-LD.zip file. From those 4,240 articles, full-text was available for only 62. The title-and-abstract annotations for those 62 articles, TAFT-dataset, are located under the pubmed-pmc folder in the JSON-LD.zip file, which also contains the full-text annotations under the folder pmc, FT-dataset. The list corresponding to articles with title-and-abstract is found in the genomics.qrels.large.pubmed.onlyRelevants.titleAndAbstract.tsv file, while those with full-text are recorded in the genomics.qrels.large.pmc.onlyRelevants.fullContent.tsv file.</p> <p>Here we include the annotations on title and abstract as well as those for full-text for all our datasets (profiles.zip). We also provide the global similarity matrices (similarity.zip).</p> <p><strong>Methods:</strong> The TA-dataset was used to calculate the Information Gain (IG) according to the UMLS semantic groups, see IG_umls_groups.PMID.xlsx. A new grouping is proposed for Biolinks, see biolinks_groups.tsv. The IG was calculated for Biolinks groups as well, IG_biolinks_groups.PMID.xlsx, showing a improvement around 5%.</p> <p>In order to assess the similarity metric regarding the cohesion of TREC-05 groups, we used Silhouette Coefficient analyses. An additional dataset Stem-TAFT-dataset was used and compared to TAFT and FT datasets.</p> <p>Biolinks groups were used to calculate a semantic group distribution score for each article in all our datasets. A semantic similarity metric based on PubMed related articles [3] is also provided; the Biolinks groups can be used to narrow the similarity to one or more selected groups. All the corresponding algorithms are open-access and available on GitHub under the license Apache-2.0, a frozen version, biotea-io-parser-master.zip, is provided here. In order to facilitate the analysis of our datasets based on the annotations as well as the distribution and similarity scores, some web-based visualization components were created. All of them open-access and available in GitHub under the license Apache-2.0; frozen versions are provided here, see files biotea-vis-annotation-master.zip, biotea-vis-similarity-master.zip, biotea-vis-tooltip-master.zip and biotea-vis-topicDistribution-master.zip. These components are brought together by biotea-vis-biolinks-master.zip. A demo is provided at http://ljgarcia.github.io/biotea-biolinks/; this demo was built on top of GitHub pages, a frozen version of the gh-pages branch is provided here, see biotea-biolinks-gh-pages.zip.</p> <p><strong>Conclusions: </strong>Biolinks assigns a weight to each semantic group based on the annotations extracted from either title-and-abstract or full-text articles. It also measures similarity for a pair of documents using the semantic information. The distribution and similarity metrics can be narrowed to a subset of the semantic groups, enabling researchers to focus on what is more relevant to them.</p> <p> </p> <p>[1] Garcia Castro, L.J., R. Berlanga, and A. Garcia, <em>In the pursuit of a semantic similarity metric based on UMLS annotations for articles in PubMed Central Open Access.</em> Journal of Biomedical Informatics, 2015. <strong>57</strong>: p. 204-218</p> <p>[2] Text Retrieval Conference 2005 - Genomics Track. <em>TREC-05 Genomics Track ad hoc relevance judgement</em>. 2005  [cited 2016 23rd August]; Available from: http://trec.nist.gov/data/genomics/05/genomics.qrels.large.txt</p> <p>[3] Lin, J. and W.J. Wilbur, <em>PubMed related articles: a probabilistic topic-based model for content similarity.</em> BMC Bioinformatics, 2007. <strong>8</strong>(1): p. 423</p>

opencc-by-4.0Feb 2017View details →
zenodo44/100

Fully differentiable, fully distributed River Discharge Prediction: data sets

<p>This repository contains the data sets used in: Scholz et al. (2025). Fully differentiable, fully distributed River Discharge Prediction.</p> <ul> <li><code>dem_1000.h5</code> based on EU-DEM v1.1, reprojected to RADOLAN grid: <a href="https://sdi.eea.europa.eu/catalogue/srv/api/records/3473589f-0854-4601-919e-2e7dd172ff50">https://sdi.eea.europa.eu/catalogue/srv/api/records/3473589f-0854-4601-919e-2e7dd172ff50</a></li> <li><code>efas.h5</code> based on EFAS historical: <a href="https://ewds.climate.copernicus.eu/datasets/efas-historical?tab=overview">https://ewds.climate.copernicus.eu/datasets/efas-historical?tab=overview</a></li> <li><code>era5_ssrd_neckar*.nc</code> based on ERA5 provided by ECMWF, reprojected to RADOLAN grid:&nbsp;<a href="https://www.ecmwf.int/en/forecasts/dataset/ecmwf-reanalysis-v5">https://www.ecmwf.int/en/forecasts/dataset/ecmwf-reanalysis-v5</a></li> <li><code>radolan_neckar_*.h5</code> based on RADOLAN rw product provided by the Deutsche Wetterdienst: <a href="https://opendata.dwd.de/climate_environment/CDC/grids_germany/hourly/radolan/">https://opendata.dwd.de/climate_environment/CDC/grids_germany/hourly/radolan/</a></li> </ul> <p>Due to copyright, the discharge data has to be downloaded manually from the Global Runoff Data Centre (<a href="https://grdc.bafg.de/">https://grdc.bafg.de/</a>), and then preprocessed with the provided <code>bafg_parser.py</code> python script. We use the following stations in our work:</p> <ul> <li>6335290: STEIN</li> <li>6335291: GAILDORF</li> <li>6335565: BAD IMNAU</li> <li>6335600: ROCKENAU SKA</li> <li>6335601: LAUFFEN</li> <li>6335602: PLOCHINGEN</li> <li>6335603: ROTTWEIL</li> <li>6335604: KIRCHENTELLINSFURT</li> <li>6335620: MOSBACH</li> <li>6335660: PFORZHEIM</li> <li>6335665: DENKENDORF</li> <li>6335671: ALTENSTEIG</li> <li>6335675: MURR</li> <li>6335676: OPPENWEILER</li> <li>6335680: SCHWABSBERG</li> <li>6335681: UNTERGRIESHEIM</li> <li>6335690: NEUSTADT</li> </ul> <p>To preprocess the discharge data, additionally the river network data "Flie&szlig;gew&auml;sser (AWGN)" provided by the Landesanstalt f&uuml;r Umwelt Baden-W&uuml;rttemberg (LUBW) is required:&nbsp;<a href="https://rips-metadaten.lubw.de/trefferanzeige?docuuid=7251515f-6aed-4555-8319-ab6314155ab1">https://rips-metadaten.lubw.de/trefferanzeige?docuuid=7251515f-6aed-4555-8319-ab6314155ab1</a></p> <p>&nbsp;</p>

opencc-by-4.0Oct 2024View details →
zenodo44/100

CLRD-GLPS: A Long-term Seasonal Dataset of Ruminant Livestock Distribution in China's Grazing Production Systems (2000-2021) Using Stacking-based Interpretable Machine Learning

<p>Advanced computational methods integrating ensemble learning with interpretable machine learning are essential for precision livestock management under increasing environmental constraints and food security pressures. This study develops a novel stacking-based interpretable machine learning (IML) framework that combines multiple algorithms with SHAP analysis techniques to generate the China's Long-term Ruminant Livestock Distribution in Grazing Livestock Production Systems (CLRD-GLPS) dataset. Our computational approach addresses critical challenges in livestock distribution modelling: livestock segmentation and spatial prediction accuracy. The framework integrates Random Forest, XGBoost, CatBoost, LightGBM, and Extra Trees through a two-layer stacking architecture, enhanced with SHAP (Shapley Additive Explanations) analysis for model interpretability. We also implemented interpretable machine learning for livestock production system segmentation to distinguish grazing from total livestock populations. The stacking ensemble demonstrated superior performance over individual algorithms, achieving R&sup2; values of 0.954-0.961 for cattle and 0.896-0.901 for sheep and goats, with improvements of up to 8.3% compared to best performance single-model approaches. Multi-scale validation confirmed computational robustness: livestock segmentation achieved R&sup2; = 0.80 at county level, while independent city-level validation of CLRD-GLPS datasets yielded R&sup2; = 0.76-0.80. SHAP interpretability analysis revealed distinct environmental drivers, with vegetation indices and topography primarily influencing cattle distribution, while snow conditions and elevation dominated sheep and goat patterns. This computational framework advances livestock distribution modelling through enhanced prediction accuracy, model stability, and interpretability, while the CLRD-GLPS dataset provides essential spatial-temporal information for rangeland sustainability assessments and evidence-based livestock management policies. This dataset is supported by the Second Tibetan Plateau Scientific Expedition and Research Program (STEP, grant no. 2019QZKK0906).</p>

opencc-by-4.0Nov 2024View details →
zenodo44/100

Mapping the global distribution of C4 vegetation using observations and optimality theory

<p>This dataset includes annual C4 vegetation distribution and its uncertainty from 2001 to 2019. We also provide the distribution of C4 natural grasses and C4 crops during the same period, as well as the code and interim dataset to generate the main figures. Please refer to manuscript for more details:</p> <p>Luo, X., Zhou, H., Satriawan, T.W., Tian, J., Zhao, R., Keenan, T.F., Griffith, D. M., Sitch, S. Smith, N.G. &amp; Still, C.J. (2024). Mapping the global distribution of C4 vegetation using observations and optimality theory.&nbsp;<em>Nature Communications.</em> https://doi.org/10.1038/s41467-024-45606-3.</p> <p><strong>Update (Nov 2023): </strong>we have updated the observational constraint from a linear model to a non-linear model - logistic curve, to better depict how C4 photosynthetic advantage translates into C4 grass coverage changes (C4_distribution_NUS_v2.2.nc).</p> <p><strong>Update (August&nbsp;2023):&nbsp;</strong>we corrected the issue caused by a bias in the remote sensing grassland base map, and released the version 2 of the C4 vmap (C4_distribution_NUS_v2.nc).</p> <p><strong>Update (June 2023):&nbsp;</strong>we noticed there is a critical issue in the version 1 of our C4 map, due to the quality of remote sensing grassland base map used. We are now working on providing a new version (V2) in the next few months (Jun 2023).</p>

opencc-by-4.0Jan 2024View details →
zenodo44/100

An updated modular set of synthetic spectral energy distributions for young stellar objects

<p>These are the models released with the following publication:</p> <p><strong><em>An updated modular set of synthetic spectral energy distributions for young stellar objects</em></strong> (<a href="https://ui.adsabs.harvard.edu/abs/2024ApJ...961..188R/abstract" target="_blank" rel="noopener">Richardson et al. 2024</a>).</p> <p>This is a set of young stellar object (YSO) models with associated spectral energy distributions (SEDs) calculated through radiative transfer. It is a significant update to the data published alongside Robitaille (2017, R17). It contains the parameters shaping each model and adds the newly calculated parameters of envelope mass, average dust temperature, disk stability, and line-of-sight extinction. It also makes explicit quantities, such as source luminosity, that were left implicit in the previous release. This set also convolves the SEDs with several new filters, primarily those on the James Webb Space Telescope, and adds a script to facilitate convolution of these models with additional filters as desired by users. All data included in Version 1.1 of the R17 set (the most recent) are included here.</p> <p>Like their predecessors, these models are versioned. Updates will be released as more models are completed or other changes are made.</p> <p>Files unzip to r+24_models-{version}/{geometry}. "files.tar.gz" contains scripts for SED convolution and main sequence comparison, the opacity to absorption of dust used in the radiative transfer calculations, main sequence T/L values used for results in the accompanying work, and reference material for the contents of the dataset and latest version.</p> <p>The primary use of these models is as templates for SED fitting. The R17 models were structured for use with the <a href="https://sedfitter.readthedocs.io/en/stable/" target="_blank" rel="noopener">sedfitter</a> python package, which enables fitting and analysis of the fit results. For a version of sedfitter which accommodates the new additions, use&nbsp;<a href="https://github.com/richardson-t/sedfitter/tree/dev" target="_blank" rel="noopener">this fork</a>.</p>

opencc-by-4.0Sep 2023View details →
zenodo44/100

Adult baobab trees's distribution map across Sahel

<p><span>The baobab tree (<em>Adansonia digitata</em> <em>L.</em>) is an integral part of rural livelihoods throughout the African continent. However, the combined effects of climate change and increasing global demand for baobab products are currently exerting pressure on the sustainable utilization of these resources. Here we employ sub-meter resolution satellite imagery to identify nearly 3</span><span>&nbsp;million baobab trees in the Sahel, a dryland region of 1.5 million km<sup>2</sup>. This achievement is considered an essential step towards improving valuable woody species' management and monitoring system. To prevent mismanagement of this specific tree species, we aggregated every single adult baobab tree map to<span>&nbsp;5 <span>&times; </span>5 km grids. We also classified the baobab trees using the tree crown diameters( small: 3-9m; medium 9m-13m; large: &gt;13m).&nbsp; The baobab tree count map is also available for this three different size classes.</span></span></p>

opencc-by-4.0Jan 2024View details →
zenodo44/100

Distributable, Metabolic PET Reporting of Tuberculosis

<p>Tuberculosis remains a large global disease burden for which treatment regimens are protracted and monitoring of disease activity difficult. Existing detection methods rely almost exclusively on bacterial culture from sputum which limits sampling to &nbsp;organisms on the pulmonary surface. Advances in monitoring tuberculous lesions have utilized the common glucoside [18 F]FDG, yet lack specificity to the causative pathogen Mycobacterium tuberculosis (Mtb) and so do not directly correlate with &nbsp;pathogen viability. Here we show that a close mimic that is also positron-emitting of the non-mammalian Mtb disaccharide trehalose &ndash; 2-[ 18 F]fluoro-2-deoxytrehalose ([18 F]FDT) &ndash; is a mechanism-based reporter of Mycobacteria-selective enzyme activity in vivo. Use of [18 F]FDT in the imaging of Mtb in diverse models of disease, &nbsp;including non-human primates, successfully co-opts Mtb-specific processing of trehalose to allow the specific imaging of TB-associated lesions and to monitor the effects of treatment. A pyrogen-free, direct enzyme-catalyzed process for its radiochemical synthesis allows the ready production of [18 F]FDT from the most globally-abundant organic 18F-containing molecule, [18 F]FDG. The full, pre-clinical validation of both production method and [18 F]FDT now creates a new, bacterium selective, clinical diagnostic candidate for clinical evaluation. We anticipate that this distributable technology to generate clinical-grade [18 F]FDT directly from the widelyavailable clinical reagent [18 F]FDG, without need for either custom-made radioisotope generation or specialist chemical methods and/or facilities, could now usher in global, democratized access to a TB-specific PET tracer.</p>

opencc-by-4.0Apr 2024View details →
zenodo44/100

Supplementary materials for "Experimental verification for self-organization process on the spatial distribution and edifice size of rootless cone"

<p>This is a ReadMe for the supplementary material for "Experimental verification for self-organization process on the spatial distribution and edifice size of rootless cone" written by Rina Noguchi and Wataru Nakagawa.</p> <p>----------------------<br>[ReadMe.txt]<br>ReadMe text file.</p> <p>[FigS1.png]<br>This figure is a supplementary figure which appeared as "Figure S1" in the main text.<br>Caption: Figure S1. &nbsp;Examples of conduits (dashed green lines) and loser conduits (solid magenta lines) were observed in the experiments with original and contrast-enhanced images.</p> <p>[FigS2.png]<br>This figure is a supplementary figure which appeared as "Figure S2" in the main text.<br>Caption: Figure S2. &nbsp;Relationships between the thickness of poured heated syrup and (A) mass losses caused by baking soda decomposition, (B) number of conduits, (C) total conduit area, (D) average conduit area, (E) number of failed conduits, and (F) sum number of conduits and failed conduits. Each plot and error bar represents the average and standard deviation in three repeated experiments, respectively. The red plots and error bars show the 350 g of heated syrup case, which performed ten repeated experiments to verify the reproducibility. Note that horizontal error bars are derived from the difficulty of strict heated syrup-pouring control.</p> <p>[Experimental_datasheet.xlsx]<br>This EXCEL file includes two sheets: a mass loss change log and a summary of experimental results.</p> <p>[movie/SSS_X_x15.mp4]<br>These MP4 files are fast-forward movies (x15) for each experiment. SSS = the amount of poured hearty syrup (g), and X = round in each condition.<br>----------------------</p> <p>For more details, please refer to a research paper "Experimental verification for self-organization process on the spatial distribution and edifice size of rootless cone".</p> <p>If you have any questions, please send an e-mail to:<br>r-noguchi@env.sc.niigata-u.ac.jp<br>or<br>flugel555@gmail.com<br>.<br>(R. Noguchi)</p>

opencc-by-4.0Jul 2024View details →
zenodo44/100

Distributional data for Lined Seedeaters (Sporophila lineola) and Saffron Finches (Sicalis flaveola)

<p><span>Distributional data for Lined Seedeaters (Sporophila lineola) and Saffron Finches (Sicalis flaveola). Data compiled from the literature, zoological collections, and community science platforms.&nbsp;</span></p>

opencc-by-4.0Oct 2024View details →
zenodo44/100

Distribution and Characteristics of Lightning-Ignited Wildfires in Boreal Forests - the BoLtFire database

<p>This repository holds a dataset of lightning-ignited wildfires across the boreal biome. The BoLtFire dataset covers the period 2012 to 2022 and encompasses 6,902 fires - 4,201 in Eurasia and 2,701 in North America.</p> <p>The layers included in this dataset are: FireID, StartDate, EndDate, FireYdear, AreaHa (burned area), ClassSize, BiomeName, EcoBiome, EcoName, EcoID, Realm, LCDN (Land cover number), LCName (land cover name), Country, Continent, HoldoverD (days), HoldoverRD (holdover rounded), IgnLat (Ignition location Latitude), IgnLong (Ignition Location Longitude), DisPol (Distance of the ignition location to the fire perimeter if it is located outside the polygon), and PerCheck (designates if the ignition location is within the fire perimeter or oustide the perimeter).</p> <p>&nbsp;</p> <p>The datasets are available per continent (North America, Europe, and Asia) as shapefiles. The spatial reference system is Global LANd Cover mapping and Estimation (GLANCE) Grids - Version 01 CRS.</p> <p>&nbsp;</p> <p>*Please note: Versions 1 and 2 are missing LIW from Canada between 2021-2022.&nbsp;</p>

opencc-by-4.0Oct 2024View details →
zenodo44/100

Global patterns of soil organic carbon distribution in the 20–100 cm soil profile for different ecosystems: A global meta-analysis

<p><span><span>&nbsp;</span></span><span>The file named <span>&ldquo;</span>Rawdata.xlsx<span>&rdquo;</span> contains data sourced from the literature.<span> The file name is &ldquo;GE_&beta;.tif<span>&rdquo;</span><span>,</span></span></span><span><span> GE represents</span></span><span> global ecosystems, which including cropland (CL), grassland (GL), and forestland (FL). &ldquo;FL_&beta;.tif&rdquo; represents the spatial distribution of &beta; for forestland at 20-100 cm depth. The file name is &ldquo;GE_d_SOCD.tif&rdquo;, where SOCD represents soil organic carbon density, d represents soil depth, for example, &ldquo;FL_20-100_SOCD.tif&rdquo; represents the spatial distribution of SOCD for forestland at 20-100 cm depth.</span></p>

opencc-by-4.0Mar 2024View details →
zenodo44/100

Distributed Predictive Drone Swarms in Cluttered Environments

<p>This folder contains data, videos, and supplementary material for the article titled &quot;Distributed Predictive Drone Swarms in Cluttered Environments&quot;.</p> <p>In the article, we present a Distributed Model Predictive Control (DMPC) algorithm for drone swarm navigation in two types of cluttered environments, i.e., a forest and a funnel-like environment.&nbsp;<br> &nbsp;</p> <p>The material in `zenodo_upload` is organized as follows.<br> 1. a data folder, with the logs of simulation and hardware experiments;<br> 2. an analysis folder, with Matlab scripts that analyze the logs in the data folder;<br> 3. a plotting folder, with Matlab functions used by the analysis scripts;<br> 4. an mp4 video file, on simulation and hardware experiments;<br> 5. a pdf, with supplementary materials.<br> <br> The `qp_swarm` folder contains MATLAB code for simulation experiments.</p>

opencc-by-4.0Aug 2021View details →
zenodo44/100

Data related to publication "Coherent phase transfer for real-world twin-field quantum key distribution; Supplementary Information"

<p>These files contains datasets from which the Figures appearing in the Supplementary Information have been calculated.&nbsp;</p> <p>Description of datasets:</p> <p>Datasets related to SupplFig1 contain two columns: Frequency in Hz and phase noise in rad^2/Hz</p> <p>Data_SupplFig1_stabilised_fringes: psd of the phase noise calculated from the interference fringes in a stabilised condition</p> <p>Data_SupplFig1_unstabilised_fringes: psd of the phase noise calculated from the interference fringes in an unstabilised condition</p> <p>Data_SupplFig1_roundtrip_sensing_laser: psd of the sensing laser signal after a round-trip in the interferometer, calculated&nbsp;from self-heterodyne beatnote</p> <p>Data_SupplFig1_differential_roundtrip_sensing_vs_reference_laser: psd of the difference between the round-trip self-heterodyne beatnotes at the sensing and reference laser wavelengths</p> <p>Datasets related to SupplFig2 contain two columns: time in seconds and normalised intensity (calculated as detailed in the main publication).</p> <p>Data_SupplFig2_High_power_PD_free_evol: normalised intensity of the interference signal&nbsp;obtained with classical power level at the source. This trace was recorded with&nbsp;a photodiode when no artificial phase drift was applied</p> <p>Data_SupplFig2_High_power_PD_phase_drift:&nbsp; normalised intensity of the interference signal&nbsp;obtained with classical power level at the source. This trace was recorded with&nbsp;a photodiode when an artificial phase drift was applied (8pi/s)</p> <p>Data_SupplFig2_High_power_SPD_free_evol:&nbsp;normalised intensity of the interference signal&nbsp;obtained with classical power level at the source. This trace was recorded on an SPD (after suitable attenuation) when no&nbsp;artificial phase drift was applied&nbsp;</p> <p>Data_SupplFig2_High_power_SPD_phase_drift:&nbsp;normalised intensity of the interference signal&nbsp;obtained with classical power level at the source. This trace was recorded on an SPD (after suitable attenuation) when an artificial phase drift was applied (8pi/s)</p> <p>Data_SupplFig2_Attenuated_SPD_free_evol:&nbsp;normalised intensity of the interference signal&nbsp;obtained with attenuated beams at the source. This trace was recorded on an SPD when no&nbsp;artificial phase drift was applied&nbsp;</p> <p>Data_SupplFig2_Attenuated_SPD_phase_drift:&nbsp;:&nbsp;normalised intensity of the interference signal&nbsp;obtained with attenuated beams at the source. This trace was recorded on an SPD when an artificial phase drift was applied (8pi/s)</p> <p>&nbsp;</p>

opencc-by-4.0Nov 2021View details →
zenodo44/100

Updated distribution and conservation perspectives of marmosine opossums from Colombia

<p>These maps are the results of ecological niche modeling (via MaxEnt) and expert&#39;s opinions. Models are based on localities from recent taxonomic reviews, catalogues, museum specimens, and curated GBIF data. Models where tuned specifically for each species and evaluated in different modeling areas (see original publication for details). These maps represent the potential distribution ranges of the Marmosini species of Colombia, but were individually adjusted based on biogeographic barriers (see article for details).</p>

opencc-by-4.0May 2021View details →
zenodo44/100

Supplementary data for publication Global distribution of mcr gene variants in 214K metagenomic samples

<p># Supplementary data for the manuscript &quot;Global distribution of mcr gene variants in 214,095 metagenomic samples&quot;</p> <p>SD1_mapped_runids.csv : tab-separated file with columns of run_accessions downloaded from ENA and whether the metagenome were positive for at least one of the mcr genes.</p> <p>SD2_mcr_df.csv : compositional table of mcr-positive metagenomes with associated metadata (collection_year, country, and host) for each run_accession, as well as mapping results.</p> <p>SD3_mcr_contigs.fa : FASTA file with contigs carrying mcr genes. The header contains the run_accession ID.</p> <p>SD4_aldex2_results.csv: CSV file containing ALDEx2 results. The columns are as follows:<br> * group: metadata category (year, country or host). If the column contains more than one label, e.g., &quot;Denmark - 2020 - Pigs&quot;, significance is tested within Danish pig samples from 2020.<br> * rab.all:&nbsp; median clr value for all samples in the feature<br> * rab.win.conditionA:&nbsp; median clr value for the condition A of samples<br> * rab.win.conditionB: median clr value for the condition B of samples<br> * diff.btw: median difference in clr values between A and B conditions<br> * diff.win: median of the largest difference in clr values within A and B conditions<br> * effect : median effect size: diff.btw / max(diff.win) for all instances<br> * overlap : proportion of effect size that overlaps 0 (i.e. no effect)<br> * we.ep: Expected P value of Welch&rsquo;s t test<br> * we.eBH: Expected Benjamini-Hochberg corrected P value of Welch&rsquo;s t test<br> * wi.ep: Expected P value of Wilcoxon rank test<br> * wi.eBH: Expected Benjamini-Hochberg corrected P value of Wilcoxon test<br> * parts: gene name<br> * conditionA: label of condition A that is compared against condition B<br> * conditionB: label of condition B that is compared against condition A<br> * conditions.A.vs.B: label to explain condition A compared against condition B<br> NOTE: see for more explanation of the output of ALDEx2 https://www.bioconductor.org/packages/release/bioc/vignettes/ALDEx2/inst/doc/ALDEx2_vignette.html#5_ALDEx2_outputs</p> <p>SD5: Multi-VCF file containing SNP information on mcr alleles. Can be used to construct consensus sequences.</p> <p>SD6: FASTA file containing all unique consensus sequences reported in the manuscript.</p> <p>SD7: CSV file with an overview of which metagenome contains which unique consensus sequence.</p>

opencc-by-4.0Feb 2022View details →
zenodo44/100

Data set for the article "Tides, topography, and seagrass cover controls on the spatial distribution of Pinna nobilis on a coastal lagoon tidal flat"

<p>Data set includes:&nbsp;coordinates of the GNSS points (reference system WGS84 UTM33N);&nbsp;density of P. nobilis&nbsp;and cover of C.nodosa&nbsp;detected in the orthophoto in the 25m<sup>2</sup> cells;&nbsp;tidal levels measured (and, for comparison, simulated with the hydrodynamic model) corrected with respect to the IGM datum; number of emersions and flood duration for different levels of the tidal flat; statistics. The first Excel sheet includes a detailed description of the data.</p>

opencc-by-4.0May 2021View details →
zenodo44/100

Supporting data for review article: The Global Distribution, Formation, and Fate of Mineral-Associated Soil Organic Matter Under a Changing Climate – A Trait-Based Perspective

<p>Supporting data and code for review article: Sokol N.W., Whalen E.D., Kallenbach C., Pett-Ridge J., Georgiou K.&nbsp;The Global Distribution, Formation, and Fate of Mineral-Associated Soil Organic Matter Under a Changing Climate &ndash;&nbsp;A Trait-Based Perspective. <em>Functional Ecology,&nbsp;</em>2022.</p> <p>We leveraged data from a global synthesis of&nbsp;soil fractionation measurements&nbsp;(DOI: 10.5281/zenodo.5987415). For this review article, we specifically focused on measurements of bulk and mineral-associated soil organic carbon concentrations (reported in units of gC/kg soil) and the proportion of bulk soil organic carbon that is mineral-associated (reported as a %). This subset&nbsp;also includes auxiliary data regarding climate and biome characteristics extracted from the synthesized papers; for more variables, see the original full dataset. K&ouml;ppen-Geiger climate zones were extracted from a georeferenced global database (using R package &#39;kgc&#39; v1.0.0.2) with site coordinates, where available.&nbsp;Three files are provided in this repository: (1) data file, (2) metadata file, and (3) code for manuscript figures and summary statistics.</p>

opencc-by-4.0Dec 2021View details →
zenodo44/100

PrevDistro - Preverb Distributions in Hungarian

<p>PrevDistro (Preverb Distributions) is an open-source dataset containing 41.5 million corpus occurrences of 49 preverb-verb construction types. It consists of the following columns:</p> <ul> <li>1 <em>sid</em>: ID</li> <li>2 <em>constype</em>: construction type</li> <li>3 <em>subtype</em>: construction subtype</li> <li>4 <em>prevpos</em>: preverb position</li> <li>5 <em>prev</em>: preverb</li> <li>6 <em>verb</em>: verb lemma</li> <li>7 <em>intervening</em>: intervening words (as lemmas)</li> <li>8 <em>actform</em>: actual form (the same content as in column 10, but this column is lowercase)</li> <li>9 <em>left</em>: left context</li> <li>10 <em>kwic</em>: keyword in context</li> <li>11 <em>right</em>: right context</li> <li>12 <em>docid</em>: document ID from the Hungarian Gigaword Corpus</li> <li>13 <em>title</em>: document title</li> <li>14 <em>style</em>: document style (e.g. official, press, ...)</li> <li>15 <em>region</em>: document region (e.g. Transylvania, Subcarpathia, ...)</li> <li>16 <em>year</em>: year of publication (sometimes several years can be found in one document)</li> </ul> <p>The first row stands for the header. If a cell&#39;s value is unspecified, it is marked with underscore (_).</p>

opengpl-3.0-or-laterJun 2021View details →
zenodo44/100

Crop classification dataset for testing domain adaptation or distributional shift methods

<p>In this upload we share processed crop type datasets from both France and Kenya. These datasets can be helpful for testing and comparing various domain adaptation methods. The datasets are processed,&nbsp;used, and described&nbsp;in this paper:&nbsp;<a href="https://doi.org/10.1016/j.rse.2021.112488">https://doi.org/10.1016/j.rse.2021.112488</a>&nbsp;(arXiv version: <a href="https://arxiv.org/pdf/2109.01246.pdf">https://arxiv.org/pdf/2109.01246.pdf</a>).&nbsp;</p> <p>In summary, each point in the uploaded datasets corresponds to a particular location. The label&nbsp;is the crop type grown at that location in 2017.&nbsp;The 70 processed features are based on&nbsp;Sentinel-2 satellite measurements at that location in 2017. The points in the France dataset come from 11 different departments (regions) in Occitanie, France, and the points in the Kenya dataset come from 3 different regions in Western Province, Kenya. Within each dataset there&nbsp;are&nbsp;notable shifts in the distribution of the labels and in the distribution of the features between regions. Therefore, these datasets can be helpful for testing&nbsp;for testing and comparing methods that are designed to address such distributional shifts.</p> <p>More details on the dataset and processing steps can be found in&nbsp;<a href="https://doi.org/10.1016/j.rse.2021.112488">Kluger et. al. (2021)</a>. Much of the&nbsp;processing steps were taken to deal with Sentinel-2 measurements that were corrupted by cloud cover. For users interested in the raw multi-spectral time series data and dealing with cloud cover issues on their own (rather than using the 70 processed features provided here), the raw dataset from Kenya can be found in <a href="https://openreview.net/forum?id=5HR3vCylqD">Yeh et. al. (2021)</a>, and the raw dataset from France can be made available upon request from the authors of this Zenodo upload.</p> <p>All of the data uploaded here can be found in &quot;CropTypeDatasetProcessed.RData&quot;. We also post the dataframes and tables within that .RData file&nbsp;as separate .csv&nbsp;files for users who do not have R. The contents of each R object (or&nbsp;.csv file) is described in the file &quot;Metadata.rtf&quot;.</p> <p><strong>Preferred Citation:</strong></p> <p>-Kluger, D.M., Wang, S., Lobell, D.B., 2021. Two shifts for crop mapping: Leveraging aggregate crop statistics to improve satellite-based maps in new regions. Remote Sens. Environ. 262, 112488. https://doi.org/10.1016/j.rse.2021.112488.</p> <p>-URL to this Zenodo post https://zenodo.org/record/6376160</p>

opencc-by-4.0Mar 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record