Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

76,402,788

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

76,402,788 results

Learn how ShareScore rates datasets ↗
zenodo52/100

Dataset and program scripts for the reproducibility of the hierarchical data structure file. Related to the manuscript entitled: Hierarchical Representation of Measurement Data, Metrological Uncertainty and Metadata for Calibrated Battery Tests

<p>We present an interoperable hierarchical data representation for battery tests, leading to improved scalability of data transmission and enhanced data accessibility and comprehensibility for both human interpretation and machine processing. The hierarchical data format includes the raw trace electrical measurement data, the metrological calibration and uncertainty data, the metadata such as experimental settings, instruments and software versions, as well as post-processed data such as electrochemical model fit parameters. This data representation allows repetition of the battery test under the exact same conditions such that identical results are achieved within defined error bounds. This is in line with the general F.A.I.R. data approach and provides repeatability and traceability in the battery value chain. As an application of the hierarchical data representation, we show the classification of cells as pass/fail being performed with quantitative confidence levels. We demonstrate the complete workflow of establishing the hierarchical data structure for electrochemical impedance spectroscopy (EIS), starting from metrological traceability of the calibration and uncertainty analysis towards the storage of the structured data as a single integrated file that preserves the hierarchical data format.</p>

openmit-licenseNov 2023View details →
zenodo52/100

UAV time series and tree crowns

<p>This dataset contains:</p><p>-A UAV time series of mosaicked images of a woodland in Northeast UK. Complete detaisl are given in: "Elias Fernando Berra, Rachel Gaulton, Stuart Barr, Assessing spring phenology of a temperate woodland: A multiscale comparison of ground, unmanned aerial vehicle and Landsat satellite observations, Remote Sensing of Environment, Volume 223, 2019, Pages 229-242, ISSN 0034-4257, https://doi.org/10.1016/j.rse.2019.01.010."&nbsp;</p><p>-Manual (reference) and automatic delinetaed tree crowns for the area covered by the UAV time series data. Complete details in: Elias F. Berra. Individual tree crown detection and delineation across a woodland using leaf-on and leaf-off imagery from a UAV consumer-grade camera. Journal of Applied Remote Sensing, Vol. 14, Issue 3, 034501 (July 2020). https://doi.org/10.1117/1.JRS.14.034501</p>

opencc-by-4.0Dec 2023View details →
zenodo52/100

Raw data for "Development and characterization of a non-human primate model of disseminated synucleinopathy"

<p><span>In this study, the performance and biodistribution of the retrogradely-spreading AAV9-SynA53T vector was evaluated in the NHP brain. Conducted intraparenchymal deliveries of viral suspensions in the left putamen gave rise to a disseminated synucleinopathy in a circuit-specific basis.</span></p>

opencc-by-4.0Dec 2023View details →
zenodo52/100

XPS spectra for Li intercalated few-layer MoS2 films

<h3>Description</h3> <p>Dataset for synchrotron-based x-ray photoelectron (XPS) spectra of Li doped MoS<sub>2</sub> nanofilms grown on c-plane sapphire substrate by two techniques: thermally assisted conversion (TAC) and pulsed laser deposition (PLD). Reference spectra for undoped MoS<sub>2</sub> were measured on a commercial MoS<sub>2</sub> powder sample.</p> <p>&nbsp;</p> <p><strong> Data formats</strong></p> <p>The XPS datasets are available in two formats.</p> <ol> <li><a href="https://doi.org/10.1002/sia.740130202">VAMAS</a> (ASCII ISO 14976).</li> <li>Plain text column files: Double-column files (dat) and corresponding metadata files (txt).</li> </ol>

opencc-by-4.0Dec 2023View details →
zenodo52/100

Edgar Anderson's Iris Data

<p>This famous (Fisher's or Anderson's) iris data set gives the measurements in centimeters of the variables sepal length and width and petal length and width, respectively, for 50 flowers from each of 3 species of iris. The species are Iris setosa, versicolor, and virginica.</p> <p>The <em>iris_dataset.rds</em> serialisation is a replication of datasets::iris_dataset as dataset s3 class.</p> <p>The <em>iris_dataset.csv </em>serialisation is an incomplete replication of the iris_dataset because the CSV file does not contain important semantic information; that is exported to <em>iris_dataset.json</em> (in a not standardised form) and the dataset-level metadata into the <em>iris_dataset.bib </em>BibLatex text file.</p>

opencc-by-4.0Jul 2018View details →
zenodo52/100

Supplementary Data to journal publication on 'The Foundations of the Patagonian Icefields'

<p>Partitioning and comparison of ice discharge estimates from the the Patagonian Icefields comprising associated uncertainties. For further details please refer to the notes in the individual files and/or consult the associated publication entitled 'The Foundations of the Patagonian Icefields' published in Communications Earth &amp; Environment.</p>

opencc-by-4.0Dec 2023View details →
zenodo52/100

Copper mineralization at Carajás mineral province - Brazil: geological, structural, and geophysical data

<p>Gridded geological, structural, and geophysical data at the Caraj&aacute;s mineral province. A number of known Cu occurrences are provided. This dataset is suitable for experimenting with machine learning methods.</p>

opencc-by-4.0Dec 2023View details →
zenodo52/100

ICON-LEM Ny-Ålesund low-level clouds polar night and polar day 2021/2022

<h3>Low-level clouds during the polar night and polar day simulated in ICON-LEM for Ny-&Aring;lesund&nbsp;</h3> <p>This data set was created using the ICON-LEM model with ca. 600m resolution and a diagnostic tool "microphysical wrapper". It contains the meteogram output of the Ny-&Aring;lesund column (Svalbard) and the microphysical process rates. The data was created for the polar night (Nov 2021- Feb 2022) and polar day (May - Aug 2022).&nbsp;Clouds are classified as low-level if their cloud top height (CTH) is below 2.5 km and the distance between any cloud with CTH above 2.5 km is at least 500 m higher. The data set was first used and described in the <em>publication:&nbsp;</em></p> <p>T. Kiszler, D. Ori, V. Schemann<em>. </em>(preprint) Microphysical processes involving the vapour phase dominate in simulated low-level Arctic clouds. <em>Atmospheric Physics and Chemistry, </em>https://doi.org/10.5194/egusphere-2023-2986<em><br></em></p> <p>This data is related to the repository <a href="https://github.com/TracyMcBean/Kiszler_et_al_2023_microphysics">https://github.com/TracyMcBean/Kiszler_et_al_2023_microphysics</a></p> <p><em>File description:</em></p> <p>*_PN is polar night data</p> <p>*_PD is polar day data</p> <p>LLC_<em>meteo_&lt;yyyymm&gt;_ICONv1</em>_v6.nc : Contains the meteogram variables (thermodynamics, surface variables, hydrometeors)</p> <p>LLC_wrapper_mass_&lt;yyyymm&gt;_ICONv1_v6.nc : Contains hydrometeors masses after diagnostic run of a microphysical wrapper</p> <p>LLC_wrapper_tend_&lt;yyyymm&gt;_ICONv1_v6.nc : Contains the mircophysical process rates showing the mass change per timestep&nbsp;</p> <p>low_cloud_times_v6_*.csv : Contains the date and time when a low-level cloud was detected</p>

opencc-by-4.0Dec 2023View details →
zenodo52/100

rename

<p>This dataset encompasses a comprehensive simulation of 1600 years of daily renewable electricity production and demand.It was generated with the use of a large ensemble approach, using 160 sets of 10-year climate model simulations (1600 years) (Muntjewerf et al., 2023)&nbsp;---each set representing a different possible sequence of weather under present-day climate conditions--- in combination with and energy production and demand modelling framework (van der Most et al., 2022).</p> <p>If you use this dataset in your ressearch, we kindly request that you cite the following paper:&nbsp;</p>

opencc-by-4.0Dec 2023View details →
zenodo52/100

ChinaHighNO₂: Daily Seamless 10 km Ground-Level NO₂ Dataset for China (2008–2018)

<p>ChinaHighNO<sub>2</sub>&nbsp;is part of a series of long-term, seamless, high-resolution, and high-quality datasets of air pollutants for China (i.e., ChinaHighAirPollutants, CHAP). It is generated from big data sources (e.g., ground-based measurements, satellite remote sensing products, atmospheric reanalysis, and model simulations) using artificial intelligence, taking into account the spatiotemporal heterogeneity of air pollution.</p> <p>Here is the big data-derived seamless (spatial coverage = 100%) daily, monthly, and yearly 10 km (i.e., D10K, M10K, and Y10K) ground-level NO<sub>2</sub>&nbsp;dataset for China&nbsp;<strong>from 2008 to 2018</strong>. This dataset exhibits high quality, with a cross-validation coefficient of determination (CV-R<sup>2</sup>) of 0.84, a root-mean-square error (RMSE) of 7.99 &micro;g m<sup>-3</sup>, and a mean absolute error (MAE) of 5.34 &micro;g m<sup>-3</sup> on a daily basis.</p> <p>If you use the ChinaHighNO<sub>2</sub> dataset in your scientific research, please cite the following references (Wei et al., ACP, 2023; Wei et al., EST, 2022):</p> <ul> <li> <p>Wei, J., Li, Z., Wang, J., Li, C., Gupta, P., and Cribb, M.&nbsp;<a href="https://weijing-rs.github.io/publications/Wei_et_al-ACP-2023.pdf">Ground-level gaseous pollutants (NO<sub>2</sub>, SO<sub>2</sub>, and CO) in China: daily seamless mapping and spatiotemporal variations</a>.&nbsp;<em>Atmospheric Chemistry and Physics</em>, 2023, 23, 1511&ndash;1532. https://doi.org/10.5194/acp-23-1511-2023</p> </li> <li> <p>Wei, J., Liu, S., Li, Z., Liu, C., Qin, K., Liu, X., Pinker, R., Dickerson, R., Lin, J., Boersma, K., Sun, L., Li, R., Xue, W., Cui, Y., Zhang, C., and Wang, J.&nbsp;<a href="https://weijing-rs.github.io/publications/Wei_et_al-EST-2022.pdf">Ground-level NO<sub>2</sub>&nbsp;surveillance from space across China for high resolution using interpretable spatiotemporally weighted artificial intelligence</a>.&nbsp;<em>Environmental Science &amp; Technology</em>, 2022, 56(14), 9988&ndash;9998. https://doi.org/10.1021/acs.est.2c03834</p> </li> </ul> <p><strong>Note that the ChinaHighNO<sub>2</sub> dataset&nbsp;was improved to a 1 km resolution after 2019:</strong></p> <p>&nbsp; &nbsp; &nbsp; &nbsp; all (including&nbsp;<strong>daily</strong>) data for the years after <strong>2019</strong><strong> </strong>are accessible at: <strong><a href="https://doi.org/10.5281/zenodo.4571660">https://doi.org/10.5281/zenodo.4571660</a></strong></p> <p><strong>More CHAP datasets for different air pollutants are available at: <a href="https://weijing-rs.github.io/product.html">https://weijing-rs.github.io/product.html</a></strong></p>

opencc-by-4.0Mar 2021View details →
zenodo52/100

TemStaPro Datasets

<p>This dataset contains protein sequences used to train, validate, and test binary classifiers that form TemStaPro program, which is applied&nbsp;for protein thermostability prediction with respect to nine&nbsp;temperature thresholds from&nbsp;40 to 80&nbsp;degrees Celsius using a step of five&nbsp;degrees.</p> <p>The data is given&nbsp;in files of FASTA format. Each protein sequence has a header made of three&nbsp;values separated by vertical bar&nbsp;symbols:&nbsp;organism's, to which the protein belongs, UniParc taxonomy identifier;&nbsp;UniProtKB/TrEMBL identifier of the protein sequence;&nbsp;organism's growth temperature taken from the dataset of growth temperatures of over 21 thousand organisms&nbsp;(Engqvist, 2018).</p> <p>TemStaPro-Major-30 set is composed of 12 files:</p> <ul> <li>one training</li> <li>one validation</li> <li>one imbalanced testing</li> <li>nine&nbsp;balanced samples&nbsp;of 2000 sequences from each of the balanced testing set</li> </ul> <p>TemStaPro-Minor-30 set is composed of cross-validation and testing files all balanced for 65 degrees Celsius temperature threshold.</p> <p>SupplementaryFileC2EPsPredictions.tsv file contains thermostability predictions using the default mode of TemStaPro program&nbsp;to check the thermostability of different C2EP groups.<br><br>The detailed description is given in the revised version of the corresponding paper (https://doi.org/10.1093/bioinformatics/btae157).</p> <p>If you use the data from this dataset, please cite both the paper and the DOI of the&nbsp;dataset.</p>

opencc-by-4.0Mar 2023View details →
zenodo52/100

Occurrence cubes for non-native taxa in Belgium and Europe

<p>This package contains aggregated occurrence data ("occurrence cubes") for non-native taxa in Belgium and Europe. These occurrence cubes were generated by grouping species occurrence data from the <a href="https://www.gbif.org/">Global Biodiversity Information Facility (GBIF)</a> by year (year), 1x1km spatial <a href="https://www.eea.europa.eu/en/datahub/datahubitem-view/3c362237-daa4-45e2-8c16-aaadfb1a003b">EEA reference grid</a> cell (eea_cell_code) and taxon (taxonKey or classKey). For each grouping, the number of occurrences found in GBIF (n) and the minimum <a href="http://rs.tdwg.org/dwc/terms/coordinateUncertaintyInMeters">coordinateUncertaintyInMeters</a> (min_coord_uncertainty) are provided. The provided coordinateUncertaintyInMeters of an occurrence is taken into account when assigning it to a grid cell (see <a href="https://github.com/trias-project/occ-cube-alien/blob/20201201/src/europe/2_assign_grid.Rmd#L463-L481">this code</a>). The occurrence cubes have been&nbsp;used as input data for indicators and risk modelling/mapping for the <a href="http://trias-project.be/">Tracking Invasive Alien Species (TrIAS)</a> project and are now used for monitoring the effectiveness of the early detection and rapid eradication of emerging Invasive Alien Species (IAS) for the <a href="https://www.riparias.be/">LIFE RIPARIAS</a> project.</p> <p>The occurrence cubes are built on open science principles and intended to be completely reproducible:</p> <ul> <li>The input data are publicly available on GBIF, with the download DOIs listed in the related identifiers of this package.</li> <li>The code to process the data to cubes is publicly available on GitHub at <a href="https://github.com/trias-project/occ-cube-alien">https://github.com/trias-project/occ-cube-alien</a> (version <a href="https://github.com/trias-project/occ-cube-alien/releases/tag/20240118">20240118</a>).</li> </ul> <h2>Files</h2> <ul> <li><strong>be_alientaxa_cube.csv</strong>: occurrence cube of alien taxa listed by the Global Register of Introduced and Invasive Species - Belgium (Desmet et al. 2019) (GRIIS) and limited to occurrences in Belgium (country=BE).</li> <li><strong>be_alientaxa_info.csv</strong>: taxonomic information for taxa in be_alientaxa_cube.csv.</li> <li><strong>be_classes_cube.csv</strong>: occurrence cube of all <a href="http://rs.tdwg.org/dwc/terms/class">classes</a> found in Belgium (country=BE), used to assess sampling effort bias in be_alientaxa_cube.csv.</li> <li><strong>eu_modellingtaxa_cube.csv</strong>: occurrence cube of <a href="https://github.com/trias-project/occ-cube-alien/blob/2ada0ded33c034946380b02a28cb9a8d2884d54a/references/modelling_species.tsv">selected modelling species</a> in Europe (bounding box).</li> <li><strong>eu_modellingtaxa_info.csv</strong>: taxonomic information for taxa in eu_modellingtaxa_cube.csv.</li> </ul> <h2>Acknowledgements</h2> <p>This work has been funded under the Belgian Science Policies Brain program (BelSPO BR/165/A1/TrIAS), the European Union's LIFE program (LIFE19 NAT/BE/000953 - LIFE RIPARIAS) and the European Union's Horizon Europe Research and Innovation Programme (ID No 101059592 - Biodiversity Building Blocks for Policy).</p>

opencc-zeroOct 2019View details →
zenodo52/100

Novel multi-omics deconfounding variational autoencoders can obtain meaningful disease subtyping

<h3>TCGA pan-cancer mRNA and DNA data augmented with artificial confounders utilised in "Novel multi-omics deconfounding variational autoencoders can obtain meaningful disease&nbsp;subtyping" by Zuqi Li and Sonja Katz (manuscript in preparation).</h3> <p>The following data curation steps were carried out:&nbsp;</p> <ul> <li><strong>Step 1. Download data from TCGA</strong> <ul> <li>R package `TCGAbiolinks`</li> <li>2547 patients (after step 2) with 6 cancer types: <ul> <li>BRCA (731)</li> <li>THCA (408)</li> <li>BLCA (387)</li> <li>LUSC (297)</li> <li>HNSC (412)</li> <li>KIRC (312)</li> </ul> </li> <li>mRNA expression profiles</li> <li>DNAm expression profiles</li> <li>Clinical data: <ul> <li>tumor stage: i, ia, ib, ii, iia, iib, iii, iiia, iiib, iiic, iv, iva, ivb, ivc, x</li> <li>age at diagnosis</li> <li>race: 'white', 'black or african amarican', 'asian', 'american indian or alaska native'</li> <li>gender<br><br></li> </ul> </li> </ul> </li> <li><strong>Step 2. Removal criteria</strong> <ul> <li>Patients with <ul> <li>NA or 'not reported' clinical data</li> <li>race 'american indian or alaska native'</li> <li>tumor stage x</li> </ul> </li> <li>mRNA and DNAm probes with <ul> <li>0 variance across all included patients</li> <li>not shared across all cancer types</li> <li>with missing values<br><br></li> </ul> </li> </ul> </li> <li>&nbsp;<strong>Step 3. Encode clinical vairables and save datasets</strong> <ul> <li>mRNA dataset: 2547 patients x 58,456 mRNAs</li> <li>DNAm dataset: 2547 patients x 232,088 DNAm</li> <li>clinic dataset: 2547 patients x 6 variables<br>&nbsp; &nbsp; 1. patient ID<br>&nbsp; &nbsp; 2. tumor stage: 1, 1, 1, 2, 2, 2, 3, 3, 3, 3, 4, 4, 4, 4<br>&nbsp; &nbsp; 3. age at diagnosis<br>&nbsp; &nbsp; 4. race: asian(1), black or african amarican(2), white(3)<br>&nbsp; &nbsp; 5. gender: female(0), male(1)<br>&nbsp; &nbsp; 6. cancer type: BRCA(1), THCA(2), BLCA(3), LUSC(4), HNSC(5), KIRC(6)<br>&nbsp; &nbsp;&nbsp;</li> </ul> </li> <li><strong>&nbsp;Step 4. Pre-process the datasets</strong> <ul> <li>mRNA dataset: '<em>TCGA_mRNAs_processed.csv'</em><br> <ul> <li>Take the 2000 mRNAs with highest variance</li> <li>Rescale every feature to [0,1]</li> <li>--&gt; 2547 patients x 2000 mRNAs</li> </ul> </li> <li>DNAm dataset: <em>'TCGA_DNAm_processed.csv'</em><br> <ul> <li>Take the 2000 DNAm with highest variance</li> <li>Rescale every feature to [0,1]</li> <li>--&gt; 2547 patients x 2000 DNAm</li> </ul> </li> <li>clinic dataset:<em> 'TCGA_clinic.csv'<br><br></em></li> </ul> </li> <li><strong>Step 5. Simulate confounders (instructions can be found in Methods section of manuscript)</strong> <ul> <li>Linear confounder: <ul> <li><em>'TCGA_confounder_linear.csv' -</em> linear confounding classes<em><br></em></li> <li><em>'TCGA_DNAm_confounded_linear.csv' </em>- linearly confounded DNAm data<em><br></em></li> <li><em>'TCGA_mRNA2_confounded_linear.csv'&nbsp;</em> - linearly confounded mRNA data<em><br></em></li> </ul> </li> <li>Squared confounder <ul> <li><em>'TCGA_confounder.csv' -</em> squared confounding classes<em><br></em></li> <li><em>'TCGA_DNAm_confounded.csv' </em>- squared confounded DNAm data<em><br></em></li> <li><em>'TCGA_mRNA2_confounded.csv'&nbsp;</em> - squared confounded mRNA data</li> </ul> </li> <li>Categorical confounder&nbsp; <ul> <li><em>'TCGA_confounder_categ2.csv' -</em> categorical confounding classes<em><br></em></li> <li><em>'TCGA_DNAm_confounded_categ2.csv' </em>- categorically confounded DNAm data<em><br></em></li> <li><em>'TCGA_mRNA2_confounded_categ2.csv'&nbsp;</em> - categorically&nbsp; confounded mRNA data</li> </ul> </li> <li>Multiple confounders - combined effect (linear + squared + categorical)<br> <ul> <li><em>'TCGA_confounder_multi.csv' -</em> confounding classes for combined effect<em><br></em></li> <li><em>'TCGA_DNAm_confounded_multi.csv' </em>- DNAm data with combined effect<em><br></em></li> <li><em>'TCGA_mRNA2_confounded_multi.csv'&nbsp;</em> - mRNA data&nbsp;with combined effect</li> </ul> </li> </ul> </li> </ul> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Dec 2023View details →
zenodo52/100

Datasets of "Carbide coating on nickel to enhance the stability of supported metal nanoclusters" Nanoscale, 2022, 14, 3589-3598

<p>These are the datasets related to the publication &quot;Carbide coating on nickel to enhance the stability of supported metal nanoclusters&quot;, Nanoscale, 2022, 14, 3589-3598 (<a href="https://doi.org/10.1039/D1NR06485A">https://doi.org/10.1039/D1NR06485A</a>). They are saved as NeXus/HDF5 files according to the nxstm NeXus application definition (<a href="https://doi.org/10.5281/zenodo.5792930">https://doi.org/10.5281/zenodo.5792930</a>).</p>

opencc-by-4.0Aug 2022View details →
zenodo52/100

ColoPola: A dataset of colorectal cancer polarimetric images (Mueller matrix elements) for colorectal cancer detection

<p><strong>ColoPola</strong> dataset is <strong>Colo</strong>rectal cancer <strong>Pola</strong>rimetric images dataset</p> <p>The dataset consists of 572 slices (specimens) with 20,592 images, 284 slices of which were designated as cancer samples and 288 as normal samples.</p> <p>Each sample has 36 polarimetric images (i.e., HH, HV, HP, HM, HR, HL, VH, VV, VP, VM, VR, VL, PH, PV, PP, PM, PR, PL, MH, MV, MP, MM, MR, ML, RH, RV, RP, RM, RR, RL, LH, LV, LP, LM, LR, and LL).</p> <p>Each folder in the <strong>ColoPola</strong> dataset consists of 36 polarimetric images. Each image is 1280x1024 pixels in size and was created in the TIF file format (HH.tif, HV.tif, ..., LL.tif).&nbsp;</p>

opencc-zeroNov 2023View details →
zenodo52/100

Fruit, seed dispersal, and life history traits of tropical rainforest trees of the Anamalai Hills, Western Ghats, India

<p>This dataset contains compiled Fruit, seed dispersal, and life history traits of tropical rainforest trees of the Anamalai Hills, Western Ghats, India. The list of species included are mainly from the following two related publications:<br>- Muthuramkumar, S., Ayyappan, N., Parthasarathy, N., Mudappa, D., Raman, T.R.S., Selwyn, M.A. and Pragasan, L.A. (2006), <a href="https://doi.org/10.1111/j.1744-7429.2006.00118.x">Plant Community Structure in Tropical Rain Forest Fragments of the Western Ghats, India</a>. <em>Biotropica</em>, 38: 143-160. https://doi.org/10.1111/j.1744-7429.2006.00118.x<br>- Osuri, A., Chakravarthy, D., Mudappa, D., Raman, T., Ayyappan, N., Muthuramkumar, S., &amp; Parthasarathy, N. (2017). <a href="http://httpd//doi.org/10.1017/S0266467417000219">Successional status, seed dispersal mode and overstorey species influence tree regeneration in tropical rain-forest fragments in Western Ghats, India</a>. <em>Journal of Tropical Ecology</em>, 33(4), 270-284. doi:10.1017/S0266467417000219<br>The present dataset is an expanded and updated version of the related dataset available at <a href="https://doi.org/10.5061/dryad.vd0nn">https://doi.org/10.5061/dryad.vd0nn</a><br>&nbsp;<br>Species traits information was collated from <a href="http://www.biotik.org/">BIOTIK (http://www.biotik.org/</a>), <a href="http://www.flowersofindia.net/">Flowers of India (http://www.flowersofindia.net/)</a>, India Biodiversity Portal (http://indiabiodiversity.org/), <a href="https://doi.org/10.5061/dryad.234/1">Global wood density database (https://doi.org/10.5061/dryad.234/1)</a> and <a href="https://doi.org/10.1017/S0266467417000219">Osuri et al. (2014): https://doi.org/10.1017/S0266467417000219</a>. We also referred to the following previous studies that provided information on the successional status of rain-forest species in the Western Ghats (Chetana 2013, Pascal 1988, Raman et al. 2009, Sreejith 2005).</p> <p><strong>References:</strong><br>CHETANA, H. C. 2013. Assessing the ecological processes in abandoned tea plantations and its implication for ecological restoration in the Western Ghats, India. PhD thesis, Manipal University.<br>OSURI, A. M., KUMAR, V. S. &amp; SANKARAN, M. 2014. Altered stand structure and tree allometry reduce carbon storage in evergreen forest fragments in India&rsquo;s Western Ghats. <em>Forest Ecology and Management </em>329: 375&ndash;383.<br>PASCAL, J. P. 1988. <em>Wet evergreen forests of the Western Ghats of India: Ecology, structure, floristic composition and succession</em>. Institut Fran&ccedil;ais de Pondich&eacute;ry, Pondicherry.<br>RAMAN, T. R. S., MUDAPPA, D. &amp; KAPOOR, V. 2009. Restoring rainforest fragments: survival of mixed-native species seedlings under contrasting site conditions in the Western Ghats, India. <em>Restoration Ecology</em> 17:137&ndash;147.<br>SREEJITH, K. A. 2005. Ecological and ecophysiological studies on the successional status of tree seedlings in tropical wet evergreen and semi-evergreen forests of Kerala. PhD thesis, Forest Research Institute, Dehradun.</p> <p><strong>Geographic Coverage:</strong><br>1. Location/Study Area: Valparai Plateau, Tamil Nadu, India; Anamalai Tiger Reserve, Tamil Nadu, India<br>2. GPS coordinates: Valparai Plateau (10&deg;15'- 10&deg;22'N, 76&deg;52' - 76&deg;59'E); Anamalai Tiger Reserve (10&deg;12' - 10&deg;35'N, 76&deg;49' - 77&deg;24'E)</p> <p><strong>Temporal Coverage:</strong><br>1. Begins: 2003-03-01 (Year, Month, Day)<br>2. Ends: 2024-02-10 (Year, Month, Day)</p> <p>Besides the <strong>README.txt</strong> file, the dataset includes the following comma-delimited text (csv) file with the data in columns as explained below:</p> <p><strong>Anamalai_tree_traits_2024.csv</strong></p> <p><strong>spec_name_ORIG:</strong> Scientific name of the species used during the data collection<br><strong>genus:</strong> Genus of the taxon<br><strong>specificEpithet:</strong> Specific epithet of the taxon in the Latin binomial name<br><strong>Accept_name_WFO:</strong> Updated scientific name of the species as in Plants of the World Online (POWO, https://powo.science.kew.org/)<br><strong>Habit:</strong> life form of the species(tree/shrub/cane/palm)<br><strong>Distribution:</strong> Distribution of the species in the study area (Native/Endemic/Introduced)<br><strong>IUCN_status:</strong> IUCN status of the species (CR-Critically Endangered,DD-Data deficient,EN-Endangered,LC-Least Concern,NT-Near Threatened,VU-Vulnerable,NA-Unknown)<br><strong>Wden_final:</strong> Wood density value assigned for the species (g cm^-3); NA - not available; sourced from Global wood density database (https://doi.org/10.5061/dryad.234/1)<br><strong>wd_level:</strong> Level in which the wood density value belongs (Species - wood density value is from species level; genus - wood density value assigned is the genus level average value)<br><strong>fruit_type:</strong> Morphological type of fruit<br><strong>fleshy_dry:</strong> Whether fruit is a dry fruit or fleshy, with aril or other parts&nbsp;<br><strong>seed_size:</strong> Species seed size: L = Large (&gt;3 cm); M = Medium (1-3 cm); S = Small (&lt;1 cm)<br><strong>disperser:</strong> Categories indicating seed dispersal mode: Bird, mammal, bird and mammal (Mammal_bird), gravity, wind, or unknown<br><strong>habitat:</strong> Habitat affinity category: EG_edg - evergreen forest edge; EG_for - evergreen forest; Dec_for - deciduous forest; Int &ndash; Introduced species; Unknown &ndash; Unknown<br><strong>habt_new:</strong> Habitat affinity new category: Mature &ndash; mature forest; Secondary &ndash; secondary forest, NA - unknown/Introduced species<br><strong>ad_ht:</strong> Species maximum adult height (m)</p>

opencc-by-4.0Feb 2024View details →
zenodo52/100

DNA Origami Raw AFM Data - NanoLocz: Image analysis platform for AFM, high-speed AFM and localization AFM

<p>The data file is in the original ARIS data format as captured on a Cypher VRS1250 AFM (Oxford Instruments)<br><br><br></p>

opencc-by-4.0Dec 2023View details →
zenodo52/100

The International Soundscape Database: An integrated multimedia database of urban soundscape surveys -- questionnaires with acoustical and contextual information

<h1>Introduction</h1> <p>The International Soundscape Database contains the results of a series of soundscape assessment campaigns carried out across Europe and China. The data collection process was conducted according to the <a href="https://www.mdpi.com/2076-3417/10/7/2397">SSID Protocol [1]</a> which integrates in situ questionnaires about users' soundscape experience, with binaural recordings, sound level meter readings, and 360 degree video. The core of this database are individual soundscape questionnaires collected for 3,500+ participants completed in situ in cities across Europe and China, and the psychoacoustic analysis of 30s binaural recordings which can be matched up to each questionnaire.</p> <p>The SSID Protocol was based on the ISO 12913&nbsp;standard for soundscape data collection [2]. For more information on the specifics of how this data is collected, please see [1].</p> <p>It is the intention that this dataset be added to and augmented with new locations, cities, and contexts in the future. This will be done both by the SSID team at University College London, but we also strongly welcome contributions from other researchers and practicioners. If a soundscape assessment is collected according to the SSID Protocol, it can be integrated with the rest of the database to form a large, cohesive, and ever-growing database of soundscape assessments.&nbsp;</p> <h2>Analysis</h2> <p>Code for exploring and analysing this dataset is included as part of the <a href="https://soundscapy.readthedocs.io/en/latest/">Soundscapy package</a>.</p> <h2>Included Files</h2> <p>This dataset incorporates surveys taken in multiple urban public spaces across several cities in Europe and China. These urban spaces include places like parks, urban squares, green spaces, and market streets. At each location, up to 100 questionnaires were collected over a series of multi-hour long sessions. Therefore the data is organised by LocationID, then SessionID, then GroupID.</p> <p>The basic directory structure and contents can be found below.&nbsp;</p> <h3>Survey Data (.csv)</h3> <p>'ISD v1.0 Data.csv' organises the data according to the labels given above.</p> <h3>Survey Metadata (.xlsx)</h3> <p>In addition a metadata file ('ISD v1.0 Metadata.xlsx') with photos and descriptions of each of the locations is provided. This metadata file also includes Data Dictionaries for each of the survey instrument versions included. These data dictionaries document precisely the questions asked and the available reponse labels and coding, along with the relevant translations.</p> <h3>Psychoacoustic Analysis (.csv)</h3> <p>The compiled csv file is formatted with a row for each individual participant's questionnaire response, then includes the psychoacoustic analysis of the 30s binaural recording taken while the participant was completing the questionnaire. Details about the psychoacoustic analyses is given in the 'Acoustic Settings' tab in the metadata file.</p> <p>The compiled survey and psychoacoustic analysis data is contained in 'ISD v1.0 Data.csv'. This is compiled from raw survey data files contained in 'Survey_Data', with individual cleaned survey and psychoacoustic data files included in 'Survey_Data/Interim_&lt;date&gt;'. The scripts for compiling this data are included in 'Scripts/'.</p> <h3>Sound Level Meter logs (.xlsx)</h3> <p>'SLM_&lt;city&gt;/' folders include session-long (i.e. ~3hrs) sound level meter log data in.xlsx files for each SessionID.</p> <h3>Binaural Recordings (32-bit floating point .wav)</h3> <p>'WAV_&lt;city&gt;/' folders include the ~30s binaural recordings in 32 bit floating point .wav format. Within each city folder are a set of LocationID folders containing their associated recordings. The wav files are titled with its GroupID, which is matched to the corresponding survey GroupIDs.&nbsp;</p> <h3>Cleaning and Compilation Scripts (.py)</h3> <p>Python code for cleaning and compiling the data from the raw survey data (within Survey_Data/source_data) are provided. These can be run within the provided demo notebook, or from the terminal by calling 'python -m ISDv1_main' with the relevant arguments. See the README.md file in this directory for more information.</p> <pre><code><br>├── ISD v1.0 Data.csv ├── ISD v1.0 Metadata.xlsx ├── SLM_Granada │ ├── CampoPrincipe1_SLM.xlsx │ ├── ... ├── SLM_Groningen │ └── Noorderplantsoen1_SLM.xlsx ├── SLM_etc ├── Scripts │ ├── ISDcleanDemo.ipynb │ ├── ISDcleaning.py │ ├── ISDpsycho.py │ ├── ISDv1_main.py │ ├── README.md │ └── pyproject.toml ├── Survey_Data │ ├── Interim_2024-02-08_cleaned │ └── source_data ├── WAV_Granada_1 │ ├── CampoPrincipe │ ├── ... ├── WAV_etc</code></pre> <p><strong>Citation</strong>: If you use the ISD or part of it, please cite our paper describing the data collection protocol [1] and this dataset itself.</p> <p><strong>License and reuse</strong>: All ISD recordings are provided under the Creative Commons Attribution 4.0 International (CC BY 4.0) License and are free to use. We encourage other researchers to replicate the SSID protocol and contribute new locations to the dataset. We also encourage the use of these recordings and the perceptual data for further soundscape research purposes. Please provide the proper attribution and get in touch with the authors if you would like to contribute new data or for any other collaborations.</p> <p>&nbsp;</p> <p>[1] Mitchell A, Oberman T, Aletta F, Erfanian M, Kachlicka M, Lionello M, Kang J. The Soundscape Indices (SSID) Protocol: A Method for Urban Soundscape Surveys&mdash;Questionnaires with Acoustical and Contextual Information. <em>Applied Sciences</em>. 2020; 10(7):2397. <a href="https://www.mdpi.com/2076-3417/10/7/2397">https://doi.org/10.3390/app10072397&nbsp;</a></p> <p>[2]&nbsp;ISO/TS 12913-2:2018 (2018). &ldquo;Acoustics &ndash; Soundscape &ndash; Part 2: Data collection and reporting requirements&rdquo; International Organization for Standardization, Geneva, Switzerland, 2018</p> <p>[3] Mitchell A, Oberman T, Aletta F, Kachlicka M, Lionello M, Erfanian M, Kang J. Investigating Urban Soundscapes of the COVID-19 Lockdown: A predictive soundscape modeling approach.<em>&nbsp;Journal of the Acoustical Society of America</em>. 2021.</p>

opencc-by-4.0Feb 2024View details →
zenodo52/100

Map of the temples of Allat in the Near East

<p>This map presents the localisation spots of the temples of Allat, located only in the Near East. The map is a result of data collection and building the database in the NodeGoat software. This map is based also on the information from the epigraphical and archaeological sources.</p>

opencc-by-4.0Dec 2023View details →
zenodo52/100

Corpus of critical citations contexts

<p>We present here a corpus of 505 critical citation contexts, i.e. a set of sentences or propositions that contain at least one citation of a study towards which the author(s) has/have a negative opinion. Those contexts come from other existing annotated corpora, from our readings about critical citation and disagreement in science, and from contexts manually annotated by native speakers of English. We have re-annotated all those contexts in order to be sure that they match our definition of critical citations. This corpus can be helpful to train tools dedicated to the automatic retrieval of critical citations. English (2024-02-20)</p>

opencc-by-4.0Feb 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record