Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

6,025

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

6,025 results for “science”

Learn how ShareScore rates datasets ↗
zenodo44/100

X-ray Fluorescence Mapping dataset for use in Heritage Science

<p>The following data sets were collected to support the potential uses of opensource data in the context of digital humanities and heritage sciences. &nbsp;</p> <p>This proposed experiment is conducted by the UCL Institute for Sustainable Heritage in collaboration with the Centre for Digital Humanities. Imaging methods including Photography, Multispectral Imaging, Hyperspectral Imaging and Xray Fluorescence Mapping have been collected along with the complete readout metadata of the instrumentation. &nbsp;</p> <p>We hope that you find the data helpful, and we welcome you to use the data in any way you wish, for all and any analysis development purposes. For us to build upon this research, we ask that in return you would be willing to share in some regard&nbsp;your experiences in using open-source data, using our data,&nbsp;successes and issues. &nbsp;</p> <p>If you would be willing to engage with us in this endeavor, please feel free to contact us so that we may be able to follow up with you. &nbsp;</p> <p>Other Data sets available&nbsp;<a href="https://zenodo.org/record/7319696#.Y3NuOXbP2Uk">Here</a></p> <p>E: <a href="mailto:molly.fort.21@ucl.ac.uk">molly.fort.21@ucl.ac.uk</a>&nbsp;</p> <p>Object Paradata; &nbsp;</p> <ul> <li><strong>Postcard &ndash; c. Early 1900&#39;s &nbsp;</strong></li> <li><strong>Language &ndash; Eng.&nbsp;</strong></li> <li><strong>Materials &ndash; colour print on card, metallic leafing.&nbsp;</strong></li> <li><strong>Front transcription - &nbsp;</strong></li> <li><strong>&nbsp;&lsquo;Greetings&rsquo;&nbsp;</strong></li> <li><strong>&nbsp;&lsquo;May your Birthday bring you Peace &amp; perfect Happiness, Golden hopes &amp; Love of Friends, And every Happiness this world can send.&rsquo;&nbsp;</strong></li> <li><strong>Object Dimensions &ndash; 138mm X 88mm</strong></li> </ul> <p>The postcard is an item of ephemera donated to the UCLDH Digitisation Suite by Prof Melissa Terras, for teaching and training purposes in 2015.</p> <p>This folder contains:</p> <p>X-ray fluorescence imaging map collected from a&nbsp;<a href="https://www.bruker.com/en/products-and-solutions/elemental-analyzers/micro-xrf-spectrometers/m4-tornado-plus.html">Bruker M4+ Tornado Micro-XRF System</a></p> <ul> <li>Postcardmap1.bcf - Bruker composite file containing full fluorescence spectral and mapping data with some other information. Can be read by Bruker software or using freely dowloadable&nbsp;<a href="https://hyperspy.org/">Hyperspy</a>.</li> <li>EDX.hdf5 - full fluorescence spectral and mapping data in open format, created and readable via. Hyperspy.</li> <li>Postcard_data.txt - metadata saved in ASCII format by Bruker system.</li> <li>postcardmap1_*.png - individual png images of mapped elements, listed in filename.&nbsp;</li> </ul> <p>&nbsp;</p>

opencc-by-4.0Nov 2022View details →
zenodo44/100

Skills of science inquiry in citizen science - Datasets

<p>The dataset provides results of a prediction task aimed at predicting the presence of science inquiry skills in CS project descriptions. Only 2939 English project descriptions from the database were used for prediction and the results indicated 438 projects (around 15%) consist of one or more skills of science inquiry that we were interested in. In total 20 different types of science inquiry skills were considered for this study.&nbsp;</p> <p>See further detail in D2.2 section 7.4.</p> <p><strong>Content and grouping:&nbsp;</strong></p> <ul> <li> <p>The dataset contains the following details: Platform ID (from which platform the CS project descriptions were retrieved),&nbsp; Project Title (Name of the CS project), How many times each science inquiry skill (keywords) was mentioned in each project description.</p> </li> </ul>

opencc-by-4.0Nov 2022View details →
zenodo44/100

[Dataset] Does Volunteer Engagement Pay Off? An Analysis of User Participation in Online Citizen Science Projects

<p>Corresponding dataset for the publication &quot;Does Volunteer Engagement Pay Off? An Analysis of User Participation in Online Citizen Science Projects&quot;, a conference paper for the conference&nbsp;CollabTech 2022:&nbsp;<a href="https://link.springer.com/book/10.1007/978-3-031-20218-6">Collaboration Technologies and Social Computing</a>&nbsp;and&nbsp;published as part of the&nbsp;<a href="https://link.springer.com/bookseries/558">Lecture Notes in Computer Science</a>&nbsp;book series (LNCS,volume 13632) <a href="https://link.springer.com/chapter/10.1007/978-3-031-20218-6_5">here</a>. Usernames have been anonymised.</p> <p>The structure of the&nbsp;dataset is as follows:</p> <p><strong>Annotations</strong>&nbsp;</p> <p><em>List of annotations made per day for each of the analysed projects.</em></p> <p><code>annotations.csv&nbsp;</code></p> <p><strong>Comments&nbsp;</strong></p> <p><em>Total list of comments with several data fields (i.e., comment id, text, reply_user_id)</em></p> <p><code>comments.csv&nbsp;</code></p> <p><strong>Rolechanges</strong>&nbsp;</p> <p><em>List of roles per user to determine number of role changes&nbsp;</em></p> <p><code>478_rolechanges.csv</code></p> <p><code>1104_rolechanges.csv</code></p> <p><code>...</code></p> <p><strong>Totalnetworkdata</strong>&nbsp;</p> <p><em>Network data (edge and node sets) for the given projects (without time slices).</em></p> <p>Edges&nbsp;</p> <ul> <li> <p><code>478_edges.csv</code></p> </li> <li> <p><code>1104_edges.csv</code></p> </li> </ul> <p>Nodes&nbsp;</p> <ul> <li> <p><code>478_nodes.csv</code>&nbsp;</p> </li> <li> <p><code>1104_nodes.csv</code>&nbsp;</p> </li> </ul> <p><strong>Trajectories</strong>&nbsp;</p> <p><em>Network data (edge and node sets) for the given projects and all time slices (Q1&nbsp;2016 - Q4 2021)</em></p> <p>478&nbsp;</p> <ul> <li>Edges&nbsp; <ul> <li> <p><code>edges_4782016_q1.csv</code></p> </li> <li> <p><code>edges_4782016_q2.csv</code></p> </li> <li> <p><code>edges_4782016_q3.csv</code></p> </li> <li> <p><code>edges_4782016_q4.csv</code></p> </li> </ul> </li> <li> <p>...</p> </li> <li>Nodes&nbsp; <ul> <li><code>nodes_4782016_q1.csv</code></li> <li> <p><code>nodes_4782016_q4.csv</code></p> </li> <li> <p><code>nodes_4782016_q3.csv</code></p> </li> <li> <p><code>nodes_4782016_q2.csv</code></p> </li> <li> <p><code>...</code></p> </li> </ul> </li> </ul> <p>&nbsp;</p> <p>1104&nbsp;</p> <ul> <li> <p>Edges&nbsp;</p> <ul> <li> <p><code>...</code></p> </li> </ul> </li> <li> <p>Nodes&nbsp;</p> <ul> <li> <p><code>...</code></p> </li> </ul> </li> <li> <p><code>...</code></p> </li> </ul> <p>&nbsp;</p>

opencc-by-4.0Nov 2022View details →
zenodo44/100

Datasets containing the results from the analysis on SDGS and eHealth inside the Citizen Science Community on Twitter

<p>This datasets contain the results from our analyses of the Citizen Science Community on Twitter. These analyses have been done to better understand the discussion about SDGs, eLearning&nbsp;and eHealth.</p> <p><strong>T</strong>he purpose of sharing these datasets&nbsp;is to provide the basis to reproduce&nbsp;the results reported in the associated deliverable. These files are not raw data, since due to privacy concerns we can not share personal information from Twitter.</p> <p><strong>dominant_topics_anonym.xlsx</strong>: Excel datasheet. This dataset contians the distribution of the most discussed topics inside the SDGs discussion.</p> <p><strong>Edges_Hashtag_connected.csv</strong>:&nbsp;&nbsp;CSV file. This dataset contains the edges to build the network of connected hashtags.&nbsp;This edges can be used to build a network and explore the connections or to statiscally analyse the results.</p> <p><strong>hashtags.csv</strong>: CSV file. This dataset contains the results of the most used hashtags in the analysis about eLearning.&nbsp;<br> &nbsp;</p> <p><strong>hashtags_treemap_health.xlsx</strong>: Excel datasheet. This dataset contains the results of the most frequent hashtags in the eHealth analysis.</p> <p><strong>ldavis_prepared_ieee17.html</strong>: HTML file. This file contains the Intertopic distance map and most salient terms from the topic modelling analysis done in the SDGs conversation study.</p> <p><strong>Most_retweeted_accounts.xlsx</strong>: Excel datasheet. This dataset contains the top 20 users that receive more retweets in the conversation around eHealth. The column called&nbsp;Indegree refers to the topological value calculated from the network of retweets. This indegree is equivalent to the number of retweets received. On the other hand, Outdegree is the opposite, so number of retweets given to others.</p> <p><strong>Most_retweeting_account.xlsx</strong>: Excel datasheet. This dataset presents the opposite part of the previous one, the accounts that retweet the most from the eHealth analysis. The columns contain the same indicators: Indegree and Outdegree.</p> <p><strong>sdgs_count_publish.csv</strong>: CSV file. This dataset contains the number of tweets assigned to the different SDGs from the analysis done on the conversation about these Goals.</p> <p><strong>sdgs_tweets_sdgsaccess.xlsx</strong>: Excel datasheet. Same file as the previous one in other format to ease the handling in Excel.</p> <p><strong>top_hash_health.xlsx</strong>: Excel datasheet. The most used hashtags inside the conversation about eHealth.</p> <p><strong>topics_tweets_sdgsaccess.xlsx</strong>: Excel datasheet. Tweets by topic extracted using Machine Learning in the SDGs analysis.</p> <p>&nbsp;</p> <p>This repository will receive updates in the future in order to present all the data available and publishable from the different analysis that were described.</p>

opencc-by-4.0Nov 2022View details →
zenodo44/100

Using Hydroclimate Modeling and Social Science to Enhance Flood Resilience on Lake Ontario through the Climate Smart Communities Program

<p>This repository contains several&nbsp;data products associated with the New York Sea Grant project R/CHD-15 entitled <em>Using Hydroclimate Modeling and Social Science to Enhance Flood Resilience on Lake Ontario through the Climate Smart Communities Program.</em><strong><em> </em></strong>These products include:</p> <p>1.&nbsp;Estimates of the 25-year, 50-year, and 100-year flood&nbsp;across&nbsp;the New York coastline of Lake Ontario. These design events (reported in feet) are for still water levels that take into account both average water levels across the lake as well as local variations in water level due to storm surge. Wave setup and wave run-up&nbsp;are not considered in these design events. The design events&nbsp;incorporate the effects of water level regulation and the potential impacts of climate change on water supplies to Lake Ontario, and they are tailored for&nbsp;79 unique locations along the shoreline (identified based on longitude and latitude). These flood levels are presented in an online flood risk assessment tool at:&nbsp;https://kts48.users.earthengine.app/view/lake-ontario-water-level-scenarios</p> <p>2. Protocols and summary of results for a series of focus groups and structured telephone interviews with local officials from communities along the Lake Ontario shoreline to assess barriers to participation in the&nbsp;New York State Climate Smart Communities Program.</p> <p>3.&nbsp; A Crosswalk between activities and administrative requirements of the New York State Climate Smart Communities Program and other federal and state flood resiliency programs.&nbsp;</p> <p>4. A final report summarizing the products above.&nbsp;</p>

opencc-by-4.0Nov 2022View details →
zenodo44/100

Availability of information on citizen science activities, checked against the Activities & Dimensions Grid of Citizen Science on the basis of some projects

<p>The research resulting in this report aimed at answering the following questions:</p> <ul> <li> <p>Which information on citizen science activities is online available that matches the Activity &amp; Dimension Grid of Citizen Science or goes beyond it?&nbsp;&nbsp;</p> </li> <li> <p>Is there any contradictory information?</p> </li> <li> <p>What can be the reason for the availability or non-availability of information about citizen science activities?</p> </li> <li> <p>How does/could this impact on the CS Track&rsquo;s recommendations?</p> </li> </ul> <p>The corresponding dataset consists of the results of a keyword-based search in the WP2 project database. The information retrieval resulted in 3318 projects on which information is available in German or English.</p> <p>More information on this research can be found in D2.2 section 3.2.</p>

opencc-by-4.0Nov 2022View details →
zenodo44/100

arXiv abstracts and titles from 1,469 single-authored papers (100 unique authors) in computer science

<p>This dataset is meant to be used for experiments of Authorship Analysis. The dataset&nbsp;consists of abstracts of single-author papers from arXiv crawled using the arXiv&#39;s API by querying a list of computer-science-related keywords (&quot;deep learning&quot;, &quot;machine learning&quot;, &quot;information retrieval&quot;, &quot;computer science&quot;, &quot;data mining&quot;, &quot;support vector&quot;, &quot;logistic regression&quot;, &quot;artificial intelligence&quot;, &quot;supervised learning&quot;&#39;).&nbsp;The corpus somehow follows a power-law distribution, with few prolific authors and many authors accounting for very few papers each: we retained authors with at&nbsp;least 10 papers, resulting in a total of 1,469 documents from 100 authors. The most prolific authors (Peter D. Turney and Subhash Kak) have 34 abstracts to their names, the 10 most prolific authors have written 22 or more articles, while 50% of the authors have no more than 12 abstracts to their names. In order to divide the corpus into a training set and a test set we perform a stratified split, with the production of each author being split into a training set (70%) and a test set (30%). We use these documents as examples of &quot;scientific communication&quot;, characterised by a precise and compact style, with an abundance of technical terminology.</p>

opencc-by-4.0Dec 2022View details →
zenodo44/100

Data for: Luo et al., Expiratory aerosol pH: the overlooked driver of airborne virus inactivation, Environmental Science and Technology, 10.1021/acs.est.2c05777

<p><strong>Experimental data </strong></p> <p>This folder contains the experimental data to the figures shown in the main manuscript and Supporting Information.</p> <p>Figures 1 and S3 (inactivation curves for IAV, SARS-CoV-2 and HCoV-229E)</p> <p>Figure 1 (rate constants)</p> <p>Figure 2 (EDB analysis of SLF)</p> <p>Figure S1A (zetasizer analysis to measure virus aggregation)</p> <p>Figure S1B (renilla and plaque assay data for viruses exposed to pH 5, 6 and 7)</p> <p>Figure S4A (EDB analysis of different SLF samples; raw data)</p> <p>Figure S4Amean&nbsp;(EDB analysis of different SLF samples; mean values)</p> <p>Figure S5 (EDB analysis of nasal mucus)</p> <p>Figure S8 (EDB analysis of&nbsp;slow crystal growth stage of SLF and nasal mucus)</p> <p>Figure S13 and S14 (literature data on inactivation of IAV and SARS-CoV-2 in aerosol particles)</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Dec 2022View details →
zenodo44/100

Seeing Through the "Science Eyes" of the ExoMars Rover - Supplementary Material

<p>Simulated views from ExoMars PanCam instrument to assist operations planning.</p> <p>Described in more detail in the linked journal article.</p>

opencc-by-4.0Mar 2020View details →
zenodo44/100

PLOS Open Science Indicators & Zotero Romania Metadata

<p>Matched metadata from PLOS Open Science Indicators and Zotero export</p>

opencc-by-4.0Dec 2022View details →
zenodo44/100

CS Track Citizen Science Survey Data 2021

<p>CS Track is launching a survey to gather citizen scientists&rsquo; (16 year old and older) perspectives on activities and forms of participation, learning and knowledge-building in citizen science (CS) projects. The aim of CS Track is to broaden our knowledge about CS and the impact CS activities can have. CS Track will do this by investigating a large and diverse set of CS activities, disseminating best practices and formulating knowledge-based policy recommendations in order to maximise the potential benefits of CS activities for individual citizens, organisations and society. This multi-perspective approach will allow us to shed light on the role of citizen science in society and social attitudes and emerging cultures in communities that engage with science and technology challenges.</p> <p><strong>CSTrack_Citizen_Science_Survey_Data_Final_Anon.csv</strong>: CSV File. CS Track Citizen Science Survey Data in a CSV file.</p> <p><strong>CSTrack_Citizen_Science_Survey_Data_Final_Anon.xlsx</strong>: Excel datasheet. Same file as previous in an Excel datasheet.</p> <p><strong>CSTrack_Citizen_Science_Survey_Final.pdf</strong>: PDF-file. CS Track Citizen Science Survey.</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Dec 2022View details →
zenodo44/100

Spectral data associated to the publication: "Near-infrared reflectance spectroscopy of sublimating salty ice analogues. Implications for icy moons" by R. Cerubini et al. (Planetary and Space Science 211, 2022)

<p>This is the complete set of experimental NIR reflectance data collected by R. Cerubini and co-authors for the article &quot;Near-infrared reflectance spectroscopy of sublimating salty ice analogues. Implications for icy moons&quot; published in Planetary and Space Science 211 (2022). doi: https://doi.org/10.1016/j.pss.2021.105391.</p> <p>The article itself is published in open-access and provides the methodology for the spectral aquisitions, discussion of the errors and uncertainties, analysis of the spectra and implications for the composition of Solar System surfaces.</p> <p>The data are contained in ASCII files (columns separated by comma). The first column is the wavelength (in micrometers) and the other columns contain the reflectance data (in unit of reflectance factor). The different compositions are indicated in the filenames and correspond directly to the figures in the published paper.</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Jan 2023View details →
zenodo44/100

MuSpinSim data files for Galaxy materials science tutorials

<p>This is a training dataset for use in Galaxy materials science tutorials. These files can be compared to the output of simulations by MuSpinSim for dissipation of muon spins.</p> <p>The files included&nbsp;are:</p> <ul> <li><strong>dissipation_theory.dat:</strong> theoretical values formatted as a MuSpinSim output</li> <li><strong>experiment.dat:</strong> mock experimental values formatted as a MuSpinSim output</li> </ul>

opencc-by-4.0Jan 2023View details →
zenodo44/100

VABB-SHW: Dataset of Flemish Academic Bibliography for the Social Sciences and Humanities (edition 12)

<p>This dataset contains the twelfth edition of the <em>Flemish Academic Bibliography for the Social Sciences and Humanities (VABB-SHW)</em>, a database of academic publications from the social sciences and humanities authored by researchers affiliated to Flemish universities (<a href="https://www.ecoom.be/en/data-collections/vabb-shw">more information</a>). Publications in the database are used as one of the parameters of the Flemish performance-based research funding system. Only approved publications are included in this dataset.</p>

opencc-by-4.0Dec 2021View details →
zenodo44/100

Data from: CoAct Citizen Science chatbot explores social support networks in mental health based on lived experiences

<p>A data set on lived experiences in the context of social support in mental health, created within a Citizen Social Science project.&nbsp;</p> <p><br> Societies around the world increasingly encounter wicked and complex problems, such as those related to mental health, environmental justice, and youth employment. <strong>CoAct as a EU-funded global effort</strong> addresses these problems by deploying Citizen Social Science.&nbsp;</p> <p>&nbsp;</p> <p><strong>Citizen Social Science</strong> is understood here as participatory research co-designed and directly driven by citizen groups sharing a social concern. This methodology wants to give citizen groups an equal &lsquo;seat at the table&rsquo; through <strong>active participation in research</strong>, from the design to the interpretation of the results and their transformation into concrete actions. Citizens thus act as <strong>co-researchers</strong> and are recognised as in-the-field competent experts.&nbsp;</p> <p>&nbsp;</p> <p>In Barcelona, a group of <strong>32 co-researchers</strong> work together with the OpenSystems group, Universitat de Barcelona, the Catalan Federation of Mental Health (Federaci&oacute; Salut Mental Catalunya), and with the help of many others on a better understanding of informal <strong>social support networks in mental health</strong> in the project <em>CoActuem per la Salut Mental</em> (lit. &ldquo;We act together for mental health&rdquo;). The co-researchers, who are either persons with a personal history of mental health problems or are family members of the latter, contributed their <strong>personal experiences related to social support</strong> in the form of <strong>222 micro-stories</strong>, each shorter than 400 characters, and most accompanied by an illustration by Pau Badia.</p> <p>&nbsp;</p> <p>Those micro-stories form the heart of the first co-created Citizen Science chatbot, the code of which is open on <a href="https://github.com/Chaotique/CoActuem_per_la_Salut_Mental_Chatbot.git">https://github.com/Chaotique/CoActuem_per_la_Salut_Mental_Chatbot.git</a> . The <strong>Telegram chatbot</strong> sends them to participants <strong>on a daily basis over the course of a year</strong> and asks them either, whether they and/ or their close surrounding lived this experience, too (stories of type C), or, how they would or would have reacted in the presented situation (stories of type T). The answers of each participant can be contrasted with the individual participants&rsquo; answer to a 32-questions <strong>socio-demographic survey</strong>. Further, the timing of the messages is included to allow for a broader analysis.&nbsp;&nbsp;</p> <p>&nbsp;</p> <p>The chatbot is still running, hence this data set will still be updated. For further information on the project <strong>CoAct</strong>, see <a href="https://coactproject.eu/">https://coactproject.eu/</a>. For further details on the co-creation process and purpose of the chatbot <strong>CoActuem per la Salut Mental</strong>, take a look on <a href="https://coactuem.ub.edu/">https://coactuem.ub.edu/</a>. Please direct your questions regarding the data set to <strong>coactuem[at]ub.edu</strong>.</p> <p>&nbsp;</p> <p><strong>Acknowledgements</strong></p> <p>The CoAct project has received funding from the European Union&#39;s Horizon 2020 research and innovation programme under grant agreement number 873048. We especially thank the co-researchers for the passion and time invested.</p>

opencc-by-4.0Feb 2023View details →
zenodo44/100

Materials Science Optimization Benchmark Dataset for Multi-Objective, Multi-Fidelity Optimization of Hard-Sphere Packing Simulations

<p>Benchmarks are an essential driver of progress in scientific disciplines. Ideal benchmarks mimic real-world tasks as closely as possible, where insufficient difficulty or applicability can stunt growth in the field. Benchmarks should also have sufficiently low computational overhead to promote accessibility and repeatability. The goal is then to win a &ldquo;Turing test&rdquo; of sorts by creating a surrogate model that is indistinguishable from the ground truth observation (at least within the dataset bounds that were explored), necessitating a large amount of data. In the fields of materials science and chemistry, industry-relevant optimization tasks are often hierarchical, noisy, multi-fidelity, multi-objective, high-dimensional, and non-linearly correlated while exhibiting mixed numerical and categorical variables subject to linear and non-linear constraints. To complicate matters, unexpected, failed simulation or experimental regions may be present in the search space. In this study, 494498 random hard-sphere packing simulations representing 206 CPU days worth of computational overhead were performed across nine input parameters with linear constraints and two discrete fidelities each with continuous fidelity parameters and results were logged to a free-tier shared MongoDB Atlas database. Two core tabular datasets resulted from this study: 1. a failure probability dataset containing unique input parameter sets and the estimated probabilities that the simulation will fail at each of the two steps, and 2. a regression dataset mapping input parameter sets (including repeats) to particle packing fractions and computational runtimes for each of the two steps. These two datasets are used to create a surrogate model as close as possible to running the actual simulations by incorporating simulation failure and heteroskedastic noise. For the regression dataset, percentile ranks were computed within each of the groups of identical parameter sets to enable capturing heteroskedastic noise. This is in contrast with a more traditional approach that imposes a-priori assumptions such as Gaussian noise e.g., by providing a mean and standard deviation. A similar approach can be applied to other benchmark datasets to bridge the gap between optimization benchmarks with low computational overhead and realistically complex, real-world optimization scenarios.</p> <p>For usage instructions, see&nbsp;https://matsci-opt-benchmarks.readthedocs.io/.</p>

opencc-zeroMar 2023View details →
zenodo44/100

The Effect of Soundscape Composition on Bird Vocalization Classification in a Citizen Science Biodiversity Monitoring Project

<p>This archive includes sound clips (.wav files) and associated mel-scale spectrograms of bird vocalizations for 54 species in Sonoma County, California, USA. These data were used for training and validating convolutional neural network (CNN) models for bird species detection. We also include xeno-canto training and validation mel spectrograms&nbsp;used to pretrain CNNs. Details on these data are explained in the paper by Clark et al. (2023) titled &quot;The effect of soundscape composition on bird vocalization classification in a citizen science biodiversity monitoring project&quot;. These data are available for use without restrictions, with no warranty on data quality or utility for a given application. We request that any work that does use these data cite the Clark et al. (2023) paper.<br> <br> Clark, M.L., Salas, L., Baligar, S., Quinn, C., Snyder, R.L., Leland, D., Schackwitz, W., Goetz, S.J., Newsam, S. (2023). The effect of soundscape composition on bird vocalization classification in a citizen science biodiversity monitoring project. <em>Ecological Informatics</em>.&nbsp;<a href="https://doi.org/10.1016/j.ecoinf.2023.102065">https://doi.org/10.1016/j.ecoinf.2023.102065</a></p> <p>Associated code for training CNN models,&nbsp;performing inference, and applying post-classification corrections can be found in the GitHub archive&nbsp;<a href="https://github.com/pointblue/Soundscapes2Landscapes/tree/master/CNN_Bird_Species">https://github.com/pointblue/Soundscapes2Landscapes/tree/master/CNN_Bird_Species</a></p> <p>Raw sound data from the Soundscapes to Landscapes project are available upon request: Dr. Matthew Clark, matthew.clark@sonoma.edu</p> <p>These data were collected as part of the&nbsp;Soundscapes to Landscapes project (<a href="https://soundscapes2landscapes.org/">soundscapes2landscapes.org</a>),&nbsp;funded by NASA&rsquo;s Citizen Science for Earth Systems Program (CSESP) 16-CSESP 2016-0009 under cooperative agreement 80NSSC18M0107.<br> <br> ----------------------------<br> This depository&nbsp;includes the following archives:</p> <ul> <li> <p>mel_specs.zip: contains 2-sec mel spectrograms split into training (&ldquo;tr&rdquo;), validation (&ldquo;val&rdquo;), testing (&ldquo;test&rdquo;) data for each target bird species (n = 54) used to fine-tune the CNNs. Select spectrogram files are appended with &ldquo;aug&rdquo; if they are augmented versions for the training data.</p> </li> <li> <p>wav.zip: contains the associated wav-format sound recordings used to generate the training, validation, testing mel spectrograms found in mel_specs.zip.</p> </li> <li> <p>Xeno-canto_pretrain.tar: contains 2-sec mel spectrograms split into training and validation data for 40 bird species used for CNN pre-training that were generated using a warbleR segmentation methodology described in the paper. The sound files used to generate these mel spectrograms came from the Kaggle competition,&nbsp;<a href="https://www.kaggle.com/datasets/imoore/xenocanto-bird-recordings-dataset">https://www.kaggle.com/datasets/imoore/xenocanto-bird-recordings-dataset</a><br> Mel spectrogram naming reflects the XC number used for cataloging on Xeno-canto in the format XC123456_2.png. The six numbers following the XC characters can be used to search for unique recordings on Xeno-canto (<a href="https://xeno-canto.org/">https://xeno-canto.org/</a>) using the search query &ldquo;nr:123456&rdquo; in the search tool or queried using the Xeno-canto API (<a href="https://xeno-canto.org/explore/api">https://xeno-canto.org/explore/api</a>). Unique recording names can be extracted from the mel spectrogram filenames.</p> </li> <li> <p>soundscape_test_wavs.zip: the wav-format&nbsp;sound recordings&nbsp;used to perform soundscape testing.</p> </li> </ul>

opencc-by-4.0Mar 2023View details →
zenodo44/100

Content Analysis of Canada's Science Based Departments and Agencies Open Science Action Plans

<p>Dataset and codebook for a content analysis of Science-based departments and agencies open science action plans in response to the Government of Canada&#39;s Roadmap for Open Science.&nbsp;</p>

opencc-by-4.0Apr 2023View details →
zenodo44/100

Designing Atmospheres: Theory and Science Symposium

<p>This dataset is an output of the &lsquo;Designing Atmospheres: Theory and Science&rsquo; Symposium (ATS), an Interfaces event of the Academy of Neuroscience for Architecture (ANFA), sponsored by the EU&rsquo;s Horizon 2020 MSCA Program &mdash; RESONANCES Project, the Perkins Eastman Studio, and the KSTATE APDesign. The symposium was hosted in the College of Architecture, Planning and Design (APDesign), Kansas State University, Manhattan (Kansas, USA), on March 28, 2023. Speakers: Kory Beighle (Kansas State University), Elisabetta Canepa (University of Genoa | Kansas State University), Bob Condia (Kansas State University), Zakaria Djebbara (Aalborg University | TU Berlin), and Harry Francis Mallgrave (Illinois Institute of Technology).</p> <p>&nbsp;</p> <p>Recent advances in science confirm many of the architects&rsquo; deep-rooted intuitions, improving knowledge about the perception of space and the meaning of architectural and urban design.&nbsp;The symposium &lsquo;Designing Atmospheres: Theory and Science&lsquo; presented to an audience of students, educators, architects, and scientists a conversation about the experience of design and building, specifically speaking to the significance of atmospheres, affordances, and emotions.</p> <p>&nbsp;</p> <p>This&nbsp;dataset is made of seven&nbsp;files:<br> no. 1 dataset summary (.pdf)<br> no. 1 symposium poster (.pdf)<br> no. 5 videos containing speakers&rsquo; presentations (.mp4)</p> <p>&nbsp;</p> <p>Recorded videos of each lecture are also available on the RESONANCES project website (www.resonances-project.com/harvest) and its YouTube channel (@resonancesproject5777).</p>

opencc-by-4.0May 2023View details →
zenodo44/100

Data from: Trends in butterfly populations in UK gardens – new evidence from citizen science monitoring

<p>This data package describes the annual abundance indices and trend estimates for 22 butterfly&nbsp;species in UK gardens for the period 2007-2020.</p> <p>These data form the basis of the results presented in:&nbsp;Plummer, K.E.,&nbsp;Dadam, D.,&nbsp;Brereton, T.,&nbsp;Dennis, E.B.,&nbsp;Massimino, D.,&nbsp;Risely, K.&nbsp;et al. (2023)&nbsp;Trends in butterfly populations in UK gardens&mdash;New evidence from citizen science monitoring.&nbsp;<em>Insect Conservation and Diversity</em>,&nbsp;1&ndash;&nbsp;13. Available from:&nbsp;<a href="https://doi.org/10.1111/icad.12645">https://doi.org/10.1111/icad.12645</a></p> <p>Please refer to the paper for an explanation of the underlying BTO Garden BirdWatch (GBW) data and modelling protocols used to produce the datasets included here.</p> <p>We would also greatly appreciate if you could fill out&nbsp;<a href="https://forms.gle/DCc58VXpdmqnTmTk8" target="_blank" rel="noopener">this very short form</a> to tell us how you intend to use these data. Thanks in advance!</p>

opencc-by-4.0May 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record