Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

188

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

188 results for “Big data”

Learn how ShareScore rates datasets ↗
zenodo48/100

ExcapeDB: An integrated large scale dataset facilitating Big Data analysis in chemogenomics

<p>ExcapeDB: An integrated large scale dataset facilitating Big Data analysis in chemogenomics</p> <p>Supplementary file (full dataset download)</p> <p>- v2 with SMILES errors fixed&nbsp; (19.01.2019)</p>

opencc-by-sa-4.0Nov 2016View details →
zenodo44/100

User requirements of Big Earth Data - Survey 2019

<p>The survey was conducted between November 2018 and May 2019 with the aim to find out how users working with large volumes of environmental data interact with data, what challenges they face and how they would like to use cloud-based data services in the future.</p> <p>The term Big Earth Data in this context refers to digital information about Earth, including observations, imagery, derived higher-level products, forecasts and analyses produced by computer models.</p> <p>The survey was conducted in collaboration with the European Centre for Medium-Range Weather Forecasts (ECMWF) and as part of a PhD thesis on &quot;Big Data technologies for environmental and climate data&quot; at University of Marburg, Germany.</p> <p>The results are published in form of two articles:</p> <ul> <li>Wagemann, J., Siemen, S., Seeger, B. and J. Bendix (2021): Users of open Big Earth data - An analysis of the current state. Computers and Geosciences 2021. <a href="https://doi.org/10.1016/j.cageo.2021.104916">doi:10.1016/j.cageo.2021.104916</a></li> <li>Wagemann, J. Siemen, S., Seeger, B. and J. Bendix (2021): A user&nbsp;perspective on future cloud-based services for Big Earth data. International Journal of Digital Earth 2021. doi:&nbsp;<a href="http://doi.org/10.1080/17538947.2021.1982031">10.1080/17538947.2021.1982031</a></li> </ul> <p>&nbsp;</p>

opencc-by-4.0Oct 2020View details →
zenodo44/100

Data for "Phenotypic responses to climate change are significantly dampened in big-brained birds"

<p>Anthropogenic climate change is rapidly altering local environments and threatening biodiversity throughout the world. Although many wildlife responses to this phenomenon appear largely idiosyncratic, a wealth of basic research on this topic is enabling the identification of general patterns across taxa. Here we expand those efforts by investigating how avian responses to climate change are affected by the ability to cope with ecological variation through behavioral flexibility (as measured by relative brain size). After accounting for the effects of phylogenetic uncertainty and interspecific variation in adaptive potential, we confirm that although climate warming is generally correlated with major body size reductions in North American migrants, these responses are significantly weaker in species with larger relative brain sizes. Our findings suggest that cognition can play an important role in organismal responses to global change by actively buffering individuals from the environmental effects of warming temperatures.</p>

opencc-by-4.0Jan 2022View details →
zenodo44/100

Fine-scale population spatialization data of China in 2018 based on real location-based big data

<p><strong>This data contains&nbsp;a geospatial population raster layer in GeoTIFF format with 1*1 km resolution&nbsp;for 31 provincial regions (2851 counties) of China in 2018 (pop2018.tif). It also provides the Tencent positioning data in 2018 (TN_hSum2018.tif), the table of statistical population of 2851 counties (statistical_population_2018_china_county.xls) and its vector map (statisitcal_pop.shp) and codes (code.docx).</strong></p>

opencc-by-4.0Apr 2021View details →
zenodo44/100

Big Data to Knowledge (BD2K) Training Coordinating Center (TCC) Educational Resource Discovery Index (ERuDIte) as Linked Data

<p>This is a release of the Big Data to Knowledge (BD2K) Training Coordinating Center (TCC) Educational Resource Discovery Index (ERuDIte)&nbsp;as Linked Data.<br> <br> ERuDIte contains over 11,000 training resources on data science including courses (MOOCs), video tutorials, conference talks, and other materials. The metadata of these resources is described uniformly using schema.org. In addition, we use machine learning techniques to tag each resource with concepts from the Data Science Education Ontology (DSEO), which we developed to further describe the contents of the training resources. Resource relevance and tags are curated by experts to ensure high quality. Finally, we map the references to people and organizations in the learning resource metadata to entities in DBpedia, DBLP, and ORCID, thus embedding our collection in the web of linked data. Our collection is continually growing. We hope that ERuDIte will provide a framework to foster open linked educational resources on the web.<br> <br> &nbsp;Distributed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License (https://creativecommons.org/licenses/by-nc-sa/4.0/)</p>

openother-openMay 2018View details →
zenodo44/100

sohamphanseiitb/BIG_Data_5MSEC: BIG Data Analysis of NASA's 5 Millennium Solar Eclipse Database

<p>Solar eclipses are a topic of interest among astronomers, astrologers and the general public as well. There were and will be about 11898 eclipses in the 5 millennia from 2000 BC to 3000 AD. Data visualization and regression techniques offer a deep insight into how various parameters of a solar eclipse are related to each other. Physical models can be verified and can be updated based on the insights gained from the analysis.</p> <p>The study covers the major aspects of data analysis including data cleaning, pre-processing, EDA, distribution fitting, regression and machine learning based data analytics. We provide a cleaned and usable database ready for EDA and statistical analysis.</p>

openother-openDec 2021View details →
edi44/100

Decomposition of Microstegium vimineum litter, plants grew through the Big Oaks National Wildlife Refuge in 2019. Litter used in this experiment naturally senesced in the fall 2019, decomposition data collected through 2020. Plants were infected or not-infected with the foliar fungal pathogen Bipolaris gigantea during the 2019 growing season.

Decomposition of plant litter, facilitated primarily by microbial decomposers, plays a critical role in biogeochemical cycling and ecosystem function. Emerging pathogens have the potential to impact litter decomposition by altering the chemical composition and associated microbial community of host tissue. Here, we compared litter decomposition of the invasive grass Microstegium vimineum collected from sites with Bipolaris leaf spot symptoms and sites with no apparent disease symptoms in a common garden experiment. Our results revealed that leaf tissue from litter from non-infected sites decomposed more rapidly through the spring than litter from infected sites. Differences in fungal composition between infected and non-infected litter at the start of the experiment largely persisted through the summer. Our work demonstrates that pathogen colonization may facilitate the persistence of infected host litter, potentially slowing the return of nutrients to the environmental pool while also promoting the survival and dispersal of primary inoculum the following season.

openCC (other)Jun 2023View details →
zenodo40/100

From a Monolithic Big Data System to a Microservices Event-Driven Architecture

<p>[Context] Data-intensive systems, a.k.a. big data systems (BDS), are software systems that handle a large volume of data in the presence of performance quality attributes, such as scalability and availability. Before the advent of big data management systems (e.g. Cassandra) and frameworks (e.g. Spark), organizations had to cope with large data volumes with custom-tailored solutions. In particular, a decade ago, Tecgraf/PUC-Rio developed a system to monitor truck fleet in real-time and proactively detect events from the positioning data received. Over the years, the system evolved into a complex and large obsolescent code base involving a hard maintenance process. [Goal] We report our experience on replacing a legacy BDS with a microservice-based event-driven system. [Method] We applied action research, investigating the reasons that motivate the adoption of a microservice-based event-driven architecture, intervening to define the new architecture, and documenting the challenges and lessons learned. [Results] We perceived that the resulting architecture enabled easier maintenance and fault-isolation. However, the myriad of technologies and the complex data flow were perceived as drawbacks. Based on the challenges faced, we highlight opportunities to improve the design of big data reactive systems. [Conclusions] We believe that our experience provides helpful takeaways for practitioners modernizing systems with data-intensive requirements.</p>

opencc-by-4.0Jan 2020View details →
zenodo40/100

Some Notable Records of Mayflies (Insecta: Ephemeroptera) from Big Rivers in Indiana: Supporting Data

<p>Significant records from July &amp; August 2019 fieldwork on big rivers in Indiana, United States.</p>

opencc-by-4.0Jun 2020View details →
dryad40/100

Data from: Not afraid of the Big Bad Wolf: calls from large predators do not silence mesopredators

<p>Large predators are known to shape the behavior and ecology of sympatric predators via conflict and competition, with mesopredators thought to avoid large predators, while dogs suppress predator activity and act as guardians of human property. However, interspecific communication between predators has not been well-explored and this assumption of avoidance may oversimplify the responses of the species involved. We explored the acoustic activity of three closely related sympatric canids: wolves <em>Canis lupus</em>, coyotes <em>Canis latrans</em>, and dogs <em>Canis familiaris</em>. These species have an unbalanced triangle of risk: coyotes, as mesopredators, are at risk from both apex-predator wolves and human-associated dogs, while wolves fear dogs, and dogs may fear wolves as apex predators or challenge them as intruders into human-allied spaces. We predicted that risk perception would dictate vocal response with wolves and dogs silencing coyotes as well as dogs silencing wolves. Dogs, in their protective role of guarding human property, would respond to both. Eleven passive acoustic monitoring devices were deployed across 13 nights in Central Wisconsin, and we measured the responses of each species to naturally occurring heterospecific vocalizations. Against our expectation, silencing did not occur. Instead, coyotes were not silenced by either species: when hearing wolves, coyotes responded at greater than chance rates and when hearing dogs, coyotes did not produce fewer calls than chance rates. Similarly, wolves responded at above chance rates to coyotes and at chance rates when hearing dogs. Only the dogs followed our prediction and responded at above chance rates in response to both coyotes and wolves. Thus, instead of silencing their competitors, canid vocalizations elicit responses from them suggesting the existence of a complex heterospecific communication network.</p>

opencc-zeroFeb 2024View details →
zenodo40/100

Fig. 3 in Helping Protists to Find Their Place in a Big Data World

Fig. 3. Selected component of Catalogue of Life (URL 17), showing that the Family Cyrtolophosidae is classified in three locations (with variant spellings).

opencc-by-4.0Dec 2014View details →
zenodo40/100

Fig 2 in Helping Protists to Find Their Place in a Big Data World

Fig 2. Reconciliation of alternative names for the same taxon (an invasive diatom species). The diagram shows three classes of 'names': scientific names, vernacular names, and surrogates or strings that act in the same way as names (sequence data in this example). Gomphonema vulgare and Echinella geminata were applied independently to the same species and are heterotypic synonyms. The reconciliation group includes the homotypic synonyms (Echinella geminata and Didymosphenia geminata), and the lexical variants of all names. Reconciliation groups allow computer-based queries initiated with one name to be answered with information associated with all names.

opencc-by-4.0Dec 2014View details →
zenodo40/100

Fig. 1 in Helping Protists to Find Their Place in a Big Data World

Fig. 1. Modular' model for the infrastructure of a big data world. a – Within a module, nodes obtain content from one or more sources, normalize, enrich, and deliver it to end users. Annotation systems allow users to advise the source and nodes as to the quality of content. b – Nodes interconnect in anarchic ways that allow for evolution and expanding functionality.

opencc-by-4.0Dec 2014View details →
zenodo40/100

Spatiotemporal dataset of dengue influencing factors in Brazil based on geospatial big data cloud computing

<p>We produced a spatiotemporal dataset of dengue influencing factors in Brazil based on geospatial big data cloud computing from 2001-2024.</p> <p>GDP and building surface area are yearly data.</p> <p>PDSI is monthly data.</p>

opencc-by-4.0Nov 2024View details →
zenodo40/100

Model weights for a Weather4cast 2021 Challenge IEEE Big Data Cup Stage solution

<p>This repository contains the pre-trained model weights for the TensorFlow/Keras models used in the <a href="https://www.iarai.ac.at/weather4cast/2021-competition/challenge/">Weather4cast 2021 Challenge IEEE Big Data Cup Stage</a> by the team &quot;antfugue&quot;. The model code can be found in <a href="https://github.com/jleinonen/weather4cast-bigdata">https://github.com/jleinonen/weather4cast-bigdata</a> along with instructions on where to extract the weights.</p>

opencc-by-4.0Nov 2021View details →
zenodo40/100

Probabilistic simulation of big climate data for robust quantification of changes in compound hazard events

<p>Data, code and supplementary Figures for paper &quot;Probabilistic simulation of big climate data for robust quantification of changes in compound hazard events&quot;.</p>

opencc-by-4.0Feb 2022View details →
zenodo40/100

Automotive CAN bus data: An Example Dataset from the AEGIS Big Data Project

<p>Here you find an example research data dataset for the automotive demonstrator within the &quot;AEGIS -&nbsp; Advanced Big Data Value Chain for Public Safety and Personal Security&quot; big data project, which has received funding from the European Union&rsquo;s Horizon 2020 research and innovation programme under grant agreement No 732189. The time series data has been collected during trips conducted by three drivers driving the same vehicle in Austria.</p> <p>The dataset contains 20Hz sampled CAN bus data from a passenger vehicle, e.g. WheelSpeed FL (speed of the front left wheel), SteerAngle (steering wheel angle), Role, Pitch, and accelerometer values per direction.</p> <p>GPS data from the vehicle (see signals &#39;Latitude_Vehicle&#39; and &#39;Longitude_Vehicle&#39; in h5 group &#39;Math&#39;) and GPS data from the IMU device (see signals &#39;Latitude_IMU&#39;, &#39;Longitude_IMU&#39; and &#39;Time_IMU&#39; in h5 group &#39;Math&#39;) are included. However, as it had to be exported with single-precision, we lost some precision for those GPS values.</p> <p>&nbsp;</p> <p>For data analysis we use R and R Studio (https://www.rstudio.com/) and the library h5.</p> <p>e.g. check file with R code:</p> <p>library(h5)</p> <p>f &lt;- h5file(&quot;file path/20181113_Driver1_Trip1.hdf&quot;)</p> <p>summary(f[&quot;CAN/Yawrate1&quot;][,])</p> <p>summary(f[&quot;Math/Latitude_IMU&quot;][,])</p> <p>h5close(f)</p>

opencc-by-4.0Jul 2019View details →
zenodo40/100

Dataset for: "Big data suggest strong constraints of linguistic similarity on adult language learning"

<p>This dataset is adapted from raw data with fully anonymized results on the State Examination of Dutch as a Second Language. This exam is officially administred by the Board of Tests and Examinations (College voor Toetsen en Examens, or CvTE). See cvte.nl/about-cvte. The Board of Tests and Examinations is mandated by the Dutch government.</p> <p>The article accompanying the dataset:</p> <p>Schepens, Job, Roeland van Hout, and T. Florian Jaeger. &ldquo;Big Data Suggest Strong Constraints of Linguistic Similarity on Adult Language Learning.&rdquo; <em>Cognition</em> 194 (January 1, 2020): 104056. <a href="https://doi.org/10.1016/j.cognition.2019.104056">https://doi.org/10.1016/j.cognition.2019.104056</a>.</p> <p>Every row in the dataset represents the first official testing score of a unique learner.<br> The columns contain the following information as based on questionnaires filled in at the time of the exam:</p> <p>&quot;L1&quot; - The first language of the learner<br> &quot;C&quot; - The country of birth<br> &quot;L1L2&quot; - The combination of first and best additional language besides Dutch<br> &quot;L2&quot; - The best additional language besides Dutch<br> &quot;AaA&quot; - Age at Arrival in the Netherlands in years (starting date of residence)<br> &quot;LoR&quot; - Length of residence in the Netherlands in years<br> &quot;Edu.day&quot; - Duration of daily education (1 low, 2 middle, 3 high, 4 very high). From 1992 until 2006, learners&#39; education has been measured by means of a side-by-side matrix question in a learner&#39;s questionnaire. Learners were asked to mark which type of education they have had (elementary, secondary, or tertiary schooling) by means of filling in for how many years they have been enrolled, in which country, and whether or not they have graduated. Based on this information we were able to estimate how many years learners have had education on a daily basis from six years of age onwards. Since 2006, the question about learners&#39; education has been altered and it is asked directly how many years learners have had formal education on a daily basis from six years of age onwards. Possible answering categories are: 1) 0 thru 5 years; 2) 6 thru 10 years; 3) 11 thru 15 years; 4) 16 years or more. The answers have been merged into the categorical answer.<br> &quot;Sex&quot; - Gender<br> &quot;Family&quot; - Language Family<br> &quot;ISO639.3&quot; - Language ID code according to Ethnologue<br> &quot;Enroll&quot; - Proportion of school-aged youth enrolled in secondary education according to the World Bank. The World Bank reports on education data in a wide number of countries around the world on a regular basis. We took the gross enrollment rate in secondary schooling per country in the year the learner has arrived in the Netherlands as an indicator for a country&#39;s educational accessibility at the time learners have left their country of origin.<br> &quot;STEX_speaking_score&quot; - The STEX test score for speaking proficiency.<br> &quot;Dissimilarity_morphological&quot; - Morphological similarity<br> &quot;Dissimilarity_lexical&quot; - Lexical similarity<br> &quot;Dissimilarity_phonological_new_features&quot; - Phonological similarity (in terms of new features)<br> &quot;Dissimilarity_phonological_new_categories&quot; - Phonological similarity (in terms of new sounds)</p> <p><br> A few rows of the data:</p> <p>&quot;L1&quot;,&quot;C&quot;,&quot;L1L2&quot;,&quot;L2&quot;,&quot;AaA&quot;,&quot;LoR&quot;,&quot;Edu.day&quot;,&quot;Sex&quot;,&quot;Family&quot;,&quot;ISO639.3&quot;,&quot;Enroll&quot;,&quot;STEX_speaking_score&quot;,&quot;Dissimilarity_morphological&quot;,&quot;Dissimilarity_lexical&quot;,&quot;Dissimilarity_phonological_new_features&quot;,&quot;Dissimilarity_phonological_new_categories&quot;<br> &quot;English&quot;,&quot;UnitedStates&quot;,&quot;EnglishMonolingual&quot;,&quot;Monolingual&quot;,34,0,4,&quot;Female&quot;,&quot;Indo-European&quot;,&quot;eng &quot;,94,541,0.0094,0.083191,11,19<br> &quot;English&quot;,&quot;UnitedStates&quot;,&quot;EnglishGerman&quot;,&quot;German&quot;,25,16,3,&quot;Female&quot;,&quot;Indo-European&quot;,&quot;eng &quot;,94,603,0.0094,0.083191,11,19<br> &quot;English&quot;,&quot;UnitedStates&quot;,&quot;EnglishFrench&quot;,&quot;French&quot;,32,3,4,&quot;Male&quot;,&quot;Indo-European&quot;,&quot;eng &quot;,94,562,0.0094,0.083191,11,19<br> &quot;English&quot;,&quot;UnitedStates&quot;,&quot;EnglishSpanish&quot;,&quot;Spanish&quot;,27,8,4,&quot;Male&quot;,&quot;Indo-European&quot;,&quot;eng &quot;,94,537,0.0094,0.083191,11,19<br> &quot;English&quot;,&quot;UnitedStates&quot;,&quot;EnglishMonolingual&quot;,&quot;Monolingual&quot;,47,5,3,&quot;Male&quot;,&quot;Indo-European&quot;,&quot;eng &quot;,94,505,0.0094,0.083191,11,19</p>

opencc-by-4.0Aug 2019View details →
zenodo40/100

Industrial Big Data Innovation Platform SCADA Dataset

<p>This is the SCADA dataset from the first Industrial Big Data Innovation Competition, for more information, visit: <a href="https://www.industrial-bigdata.com/Challenge/title?competitionId=LEIREZMM8TT5VBU0TLJ61FPAI6WWJOJY&amp;type=">数境创新大赛平台(industrial-bigdata.com)。</a></p>

opencc-by-4.0Aug 2024View details →
zenodo40/100

V 1.0 Dataset for "Emergence and Evolution of Big Data Research: A 30-year (1993-2022) Scientometric Analysis of The Knowledge Field"

<p>This dataset includes the bibliometric data used in the scientometric analysis of the field of big data research over a 30-year period (1993-2022). The data was collected from the Scopus database, and contains information on 70,163 articles and 315,235 author keywords. The dataset is structured by 17 interrelated data categories that trace the conceptual emergence and evolution of the big data field, focusing on keyword co-occurrences, disciplinary distributions, and the temporal growth of publications. This dataset supports the analyses presented in the related manuscript.</p>

opencc-by-4.0Aug 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record