Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

29

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

29 results for “Large-scale Assessment”

Learn how ShareScore rates datasets ↗
zenodo56/100

Dataset for algorithmic thinking skills assessment: Results from the virtual CAT large-scale study in Swiss compulsory education

<p><strong>Overview</strong><br>This dataset was collected during a main study that evaluated the virtual Cross Array Task (CAT) platform as an assessment tool for algorithmic thinking (AT) skills among K-12 students in Swiss compulsory education.<br>As algorithmic thinking becomes increasingly vital in our digital age, this study bridges the gap between traditional assessments and the needs of today's learners by introducing a digital platform. The virtual CAT, a digital adaptation of an unplugged assessment activity, offers scalable, automated assessments with reduced human intervention.</p> <p><strong>Study Context, Location and Participants</strong><br>To comprehensively investigate algorithmic competencies within compulsory education, exploring their variations and determining the factors influencing them, in Spring 2023 we conducted an experimental study with the virtual CAT's.<br>The sample comprises 129 students (65 girls and 64 boys), selected from nine classes across five public schools in Ticino and Solothurn cantons.</p> <p><strong>Data Collection</strong><br>During the data collection process, session and participant details were manually recorded by the administrator. <br>Each session has been assigned a unique identifier, and specific details, such as the date, canton, school name and type, and the students&rsquo; HarmoS grade (HG) level, have been recorded.&nbsp;<br>Student information are limited to sex and date of birth, with birth dates used to calculate ages, a significant factor in our demographic analysis. <br>To protect student privacy, unique identifiers have been assigned to each participant, keeping the data anonymous and secure. <br>The assessment tool automatically tracked all user interaction within the platform.<br>All data collected have been pseudonymised, aligning with prevailing open science practices in Switzerland (SNSF, 2021).&nbsp;<br>Data collection was integrated into a validation module of the app.&nbsp;</p> <p><strong>Data Features</strong><br>The dataset comprises the following files:</p> <ul> <li>STUDENTS_SESSIONS.csv</li> <li>RESULTS.csv</li> <li>LOGS.csv</li> <li>CANTONS.csv</li> <li>ALGORITHMS.csv</li> </ul> <p>These files collectively provide insights into the algorithmic actions of the students, demographic details, session logs, results, and more.</p> <p><strong>Usage &amp; Ethics</strong><br>In the spirit of open science, this dataset is made available to the public after meticulous anonymisation to ensure all participants' privacy and ethical treatment.&nbsp;<br>Initial authorisations were secured from school administrators, teachers, and parents.&nbsp;<br>Detailed communication regarding the study's nature, data handling, and objectives was transparently shared with all stakeholders.</p> <p><strong>REFERENCES</strong></p> <p><strong>[1]</strong>&nbsp;A. Piatti, G. Adorni, L. El-Hamamsy, L. Negrini, D. Assaf, L. Gambardella &amp; F. Mondada. (2022). The CT-cube: A framework for the design and the assessment of computational thinking activities. Computers in Human Behavior Reports, 5, 100166.&nbsp;<a href="https://doi.org/10.1016/j.chbr.2021.100166">https://doi.org/10.1016/j.chbr.2021.100166</a></p> <p><strong>[2]</strong>&nbsp;Adorni, G., &amp; Piatti, S., &amp; Karpenko, V. (2023). virtual CAT: An app for algorithmic thinking assessment within Swiss compulsory education. Zenodo Software.&nbsp;<a href="https://doi.org/10.5281/zenodo.10027851">https://doi.org/10.5281/zenodo.10027851</a>&nbsp;On GitHub:&nbsp;<a href="https://github.com/GiorgiaAuroraAdorni/virtual-CAT-app/">https://github.com/GiorgiaAuroraAdorni/virtual-CAT-app/</a></p> <p><strong>[3]</strong>&nbsp;Adorni, G., &amp; Karpenko, V. (2023). virtual CAT programming language interpreter. Zenodo Software.&nbsp;<a href="https://doi.org/10.5281/zenodo.10016535">https://doi.org/10.5281/zenodo.10016535</a>&nbsp;On GitHub:&nbsp;<a href="https://github.com/GiorgiaAuroraAdorni/virtual-CAT-programming-language-interpreter/">https://github.com/GiorgiaAuroraAdorni/virtual-CAT-programming-language-interpreter/</a></p> <p><strong>[4]</strong>&nbsp;Adorni, G., &amp; Karpenko, V. (2023). virtual CAT data infrastructure. Zenodo Software.&nbsp;<a href="https://doi.org/10.5281/zenodo.10015011">https://doi.org/10.5281/zenodo.10015011</a>&nbsp;On GitHub:&nbsp;<a href="https://github.com/GiorgiaAuroraAdorni/virtual-CAT-data-infrastructure">https://github.com/GiorgiaAuroraAdorni/virtual-CAT-data-infrastructure</a></p> <p>&nbsp;</p>

opencc-by-4.0Apr 2024View details →
zenodo40/100

Data supporting "Large-scale citizen science programs can support ecological and climate change assessments"

<p>Text file of phenology observations pulled from the USA National Phenology Network&#39;s database (www.usanpn.org) and used in this analysis.&nbsp;</p>

opencc-by-4.0Jan 2022View details →
zenodo40/100

Assessment of the acoustic adaptation hypothesis in frogs using large-scale citizen science data

<p>This is the data required to reproduce the results of the manuscript &quot;Assessment of the acoustic adaptation hypothesis in frogs using large-scale citizen science data&quot;, including measurements of tree canopy cover extracted from the Global Forest Cover Change dataset (Townshend 2016).</p> <p>&nbsp;</p> <p><strong>Reference</strong></p> <p>Gillard, G. L. &amp;&nbsp;Rowley, J. J. L. (2023). Assessment of the acoustic adaptation hypothesis in frogs using large-scale citizen science data.&nbsp;<em>Journal of Zoology</em>. [In publication].</p> <p>&nbsp;</p> <p><strong>Global Forest Cover Change Dataset</strong></p> <p>Townshend J. 2016. Global Forest Cover Change (GFCC) Tree Cover Multi-Year Global 30 m V003 [Data set]. NASA EOSDIS Land Processes DAAC. Accessed June 22, 2022. doi:10.5067/MEaSUREs/GFCC/GFCC30TC.003.Townshend J. 2016. Global Forest Cover Change (GFCC) Tree Cover Multi-Year Global 30 m V003 [Data set]. NASA EOSDIS Land Processes DAAC. Accessed June 22, 2022. doi:10.5067/MEaSUREs/GFCC/GFCC30TC.003.</p>

opencc-by-4.0May 2023View details →
dryad36/100

A large-scale assessment of plant dispersal mode and seed traits across human-modified Amazonian forests

1. Quantifying the impact of habitat disturbance on ecosystem function is critical for understanding and predicting the future of tropical forests. Many studies have examined post-disturbance changes in animal traits related to mutualistic interactions with plants, but the effect of disturbance on plant traits in diverse forests has received much less attention. 2. Focusing on two study regions in the eastern Brazilian Amazon, we used a trait-based approach to examine how seed dispersal functionality within tropical plant communities changes across a landscape-scale gradient of human modification, including both regenerating secondary forests and primary forests disturbed by burning and selective logging. 3. Surveys of 230 forest plots recorded 26,533 live stems from 846 tree species. Using herbarium material and literature, we compiled trait information for each tree species, focusing on dispersal mode and seed size. 4. Disturbance reduced tree diversity and increased the proportion of lower wood-density and smaller-seeded tree species in study plots. Unexpectedly, disturbance also increased the proportion of stems with seeds that are ingested by animals and reduced those dispersed by other mechanisms (e.g. wind). Older secondary forests had functionally similar plant communities to the most heavily disturbed primary forests. 5. <i>Synthesis</i>. Anthropogenic disturbance has major effects on the seed traits of tree communities, with implications for mutualistic interactions with animals. The higher importance of animal-mediated seed dispersal in disturbed and recovering forests highlights the importance of avoiding defaunation or promoting faunal recovery. The changes in mean seed width suggest larger vertebrates hold especially important functional roles in these human-modified forests. Monitoring fruit and seed traits can provide a valuable indicator of ecosystem condition, emphasising the importance of developing a comprehensive plant traits database for the Amazon and other biomes.

opencc-zeroJan 2020View details →
dryad36/100

A large-scale assessment of ant diversity across the Brazilian Amazon Basin: integrating geographic, ecological, and morphological drivers of sampling bias

<p>Tropical ecosystems are often biodiversity hotspots, and invertebrates represent the main underrepresented component of diversity in large-scale analyses. This problem is partly related to the scarcity of data widely available to conduct these studies and the lack of systematic organization of knowledge about invertebrates' distributions in biodiversity hotspots. Here, we introduce and analyze a comprehensive data compilation of Amazonian ant diversity. Using records from 1817 to 2020 from both published and unpublished sources, we describe the diversity and distribution of ant species in the Brazilian Amazon Basin. Further, using high-definition images and data from taxonomic publications, we build a comprehensive database of morphological traits for the ant species that occur in the region. In total, we recorded 1,067 nominal species in the Brazilian Amazon Basin, with sampling locations strongly biased by access routes, urban centers, research institutions, and major infrastructure projects. Large areas where ant sampling is non-existent represent about 52% of the basin and are concentrated mainly in the North, Southeastern, and Western Brazilian Amazon. We found that distance to roads is the main driver of ant sampling in the Amazon. Contrary to our expectations, morphological traits had lower predictive power in predicting sample bias than purely geographic variables. However, when geographic predictors were controlled, habitat stratum and traits contribute to explain the remaining variance. More species were recorded in better-sampled areas, but species richness estimation models suggest that areas in South Amazonian edge forests are associated with especially high species richness. Our results represent the first trait-based, large-scale study for insects in Amazonian forests and a starting point for macroecological studies focusing on insect diversity in the Amazon Basin.</p>

opencc-zeroMay 2022View details →
zenodo36/100

Assessing Vaccine-Related Content for Journalistic Quality: A large-scale dataset and article repository

<p>This dataset was produced through a collaboration with the NSF-funded <a href="https://artt.cs.washington.edu/">ARTT project</a> (led by Hacks/Hackers) and <a href="https://overtone.ai/">Overtone</a>. It consists of 1,000 vaccine-related articles, pulled from a wide variety of news media sources, with associated scores based on their journalistic quality. The scores were provided through Overtone&rsquo;s algorithm, and range from one (low-quality or low informational value add) to five (high-quality or high informational value add). Articles were sourced from&nbsp;traditional journalism outlets (news and news-leaning websites), as well as non-journalistic sources of vaccine information, such as governmental websites, healthcare and NGO websites, and medical journals.&nbsp;Given the algorithm&rsquo;s focus on editorial content, as opposed to other metrics such as author, outlet, or engagement, analyzing a diverse set of article types allowed the research team&nbsp;to examine how different styles of vaccine-related content measured against traditional journalistic quality standards.<strong>&nbsp;</strong>Therefore, this dataset provides a unique insight into the spectrum of vaccine reporting, and serves as a contribution to the field of&nbsp;automated quality assessment.&nbsp;</p> <p><strong>About ARTT:</strong> The Analysis and Response Toolkit for Trust (ARTT) project is focused on helping people engage in trust-building ways when discussing vaccine efficacy and other topics online.&nbsp;</p> <p><strong>About Overtone:</strong>&nbsp;Overtone has built a Natural Language Processing algorithm that finds and sorts online content by its intrinsic qualities, rather than clicks or shares. Their AI assesses texts for journalistic signals that demonstrate human effort.</p> <p>For any questions about this dataset, please contact artt@hackshackers.com.</p>

opencc-by-nc-nd-1.0Aug 2022View details →
zenodo36/100

Large-Scale Multipurpose Benchmark Datasets For Assessing Data-Driven Deep Learning Approaches For Water Distribution Networks

<p>&nbsp;</p> <div> <div><a href="https://arxiv.org/search/cs?searchtype=author&amp;query=Tello,+A">Andres Tello*</a><em>, </em><a href="https://arxiv.org/search/cs?searchtype=author&amp;query=Truong,+H">Huy Truong*</a>, <a href="https://arxiv.org/search/cs?searchtype=author&amp;query=Lazovik,+A">Alexander Lazovik</a>, <a href="https://arxiv.org/search/cs?searchtype=author&amp;query=Degeler,+V">Victoria Degeler</a>. Large-Scale Multipurpose Benchmark Datasets For Assessing Data-Driven Deep Learning Approaches For Water Distribution Networks. Engineering Proceedings. 2024; 69(1):50. <a href="https://doi.org/10.3390/engproc2024069050">https://doi.org/10.3390/engproc2024069050</a></div> <br> <div>(*) Both authors contributed equally.<br><br></div> <h2>Update</h2> <div>(04/09/2024): Citation is updated.<br>We have added headers for CSVs and auxiliary data (duration time, edge list, ordered names.. ) in the configuration file (JSON format). As such, corresponding INP files can be omitted when working with this version.&nbsp;<br>The EXN network has been included in this version, so the total number of processed networks is 11.<br>For more details, please read ZENODO_README.md.</div> <h2>Contact</h2> <div>For dataset-related questions: <a href="mailto:h.c.truong@rug.nl" target="_blank" rel="noopener">Huy Truong</a></div> <br> <div>For data acquisition: <a href="mailto:a.tello@rug.nl" target="_blank" rel="noopener">Andres Tello</a></div> <br> <div>If you use this dataset, please cite:</div> <blockquote>@article{tello2024largescale,<br>&nbsp; &nbsp; AUTHOR = {Tello, Andr&eacute;s and Truong, Huy and Lazovik, Alexander and Degeler, Victoria},<br>&nbsp; &nbsp; TITLE = {Large-Scale Multipurpose Benchmark Datasets for Assessing Data-Driven Deep Learning Approaches for Water Distribution Networks},<br>&nbsp; &nbsp; JOURNAL = {Engineering Proceedings},<br>&nbsp; &nbsp; VOLUME = {69},<br>&nbsp; &nbsp; YEAR = {2024},<br>&nbsp; &nbsp; NUMBER = {1},<br>&nbsp; &nbsp; ARTICLE-NUMBER = {50},<br>&nbsp; &nbsp; URL = {https://www.mdpi.com/2673-4591/69/1/50},<br>&nbsp; &nbsp; ISSN = {2673-4591},<br>&nbsp; &nbsp; DOI = {10.3390/engproc2024069050}<br>}</blockquote> </div>

opencc-by-4.0May 2024View details →
zenodo36/100

A Large-scale Benchmark for Technical Debt Assessment

<p>This repository hosts the supplementary material of the paper entitled "A Large-scale Benchmark for Technical Debt Assessment". It comprises a very large-scale study on the Technical Debt of more than 55 million open-source Java files.</p>

opencc-by-4.0Oct 2023View details →
dryad36/100

A large-scale assessment of ant diversity across the Brazilian Amazon Basin: integrating geographic, ecological, and morphological drivers of sampling bias

Open the record for dataset details and reuse information.

publicJun 2022View details →
dryad36/100

A large-scale assessment of plant dispersal mode and seed traits across human-modified Amazonian forests

Open the record for dataset details and reuse information.

publicFeb 2020View details →
dryad32/100

Data from: A new integrative framework for large-scale assessments of biodiversity and community dynamics, using littoral gastropods and crabs of British Columbia, Canada

Improving our understanding of species responses to environmental changes is an important contribution ecologists can make to facilitate effective management decisions. Novel synthetic approaches to assessing biodiversity and ecosystem integrity are needed, ideally including all species living in a community and the dynamics defining their ecological relationships. Here we present and apply an integrative approach that links high-throughput, multi-character taxonomy with community ecology. The overall purpose is to enable the coupling of biodiversity assessments with investigations into the nature of ecological interactions in a community-level data set. We collected 1,195 gastropods and crabs in British Columbia. First, the General mixed Yule-coalescent (GMYC) and the Poisson Tree Processes (PTP) methods for proposing primary species-hypotheses based on cox1 sequences were evaluated against an integrative taxonomic framework. We then used data on the geographic distribution of delineated species to test species co-occurrence patterns for non-randomness using community-wide and pairwise approaches. Results showed that PTP generally outperformed GMYC and thus constitutes a more effective option for producing species-hypotheses in community-level datasets. Non-random species co-occurrence patterns indicative of ecological relationships or habitat preferences were observed for grazer gastropods, whereas assemblages of opportunistic omnivorous gastropods and crabs appeared influenced by random processes. Species-pair associations were consistent with current ecological knowledge, thus suggesting that applying community assembly within a large taxonomical framework constitutes a valuable tool for assessing ecological interactions. Combining phylogenetic, morphological and co-occurrence data enabled an integrated view of communities, providing both a conceptual and pragmatic framework for biodiversity assessments and investigations into community dynamics.

opencc-zeroDec 2015View details →
dryad32/100

Data from: Automated DNA-based plant identification for large-scale biodiversity assessment

Rapid degradation of tropical forests urges to improve our efficiency in large-scale biodiversity assessment. DNA-barcoding can assist greatly in this task, but commonly used phenetic approaches for DNA-based identifications rely on the existence of comprehensive reference databases, which are infeasible for hyperdiverse tropical ecosystems. Alternatively, phylogenetic methods are more robust to sparse taxon sampling but time-consuming, while multiple alignment of species-diagnostic, typically length-variable markers can be problematic across divergent taxa. We advocate the combination of phylogenetic and phenetic methods for taxonomic assignment of DNA-barcode sequences against incomplete reference databases such as GenBank, and we developed a pipeline to implement this approach on large-scale plant diversity projects. The pipeline workflow includes several steps: database construction and curation, query sequence clustering, sequence retrieval, distance calculation, multiple alignment and phylogenetic reconstruction. We describe the strategies used to establish these steps and the optimisation of parameters to fit the selected psbA-trnH marker. We tested the pipeline using infertile plant samples and herbivore diet sequences from the highly threatened Nicaraguan seasonally dry forest and exploiting a valuable purpose-built resource: a partial local reference database of plant psbA-trnH. The selected methodology proved efficient and reliable for high-throughput taxonomic assignment, and our results corroborate the advantage of applying 'strict' tree-based criteria to avoid false positives. The pipeline tools are distributed as the scripts suite 'BAGpipe' (pipeline for Biodiversity Assessment using GenBank data), which can be readily adjusted to the purposes of other projects and applied to sequence-based identification for any marker or taxon.

opencc-zeroDec 2013View details →
zenodo32/100

Supporting Information: Environmental benefits of large-scale second-generation bioethanol production in the EU: An integrated supply chain network optimization and Life Cycle Assessment approache

<p>This supporting information provides all input data of the model, the assumptions made and the literature and database references for the publication<em> Environmental benefits of large-scale second-generation bioethanol production in the EU: An integrated supply chain network optimization and Life Cycle Assessment approache</em>. The environmental data is based on life cycle assessments, with full information on the life cycle inventory and the results of the life cycle impact assessment based on the ReCiPe method. It also includes detailed results for all objective functions in all scenarios (optimization of 18 midpoints, 3 endpoints, and economic optimization in 5 tax scenarios and 2 feedstock scenarios), and detailed results of the sensitivity analysis and Pareto optimization.</p>

opencc-by-4.0Aug 2020View details →
ClinicalTrials.gov32/100

A 3-year Cross-sectional Assessment of the Tianjin Child and Adolescent Large-scale Eye Study

ClinicalTrials.gov study NCT05967702. IPD Sharing: NO. Countries: 1. Publications: 3.

closedIPD-NOFeb 2026View details →
dryad32/100

Data from: Is environmental legislation conserving tropical stream faunas? a large-scale assessment of local, riparian and catchment-scale influences on Amazonian stream fish

Open the record for dataset details and reuse information.

publicSep 2018View details →
dryad32/100

Data from: A new integrative framework for large-scale assessments of biodiversity and community dynamics, using littoral gastropods and crabs of British Columbia, Canada

Open the record for dataset details and reuse information.

publicMar 2016View details →
dryad32/100

Data from: Automated DNA-based plant identification for large-scale biodiversity assessment

Open the record for dataset details and reuse information.

publicMar 2014View details →
dryad32/100

Data from: A comprehensive large-scale assessment of fisheries bycatch risk to threatened seabird populations

Open the record for dataset details and reuse information.

publicMay 2019View details →
dryad28/100

Data from: Large-scale fungal diversity assessment in the Andean Yungas forests reveals strong community turnover among forest types along an altitudinal gradient

The Yungas, a system of tropical and subtropical montane forests on the eastern slopes of the Andes, are extremely diverse and severely threatened by anthropogenic pressure and climate change. Previous mycological works focused on macrofungi (e.g., agarics, polypores) and mycorrhizae in Alnus acuminata forests, while fungal diversity in other parts of the Yungas has remained mostly unexplored. We carried out Ion Torrent sequencing of ITS2 rDNA from soil samples taken at 24 sites along the entire latitudinal extent of the Yungas in Argentina. The sampled sites represent the three altitudinal forest types: the piedmont (400–700 masl), montane (700–1500 masl), and montane cloud (1500–3000 masl) forests. The deep sequence data presented here (i.e. 4 108 126 quality-filtered sequences) indicate that fungal community composition correlates most strongly with elevation, with many fungi showing preference for a certain altitudinal forest type. For example, ectomycorrhizal and root endophytic fungi were most diverse in the montane cloud forests, particularly at sites dominated by Alnus acuminata, while the diversity values of various saprobic groups were highest at lower elevations. Despite the strong altitudinal community turnover, fungal diversity was comparable across the different zonal forest types. Besides elevation, soil pH, N, P, and organic matter contents correlated with fungal community structure as well, although most of these variables were co-correlated with elevation. Our data provide an unprecedented insight into the high diversity and spatial distribution of fungi in the Yungas forests.

opencc-zeroDec 2013View details →
zenodo28/100

R code: Large-scale assessment of bird data quality in a citizen science platform

<p>R code archived to Zenodo for publication</p>

opencc-by-4.0Jan 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record