Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
486
datasets available to search
ShareScore release 0.7.1
Dataset results
486 results for “curation”
FIGURE 4 in The Megachilidae (Hymenoptera, Apoidea, Apiformes) of the Democratic Republic of Congo curated at the Royal Museum for Central Africa (RMCA, Belgium)
FIGURE 4. Taxonomic and temporal patterns of Megachilidae specimen records in the RMCA collection (Tervuren, Belgium). A. Rank abundance plot and associated data for the 10 most frequently collected species showing that the top 8 species make up to 50% of the dataset. The distribution of species along the curve shows a very long "tail", implying that many of the species curated are only represented by a low or a very low number of species. At the extreme, 26 species of Megachilidae were represented in the RMCA collections only as type specimens (holotype and paratype); B. Specimen accumulation curve through time (1905-1978) for the seven groups of species in the family Megachilidae showing that the best represented groups are Megachile sensu lato, Gronoceras, Lithurgini and Anthidiini (by decreasing order of magnitude). It is noteworthy that Megachile sensu lato has been regularly collected and in higher numbers during the period 1905-1920, and that all the above groups were collected in larger numbers during the decade 1930-40 compared to other groups of Megachilidae.
BOLD release 5 April 2024, curated with pipeline commit d7f034f1e14f15daed708bb3e0e8ddf1b50e0249
Open the record for dataset details and reuse information.
Test Data for iwc Pre-curation PretextMap generation workflow
Open the record for dataset details and reuse information.
CofeXHug: A curated dataset of HuggingFace pre-trained models exploited in the GitHub ecosystem
<p> Pre-trained models (PTMs) are becoming increasingly popular in the software engineering community. Their usage is facilitated by model repositories, e.g., HuggingFace, which collect, store, and maintain a wide range of PTMs. However, the actual adoption of these models in real-world projects is still an open question. In particular, many of them are used in toy projects or simply as a mirror for the HF repository. Thus, we see the need for a curated codebase related to PTMs to support developers and practitioners who are interested in using them in their projects.<br>This artifact contains CodeXHug, a curated dataset of HuggingFace PTMs exploited in the GitHub ecosystem. Starting from the latest HF dump, we first conduct a data curation to collect PTMs with a tag and a model card. Then, the GitHub platform has been queried to find actual usages of the identified PTMs, resulting in 7,325 different models and 372,063 Python files. We also present a statistical analysis of the dataset, highlighting the most popular PTMs and the most common tasks for which they are used. Finally, we discuss the research opportunities enabled by CodeXHug and the implications of our findings for the software engineering community.</p>
Erkomaishvili Dataset: A Curated Corpus of Traditional Georgian Vocal Music for Computational Musicology
<p><strong>Abstract</strong></p> <p>The analysis of recorded audio material using computational methods has received increased attention in ethnomusicological research. We present a curated dataset of traditional Georgian vocal music for computational musicology. The corpus is based on historic tape recordings of three-voice Georgian songs performed by the the former master chanter Artem Erkomaishvili. In this article, we give a detailed overview on the audio material, transcriptions, and annotations contained in the dataset. Beyond its importance for ethnomusicological research, this carefully organized and annotated corpus constitutes a challenging scenario for music information retrieval tasks such as fundamental frequency estimation, onset detection, and score-to-audio alignment. The corpus is publicly available and accessible through score-following web-players.</p> <p><strong>License</strong></p> <p>This work is licensed under the Creative Commons Attribution-NonCommercial 4.0 International License. To view a copy of this license, visit http://creativecommons.org/licenses/by-nc/4.0/ or send a letter to Creative Commons, PO Box 1866, Mountain View, CA 94042, USA.</p> <p><strong>Copyright of Audio (wav)</strong></p> <p>Ministry of Culture, Sports and Youth of Georgia<br> Legal Entity of Public Law<br> Vano Sarajishvili Tbilisi State Conservatoire (TSC)<br> 8-10, GRIBOEDOV St, TBILISI 0108, GEORGIA Tel. / fax :(+995 32) 2 999 144,<br> www.tsc.edu.ge E-mail: info@tsc.edu.ge; inter@tsc.edu.ge</p> <p>We thank the rector of TSC, Nana Sharikadze, for the permission to publish the recordings along with our annotations on Zenodo.</p> <p><strong>Copyright of Annotations (csv)</strong></p> <p>Sebastian Rosenzweig^1, Frank Scherbaum^2, David Shugliashvili^3, Vlora Arifi-Müller^1, and Meinard Müller^1<br> ^1: International Audio Laboratories Erlangen, Germany<br> ^2: University of Potsdam, Germany<br> ^3: Tbilisi State Conservatoire, Georgia</p> <p>The provided digital sheet music in MusicXML-format is based on the transcriptions by David Shugliashvili as published in the book:</p> <p>David Shugliashvili<br> Georgian Church Hymns, Shemokmedi School<br> Georgian Chanting Foundation, 2014.</p> <p><strong>References</strong></p> <p>If you use the Erkomaishvili dataset in your research, please cite:</p> <p>Sebastian Rosenzweig, Frank Scherbaum, David Shugliashvili, Vlora Arifi-Müller, and Meinard Müller<br> Erkomaishvili Dataset: A Curated Corpus of Traditional Georgian Vocal Music for Computational Musicology<br> Transactions of the International Society for Music Information Retrieval (TISMIR), 3(1): 31–41, 2020.</p>
Aligned and curated mtDNA sequences from: Ancient DNA reveals interstadials as a driver of common vole population dynamics during the last glacial period
<p><strong><span>Aim: </span></strong><span>Many species experienced population turnover and local extinction during the Late Pleistocene. In the case of megafauna, it remains challenging to disentangle climate change and the activities of Palaeolithic hunter-gatherers as the main cause. In contrast, the impact of humans on rodent populations </span><span>is likely to be negligible. This study investigated which climatic and/or environmental factors affect the population dynamics of the common vole. </span><span>This temperate rodent is widespread across Europe and was one of the most abundant small mammal species throughout the Late Pleistocene.</span></p> <p><span><strong>Location:</strong> </span><span>Europe</span></p> <p><strong><span>Taxon: </span></strong><span>Common vole (<em>Microtus arvalis</em>)</span></p> <p><strong><span>Methods: </span></strong><span>We generated a dataset comprised of a 4.2-kb-long fragment of mitochondrial DNA (mtDNA) from 148 ancient and 51 modern specimens sampled from multiple localities across Europe and covering the last 60 thousand years (ka). We used Bayesian inference to reconstruct their phylogenetic relationships and to estimate the age of the specimens that were not directly dated.</span></p> <p><span><strong>Results:</strong> </span><span>We estimated the time to the most recent common ancestor of all last glacial and extant common vole lineages to be 90 ka ago and the divergence of the main mtDNA lineages present in extant populations to between 55 and 40 ka ago, which is earlier than previous estimates. </span><span>We detected several lineage turnovers in Europe during the period of high climate variability at the end of Marine Isotope Stage 3 (MIS 3; 57–29 ka ago) in addition to those found previously around the Pleistocene/Holocene transition.</span><span> </span><span>In contrast, data from the Western Carpathians suggest continuity throughout the Last Glacial Maximum (LGM), even at high latitudes.</span></p> <p><strong><span>Main conclusions: </span></strong><span>The main factor affecting the common vole populations during the last glacial period was the decrease in open habitat during the interstadials, whereas </span><span>climate </span><span>deterioration </span><span>during</span><span> the LGM had little impact on population dynamics. This suggests that the rapid environmental change rather than other factors was the major force shaping the histories of the Late Pleistocene faunas.</span></p>
Virtual Curation Lab table at MAAC 2022
Quick scan of exhibits table done at the Middle Atlantic Archaeological Conference on 26 March 2022 with the Scaniverse app. Source: Objaverse 1.0 / Sketchfab
BOLD release 5 April 2024, curated with pipeline commit aeda07ed466cf326bfc2cd6f7897f9f413480ac0
Open the record for dataset details and reuse information.
DataverseNO curation workflow used by the Research data team at the University Library of Bergen
<p>DataverseNO curation workflow used by the Research data team at the University Library of Bergen.</p> <p>Original Google Slides are available here: https://docs.google.com/presentation/d/1fxpAnvca7jisqjb0oqD1L3N3Kxj4hrpn_ywxSEl6-RU/edit?usp=sharing</p>
BOLD release 5 April 2024, curated with pipeline commit 3c74c5e91daa2ddccce41d38881edc2e242e9c81
Open the record for dataset details and reuse information.
Sociotechnical Dynamics in Open Source Smart Contract Repositories: An Exploratory Data Analysis of Curated High Market Value Projects
<p>This is the replication package for the paper “Sociotechnical Dynamics in Open Source Smart Contract Repositories: An Exploratory Data Analysis of Curated High Market Value Projects”.</p> <p>In project_curation_selection, there is the curation process of the 100 selected projects including the identification of GitHub repositories and classification of evolution scenarios. </p> <p>In distribution_commits_issues_contributors_market_value_before_after_deploy, data collection from GitHub projects includes the distribution of total commits, contributors, and issues before and after deployment of each investigated project. </p> <p>In analysis_commit_messages, there is qualitative analysis of commit message content from all investigated projects. </p> <p>In the analysis_contributors section, the data focuses on analyzing the profiles of each GitHub contributor involved in the investigated projects.</p> <p>In analysis_market_value_by_project, data refers to the market value and volume of each investigated project. </p> <p>In codes, there are scripts used to obtain the analyzed data.</p> <p> </p>
Curated protostome sequences to validate deuterostome specific orthologous groups
<p>This dataset contains protostome peptide sequences that have been used to test the validity of deuterostome specific orthologous groups. BLAST was used to find potential homologous protostome sequences that could invalidate the specificity of the deuterostome orthogroups in question.</p>
World Ocean Database Dataset Curated for Analyzing ENSO-Driven Tropical Pacific Oxygen Variability
<p><strong>Description of folders</strong></p> <p><em>- ENSO index time series</em></p> <p>A file containing time series of the Oceanic Niño Index (ONI) is located in the following folder: ENSOtimeseries. ONI time series data was downloaded from http://origin.cpc.ncep.noaa.gov/products/analysis_monitoring/ensostuff/ONI_v5.php on June 4, 2018.</p> <p><em>- Exclusive economic zones (EEZs)</em></p> <p>A shapefile of the world's EEZs, downloaded from http://marineregions.org/downloads.php, is located in the following folder: EEZs.</p> <p><em>- World Ocean Database</em></p> <p><em>In situ</em> profiles of near-surface (0-700 m) temperature, salinity, and O<sub>2</sub> collected between January 1955 and May 2017 were downloaded from the World Ocean Database (WOD) at<a href="https://www.nodc.noaa.gov/OC5/SELECT/dbsearch/dbsearch.html"> </a><a href="https://www.nodc.noaa.gov/OC5/SELECT/dbsearch/dbsearch.html">https://www.nodc.noaa.gov/OC5/SELECT/dbsearch/dbsearch.html</a> on June 25th, 2017 and binned onto a monthly mean, 5-by-5 degree horizontal grid or grouped by EEZ.</p> <p>The binned monthly mean, 5-by-5 degree horizontal gridded data are in the following folders: WODglobalgridded, WODtropicalpacificonlygridded.</p> <p>The EEZ-grouped profiles are in the following folder: WODrawprofs.</p> <p>----------------------------------------------------------------------------------------------------------------------------------------------------------------</p> <p><strong>Associated publication</strong></p> <p>This dataset was used to generate analyses and figures in the following publication:</p> <p>Leung, S., Thompson, L., McPhaden, M. J., & Mislan, K. A. S. (2019). ENSO drives near-surface oxygen and vertical habitat variability in the tropical Pacific. <em>Environmental Research Letters</em>.</p> <p>----------------------------------------------------------------------------------------------------------------------------------------------------------------</p> <p><strong>Associated code</strong></p> <p>After downloading this dataset, run the associated MATLAB code at the following link to generate the figures and analyses in the above publication:</p> <p>http://doi.org/10.5281/zenodo.2648131</p>
Singing Insects of North America: SINA maps curated
removed taxa w/o scientific name
Towards the conservation of Brazilian legumes: Summary, land cover and use analytics, and a comparative analysis of locations count methods (curated vs. automated [buffer-dissolution])
<p>This dataset presents key findings from the study "<strong>Automating and Enhancing Species Extinction Risk Assessments with Historical Land Use and Land Cover Data</strong>."</p> <p>The file `<em>summary-threatened-legume-species.ods</em>` provides a summary of all threatened species of the Leguminosae family native to Brazil, including IUCN category and criteria, location counts, and the year of the latest assessment, sourced from official records. Additionally, it includes calculated values for Area of Occupancy (AOO) and Extent of Occurrence (EOO), along with trend data indicating natural area change rates (decline [positive number, red] or growth [negative number, green]) within both AOO and EOO. Location counts are further detailed across various buffer radii (1-5 km), utilizing a buffer-dissolution method for automated location counting. Each buffer radius is analyzed to assess AOO and EOO decline, where '1' indicates a species is threatened and '0' indicates it is not. This data provides an efficient method for screening threatened species under criterion B of the IUCN Red List guidelines.</p> <p>The file `<em>overlay-analysis.ods</em>` contains overlay analysis results for AOO and EOO of each species using MapBiomas land use and land cover (LULC) data (specifically MapBiomas Brazil, collection 7.1) from 1985 to 2021, covering all threatened legume species. This file provides both absolute area in square kilometers and percentages for each LULC class. The overlay analysis results support estimates of growth and decline trends for each LULC class.</p> <p>The file `<em>trend-analysis.ods</em>` presents results of annual rate estimates from trend analysis across LULC classes, and including both natural and anthropic groupings. A complete JSON database with p-values and R² values is provided in `<em>trend-analysis.json</em>`.</p> <p>This approach, combining all results for each species in a comprehensive, merged dataset, allows for effective filtering and ranking of the most threatened species as well as identification of their primary threats.</p> <p>We recommend opening the ODS files with LibreOffice, as Microsoft Excel may experience issues parsing decimal formats accurately.</p> <p>More detailed maps and graphs are available at <a title="LULC-MapBiomas-Leguminosae" href="https://github.com/lsbjordao/LULC-MapBiomas-Leguminosae" target="_blank" rel="noopener">https://github.com/lsbjordao/LULC-MapBiomas-Leguminosae</a>.</p>
Figure 1 in A new taxonomist-curated reference library of DNA barcodes for Neotropical electric fish (Teleostei: Gymnotiformes)
Figure 1. Percentages of 934 BOLD systems CO1 sequences (A) and 1015 GenBank CO1 sequences (B) categorized with respect to comparisons with reference library sequences. Minimum interspecific pairwise genetic divergence thresholds of 1.0% and 2.0% distinguish conspecifics (≤1.0/2.0%) and heterospecifics (>1.0/2.0%).
Supplementary material 1 from: Bourret A, Nozères C, Parent E, Parent GJ (2023) Maximizing the reliability and the number of species assignments in metabarcoding studies using a curated regional library and a public repository. Metabarcoding and Metagenomics 7: e98539. https://doi.org/10.3897/mbmg.7.98539
Creation of Gulf of St. Lawrence regional library (GSL-rl) and creation of an eDNA metabarcoding dataset
SCTrans Curated Scenarios
<p>Scenarios curated by SCTrans. The structure of the zip file is:</p> <pre><code class="language-bash">SCTrans-data ├── LGSVL-environment ├── LGSVL-VSE └── OpenSCENARIO</code></pre> <p>And each directory contains:</p> <ul> <li>LGSVL-environment: LGSVL map assets</li> <li>LGSVL-VSE: LGSVL VSE scenarios in JSON format</li> <li>OpenSCENARIO: OpenSCENARIO and OpenDRIVE files, i.e., *.xosc and *.xodr </li> </ul> <p> </p> <p>Hello</p>
Randomized Study of Adjuvant Radiotherapy After Curative Resection of HCC With Narrow Margin (RAISE)
ClinicalTrials.gov study NCT03732105. IPD Sharing: Not stated. Countries: 1. Publications: 3.
Curative Efficacy of Secondary Prevention for Patients With Ischemic Stroke Through Syndrome Differentiation of TCM
ClinicalTrials.gov study NCT02334969. IPD Sharing: NO. Countries: 1. Publications: 9.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.