Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

486

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

486 results for “curation”

Learn how ShareScore rates datasets ↗
dryad36/100

Dataset from: A curated DNA barcode reference library for parasitoids of northern European cyclically outbreaking geometrid moths

<p><span>Large areas of forests are annually damaged or destroyed by outbreaking insect pests. Understanding the factors that trigger and terminate such population eruptions has become crucially important, as plants, plant-feeding insects, and their natural enemies may respond differentially to the ongoing changes in the global climate. In northernmost Europe, climate-driven range expansions of the geometrid moths <em>Epirrita autumnata</em> and <em>Operophtera brumata</em> have resulted in overlapping and increasingly severe outbreaks. Delayed density-dependent responses of parasitoids are a plausible explanation for the ten-year population cycles of these moth species, but the impact of parasitoids on geometrid outbreak dynamics is unclear due to a lack of knowledge on the host ranges and prevalences of parasitoids attacking the moths in nature. To overcome these problems, we reviewed the literature on parasitism in the focal geometrid species in their outbreak range, and then constructed a DNA barcode reference library for all relevant parasitoid species based on reared specimens and sequences obtained from public databases. The combined parasitoid community of<em> E. autumnata</em> and <em>O. brumata</em> consists of 32 hymenopteran species, all of which can be reliably identified based on their barcode sequences. The curated barcode library presented here opens up new opportunities for estimating the abundance and community composition of parasitoids across populations and ecosystems based on mass barcoding and metabarcoding approaches. Such information can be used for elucidating the role of parasitoids in moth population control, possibly also for devising methods for reducing the extent, intensity, and duration of outbreaks.</span></p>

opencc-zeroOct 2022View details →
zenodo36/100

Long-Term Survival and Curative-Intent Treatment in Hepatitis B or C Virus-Associated Hepatocellular Carcinoma Patients Diagnosed during Screening

<p>We uploaded the image of the mansucript <strong>Long-Term Survival and Curative-Intent Treatment in Hepatitis B or C Virus-Associated Hepatocellular Carcinoma Patients Diagnosed during Screening </strong>accepted on Biology.</p>

opencc-by-4.0Oct 2022View details →
zenodo36/100

Bilateral Cultural Agreements, 1935-1972: a Curated International Text Corpus

<p>This is a curated text corpus composed of the texts of bilateral general cultural agreements signed between 1935 and 1972 and deposited with either the League of Nations Treaty Service (LTS) or United Nations Treaty Service (UNTS) and published by those organizations. All texts are in English, either as an original language or in the translation provided by these treaty services.</p> <p>The complete list of bilateral treaties included in this corpus is <a href="https://zenodo.org/record/7361913">available here</a>. The source documents can be found online in PDF form through the <a href="https://treaties.un.org/Pages/Content.aspx?path=DB/UNTS/pageIntro_en.xml">UNTS website</a>. We were able to identify this selection of agreements thanks to the electronic World Treaty Index (<a href="http://worldtreatyindex.com">worldtreatyindex.com</a>), or eWTI [1]. An edited version of the portions of the eWTI related to cultural agreements is available at: <a href="https://zenodo.org/record/5159745">zenodo.org/record/5159745</a>. Discussion of the principles of selection behind this corpus, including a definition of &quot;general cultural agreements,&quot; can be found in this article (<a href="https://doi.org/10.1080/07075332.2022.2048051">B. G. Martin, &quot;The Rise of the Cultural Treaty,&quot; <em>International History Review</em>, 2022</a>); as well as in <a href="https://github.com/benjamingmartin/the_culture_of_international_relations/wiki/eWTI-Notes:-Selection-Decisions-Log-2020,-2021">this log</a>, which documents the decision-making process. The 464 agreements collected in this corpus represent almost exactly half of the total number of general cultural agreements signed between January 1935 and December 1972 (49.88% of 931). The remainder were not deposited with the League of Nations or the UN; some were published in national treaty collections.</p> <p>I assembled this corpus, together with Roger M&auml;hler and Andreas Marklund of <a href="https://www.umu.se/en/humlab/">Humlab, Ume&aring; University</a>, between 2018 and 2020, as part of the research project &quot;The Culture of International Society,&quot; which was funded by a generous grant from Riksbankens Jubileumsfond (the Swedish Foundation for Humanities and Social Sciences, P16-0900:1). Sincere thanks to Roger and Andreas, as well as to Milada Jamroskovic, who conducted many hours of painstaking editing.</p> <p>Read about the project at: <a href="http://www.benjamingmartin.com/cultintsoc">benjamingmartin.com/cultintsoc</a>.</p> <p>Code related to the project, including text analysis tools designed for this corpus, is available at <a href="https://github.com/benjamingmartin/the_culture_of_international_relations">the project&#39;s GitHub page</a>.</p> <p>Please contact me with questions or suggestions: benjamin [dot] martin [at] idehist.uu.se</p> <p>&nbsp;</p> <p>[1] Paul Poast, Michael James Bommarito and Daniel Martin Katz, &lsquo;The Electronic World Treaty Index: Collecting the Population of International Agreements in the 20th Century&rsquo;, SSRN Scholarly Paper (Rochester, NY: Social Science Research Network, February 15, 2010), online at: <a href="http://papers.ssrn.com/abstract=2652760">http://papers.ssrn.com/abstract=2652760</a>.</p>

opencc-by-4.0Nov 2022View details →
zenodo36/100

Curated Dataset of Association Constants Between a Cyclodextrin and a Guest for Machine Learning: Raw Data and Generation Script

<p>Determining the association constant between a cyclodextrin and a guest molecule is an important task for various applications in various industrial and academical fields. However, such a task is time consuming, tedious and requires samples of both molecules. A significant number of association constants and relevant data is available from the literature. The availability of data makes the use of machine learning techniques to predict association constants possible. However, such data is mainly available from tables in articles or appendices. It is necessary to make them available in a computer friendly format and to curate them. Furthermore, the raw data need to be enriched with physicochemical information about each molecule and when such information does not allow to discriminate molecules, some additional data is needed. We present a dataset built from data gathered from the literature. The dataset contains both the original raw data from the articles and the enriched ones. We also provide the scripts used to curate and enrich the raw data.</p>

openbsd-3-clauseJan 2023View details →
dryad36/100

Aligned and curated mtDNA sequences from: Ancient DNA of narrow-headed voles reveals common features of the Late Pleistocene population dynamics in cold-adapted small mammals

<p><span>Narrow-headed vole, together with collared lemming and common vole, was the most abundant small mammal species across Eurasian Late Pleistocene steppe-tundra environments. Previous ancient DNA studies of </span><span>the latter</span><span> </span><span>two</span><span> revealed dynamic past population histories shaped by climatic fluctuations. To investigate the extent to which species with similar adaptations share common evolutionary </span><span>histories,</span><span> we generated a dataset comprising mitochondrial genomes of 139 ancient and 6 modern narrow-headed voles from multiple sites across Europe and north-</span><span>western</span><span> Asia and covering the last ca. 100 thousand years (ka). We inferred Bayesian time-aware phylogenies using 11 </span><span>radiocarbon-dated</span><span> samples for calibration of the molecular clock. We found that across the three </span><span>species,</span><span> divergence of the main mtDNA lineages occurred during Marine Isotope Stages (MIS) 7 and MIS 5, suggesting a common response </span><span>of species adapted to open habitat to the interglacial environments. </span><span>In European narrow-headed voles, we identified multiple </span><span>time-structured</span><span> mtDNA lineages, implying lineage turnovers. Timing of some of these turnovers was synchronous across all three </span><span>species,</span><span> allowing us to identify the main drivers of the Late Pleistocene dynamics of steppe- and cold-adapted species.</span></p>

opencc-zeroFeb 2023View details →
zenodo36/100

Wasmizer: Curating WebAssembly-driven Projects on GitHub

<p>This is the dataset and artifact that accompanies the MSR 2023 paper titled: &quot;Wasmizer: Curating WebAssembly-driven Projects<br> on GitHub&quot;</p>

opencc-by-4.0Mar 2023View details →
zenodo36/100

A large expert-curated cryo-EM image dataset for machine learning protein particle picking

<p>Cryo-electron microscopy (cryo-EM) is a powerful technique for determining the structures of biological macromolecular complexes. Picking single-protein particles from cryo-EM micrographs is a crucial step in reconstructing protein structures. However, the widely used template-based particle picking process is labor-intensive and time-consuming. Though machine learning and artificial intelligence (AI) based particle picking can potentially automate the process, its development is hindered by lack of large, high-quality labelled training data. To address this bottleneck, we present CryoPPP, a large, diverse, expert-curated cryo-EM image dataset for protein particle picking and analysis. It consists of labelled cryo-EM micrographs (images) of 34 representative protein datasets selected from the Electron Microscopy Public Image Archive (EMPIAR). The dataset is 2.6 terabytes and includes 9,893 high-resolution micrographs with labelled protein particle coordinates. The labelling process was rigorously validated through 2D particle class validation and 3D density map validation with the gold standard. The dataset is expected to greatly facilitate the development of both AI and classical methods for automated cryo-EM protein particle picking.</p>

opencc-by-4.0May 2023View details →
zenodo36/100

Triplets of core transmembrane alpha-helical seqments consecutive in sequence from a curated subset of OPM pdb files.

<p>PDB structure files of triplets of transmembrane alpha-helical segments consecutive in sequence from a curated subset of alpha-helical polytopic membrane proteins from OPM. The structures originate from representatives with less than 40% sequence identity and at least 2.7 Angstrom resolution. All together 332 triplet structures are present from 94 membrane proteins. The PDB ID of each file is the first four characters of the file name.</p>

opencc-by-4.0Jun 2023View details →
dryad36/100

Creating, curating, and evaluating a mitogenomic reference database to improve regional species identification using environmental DNA

<p><span>Species detection using eDNA is revolutionizing global capacity to monitor biodiversity. However, the lack of regional, vouchered, genomic sequence information—especially sequence information that includes intraspecific variation—creates a bottleneck for management agencies wanting to harness the complete power of eDNA to monitor taxa and implement eDNA analyses. eDNA studies depend upon regional databases of mitogenomic sequence information to evaluate the effectiveness of such data to detect and identify taxa. We created the Oregon Biodiversity Genome Project to create a database of complete, nearly error-free mitogenomic sequences for all of Oregon's fishes. We have successfully assembled the complete mitogenomes of 313 specimens of freshwater, anadromous, and estuarine fishes representing 24 families, 55 genera, and 129 </span><span>species and lineages. Comparative analyses of these sequences illustrate that many regions of the mitogenome are taxonomically informative, that the short (~150 bp) mitochondrial "barcode" regions typically used for eDNA assays do not consistently diagnose for species, and that complete single or multiple genes of the mitogenome are preferable for identifying Oregon's fishes. This project provides a blueprint for other researchers to follow as they build regional databases, illustrates the taxonomic value and limits of complete mitogenomic sequences, and offers clues as to how current eDNA assays and environmental genomics methods of the future can best leverage this information.</span></p>

opencc-zeroJun 2023View details →
zenodo36/100

Phishing Email Curated Datasets

<p>We have curated 11 datasets spanning from 1995 to 2022.</p> <p>If you use this datasets, please cite:<br>1. A. I. Champa, M. F. Rabbi, and M. F. Zibran, &ldquo;Why phishing emails escape detection: A closer look at the failure points,&rdquo; in&nbsp;<em>12th Interna- tional Symposium on Digital Forensics and Security (ISDFS)</em>, 2024, pp. 1&ndash;6.<br>2. A. I. Champa, M. F. Rabbi, and M. F. Zibran, &ldquo;Curated datasets and feature analysis for phishing email detection with machine learning,&rdquo; in 3rd IEEE International Conference on Computing and Machine Intelligence (ICMI), 2024, pp. 1&ndash;7.</p> <p><br>Bibtext:<br>1. @inproceedings{champa2024phishing,<br>&nbsp; title={Why Phishing Emails Escape Detection: A Closer Look at the Failure Points},<br>&nbsp; author={Champa, Arifa I and Rabbi, Fazle and Zibran, Minhaz F},<br>&nbsp; booktitle={2024 12th International Symposium on Digital Forensics and Security (ISDFS)},<br>&nbsp; pages={1--6},<br>&nbsp; year={2024},<br>&nbsp; organization={IEEE}<br>}</p> <p>2. @inproceedings{champa2024curated,<br>&nbsp; title={Curated Datasets and Feature Analysis for Phishing Email Detection with Machine Learning},<br>&nbsp; author={Champa, Arifa I and Rabbi, Md Fazle and Zibran, Minhaz F},<br>&nbsp; booktitle={3rd IEEE International Conference on Computing and Machine Intelligence (ICMI)},<br>&nbsp; pages = {1--7},<br>&nbsp; year={2024}<br>}</p>

opencc-by-4.0Sep 2023View details →
ClinicalTrials.gov36/100

Adjuvant Tislelizumab Plus Lenvatinib for Patients at High-risk of HCC Recurrence After Curative Resection or Ablation

ClinicalTrials.gov study NCT05910970. IPD Sharing: YES. Countries: 1. Publications: 5.

controlledIPD-YESFeb 2026View details →
ClinicalTrials.gov36/100

The Curative Effect of the Length of the Jejunum Exclusion in Grstric Bypass Surgery for Type 2 Diabetes

ClinicalTrials.gov study NCT02667652. IPD Sharing: NO. Countries: 1. Publications: 5.

closedIPD-NOFeb 2026View details →
ClinicalTrials.gov36/100

Ph II Adjuvant Carboplatin/Docetaxel in Curatively Resected Stage I-IIIA NSCLC

ClinicalTrials.gov study NCT00280735. IPD Sharing: NO. Countries: 1. Publications: 1.

closedIPD-NOFeb 2026View details →
ClinicalTrials.gov36/100

Angiogenic and EGFR Blockade With Curative Chemoradiation for Advanced Head and Neck Cancer

ClinicalTrials.gov study NCT00140556. IPD Sharing: Not stated. Countries: 1. Publications: 8.

restrictedIPD-UNDECIDEDFeb 2026View details →
ClinicalTrials.gov36/100

Preventive Versus Curative Treatment of Fluid Overload

ClinicalTrials.gov study NCT04050007. IPD Sharing: YES. Countries: 1. Publications: 1.

controlledIPD-YESFeb 2026View details →
ClinicalTrials.gov36/100

Curative Effect Study of Endostatin Combined With Chemoradiotherapy to Non-small-cell Lung Cancer

ClinicalTrials.gov study NCT01218594. IPD Sharing: Not stated. Countries: 1. Publications: 1.

restrictedIPD-UNDECIDEDFeb 2026View details →
ClinicalTrials.gov36/100

Phase II Study of Pembrolizumab After Curative Intent Treatment for Oligometastatic Non-Small Cell Lung Cancer

ClinicalTrials.gov study NCT02316002. IPD Sharing: Not stated. Countries: 1. Publications: 1.

restrictedIPD-UNDECIDEDFeb 2026View details →
ClinicalTrials.gov36/100

Initial Attack on Latent Metastasis Using TAS-102 for ct DNA Identified Colorectal Cancer Patients After Curative Resection

ClinicalTrials.gov study NCT04457297. IPD Sharing: Not stated. Countries: 2. Publications: 0.

restrictedIPD-UNDECIDEDFeb 2026View details →
dryad36/100

A curated database of milky sea observations from 1600 to present

Open the record for dataset details and reuse information.

publicMar 2025View details →
dryad36/100

Creating, curating, and evaluating a mitogenomic reference database to improve regional species identification using environmental DNA

Open the record for dataset details and reuse information.

publicJun 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record