Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
3,996
datasets available to search
ShareScore release 0.7.1
Dataset results
3,996 results for “registry”
Ethnic and Migrant Minorities (EMM) Survey Registry: All metadata records
<p>The <a href="https://ethmigsurveydatahub.eu/emmregistry/">Ethnic and Migrant Minorities (EMM) Survey Registry</a> is a free online tool that allows users to search for and learn about existing quantitative surveys undertaken with EMM (sub)populations conducted in 34 European countries, from 2000 onwards, through compiled survey-level metadata.</p> <p>The first version was produced by a team led by CEE (Sciences Po, CNRS) and jointly funded through the COST Action 16111 – ETHMIGSURVEYDATA (a network of more than 200 European researchers active in the ethnic and migration studies field), the Horizon 2020 infrastructure project SSHOC (within Task 9.2 on Ethnic and Migration Studies, within Work Package 9 on Data Communities) and the project FAIRETHMIGQUANT (an Open Science project funded by the French Agence Nationale de la Recherche, ANR).</p> <p>This specific record includes the metadata for 2,120 survey records as .dta, .sav and .csv files published on the Registry, as of 31.07.2025.</p>
Schedatura dei notai dell'Italia meridionale e insulare dei secc. XIII-XV di cui si conservano i rispettivi registri
<p>L’obiettivo della schedatura dei notai nell'ambito del progetto NotMed (EL NOTARIAT PÚBLIC EN LA MEDITERRÀNIA OCCIDENTAL: ESCRIPTURA, INSTITUCIONS, SOCIETAT I ECONOMIA (SEGLES XIII-XV) - Ministerio de Ciencia e Innovación. PID2019-105072GB-I00 - <a href="https://www.ub.edu/notmed/">https://www.ub.edu/notmed/</a>) era quello di conoscere il numero di volumi in legatura (protocolli notarili, bastardelli, etc.) esistenti nell’Italia meridionale e insulare per i secoli medievali e di creare una base per ulteriori ricerche.</p> <p>Hanno contribuito:</p> <p>Giuliano Capriolo, Andrea Casalboni, Gemma Teresa Colesanti, Martina Del Popolo, Corinna Drago, Alessandro Gaudiero, Antonio Macchione, Eleni Sakellariou, Daniela Santoro, Vera Isabell Schwarz-Ricci, Chiara Sciarroni, Alessandro Soddu, Maria Elisabetta Vendemia, Elisa Turrisi e Maurizio Vesco.</p> <p>NB.</p> <ul> <li> Nella dicitura “volumi in legatura” rientrano sia veri e proprio protocolli notarili sia bastardelli sia fascicoli rilegati.</li> <li> Il limite cronologico è l’anno 1500, tuttavia nei casi di notai che iniziano a rogare nella seconda metà del ‘400 sono confluiti nel censimento anche i registri dei primi decenni del ‘500.</li> <li> Per ogni notaio è stata compilata una singola scheda, tranne in due casi nei quali i protocolli si conservano in due istituzioni diverse.</li> <li> I volumi miscellanei sono stati conteggiati e schedati con una nota specifica inserita nel campo commento.</li> <li> È da tener presente che la base di rilevamento è eterogenea: alcune indicazioni si basano sull’esame autoptico del materiale, altre sulle indicazioni dell’inventario on line dell’archivio o su lavori pubblicati in precedenza. Per questo motivo si consiglia di consultare sempre le osservazioni del compilatore nel campo commento e le indicazioni sulla fonte dell’informazione.</li> </ul>
The Semantic Turkey metadata registry ontology
<p>An application profile of DCAT combining it with other metadata vocabularies (e.g. VoID, DCTERMS, LIME) to meet requirements elicited in various use cases of the Semantic Web platform Semantic Turkey</p>
Event Registry dataset with multiple extracted features (both sparse and dense)
<p>This is a republication of the Event Registry dataset originaly published by:</p> <p>Rupnik, Jan, Andrej Muhic, Gregor Leban, Primoz Skraba, Blaz Fortuna, et Marko Grobelnik. 2016. « News Across Languages - Cross-Lingual Document Similarity and Event Tracking ». <em>Journal of Artificial Intelligence Research</em> 55 (janvier): 283‑316. <a href="https://doi.org/10.1613/jair.4780">https://doi.org/10.1613/jair.4780</a>.</p> <p>And reorganised for document tracking by:</p> <p>Miranda, Sebastião, Artūrs Znotiņš, Shay B. Cohen, et Guntis Barzdins. 2018. « Multilingual Clustering of Streaming News ». In <em>2018 Conference on Empirical Methods in Natural Language Processing</em>, 4535‑44. Brussels, Belgium: Association for Computational Linguistics. <a href="https://www.aclweb.org/anthology/D18-1483/">https://www.aclweb.org/anthology/D18-1483/</a>.</p> <p>In this dataset, we provide multiple features extracted from the text itself. <strong>Please note the text is missing from the dataset published in the CSV format for copyright reasons. You can download the original datasets and manually add the missing texts from the original publications.</strong></p> <p>Features are extracted using:</p> <p>- A corpus of reference articles in multiple languages languages for TF-IDF weighting. (<em>features_news</em>) [1]</p> <p>- A corpus of tweets reporting news for TF-IDF weighting. (<em>features_tweets)</em> [1]</p> <p>- A S-BERT model [2] that uses <em>distiluse-base-multilingual-cased-v1 </em>(called <em>features_use</em>) [3]</p> <p>- A S-BERT model [2] that uses <em>paraphrase-multilingual-mpnet-base-v2 </em>(called <em>features_mpnet</em>) [4]</p> <p><strong>References:</strong></p> <p>[1]: Guillaume Bernard. (2022). Resources to compute TF-IDF weightings on press articles and tweets (1.0) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.6610406</p> <p>[2]: Reimers, Nils, et Iryna Gurevych. 2019. « Sentence-BERT: Sentence Embeddings Using Siamese BERT-Networks ». In <em>Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)</em>, 3982‑92. Hong Kong, China: Association for Computational Linguistics. <a href="https://doi.org/10.18653/v1/D19-1410">https://doi.org/10.18653/v1/D19-1410</a>.</p> <p>[3]: https://huggingface.co/sentence-transformers/distiluse-base-multilingual-cased-v1</p> <p>[4]: https://huggingface.co/sentence-transformers/paraphrase-multilingual-mpnet-base-v2</p>
Event Registry events associated to Wikidata entities
<p>This is a set of labels that associate the cluster and events ids of the Event Registry dataset with entities in the Wikidata KG. This can be used in order to analyse the event description completeness in news articles, evaluate a query based document retrieval on Event Registry, etc.</p>
S31 | WRTMSD | Wiley Registry of Tandem Mass Spectral Data, MSforID
<p>This is the collection associated with list S31 WRTMSD on the NORMAN Suspect List Exchange.</p> <p><a href="https://www.norman-network.com/?q=suspect-list-exchange">https://www.norman-network.com/?q=suspect-list-exchange</a></p> <p>S31</p> <p>WRTMSD</p> <p><strong>Wiley Registry of Tandem Mass Spectral Data, MSforID</strong></p> <p>WRTMSD <a href="https://www.norman-network.com/sites/default/files/files/suspectListExchange/210119Update/WRTMSD_wDTXSIDs_24012019.csv">CSV</a>, <a href="https://www.norman-network.com/sites/default/files/files/suspectListExchange/210119Update/WRTMSD_wDTXSIDs_24012019.xlsx">XLSX</a> (24/01/2019)</p> <p>CompTox <a href="https://comptox.epa.gov/dashboard/chemical_lists/wrtmsd">WRTMSD List</a></p> <p>WRTMSD <a href="https://www.norman-network.com/sites/default/files/files/suspectListExchange/210119Update/WRTMSD_InChIKeys.txt">InChIKeys </a>(24/01/2019)</p> <p>The "Wiley Registry of Tandem Mass Spectral Data, MSforID" contains high-quality tandem MSacquired on a QqTOF instrument, developed by Herbert Oberacher (Medical University of Innsbruck, Austria). More information at <a href="http://www.msforid.com/">www.msforid.com</a> </p>
MiRoR11 - P3 - Primary outcomes extracted from registries
<p>Two sets (original and deduplicated) of primary outcomes from trial registries</p>
PARESv3 : PArish REgistry Survey − Historical Census Table Dataset (19th, 20th centuries) − France
<h2>PARES Dataset v3</h2> <p>PARES (PArish REcord Survey) contains<strong> 535 images of handwritten census tables</strong> for years ranging from around <strong>1650 A.D. until 1850 A.D.</strong>.They come from two <strong>French cities</strong>, Vic-sur-Seille (French department of Moselle) and Echevronne (French department of Côte d'Or). While they mention very ancient times, the documents are handwritten transcriptions of even older documents and are quite recent, copied from original documents during the 1950's and 1960's for demographic studies led by the INED in France (<em>Institut National des études démographiques</em> − National Institute for Demographic Studies). These copies were made by only a few different writers.</p> <p>In this updated version of the dataset, each table row has been carefully annotated and transcribed. Please note that for each row transcription, we have specified the attribute to which each value corresponds.</p> <p>We published a paper, <a href="https://link.springer.com/article/10.1007/s10032-025-00531-z">The PARES Database: Information Extraction over Historical Parish Records,</a> in which we better describe the dataset and the tasks it's possible to run on it.</p> <p> </p>
Event Registry dataset with multiple extracted features (both sparse and dense) and degraded by OCR
<p>This is the same dataset as:</p> <p>Guillaume Bernard. (2022). Event Registry dataset with multiple extracted features (both sparse and dense) (1.0) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.6630367</p> <p>But with texts degraded by OCR as described in:</p> <p>Guillaume Bernard. (2022). Event Registry dataset texts with OCR degradations and synthesised segmentation (1.0) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.6631305</p>
Event Registry titles dataset with multiple extracted features (both sparse and dense) and degraded by OCR
<p>This is the same dataset as:</p> <p>Guillaume Bernard. (2022). Event Registry titles only dataset with multiple extracted features (both sparse and dense) (1.0) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.6630447</p> <p>But with texts degraded by OCR as described in:</p> <p>Guillaume Bernard. (2022). Event Registry titles dataset texts with OCR degradations (1.0) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.6630828</p>
Hypertension and Diabetes Registry Dataset in Addis Ababa, Ethiopia
<p>Dataset collected from the “hypertension and diabetes” national registry of the ministry of health of Ethiopia.<br> </p>
Longitudinal observational study of pediatric patients with primary brain tumors: establishment of a hospital-based registry.
<p>Although tumors of the central nervous system (CNS) represent 2 % of all malignancies in general, they cause a disproportionately large morbidity and mortality and are the second most common form of cancer in children and the major solid tumor in childhood in the U.S., occurring in 21.3% of all children with malignant disease. The treatment of brain tumors in children and adolescents has evolved significantly in recent decades. Nowadays, most children with a diagnosis of brain tumor are treated properly and achieve prolonged survival. In order to obtain an overview of the impact of brain tumors, specialized registries, which provide information on all types of brain tumors, have emerged in several countries. Following on the pioneering Japanese and American experiences of specialized national records of brain tumors, other specialized registries were opened in European countries. This project aims to initiate a registry of the epidemiological profile of patients treated for CNS tumors in the Pediatric Cancer Center (CPC) of our hospital from January 2000 to December 2013, at diagnosis and during follow-up, updating information periodically. This data will be recorded in an electronic database capable of storing, retrieving and presenting information of interest. Prospectively recorded epidemiological data of patients diagnosed from January 2014 will be accrued, maintaining the database active to continuously record information on patients with CNS tumors treated in the CPC HIAS. Thus, creating a hospital registry of pediatric patients with CNS tumors. To this end, an instrument of data collection will be created using Google Apps (Google Inc., 2014), a digital platform with capacity for storage, creation and editing of documents and collaboration in real time over the cloud.</p>
FREYA Webinar on the PID Services Registry
<p>Recording of the FREYA webinar on the PID Services Registry where you can find services related to Persistent Identifiers (PIDs). Visit the registry at <a href="https://www.youtube.com/redirect?redir_token=QUFFLUhqbWdDUXcwTWZZREZWTFpGYlZfZjBTQmNRcnlqd3xBQ3Jtc0tsZTNMaUJvaEFFc2NWN3hTbE81bEFsZzh5eUxONWJPdVA1OWdFYWdXSGhlY090bzBibFNqWUF6dmJ3VTI1NmVwY2w2cHlIU1JZX29SSTdVOE5Bd1BUQUplTWYwRS1OLUVlUXNrRXBwMHlYc2pZblVWYw%3D%3D&event=video_description&v=MTM8d0YbfnY&q=https%3A%2F%2Fpidservices.org">https://pidservices.org</a>.</p>
Container Registry Benchmark experiments measurements and trace workload samples
<p>Measurements for experiments using Container Registry Benchmark, CReB. 4 experiments: Long running, small experiment stress mode, small experiment delay mode, and large workload experiment.</p> <p> </p> <p>Structure:</p> <ol> <li><strong>full-measurements-long-running-pull.csv : </strong>measurements for long running pull experiment</li> <li><strong>full-measurements-long-running-push.csv: </strong>measurements for long running push experiment</li> <li><strong>result-bug-analysis.zip: </strong>results from bug analysis of trace replayer</li> <li><strong>results-1hr-experiment.zip: </strong>measurements for the large experiment (4 registries)</li> <li><strong>results-small-delay.zip: </strong>measurements for the delay mode, small experiment with real workload</li> <li><strong>results-small-stress.zip: </strong>measurements for the stress mode, small experiment with real workload</li> <li><strong>traces.zip: </strong>traces used for pen-and-paper experiment, 1 hour sample, and the trace used for small experiment (selected are first 405 requests)</li> </ol>
Nordic trial reporting project: Raw data from EU Clinical Trials Registry (EUCTR) and ClinicalTrials.gov
<p>Uploaded on behalf of the author team for the research project "<strong>Systematic evaluation of clinical trial reporting at medical universities and university hospitals in the Nordic countries</strong>".</p><p><strong>Raw data</strong> from EU Clinical Trials Registry (EUCTR) and ClinicalTrials.gov:</p><p><strong>EUCTR</strong>: We retrieved the latest dataset for the EU Trials Tracker of EUCTR trials on Nov 27, 2022, reflecting data from Nov 7, 2022 (1,2). We also used a custom web scraper that automatically extracts data from EUCTR country protocols and results sections (variables described in Appendix Table 2), developed by the EU Trials Tracker team (2).<br>References: <br>1. Goldacre B, DeVito NJ, Heneghan C, Irving F, Bacon S, Fleminger J, Curtis H. Compliance with requirement to report results on the EU Clinical Trials Register: cohort study and web resource. BMJ. 2018 Sep 12;362:k3218.<br>2. EU Trials Tracker — Who's not sharing clinical trial results? [Internet]. [cited 2022 Aug 30]. Available from: http://eu.trialstracker.net/</p><p><strong>ClinicalTrials.gov</strong>: We downloaded the complete Aggregate Analysis of ClinicalTrials.gov dataset (AACT, http://aact.ctti-clinicaltrials.org/) on Nov 27, 2022, reflecting data from Nov 9, 2022. </p><p>See our GitHub and preregistered protocol for more details:<br>https://github.com/cathrineaxfors/nordic-trial-reporting<br>https://osf.io/97qkv/</p>
Presentation of The Australian Companion Animal Registry of Cancers (ACARCinom)
<p>With support from the Australian Research Data Commons (ARDC) through the Australian Data Partnership program, ACARCinom is the first Australia-wide registry of animal cancer occurrences that addresses the gaps in veterinary cancer data registries. ACARCinom aims to make a positive impact on cancer research for our pets. Having reliable data is crucial for understanding the patterns of cancer and for evaluating treatments in both animals and humans.</p><p>Five university veterinary schools and Australia's 2 leading veterinary pathology providers are partnering in the ACARCinom project: The University of Queensland, Queensland University of Technology, University of Sydney, Gribbles Veterinary Pathology, IDEXX, University of Adelaide, Murdoch University</p><p>By uniting the expertise and resources of these institutions, ACARCinom is poised to make significant advancements in understanding and combating cancer in dogs and cats. This project represents a remarkable collaboration that harnesses the power of data to unlock new insights and drive progress in the field of veterinary oncology.</p><p>This video explains how the ACARCinom Dashboard works and what its functionalities are. You can have access to the ACARCinom database at the following link: <a href="https://www.acarcinom.org.au/">acarcinom.org.au</a></p>
MaStR power unit registry for eGon-data (legacy)
<p>The data set contains selected data from the <a href="https://www.marktstammdatenregister.de/MaStR">Marktstammdatenregister</a> for eGon-data pipeline.<br><strong>Note:</strong> This dataset is legacy and should be only used for compatibility reasons in eGon-data.</p> <p><strong>Dump version: 2021-04-30</strong></p> <p>The data set has been downloaded and processed with a very early version of the software <a href="https://github.com/OpenEnergyPlatform/open-MaStR">open-MaStR</a>.</p> <p>License information:<br>Marktstammdatenregister - © Bundesnetzagentur für Elektrizität, Gas, Telekommunikation, Post und Eisenbahnen | <a href="https://www.govdata.de/dl-de/by-2-0">DL-DE-BY-2.0</a></p>
MaStR power unit registry for eGon-data
<p>The data set contains selected data from the <a href="https://www.marktstammdatenregister.de/MaStR">Marktstammdatenregister</a> for eGon-data pipeline.</p> <p><strong>Dump version: 2022-11-17</strong></p> <p>The data set has been downloaded and processed with the software <a href="https://github.com/OpenEnergyPlatform/open-MaStR">open-MaStR</a>.<br>Compared to the <a href="../doi/10.5281/zenodo.6807425">official open-MaStR releases</a> a custom format postprocessing has been applied.</p> <p>License information:<br>Marktstammdatenregister - © Bundesnetzagentur für Elektrizität, Gas, Telekommunikation, Post und Eisenbahnen | <a href="https://www.govdata.de/dl-de/by-2-0">DL-DE-BY-2.0</a></p>
Diabetes-related excess mortality in Mexico: a comparative analysis of national death registries between 2017-2019 and 2020
<p>Dataset to replicate the article "Diabetes-related excess mortality in Mexico: a comparative analysis of national death registries between 2017-2019 and 2020" available in MedRxiv at: https://www.medrxiv.org/content/10.1101/2022.02.24.22271337v1</p> <p>Code available at: https://github.com/oyaxbell/diabetes_excess</p>
Research Organization Registry (ROR) Data in Sqlite Database
<p>Version 1.1 of the Research Organization Registry data translated into a SQLite database using ROR2DB (https://github.com/Metadata-Game-Changers/ROR2DB).</p> <p>This version adds country, country_code, and status to the ror table.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.