Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
15,702
datasets available to search
ShareScore release 0.7.1
Dataset results
15,702 results for “history”
Herbarium specimen image of Linnaea borealis L., part of the collection of Natural History Museum, University of Tartu
<p>Part of a training dataset of scanned herbarium specimens. The data paper and a summary landing page will be published on Zenodo as it gets published.<br> <br> Content of this deposition:<br> <br> - A JSON-LD datafile listing the label data associated with this herbarium specimen. The Darwin and Dublin Core data standards are used for most values.<br> - Two JPEG image files of the scanned herbarium sheet, one of the specimen itself and one of its label (indicated by_lab).<br> - Two lossless TIFF images from which the JPEG images have been derived.</p>
Herbarium specimen image of Lobelia dortmanna L., part of the collection of Natural History Museum, University of Tartu
<p>Part of a training dataset of scanned herbarium specimens. The data paper and a summary landing page will be published on Zenodo as it gets published.<br> <br> Content of this deposition:<br> <br> - A JSON-LD datafile listing the label data associated with this herbarium specimen. The Darwin and Dublin Core data standards are used for most values.<br> - Two JPEG image files of the scanned herbarium sheet, one of the specimen itself and one of its label (indicated by_lab).<br> - Two lossless TIFF images from which the JPEG images have been derived.</p>
Herbarium specimen image of Allium angulosum L., part of the collection of Natural History Museum, University of Tartu
<p>Part of a training dataset of scanned herbarium specimens. The data paper and a summary landing page will be published on Zenodo as it gets published.<br> <br> Content of this deposition:<br> <br> - A JSON-LD datafile listing the label data associated with this herbarium specimen. The Darwin and Dublin Core data standards are used for most values.<br> - Two JPEG image files of the scanned herbarium sheet, one of the specimen itself and one of its label (indicated by_lab).<br> - Two lossless TIFF images from which the JPEG images have been derived.</p>
Herbarium specimen image of Asplenium ruta-muraria L., part of the collection of Natural History Museum, University of Tartu
<p>Part of a training dataset of scanned herbarium specimens. The data paper and a summary landing page will be published on Zenodo as it gets published.<br> <br> Content of this deposition:<br> <br> - A JSON-LD datafile listing the label data associated with this herbarium specimen. The Darwin and Dublin Core data standards are used for most values.<br> - Two JPEG image files of the scanned herbarium sheet, one of the specimen itself and one of its label (indicated by_lab).<br> - Two lossless TIFF images from which the JPEG images have been derived.</p>
CLDF dataset derived from Lee's "Sketch of Language History in the Korean Peninsula" from 2015
<p>Cite the source of the dataset as:</p> <blockquote> <p>Lee, Sean (2015). A Sketch of Language History in the Korean Peninsula. PLoS ONE 10(5): e0128448. doi:10.1371/journal.pone.0128448</p> </blockquote>
D-PLACE dataset derived from Bertolo et al. 2023 'Cross-cultural music corpus: The Expanded Natural History of Song Discography'
<p>Cite the source of the dataset as:</p> <blockquote> <p>Mila Bertolo, Martynas Snarskis, Manvir Singh, & Samuel Mehr. (2023, August 8). Cross-cultural music corpus: The Expanded Natural History of Song Discography. Zenodo. https://doi.org/10.5281/zenodo.8378337</p> </blockquote>
Mesquita et al, 2015: Life History Data
Daniel O. Mesquita, Guarino R. Colli, Gabriel C. Costa, Taís B. Costa, Donald B. Shephard, Laurie J. Vitt, and Eric R. Pianka. 2015. Life history data of lizards of the world. Ecology 96:594. <p></p>http://dx.doi.org/10.1890/14-1453.1<p></p>
Mesquita et al, 2015: Life history data of lizards of the world (1014) DwCA
Daniel O. Mesquita, Guarino R. Colli, Gabriel C. Costa, Taís B. Costa, Donald B. Shephard, Laurie J. Vitt, and Eric R. Pianka. 2015. Life history data of lizards of the world. Ecology 96:594. <p></p>http://dx.doi.org/10.1890/14-1453.1<p></p>Daniel O. Mesquita, Guarino R. Colli, Gabriel C. Costa, Taís B. Costa, Donald B. Shephard, Laurie J. Vitt, and Eric R. Pianka. 2015. Life history data of lizards of the world. Ecology 96:594. <p></p>http://dx.doi.org/10.1890/14-1453.1
Morgan Ernest, 2003: Life history characteristics of placental non-volant mammals
S. K. Morgan Ernest. 2003. Life history characteristics of placental non-volant mammals. Ecology 84:3402.<p></p>S. K. Morgan Ernest. 2003. Life history characteristics of placental non-volant mammals. Ecology 84:3402.
Natural history specimens collected and/or identified and deposited.
Natural history specimen data collected and/or identified by John Howieson, <a href="https://orcid.org/0000-0003-1378-293X">https://orcid.org/0000-0003-1378-293X</a>. Claims or attributions were made on Bionomia, <a href="http://bionomia.net">https://bionomia.net</a> using specimen data from the Global Biodiversity Information Facility, <a href="https://gbif.org">https://gbif.org</a>.
Genome data and resources on the recombination landscape and population history of the harlequin fly
<p>This dataset contains phased vcf files of <em>Chironomus riparius, </em>ouput files of RepeatMasker, MELT, RepeatOBserver, MSMC2, iSMC and bedtools, such as supporting files. </p> <p>For further details also check the GitHub page: <a href="https://github.com/lpettrich/Crip_Recombination_PopHistory_Cla_2024" target="_blank" rel="noopener">https://github.com/lpettrich/Crip_Recombination_PopHistory_Cla_2024</a></p> <ul> <li><strong>phased-vcfs: </strong>Artificially phased vcf-files of five populations with four individuals each. Needed to generate multihetsep files. Input files for iSMC.<br> <ul> <li>Hesse in Germany = MG</li> <li>Rhône-Alpes in France = MF</li> <li>Lorraine in France = NMF</li> <li>Piemont in Italy = SI</li> <li>Andalusia in Spain = SS</li> </ul> </li> <li><strong>multihetsep-files: </strong>Created with msmc-tools. Input files for MSMC2. </li> <li><strong>RepeatMasker: </strong>Raw output of RepeatMasker run. Summary file and file with filtered <em>Cla</em>-element (a transposable element) included.<strong><br></strong></li> <li><strong>MELT: </strong>MELT ouput with added info on population and numbered insertions reflecting all 441 detected <em>Cla </em>insertions.<strong><br></strong></li> <li><strong>RepeatOBserver: </strong>Summary files on centromere predictions based on histograms and Shannon Diversity from RepeatOBserver. Genome-wise Shannon Diversity per chromosome included. <strong><br></strong></li> <li><strong>MSMC2: </strong>Raw ouput of combined cross-coalescence and mean values if MSMC2 per populations. <strong><br></strong></li> <li><strong>iSMC: </strong>Recombination rate rho in 10 kb windows and 100 kb windows along the genome. <strong><br></strong></li> <li><strong>bedtools closest ismc 10 kb: </strong>Bedtools closest analysis of the distance of the next <em>Cla</em>-element to the recombination rate rho in 10 kb windows.<strong><br></strong></li> <li><strong>bedtools closest ismc 100 kb: </strong>Bedtools closest analysis of the distance of the next <em>Cla</em>-element to the recombination rate rho in 100 kb windows.</li> <li><strong>input-files figures: </strong>Supporting files needed to create figures.<strong><br></strong></li> </ul>
The Acoustics of Ely Cathedral's Lady Chapel: a study of its changes throughout history
<p>This repository contains the data set related to the conference paper “The Acoustics of Ely Cathedral’s Lady Chapel: a study of its changes throughout history”, published in the "I3DA 2021 International Conference" and available at: DOI: tbc</p> <p>This dataset contains the B-format Room Impulse Responses (RIR) in the Waveform Audio File standard Format (.wav) measured and simulated at a number of source-receiver combinations in the Ely Cathedral's Lady Chapel, used for the acoustical analysis performed as part of the CATHEDRAL ACOUSTICS project (CA-MRIR-EC-LC and CA-SRIR-EC-LC folders respectively).</p> <p>LICENCE.txt, METADATA.txt and README.txt contain a brief description of the folder contents, authors, and other useful information.</p> <p>Details on the acoustic measurement campaigns and acoustic simulations can be found in the conference paper.</p> <p>Please cite both the conference paper and the dataset if used.</p> <p>--------------------------------------------------</p> <p>This work is license under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License (CC BY-NC-SA 4.0). (see https://creativecommons.org/licenses/by-nc-sa/4.0/)</p> <p>--------------------------------------------------</p> <p>Dataset curated by Lidia Álvarez-Morales, Theatre, Film, Television and Interactive Media Department, University of York.<br> Contact: lidiaalvarezmorales@gmail.com, lidia.alvarezmorales@york.ac.uk.</p> <p>-------------------------------------------------</p> <p>Funding was provided by the European Union’s Horizon 2020 research and innovation programme (http://dx.doi.org/10.13039/501100007601) under the Marie Sklodowska-Curie grant agreement No 797586</p>
A dataset of GitHub Actions workflow histories
<p>This replication package accompagnies the dataset and exploratory empirical analysis reported in the paper "A dataset of GitHub Actions workflow histories" published in the IEEE MSR 2024 conference. (The Jupyter notebook can be found in previous version of this dataset).</p> <p><em><strong>Important notice :</strong> It looks like Zenodo is compressing gzipped files two times without notice, they are "double compressed". So, when you download them they should be named : <code>x.gz.gz</code> instead of <code>x.gz</code>. Notice that the provided MD5 refers to the original file. </em></p> <p><em><strong>2025-10-09 update: update repositories list and observation period</strong>. We now have 3M+ workflows from 49.2K+ repositories. We consider repositories with at least one commit after August 25th, 2024, and they were pulled on August 25th-26th, 2025.</em></p> <p>2025-04-15 update: fix missing metadata and minor notation bug. (unchanged observation period)</p> <p><strong>2024-10-25 update: update repositories list and observation period</strong>. <em>We now have 2.3M+ workflows from 43.3K+ repositories. We consider repositories with at least one commit after January 1st, 2024, and they were pulled on October 7th, 2024.</em></p> <p>2024-07-09 update: fix sometimes invalid <code>valid_yaml</code> flag.</p> <p>2024-04-30: initial version</p> <p>The dataset was created as follow : </p> <ol> <li>First, we used GitHub SEART (on August 25th, 2025) to get a list of every non-fork repositories created at least one year before. having at least 300 commits and at least 100 stars where at least one commit was made in the last year. (The goal of these filter is to exclude experimental and personnal repositories).</li> <li>We checked if a folder <code>.github/workflows</code> existed. We filtered out those that did not contained this folder and pulled the others (on August 25th-26th, 2025).</li> <li>We applied the tool <code>gigawork</code> (version 1.4.2) to extract every files from this folder. The exact command used is <code>python batch.py -d /ourDataFolder/repositories -e /ourDataFolder/errors -o /ourDataFolder/output -r /ourDataFolder/repositories_everything.csv.gz -- -w /ourDataFolder/workflows_auxiliaries</code>. (The script <code>batch.py</code> can be found <a href="https://github.com/cardoeng/gigawork/blob/master/scripts/batch.py" target="_blank" rel="noopener">on GitHub</a>).</li> <li>We concatenated every files in <code>/ourDataFolder/output</code> into a csv (using <code>cat headers.csv output/*.csv > workflows_auxiliaries.csv</code> in <code>/ourDataFolder</code>) and compressed it.</li> <li>We added the column <code>uid</code> via a script available <a href="https://github.com/cardoeng/gigawork/blob/master/scripts/uid.py">on GitHub.</a></li> <li>Finally, we archived the folder with pigz <code>/ourDataFolder/workflows</code> (<code>tar -c --use-compress-program=pigz -f workflows_auxiliaries.tar.gz /ourDataFolder/workflows</code>)</li> </ol> <p>Using the extracted data, the following files were created :</p> <ol> <li><code>workflows.tar.gz</code> contains the dataset of GitHub Actions workflow file histories.</li> <li><code>workflows_auxiliaries.tar.gz</code> is a similar file containing also auxiliary files.</li> <li><code>workflows.csv.gz</code> contains the metadata for the extracted workflow files.</li> <li><code>workflows_auxiliaries.csv.gz</code> is a similar file containing also metadata for auxiliary files.</li> <li><code>repositories.csv.gz</code> contains metadata about the GitHub repositories containing the workflow files. These metadata were extracted using the SEART Search tool. </li> </ol> <p>The metadata is separated in different columns:</p> <ol> <li><code>repository</code>: The repository (author and repository name) from which the workflow was extracted. The separator "/" allows to distinguish between the author and the repository name</li> <li><code>commit_hash</code>: The commit hash returned by git</li> <li><code>author_name</code>: The name of the author that changed this file</li> <li><code>author_email</code>: The email of the author that changed this file</li> <li><code>committer_name</code>: The name of the committer</li> <li><code>committer_email</code>: The email of the committer</li> <li><code>committed_date</code>: The committed date of the commit</li> <li><code>authored_date</code>: The authored date of the commit</li> <li><code>file_path</code>: The path to this file in the repository</li> <li><code>previous_file_path</code>: The path to this file before it has been touched</li> <li><code>file_hash</code>: The name of the related workflow file in the dataset</li> <li><code>previous_file_hash</code>: The name of the related workflow file in the dataset, before it has been touched</li> <li><code>git_change_type</code>: A single letter (A,D, M or R) representing the type of change made to the workflow (Added, Deleted, Modified or Renamed). This letter is given by <code>gitpython</code> and provided as is. </li> <li><code>valid_yaml</code>: A boolean indicating if the file is a valid YAML file.</li> <li><code>probably_workflow</code>: A boolean representing if the file contains the YAML key <code>on</code> and <code>jobs</code>. (Note that it can still be an invalid YAML file).</li> <li><code>valid_workflow</code>: A boolean indicating if the file respect the syntax of GitHub Actions workflow. A freely available JSON Schema (used by gigawork) was used in this goal.</li> <li><code>uid</code>: Unique identifier for a given file surviving modifications and renames. It is generated on the addition of the file and stays the same until the file is deleted. Renamings does not change the identifier.</li> </ol> <p>Both <code>workflows.csv.gz</code> and <code>workflows_auxiliaries.csv.gz</code> are following this format.</p>
The Natural History Museum's collection of Dalbergia, Pterocarpus and the Phaseolinae subtribe
<p>In 2018 the Natural History Museum (NHMUK, herbarium code: BM) undertook a pilot digitisation project together with the Royal Botanic Gardens Kew (project Lead) and the Royal Botanic Garden Edinburgh to collectively digitise non-type herbarium material of the subtribe <em>Phaseolinae</em> and the genera <em>Dalbergia </em>L.f. and <em>Pterocarpus</em> Jacq. (rosewoods and padauk), all from the economically important family of legumes (<em>Leguminosae</em> or <em>Fabaceae</em>). </p> <p>These taxonomic groups were chosen for two case studies using the herbarium collections to support the aims of the UK’s Department for Environment Food & Rural Affairs (DEFRA)-allocated, Official Development Assistance (ODA) funding: study 1 - to support the development of dry beans as a sustainable and resilient crop; study 2 - to aid conservation and sustainable use of rosewoods and padauk.</p> <p>We present the images and metadata for 11,222 NHMUK specimens. This includes label transcription and georeferencing, along with summary data on geographic, taxonomic, collector and temporal coverage. We also provide timings and the methodology for our transcription and georeferencing protocols. Approximately 35% of specimens digitised were collected in ODA-listed countries, in tropical Africa, but also in south east Asia and South America.</p>
The dataset of the Global Collections survey of natural history collections
<p>From 2016 to 2018, we surveyed the world’s largest natural history museum collections to begin mapping this globally distributed scientific infrastructure. The resulting dataset includes 73 institutions across the globe. It has:</p> <ul> <li> <p>Basic institution data for the 73 contributing institutions, including estimated total collection sizes, geographic locations (to the city) and latitude/longitude, and Research Organization Registry (ROR) identifiers where available.</p> </li> <li> <p>Resourcing information, covering the numbers of research, collections and volunteer staff in each institution.</p> </li> <li> <p>Indicators of the presence and size of collections within each institution broken down into a grid of 19 collection disciplines and 16 geographic regions.</p> </li> <li> <p>Measures of the depth and breadth of individual researcher experience across the same disciplines and geographic regions.</p> </li> </ul> <p>This dataset contains the data (raw and processed) collected for the survey, and specifications for the schema used to store the data. It includes:</p> <ol> <li>A diagram of the MySQL database schema.</li> <li>A SQL dump of the MySQL database schema, excluding the data.</li> <li>A SQL dump of the MySQL database schema with all data. This may be imported into an instance of MySQL Server to create a complete reconstruction of the database.</li> <li>Raw data from each database table in CSV format.</li> <li>A set of more human-readable views of the data in CSV format. These correspond to the database tables, but foreign keys are substituted for values from the linked tables to make the data easier to read and analyse.</li> <li>A text file containing the definitions of the size categories used in the collection_unit table.</li> </ol> <p>The global collections data may also be accessed at<a href="https://rebrand.ly/global-collections"> https://rebrand.ly/global-collections</a>. This is a preliminary dashboard, constructed and published using Microsoft Power BI, that enables the exploration of the data through a set of visualisations and filters. The dashboard consists of three pages:</p> <p><strong>Institutional profile:</strong> Enables the selection of a specific institution and provides summary information on the institution and its location, staffing, total collection size, collection breakdown and researcher expertise.</p> <p><strong>Overall heatmap:</strong> Supports an interactive exploration of the global picture, including a heatmap of collection distribution across the discipline and geographic categories, and visualisations that demonstrate the relative breadth of collections across institutions and correlations between collection size and breadth. Various filters allow the focus to be refined to specific regions and collection sizes.</p> <p><strong>Browse:</strong> Provides some alternative methods of filtering and visualising the global dataset to look at patterns in the distribution and size of different types of collections across the global view.</p>
Orogens of Big Sky Country: Reconstructing the Deep-Time Tectonothermal History of the Beartooth Mountains, Montana and Wyoming, USA (Supporting Information)
<p>Supporting datasets for Ronemus et al., "Orogens of Big Sky Country: Reconstructing the Deep-Time Tectonothermal History of the Beartooth Mountains, Montana and Wyoming, USA" in review at <em>Tectonics </em>as of August 16, 2022.</p> <p>Dataset S1. Detailed analytical settings and data for zircon U-Pb geochronology</p> <p>Dataset S2. Biotite <sup>40</sup>Ar/<sup>39</sup>Ar analytic results</p> <p>Dataset S3. Zircon (U-Th)/He analytic results</p> <p>Dataset S4. Apatite and zircon grain photomicrographs with measurements</p> <p>Dataset S5. Apatite (U-Th-Sm)/He analytic results</p> <p>Dataset S6. QTQt input and results files</p> <p> </p> <p>Datasets S1, S2, S3, and S5 are also available in the Tectonics submission supporting information.</p>
Data for "Ecosystem size filters life-history strategies to shape community assembly in lakes"
<p>Dataset 1. List of 71 fish species collected from north temperate lakes in Wisconsin USA. Data include critical life-history data used for strategy classifications according to Winemiller and Rose (1992), principal component scores, and strategy classification according to the cluster analysis.</p> <p>Dataset 2. Species occurrence data in all study lakes along with results from the 'soft classification" according to Euclidean distance.</p> <p>Dataset 3. Limnological and fish community characteristics of study lakes including species richness, lake area, estimated lake volume, and convex hull statistics for the overall fish community and each life-history strategy type.</p>
Mitochondrial genome sequencing and analysis of the invasive Microstegium vimineum: a resource for systematics, invasion history, and management
<p>Table S1: Accession data for Microstegium samples included in this study.</p> <p>File S1: Alignment of Mitochondrial CDS for Poales mitochondrial sequences.</p> <p>File S2: SNP data for Microstegium vimineum mitochondrial variants.</p> <p>Figure S1: Transposable element content in the Microstegium vimineum mitogenome.</p> <p>Figure S2: Summary of Kraken2 output.</p> <p> </p>
Precipitation and fire history at landslide sites
<p>These data include precipitation and burned area histories for events listed in the NASA Global Landslide Catalog. Each landslide includes a location uncertainty estimate. Precipitation values are the mean of all values within the uncertainty radius, while the fraction burned is computed for burned area.</p> <p>These data are intended to be used with the an RMarkdown notebook available at <a href="http://doi.org/10.5281/zenodo.7653683">this GitHub repository</a></p>
1805-1898 Census Records of Lausanne : a Long Digital Dataset for Demographic History
<p><strong>Context. </strong>This historical dataset stems from the project of automatic extraction of 72 census records of Lausanne, Switzerland. The complete dataset covers a century of historical demography in Lausanne (1805-1898), which corresponds to 18,831 pages, and nearly 6 million cells.</p> <p><strong>Content.</strong> The data published in this repository correspond to a first release, i.e. a diachronic slice of one register every 8 to 9 years. Unfortunately, the remaining data are currently under embargo. Their publication will take place as soon as possible, and at the latest by the end of 2023. In the meantime, the data presented here correspond to a large subset of 2,844 pages, which already allows to investigate most research hypotheses.</p> <p><strong>Description. </strong>The population censuses, digitized by the <a href="https://www.lausanne.ch/vie-pratique/culture/bibliotheques-et-archives/archives.html">Archives of the city of Lausanne</a>, continuously cover the evolution of the population in Lausanne throughout the 19th century, starting in 1805, with only one long interruption from 1814 to 1831. Highly detailed, they are an invaluable source for studying migration, economic and social history, and traces of cultural exchanges not only with Bern, but also with France and Italy. Indeed, the system of tracing family origin, specific to Switzerland, allows to follow the migratory movements of families long before the censuses appeared. The bourgeoisie is also an essential economic tracer. In addition, censuses extensively describe the organization of the social fabric into family nuclei, around which gravitate various boarders, workers, servants or apprentices, often living in the same apartment with the family.</p> <p><strong>Production. </strong>The structure and richness of censuses have also provided an opportunity to develop automatic methods for processing structured documents. The processing of censuses includes several steps, from the identification of text segments to the restructuring of information as digital tabular data, through Handwritten Text Recognition and the automatic segmentation of the structure using neural networks. Please note that the detailed extraction methodology, as well as the complete evaluation of performance and reliability is published in:</p> <ul> <li>Petitpierre R., Rappo L., Kramer M. (2023). <em>An end-to-end pipeline for historical censuses processing</em>. International Journal on Document Analysis and Recognition (IJDAR). doi: <a href="https://doi.org/10.1007/s10032-023-00428-9">10.1007/s10032-023-00428-9</a></li> </ul> <p><strong>Data structure.</strong> The data are structured in rows and columns, with each row corresponding to a household. Multiple entries in the same column for a single household are separated by vertical bars ⟨|⟩. The center point ⟨·⟩ indicates an empty entry. For some columns (e.g., street name, house number, owner name), an empty entry indicates that the last non-empty value should be carried over. The page number is in the last column.</p> <p><strong>Liability. </strong>The data presented here are not curated nor verified. They are the raw results of the extraction, the reliability of which was thoroughly assessed in the above-mentioned publication. We insist on the fact that for any reuse of this data for research purposes, the implementation of an appropriate methodology is necessary. This may typically include string distance heuristics, or statistical methodologies to deal with noise and uncertainty.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.