Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

11

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

11 results for “palaeoproteomics”

Learn how ShareScore rates datasets ↗
zenodo44/100

Hominid Palaeoproteomic Reference Dataset

<p>This dataset contains the &#39;Hominid Palaeoproteomic Reference Dataset&#39;.We used PaleoProPhyler ( https://github.com/johnpatramanis/Proteomic_Pipeline )&nbsp; to generate a palaeoproteomic reference dataset of protein sequences from ancient and present-day hominids. Using the first two modules of PaleoProPhyler, we translated 176 publicly available whole genomes from extant and extinct hominid groups.</p> <p>We also translated 8 ancient hominin genomes from VCF files, including those of 3 Neanderthals and one Denisovan. Since the dataset is tailored for palaeoproteomic analyses, we chose&nbsp;to translate proteins that have previously been&nbsp;reported as present in either teeth or bones. We compiled a list of 1,696 proteins from previous works and successfully translated 1,543 of them.&nbsp;For each protein, both the canonical and all alternative protein coding isoforms were translated, leading to a total of around 10,058 protein sequences for each individual in the dataset.</p> <p>Details on the processing of the sequences can be found in the supplementary materials of PaleoProPhyler ( https://github.com/johnpatramanis/Proteomic_Pipeline/blob/main/GitHub_Tutorial/Supplementary.pdf ). The full list of the proteins translated can be found here:&nbsp;https://github.com/johnpatramanis/Proteomic_Pipeline/blob/main/Reference_Protein_List.txt and a table with information on each sample included in the dataset can be found here:&nbsp;https://github.com/johnpatramanis/Proteomic_Pipeline/blob/main/Reference_Sample_List.csv</p> <p>&nbsp;</p> <p>Content:&nbsp;</p> <p>The zipped file contains 5 files: one txt file,&nbsp;two fasta files as well as two additional folders:</p> <p>-&nbsp;&nbsp;PalaeoProPhyler_Publication_Data_for_Tree.fa contains all of the sequences used to generate the phylogenetic tree presented at PalaeoProPhylers manuscript.</p> <p>-&nbsp;ALL_PROT_REFERENCE.fa contains all of the sequences generated as part of the Hominid Palaeoproteomic Reference Dataset described above, all in a single fasta.</p> <p>-&nbsp;PER_PROTEIN is a folder containing one fasta file for each protein within the&nbsp;Hominid Palaeoproteomic Reference Dataset. Each protein fasta file has the sequences of all individuals for that particular protein.</p> <p>-&nbsp;PER_SAMPLE is a folder containing one fasta file for each sample/individual within the&nbsp;Hominid Palaeoproteomic Reference Dataset. Each sample fasta file has the sequences of all proteins for that particular sample.</p> <p>-Reference_Protein_List.txt is a txt file containing two columns. The first column is a list of all the proteins selected to be translated. The second column describes where each of these proteins&nbsp;was mentioned or identified. If a protein was identified in a publication, the title of the publication is given. If a protein was&nbsp;identified in one of the publications of our group (E.Cappellini group) the identifier &#39;our samples&#39; is given. If multiple publications supported a protein they are all given and seperated by comma.</p> <p>&nbsp;</p> <p>~ NOTE (!) ~</p> <p>Depending on which samples you use we highly encourage you to cite the original publication(s) from which we got the DNA data and translate the proteins from:</p> <p>&nbsp;</p> <p>Modern Humans:</p> <p>M Byrska-Bishop et al. &ldquo;High Coverage Whole Genome Sequencing of the Expanded 1000 Genomes Project Cohort Including 602 Trios. bioRxiv. 2021&rdquo;.</p> <p>Non-human great apes:</p> <p>Javier Prado-Martinez et al. &ldquo;Great ape genetic diversity and population history&rdquo;. In: Nature 499.7459 (2013), pp. 471&ndash;475.</p> <p>Pongos:</p> <p>Alexander Nater et al. &ldquo;Morphometric, behavioral, and genomic evidence for a new orangutan species&rdquo;. In: Current Biology 27.22 (2017), pp. 3487&ndash;3498.</p> <p>Neanderthal, Denisovan and other ancient anatomically modern humans:</p> <p>Kay Pr&uml;ufer et al. &ldquo;A high-coverage Neandertal genome from Vindija Cave in Croatia&rdquo;. In: Science 358.6363 (2017), pp. 655&ndash;658.</p> <p>Fabrizio Mafessoni et al. &ldquo;A high-coverage Neandertal genome from Chagyrskaya Cave&rdquo;. In: Proceedings of the National Academy of Sciences 117.26 (2020), pp. 15132&ndash;15136.</p>

opencc-by-4.0Nov 2022View details →
zenodo44/100

Tryps-IN: A streamlined palaeoproteomics workflow enables ZooMS analysis of 10,000-year-old petrous bones from Jordan rift-valley

<p>Poor preservation of collagen in dry and/or arid environments has hindered the application of Zooarchaeology by mass spectrometry (ZooMS) analysis in many regions of the world, and as a result many zooarchaeological investigations have relied exclusively on the morphological assessment of fragmentary remains, due to the inadequate preservation of biomolecules. The climatic conditions of Southwest Asia include extreme temperature fluctuations unconducive to preservation of proteins and DNA. We performed zooarchaeological analysis of remains from the 10,000-year-old site of Shkārat Msaied in Jordan and sub-sampled twenty-eight petrous bones, the hardest bone in the mammalian skeleton, for species identification by ZooMS. Using an unconventional and simplified extraction protocol we call Tryps-IN, in which digestion was performed without removal of the demineralising EDTA, we taxonomically identified several fragments, outperforming the established ZooMS work-flow. A subset of identifications was subsequently confirmed using liquid chromatography coupled to tandem mass spectrometry (LC-MS/MS) protein sequencing. The new methodology presented here opens the possibility of further bioarchaeological investigation of other fragmentary faunal assemblages within this region of archaeological significance.&nbsp;</p>

opencc-by-4.0Apr 2023View details →
zenodo44/100

Proboscidean Palaeoproteomic Reference Dataset

<p>This entry contains the &#39;Proboscidean Palaeoproteomic Reference Dataset&#39;.</p> <p>We used PaleoProPhyler ( https://github.com/johnpatramanis/Proteomic_Pipeline )&nbsp;to generate a palaeoproteomic reference dataset of protein sequences from ancient and present-day Proboscidae. Using the first two modules of PaleoProPhyler, we translated more than 35 publicly available whole genomes from extant and extinct species. Details on the processing of the sequences can be found below.</p> <p>&nbsp;</p> <p>Which individuals / species are included?</p> <p>The full list of individuals, the original fastq repository location and the species included in the dataset are contained within the tab seperated file &#39;METATADATA.txt&#39;, that also contains headers. Most individuals of the dataset&nbsp;belong to one of these 3 species: <em>Loxodonta africana</em>,<em> Elephas maximus</em>,&nbsp;<em>Mammuthus&nbsp; primigenius.&nbsp;</em></p> <p>&nbsp;</p> <p>Which Proteins are included?</p> <p>We compiled a small initial&nbsp;list of 262 proteins that had been indentified in either teeth, bone or items made out of ivory. For each protein, both the canonical and all alternative protein coding isoforms (based on the Loxodonta africana reference proteome of Ensembl)&nbsp;were translated, leading to more than 350&nbsp;unique protein sequences for each individual in the dataset. The protein list is available in the file &#39;proteins.txt&#39;</p> <p>&nbsp;</p> <p>How were the proteins translated/generated?</p> <p>All genetic data were downloaded from ENA (https://www.ebi.ac.uk/ena/browser/home) as fastq files and mapped onto LoxAfr3, which is the latest annotated African elephant genome in Ensembl. The scripts used for the mapping are available here:&nbsp;https://github.com/johnpatramanis/Mapping_Scripts . We used the resulting bam files as input for PaleoProPhyler&#39;s module 1 &amp; 2 , using LoxAfr3 as the reference proteome.</p> <p>Other files included in the zip folder:</p> <p>-&nbsp;ALL_PROT_REFERENCE.fa contains all of the sequences generated as part of the Proboscidean Palaeoproteomic Reference Dataset described above</p> <p>-&nbsp;PER_PROTEIN is a folder containing one fasta file for each protein within the&nbsp;Proboscidean Palaeoproteomic Reference Dataset, each protein fasta file has the sequences of all individuals for that particular protein</p> <p>-&nbsp;PER_SAMPLE is a folder containing one fasta file for each sample/individual within the&nbsp;Proboscidean Palaeoproteomic Reference Dataset, each sample fasta file has the sequences of all proteins for that particular sample.</p>

opencc-by-4.0Apr 2023View details →
dryad40/100

The name of the game: Palaeoproteomics and radiocarbon dates further refine the presence and dispersal of caprines in Eastern and Southern Africa

<p><span>We report the first large-scale palaeoproteomics research on eastern and southern African zooarchaeological samples, thereby refining our understanding of early caprine (sheep and goat) pastoralism in Africa. Assessing caprine introductions is a complicated task because of their skeletal similarity to endemic wild bovid species and the sparse and fragmentary state of relevant archaeological remains. Palaeoproteomics has previously proved effective in clarifying species attributions in African zooarchaeological materials, but few comparative protein sequences of wild bovid species have been available. Using newly generated collagen type I sequences for wild species, as well as previously published sequences, we assess species attributions for elements originally identified as caprine or "unidentifiable bovid" from seventeen eastern and southern African sites that span seven millennia. We identified over 70% of the archaeological remains and the direct radiocarbon dating of domesticate specimens allows refinement of the chronology of caprine presence in both African regions. These results thus confirm earlier occurrences in eastern Africa and the systematic association of domesticated caprines with wild bovids at all archaeological sites. The combined biomolecular approach highlights repeatability and accuracy of the methods for conclusive contribution in species attribution of archaeological remains in dry African environments.</span></p>

opencc-zeroNov 2023View details →
dryad40/100

Increasing sustainability in palaeoproteomics by optimizing digestion times for large-scale archaeological bone analyses

<p>Palaeoproteomic analysis of skeletal proteomes is used to provide taxonomic identifications for an increasing number of archaeological specimens. The success rate depends on a range of taphonomic factors and differences in the extraction protocols employed. By analyzing 12 archaeological bone specimens from two archaeological sites, we demonstrate that reducing digestion duration from 18 to 3 hours has no measurable impact on the obtained taxonomic identifications. Peptide marker recovery, COL1 sequence coverage, or proteome complexity are also not significantly impacted. Although we observe minor differences in sequence coverage and glutamine deamidation, these are not consistent across our dataset. A 6-fold reduction in digestion time reduces electricity consumption, and therefore CO<sub>2</sub> emission intensities. We furthermore demonstrate that working in 96-well plates further reduces electricity consumption by 60%, in comparison to individual microtubes. Reducing digestion time therefore has no impact on the taxonomic identifications, while reducing the environmental impact of palaeoproteomic projects.</p>

opencc-zeroApr 2024View details →
zenodo40/100

A comparative study of commercially available, minimally invasive, sampling methods on Early Neolithic humeri analysed via palaeoproteomics

Open the record for dataset details and reuse information.

opencc-by-4.0Apr 2024View details →
dryad40/100

Increasing sustainability in palaeoproteomics by optimizing digestion times for large-scale archaeological bone analyses

Open the record for dataset details and reuse information.

publicApr 2024View details →
dryad40/100

The name of the game: Palaeoproteomics and radiocarbon dates further refine the presence and dispersal of caprines in Eastern and Southern Africa

Open the record for dataset details and reuse information.

publicNov 2023View details →
zenodo36/100

Tryps-IN: A streamlined palaeoproteomics workflow enables ZooMS analysis of 10,000-year-old petrous bones from Jordan rift-valley

<p>Poor preservation of collagen in dry and/or arid environments has hindered the application of Zooarchaeology by mass spectrometry (ZooMS) analysis in many regions of the world, and as a result many zooarchaeological investigations have relied exclusively on the morphological assessment of fragmentary remains, due to the inadequate preservation of biomolecules. The climatic conditions of Southwest Asia include extreme temperature fluctuations unconducive to preservation of proteins and DNA. We performed zooarchaeological analysis of remains from the 10,000-year-old site of Shkārat Msaied in Jordan and sub-sampled twenty-eight petrous bones, the hardest bone in the mammalian skeleton, for species identification by ZooMS. Using an unconventional and simplified extraction protocol we call Tryps-IN, in which digestion was performed without removal of the demineralising EDTA, we taxonomically identified several fragments, outperforming the established ZooMS work-flow. A subset of identifications was subsequently confirmed using liquid chromatography coupled to tandem mass spectrometry (LC-MS/MS) protein sequencing. The new methodology presented here opens the possibility of further bioarchaeological investigation of other fragmentary faunal assemblages within this region of archaeological significance.&nbsp;</p>

opencc-by-4.0Apr 2023View details →
dryad32/100

Data from: Palaeoproteomics resolves sloth phylogeny

The living tree sloths Choloepus and Bradypus are the only remaining members of Folivora, a major xenarthran radiation that occupied a wide range of habitats in many parts of the western hemisphere during the Cenozoic, including both continents and the West Indies. Ancient DNA evidence has played only a minor role in folivoran systematics, as most sloths lived in places not conducive to genomic preservation. Here we utilize collagen sequence information, both separately and in combination with published mitochondrial DNA evidence, to assess the relationships of tree sloths and their extinct relatives. Results from phylogenetic analysis of these datasets differ substantially from morphology-based concepts: Choloepus groups with Mylodontidae, not Megalonychidae; Bradypus and Megalonyx pair together as megatherioids, while monophyletic Antillean sloths may be sister to all other folivorans. Divergence estimates are consistent with fossil evidence for mid-Cenozoic presence of sloths in the West Indies and an early Miocene radiation in South America.

opencc-zeroDec 2018View details →
dryad32/100

Data from: Palaeoproteomics resolves sloth phylogeny

Open the record for dataset details and reuse information.

publicJun 2019View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record