Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
486
datasets available to search
ShareScore release 0.7.1
Dataset results
486 results for “curation”
United States LEMIS wildlife trade data curated by EcoHealth Alliance
<p>Shared here are United States Fish and Wildlife Service (USFWS) Law Enforcement Management Information System (LEMIS) data on wildlife and wildlife product imports into the United States. This data was obtained via Freedom of Information Act (FOIA) requests by EcoHealth Alliance.</p> <p>Data were curated, cleaned, and made accessible via an R package interface: <a href="https://github.com/ecohealthalliance/lemis">https://github.com/ecohealthalliance/lemis</a>.</p> <p>Additionally, a summary of a portion of the data can be found in Smith et al. 2017, <em>EcoHealth </em>(<a href="https://doi.org/10.1007/s10393-017-1211-7">https://doi.org/10.1007/s10393-017-1211-7</a>).</p> <p>l<strong>emis_2000_2014_cleaned.csv</strong>: This file represents the compiled, cleaned LEMIS data from 2000-2014. This data is identical to the version 1.1.0 dataset available through the <strong>lemis </strong>R package.</p> <p><strong>lemis_codes.csv</strong>: Full values for all coded values used in the LEMIS data. Identical to the output from the <strong>lemis </strong>R package function "lemis_codes()".</p> <p><strong>lemis_metadata.csv</strong>: Data fields and field descriptions for all variables in the LEMIS data. Identical to the output from the <strong>lemis </strong>R package function "lemis_metadata()".</p> <p><strong>raw_data.zip</strong>: This archive contains all of the raw LEMIS data files that are processed and cleaned with the code contained in the 'data-raw' subdirectory of the <strong>lemis </strong>R package repository.</p>
AusTraits: a curated plant trait database for the Australian flora
<p>AusTraits is a transformative database, containing measurements on the traits of Australia's plant taxa, standardised from hundreds of disconnected primary sources. So far, data have been assembled from > 300 distinct sources, describing > 500 plant traits and > 34,000 taxa.</p> <p>To handle the harmonising of diverse data sources, we use a reproducible workflow to implement the various changes required for each source to reformat it suitable for incorporation in AusTraits. Such changes include restructuring datasets, renaming variables, changing variable units, changing taxon names. While this repository contains the harmonised data, the raw data and code used to build the resource are also available on the project's GitHub repository, <a href="https://github.com/traitecoevo/austraits.build/">https://github.com/traitecoevo/austraits.build/</a>.</p> <p>Further information on the project is available at the project website <a href="https://austraits.org">austraits.org</a> and in the associated publication (see below).</p> <p><strong>CONTRIBUTORS</strong></p> <p>The project is jointly led by Dr Daniel Falster (UNSW Sydney), Dr Rachael Gallagher (Western Sydney University), Dr Elizabeth Wenk (UNSW Sydney), and Dr Hervé Sauquet (Royal Botanic Gardens and Domain Trust Sydney), with input from > 300 contributors from over > 100 institutions (see full list above). The project was initiated by Dr Rachael Gallagher and Prof Ian Wright while at Macquarie University.</p> <p>We are grateful to the following institutions for contributing data Australian National Botanic Garden, Brisbane Rainforest Action and Information Network, Kew Botanic Gardens, National Herbarium of NSW, Northern Territory Herbarium, Queensland Herbarium, Western Australian Herbarium, South Australian Herbarium, State Herbarium of South Australia, Tasmanian Herbarium, Department of Environment Land Water and Planning Victoria and the Royal Botanic Gardens Victoria.</p> <p>AusTraits has been supported by investment from the Australian Research Data Commons (ARDC), via their "Transformative data collections" (https://doi.org/10.47486/TD044) and "Data Partnerships" (https://doi.org/10.47486/DP720, https://doi.org/10.47486/DP720A) programs; and grants from the Australian Research Council (FT160100113, DE170100208, FT100100910) and Macquarie University, The ARDC is enabled by National Collaborative Research Investment Strategy (NCRIS).</p> <p><strong>ACCESSING AND USE OF DATA</strong></p> <p>The compiled AusTraits database is released under an open source licence (CC-BY), enabling re-use by the community.</p> <p>A requirement of use is that users cite the AusTraits resource paper, which includes all contributors as co-authors:</p> <blockquote> <p>Falster, Gallagher et al (2021) <em>AusTraits, a curated plant trait database for the Australian flora</em>. Scientific Data 8: 254, <a href="https://doi.org/10.1038/s41597-021-01006-6">https://doi.org/10.1038/s41597-021-01006-6</a></p> </blockquote> <p>In addition, we encourage users you to cite the original data sources, wherever possible.</p> <p>Note that under the license data may be redistributed, provided the attribution is maintained.</p> <p>The downloads below provide the data in two formats:</p> <ul> <li>austraits-X.X.X.zip: data in plain text format (.csv, .bib, .yml files). Suitable for anyone, including those using Python.</li> <li>austraits-X.X.X.rds: data as compressed R object. Suitable for users of R (see below).</li> <li> <div>austraits-X.X.X-flattened.rds: contains a flattened version of the dataset for direct loading in R; all data tables are joined into a wider format</div> </li> <li> <div>austraits-X.X.X-flattened.parquet: contains a flattened version of the dataset in parquet format; all data tables are joined into a wider format </div> </li> </ul> <p>For R users, access and manipulation of data is assisted with the <a href="http://github.com/traitecoevo/austraits">austraits R package</a>. The package can both download data and provides examples and functions for running queries.<br><br><strong>STRUCTURE OF AUSTRAITS</strong></p> <p>The compiled AusTraits database contains a series of relational tables and files. These elements include all the data, contextual information submitted with each contributed datasets, database schema, and trait definitions. The file dictionary.html provides the same information in textual format. Similar information is available at <a href="https://traitecoevo.github.io/traits.build-book/">https://traitecoevo.github.io/traits.build-book/</a>.</p> <p><strong>CONTRIBUTING</strong></p> <p>We envision AusTraits as an on-going collaborative community resource that:</p> <ol> <li>Increases our collective understanding the Australian flora;</li> <li>Facilitates accumulation and sharing of trait data;</li> <li>Builds a sense of community among contributors and users; and</li> <li>Aspires to fully transparent and reproducible research of the highest standard.</li> </ol> <p>As a community resource, we are very keen for people to contribute. Assembly of the database is managed on GitHub at <a href="https://github.com/traitecoevo/austraits.build/">https://github.com/traitecoevo/austraits.build/</a>.</p> <p>Here are some of the ways you can contribute:</p> <p><strong>Reporting Errors</strong>: If you notice a possible error in AusTraits, please <a href="https://github.com/traitecoevo/austraits.build/issues">post an issue on GitHub</a>.</p> <p><strong>Refining documentation:</strong> We welcome additions and edits that make using the existing data or adding new data easier for the community.</p> <p><strong>Contributing new data</strong>: We gladly accept new data contributions to AusTraits. See full instructions on how to contribute at <a href="https://github.com/traitecoevo/austraits.build/">https://github.com/traitecoevo/austraits.build/</a>.</p>
BioASQ-QA: A manually curated corpus for Biomedical Question Answering
<p>The BioASQ question answering (QA) benchmark dataset contains questions in English, along with golden standard (reference) answers and related material. The dataset has been designed to reflect real information needs of biomedical experts and is therefore more realistic and challenging than most existing datasets. Furthermore, unlike most previous QA benchmarks that contain only exact answers, the BioASQ-QA dataset also includes ideal answers (in effect summaries), which are particularly useful for research on multi-document summarization. The dataset combines structured and unstructured data. The material linked with each question comprise documents and snippets, which are useful for Information Retrieval and Passage Retrieval experiments, as well as concepts that are useful in concept-to-text Natural Language Generation. Researchers working on paraphrasing and textual entailment can also measure the degree to which their methods improve the performance of biomedical QA systems. Last but not least, the dataset is continuously extended, as the BioASQ challenge is running and new data are generated.</p>
Curated list of HAR datasets
<p>A curated list of <em>preprocessed</em> & <em>ready to use under a minute</em> Human Activity Recognition datasets.</p> <p>All the datasets are preprocessed in <a href="https://www.hdfgroup.org/solutions/hdf5/">HDF5</a> format, created using the <a href="http://www.h5py.org">h5py</a> python library. Scripts used for data preprocessing are provided as well (Load.ipynb and load_jordao.py)</p> <p>Each HDF5 file contains at least the keys:</p> <ul> <li><code>x</code> a single array of size <code>[sample count, temporal length, sensor channel count]</code>, contains the actual sensor data. Metadata contains the names of individual sensor channel count. All samples are zero-padded for constant length in the file, original lengths before padding available under the <code>meta</code> keys.</li> <li><code>y</code> a single array of size <code>[sample count]</code> with integer values for target classes (zero-based). Metadata contains the names of the target classes.</li> <li><code>meta</code> contain various metadata, depends on the dataset (original length before padding, subject no., trial no., etc.)</li> </ul> <p>Usage example</p> <pre><code>import h5py with h5py.File(f'data/waveglove_multi.h5', 'r') as h5f: x = h5f['x'] y = h5f['y']['class'] print(f'WaveGlove-multi: {x.shape[0]} samples') print(f'Sensor channels: {h5f["x"].attrs["channels"]}') print(f'Target classes: {h5f["y"].attrs["labels"]}') first_sample = x[0] # Output: # WaveGlove-multi: 10044 samples # Sensor channels: ['acc1-x' 'acc1-y' 'acc1-z' 'gyro1-x' 'gyro1-y' 'gyro1-z' 'acc2-x' # 'acc2-y' 'acc2-z' 'gyro2-x' 'gyro2-y' 'gyro2-z' 'acc3-x' 'acc3-y' # 'acc3-z' 'gyro3-x' 'gyro3-y' 'gyro3-z' 'acc4-x' 'acc4-y' 'acc4-z' # 'gyro4-x' 'gyro4-y' 'gyro4-z' 'acc5-x' 'acc5-y' 'acc5-z' 'gyro5-x' # 'gyro5-y' 'gyro5-z'] # Target classes: ['null' 'hand swipe left' 'hand swipe right' 'pinch in' 'pinch out' # 'thumb double tap' 'grab' 'ungrab' 'page flip' 'peace' 'metal'] </code></pre> <p>Current list of datasets:</p> <ul> <li>WaveGlove-single (waveglove_single.h5)</li> <li>WaveGlove-multi (waveglove_multi.h5)</li> <li>uWave (uwave.h5)</li> <li>OPPORTUNITY (opportunity.h5)</li> <li>PAMAP2 (pamap2.h5)</li> <li>SKODA (skoda.h5)</li> <li>MHEALTH (non overlapping windows) (mhealth.h5)</li> <li>Six datasets with all four predefined train/test folds<br> as preprocessed by Jordao et al. originally in <a href="https://github.com/arturjordao/WearableSensorData">WearableSensorData</a><br> (FNOW, LOSO, LOTO and SNOW prefixed .h5 files)</li> </ul>
Curated mode-of-action data and effect concentrations for chemicals relevant for the aquatic environment
<p>Chemicals in the aquatic environment can be harmful to organisms and ecosystems. Knowledge on effect concentrations as well as on mechanisms and modes of interaction with biological molecules and signaling pathways is necessary to perform chemical risk assessment and identify toxic compounds. To this end, we developed criteria and a pipeline for harvesting and summarizing effect concentrations from the US ECOTOX database for the three aquatic species groups algae, crustaceans, and fish and researched the modes of action of more than 3,300 environmentally relevant chemicals in literature and databases. We provide a curated dataset ready to be used for risk assessment based on monitoring data and the first comprehensive collection and categorization of modes of action of environmental chemicals. Authorities, regulators, and scientists can use this data for the grouping of chemicals, the establishment of meaningful assessment groups, and the development of <em>in vitro</em> and <em>in silico</em> approaches for chemical testing and assessment.</p> <p> </p> <p> </p>
A Curated Gene and Biological System Annotation of Adverse Outcome Pathways Related to Human Health
<p>Adverse Outcome Pathways (AOPs) are multi-scale models of biological mechanisms connecting molecular initiating events to adverse outcomes through measurable key events. AOPs can guide the use and development of new approach methodologies (NAMs) aimed at reducing animal experimentation in chemical safety assessment. Here, we present a comprehensive molecular annotation of AOPs relevant to human health to embed the AOP framework into molecular data interpretation, which supports the development and application of novel AOP-based approaches in biomedical research.</p> <p>Please cite the following publication alongside this Zenodo entry when using the data:</p> <p>Saarimäki, L.A., Fratello, M., Pavel, A. <em>et al.</em> A curated gene and biological system annotation of adverse outcome pathways related to human health. <em>Sci Data</em> <strong>10</strong>, 409 (2023). https://doi.org/10.1038/s41597-023-02321-w</p>
A pangenome-guided manually curated library of transposable elements for Zymoseptoria tritici
<p>A manually-curated TE consensus library generated using a panel of 19 reference genomes for <em>Zymoseptoria tritici</em><sup>1-3</sup> along with reference genome assemblies for the sister species <em>Z. ardabiliae</em>, <em>Z. brevis</em>, <em>Z. pseudotritici</em>, and <em>Z. passerinii<sup>4</sup></em>. </p> <p> </p> <p><strong>Methods</strong></p> <p>Putative TE consensus sequences were first obtained by annotating all 23 genome assemblies<sup>1–4</sup> with Earl Grey with default settings (v3.0; <a href="https://github.com/TobyBaril/EarlGrey">https://github.com/TobyBaril/EarlGrey</a>)<sup>5,6</sup>. Consensus sequences generated from each reference genome were clustered using CD-Hit-Est (v4.8.1)<sup>7,8</sup> to group sequences with 90% similarity across 80% of the longer sequence length (<em>-n 8 -d 0 -aL 0.8 -c 0.90 -G 0 -g 1 -b 500 -r 1</em>) to reduce redundancy whilst preventing the collapsing of chimeric sequences. Consensus sequences <100bp were removed, as these are unlikely to represent true TE sequences. Each consensus sequence was then subject to manual curation as described by Goubert et al. (2022)<sup>9</sup>. Briefly, genomic copies of each TE were obtained using a “BLAST, Extract, Extend” process to recover genomic copies from each of the 23 reference genome assemblies with 1,000 flanking bases at either end<sup>9,10</sup>. For families with >100 BLASTN hits, the 25 longest hits were selected, along with 75 random hits. Multiple alignments were generated for each putative TE family using MAFFT (v7.505) with the --auto flag<sup>11</sup>. Columns composed of >=80% gaps were removed with T-COFFEE (v13.45.0.4846264)<sup>12</sup>. Subsequently, all sequence alignments were manually curated to define TE boundaries and remove regions of low conservation and rare insertions. Following manual curation, new majority-rule consensus sequences were generated with EMBOSS (v6.6.0.0) cons<sup>13</sup>. TE-Aid (<a href="https://github.com/clemgoub/TE-Aid/">https://github.com/clemgoub/TE-Aid/</a>) was used to aid visual inspection and to identify diagnostic features for classification of extended consensus sequences. Following this, TIRs were recorded if present, and nhmmscan (HMMER v3.3.2)<sup>14</sup> was used to identify homology to known curated elements in Dfam (v3.7). Combining this information, each TE consensus sequence was manually classified using available information following the naming convention ‘>ZymTri_2023_family_[n]#[Classification]/[Family]’. Consensus sequences classified with low confidence have a ‘?’ added to the name, as well as the string ‘_LowConf’. To reduce redundancy in the final TE library, sequences were clustered to the family-level using the 80-80-80 rule implemented in CD-hit-est<sup>9,15 </sup>(<em>-d 0 -aS 0.8 -c 0.8 -G 0 -g 1 -b 500 -r 1</em>). The representative sequence for each cluster was manually selected to select the sequence with the highest classification confidence, also defined as the ‘most intact consensus’. Chimeric sequences erroneously clustered were manually separated to retain sequences for the chimeric TE and the individual elements that generated the chimer.</p> <p> </p> <p><strong>References</strong></p> <p>1. Badet, T., Oggenfuss, U., Abraham, L., McDonald, B. A. & Croll, D. A 19-isolate reference-quality global pangenome for the fungal wheat pathogen Zymoseptoria tritici. <em>BMC Biol.</em> <strong>18</strong>, 12 (2020).</p> <p>2. Goodwin, S. B. <em>et al.</em> Finished genome of the fungal wheat pathogen Mycosphaerella graminicola reveals dispensome structure, chromosome plasticity, and stealth pathogenesis. <em>PLoS Genet.</em> <strong>7</strong>, e1002070 (2011).</p> <p>3. Plissonneau, C., Hartmann, F. E. & Croll, D. Pangenome analyses of the wheat pathogen Zymoseptoria tritici reveal the structural basis of a highly plastic eukaryotic genome. <em>BMC Biol.</em> <strong>16</strong>, 5 (2018).</p> <p>4. Feurtey, A. <em>et al.</em> Genome compartmentalization predates species divergence in the plant pathogen genus Zymoseptoria. <em>BMC Genomics</em> <strong>21</strong>, 588 (2020).</p> <p>5. Baril, T., Imrie, R. M. & Hayward, A. Earl Grey: a fully automated user-friendly transposable element annotation and analysis pipeline. (2022) doi:10.21203/rs.3.rs-1812599/v1.</p> <p>6. Baril, T., Galbraith, J. & Hayward, A. <em>Earl Grey</em>. (Zenodo, 2023). doi:10.5281/ZENODO.8116025.</p> <p>7. Li, W. & Godzik, A. Cd-hit: a fast program for clustering and comparing large sets of protein or nucleotide sequences. <em>Bioinformatics</em> <strong>22</strong>, 1658–1659 (2006).</p> <p>8. Fu, L., Niu, B., Zhu, Z., Wu, S. & Li, W. CD-HIT: accelerated for clustering the next-generation sequencing data. <em>Bioinformatics</em> <strong>28</strong>, 3150–3152 (2012).</p> <p>9. Goubert, C. <em>et al.</em> A beginner’s guide to manual curation of transposable elements. <em>Mob. DNA</em> <strong>13</strong>, 7 (2022).</p> <p>10. Camacho, C. <em>et al.</em> BLAST+: Architecture and applications. <em>BMC Bioinformatics</em> <strong>10</strong>, 1–9 (2009).</p> <p>11. Katoh, K. & Standley, D. M. MAFFT multiple sequence alignment software version 7: Improvements in performance and usability. <em>Mol. Biol. Evol.</em> <strong>30</strong>, 772–780 (2013).</p> <p>12. Notredame, C., Higgins, D. G. & Heringa, J. T-coffee: a novel method for fast and accurate multiple sequence alignment. <em>J. Mol. Biol.</em> <strong>302</strong>, 205–217 (2000).</p> <p>13. Rice, P., Longden, L. & Bleasby, A. EMBOSS: The European Molecular Biology Open Software Suite. <em>Trends Genet.</em> <strong>16</strong>, 276–277 (2000).</p> <p>14. Wheeler, T. J. & Eddy, S. R. nhmmer: DNA homology search with profile HMMs. <em>Bioinformatics</em> <strong>29</strong>, 2487–2489 (2013).</p> <p>15. Wicker, T. <em>et al.</em> A unified classification system for eukaryotic transposable elements. <em>Nat. Rev. Genet.</em> <strong>8</strong>, 973–982 (2007).</p>
Plumes and Blooms: Curated oceanographic and phytoplankton pigment observations
These data come from the Plumes and Blooms project (PnB), which has conducted approximately monthly 1-day oceanographic cruises since August 1996. The data included here encompass all PnB cruises from August 1996 through December 2018. Data are two data tables: one table includes conductivity-temperature-depth profiles (CTD) and derived physical parameters, the other table includes discrete seawater samples for various biological and biogeochemical parameters. Details are available in Catlett et al., in prep. References: Catlett, D., D. A. Siegel, R. D. Simons, N. Guillocheau, F. Henderikx-Freitas, C. S. Thomas.2021. Diagnosing seasonal to multi-decadal phytoplankton group dynamics in a highly productive coastal ecosystem, Progress in Oceanography. 197. https://doi.org/10.1016/j.pocean.2021.102637.
The Curated Courier: Digital Text Corpora from the UNESCO Courier (1948–2020)
<p>Founded in 1948 as the official magazine of the United Nations Educational, Scientific and Cultural Organization, <i>The UNESCO Courier</i> represents an extraordinary resource for research on global themes in the humanities. The complete <a href="https://en.unesco.org/courier/archives">archive of the magazine</a> is available in PDF form through UNESCO. These files make it possible for users anywhere to read individual issues, but it does not allow for full-text searching, much less any of the computational text analysis methods that have recently made important advances in humanities research.</p><p>The Curated Courier 1.0 is a package of digital text corpora, text analysis tools, and supplementary materials that makes the complete archive of <i>The UNESCO Courier</i> from 1948 to 2020 machine-readable, accessible, and reusable for digital text analysis. </p><p>Here on Zenodo we publish two <i>Courier</i> corpora. The first corpus (curated_courier_article_corpus) consists of the texts of all articles published in the English-language edition of <i>The UNESCO Courier</i> between 1948 and 2020. For this corpus we have extracted and reconstructed the complete text of all articles, for example by pulling together non-contiguous pages where necessary and by removing non-article text (masthead, photo captions, letters to the editor, and so on). We have linked each article to a comprehensive curated metadata index, included in the download (document_index.csv).</p><p>The second corpus (curated_issues) compiles the complete text of all <i>Courier</i> issues (English-language edition), 1948-2020. To prepare this corpus we extracted text from <a href="https://en.unesco.org/courier/archives">the PDFs that UNESCO has made available</a>, used multiple modes of OCR, and rendered each issue as a simple text file. Our test of the OCR quality finds an average error rate of 0.7 %, which should be considered good quality.</p><p>Working data from the process can be found in our <a href="https://github.com/inidun/tagged_courier">GitHub repository "tagged Courier."</a> The products, text analysis tools, and additional documentation are in the <a href="https://github.com/inidun/curated_courier">repository "Curated Courier."</a></p><p>The text of <i>The UNESCO Courier</i> is <a href="https://courier.unesco.org/en/about">available in Open Access</a> under the Attribution-ShareAlike 3.0 IGO (CC-BY-SA 3.0 IGO) license, in the context of <a href="https://en.unesco.org/open-access/">UNESCO's open access publications policy</a>. This dataset is published under the most recent version of the same license: Attribution-ShareAlike 4.0 International (<a href="https://creativecommons.org/licenses/by-sa/4.0/deed.en">CC BY-SA 4.0 Deed</a>).</p><p>These datasets was developed as part of the research project "International Ideas at UNESCO: Digital Approaches to Global Conceptual History" (INIDUN), led by Benjamin G. Martin at Uppsala University and funded by a grant from the Swedish Research Council (Vetenskapsrådet), 2020-2024. For more information, see: <a href="https://inidun.github.io">https://inidun.github.io</a>, as well as the<a href="https://github.com/inidun"> project repository on GitHub</a>, which includes documentation and files related to the curating process.</p>
pofatu/pofatu-data: Pofatu, a curated and open-access database for geochemical sourcing of archaeological materials
<p>Geochemical fingerprinting of artefacts and sources has proven to be the most effective way to use material evidence in order to reconstruct strategies of raw material procurement, exchange systems, and mobility patterns among past societies. In order to facilitate access to this growing body of data and to promote comparability and reproducibility in provenance studies, we designed Pofatu, the first online and open-access database presenting geochemical compositions and contextual information for archaeological sources and artefacts.</p> <p>The data repository includes a compilation of geochemical data and supporting analytical metadata, as well as the archaeological provenance and context for each sample. All information on Samples related to sources and artefacts can be accessed on this platform or downloaded from Zenodo or GitHub.</p> <p>While most prehistoric quarries and surface procurement sources used in the past have yet to be identified, provenance studies must also integrate wide and reliable geological data. For this reason, we advise Pofatu users to also consult other open-access repositories focusing specifically on geological samples, such as GeoRoc and EarthChem.</p>
Dataset - Papyrus 05.4 - A large scale curated dataset aimed at bioactivity predictions
<div> <p>This repository contains the version 05.5 of the Papyrus dataset, an aggregated dataset of small molecule bioactivities, as described in the article "Papyrus - A large scale curated dataset aimed at bioactivity predictions" <a href="https://doi.org/10.26434/chemrxiv-2021-1rxhk">https://doi.org/10.26434/chemrxiv-2021-1rxhk</a>.</p> <p> </p> <p>With the ongoing rapid growth of publicly available ligand-protein bioactivity data, there is a trove of valuable data that can be used to train a plethora of machine learning algorithms. However, not all data is equal in terms of size and quality and a significant portion of researchers’ time is needed to adapt the data to their needs. On top of that, finding the right data for a research question can often be a challenge on its own. To meet these challenges we have constructed the Papyrus dataset. Papyrus is comprised of around 60 million datapoints. This dataset contains multiple large publicly available datasets such as ChEMBL and ExCAPE-DB combined with several smaller datasets containing high-quality data. The aggregated data has been standardised and normalised in a manner that is suitable for machine learning. We show how data can be filtered in a variety of ways and also perform some example quantitative structure-activity relationship analyses and proteochemometric modelling. Our ambition is that this pruned data collection constitutes a benchmark set that can be used for constructing predictive models, while also providing a solid baseline for related research.</p> </div>
Accompanying data - Papyrus - A large scale curated dataset aimed at bioactivity predictions
<p><strong>Addition of supporting files:<br>- </strong>LICENSE.txt<strong><br>- </strong>data_types.json<strong><br>- </strong>data_size.json</p> <p> </p> <p><strong>Fixed version of Papyrus++ 05.5:<br>- In the previous 05.5 version </strong>data was incorrectly duplicated based on assay type. This resulted in unintended data augmentation.<br><strong>- In this fixed 05.5 version</strong> the duplicates have been eliminated, now reporting the correct amount of data per assay type.</p> <p> </p> <p>This repository contains the version 05.5 of the Papyrus dataset, an aggregated dataset of small molecule bioactivities, as described in the article "Papyrus - A large scale curated dataset aimed at bioactivity predictions" <a href="http://doi.org/10.1186/s13321-022-00672-x">http://doi.org/10.1186/s13321-022-00672-x</a>.</p> <p> </p> <p>With the ongoing rapid growth of publicly available ligand-protein bioactivity data, there is a trove of valuable data that can be used to train a plethora of machine learning algorithms. However, not all data is equal in terms of size and quality and a significant portion of researchers’ time is needed to adapt the data to their needs. On top of that, finding the right data for a research question can often be a challenge on its own. To meet these challenges we have constructed the Papyrus dataset. Papyrus is comprised of around 60 million datapoints. This dataset contains multiple large publicly available datasets such as ChEMBL and ExCAPE-DB combined with several smaller datasets containing high-quality data. The aggregated data has been standardised and normalised in a manner that is suitable for machine learning. We show how data can be filtered in a variety of ways and also perform some example quantitative structure-activity relationship analyses and proteochemometric modelling. Our ambition is that this pruned data collection constitutes a benchmark set that can be used for constructing predictive models, while also providing a solid baseline for related research.</p>
Understanding the Publish-Review-Curate (PRC) Model of Scholarly Communication - Data and Code
<p>Summary data for the number of articles submitted to publish-review-curate platforms as of August 2024 (Figure 1) [Update 14 Nov 2024: Added JMIRx. Data still from August 2024]</p> <p>Summary data for the number of articles reviewed by review platforms (Figure 2)</p> <p>Analysis code to produce Figures 1 and 2</p> <p>Code to extract articles for inclusion in data</p>
Curated Estonian National Bibliography - persons
<p>This curated dataset is derived from the persons authority file of the Estonian National Bibliography (ENB), a comprehensive catalog of publications written in Estonian, published in Estonia, or focusing on Estonian culture and people. Designed for computational analysis, this dataset adapts the original authority file for research and cultural exploration. Through a systematic process of filtering, cleaning, and harmonizing, the ENB dataset is presented in a streamlined tabular format that retains rich metadata while improving accessibility. Fields selected for inclusion are harmonized and, where possible, linked to external sources, offering an optimized and reproducible resource for historical, cultural, and bibliographic research.</p>
Curated Estonian National Bibliography - books
<p>This curated dataset is derived from the books subset of the Estonian National Bibliography (ENB), a comprehensive catalog of publications written in Estonian, published in Estonia, or focusing on Estonian culture and people. Designed for computational analysis, this dataset adapts the original catalog for research and cultural exploration.</p> <p>Through a systematic process of filtering, cleaning, and harmonizing, the ENB dataset is presented in a streamlined tabular format that retains rich metadata while improving accessibility. Fields selected for inclusion are harmonized and, where possible, linked to external sources, offering an optimized and reproducible resource for historical, cultural, and bibliographic research.</p>
ChemTastesDB: A Curated Database of Molecular Tastants
<p><em><strong>ChemTastesDB</strong></em> is a database that includes curated information of 4075 molecular tastants. <strong><em>ChemTastesDB</em></strong> is distributed to the scientific community to expand the information of molecular tastants, which could assist the analysis of the relationships between molecular structure and taste, as well as <em>in</em> <em>silico</em> (QSAR/QSPR) studies for taste prediction. Examples of QSPR approaches for the prediction of molecular taste are given in the following publication: <em>Rojas, C., Abril-González, M., Ballabio, D. & García, F. (2025). ChemTastesPredictor: An ensemble of machine learning classifiers to predict the taste of molecular tastants. Chemometrics and Intelligent Laboratory Systems. 261, 105380. <a href="https://doi.org/10.1016/j.chemolab.2025.105380">https://doi.org/10.1016/j.chemolab.2025.105380</a>.</em></p> <p>The 4075 molecular tastants are categorized into one of the five basic tastes (sweet, bitter, umami sour and salty), as well as to other classes related to non-basic tastes (tasteless, non-sweet, non-bitter, multitaste and miscellaneous). The molecules are categorized into following ten classes: sweet (1313), bitter (1615), umami (220), sour (49), salty (16), multitaste (179), tasteless (232), non-sweet (304), non-bitter (28), and miscellaneous (119).</p> <p><strong><em>ChemTastesDB</em></strong> provides the following information for each molecule: name, PubChem CID, CAS registry number, canonical SMILES string, class taste and the reference to the scientific sources from where data were retrieved. In addition, the molecular structure in the HyperChem (<em>.hin</em>) format of each compound is provided.</p> <p>This is version 2.1 of the <em><strong>ChemTastesDB</strong></em>. In this new version, 1131 newly curated compounds were added. These new molecules were retrieved from 52 new bibliographic references.</p>
Fatiando a Terra data v1.0.0: A curated collection of open geophysics data for tutorials and documentation
<p>This repository holds curated sample datasets that can be used in the documentation and tutorials of the <a href="https://www.fatiando.org/">Fatiando a Terra</a> project. All datasets are cleaned and formatted versions of openly available data under permissive licenses or in the public domain.</p> <p>More information about datasets and the code for cleaning, formatting, and preprocessing the data can be found at: <a href="https://github.com/fatiando/data">https://github.com/fatiando/data</a></p> <p>See the README.md file for information on data sources and their original licenses.</p> <p><strong>NOTE:</strong> This collection uses <a href="https://semver.org/">semantic versioning</a> (i.e., MAJOR.MINOR.BUGFIX). Major releases mean that backwards incompatible changes were made to the data. Minor releases add new data without changing existing files. Bug fix releases fix errors in a previous release that makes the data unusable. Changes to the current data files will always be published as a major release unless the file(s) in the previous release was unusable/corrupted.</p>
Manually curated transcriptomics data collection for toxicogenomic assessment of engineered nanomaterials
<p>Toxicogenomics (TGx) approaches are increasingly applied to gain insight into the possible toxicity mechanisms of engineered nanomaterials (ENMs). Omics data can be valuable to elucidate the mechanism of action of chemicals and develop predictive models in toxicology. While vast amounts of transcriptomics data from ENM exposures have already been accumulated, a unified, easily accessible and reusable collection of transcriptomics data for ENMs is currently lacking. In an attempt to improve the FAIRness of already existing transcriptomics data for nanomaterials, we curated a collection of homogenized transcriptomics data from human, mouse and rat ENM exposures <em>in vitro</em> and <em>in vivo</em>.</p>
Sharkipedia: A Curated Open Access Database of Shark and Ray Life History Traits and Abundance Time-series
<p>This dataset represent the intial launch of Sharkipedia: a curated open access database of shark and ray life history traits and abundance time-series. A curated database of shark and ray biological data is increasingly necessary both to support fisheries management and conservation efforts, and to test the generality of hypotheses of vertebrate macroecology and macroevolution. Sharks and rays are one of the most charismatic, evolutionary distinct, and threatened lineages of vertebrates, comprising around 1,250 species. To accelerate shark and ray conservation and science, we developed Sharkipedia as a curated open-source database and research initiative to make all published biological traits and population trends accessible to everyone. Sharkipedia hosts information on 58 life history traits from 264 sources, for 170 species, from 39 families, and 12 orders related to length (n=9 traits), age (8), growth (12), reproduction (19), demography (5), and allometric relationships (5), as well as 871 population time-series from 202 species. Sharkipedia relies on the backbone taxonomy of the IUCN Red List and the bibliography of Shark-References. Sharkipedia has profound potential to support the rapidly growing data demands of fisheries management, international trade regulation as well as anchoring vertebrate macroecology and macroevolution.</p>
EukRibo: a manually curated eukaryotic 18S rDNA reference database
<p>EukRibo is a manually curated database of reference small-subunit ribosomal RNA gene (18S rDNA) sequences of eukaryotes, specifically aimed at taxonomic annotation of high-throughput metabarcoding datasets. Unlike other reference databases of ribosomal genes, it is not meant to exhaustively capture all publicly available 18S rDNA sequences from the INSDC repositories, but to represent a subset of highly trustable sequences covering the whole known diversity of eukaryotes, with a focus on protists, manually verified taxonomic identifications, and relatively low genetic redundancy.</p> <p>EukRibo is part of a suite of public resources generated by the UniEuk project (www.unieuk.org), which are all designed to follow a common taxonomic framework for maximal interoperability. The high level of taxonomic accuracy of EukRibo, together with a newly designed, phylogenetically-informed annotation approach, allow high confidence in the taxonomic annotation of environmental metabarcodes, as well as identification of new eukaryotic diversity at various taxonomic levels using a connected components approach.</p> <p>* * *</p> <p>Accompanying preprint available at <a href="https://doi.org/10.1101/2022.11.03.515105">https://doi.org/10.1101/2022.11.03.515105</a>.</p> <p>* * *</p> <p><strong>EukRibo ReadMe file, versions 1 and 2</strong></p> <p>Each EukRibo release consists of <strong>4 files</strong>:<br> - a <strong>tsv table </strong>containing the taxonomic and other information about the 18S rDNA sequences included in the release<br> - a <strong>fasta file </strong>containing the <strong>full sequences </strong>as retrieved from the INSDC repositories (NCBI, EMBL-EBI/ENA, DDBJ)<br> - a <strong>fasta file </strong>containing the <strong>variable region V4 </strong>extracted from all these sequences (based on the fragment amplified with the Tara-Oceans V4 primers)<br> - a <strong>fasta file </strong>containing the <strong>variable region V9 </strong>extracted from the subset of sequences where it is present (based on the fragment amplified with the Tara-Oceans V9 primers)</p> <p>The primary goal of EukRibo was to be used to annotate the EukBank meta-dataset of available V4 metabarcoding datasets, and therefore all sequences included in EukRibo contain the variable region V4.<br> Only a subset of these sequences (about 75%) also contain the variable region V9; this is because many 18S rDNA sequences in the INSDC repositories stop before the V9 fragment.</p> <p>Sequences with slightly incomplete V4 or V9 fragments were kept if phylogenetically useful - i.e. if they are the only available representatives of a certain taxonomic lineage.<br> <strong>V4 </strong>We allowed up to 50 missing positions in the relatively conserved area at the 5' end of the V4 fragment (for an average fragment length of about 380 bp); no sequence incomplete at the 3' end of the V4 fragment is included.<br> <strong>V9 </strong>We allowed up to 30 missing positions in the relatively conserved area at the 3' end of the V9 fragment (for an average length of about 135 bp); no sequence incomplete at the 5' end of the V9 fragment is included.<br> We allowed a higher proportion of missing positions for the V9 region because being more conservative would imply losing too many sequences, including entire taxonomic lineages.</p> <p><strong>Version 1 of EukRibo</strong><br> This is the starting version of EukRibo that was used for the taxonomic annotation of the EukBank dataset, with taxonomy strings that were fixed as of October 2020.<br> - Contains 46,345 sequences with a sufficiently complete V4 region; 46,299 with the actual complete V4 region and 46 (about 0.1%) with missing positions at the 5' end.<br> - Of these, 34,438 also include a sufficiently complete V9 region; 23,226 with the actual complete V9 region and 11,206 (about 33%) with missing positions at the 3' end.</p> <p><strong>Version 2 of EukRibo</strong><br> This is a version of EukRibo that was made taxonomically compatible with version 3 of the EukProt database (<a href="https://doi.org/10.1101/2020.06.30.180687">https://doi.org/10.1101/2020.06.30.180687</a>), with taxonomic revisions as of July 2022 as well as additional information on the included selection of sequences that was not provided in the tsv file of version 1.<br> - Contains the exact same selection of sequences as in version 1, with the addition of genus <em>Meteora</em>, the last remaining known supergroup-level eukaryotic lineage for which an 18S rDNA was not previously available. (The <em>Meteora </em>sequence contains the full V4 fragment but does not include a sufficiently complete V9 fragment.)<br> - Only 34,432 sequences with a sufficiently complete V9 region are now retained because of 6 previously unrecognised chimeric sequences where the V9 fragment does not originate from the same organism as the V4 fragment.</p> <p><strong>Files in EukRibo version 1</strong>:<br> 46345_EukRibo.tsv.gz<br> 46345_EukRibo_full_seqs.fas.gz<br> 46345_EukRibo_V4.fas.gz<br> 34438_EukRibo_V9.fas.gz</p> <p>The tsv file contains 6 columns:<br> <strong>gb_accession </strong>- INSDC accession number of the sequence<br> <strong>supergroup</strong>, <strong>taxogroup1</strong>, <strong>taxogroup2 </strong>- binning of the taxa into strictly monophyletic clades of evolutionary and/or ecological significance<br> <strong>UniEuk_taxonomy_string </strong>- full UniEuk-compatible taxonomic annotation of the sequence<br> - an unlimited number of levels is allowed (going down to strain for isolated organisms or to clone for environmental sequences)<br> - informal names are used for phylogenetically supported clades without formal name<br> <strong>V9 </strong>- presence ('Y') or absence ('N') of a sufficiently complete V9 fragment in the sequence</p> <p><strong>Files in EukRibo version 2</strong>:<br> 46346_EukRibo-02.tsv.gz<br> 46346_EukRibo-02_full_seqs.fas.gz<br> 46346_EukRibo-02_V4.fas.gz<br> 34432_EukRibo-02_V9.fas.gz</p> <p>The tsv file now contains 12 columns:<br> <strong>gb_accession</strong>, <strong>supergroup</strong>, <strong>taxogroup1</strong>, <strong>taxogroup2</strong>, <strong>UniEuk_taxonomy_string</strong><br> - same columns as in version 1<br> <strong>alternative_strain_names </strong>(new) - provides alternative strain/isolate names when known to help cross-linking genetic data coming from the same organism<br> <strong>V4 </strong>(new) - indicates whether the V4 fragment is complete ('yes - complete') or missing positions at the 5' end ('yes - partial')<br> <strong>V9 </strong>(emended content) - now contains more precise information than in version 1 about whether it is complete ('yes - complete'), missing positions at the 3' end ('yes - partial'), or was excluded, and the 6 possible reasons why ('no - missing', 'no - too incomplete', 'no - chimera', 'no - bad quality', 'no - deletion in V9', 'no - Ns in V9')<br> <strong>EukProt_ID_same_strain </strong>(new) - accession of EukProt datasets from the same isolate<br> <strong>EukProt_ID_different_strain </strong>(new) - accession of EukProt datasets from a different isolate of the same species<br> <strong>columns_modified_since_previous_version </strong>(new) - lists all of the 6 pre-existing columns that have a modified content compared to version 1<br> <strong>remarks </strong>(new) - additional information such as presence of an intron in the V9 fragment, taxonomic identity of the two parts of chimeric sequences, or the presence of Ns or a deletion in the V4 or the V9 fragment (but insufficient to warrant exclusion)</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.