Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

31

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

31 results for “Biodiversity informatics”

Learn how ShareScore rates datasets ↗
zenodo24/100

Figure 2 from: Baker E, Johnson K, Young J (2011) The future of the past in the present: biodiversity informatics and geological time. ZooKeys 150: 397-405. https://doi.org/10.3897/zookeys.150.2350

Figure 2 - Nannotax Screenshot

opencc-by-4.0Nov 2011View details →
zenodo24/100

Figure 6 from: Baker E, Johnson K, Young J (2011) The future of the past in the present: biodiversity informatics and geological time. ZooKeys 150: 397-405. https://doi.org/10.3897/zookeys.150.2350

Figure 6 - When there is no overlapping time periods the intersect is undefined.

opencc-by-4.0Nov 2011View details →
zenodo24/100

Figure 5 from: Remsen D (2016) The use and limits of scientific names in biological informatics. In: Michel E (Ed.) Anchoring Biodiversity Information: From Sherborn to the 21st century and beyond. ZooKeys 550: 207–223. https://doi.org/10.3897/zookeys.550.9546

Figure 5 - A polyseme is a single name referring to more than one overlapping or included concept.

opencc-by-4.0Jan 2016View details →
zenodo24/100

Figure 3 from: Remsen D (2016) The use and limits of scientific names in biological informatics. In: Michel E (Ed.) Anchoring Biodiversity Information: From Sherborn to the 21st century and beyond. ZooKeys 550: 207–223. https://doi.org/10.3897/zookeys.550.9546

Figure 3 - The semiotic triangle describes how names communicate meaning.

opencc-by-4.0Jan 2016View details →
zenodo24/100

Figure 2 from: Remsen D (2016) The use and limits of scientific names in biological informatics. In: Michel E (Ed.) Anchoring Biodiversity Information: From Sherborn to the 21st century and beyond. ZooKeys 550: 207–223. https://doi.org/10.3897/zookeys.550.9546

Figure 2 - All information of a species is linked by a name.

opencc-by-4.0Jan 2016View details →
zenodo24/100

Figure 4 from: Remsen D (2016) The use and limits of scientific names in biological informatics. In: Michel E (Ed.) Anchoring Biodiversity Information: From Sherborn to the 21st century and beyond. ZooKeys 550: 207–223. https://doi.org/10.3897/zookeys.550.9546

Figure 4 - Precision vs. recall in search results.

opencc-by-4.0Jan 2016View details →
zenodo24/100

Figure 1 from: Remsen D (2016) The use and limits of scientific names in biological informatics. In: Michel E (Ed.) Anchoring Biodiversity Information: From Sherborn to the 21st century and beyond. ZooKeys 550: 207–223. https://doi.org/10.3897/zookeys.550.9546

Figure 1 - Scientific names label information about species.

opencc-by-4.0Jan 2016View details →
zenodo24/100

Data set to train a natural language classifier able to differentiate between 15 topics relevant to biodiversity informatics

<p><strong>Scope and size</strong><br> This data set is used to train a natural language processing classifier. The classifier shall be able to differentiate between 15 topics relevant to biodiversity informatics. The list of relevant topics was adapted from Searls (2012).</p> <p>The data set was split into training data, testing data (for tweaking and unit-testing the classifier) and validation data. Each data set is stored as PDF files in a separate directory.</p> <ul> <li>Training data (5494 pages)</li> <li>Test data (977 pages)</li> <li>Validation data (215 pages)</li> </ul> <p><strong>Data sources and licenses</strong><br> Details about the licenses for each data set can be found in the corresponding directories.</p> <ul> <li>Training data was compiled from MIT OpenCourseWare resources provided by MIT under a Creative Commons BY-NC-SA License.</li> <li>Testing data was compiled from MIT OpenCourseWare exams, provided by MIT under a Creative Commons License BY-NC-SA.</li> <li>Validation data was compiled from Wikipedia, provided under a Creative Commons License by Wikipedia editors and contributors.</li> </ul> <p><strong>Topic references</strong><br> Each topic references one or more MIT OpenCourseWare courses:</p> <ul> <li>Algorithms (Demaine, and Devadas, 2011)</li> <li>Artificial Intelligence (Winston, 2010)</li> <li>Building Dynamic Websites (Abelson, and Greenspun, 2003)</li> <li>Computational Biology (Kellis, 2015)</li> <li>Computer Graphics (Matusik, and Durand, 2012)</li> <li>Computer Science and Programming (Bell, Grimson, and Guttag, 2016)</li> <li>Databases (Madden, Morris, Stonebraker, and Curino, 2010)</li> <li>Data Structures (Demaine, 2012)</li> <li>Digital Image Processing (Clifford, Fisher, Greenberg, and Wells, 2007; Golland, 2005)</li> <li>Machine Learning (Singh, Jaakkola, and Mohammad, 2006)</li> <li>Machine Structures (Morris, and Madden, 2009)</li> <li>Natural Language Processing (Berwick, 2003; Collins, and Barzilay, 2005)</li> <li>Parallel Computing (Edelman, 2011)</li> <li>Software Engineering (Jackson, and Devadas, 2005)</li> <li>Structure and Interpretation of Computer Programs (Miller, and Goldman, 2016).</li> </ul> <p><strong>References</strong></p> <p>Harold Abelson, and Philip Greenspun. 6.171 Software Engineering for Web Applications. Fall 2003. Massachusetts Institute of Technology: MIT OpenCourseWare, https://ocw.mit.edu. License: Creative Commons BY-NC-SA.</p> <p>Ana Bell, Eric Grimson, and John Guttag. 6.0001 Introduction to Computer Science and Programming in Python. Fall 2016. Massachusetts Institute of Technology: MIT OpenCourseWare, https://ocw.mit.edu. License: Creative Commons BY-NC-SA.</p> <p>Robert Berwick. 6.863J Natural Language and the Computer Representation of Knowledge. Spring 2003. Massachusetts Institute of Technology: MIT OpenCourseWare, https://ocw.mit.edu. License: Creative Commons BY-NC-SA.</p> <p>Gari Clifford, John Fisher, Julie Greenberg, and William Wells. HST.582J Biomedical Signal and Image Processing. Spring 2007. Massachusetts Institute of Technology: MIT OpenCourseWare, https://ocw.mit.edu. License: Creative Commons BY-NC-SA.</p> <p>Michael Collins, and Regina Barzilay. 6.864 Advanced Natural Language Processing. Fall 2005. Massachusetts Institute of Technology: MIT OpenCourseWare, https://ocw.mit.edu. License: Creative Commons BY-NC-SA.</p> <p>Erik Demaine. 6.851 Advanced Data Structures. Spring 2012. Massachusetts Institute of Technology: MIT OpenCourseWare, https://ocw.mit.edu. License: Creative Commons BY-NC-SA.</p> <p>Erik Demaine, and Srini Devadas. 6.006 Introduction to Algorithms. Fall 2011. Massachusetts Institute of Technology: MIT OpenCourseWare, https://ocw.mit.edu. License: Creative Commons BY-NC-SA.</p> <p>Alan Edelman. 18.337J Parallel Computing. Fall 2011. Massachusetts Institute of Technology: MIT OpenCourseWare, https://ocw.mit.edu. License: Creative Commons BY-NC-SA.</p> <p>Polina Golland. 6.881 Representation and Modeling for Image Analysis. Spring 2005. Massachusetts Institute of Technology: MIT OpenCourseWare, https://ocw.mit.edu. License: Creative Commons BY-NC-SA.</p> <p>Daniel Jackson, and Srini Devadas. 6.170 Laboratory in Software Engineering. Fall 2005. Massachusetts Institute of Technology: MIT OpenCourseWare, https://ocw.mit.edu. License: Creative Commons BY-NC-SA.</p> <p>Manolis Kellis. 6.047 Computational Biology. Fall 2015. Massachusetts Institute of Technology: MIT OpenCourseWare, https://ocw.mit.edu. License: Creative Commons BY-NC-SA.</p> <p>Robert Miller, and Max Goldman. 6.005 Software Construction. Spring 2016. Massachusetts Institute of Technology: MIT OpenCourseWare, https://ocw.mit.edu. License: Creative Commons BY-NC-SA.</p> <p>Samuel Madden, Robert Morris, Michael Stonebraker, and Carlo Curino. 6.830 Database Systems. Fall 2010. Massachusetts Institute of Technology: MIT OpenCourseWare, https://ocw.mit.edu. License: Creative Commons BY-NC-SA.</p> <p>Wojciech Matusik, and Fr&eacute;do Durand. 6.837 Computer Graphics. Fall 2012. Massachusetts Institute of Technology: MIT OpenCourseWare, https://ocw.mit.edu. License: Creative Commons BY-NC-SA.</p> <p>Morris, Robert, and Madden, Samuel. 6.033 Computer System Engineering. Spring 2009. Massachusetts Institute of Technology: MIT OpenCourseWare,&nbsp; http://hdl.handle.net/1721.1/118791. License: Creative Commons BY-NC-SA.</p> <p>David B. Searls. An online bioinformatics curriculum. 2012. PLoS computational biology, 8(9), p.e1002632.</p> <p>Rohit Singh, Tommi Jaakkola, and Ali Mohammad. 6.867 Machine Learning. Fall 2006. Massachusetts Institute of Technology: MIT OpenCourseWare, https://ocw.mit.edu. License: Creative Commons BY-NC-SA.</p> <p>Patrick Winston. 6.034 Artificial Intelligence. Fall 2010. Massachusetts Institute of Technology: MIT OpenCourseWare, https://ocw.mit.edu. License: Creative Commons BY-NC-SA.</p> <p>&nbsp;</p>

opencc-by-nc-sa-4.0Dec 2018View details →
zenodo20/100

Figure 3 in Optimizing biodiversity informatics to improve information flow, data quality, and utility for science and society

Figure 3. Examples of ways in which aggregators can make uncertainties and biases visually available to users of Primary Biodiversity Data. Such information can be employed to filter data and to quantify and correct for biases in sampling effort, respectively. (a) Georeferenced localities of a given species are simply plotted in geographic space (black dots; current practice). (b) Those same localities appear using symbologies that provide additional information; a hazy cloud indicates the radius of error for localities holding information regarding uncertainty of the georeference, and localities lacking such data appear only as hollow black circles. (c) Information appears that reflects the results of sampling effort, by showing in gray the georeferenced localities for all species belonging to a more inclusive target group (i.e., all species detected with the same techniques as the species of interest; conventions the same as in b). Note that the right-hand side of the study region lacks records for any species of the target group, suggestive of very low sampling effort there.

opennotspecifiedSep 2020View details →
zenodo20/100

Figure 2 in Optimizing biodiversity informatics to improve information flow, data quality, and utility for science and society

Figure 2. Use of individual and collective Stable Unique Identifiers (e.g., DOIs) in biodiversity informatics. (a) Individual Stable Unique Identifier (I-SUI) allows linking diverse data domains for a given organism. In this example, an I-SUI links the voucher specimen and associated Primary Biodiversity Data (e.g., date and locality) of an individual mammal to information regarding various aspects of molecular- to population-level biology. (b) Collective Stable Unique Identifier (C-SUI) denotes a set (i.e., a list) of individual identifiers. For example, a C-SUI could indicate the n individual records used in a given analysis.

opennotspecifiedSep 2020View details →
zenodo20/100

Figure 1 in Optimizing biodiversity informatics to improve information flow, data quality, and utility for science and society

Figure 1. Simplified overview of the interactions and flow of data among providers, aggregators, and users in biodiversity informatics. Numbers indicate the typical order of actions: 1. Aggregator receives data uploads (and periodic updates) from providers; 2. User makes a data query to aggregator's online portal; 3. Aggregator responds to query by making data available on portal (for viewing and/or download). Note that by querying a single aggregator, a user can receive data from multiple providers. Additionally, multiple intermediate aggregators typically exist, feeding into the largest ones most commonly consulted by users (e.g., GBIF).

opennotspecifiedSep 2020View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record