Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

45,411

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

45,411 results for “collection”

Learn how ShareScore rates datasets ↗
zenodo44/100

Islam West Africa Collection (IWAC)

<p>Directed by&nbsp;<a href="https://www.frederickmadore.com/" target="_blank" rel="noopener">Fr&eacute;d&eacute;rick Madore</a>, the&nbsp;<a href="https://islam.zmo.de/s/westafrica/" target="_blank" rel="noopener"><em>Islam West Africa Collection</em>&nbsp;(IWAC)</a> is a collaborative, open-access digital database that currently contains over 5,000 archival documents, newspaper articles, Islamic publications of various kinds, audio and video recordings, and photographs on Islam and Muslims in Burkina Faso, Benin, Niger, Nigeria, Togo and C&ocirc;te d'Ivoire. Most of the documents are in French, but some are also available in Hausa, Arabic, Dendi, and English. The site also indexes over 800 references to relevant books, book chapters, book reviews, journal articles, dissertations, theses, reports and blog posts. This project, hosted by the&nbsp;<a href="https://www.zmo.de/en" target="_blank" rel="noopener">Leibniz-Zentrum Moderner Orient (ZMO)</a>&nbsp;and funded by the Berlin Senate Department for Science, Health and Care, is a continuation of the award-winning&nbsp;<a href="https://web.archive.org/web/20231207083222/https://islam.domains.uflib.ufl.edu/s/bf/page/home" target="_blank" rel="noopener"><em>Islam Burkina Faso Collection</em></a>&nbsp;created in 2021 in collaboration with&nbsp;<a href="https://librarypress.domains.uflib.ufl.edu/" target="_blank" rel="noopener">LibraryPress@UF</a>.</p> <p>This dataset contains all the metadata of the items in the Collection, the Jupyter notebooks that were used to create the visualisations that showcase the possibilities of&nbsp;<a href="https://islam.zmo.de/s/westafrica/page/digital-humanities" target="_blank" rel="noopener">digital humanities</a> with the IWAC, and a copy of the spreadsheets that were used to create the&nbsp;<a href="https://islam.zmo.de/s/westafrica/page/exhibits" target="_blank" rel="noopener">digital exhibits</a>&nbsp;using&nbsp;<a href="https://timeline.knightlab.com/" target="_blank" rel="noopener">Timeline JS</a>.</p>

opencc-by-4.0Dec 2023View details →
zenodo44/100

Conductivity–Temperature–Depth (CTD) and dissolved oxygen profile data from shipboard surveys collected within Olympic Coast National Marine Sanctuary, 2005-2023

<p>This data set includes Conductivity-Temperature-Depth (CTD) and dissolved oxygen profile data that were collected along Washington State&rsquo;s outer coast within Olympic Coast National Marine Sanctuary towards the northernmost extent of the California Current System. Measurements were made at fourteen hydrographic stations during mooring deployment, recovery, and maintenance cruises between the months of May and October from 2005&ndash;2023. The 792 CTD profiles were acquired using Sea-Bird Scientific 19 SeaCAT or 19plus SeaCAT CTD profilers with associated SBE-43 (Sea-Bird Electronics) or Beckman or YSI-type (Yellow Springs Instruments) dissolved oxygen sensors. The data were processed via Sea-Bird Scientific&rsquo;s SBE Data Processing application using six of the modules in the following order: <em>Data Conversion, Filter, Align CTD, Loop Edit, Derive, and Bin Average</em>. These processing steps and associated methods are the same as those used to process CTD data that make up the&nbsp;<a href="../records/5814071">Newport Hydrographic Line time series</a> located off the central Oregon coast thus allowing for a direct comparison between the two regions.</p> <table> <tbody> <tr> <td><strong>Station Name &nbsp;&nbsp;</strong></td> <td><strong>Latitude</strong></td> <td><strong>Longitude</strong></td> <td><strong>Water Depth (m, MLLW)</strong></td> </tr> <tr> <td><strong>Makah Bay (MB)</strong></td> <td>&nbsp;</td> <td>&nbsp;</td> <td>&nbsp;</td> </tr> <tr> <td>MB015</td> <td>48.3254oN</td> <td>124.6768oW</td> <td>15</td> </tr> <tr> <td>MB042</td> <td>48.3240oN</td> <td>124.7354oW</td> <td>42</td> </tr> <tr> <td><strong>Cape Alava (CA)</strong></td> <td>&nbsp;</td> <td>&nbsp;</td> <td>&nbsp;</td> </tr> <tr> <td>CA015</td> <td>48.1663oN</td> <td>124.7568oW</td> <td>15</td> </tr> <tr> <td>CA042</td> <td>48.1660oN</td> <td>124.8234oW</td> <td>42</td> </tr> <tr> <td>CA065&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</td> <td>48.1659oN</td> <td>124.8949oW</td> <td>65</td> </tr> <tr> <td><strong>Teahwhit Head (TH)</strong></td> <td>&nbsp;</td> <td>&nbsp;</td> <td>&nbsp;</td> </tr> <tr> <td>TH015</td> <td>47.8761oN</td> <td>124.6195oW</td> <td>15</td> </tr> <tr> <td>TH042</td> <td>47.8762oN</td> <td>124.7334oW</td> <td>42</td> </tr> <tr> <td>TH065&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</td> <td>47.8767oN</td> <td>124.7967oW</td> <td>65</td> </tr> <tr> <td><strong>Kalaloch (KL)</strong></td> <td>&nbsp;</td> <td>&nbsp;</td> <td>&nbsp;</td> </tr> <tr> <td>KL015</td> <td>47.6008oN</td> <td>124.4284oW</td> <td>15</td> </tr> <tr> <td>KL027</td> <td>47.5946oN</td> <td>124.4971oW</td> <td>27</td> </tr> <tr> <td>KL050&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</td> <td>47.5933oN</td> <td>124.6112oW</td> <td>50</td> </tr> <tr> <td><strong>Cape Elizabeth (CE)</strong></td> <td>&nbsp;</td> <td>&nbsp;</td> <td>&nbsp;</td> </tr> <tr> <td>CE015</td> <td>47.3568oN</td> <td>124.3481oW</td> <td>15</td> </tr> <tr> <td>CE042</td> <td>47.3531oN</td> <td>124.4887oW</td> <td>42</td> </tr> <tr> <td>CE065&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</td> <td>&nbsp;47.3528oN</td> <td>124.5669oW</td> <td>65</td> </tr> </tbody> </table>

opencc-by-4.0Jan 2024View details →
zenodo44/100

Classical Philology Syndication Feed Collection

<p>This feed collection allows gathering information on recent publications, scientific books and journals, in the field of Classical Philology and related disciplines. The OPML file (https://en.wikipedia.org/wiki/OPML) can be imported in programs with news aggregation functionality, f.e. Thunderbird.</p>

opencc-zeroNov 2023View details →
zenodo44/100

Collective Strong Coupling Modifies Aggregation and Solvation - Dataset

<p>Dataset to complement "Collective Strong Coupling Modifies Aggregation and Solvation" - includes output and cube files obtained using the <a href="https://etprogram.org/">eT program</a>, an open source electronic (and molecular-polaritonic) structure program.</p> <p>See the paper at <a title="DOI URL" href="https://doi.org/10.1021/acs.jpclett.3c03506">https://doi.org/10.1021/acs.jpclett.3c03506</a></p>

opencc-by-4.0Jan 2024View details →
zenodo44/100

RAYUELA - Open Data - Data collected through a serious game created to identify patterns and profiles of young potential victims/perpetrators of cybercrimes.

<p>The data of this dataset have been collected in the pilots carried out by the RAYUELA project in different countries of the European Union. The participants are minors and the game sessions have been carried out in schools and summer camps in a supervised way.</p>

opencc-by-4.0Sep 2023View details →
zenodo44/100

Leaf samples of three common plant species collected in seven LandKlif quadrants

<p><span>Leaf samples of Acer pseudoplatanus, Dactylis glomerata and Potentilla reptans were collected in seven LandKlif quadrants along a climate gradient in summer 2020. In each quadrant, leaves were sampled in two habitats (forest and open landscape). In each habitat, three leaves from seven individuals of each species were collected. Specific leaf area (SLA) and leaf dry matter content (LDMC) of each leaf sample were determined in the lab. In addition, nitrogen content was measured at the level of individuals. This dataset contains the mean SLA and LDMC of the leaves sampled from each individual, as well as information on the site where they were collected.</span></p> <p><span>LandKlif is funded by the Bavarian State Ministry of Science and the Arts within the Bavarian Climate Research Network (bayklif). &nbsp;Within the five year funding period of bayklif, five interdisciplinary senior research associations and five junior research groups are be financed with a total sum of 18 million Euro. LandKliF, as one of the five interdisciplinary senior research associations, addresses the effects of climate change on biodiversity and ecosystem services in semi-natural, agricultural and urban landscapes.</span></p>

opencc-by-4.0Mar 2024View details →
zenodo44/100

A collection of annotated soundscape recordings from western Kenya

<p>This collection contains 35 soundscape recordings of 32 hours total duration, which have been annotated with 10,294 labels for 176 different bird species from western Kenya. The data were recorded in 2021 and 2022 west and southwest of Lake Baringo in Baringo County, Kenya. This collection has partially been featured as test data in the 2023 BirdCLEF competition and can primarily be used for training and evaluation of machine learning algorithms.</p> <p><strong>Data collection</strong></p> <p>For this collection, AudioMoths and SWIFT recording units were deployed at multiple locations west and southwest of Lake Baringo, Baringo County, Kenya between Dezember 2021 and February 2022. Recording locations cover a variety of habitats from open grasslands to semi-arid scrubland and mountain forests. Recordings were originally sampled at 48 kHz and converted to MP3 for faster file transfer. For publication, all files were resampled to 32 kHz and converted to FLAC.</p> <p><strong>Sampling and annotation protocol</strong></p> <p>A total of 32 hours of audio from various sites west and southwest of Lake Baringo were selected for annotation. Annotators were tasked with identifying and labeling each bird call they could discern, excluding any calls that were too weak or indiscernible. The annotation process was carried out using Audacity. Provided labels mark the center of each bird call. In this collection, we use eBird species codes as labels, following the 2021 eBird taxonomy (Clements list). Parts of this dataset have previously been used in the 2023 BirdCLEF competition.&nbsp;</p> <p><strong>Files in this collection</strong></p> <p>Audio recordings can be accessed by downloading and extracting the &ldquo;soundscape_data.zip&rdquo; file. Soundscape recording filenames contain a sequential file ID, recording date and timestamp in EAT (UTC+3). As an example, the file &ldquo;KEN_001_20211207_153852.flac&rdquo; has sequential ID 001 and was recorded on December 7th 2021 at 15:38:52 EAT. Ground truth annotations are listed in &ldquo;annotations.csv&rdquo; where each line specifies the corresponding filename, start and end time in seconds, and an eBird species code. These species codes can be assigned to scientific and common name of a species with the &ldquo;species.csv&rdquo; file. The approximate recording location with longitude and latitude can be found in the &ldquo;recording_location.txt&rdquo; file.</p> <p><strong>Acknowledgements</strong></p> <p>Compiling this extensive dataset was a major undertaking, and we are very thankful to the domain experts who helped to collect and manually annotate the data for this collection. In particular, our thanks go to Francis Cherutich for setting up recording units, collecting and annotating data, and to Alain Jacot for assisting in programming the units and transporting the recorders to Kenya.</p>

opencc-by-4.0Apr 2024View details →
zenodo44/100

A Dataset from Collected Events through the INCENTIVE Community Insider

<p>This dataset contains the collected events related to Citizen Science and public engagement, harvested from trusted online sources worldwide and provided by the&nbsp;<a href="https://toolkit.incentive-project.eu/community-insider/news-events">Community Insider</a> of the INCENTIVE Digital Toolkit.</p>

opencc-by-4.0Jan 2024View details →
zenodo44/100

MATEdb2, a Collection of High-Quality Metazoan Proteomes across the Animal Tree of Life to Speed Up Phylogenomic Studies

<p>Recent advances in high-throughput sequencing have exponentially increased the number of genomic data available for animals (Metazoa) in the last decades, with high-quality chromosome-level genomes being published almost daily. Nevertheless, generating a new genome is not an easy task due to the high cost of genome sequencing, the high complexity of assembly, and the lack of standardized protocols for genome annotation. The lack of consensus in the annotation and publication of genome files hinders research by making researchers lose time in reformatting the files for their purposes but can also reduce the quality of the genetic repertoire for an evolutionary study. Thus, the use of transcriptomes obtained using the same pipeline as a proxy for the genetic content of species remains a valuable resource that is easier to obtain, cheaper, and more comparable than genomes. In a previous study, we presented the Metazoan Assemblies from Transcriptomic Ensembles database (MATEdb), a repository of high-quality transcriptomic and genomic data for the two most diverse animal phyla, Arthropoda and Mollusca. Here, we present the newest version of MATEdb (MATEdb2) that overcomes some of the previous limitations of our database: (i) we include data from all animal phyla where public data are available, and (ii) we provide gene annotations extracted from the original GFF genome files using the same pipeline. In total, we provide proteomes inferred from high-quality transcriptomic or genomic data for almost 1,000 animal species, including the longest isoforms, all isoforms, and functional annotation based on sequence homology and protein language models, as well as the embedding representations of the sequences. We believe this new version of MATEdb will accelerate research on animal phylogenomics while saving thousands of hours of computational work in a plea for open, greener, and collaborative science.</p>

opencc-by-4.0Nov 2024View details →
zenodo44/100

UC Santa Barbara Invertebrate Zoology Collection (UCSB-IZC) Data Archive and Biodiversity Dataset Graph hash://md5/10663911550bb52a0f5741993f82db9d hash://sha256/80c0f5fc598be1446d23c95141e87880c9e53773cb2e0b5b54cb57a8ea00b20c

<p>A biodiversity dataset graph: UCSB-IZC</p> <p>The intended use of this archive is to facilitate (meta-)analysis of the UC Santa Barbara Invertebrate Zoology Collection (UCSB-IZC). UCSB-IZC is a natural history collection of invertebrate zoology at Cheadle Center of Biodiversity and Ecological Restoration, University of California Santa Barbara.</p> <p>This dataset provides versioned snapshots of the UCSB-IZC network as tracked by Preston [2,3] between 2021-10-08 and 2021-11-04 using [preston track &quot;https://api.gbif.org/v1/occurrence/search/?datasetKey=d6097f75-f99e-4c2a-b8a5-b0fc213ecbd0&quot;].</p> <p>This archive contains 14349 images related to 32533 occurrence/specimen records. See included sample-image.jpg and their associated meta-data sample-image.json [4].</p> <p>The images were counted using:</p> <p>$ preston cat hash://sha256/80c0f5fc598be1446d23c95141e87880c9e53773cb2e0b5b54cb57a8ea00b20c\<br> &nbsp;| grep -o -P &quot;.*depict&quot;\<br> &nbsp;| sort\<br> &nbsp;| uniq\<br> &nbsp;| wc -l</p> <p>And the occurrences were counted using:</p> <p>$ preston cat hash://sha256/80c0f5fc598be1446d23c95141e87880c9e53773cb2e0b5b54cb57a8ea00b20c\<br> &nbsp;| grep -o -P &quot;occurrence/([0-9])+&quot;\<br> &nbsp;| sort\<br> &nbsp;| uniq\<br> &nbsp;| wc -l</p> <p>The archive consists of 256 individual parts (e.g., preston-00.tar.gz, preston-01.tar.gz, ...) to allow for parallel file downloads. The archive contains three types of files: index files, provenance files and data files. Only two index and provenance files are included and have been individually included in this dataset publication. Index files provide a way to links provenance files in time to establish a versioning mechanism.</p> <p>To retrieve and verify the downloaded UCSB-IZC biodiversity dataset graph, first download preston-*.tar.gz. Then, extract the archives into a &quot;data&quot; folder. Alternatively, you can use the Preston [2,3] command-line tool to &quot;clone&quot; this dataset using:</p> <p>$ java -jar preston.jar clone --remote https://archive.org/download/preston-ucsb-izc/data.zip/,https://zenodo.org/record/5557670/files,https://zenodo.org/record/5660088/files/</p> <p>After that, verify the index of the archive by reproducing the following provenance log history:</p> <p>$ java -jar preston.jar history<br> &lt;urn:uuid:0659a54f-b713-4f86-a917-5be166a14110&gt; &lt;http://purl.org/pav/hasVersion&gt; &lt;hash://sha256/d5eb492d3e0304afadcc85f968de1e23042479ad670a5819cee00f2c2c277f36&gt; .<br> &lt;hash://sha256/80c0f5fc598be1446d23c95141e87880c9e53773cb2e0b5b54cb57a8ea00b20c&gt; &lt;http://purl.org/pav/previousVersion&gt; &lt;hash://sha256/d5eb492d3e0304afadcc85f968de1e23042479ad670a5819cee00f2c2c277f36&gt; .</p> <p>To check the integrity of the extracted archive, confirm that each line produce by the command &quot;preston verify&quot; produces lines as shown below, with each line including &quot;CONTENT_PRESENT_VALID_HASH&quot;. Depending on hardware capacity, this may take a while.</p> <p>$ java -jar preston.jar verify<br> hash://sha256/ce1dc2468dfb1706a6f972f11b5489dc635bdcf9c9fd62a942af14898c488b2c&nbsp;&nbsp;&nbsp; file:/home/jhpoelen/ucsb-izc/data/ce/1d/ce1dc2468dfb1706a6f972f11b5489dc635bdcf9c9fd62a942af14898c488b2c&nbsp;&nbsp;&nbsp; OK&nbsp;&nbsp;&nbsp; CONTENT_PRESENT_VALID_HASH&nbsp;&nbsp;&nbsp; 66438&nbsp;&nbsp;&nbsp; hash://sha256/ce1dc2468dfb1706a6f972f11b5489dc635bdcf9c9fd62a942af14898c488b2c<br> hash://sha256/f68d489a9275cb9d1249767244b594c09ab23fd00b82374cb5877cabaa4d0844&nbsp;&nbsp;&nbsp; file:/home/jhpoelen/ucsb-izc/data/f6/8d/f68d489a9275cb9d1249767244b594c09ab23fd00b82374cb5877cabaa4d0844&nbsp;&nbsp;&nbsp; OK&nbsp;&nbsp;&nbsp; CONTENT_PRESENT_VALID_HASH&nbsp;&nbsp;&nbsp; 4093&nbsp;&nbsp;&nbsp; hash://sha256/f68d489a9275cb9d1249767244b594c09ab23fd00b82374cb5877cabaa4d0844<br> hash://sha256/3e70b7adc1a342e5551b598d732c20b96a0102bb1e7f42cfc2ae8a2c4227edef&nbsp;&nbsp;&nbsp; file:/home/jhpoelen/ucsb-izc/data/3e/70/3e70b7adc1a342e5551b598d732c20b96a0102bb1e7f42cfc2ae8a2c4227edef&nbsp;&nbsp;&nbsp; OK&nbsp;&nbsp;&nbsp; CONTENT_PRESENT_VALID_HASH&nbsp;&nbsp;&nbsp; 5746&nbsp;&nbsp;&nbsp; hash://sha256/3e70b7adc1a342e5551b598d732c20b96a0102bb1e7f42cfc2ae8a2c4227edef<br> hash://sha256/995806159ae2fdffdc35eef2a7eccf362cb663522c308aa6aa52e2faca8bb25b&nbsp;&nbsp;&nbsp; file:/home/jhpoelen/ucsb-izc/data/99/58/995806159ae2fdffdc35eef2a7eccf362cb663522c308aa6aa52e2faca8bb25b&nbsp;&nbsp;&nbsp; OK&nbsp;&nbsp;&nbsp; CONTENT_PRESENT_VALID_HASH&nbsp;&nbsp;&nbsp; 6147&nbsp;&nbsp;&nbsp; hash://sha256/995806159ae2fdffdc35eef2a7eccf362cb663522c308aa6aa52e2faca8bb25b</p> <p>Note that a copy of the java program &quot;preston&quot;, preston.jar, is included in this publication. The program runs on java 8+ virtual machine using &quot;java -jar preston.jar&quot;, or in short &quot;preston&quot;.</p> <p>Files in this data publication:</p> <p>--- start of file descriptions ---</p> <p>-- description of archive and its contents (this file) --<br> README</p> <p>-- executable java jar containing preston [2,3] v0.3.1. --<br> preston.jar</p> <p>-- preston archive containing UCSB-IZC (meta-)data/image files, associated provenance logs and a provenance index --<br> preston-[00-ff].tar.gz</p> <p>-- individual provenance index files --<br> 2a5de79372318317a382ea9a2cef069780b852b01210ef59e06b640a3539cb5a</p> <p>-- example image and meta-data --<br> sample-image.jpg (with hash://sha256/916ba5dc6ad37a3c16634e1a0e3d2a09969f2527bb207220e3dbdbcf4d6b810c)<br> sample-image.json (with hash://sha256/f68d489a9275cb9d1249767244b594c09ab23fd00b82374cb5877cabaa4d0844)</p> <p>--- end of file descriptions ---</p> <p><br> References</p> <p>[1] Cheadle Center for Biodiversity and Ecological Restoration (2021). University of California Santa Barbara Invertebrate Zoology Collection. Occurrence dataset https://doi.org/10.15468/w6hvhv accessed via GBIF.org on 2021-11-04 as indexed by the Global Biodiversity Informatics Facility (GBIF) with provenance hash://sha256/d5eb492d3e0304afadcc85f968de1e23042479ad670a5819cee00f2c2c277f36 hash://sha256/80c0f5fc598be1446d23c95141e87880c9e53773cb2e0b5b54cb57a8ea00b20c.<br> [2] https://preston.guoda.bio, https://doi.org/10.5281/zenodo.1410543 .<br> [3] MJ Elliott, JH Poelen, JAB Fortes (2020). Toward Reliable Biodiversity Dataset References. Ecological Informatics. https://doi.org/10.1016/j.ecoinf.2020.101132<br> [4] Cheadle Center for Biodiversity and Ecological Restoration (2021). University of California Santa Barbara Invertebrate Zoology Collection. Occurrence dataset https://doi.org/10.15468/w6hvhv accessed via GBIF.org on 2021-10-08. https://www.gbif.org/occurrence/3323647301 . hash://sha256/f68d489a9275cb9d1249767244b594c09ab23fd00b82374cb5877cabaa4d0844 hash://sha256/916ba5dc6ad37a3c16634e1a0e3d2a09969f2527bb207220e3dbdbcf4d6b810c</p>

opencc-zeroNov 2021View details →
zenodo44/100

Fatiando a Terra data v1.0.0: A curated collection of open geophysics data for tutorials and documentation

<p>This repository holds curated sample datasets that can be used in the documentation and tutorials of the <a href="https://www.fatiando.org/">Fatiando a Terra</a> project. All datasets are cleaned and formatted versions of openly available data under permissive licenses or in the public domain.</p> <p>More information about datasets and the code for cleaning, formatting, and preprocessing the data can be found at: <a href="https://github.com/fatiando/data">https://github.com/fatiando/data</a></p> <p>See the README.md file for information on data sources and their original licenses.</p> <p><strong>NOTE:</strong> This collection uses <a href="https://semver.org/">semantic versioning</a> (i.e., MAJOR.MINOR.BUGFIX). Major releases mean that backwards incompatible changes were made to the data. Minor releases add new data without changing existing files. Bug fix releases fix errors in a previous release that makes the data unusable. Changes to the current data files will always be published as a major release unless the file(s) in the previous release was unusable/corrupted.</p>

opencc-by-4.0Nov 2021View details →
zenodo44/100

Cambridge butterfly wing collection - Ecuador, August 2019

<p>EN: This upload contains photographs taken by Annalie Barker and Joana Meier at&nbsp;the University of Cambridge, from a butterfly wing collection from Ecuador (August 2019), in collaboration with Caroline Bacquet (IKIAM). Individual sample names can be found in the information sheet. Further Information on individual samples from the Butterfly Genetics Group Collection can be found on the public database Earthcape (<a href="https://heliconius.ecdb.io/">click here for the database</a>, and <a href="https://heliconius.zoo.cam.ac.uk/databases/earthcape-specimen-database/">here for FAQ</a>). &nbsp;Please contact Joana Meier (jm2276[at]cam.ac.uk) or Chris Jiggins (c.jiggins[at]zoo.cam.ac.uk) for further information.</p> <p>&nbsp;</p> <p>ES: Este repositorio contiene fotograf&iacute;as tomadas por Annalie Barker y Joana Meier en la Universidad de Cambridge, de mariposas de Ecuador (Agosto 2019), en colaboraci&oacute;n con Caroline Bacquet (IKIAM, Ecuador). Puede encontrar informaci&oacute;n sobre muestras individuales de Butterfly Genetics Group Collection en la base de datos p&uacute;blica Earthcape (<a href="https://heliconius.ecdb.io/">haga clic aqu&iacute; para la base de datos</a>, y <a href="https://heliconius.zoo.cam.ac.uk/databases/earthcape-specimen-database/">aqu&iacute; para preguntas frecuentes</a>) Por favor, p&oacute;ngase en contacto con Joana Meier (jm2276 [arroba] cam.ac.uk) or Chris Jiggins (c.jiggins [arroba] zoo.cam.ac.uk) con sus preguntas o peticiones.</p>

opencc-by-4.0Nov 2021View details →
zenodo44/100

Can Artificial Intelligence help in the study of vegetative growth dynamics from herbarium collections? An evaluation of the tropical flora of the French Guiana forest

<p>Dataset was used for the article &quot;Can Artificial Intelligence help in the study of vegetative growth dynamics from herbarium collections? An evaluation of the tropical flora of the French Guiana forest&quot;.</p> <p>The related work proposes to study to what extent the use of automated visual analysis techniques, based on deep learning, can help not only to detect relatively rare vegetative structures in herbarium collections but also to automatically classify them by type of growing shoot (continuous or rhythmic).</p> <p>Abstract of the paper:</p> <p>A better knowledge of tree vegetative growth patterns and their relationship to environmental variables is crucial in understanding forest growth dynamics and how climate change may affect them. Generally less studied than reproductive structures, the phenology of tree vegetative growth mainly focuses on the analysis of growing shoots, from vegetative buds development to leaf fall. This growth process usually strongly differs between temperate and tropical regions. In temperate regions, this pattern is quite well known. Low winter temperatures impose a stop of the vegetative growth shoots and lead to the typical expression of an annual growth cycle for the vast majority of tree species. In moist tropical regions, on the other hand, the seasonality is much less marked. In addition, these regions contain a much wider variety of tree species. These two aspects lead to a tremendous diversity of phenological patterns that are still poorly known and understood. In particular, not much is known on the periodicity and timing of growth at individual trees, population, or community levels.</p> <p>The work carried out in this study aims to advance knowledge in this area, focusing more particularly on herbarium scans, as herbarium collections offer the promise of monitoring plant phenology over long time periods. However, such a study requires the ability to detect a sufficiently large number of growing shoots in herbarium collections to draw statistically relevant conclusions, which can be very costly if the work is done manually. Furthermore, herbarium collections traditionally focus on reproductive organs, and herbarium specimens showing growing shoots are pretty rare.</p> <p>We propose in this paper to study to what extent the use of automated visual analysis techniques, based on deep learning, can help not only to detect these relatively rare vegetative structures in herbarium collections but also to automatically classify them by type of growing shoot (continuous or rhythmic). Our results show the relevance of using herbarium data for vegetative phenology research, as well as the potential of deep learning approaches for growth shoot detection.</p>

opencc-by-4.0Dec 2021View details →
zenodo44/100

Antarctic Circumnavigation Expedition sample log: samples collected in the Southern Ocean during the austral summer of 2016/17.

<p><strong>Dataset abstract</strong></p> <p>The Antarctic Circumnavigation Expedition (ACE) spent 90 days circumnavigating Antarctica on the R/V Akademik Tryoshnikov during the austral summer of 2016/17. This dataset provides a record of the samples that were collected during the expedition.</p> <p><strong>Dataset contents</strong></p> <ul> <li>ace_sample_log.csv, data file, comma-separated values</li> <li>README.txt, metadata, text</li> <li>data_file_header.txt, metadata, text</li> </ul> <p><strong>Dataset license</strong></p> <p>This sample log is made available under a Creative Commons Attribution 4.0 International License (CC BY 4.0) whose full text can be found at https://creativecommons.org/licenses/by/4.0/</p>

opencc-by-4.0Dec 2019View details →
zenodo44/100

Spatially gridded cross-shelf hydrographic sections and monthly climatologies from shipboard survey data collected along the Newport Hydrographic Line, 1997-2021

<p>This data set, described in detail in <a href="https://www.sciencedirect.com/science/article/pii/S2352340922001342">Risien et al. (2022)</a>, contains Newport Hydrographic Line station data; gridded, cross-shelf hydrographic sections; and derived monthly climatologies for temperature, practical salinity, potential density, spiciness, and dissolved oxygen. It consists of CSV (Comma Separated Values) files (<em>newport_hydrographic_line_station_data</em><em>.</em><em>zip</em>) that contain CTD observations collected at the seven hydrographic stations located 1, 3, 5, 10, 15, 20 and 25 nautical miles west of Newport, Oregon between March 1997 and July 2021. Additionally, the data set contains three NetCDF files that follow CF (Climate and Forecast) metadata conventions: <em>newport_hydrographic_line_gridded_sections</em><em>.nc</em> contains observations gridded to a 0.01<sup>o</sup> x 1 dbar longitude - pressure grid to create cross-shelf hydrographic sections for each of the five variables for each cruise. <em>newport_hydrographic_line_gridded_section_climatologies</em><em>.nc</em> contains climatological hydrographic sections, calculated using harmonic analysis over the 24-year period March 1997 to February 2021 and reported here for the middle of each month, and <em>newport_hydrographic_line_gridded_section_coefficients.nc</em> contains the associated linear regression model coefficients for all five variables. From the regression coefficients, users can construct seasonal cycles at any location in the gridded section with a temporal resolution that best suits their specific needs. Finally, this data set includes example MATLAB and R scripts that show how to read the data files, plot&nbsp;cross-shelf hydrographic sections, and calculate daily and monthly&nbsp;climatologies using the&nbsp;regression coefficients.</p>

opencc-by-4.0Jan 2022View details →
zenodo44/100

Collected metadata for masses in the early Cappella Sistina choirbooks

<p>This dataset contains a collection of metadata about the early Cappella Sistina choirbooks, once belonging to the Papal Chapel. It presents the mass compositions contained in the manuscript sources V-CVbav MS 14, 23, 35, 41, 49, 51, 63, 197, 64 and concordant sources of these.</p> <p>The data has been collected in the year 2015 and hasn&#39;t been double-checked by any other person. Even though the data was collected as conscientiously as possible, it is sometimes incomplete and errors are also possible. The data as well as the data structure has been collected mostly in German. References for this metadata collection can be found under <code>./references</code>.</p> <p>This data has been used already in:</p> <ul> <li>Plaksin, Anna Viktoria Katrin: &quot;Modelle zur computergest&uuml;tzten Analyse von &Uuml;berlieferungen der Mensuralmusik. Emprische Textforschung im Kontext phylogenetischer Verfahren.&quot;, M&uuml;nster, 2021 (Schriften zur Musikwissenschaft aus M&uuml;nster 27), online: <a href="http://nbn-resolving.de/urn:nbn:de:hbz:6-59029717067">urn:nbn🇩🇪hbz:6-59029717067</a>, DOI: <a href="https://doi.org/10.26083/tuprints-00017211">10.26083/tuprints-00017211</a></li> </ul> <p>The repo for this dataset can be found here: <a href="https://github.com/annplaksin/earlyCappSistMasses">https://github.com/annplaksin/earlyCappSistMasses</a></p> <p>This release contains the slightly updated version that has been used in <a href="https://doi.org/10.26083/tuprints-00017211">10.26083/tuprints-00017211</a>:</p> <ul> <li>SQL dump has been changed from MySQL to SQLite dialect</li> <li>Views have been updated</li> <li>PKs have been updated to ensure database-wide uniqueness (in a very basic way).</li> </ul> <p><strong>Contents</strong></p> <ul> </ul> <p>The SQLite database set is found under <code>./sql</code>. It contains not only the tables but various views for a better overview. For opening and viewing the sql data, tools like e.g. the <a href="https://sqlitebrowser.org/">DB Browser for SQLite</a> can be used.</p> <p>The <code>./csv</code> folder contains csv exports of the tables and views.</p> <p>In <code>./networkData</code> contains the relations of works and sources as bipartite network in <code>.csv</code> format. The edges are available as edge list or as adjacency matrix. The node list contains very basic information necessary for identification. Please note, this is the raw exported data for analysis.</p>

opencc-by-4.0Feb 2022View details →
zenodo44/100

A database of physical therapy exercises with variability of execution collected by wearable sensors

<p>The PHYTMO database contains data from physical therapy exercises and gait variations recorded with magneto-inertial sensors, including information from an optical reference system. PHYTMO includes the recording of 30 volunteers, aged between 20 and 70 years old. A total amount of 6 exercises and 3 gait variations commonly prescribed in physical therapies were recorded. The volunteers performed two series with a minimum of 8 repetitions in each one. Four magneto-inertial sensors were placed on the lower-or upper-limbs for the recording of the motions together with passive optical reflectors.&nbsp;The files include&nbsp;the specifications of the inertial sensors and the cameras. The database includes magneto-inertial data (linear acceleration, turn rate and magnetic field), together with a highly accurate location and orientation in the 3D space provided by the optical system (errors are lower than 1mm). The database files were stored in CSV format to ensure usability with common data processing software. The main aim of this dataset is the availability of inertial data for two main purposes: the analysis of different techniques for the identification and evaluation of exercises monitored with inertial wearable sensors and the validation of inertial sensor-based algorithms for human motion monitoring that obtains segments orientation in the 3D space. Furthermore, the database stores enough data to train and evaluate Machine Learning-based algorithms. The age range of the participants can be useful for establishing age-based metrics for the exercises evaluation or the study of differences in motions between different aged groups. Finally, the MATLAB function <em>features_extraction</em>, developed by the authors, is also given.&nbsp;This function splits signals using a sliding window, returning its segments, and extract signal features, in the time and frequency domains, based on prior studies of the literature.</p>

opencc-by-4.0Feb 2022View details →
zenodo44/100

Single-crystal X-ray diffractometry data for a sample of [Cu(HF₂)(pyrazine)₂]PF₆ collected on beamline I19-2 at Diamond Light Source

<p>Single-crystal X-ray diffractometry data for a sample of [Cu(HF₂)(pyrazine)₂]PF₆.</p> <p>These data were collected at Diamond Light Source, on&nbsp;beamline I19 (experiments hutch 2), on 2022-01-30, and are particularly useful for testing data reduction routines. They are known to produce good merging statistics and final structure refinement.</p> <p>The sample was prepared as follows:<br> Ammonium hexafluorophosphate (NH₄PF₆) (0.310&nbsp;g, 1.9&nbsp;mmol), ammonium hydrogen difluoride ((NH₄)HF₂) (0.109&nbsp;g, 1.9&nbsp;mmol) and pyrazine (C₄H₄N₂) (0.300&nbsp;g, 3.7&nbsp;mmol) were dissolved in 5&nbsp;mL of deionised water. The obtained colourless solution was slowly added to a blue solution of copper(II) nitrate prepared by dissolving copper(II) nitrate hemipentahydrate (Cu(NO₃)₂&nbsp;&middot;&nbsp;2.5(H₂O)) (0.425&nbsp;g, 1.8&nbsp;mmol) in 5&nbsp;mL of deionised water. The solutions were mixed in a plastic beaker at room temperature. The formation of blue crystals of [Cu(HF₂)(pyrazine)₂]PF₆ on the side of the beaker started after few seconds and continued for about 24&nbsp;hours during which the sealed beaker was not moved.</p> <p>The sample was measured at room temperature and the illuminating beam had a wavelength of 0.4859 &Aring; (25.52 keV).</p> <p>Beamline I19-2 at Diamond Light Source, a four-circle &kappa;-geometry diffractometer (see <a href="https://onlinelibrary.wiley.com/doi/10.1107/97809553602060000936">[Kern 2019]</a>) with an undulator source, is described in <a href="https://doi.org/10.1107/S0909049512008801">[Nowell 2012]</a> but has since been upgraded to use a Dectris Eiger2&nbsp;X&nbsp;4M CdTe hybrid photon counting detector. The data are written in the <a href="https://manual.nexusformat.org/classes/applications/NXmx.html">NXmx variant</a> of the <a href="https://www.nexusformat.org/">NeXus format</a>, and so include metadata with a functionally complete description of the diffractometer.</p> <p>Inventory of data:</p> <ul> <li><strong><code>01_CuHF2pyz2PF6b_Phi.tar.xz</code></strong><br> A single 1750-image 350&deg; &phi; rotation scan from -175&deg; to 175&deg; with 0.2&deg; rotation per image, an exposure time of 0.1&nbsp;s per image, &omega;&nbsp;=&nbsp;-90&deg;, &kappa;&nbsp;=&nbsp;0&deg; and 2&theta;&nbsp;=&nbsp;0&deg;.</li> <li><strong><code>02_CuHF2pyz2PF6b_2T.tar.xz</code></strong><br> A single 1750-image 350&deg; &phi; rotation scan from -175&deg; to 175&deg; with 0.2&deg; rotation per image, an exposure time of 0.1&nbsp;s per image, &omega;&nbsp;=&nbsp;-90&deg;, &kappa;&nbsp;=&nbsp;0&deg; and 2&theta;&nbsp;=&nbsp;20&deg;.</li> <li><strong><code>03_CuHF2pyz2PF6b_P_O.tar.xz</code></strong><br> Two sequential rotation scans: <ul> <li><strong><code>CuHF2pyz2PF6b_P_O_01.nxs</code></strong><br> A 1750-image 350&deg; &phi; scan from -175&deg; to 175&deg; with &omega;&nbsp;=&nbsp;-90&deg;, &kappa;&nbsp;=&nbsp;0&deg; and 2&theta;&nbsp;=&nbsp;0&deg;.</li> <li><strong><code>CuHF2pyz2PF6b_P_O_02.nxs</code></strong><br> A 600-image 120&deg; &omega; scan from -125&deg; to -5&deg; with &phi;&nbsp;=&nbsp;-90&deg;, &kappa;&nbsp;=&nbsp;45&deg; and 2&theta;&nbsp;=&nbsp;0&deg;.</li> </ul> Both scans had 0.2&deg; rotation per image and an exposure time of 0.1&nbsp;s per image.</li> </ul> <p>The same sample was used for all these measurements. Throughout, the sample-to-detector distance was 85&nbsp;mm and the beam was attenuated to 0.2% of its full intensity.</p> <p>For each rotation scan, the data comprise a single top-level NXmx-format NeXus file named <code>&lt;filename&gt;.nxs</code>, one or more image files named <code>&lt;filename&gt;_00000n.h5</code>, where <code>n</code> is a numeral, and a single detector metadata file named <code>&lt;filename&gt;_meta.h5</code>. The NeXus file contains an HDF5 virtual data set that links to the data in the image file(s), and several HDF5 external links to data in the detector metadata file.</p> <p>For internal reference of Diamond Light Source staff, these data were collected as part of commissioning visit CM31144-1. Some file names and corresponding HDF5 link targets have been altered from their original names for consistency with the file contents.</p>

opencc-by-4.0Mar 2022View details →
zenodo44/100

Manually curated transcriptomics data collection for toxicogenomic assessment of engineered nanomaterials

<p>Toxicogenomics (TGx) approaches are increasingly applied to gain insight into the possible toxicity mechanisms of engineered nanomaterials (ENMs). Omics data can be valuable to elucidate the mechanism of action of chemicals and develop predictive models in toxicology. While vast amounts of transcriptomics data from ENM exposures have already been accumulated, a unified, easily accessible and reusable collection of transcriptomics data for ENMs is currently lacking. In an attempt to improve the FAIRness of already existing transcriptomics data for nanomaterials, we curated a collection of homogenized transcriptomics data from human, mouse and rat ENM exposures <em>in vitro</em> and <em>in vivo</em>.</p>

opencc-by-4.0Jul 2020View details →
zenodo44/100

Collective Variable for Metadynamics Derived from AlphaFold Output

<p>AlphaFold is the state of the art method for prediction of 3D structures of proteins from the amino acid sequence by neural networks. One of the outputs of AlphaFold is a probability profile of inter-residue distances for all residue pairs. We used this profile to evaluate any conformation of the studied protein to express its compliance with the AlphaFold prediction. This value can be used as a collective variable in metadynamics or parallel tempering metadynamics to accelerate protein folding in a molecular simulation. We applied this approach on folding of mini-proteins Trp-cage and beta hairpin. See V. Spiwok, M. Krečka &amp; A. Křenek: <a href="http://doi.org/10.3389/fmolb.2022.878133">Collective Variable for Metadynamics Derived from AlphaFold Output</a> <em>Frontiers in Molecular Biosciences</em> <strong>9</strong> 878133 (2022) DOI: 10.3389/fmolb.2022.878133.</p>

opencc-by-4.0Feb 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record