Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
40,091
datasets available to search
ShareScore release 0.9.0
Dataset results
40,091 results for “recordings”
Migration Drivers Data Inventory Records
<p>This inventory includes metadata on various quantitative sources of information on migration drivers that can be used for modelling purposes. Additionally, the inventory includes information on articles that used those quantitative sources such as the statistical effect found in their analysis.</p>
Underwater sounds, including killer whale and humpback whale vocalizations, recorded in northern Norway in January 2023
<p>Dataset of underwater acoustic recordings obtained during the expedition “Orcalize” that took place in Skjervøy in northern Norway from 29<sup>th</sup> December 2022 till 6<sup>th</sup> January 2023. The data contains vocalizations from killer whales and songs from humpback whales which gather in the local fjords during the winter months to feed on herring. We recorded in the band of 20 Hz – 60 kHz with calibrated hydrophones arranged in a compact tetrahedral array that we deployed over board of a motorboat. In total we provide 16 files of continuous recordings with duration from several minutes to over one hour. The total dataset is about 7 hours 37 minutes long and the memory size is 62.8 GB. See the file info.pdf for more information.</p>
Azcorra2023 - Fiber photometry recordings (pre-processed to get DF/F)
<p>Pre-processed raw data from fiber photometry recordings of different subtypes (Vglut2+, Calb1+, Anxa1+ and Aldh1a1+ as well as DAT+) SNc dopamine neurons labelled with GCaMP6f, as used in Azcorra et al. Nat Neuro 2023. This dataset has been pre-processed to calculate DF/F from the raw data (see below for code and raw data), which are then normalized from 0 to 1 (un-normalized DF/F data can be recovered using the 'norm' value included in the dataset). This dataset also includes metadata for each recording (recording location, mouse sex...). </p> <p>The code used to generate this pre-processed data from raw data is available on GitHub (<a href="https://github.com/DombeckLab/Azcorra2023/releases/tag/Azcorra2023">https://github.com/DombeckLab/Azcorra2023/releases/tag/Azcorra2023</a>) and Zenodo (DOI: 10.5281/zenodo.7872052, <a href="https://zenodo.org/record/7872052">https://zenodo.org/record/7872052</a>). The original raw data has been deposited on Zenodo (DOI: 10.5281/zenodo.7871634, <a href="https://zenodo.org/record/7871634">https://zenodo.org/record/7871634</a>). The code necessary to analyze this data and generate the figures shown in the manuscript is is found in that same GitHub repository as the pre-processing code above.</p>
Accessible Oceans: Auditory Display. Longterm Axial Seamount Inflation Record
<p>The thirteen tracks make up an auditory display of the Longterm Axial Seamount Inflation Record. The tracks in the auditory display are comprised of data sonifications and contextual audio supports (dialogue, auditory icons, and music). You may <a href="https://samply.app/p/MViV0dJLZjJpFXEHN8EA">listen online here</a>.</p> <p>The display leverages data from NOAA PMEL that extend the record of the National Science Foundation (NSF) Ocean Observatories Initiative (OOI) data back to 1997. This audio display focuses on the long-term pattern observed by bottom pressure recorders where the seafloor inflates (lifts), then an eruption event occurs, and the seafloor drops.</p> <p>The “Accessible Oceans” AISL Pilots and Feasibility study aims to inclusively design auditory displays that support the perception and understanding of ocean data in informal learning environments (ILEs). More can be found on the project website: <a href="https://accessibleoceans.whoi.edu/">https://accessibleoceans.whoi.edu/</a></p>
Extracellular recordings and juxtacellular labelling with glass electrodes in the mouse medial septum and hippocampus
<p>This repository contains MAT files consisting of simultaneously recorded mouse medial septal and hippocampal local field potentials (20 kHz sampling rates) and spikes from single medial septal cells. Data were recorded with glass electrodes during spontaneous movement and rest periods, followed by juxtacellular labelling of the medial septal cell. Text files of the spike times and detected hippocampal CA1 theta (5-12 Hz) oscillation trough times are associated with each MAT file.</p> <p>The files are organised by cell (neuron) name. For further details, see the CSV file included with the dataset. These recorded and labelled single cells were originally reported in Joshi et al 2017, Viney et al 2018, and Salib et al 2019.</p> <p>Each MAT file contains the following channels, exported from the original Spike2 (smr) recording files:</p> <p>(1) Details of the recording</p> <p>(2) Detected spikes (in seconds) from the single medial septal cell</p> <p>(3) Movement detection (eg. accelerometer or rotary encoder)</p> <p>(4) Local field potential (medial septum), in mV</p> <p>(5) Local field potential (hippocampal CA1), in mV; see CSV file for precise location (e.g. within stratum pyramidale)</p> <p>This dataset is made available under a Creative Commons Attribution 4.0 International (CC BY 4.0) license: If you share or adapt these data you must give appropriate credit, provide a link to the license, and indicate if changes were made.</p>
Dataset for "Fast creation of data-driven low-order predictive cardiac tissue excitation models from recorded activation patterns"
<p>This archive contains the source code and data sets presented in the publication "Fast creation of data-driven low-order predictive cardiac tissue excitation models from recorded activation patterns".</p> <p>Kabus, D., De Coster, T., de Vries, A. A., Pijnappels, D. A., & Dierckx, H. (2024). Fast creation of data-driven low-order predictive cardiac tissue excitation models from recorded activation patterns. <em>Computers in Biology and Medicine</em>, 107949. <a title="Persistent link using digital object identifier" href="https://doi.org/10.1016/j.compbiomed.2024.107949" target="_blank" rel="noreferrer noopener"><span>https://doi.org/10.1016/j.compbiomed.2024.107949</span></a></p>
Guinea baboon vocalizations dataset automatically extracted with a deep neural network from natural audio recordings
<p><strong>Abstract</strong></p> <p>The data collection process consisted of continuously recording during one month a group of Guinea baboons living in semi-liberty at the CNRS primatology center in Rousset-sur-Arc (France). Two microphones we placed nearby their enclosure to continuously record the sounds produced by the group. A convolutional neural network (CNN) was used on these large and noisy audio recordings to automatically extract segments of sound containing a baboon vocal production by following the method of <a href="https://arxiv.org/abs/2302.07640">Bonafos et al. (2023)</a>. The resulting dataset consists of one-second to several-minute wav files of automatically detected vocalizations segments. The dataset thus provides a wide range of baboon vocalizations produced at all times of the day. It can be used to study vocal productions of non-human primates, their repertoire, their distribution over the day, their frequency, and their heterogeneity. In addition to the analysis of animal communication, the dataset can also be used as a learning base for sound classification models.</p> <p> </p> <p><strong>Data acquisition</strong></p> <p>The data are audio recordings of baboons. The recordings were made with a H6 Zoom recorder, using the included XYH-6 stereo microphone. The sample size is 44100 Hertz, 16 bits. The microphones were placed in the vicinity of the enclosure for one month and recorded continuously on a PC computer. A CNN passed over the data with a sliding window of 1 second and an overlap of 80% to detect the vocal productions of the baboons. The dataset consists of the segments predicted by the CNN to contain a baboon vocalization. Windows containing signal less than one second apart were merged into a single vocalization.</p> <p> </p> <p><strong>Data source location</strong></p> <ul> <li>Institution: CNRS, Primate Facility</li> <li> <p>City/Town/Region: Rousset-sur-Arc</p> </li> <li> <p>Country: France</p> </li> <li> <p>Latitude and longitude for collected samples/data: 43.47033535251509, 5.6514732876668905</p> </li> </ul> <p> </p> <p><strong>Value of the data</strong></p> <ul> <li> <p>This dataset is relatively unique in terms of the quantity of vocalizations available.</p> </li> <li> <p>This massive dataset can be very useful to two types of scientific communities: experts in primatology who study the vocal productions of non-human primates, and experts in data science and audio signal processing.</p> </li> <li> <p>The machine learning research community has at its disposal a database of several dozen hours of animal vocalizations, which will make it possible to build up a large learning base, very useful for Environemental Sound Recognition tasks, for example.</p> </li> </ul> <p> </p> <p><strong>Objective</strong></p> <p>This dataset is a follow-up of two studies on the vocal productions of Guinea baboons (Papio papio) in which we carried out analyses of their vocal productions on the basis of a relatively large vocalization sample containing around 1300 vocalizations (<a href="https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0169321">Boë, Berthommier, Legou, Captier, Kemp, Sawallis, Becker, Rey, & Fagot, 2017</a>; <a href="https://hal.science/hal-01649539">Kemp, Rey, Legou, Boë, Berthommier, Becker, & Fagot, 2017</a>). The aim was to collect a larger database using the technique of deep convolutional neural networks in order to 1) automatically detect vocal productions in a large continuous audio recording and 2) perform a categorization of these vocalizations on a more massive sample. A description of the pipeline that enabled these automatic detections and categorizations is given in <a href="https://arxiv.org/abs/2302.07640">Bonafos, Pudlo, Freyermuth, Legou, Fagot, Tronçon, & Rey (2023)</a>.</p> <p> </p> <p><strong>Data description</strong></p> <p>The data is a set of audio files in wav format. They are at least one second long (the size of the window), up to several minutes, if several windows are consecutively predicted as containing signal. Moreover, we add the labeled data we used to train the CNN which did the prediction. We also provide two hours of the continuous recordings to have an idea of the continuous recordings and test the code of the paper provided on <a href="https://gitlab.com/papers4375727/detection-and-classification-of-vocal-productions">gitlab</a>.</p> <p>In addition, there is a database in csv format listing all the vocalizations, the day and time of their production, and the prediction probabilities of the model.</p> <p> </p> <p><strong>Experimental design, materials and methods</strong></p> <p>The original recordings represent one month of continuous audio recording. Seven hours of this month were manually labelled. They were segmented and labelled according to whether or not there was a monkey vocalization (i.e., noise or vocalization) and, if there was a vocalization, according to the type of vocalization (6 possible classes: bark, copulation grunt, grunt, scream, yak, wahoo). These manually labelled data were used as a training set for a CNN, which was automatically trained following the pipeline of Bonafos et al. (2023). This model was then used to automatically detect and classify vocalization during the whole month of audio recording. It processes the data in the same way when predicting new data as it does when training. It uses a sliding window of one second with an overlap of 80%. It does not take into account information from previous predictions, but calculates the probability of a vocalization in each one-second window independently. It then iterates through the month. For each window, the model predicts two outputs: the probability that there is a vocalization and the probability of each class of vocalization.</p> <p>For the purpose of generating the wav files, if a window has a probability of a vocalization greater than 0.5, it is considered to contain a vocalization. If it is the first one, a vocalization is started at that moment. If the time windows that follow a vocalization also contain a vocalization, then the signal they contain is added to the first segment for which a vocalization has been detected. As soon as a one-second segment no longer contains a signal corresponding to a vocalization, the wav file is closed. If windows are predicted to contain no vocalizations, but are between two windows that contain vocalizations within 1 second of each other, then all windows are merged.</p>
Tree species, diameter, and canopy class records for 4-paired plots in Black Rock Forest, NY, since 1931.
Black Rock Forest maintains eight long-term forest monitoring plots in Cornwall, NY. Four pairs of plots were established in 1931 to compare thinning treatments to nearby control plots. Four plots are approximately 0.25 acres and the other four are 0.1 acres. Tree species, diameter at breast height, height and canopy class have been measured on all stems greater than 1 inch in diameter since 1931. Plots were revisited every five years until the 1990s and annually after 1994.
Historical Records of the H.J. Andrews Experimental Forest Program, 1947 to present
This database contains historical records of the H.J. Andrews Experimental Forest (a U.S. Forest Service property near Blue River, OR, within the Willamette National Forest) consisting of three parts: 1.) inventory of the physical records and information about locating those physical records and three collections of digital records created from the physical records: 2.) records concerning the forest itself (HJATheForest), and 3.) records of the personal history of H.J. Andrews (HJATheMan).
Long-term record of streamwater chemistry in Sycamore Creek, Arizona, USA (1977-1999)
The primary objective of this project is to understand how long-term climate variability and change influence the structure and function of desert streams via effects on hydrologic disturbance regimes. Climate and hydrology are intimately linked in arid landscapes; for this reason, desert streams are particularly well suited for both observing and understanding the consequences of climate variability and directional change. Researchers try to (1) determine how climate variability and change over multiple years influence stream biogeomorphic structure (i.e., prevalence and persistence of wetland and gravel-bed ecosystem states) via their influence on factors that control vegetation biomass, and (2) compare interannual variability in within-year successional patterns in ecosystem processes and community structure of primary producers and consumers of two contrasting reach types (wetland and gravel-bed stream reaches). This specific dataset was collected to monitor long-term changes in dissolved nutrient concentrations (e.g., nitrogen, phosphorus) and other water-quality parameters by sampling surface water.
CCE LTER process cruise, in the California Current region, event log records including date, time, position and activity for use in post-cruise data integration based on co-sampling indexes. From 2006 to 2019 CCE LTER used a locally developed event logging system. During P2107, CCE LTER started to utilize the R2R Event Logger on UNOL ships, 2006 - 2024 (ongoing).
The event logger program developed and maintained by the California Cooperative Oceanic Fisheries Investigations, SIO, program is used aboard CCE LTER process cruises to create indexes with temporal, spatial and activity information for post-cruise data integration. The event log is configured aboard the ship for the recording of sampling events by both ship crew personnel on the bridge, and research personnel in the lab. The event log is processed post-cruise to correct for various errors.
Hubbard Brook Experimental Forest: Watershed 3 well water level recordings, 2007 - ongoing
This dataset consists of groundwater levels measured within wells distributed across Watershed 3 at Hubbard Brook Experimental Forest from 2007-2020. Water levels are expressed as a depth (cm) from the soil surface. This dataset is a part of a larger project aimed at explaining the spatial and temporal variation in stream water chemistry at the headwater catchment scale using a framework based on the combined study of hydrology and soil development – hydropedology. The project will demonstrate how hydrology strongly influences soil development and soil chemistry, and in turn, controls stream water quality in headwater catchments. Understanding the linkages between hydrology and soil development can provide valuable information for managing forests and stream water quality. Feedbacks between soils and hydrology that lead to predictable landscape patterns of soil chemistry have implications for understanding spatial gradients in site productivity and suitability for species with differing habitat requirements or chemical sensitivity. Tools are needed that identify and predict these gradients that can ultimately provide guidance for land management and silvicultural decision making. Better integration between soil science, hydrology, and biogeochemistry will provide the conceptual leap needed by the hydrologic community to be able to better predict and explain temporal and spatial variability of stream water quality and understand water sources contributing to streamflow. These data were gathered as part of the Hubbard Brook Ecosystem Study (HBES). The HBES is a collaborative effort at the Hubbard Brook Experimental Forest, which is operated and maintained by the USDA Forest Service, Northern Research Station.
The Hubbard Brook Stream Ecology Record: Light, 2018 - ongoing
The Hubbard Brook Stream Ecology record is a companion dataset to the Hubbard Brook Watershed Stream and Precipitation Chemistry record. The Stream Ecology record started in 2018 and HBWatER collects ecological samples from seven gauged watersheds: Watersheds 1 through 6 and Watershed 9. HBWatER measures algal biomass, aquatic invertebrate emergence, and stream decomposition by measuring (1) chlorophyll-a on tiles and artificial moss, which approximate algal biomass growth on bare rock and bryophyte mats, (2) preserved algal biomass on artificial moss substrates in Lugol’s Iodine solution, (3) aquatic invertebrate emergence on replicate sticky traps placed above the stream, and (4) stream decomposition through leaf litter pack and cotton strip decay. To complement these ecological records, HBWaTER installed light sensors and field cameras to obtain better information about the light and stream environment daily. Three replicate light sensors that take sub-daily measurements of light level intensity are placed at each watershed at the weir pond (full-sun), and two under the canopy (partial shade). Field cameras take daily photos at noon of the stream canopy and the stream channel. While many studies at Hubbard Brook have measured algal biomass, aquatic invertebrates, and stream decomposition, they are scattered in locations across the valley, were performed at non-continuous times, and use various semi-comparable methods. The HBWatER Stream Ecology record was created to address this gap and systematically measure any long-term changes in the organisms living in the stream. The collection of HBWatER samples is currently sustained by Tammy Wooster (Cary IES) and analyses of these samples has been performed by Heather Malcom (Cary IES), Audrey Thellman (Duke), and Geoff Wilson (Cary IES). The dataset is curated and maintained by a team of researchers: Chris Solomon (Cary IES), Emma Rosi (Cary IES), and Emily Bernhardt (Duke). Current Financial Support for HBWatER is pro
Hubbard Brook Stream Ecology Record: Diatom Species Richness and Voucher Flora, 2018-2022
This dataset contains species richness data for epiphytic diatom communities collected from weir ponds in seven headwater streams within the Hubbard Brook Experimental Forest (HBEF) in New Hampshire between 2018 and 2021. Diatom samples were gathered using artificial bryophyte substrates, deployed in weir ponds to mimic natural diatom habitats. Species richness was quantified by identifying diatom taxa to the lowest possible taxonomic level, with 86 taxa spanning 43 genera recorded. This dataset represents the first comprehensive classification of diatom communities at HBEF, providing a baseline for future studies in this ecosystem. Environmental variables, including light availability, dissolved organic carbon, total dissolved nitrogen, and pH, were concurrently measured to assess their influence on diatom community composition. The light (lux) data used in this study is openly available in the EDI Data Portal at https://doi.org/10.6073/pasta/0f40b75b299494d736645d940fa2b5a4. The chlorophyll-a data and analysis methodology are available at https://doi.org/10.6073/pasta/7fa32d94240fc7780d62cb7e65eafdb2. Reach characteristics were sourced from the EDI Data Portal at https://doi.org/10.6073/pasta/3e4b95149245341d522383bba51de7c7. This study provides valuable insights into the relationships between environmental factors and diatom diversity in northern hardwood forest streams, aiding ecological monitoring and bioindicator studies. These data were gathered as part of the Hubbard Brook Ecosystem Study (HBES). The HBES is a collaborative effort at the Hubbard Brook Experimental Forest, which is operated and maintained by the USDA Forest Service, Northern Research Station.
CBS01 Capture records of (mainly) Grasshopper Sparrows on Konza Prairie
This dataset includes captures of mainly Grasshopper Sparrows (GRSP) prior to 2017, and after that, additionally many Dickcissels, Eastern Meadowlarks, Brown-headed Cowbirds and other songbirds. Each row pertains to an individual captured on a certain day. Individuals can repeat. Most captures include data on age, sex, head-bill, tarsus, wind chord, molt score, fat score, and mass. In many cases, a single feather was collected from each bird for isotopic analyses. Some individuals were measured for body composition (fat mass, lean mass, and body water) using a mobile Quantitative Magnetic Resonance (QMR) machine. Most individuals were bled in the field within 5 min of capture. The blood was chilled, centrifuged the same day, and plasma stored frozen for analyses of metabolite concentrations. Red blood cells were stored in lysis buffer for genotyping. All birds were banded with a USFWS band and many of the adults were individually marked using a unique combination of 3 plastic colored leg bands. Birds captured as independent young or nestlings banded prior to fledge were only marked with the USFWS bands. All birds were released at the location of capture. Missing values in character fields denoted by NA, and in numeric fields -999.
CBC01 Weekly record of bird species observed on Konza Prairie
Long-term monitoring of bird presence is performed on Konza Prairie. The purpose was to determine bird species phenology of occurrence on entire Konza Prairie. Data on the presence, including documented nesting, of all bird species is recorded weekly in five-year periods e.g. 1980-1984, 1985-1989, 1990-1994.
CBN01 Records of breeding activities for birds on Konza
Dates by species of documented records of breeding - either nests or dependent, fledged young - with contents of nest, nest placement information and location on Konza Prairie recorded by grid square.
Records of Sargassum horneri occurrence in the eastern Pacific
Presented here are records of the occurrence of Sargassum horneri in California, USA, and Baja California, Mexico, since 2003, the year it was first discovered in the eastern Pacific. These data and their sources were published as supplementary tables in: Marks LM, Salinas-Ruiz P, Reed DC, Holbrook SJ, Culver CS, Engle JM, Kushner DJ, Caselle JE, Freiwald J, Williams JP, Smith JR, Aguilar-Rosas LE, Kaplanis NJ (2015) Range expansion of a non-native, invasive, macroalga Sargassum horneri (Turner) C. Agardh, 1820 in the eastern Pacific. BioInvasions Records 4, DOI: 10.3391/bir.2015.4.4.02
Records of moored SeaFET pH, SeaBird CTD and oxygen at Anacapa, Santa Cruz and San Miguel Islands, California from 2012-2015
Data are pH (total scale, SeaFET), salinity, conductivity, temperature, depth (Seabird 37 CTD) and Oxygen (MicroCAT C-T-ODO) from instruments moored at three sites along the north shores of the northern Santa Barbara Channel Islands, California, USA. At Anacapa Island (ALC), in a marine reserve with kelp forest habitat; at Santa Cruz Island (PRZ), surrounded by a large shallow eelgrass bed (Zostera pacifica); and at San Miguel Island (SMN), in open water over a sandy bottom. pH sensors were deployed at all three sites in 2012, with CTD and Oxyten senosrs added to moorings at ALC and PRZ in May 2013. Benchmark samples for SeaFET sensors calibration were collected 1 to 8 times during each 2 - 3 month deployment via SCUBA, free diving or from a pier with a GO-FLOW (General Oceanics) bottle (see methods). Oxygen data were validated by Winkler titations, with concentrations reported in various units and as saturation. Data are presented in: Kapsenberg, L. and G. E. Hofmann 2016. Ocean pH time-series and drivers of variability along the northern Channel Islands, California, USA. Limnology and Oceanography. 61: 953-968. DOI: 101002/lno.10264.
Record of storm events and associated water levels for the Virginia Coast Reserve, 1980-2013
This empirical storm record for the Virginia Coast Reserve was created using a 34-year record of hourly wave hindcast data - including wave height (Hs) and wave period (Tp) - from the USACE's Wave Information Studies buoy offshore Hog Island in the Virginia Coast Reserve (Station 63183, 22 m water depth) and hourly records of water level from the nearest NOAA tide gauge (Station 8631044, Wachapreague, VA). The record includes wave and water level statistics for each event relevant for coastal modeling applications: storm start and end times, duration, total water level, still water level, as well as concurrent tidal amplitude, non-tidal residual, Hs, and Tp. The raw data is processed by first removing the 1 yr running median, which accounts for non-stationarity in wave and water level parameters due to inter-annual and decadal variability while maintaining seasonality. The median of the last 3 years is then applied to the entire time series such that the new time series is representative of the current climate. A year-by-year tidal analysis is performed to obtain the tidal amplitude and non-tidal residual. Lastly, water elevations are calculated following the run-up equations of Stockdon et al. (2006). Storm events are then extracted from the corrected time series by conditioning on Hs: events are identified as periods of 8 or more consecutive hours with deep-water significant wave heights greater than 2.1 m, which is the minimum monthly averaged wave height for periods in which waters levels exceeded the measured average dune toe elevation (1.9 m NAVD88) of barriers in the Virginia Coast Reserve. In total, we identify 282 independent sea-storm events over the 34-year record, resulting in an average of 8.3 events per year. See Reeves et al. (2021; https://doi.org/10.1029/2021GL092958) and supplementary information therein for complete details of the methodology.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.