Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

200

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

200 results for “contextualization”

Learn how ShareScore rates datasets ↗
zenodo52/100

The International Soundscape Database: An integrated multimedia database of urban soundscape surveys -- questionnaires with acoustical and contextual information

<h1>Introduction</h1> <p>The International Soundscape Database contains the results of a series of soundscape assessment campaigns carried out across Europe and China. The data collection process was conducted according to the <a href="https://www.mdpi.com/2076-3417/10/7/2397">SSID Protocol [1]</a> which integrates in situ questionnaires about users' soundscape experience, with binaural recordings, sound level meter readings, and 360 degree video. The core of this database are individual soundscape questionnaires collected for 3,500+ participants completed in situ in cities across Europe and China, and the psychoacoustic analysis of 30s binaural recordings which can be matched up to each questionnaire.</p> <p>The SSID Protocol was based on the ISO 12913&nbsp;standard for soundscape data collection [2]. For more information on the specifics of how this data is collected, please see [1].</p> <p>It is the intention that this dataset be added to and augmented with new locations, cities, and contexts in the future. This will be done both by the SSID team at University College London, but we also strongly welcome contributions from other researchers and practicioners. If a soundscape assessment is collected according to the SSID Protocol, it can be integrated with the rest of the database to form a large, cohesive, and ever-growing database of soundscape assessments.&nbsp;</p> <h2>Analysis</h2> <p>Code for exploring and analysing this dataset is included as part of the <a href="https://soundscapy.readthedocs.io/en/latest/">Soundscapy package</a>.</p> <h2>Included Files</h2> <p>This dataset incorporates surveys taken in multiple urban public spaces across several cities in Europe and China. These urban spaces include places like parks, urban squares, green spaces, and market streets. At each location, up to 100 questionnaires were collected over a series of multi-hour long sessions. Therefore the data is organised by LocationID, then SessionID, then GroupID.</p> <p>The basic directory structure and contents can be found below.&nbsp;</p> <h3>Survey Data (.csv)</h3> <p>'ISD v1.0 Data.csv' organises the data according to the labels given above.</p> <h3>Survey Metadata (.xlsx)</h3> <p>In addition a metadata file ('ISD v1.0 Metadata.xlsx') with photos and descriptions of each of the locations is provided. This metadata file also includes Data Dictionaries for each of the survey instrument versions included. These data dictionaries document precisely the questions asked and the available reponse labels and coding, along with the relevant translations.</p> <h3>Psychoacoustic Analysis (.csv)</h3> <p>The compiled csv file is formatted with a row for each individual participant's questionnaire response, then includes the psychoacoustic analysis of the 30s binaural recording taken while the participant was completing the questionnaire. Details about the psychoacoustic analyses is given in the 'Acoustic Settings' tab in the metadata file.</p> <p>The compiled survey and psychoacoustic analysis data is contained in 'ISD v1.0 Data.csv'. This is compiled from raw survey data files contained in 'Survey_Data', with individual cleaned survey and psychoacoustic data files included in 'Survey_Data/Interim_&lt;date&gt;'. The scripts for compiling this data are included in 'Scripts/'.</p> <h3>Sound Level Meter logs (.xlsx)</h3> <p>'SLM_&lt;city&gt;/' folders include session-long (i.e. ~3hrs) sound level meter log data in.xlsx files for each SessionID.</p> <h3>Binaural Recordings (32-bit floating point .wav)</h3> <p>'WAV_&lt;city&gt;/' folders include the ~30s binaural recordings in 32 bit floating point .wav format. Within each city folder are a set of LocationID folders containing their associated recordings. The wav files are titled with its GroupID, which is matched to the corresponding survey GroupIDs.&nbsp;</p> <h3>Cleaning and Compilation Scripts (.py)</h3> <p>Python code for cleaning and compiling the data from the raw survey data (within Survey_Data/source_data) are provided. These can be run within the provided demo notebook, or from the terminal by calling 'python -m ISDv1_main' with the relevant arguments. See the README.md file in this directory for more information.</p> <pre><code><br>├── ISD v1.0 Data.csv ├── ISD v1.0 Metadata.xlsx ├── SLM_Granada │ ├── CampoPrincipe1_SLM.xlsx │ ├── ... ├── SLM_Groningen │ └── Noorderplantsoen1_SLM.xlsx ├── SLM_etc ├── Scripts │ ├── ISDcleanDemo.ipynb │ ├── ISDcleaning.py │ ├── ISDpsycho.py │ ├── ISDv1_main.py │ ├── README.md │ └── pyproject.toml ├── Survey_Data │ ├── Interim_2024-02-08_cleaned │ └── source_data ├── WAV_Granada_1 │ ├── CampoPrincipe │ ├── ... ├── WAV_etc</code></pre> <p><strong>Citation</strong>: If you use the ISD or part of it, please cite our paper describing the data collection protocol [1] and this dataset itself.</p> <p><strong>License and reuse</strong>: All ISD recordings are provided under the Creative Commons Attribution 4.0 International (CC BY 4.0) License and are free to use. We encourage other researchers to replicate the SSID protocol and contribute new locations to the dataset. We also encourage the use of these recordings and the perceptual data for further soundscape research purposes. Please provide the proper attribution and get in touch with the authors if you would like to contribute new data or for any other collaborations.</p> <p>&nbsp;</p> <p>[1] Mitchell A, Oberman T, Aletta F, Erfanian M, Kachlicka M, Lionello M, Kang J. The Soundscape Indices (SSID) Protocol: A Method for Urban Soundscape Surveys&mdash;Questionnaires with Acoustical and Contextual Information. <em>Applied Sciences</em>. 2020; 10(7):2397. <a href="https://www.mdpi.com/2076-3417/10/7/2397">https://doi.org/10.3390/app10072397&nbsp;</a></p> <p>[2]&nbsp;ISO/TS 12913-2:2018 (2018). &ldquo;Acoustics &ndash; Soundscape &ndash; Part 2: Data collection and reporting requirements&rdquo; International Organization for Standardization, Geneva, Switzerland, 2018</p> <p>[3] Mitchell A, Oberman T, Aletta F, Kachlicka M, Lionello M, Erfanian M, Kang J. Investigating Urban Soundscapes of the COVID-19 Lockdown: A predictive soundscape modeling approach.<em>&nbsp;Journal of the Acoustical Society of America</em>. 2021.</p>

opencc-by-4.0Feb 2024View details →
zenodo48/100

Auditory stimuli suppress contextual fear responses in safety learning independent of a possible safety meaning

<p>This repository stores the raw data that gave rise to the study by Mombelli et al. (2024) (Title: Auditory stimuli suppress contextual fear responses in safety learning independent of a possible safety meaning; DOI: 10.3389/fnbeh.2024.1415047, Journal: Frontiers in Behavioral Neuroscience).&nbsp; Below we supply information on the provided metadata files which, in turn, refer to individual raw data files.</p> <p><strong>General structure of the repository:</strong></p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; the raw data is organized in 5 subsets defined by the figures or supplementary figures they contribute to. Each subset is documented by its own metadata file. Raw data files were compressed into ZIP archives, one per subset;</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; the metadata files listing names of the individual data files are provided in &ldquo;.csv&rdquo; format, one per data subset. Field separator: comma;</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; the dataset is accessible at the following doi: 10.5281/zenodo.13524007</p> <p>&nbsp;</p> <p><strong>Description of the non-textual data formats:</strong></p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; video recordings of animal behavior were provided as unmodified ".wmv" files created by the VideoFreeze acquisition software (Med Associates Inc). Video stream parameters: wmv3 codec, color space yuv420p, 320x240 pixels, 30 fps, bitrate 300 kb/s.</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; movement traces were obtained from the videos, as described in the Methods section (Mombelli et al., 2024).</p>

opencc-by-4.0Sep 2024View details →
zenodo48/100

TDA4ContextualEmbeddings - Public - Debug Data for the codebase of the publication "Local Topology Measures of Contextual Language Model Latent Spaces With Applications to Dialogue Term Extraction"

<p>Debug dataset for testing the <a href="https://gitlab.cs.uni-duesseldorf.de/general/dsml/tda4contextualembeddings-public">codebase</a> of the paper <a href="https://doi.org/10.18653/v1/2024.sigdial-1.31">&ldquo;Local Topology Measures of Contextual Language Model Latent Spaces With Applications to Dialogue Term Extraction&rdquo;</a> published at the 25th Meeting of the Special Interest Group on Discourse and Dialogue, Kyoto, Japan (SIGDIAL 2024).</p>

openapache2.0Nov 2024View details →
zenodo44/100

Contextual dataset from a Public Service Media

<p>Dataset from a Public Service Media (PSM) whose users are not logged in. Thus, this leads to a pure cold-start problem where each interaction is viewed as an isolated event.</p>

opencc-by-4.0Apr 2023View details →
zenodo40/100

Figure 2 in Improved local inventory and regional contextualization for anuran (Amphibia) diversity assessment at an endangered habitat in southeastern Brazil

Figure 2. Rarefaction curves based on Jackknife I species-richness estimator for records of adults, tadpoles and all life stages pooled for four canga lakes at the Quadrilátero Ferrífero region, southeastern Brazil.

opencc-by-4.0Sep 2015View details →
zenodo40/100

Datasets from the RecSys 2020 article "Carousel Personalization in Music Streaming Apps with Contextual Bandits"

<p>We publicly release&nbsp;the anonymized <em>user_features.csv</em> and <em>playlist_features.csv</em> datasets, from the music streaming platform Deezer, as described in the&nbsp;article &quot;<em>Carousel Personalization in Music Streaming Apps with Contextual Bandits&quot;</em>&nbsp;published in the proceedings of the 14th ACM Conference on Recommender Systems (<em>RecSys 2020</em>). The paper is available <a href="https://arxiv.org/abs/2009.06546">here</a>.</p> <p>These datasets are used in the&nbsp;GitHub repository <a href="https://github.com/deezer/carousel_bandits">deezer/carousel_bandits</a> to reproduce experiments from the article.</p> <p>Please cite our paper if you use our code or data in your work.</p>

opencc-by-4.0Sep 2020View details →
zenodo40/100

Contextualizing Trending Entities in News Stories

<p>This repository contains the enrichments for the dataset <a href="https://catalog.ldc.upenn.edu/LDC2008T19">The New York Times Annotated Corpus</a> developed for the paper:</p> <p>&ldquo;Marco Ponza, Diego Ceccarelli, Paolo Ferragina, Edgar Meij, Sambhav Kothari. Contextualizing Trending Entities in News Stories. In Proceedings of the 14th ACM International Conference on Web Search and Data Mining (WSDM 2021).&rdquo;</p> <p>It includes a total of 149 trends constituted by 120K entities. The goal is to retrieve a set of entities ranked with respect to their usefulness in explaining why a given trending entity is actually trending.</p> <p><strong>Format</strong></p> <p>The repository contains the enrichments in JSON format.</p> <p>The news stories of the New York Times from which these enrichments have been developed are available from <a href="https://catalog.ldc.upenn.edu/LDC2008T19">LDC</a>.</p> <p><strong>Data Splits</strong></p> <p>We perform two kinds of evaluation.</p> <ol> <li>Unsupervised evaluation, where we use the complete dataset of 149 trends as a benchmark.</li> <li>Supervised evaluation, where we train/tune our models on a training/development set and we test them on a test set.</li> </ol> <ul> <li>The training set contains 50 trends constituted by 36.3K entities from 1996 to 2000.</li> <li>The development set contains 34 trends constituted by 26.7K entities from 2000 to 2002.</li> <li>The test set contains 65 trends constituted by 57K entities from 2002 to 2007.</li> </ul> <p>Use</p> <p>Please cite the data set and the accompanying paper if you found the resources in this repository useful:</p> <p>@inproceedings{ponza2021,<br> &nbsp;&nbsp;&nbsp;&nbsp; Title = {Contextualizing Trending Entities in News Stories},<br> &nbsp;&nbsp;&nbsp;&nbsp; author = {Ponza, Marco and Ceccarelli, Diego and Ferragina, Paolo and Meij, Edgar and Kothari, Sambhav},<br> &nbsp;&nbsp;&nbsp;&nbsp; Booktitle = {Proceedings of the 14th ACM International Conference on Web Search and Data Mining},<br> &nbsp;&nbsp;&nbsp;&nbsp; Year = {2021},<br> }</p>

opencc-by-4.0Jan 2021View details →
zenodo40/100

Contextual Factors Research in Continuous Integration (CI) Projects

<p>These files include process documentation for the research on project contextual factors in Continuous Integration (CI). They cover previous research studies and survey details.&nbsp;</p>

opencc-by-4.0May 2024View details →
zenodo40/100

Summarised contextual data about metabarcoding Tara Oceans samples (2009-2013)

<p>Tab-separated values table describing the metabarcoding samples from the expedition Tara Oceans (2009-2013).</p> <p>Information such as depth, time, geographic position, size fraction, collected from <a href="https://pangaea.de/">Pangaea</a>, are listed in context_general tables. In context_stat tables, you will find a selection of physico-chemical parameters. Tara_Oceans_Pangaea_context.rds gathers all the data collected from Pangaea in a single R object.</p> <p>These tables have been built using the code here: <a href="https://gitlab.com/tara-and-friends-euk-metab/tara-oceans-metab-context/-/tree/v1.1.1" target="_blank" rel="noopener">https://gitlab.com/tara-and-friends-euk-metab/tara-oceans-metab-context/-/tree/v1.1.2</a> (v1.1.2).</p> <p>In this version 16S metabarcoding samples missing in previous versions were added in context_general.* and context_sats.*</p>

opencc-by-4.0Oct 2022View details →
zenodo40/100

Dataset for the manuscript "Crowding results from optimal integration of visual targets with contextual information"

<p>There are seven experimental datasets, two program with which data are collected, two supplemetary programs needed to run the main code and one program to analyse data.&nbsp;Two .txt files are included, where we describe how to use the stimulation and analysis programs.</p>

opencc-by-4.0Jul 2022View details →
zenodo40/100

Figure 4 in Fishers' perceptions of river resources: case study of French Guiana native populations using contextual cognitive mapping

Figure 4. – Cognitive maps focused on a minimum common overview (gray concepts and edges (incoming arrows)) and village-specific overviews (white concepts and blue (5 Amerindian villages) or brown (2 Aluku villages) edges) of threats to the fish resource and environment.

opencc-by-4.0Dec 2021View details →
zenodo40/100

Behavior-driven Load Testing Using Contextual Knowledge—Approach and Experiences

<p>This package holds supplementary material for the paper <em>Behavior-driven Load Testing Using Contextual Knowledge &mdash; Approach and Experiences</em>, Proceedings of the 10th ACM/SPEC International Conference on Software Performance (ICPE 2019).</p> <p>The supplementary material subsumes the following:</p> <ol> <li>The <em>readme.md</em> provides more details about the supplementary material.</li> <li>The <em>bdlt-grammar</em> folder holds the BDLT grammar in the Extended Backus-Naur Form (EBNF) notation (<em>bdlt-grammar.ebnf</em>) and a visualization of this EBNF using railroad diagrams (<em>index.html</em>).</li> <li>The <em>industrial</em> folder holds BDLT definitions as well as a corresponding BenchFlow test, which have been developed in the industrial case study described in the paper.</li> <li>The <em>laboratory</em> folder provides details of another laboratory study that goes beyond the level of details in the paper.</li> </ol> <p>&nbsp;</p>

opencc-by-4.0Feb 2019View details →
zenodo40/100

Obtaining Better Static Word Embeddings Using Contextual Embedding Models

<p><strong>Obtaining Better Static Word Embeddings Using Contextual Embedding Models</strong></p> <p>This repository contains the dataset of pretrained word embeddings as well as datasets used to train them, released with the following <a href="https://arxiv.org/pdf/2106.04302.pdf">paper</a>.</p> <blockquote> <p>&ldquo;Obtaining Better Static Word Embeddings Using Contextual Embedding Models&rdquo; <em>ACL</em> (2021).</p> </blockquote> <p>The wikipedia datasets were preprocessed from the wikipedia dump downloaded from <a href="http://dumps.wikimedia.org">dumps.wikimedia.org</a> under&nbsp;Creative Commons Attribution-Share-Alike 3.0 License .</p> <p>If you found the provided resources useful, please cite the above paper. Here&#39;s a BibTeX entry you may use:</p> <blockquote> <p>@inproceedings{Gupta2021ObtainingPC,<br> &nbsp; title={Obtaining Better Static Word Embeddings Using Contextual Embedding Models},<br> &nbsp; author={Prakhar Gupta and Martin Jaggi},<br> &nbsp; booktitle={ACL},<br> &nbsp; year={2021}<br> }</p> </blockquote>

opencc-by-3.0Jun 2021View details →
zenodo40/100

single-cell RNAseq data (data set 1) in the publication scFASTCORMICS: A contextualization algorithm to reconstruct metabolic multi-cell population models from single-cell RNAseq data

<p>The present dataset (dataset1) was used as input to build scFASTCORMICS models. The files correspond to the clusters identified by&nbsp;Seurat in the single-cell data from CRC samples downloaded from the GEO website&nbsp; (<strong>GSE81861). </strong></p> <p>see the protocol: scFASTCORMICS: A contextualization algorithm to reconstruct metabolic multi-cell population models from single-cell RNAseq data</p> <p>and github: https://github.com/sysbiolux/scFASTCORMICS</p> <p>For more information, version updates of the scFASTCORMICS.&nbsp;</p>

opencc-by-4.0Nov 2022View details →
zenodo40/100

Manually labeled Bird song dataset of 22 species from Xeno-canto to enhance deep learning acoustic classifiers with contextual information.

<p>Data accompanying the paper: Jeantet and Dufourq (2023). Empowering Deep Learning Acoustic Classifiers with Human-like Ability to Utilize Contextual Information for Wildlife Monitoring. <em>Ecological Informatics</em>. 77, 15749541, DOI: 10.1016/j.ecoinf.2023.102256</p> <p>&nbsp;</p> <p>Our investigation contributes to the field of deep learning and bioacoustics by highlighting the potential for improved classification performance through the incorporation of contextual information such as time and location.</p> <p>To test if spatial-temporal information can enhance deep learning classifier, we developed a subset dataset derived from Xeno-Canto that included location metadata as input alongside the spectrogram. We used this dataset with the primary purpose of creating a bird song classification task with species carefully selected to share similar vocal characteristics but from distinct geographical distributions. We only considered the recordings of category `A', corresponding to the best quality score in the database.</p> <p>The dataset contains songs of <strong>22 bird species</strong> from 5 families and genera differents. The recordings were downloaded from the Xeno-canto database in .wav format and each recording was <strong>manually annotated </strong>by labelling the start and stop time for every vocalisation occurrence using Sonic Visualiser. In total, database contained 6537 occurrences of bird songs of various length from <strong>967 file recordings</strong>. A precise description of the distribution by species and country can be found in the associated article.</p> <p>&nbsp;</p> <p>The audio files are provided in "Audio.zip" and the manually verified annotation in "Annotations.zip". The name of each file follows the following nomenclature: Family_genus_species_country of recording_date of recording_ID Xenocanto_type of song.wav/svl. The meta-data information of each file can be find in the csv file provided (Xenocanto_metadata_qualityA_selection) based on the number of the ID Xeno-canto. The annotations can be viewed using the Sonic Visualiser software. The python codes to process these files and train neural networks can be found here : github</p> <p>The files were divided into a <strong>training folder</strong> and a<strong> validation folder</strong> to train and evaluate the efficiency of each method. For each species and country, we randomly selected 70% of the downloaded recordings for the training dataset and kept the remaining 30% for validation.</p> <p><strong>Process to select the species</strong> : We selected the ten most recorded families in the Passeriformes order, the most represented order in Xeno-canto database. From each of the ten families, we again sub-samples the ten most recorded genera. For each genus, we observed the countries of the recordings and the number of available recordings per species and countries. From these observations, we made a self-selection of genera containing species with similar songs but recorded in different regions, with enough recordings available by species and country to form a dataset . At the end, 5 genus were&nbsp; selected containing 22 species. We considered only recordings associated with bird songs, specifically, within Xeno-canto we selected the `song' type. To balance the number of recordings between species of the same genus, we reduced the number of recordings for the most represented species. Thus, for each genus we calculated the average of the number of records available per species and per country and limited the number of recordings for the species/country pairs that were in greater number to this value plus two.</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Apr 2023View details →
dryad40/100

Experimental data analyzed in: Signal detection models as contextual bandits

<p>Signal detection theory (SDT) has been widely applied to identify the optimal discriminative decisions of receivers under uncertainty. However, the approach assumes that decision-makers immediately adopt the appropriate acceptance threshold, even though the optimal response must often be learned. Here we recast the classical normal-normal (and power-law) signal detection model as a contextual multi-armed bandit (CMAB). Thus, rather than starting with complete information, decision-makers must infer how the magnitude of a continuous cue is related to the probability that a signaller is desirable, while simultaneously seeking to exploit the information they acquire. We explain how various CMAB heuristics resolve the trade-off between better estimating the underlying relationship and exploiting it. Next, we determined how naïve human volunteers resolve signal detection problems with a continuous cue. As anticipated, a model of choice (accept/reject) that assumed volunteers immediately adopted the SDT-predicted acceptance threshold did not predict volunteer behaviour well. The Softmax rule for solving CMABs, with choices based on a logistic function of the expected payoffs, best explained the decisions of our volunteers but a simple midpoint algorithm also predicted decisions well under some conditions. CMABs offer principled parametric solutions to solving many classical SDT problems when decision-makers start with incomplete information.</p>

opencc-zeroMay 2023View details →
dryad40/100

When the neighborhood matters: contextual selection on seedling traits in native and non-native California grasses

<p>Plants interact extensively with their neighbors, but the evolutionary consequences of variation in neighbor identity are not well understood. Seedling traits are likely to experience selection that depends on the identity of neighbors because they influence competitive outcomes. To explore this, we evaluated selection on seed mass and emergence time in two California grasses, the native perennial <em>Stipa pulchra</em> and the non-native annual <em>Bromus diandrus</em>, in the field with six other native and non-native neighbor grasses in individual and mixed species treatments. We also <span>quantified characteristics of each neighbor treatment to further investigate factors influencing their effects on fitness and phenotypic selection. Selection favored larger seeds in both focal species and this</span> was largely independent of neighbor identity. Selection generally favored earlier emergence in both focal species, but neighbor identity influenced the strength and direction of selection on emergence time in <em>S. pulchra</em> but not <em>B. diandrus</em>. Greater light interception, higher soil moisture, and greater productivity of neighbors was associated with more intense selection for earlier emergence and larger seeds. Our findings suggest that changes in plant community composition can alter patterns of selection in seedling traits, and that these effects can be associated with measurable characteristics of the community.</p>

opencc-zeroJun 2023View details →
zenodo40/100

MalaMix dataset: contextual and metabarcoding data

<p>1. INTRODUCTION</p> <p><em>MalaMix </em>is a compiled metabarcoding dataset composed of 451 marine samples collected from a range of depths - from the surface (3m) to deep waters (as far down as 4800m). This dataset covers three ocean layers: the epi- (0-200m &ndash; including DCM), meso- (200-1000m) and bathypelagic (1000-4000m). <em>MalaMix</em> combines samples obtained during two oceanographic expeditions with similar sampling strategies: i) the Malaspina-2010 global expedition that produced 263 samples collected between December 2010 and July 2011 from 120 stations distributed along the tropical and subtropical portions (latitudes between 35&deg; N and 40&deg; S) of the Pacific, Atlantic and Indian oceans; and ii) the <em>HotMix</em> trans-Mediterranean cruise that produced 188 samples collected between April and May 2014 in 29 stations distributed along the whole Mediterranean Sea (from -5&deg; W to 33&deg; E) and the adjacent Northeast Atlantic Ocean.</p> <p>&nbsp;</p> <p><em>MalaMix</em> comprises:</p> <ul> <li>a 16S-V4V5 rRNA gene ASV table (MalaMix_16S.csv);</li> <li>an 18S-V4 rRNA gene ASV table (MalaMix_18S.csv);</li> <li>two tables of contextual metadata (MalaMix_EnvData_16S and MalaMix_EnvData_18S) including 6 standardized environmental parameters (temperature [&deg;C], salinity, fluorescence, PO<sub>4</sub><sup>3&minus; </sup>[&micro;mol L<sup>-1</sup>], NO<sub>3</sub><sup>&minus; </sup>[&micro;mol L<sup>-1</sup>], and SiO<sub>2 </sub>[&micro;mol L<sup>-1</sup>]) as well as species taxonomic and phylogenetic diversity metrics</li> <li>a table (MalaMix_FCdata.csv) with flow cytometry microbial counts [cell mL<sup>-1</sup>] and bacterial activity measurements [pmol Leu L<sup>-1</sup> h<sup>-1</sup>];</li> <li>a README file (README_Metadata.csv) describing the meaning and units of each variable column in the metadata tables.</li> </ul> <p>&nbsp;</p> <p>The raw DNA sequences are publicly available at the European Nucleotide Archive (<a href="https://www.ebi.ac.uk/ena">https://www.ebi.ac.uk/ena</a>) under accession numbers PRJEB23913 [18S rRNA genes] &amp; PRJEB25224 [16S rRNA genes] for the Malaspina surface dataset; PRJEB23771 [18S rRNA genes] &amp; PRJEB45015 [16S rRNA genes] for the Malaspina vertical profiles; PRJEB45011 [16S rRNA genes] &amp; PRJEB45014 [18S rRNA genes] for the Malaspina deep sea dataset; and PRJEB44683 [18S rRNA genes] &amp; PRJEB44474 [16S rRNA genes] for the HotMix expedition.</p> <p>&nbsp;</p> <p>Further methodological details are available here: <a href="https://www.biorxiv.org/content/10.1101/2023.01.13.523743v1">https://www.biorxiv.org/content/10.1101/2023.01.13.523743v1</a></p> <p>&nbsp;</p> <p>2. FUTURE FORMAT CHANGES</p> <p>No major changes are expected for the main general format of the database.</p> <p>3. ACKNOWLEDGMENTS</p> <p>The current dataset was generated with funds from the projects INTERACTOMICS (CTM2015-69936-P, MINECO, Spain), MicroEcoSystems (240904, RCN, Norway), MINIME (PID2019-105775RB-I00, AEI, Spain), and PID2021-125469NB-C31 (AEI, Spain), as well as DOREMI (CTM2012-34294) and HOTMIX (CTM2011-30010-C02-01 and CTM2011-30010-C02-02) of the Spanish Ministry of Economy and Innovation, co-financed with FEDER funds.</p> <p>4. COPYRIGHT NOTICE<br> <br> This database is provided &ldquo;as is&rdquo; and without any warranty of any kind, of openly available for non-commerical purposes (CC BY-NC). <strong>CC BY-NC</strong> means that users can make use of the work (including copying, distributing, adapting and building upon the work), but only for noncommercial purposes and as long as attribution is given to the creator: <a href="https://oabooks-toolkit.org/lifecycle/article/4012101-choosing-a-license">https://oabooks-toolkit.org/lifecycle/article/4012101-choosing-a-license</a></p>

opencc-by-nc-4.0Sep 2023View details →
dryad40/100

Experimental data analyzed in: Signal detection models as contextual bandits

Open the record for dataset details and reuse information.

publicMay 2023View details →
dryad40/100

When the neighborhood matters: contextual selection on seedling traits in native and non-native California grasses

Open the record for dataset details and reuse information.

publicJun 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record