Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,655
datasets available to search
ShareScore release 0.9.0
Dataset results
1,655 results for “Subset”
California Current Ecosystem site, station Ohman Region: subset of CalCOFI stations inshore and nearshore in the Southern California Bight region; CalCOFI lines 80-93, stations from shore offshore to station 70, study of chlorophyll a in units of microgramsPerLiter on a yearly timescale
The EcoTrends project was established in 2004 by Dr. Debra Peters (Jornada Basin LTER, USDA-ARS Jornada Experimental Range) and Dr. Ariel Lugo (Luquillo LTER, USDA-FS Luquillo Experimental Forest) to support the collection and analysis of long-term ecological datasets. The project is a large synthesis effort focused on improving the accessibility and use of long-term data. At present, there are ~50 state and federally funded research sites that are participating and contributing to the EcoTrends project, including all 26 Long-Term Ecological Research (LTER) sites and sites funded by the USDA Agriculture Research Service (ARS), USDA Forest Service, US Department of Energy, US Geological Survey (USGS) and numerous universities. Data from the EcoTrends project are available through an exploratory web portal (http://www.ecotrends.info). This web portal enables the continuation of data compilation and accessibility by users through an interactive web application. Ongoing data compilation is updated through both manual and automatic processing as part of the LTER Provenance Aware Synthesis Tracking Architecture (PASTA). The web portal is a collaboration between the Jornada LTER and the LTER Network Office. The following dataset from California Current Ecosystem (CCE) contains chlorophyll a measurements in microgramsPerLiter units and were aggregated to a yearly timescale.
California Current Ecosystem site, station Ohman Region: subset of CalCOFI stations inshore and nearshore in the Southern California Bight region; CalCOFI lines 80-93, stations from shore offshore to station 70, study of nitrogen from nitrate in coastal water in units of microMolesPerLiter on a yearly timescale
The EcoTrends project was established in 2004 by Dr. Debra Peters (Jornada Basin LTER, USDA-ARS Jornada Experimental Range) and Dr. Ariel Lugo (Luquillo LTER, USDA-FS Luquillo Experimental Forest) to support the collection and analysis of long-term ecological datasets. The project is a large synthesis effort focused on improving the accessibility and use of long-term data. At present, there are ~50 state and federally funded research sites that are participating and contributing to the EcoTrends project, including all 26 Long-Term Ecological Research (LTER) sites and sites funded by the USDA Agriculture Research Service (ARS), USDA Forest Service, US Department of Energy, US Geological Survey (USGS) and numerous universities. Data from the EcoTrends project are available through an exploratory web portal (http://www.ecotrends.info). This web portal enables the continuation of data compilation and accessibility by users through an interactive web application. Ongoing data compilation is updated through both manual and automatic processing as part of the LTER Provenance Aware Synthesis Tracking Architecture (PASTA). The web portal is a collaboration between the Jornada LTER and the LTER Network Office. The following dataset from California Current Ecosystem (CCE) contains nitrogen from nitrate in coastal water measurements in microMolesPerLiter units and were aggregated to a yearly timescale.
California Current Ecosystem site, station Ohman Region: subset of CalCOFI stations inshore and nearshore in the Southern California Bight region; CalCOFI lines 80-93, stations from shore offshore to station 70, study of nitrate in coastal water in units of microMolesPerLiter on a yearly timescale
The EcoTrends project was established in 2004 by Dr. Debra Peters (Jornada Basin LTER, USDA-ARS Jornada Experimental Range) and Dr. Ariel Lugo (Luquillo LTER, USDA-FS Luquillo Experimental Forest) to support the collection and analysis of long-term ecological datasets. The project is a large synthesis effort focused on improving the accessibility and use of long-term data. At present, there are ~50 state and federally funded research sites that are participating and contributing to the EcoTrends project, including all 26 Long-Term Ecological Research (LTER) sites and sites funded by the USDA Agriculture Research Service (ARS), USDA Forest Service, US Department of Energy, US Geological Survey (USGS) and numerous universities. Data from the EcoTrends project are available through an exploratory web portal (http://www.ecotrends.info). This web portal enables the continuation of data compilation and accessibility by users through an interactive web application. Ongoing data compilation is updated through both manual and automatic processing as part of the LTER Provenance Aware Synthesis Tracking Architecture (PASTA). The web portal is a collaboration between the Jornada LTER and the LTER Network Office. The following dataset from California Current Ecosystem (CCE) contains nitrate in coastal water measurements in microMolesPerLiter units and were aggregated to a yearly timescale.
California Current Ecosystem site, station Ohman Region: subset of CalCOFI stations inshore and nearshore in the Southern California Bight region; CalCOFI lines 80-93, stations from shore offshore to station 70, study of phosphorus from phosphate in coastal water in units of microMolesPerLiter on a yearly timescale
The EcoTrends project was established in 2004 by Dr. Debra Peters (Jornada Basin LTER, USDA-ARS Jornada Experimental Range) and Dr. Ariel Lugo (Luquillo LTER, USDA-FS Luquillo Experimental Forest) to support the collection and analysis of long-term ecological datasets. The project is a large synthesis effort focused on improving the accessibility and use of long-term data. At present, there are ~50 state and federally funded research sites that are participating and contributing to the EcoTrends project, including all 26 Long-Term Ecological Research (LTER) sites and sites funded by the USDA Agriculture Research Service (ARS), USDA Forest Service, US Department of Energy, US Geological Survey (USGS) and numerous universities. Data from the EcoTrends project are available through an exploratory web portal (http://www.ecotrends.info). This web portal enables the continuation of data compilation and accessibility by users through an interactive web application. Ongoing data compilation is updated through both manual and automatic processing as part of the LTER Provenance Aware Synthesis Tracking Architecture (PASTA). The web portal is a collaboration between the Jornada LTER and the LTER Network Office. The following dataset from California Current Ecosystem (CCE) contains phosphorus from phosphate in coastal water measurements in microMolesPerLiter units and were aggregated to a yearly timescale.
California Current Ecosystem site, station Ohman Region: subset of CalCOFI stations inshore and nearshore in the Southern California Bight region; CalCOFI lines 80-93, stations from shore offshore to station 70, study of phosphate in coastal water in units of microMolesPerLiter on a yearly timescale
The EcoTrends project was established in 2004 by Dr. Debra Peters (Jornada Basin LTER, USDA-ARS Jornada Experimental Range) and Dr. Ariel Lugo (Luquillo LTER, USDA-FS Luquillo Experimental Forest) to support the collection and analysis of long-term ecological datasets. The project is a large synthesis effort focused on improving the accessibility and use of long-term data. At present, there are ~50 state and federally funded research sites that are participating and contributing to the EcoTrends project, including all 26 Long-Term Ecological Research (LTER) sites and sites funded by the USDA Agriculture Research Service (ARS), USDA Forest Service, US Department of Energy, US Geological Survey (USGS) and numerous universities. Data from the EcoTrends project are available through an exploratory web portal (http://www.ecotrends.info). This web portal enables the continuation of data compilation and accessibility by users through an interactive web application. Ongoing data compilation is updated through both manual and automatic processing as part of the LTER Provenance Aware Synthesis Tracking Architecture (PASTA). The web portal is a collaboration between the Jornada LTER and the LTER Network Office. The following dataset from California Current Ecosystem (CCE) contains phosphate in coastal water measurements in microMolesPerLiter units and were aggregated to a yearly timescale.
California Current Ecosystem site, station Ohman Region: subset of CalCOFI stations inshore and nearshore in the Southern California Bight region; CalCOFI lines 80-93, stations from shore offshore to station 70, study of primary production, measured as carbon in units of gramsPerMeterSquaredPerYear on a yearly timescale
The EcoTrends project was established in 2004 by Dr. Debra Peters (Jornada Basin LTER, USDA-ARS Jornada Experimental Range) and Dr. Ariel Lugo (Luquillo LTER, USDA-FS Luquillo Experimental Forest) to support the collection and analysis of long-term ecological datasets. The project is a large synthesis effort focused on improving the accessibility and use of long-term data. At present, there are ~50 state and federally funded research sites that are participating and contributing to the EcoTrends project, including all 26 Long-Term Ecological Research (LTER) sites and sites funded by the USDA Agriculture Research Service (ARS), USDA Forest Service, US Department of Energy, US Geological Survey (USGS) and numerous universities. Data from the EcoTrends project are available through an exploratory web portal (http://www.ecotrends.info). This web portal enables the continuation of data compilation and accessibility by users through an interactive web application. Ongoing data compilation is updated through both manual and automatic processing as part of the LTER Provenance Aware Synthesis Tracking Architecture (PASTA). The web portal is a collaboration between the Jornada LTER and the LTER Network Office. The following dataset from California Current Ecosystem (CCE) contains primary production, measured as carbon measurements in gramsPerMeterSquaredPerYear units and were aggregated to a yearly timescale.
Point Count Bird Censusing Data Subset for Paper 'EFFECTS OF LAND USE AND VEGETATION COVER ON BIRD COMMUNITIES' Walker et. al
Animals utilize their environment across a range of scales, which is bounded by their extent, the broadest spatial area which organisms respond to their environment within their lifetime, and the spatial grain, the smallest area they respond to their environment (Kotlier and Wiens 1990). Within this range, organisms likely respond to their environment at a hierarchy of levels. Johnson (1980) recognizes four distinct levels of hierarchical habitat selection. At the very largest scale, first order selection, includes the entire area that an organism utilizes within its lifetime, and is also known as an organisms global home range or extent. In contrast, second order selection is an organisms local home range, or the area that it occupies within a unique ecosystem. This distinction is most apparent with migratory animals who utilize more than one distinct landscape for their survival (i.e. summer vs. winter feeding grounds), and much less so for organisms resident of one specific landscape for their entire life span. Third order selection is the selection of specific habitat patches within an ecosystem. For example, a Monarch butterfly would tend to select patches of milkweed within a prairie. And the lowest level, fourth order selection, involves the physical procurement of food within a selected patch, in our example, specific flowers within a milkweed patch, and is also known as grain. Realizing the importance of hierarchical habitat selection, it has become apparent that single-scale studies of animals responses to their environment may fail to adequately represent how that specific animal is responding to ecological parameter of interest, especially if they are not responding to the landscape at that scale (Holling 1992). The range of scales which an animal of interest is utilizing a landscape is important to determine prior to any further ecological investigation, as inappropriate scalar mismatch between organism and environment can lead to ambiguous or even dece
MeSDiCon subset for CodiEsp: MESH terms in MeSDiCon mapped to ICD10 CM and ICD10 PCS
<p>The MeSDiCon consists of a list or gazetteer of candidate names of diseases and symptoms mentioned in Spanish clinical texts. Thus MeSDiCon serves as a lexical resource or dictionary for automatic detection of disease/symptom mentions, as well as indexing or classification of medical texts with such concept types. Terms in MeSDiCon were mapped to MESH terminology.</p> <p>In this subset, we have mapped MESH codes to ICD10-CM and ICD10-PCS through UMLS Metathesaurus. Then, this resource contains diseases and symptoms terms from Spanish clinical texts mapped to MESH and ICD10.</p> <p> </p> <p><strong>Please cite if you use this dataset:</strong></p> <p>Antonio Miranda-Escalada, Aitor Gonzalez-Agirre, Jordi Armengol-Estapé and Martin Krallinger. Overview of automatic clinical coding: annotations, guidelines, and solutions for non-English clinical cases at CodiEsp track of CLEF eHealth 2020. In CLEF (Working Notes). 2020</p> <pre><code>@inproceedings{miranda2020overview, title={Overview of automatic clinical coding: annotations, guidelines, and solutions for non-english clinical cases at codiesp track of CLEF eHealth 2020}, author={Miranda-Escalada, Antonio and Gonzalez-Agirre, Aitor and Armengol-Estap{\'e}, Jordi and Krallinger, Martin}, booktitle={Working Notes of Conference and Labs of the Evaluation (CLEF) Forum. CEUR Workshop Proceedings}, year={2020} }</code></pre> <p> </p> <p><strong>File structure</strong></p> <p>TSV. Data is separated by tabs (\t). Every row of the file has the following fields:</p> <pre><code>terminology identifier translatedTerm termCount documentCount ICD10CM-code ICD10PCS-code</code></pre> <p>In case one MESH term is mapped to more than one ICD10 code, they are separated by commas.</p>
CNR global observation-based OMEGA3D quasi-geostrophic vertical and horizontal ocean currents (1993-2018): validation subset.
<p>A subset of the Global Ocean Multi Observations OMEGA3D product is provided here for comparison with independent in situ observations and model data. OMEGA3D product has been developed by the Consiglio Nazionale delle Ricerche in the framework of the Copernicus Monitoring Environment Marine Service (https://doi.org/10.25423/cmcc/multiobs_glo_phy_w_rep_015_007). It consists of 26 year of 3D quasi-geostrophic vertical and horizontal ocean currents (January 1993-December 2018) provided over a regular grid at 1/4° horizontal resolution, with 75 non-uniformly spaced vertical levels between the surface and 1500 m depth. The subset provided here includes horizontal velocities at 15 m and 1000 m depths (that can be used for validation with independent velocity estimates from SVP drifters and Argo floats displacements, respectively) and vertical velocities at 100 m depth (for comparison with model re-analyses). OMEGA3D product is delivered with a weekly sampling (representative of each Wednesday). The velocities are obtained by solving a Q-vector formulation of the Omega equation that explicitly considers the effect of both geostrophic advection and upper layer turbulent mixing as described by KPP parameterization. Omega forcings are estimated from the CMEMS observation-based ARMOR3D temperature and salinity multi-year data and ERA-Interim surface fluxes. </p>
Dataset related to article "Costimulatory Molecules and Immune Checkpoints Are Differentially Expressed on Different Subsets of Dendritic Cells."
<p>Dendritic cells (DCs) play a crucial role in initiating and shaping immune responses. The effects of DCs on adaptive immune responses depend partly on functional specialization of distinct DC subsets, and partly on the activation state of DCs, which is largely dictated by environmental signals. Fully activated immunostimulatory DCs express high levels of costimulatory molecules, produce pro-inflammatory cytokines, and stimulate T cell proliferation, whereas tolerogenic DCs express low levels of costimulatory molecules, produce immunomodulatory cytokines and impair T cell proliferation. Relevant to the increasing use of immune checkpoint blockade in cancer treatment, signals generated from inhibitory checkpoint molecules on DC surface may also contribute to the inhibitory properties of tolerogenic DCs. Yet, our knowledge on the expression of inhibitory molecules on human DC subsets is fragmentary. Therefore, in this study, we investigated the expression of three immune checkpoints on peripheral blood DC subsets, in basal conditions and upon exposure to pro-inflammatory and anti-inflammatory stimuli, by using a flow cytometric panel that allows a direct comparison of the activatory/inhibitory phenotype of DC-lineage and inflammatory DC subsets. We demonstrated that functionally distinct DC subsets are characterized by differential expression of activatory and inhibitory molecules, and that cDC1s in particular are endowed with a unique immune checkpoint repertoire characterized by high TIM-3 expression, scarce PD-L1 expression and lack of ILT2. Notably, this unique cDC1 repertoire was subverted in a group of patients with myelodysplastic syndromes included in the study. Applied to the characterization of DCs in the tumor microenvironment, this panel has the potential to provide valuable information to be used for investigating the role of DC subsets in cancer, guiding DC-targeting treatments, and possibly identifying predictive biomarkers for clinical response to cancer immunotherapy.</p>
Jaffe et al. 2020 CPR Genome Subset
<p>Subset of Candidate Phyla Radiation genomes associated with Jaffe et al. 2020 ("<strong>The rise of diversity in metabolic platforms across the Candidate Phyla Radiation")</strong>, in .fasta format.</p>
A Subset of CyberShake Ground Motion Time Series for Response History Analysis
<p>A subset of CyberShake numerically simulated ground motions that were selected and vetted for use in engineering response history analyses.</p> <p>v1.0.1: update readme file</p>
Student oriented subset of the Open University Learning Analytics dataset
<p>The Open University (OU) dataset is an open database containing student demographic and click-stream interaction with the virtual learning platform. The available data are structured in different CSV files. You can find more information about the original dataset at the following link: <a href="https://analyse.kmi.open.ac.uk/open_dataset">https://analyse.kmi.open.ac.uk/open_dataset</a>.</p> <p>We extracted a subset of the original dataset that focuses on student information. 25,819 records were collected referring to a specific student, course and semester. Each record is described by the following 20 attributes: <em> code_module, code_presentation, gender, highest_education, imd_band, age_band, num_of_prev_attempts, studies_credits, disability, resource, homepage, forum, glossary, outcontent, subpage, url, outcollaborate, quiz, AvgScore, count</em>.<br> <br> Two target classes were considered, namely Fail and Pass, combining the original four classes (Fail and Withdrawn and Pass and Distinction, respectively). The final_result attribute contains the target values.<br> <br> All features have been converted to numbers for automatic processing.<br> <br> Below is the mapping used to convert categorical values to numeric:</p> <ul> <li>code_module: 'AAA'=0, 'BBB'=1, 'CCC'=2, 'DDD'=3, 'EEE'=4, 'FFF'=5, 'GGG'=6</li> <li>code_presentation: '2013B'=0, '2013J'=1, '2014B'=2, '2014J'=3</li> <li>gender: 'F'=0, 'M'=1</li> <li>highest_education: 'No_Formal_quals'=0, 'Post_Graduate_Qualification'=1, 'HE_Qualification'=2, 'Lower_Than_A_Level'=3, 'A_level_or_Equivalent'=4</li> <li>IMBD_band: 'unknown'=0, 'between_0_and_10_percent'=1, 'between_10_and_20_percent'=2, 'between_20_and_30_percent'=3, 'between_30_and_40_percent'=4, 'between_40_and_50_percent'=5, 'between_50_and_60_percent'=6, 'between_60_and_70_percent'=7, 'between_70_and_80_percent'=8, 'between_80_and_90_percent'=9, 'between_90_and_100_percent'=10</li> <li>age_band: 'between_0_and_35'=0, 'between_35_and_55'=1, 'higher_than_55'=2</li> <li>disability: 'N'=0, 'Y'=1</li> <li>student's outcome: 'Fail'=0, 'Pass'=1</li> </ul> <p>For more detailed information, please refer to:</p> <p><br> Casalino G., Castellano G., Vessio G. (2021) Exploiting Time in Adaptive Learning from Educational Data. In: Agrati L.S. et al. (eds) Bridges and Mediation in Higher Distance Education. HELMeTO 2020. Communications in Computer and Information Science, vol 1344. Springer, Cham. <a href="https://www.google.com/url?q=https%3A%2F%2Fdoi.org%2F10.1007%2F978-3-030-67435-9_1&sa=D&sntz=1&usg=AFQjCNF7fUT9S4TcSpImSr4e_DjaLn3wtg">https://doi.org/10.1007/978-3-030-67435-9_1</a></p>
RAMP data subset, January 1 through May 31, 2019
<p>The data are a subset of data from RAMP, the Repository Analytics and Metrics Portal (<a href="http://ramp.montana.edu/">http://ramp.montana.edu/</a>), consisting of data from 35 (out of 50) participating institutional repositories (IR) from the period of January 1 through May 31, 2019. This subset represents data analyzed for a pending publication. For a description of the data collection, processing, and output methods, please see the "methods" section below.</p> <p>The 'RAMP Primer,' a Jupyter Notebook consisting of Python code for combining monthly data and generating some aggregate statistics is available from <a href="https://github.com/imls-measuring-up/ramp-documentation.git">https://github.com/imls-measuring-up/ramp-documentation.git</a>. The linked repository also includes data documentation similar to that provided in the "methods" section below, as well as a file of IR index names useful for subsetting and filtering the data.</p> <p><b>Update: </b>Version two of this dataset was uploaded on January 14, 2020. Thanks to RAMP participants, the RAMP administrators discovered an error in the daily data harvest that resulted in incomplete page-click data for roughly 15 RAMP participating repositories. Only page-click data as described below were affected, and the corrsponding CSV files have been replaced with corrected data. The country-device data were not affected and have not been changed from version 1.</p>
Data from: Generalist haemosporidian parasites are better adapted to a subset of host species in a multiple host community
Parasites that can infect multiple host species are considered to be host generalists with low host specificity. However, whether generalist parasites are better adapted to a subset of their host species remains unknown. To elucidate this possibility, we compared the variation in prevalence and infection intensity among host species of three generalist parasite lineages belonging to the morphological species Haemoproteus majoris, in a natural bird community in southern Sweden. Prevalence in each host species was confirmed by nested PCR and DNA sequencing and infection intensities were quantified using lineage-specific real-time qPCR. For two of the three lineages, we detected positive correlations between prevalence and infection intensity, indicating that these generalist parasites are better adapted to a subset of host species, which may have been more frequently encountered during the evolution of the parasite; we refer to these as main host species. For both lineages, the main host species were more phylogenetically related than expected by chance as revealed by strong phylogenetic signal in prevalence among hosts. By comparing our results with previous records of these parasites, we found that the host range of a generalist parasite can vary among different communities and may partly be shaped by the presence of other parasites. Our study reveals that generalist parasites may be specialized on a subset of their host species and it highlights the importance of considering infection intensity and host phylogeny when determining the host specificity of a parasite.
Tri-modal single cell profiling reveals a distinct pediatric CD8αα T cell subset and broad age-related molecular reprogramming across the T cell compartment
<p>Age-associated changes in the T cell compartment are well-described. However, limitations of current single- or bi-modal single-cell assays, including flow cytometry, RNA-seq, and CITE-seq, have restricted our ability to deconvolve more complex cellular and molecular changes. Here, we profile more than 300,000 single T cells from healthy children (11–13 yrs) and older adults (55–65 yrs) using the trimodal assay, TEA-seq (protein, RNA, and chromatin accessibility), revealing that molecular programming of T cell subsets shifts toward a more activated basal state with age. Naive CD4 T cells, considered relatively resistant to aging, exhibited pronounced transcriptional and epigenetic reprogramming. Moreover, we discovered a novel CD8aa T cell subset lost with age that is epigenetically poised for rapid effector responses and displays distinct inhibitory, co-stimulatory and tissue homing properties. Together, these data reveal new insights into age-associated changes in the T cell compartment that may contribute to differential immune responses.</p>
Subset of instances used in article "On solving the 1.5-dimensional cutting stock problem with heterogeneous slitting lines allocation in the steel industry"
<p>The dataset presented is part of the one used in the article "On solving the 1.5-dimensional cutting stock problem with heterogeneous slitting lines allocation in the steel industry" by María Sierra-Paradinas, Óscar Soto-Sánchez, Antonio Alonso-Ayuso, F. Javier Martín-Campo and Micael Gallego, in Computers & Industrial Engineering (2024), doi: <a title="Persistent link using digital object identifier" href="https://doi.org/10.1016/j.cie.2024.110120" target="_blank" rel="noreferrer noopener"><span>https://doi.org/10.1016/j.cie.2024.110120</span></a>.</p> <p>This paper proposes a mathematical optimisation model for a cutting stock problem in the steel industry. This problem appears in a Spanish company and the proposed model has been tested on real orders received by the company.</p> <p>The dataset presented here includes twelve instances used in the paper (the rest cannot be presented for confidentiality reasons). For each instance, the characteristics of the order and the solution obtained by the model are provided.</p>
Massive Compression for High Data Rate Macromolecular Crystallography (HDRMX): Impact on Diffraction Data and Subsequent Structural Analysis: Subset with data from 2 deposited PDBs.
<p>Diffraction data from a lysozyme crystal. Data collected at 7.5 keV at the AMX beamline, NSLS-II. 360 degrees were collected, with 0.2 deg per frame. This data set contains 2 folders; 1 from uncompressed data and 1 from data compressed using lossy compression as follow: frames were summed (2x), pixels were binned (2x) and Hcompress with level 24 was applied to uncompressed data. </p>
dogs_vs_cats_subset_kaggle
<p>This is a subset of the well known image classification dataset for cats and dogs made by kaggle.</p> <p>https://www.microsoft.com/en-us/download/details.aspx?id=54765</p> <p>This dataset contains 1000 training images (500 cats & 500 dogs), 200 validation images (100 cats/100 dogs) and 100 unlabelled test images.</p> <p>There are other similar dataset already on zenodo like this https://zenodo.org/doi/10.5281/zenodo.5226944, but that dataset is not balanced even though it claims to be.</p> <p> </p> <p> </p>
Raw IMC files for Spatial subsetting enables integrative modeling of oral squamous cell carcinoma multiplex imaging data.
<p>Raw MCD files for the Stanford cohort of oral squamous cell carcinoma patients in this publication:</p> <p>Spatial subsetting enables integrative modeling of oral squamous cell carcinoma multiplex imaging data (DOI:<span> <a href="https://doi.org/10.1016/j.isci.2023.108486" target="_blank" rel="noopener">10.1016/j.isci.2023.108486</a>).</span></p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.