Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
753
datasets available to search
ShareScore release 0.9.0
Dataset results
753 results for “metrics”
SURF Machine Metric Dataset [2019-12-29, 2020-08-07]
<p>An addition to our earlier <a href="https://zenodo.org/record/3878143">uploaded dataset.</a></p> <p>This dataset contains an update of Prometheus: it now contains 327 metrics and spans more than 7 months. Note that the rack and node names have changed with respect to the previous dataset and invalid values are now set to NaN to better indicate what is an invalid value and what is a true zero value.</p>
A Tool for Uncertainty Quantification in Reconstructing Sparse Water Quality Time Series Data to Assess Risk Metrics for Watershed Health and TMDL Analysis
<p>The uploaded file contains the input and output data which can be used to reproduce the results in the research article 'Uncertainty Quantification in Reconstruction of Sparse Water Quality Time Series: Implications for Watershed Health and Risk-Based TMDL Assessment'. Please refer to the file '<a href="https://zenodo.org/api/files/31b59cce-8eb2-4ee7-93aa-61474c6f6359/dst_2019_SJRW_TP_TDS.zip?versionId=2af2b54d-d5fb-4720-919d-de2e827595e2">dst_2019_SJRW_TP_TDS.zip'</a> for updated files..</p>
A dataset used to determine a semantic similarity metric based on UMLS for PMC-OA
<p>We have performed a series of in-silico experiments in order to determine a semantic similarity metric based on UMLS annotations for PubMed Central Open Access. Here we have stored the data used for and obtained from such experiments. We have worked with relevant and partially relevant articles from the TREC-2005 Genomics Track Collection, from now referred as the initial collection, including a total of 4240 unique PubMed articles. From those 4240 articles, only 62 had publicly available; those 62 articles correspond to the full-text collection.</p> <p>Our data comprises flat files using tabs as separators and one Excel sheet. Tab separated values always include a first row with headings:</p> <ul> <li>Stems extracted from title and abstract for articles in the initial collection. Each row contains a stem with its inverse-document-frequency (IDF) within the initial collection. Stems were calculated following the Porter algorithm (available at http://tartarus.org/martin/PorterStemmer/java.txt) <ul> <li>stems.TA.tsv</li> </ul> </li> <li>Article profiles, i.e., terms (either word stems or UMLS concepts) found in the articles with term frequency (TF) and IDF. The first two columns correspond to PubMed Identifier (PMID) and PubMed Central identifier (PMC). PMC identifier was set to 0 whenever full-text was not available. <ul> <li>profiles.TA.tsv: Profiles according word stems in title and abstract for the initial collection</li> <li>profiles.PMID.tsv: Profiles according to UMLS concpets in title and abstract for the initial collection</li> <li>profiles.PMC_TA.tsv: Profiles according to UMLS concepts in title and abstract for the full-text collection</li> <li>profiles.PMC.tsv: Profiles according to UMLS concepts in the full-text for the full-text collection</li> </ul> </li> <li>Similarity matrixes calculated on the article profiles with PubMed Related Article metric (PMRA), BM25, and Cosine. There are matrixes for terms found in title-and-abstract as well as full-text. In a similarity matrix, a reference article (an interest has been already expressed for it) correspond to a row, while the columns correspond to all the other articles for which the similarity was calculated. <ul> <li>Matrixes for our initial collection <ul> <li>similarity.PMRA.TA.profiles.TA.tsv: Similarity matrix for profiles.TA.tsv following the algorithm PMRA. This matrix is considered the baseline for further analyses</li> <li>similarity.PMRA.profiles.PMID.tsv: Similarity matrix for profiles.PMID.tsv following the algorithm PMRA</li> <li>similarity.BM25_1.2_0.75.profiles.PMID.tsv: Similarity matrix for profiles.PMID.tsv following the algorithm BM25 with k=1.2 and b=0.75</li> <li>similarity.COSINE.profiles.PMID.tsv: Similarity matrix for profiles.PMID.tsv following the algorithm Cosine</li> </ul> </li> <li>Matrixes for our full-text collection <ul> <li>similarity.PMRA.profiles.PMC_TA.tsv: Similarity matrix for profiles.PMC_TA.tsv following the algorithm PMRA</li> <li>similarity.PMRA.profiles.PMC.tsv: Similarity matrix for profiles.PMC.tsv following the algorithm PMRA</li> <li>similarity.BM25.profiles.PMC_TA.tsv: Similarity matrix for profiles.PMC_TA.tsv following the algorithm BM25 with k=1.2 and b=0.75</li> <li>similarity.BM25.profiles.PMC.tsv: Similarity matrix for profiles.PMC.tsv following the algorithm BM25 with k= 1.2 and b= 0.75</li> <li>similarity.COSINE.profiles.PMC_TA.tsv: Similarity matrix for profiles.PMC_TA.tsv following the algorithm Cosine</li> <li>similarity.COSINE.profiles.PMC.tsv: Similarity matrix for profiles.PMC.tsv following the algorithm Cosine</li> </ul> </li> </ul> </li> <li>Correlation matrixes for similarities calculated for title-and-abstract taking as reference the similarity values obtained with PMRA for word stems on title-and-abstract. <ul> <li>pearsonCorrelation.PMRA.tsv: Correlation for similarity.PMRA.profiles.PMID.tsv</li> <li>pearsonCorrelationTopic.PMRA.tsv: Correlation for similarity.PMRA.profiles.PMID.tsv discriminated by TREC topics</li> <li>pearsonCorrelation.BM25_1.2_0.75.tsv: Correlation for similarity.BM25_1.2_0.75.profiles.PMID.tsv</li> <li>pearsonCorrelationTopic.BM25_1.2_0.75.tsv: Correlation for similarity.BM25_1.2_0.75.profiles.PMID.tsv discriminated by TREC topics</li> <li>pearsonCorrelation.COSINE.tsv: Correlation for similarity.COSINE.profiles.PMID.tsv</li> <li>pearsonCorrelationTopic.COSINE.tsv: Correlation for similarity.COSINE.profiles.PMID.tsv discriminated by TREC topics</li> </ul> </li> <li>Precision and recall summaries for the similarities calculated based on title-and-abstract. <ul> <li>StatsAllSummary.xlsx: Precision and recall at a global level, i.e., without considering TREC topics. This file includes information for BM25 with multiples values for constants k and b</li> </ul> </li> </ul> <p>Visualization for correlation matrixes as well as scattered plots for full-text based similarity is available at http://ljgarcia.github.io/semsim.benchmark</p>
Cumulative Data-Level Metrics (DLM) Report - September 2015
<p>Article-Level Metrics (ALM) measure the reach and online engagement of scholarly works. This DataSite Data-Level Metrics (DLM) Server report contains the cumulative stats collected for all works through September 11, 2015. Data are generated by the Lagotto open source software. Go to the Lagotto forum for questions or comments.</p>
WikiRate metric values dataset
<p>WikiRate is a collaborative platform that enables everyone to contribute publicly available environmental, social and governance (ESG) data. This dataset contains all metric values that have been contributed to the platform.</p>
SIMpat: a synthetic benchmark for similarity metrics on patient representations
<div> <h2>Introduction</h2> <p>We used Synthea to generate six cohorts of patients with certain specified disease. Please refer to Synthea documentation for the generation process.</p> <p>We selected 6 different diseases that could be generated by Synthea, that were deemed by a medical professional as “different enough”. The goal of this simulation is to find a metric that can differentiate between patients.</p> <p>The six conditions are:</p> <ul> <li>Cerebral Palsy (SNOMED-CT code : 128188000 - Cerebral palsy (disorder))</li> <li>Colorectal Cancer (SNOMED-CT code : 93761005 - Primary malignant neoplasm of colon (disorder))</li> <li>Dialisys (SNOMED-CT code : 265764009 - Renal dialysis (procedure))</li> <li>Hypertension (SNOMED-CT code : 59621000 - Essential hypertension (disorder))</li> <li>Breast Cancer (SNOMED-CT code : 254837009 - Malignant neoplasm of breast (disorder))</li> <li>Prostate Cancer (SNOMED-CT code : 126906006 - Neoplasm of prostate (disorder))</li> </ul> <p>NB:</p> <ul> <li>Dialisys is not a disorder, but a condition, but is used here as a proxy for renal issue</li> <li>Synthea doesn’t have a module to generate prostate cancer in men, but only prostate cancer in veteran, hence this is the module used here (all men with prostate cancer are veterans)</li> </ul> <p>We use those cohort to compare the ability of 12 different distance metrics to separate patients.</p> <p>Those 12 metrics are split in three groups :</p> <p>Sementic based metrics:</p> <ul> <li>AvgEmb* method encodes text by averaging the pre-trained word embeddings of all the words present in it.</li> <li>BERT* uses bidirectional transformer based neural model to solve the task of masked language modeling.</li> <li>Universal Sentence Encoders (USE)* use transformer based encoders to encode sentences into embedding vectors.</li> <li>Embeddings from Language Models (ELMo)* uses bi-directional LSTM based encoders to encode a sentence into a fixed size representation</li> </ul> <p>Graph based metrics:</p> <ul> <li>DeepWalk* uses random walks to generate sequences of vertices (vertex sentences) which are subsequently fed to a skip-gram model to learn the embeddings corresponding to the vertices.</li> <li>Node2Vec* uses biased random walks to optimize a neighborhood preserving objective function such that the nodes which are highly interconnected and the nodes with similar roles in the graph are closer in the embedding space.</li> <li>LINE* tries to directly optimize the vertex embeddings based on one hop and two hop random walk probabilities.</li> <li>HARP* proposes a meta-strategy for embedding vertices of a graph such that they preserve the higher-order structural features.</li> <li>Bags of findings^</li> <li>Average Links^</li> <li>Average Links Weighted by Information Content (IC)^</li> <li>Path Distance weighted by IC^</li> </ul> <p>Concept followed by a * are extracted from <a href="https://proceedings.mlr.press/v116/pattisapu20a/pattisapu20a.pdf">this paper</a> and can be downloaded <a href="../records/3842143">here</a></p> <p>Concept followed by a ^ were develloped by Jean-Virgile Voegeli (SIMED)</p> </div> <div> <div> </div> <h2>Descriptive analysis of the sample</h2> <p>We will first look at the cohort that were created by Synthea. The cohorts were created using the seed 123456789 for reproducibility.</p> <p>For this first experiment, Synthea was asked to generate 100 alive individuals for each specific disease. We asked Synthea to keep only 10 years of history. Each individual was set to be between the age of 18 and 80 years old. Except for specific sex-disease such as breast cancer and prostate cancer, all cohorts contains both male and female individuals. We used the default location, which is Massachussetts.</p> <p>One important note on age. The Synthea modules sometimes specify a minimum age to onset a certain condition / disease. For example, colorectal cancer can only onset after 50 years old, and prostate cancer after 60 years old.</p> <p>Each Synthea run was set to run 10.000 times. If after 10.000 tries, the software didn’t manage to generate a patient that fit the criterion (here, a specific snomed code), the run would fail. Synthea can also generate patients that dies before the “run date”, and if this happens will simulate another patient.</p> <p>This explains why we have cohorts of more than 100 individuals but less than 100 alive individuals. We can also have in certain cases a little above 100 individuals. This is due to the fact that the synthea generator is multicore, and patients are generated simultaneously.</p> </div>
Local optima network metrics from the IEEE CEC 2024 paper "Information flow and Laplacian dynamics on local optima networks"
<p>Local optima network metrics from the IEEE CEC 2024 paper "Information flow and Laplacian dynamics on local optima networks". </p> <p>There are two CSV files: one for each of the two iterated local search confgurations used to construct the networks (low or high). In each file, a row contains information about one QAPLIB instance. Easch row contains all the metrics computed for the associated LON and also algorithm performance data on the instance. </p>
Data from: macrofaunal diversity patterns in coastal marine sediments: re-examining common metrics and methods
<p>Complex biodiversity patterns arise in marine systems due to overlapping ecological processes, including organism interactions, resource distribution, and environmental conditions. Despite the importance of documenting these patterns, describing diversity in natural ecosystems remains challenging. Here, we investigate three nearshore sub-Arctic sites to describe benthic macroinfaunal taxa and biological traits, with the ultimate aim of determining whether common diversity metrics and typical sampling efforts adequately capture community composition in these systems. First, we assess how diversity relates to sediment depth, and examine relationships among commonly used taxonomic and functional diversity indices. Second, using a power analysis, we explore how sampling effort influences the interpretation of diversity patterns in coastal systems. We report significant variation in community composition among sites, even across small spatial scales of kilometers, and find that taxonomically diverse communities do not necessarily correspond to high functional diversity. We further find that although environmental factors such as sediment depth consistently affect macroinfaunal diversity, the direction and magnitude of these relationships are site dependent. Finally, we demonstrate that typical sampling effort for coastal benthic studies may not capture macroinfaunal community composition adequately, potentially obscuring hotspots in common diversity metrics such as taxonomic or functional richness. Conversely, indices such as Simpson's diversity may be well-suited to resource-limited studies with restricted sampling capacity. Our results highlight the importance of adopting a multi-pronged approach to biodiversity assessment and determining optimal sample sizes for a wide range of marine benthic systems, particularly in the context of biodiversity monitoring for conservation purposes.</p>
Piecewise continuous sampling: a method for minimizing bias and sampling effort for estimated metrics of animal behavior
<p>Capturing qualitative features of animal behavior requires recording occurrences of behavior over time. Continuous sampling is best for capturing brief behaviors, but can be very time consuming. Instantaneous sampling can reduce the amount of labor required, but can miss short-duration behaviors. We therefore synthesized these techniques by continuously sampling during randomly scattered time intervals; a technique we call piecewise continuous sampling. To optimize and test the efficacy of this technique, we collected a continuous behavioral dataset of harvester ant workers, and then we developed a protocol to estimate the amount of sampling time necessary to reconstruct the proportion of time animals spend in different behavioral states. This protocol finds the sample size needed for the variance of the sample to converge on the variation of the population. We then divided this estimated time into equal-duration intervals that were randomly distributed across the entire continuous dataset. Finally, we calculated both time-dependent and time-independent error from this sample. We found that 4 to 16 sampling intervals minimize both types of error simultaneously. This finding was robust to differences in underlying behavior and was validated with simulations, implying that this method could be used for many types of organisms.</p>
Supplementary data frames, AlphaFold models, Normal Mode Analysis (NMA) Data, and NMA of Corresponding NMR Ensembles in the S2RCI, MD, and S2 Datasets for "Gradations in protein dynamics captured by experimental NMR are not well represented by AlphaFold2 models and other computational metrics"
<h1><strong>Changes applied to V2</strong></h1> <p>In addition to the supplementary dataframes and AlphaFold models from each dataset in V1, V2 includes the additional data outlined below.</p> <p>The <strong>S2RCI</strong> and <strong>MD</strong> datasets include comprehensive analyses of AlphaFold2 models (both before and after truncation). These datasets feature: </p> <ul> <li><strong>AlphaFold2 Models</strong>: Both original and truncated structures. </li> <li><strong>WEBnma Modes</strong>: `modes.txt` files generated from WEBnma analysis, available for both non-truncated and truncated AF2 models. </li> <li><strong>Root-Mean-Square-Fluctuations (RMSF)</strong>: Profiles calculated before and after truncation of AF2 models. </li> <li><strong>NMR Data: Normal Mode Analysis (NMA)</strong>: Performed on corresponding NMR ensembles (see below). </li> </ul> <p> </p> <p>The <strong>NMR Data</strong> of NMA in these datasets includes: </p> <ul> <li>NMR ensembles </li> <li>Individual NMR models extracted from each ensemble </li> <li>STRIDE secondary structure calculations per-individual NMR models</li> <li>RMSF profiles per-individual NMR models</li> </ul> <p>For detailed information, please refer to the `Readme.txt` file within each corresponding folder. </p> <p>The <strong>S2 dataset</strong> includes all the features listed above, except for the NMR analysis.</p>
Figure 2 in Examining metrics and magnitudes of molecular genetic differentiation used to delimit cetacean subspecies based on mitochondrial DNA control region sequences
Figure 2. Relationship between ΦST and Nei's estimate of net divergence (dA) among cetacean population, subspecies, and species pairs estimated using mitochondrial DNA control region sequence data. Specific values mentioned in the text are numbered: 1 = Neophocaena species; 2 = killer whale populations. The three green squares in the left-hand side of the figure (ΦST <0.07) represent, from bottom to top, the subspecies comparisons for S. attenuata, S. longirostris, and L. obscurus, respectively.
Figure 1 in Examining metrics and magnitudes of molecular genetic differentiation used to delimit cetacean subspecies based on mitochondrial DNA control region sequences
Figure 1. Box and whisker plots showing median and 1st and 3rd quartiles, and minimum and maximum values for six metrics of genetic divergence among cetacean population, subspecies, and species pairs estimated using mitochondrial DNA control region sequence data.
SSL Metrics Datasets
<p>SSL Metrics datasets mined from GitHub research software projects of interest.</p>
Latin American and Caribbean journals indexed in Google Scholar Metrics
<p>Dataset from a study aiming to analyze the coverage of Latin American and Caribbean journals in Google Scholar Metrics (GSM). Data from 8,205 journals from 24 countries of the region were downloaded from Latindex database. A Python script was used for automated title search and data extraction (titles, h5-index, h5-median, URLs) in GSM. For the journals not found, a manual search was carried out, with attempts by variations of the title. It was found 3,070 journals indexed in GSM, which corresponds to 37.42% of the Latindex list. The search was performed on the 2021 edition of GSM, which considers articles published between 2016 and 2020 and citations registered until July 2021. The number of all types of documents published (productivity) in the h5-index period (2016-2020) in Scopus, Journal Citation Reports, and SciELO of 1,314 journals was also identified. </p> <p>The present dataset is the result of this study, which is under peer-review in a scientific journal. </p> <p>The dataset comprises titles, h5-index; h5-median, URLs of 3,070 publications from Latin America and the Caribbean identified in Google Scholar Metrics, and the respective editorial information of the publications was extracted from Latindex</p> <p>The original language of the content was kept, mainly Spanish in the case of editorial data from Latindex. The columns descriptors are also shown in English.</p> <p>The productivity data refer to the number of all types of documents published by the journals in the period 2016-2020. Data were extracted from the InCities Journal Citation Reports, Scopus, and SciELO Citation Index (Web of Science database).</p> <p>In this version 2, only the productivity data were changed, covering a larger number of journals (1,314) and including all types of documents. Other data are the same as in the first version (https://doi.org/10.5281/zenodo.5572873).</p> <p> </p> <p> </p> <p> </p> <pre> </pre> <p> </p>
Supporting dataset for the paper : " Hydro-geomorphic metrics for high resolution fluvial landscape analysis"
<p>This repository contains all the original data supporting the results of Bernard et al., 2021: "Consistent hydro-geomorphic indicators for high resolution topographic analysis".<br> The parameter used to perform hydraulic simulations are also available.<br> </p>
Dataset#1 and Dataset#2 for Making drawings speak through mathematical metrics
<p>Dataset 1 and Dataset 2 for the paper Making drawings speak through mathematical metrics</p> <p>Figurative drawing is a skill that takes time to learn, and evolves during different childhood phases that begin with scribbling and end with representational drawing. Between these phases, it is difficult to assess when and how children demonstrate intentions and representativeness in their drawings. The marks produced are increasingly goal-oriented and efficient as the child’s skills progress from scribbles to figurative drawings. Pre-figurative activities provide an opportunity to focus on drawing processes. We applied fourteen metrics to two different datasets (N=65 and N=345) to better understand the intentional and representational processes behind drawing, and combined these metrics using principal component analysis (PCA) in different biologically significant dimensions. Three dimensions were identified: efficiency based on spatial metrics, diversity with colour metrics, and temporal sequentiality. The metrics at play in each dimension are similar for both datasets, and PCA explains 77% of the variance in both datasets. These analyses differentiate scribbles by children from those drawn by adults. The three dimensions highlighted by this study provide a better understanding of the emergence of intentions and representativeness in drawings. We have already discussed the perspectives of such findings in Comparative Psychology and Evolutionary Anthropology.</p>
Lagrangian Sequestration Efficiency Trajectories and Extracted Particle Metrics – 2000m Y1 & Y2
<p>A dataset of Lagrangian trajectories used to estimate North Atlantic sequestration efficiency and extracted metrics for the re-entrained and sequestered particles. All variables have long names and units. These files have been used for the analysis in Baker et al. ‘Biological carbon pump sequestration efficiency in the North Atlantic: a leaky or a long-term sink?’ with further information about the methodology available in the paper. Due to the size of the datasets, each DOI only contains two files. This dataset contains the 2000m particles releases for the years 1996 (Y1) and 1997 (Y2).</p>
multi-model metrics of daily sea-ice concentrations from CMIP6 models
<p>Multi-model metrics of daily mean sea-ice concentrations from 9 CMIP6 models:</p> <ul> <li>ACCESS-CM2</li> <li>CanESM5</li> <li>CNRM-ESM2.1</li> <li>EC-Earth3</li> <li>EC-Earth3-Veg</li> <li>IPSL-CM6A-LR</li> <li>MIROC6</li> <li>MRI-ESM2</li> <li>NorESM2-LM</li> </ul> <p> </p> <p>Metrics are mean, median, standard deviation, minimum and maximum sea-ice concentrations.</p>
Dataset - Associations between author-level metrics in subsequent time periods
<p>Dataset used in the project https://github.com/carolmb/associations_between_author_level_metrics.</p> <p>Author-level metrics are calculated in sequential 5-years-windows and subsequent 3-years-windows, more details are described in https://doi.org/10.1016/j.joi.2021.101218.</p>
Diverse Topologies for Evaluation of Geometric Similarity Metrics
<p>A collection of 7 datasets with each set containing 3D shapes with varying topological complexity. The datasets can be used to compare different metrics of geometric dissimilarity. Two of the datasets have topologically complex shapes that resemble designs obtained from topology optimization, a widely used design optimization method for engineering structures.</p> <p>We used this dataset for a related journal article with the following abstract: "In the early stages of engineering design, multitudes of feasible designs can be generated using structural optimization methods by varying the design requirements or user preferences for different performance objectives. Data mining such potentially large datasets is a challenging task. An unsupervised data-centric approach for exploring designs is to find clusters of similar designs and recommend only the cluster representatives for review. Design similarity can be defined not only on a purely functional level but also based on geometric properties, such as size, shape, and topology. While metrics such as chamfer distance measure the geometrical differences intuitively, it is more useful for design exploration to use metrics based on <em>geometric features</em>, which are extracted from high-dimensional 3D geometric data using dimensionality reduction techniques. If the Euclidean distance in the <em>geometric features</em> is meaningful, the features can be combined with performance attributes resulting in an aggregate feature vector that can potentially be useful in design exploration based on both geometry and performance. We propose a novel approach to evaluate such derived metrics by measuring their similarity with the metrics commonly used in 3D object classification. Furthermore, we measure clustering accuracy, which is a state-of-the-art unsupervised approach to evaluate metrics. For this purpose, we use a labeled, synthetic dataset with topologically complex designs. From our results, we conclude that Pointcloud Autoencoder is promising in encoding geometric features and developing a comprehensive design exploration method."</p> <p>For each dataset, shapes/designs are saved as surface mesh files (extension: stl) and point cloud files (extension: ply) in the folders "stls" and "plys" respectively. A brief description of the 7 different datasets is in the following table. For each dataset, the designs are named using numbers starting from 0, e.g., “0.stl, 1.stl, …, 19.stl” in the folder for the surface mesh files. Some of the datasets are labeled, i.e., each design belongs to a class. In a labeled dataset, all classes have the same number of designs, and the designs are named in the order of their class. For example, a labeled dataset with 4 designs and 2 classes contains files whose names start with {0, 1, 2, 3} where the designs {0, 1} belong to class 1, and {2, 3} belong to class 2.</p> <table> <thead> <tr> <th scope="col">Dataset name</th> <th scope="col">Directory name</th> <th scope="col">Number of designs</th> <th scope="col">Number of classes</th> </tr> </thead> <tbody> <tr> <td>Beam-rotation</td> <td>"rotate_beam"</td> <td>20</td> <td>None</td> </tr> <tr> <td>Beam-elongation</td> <td>"elongate_beam"</td> <td>20</td> <td>None</td> </tr> <tr> <td>Beam-translation</td> <td>"move_beam"</td> <td>20</td> <td>None</td> </tr> <tr> <td>Three cube trusses</td> <td>"three_cube_truss"</td> <td>150</td> <td>6</td> </tr> <tr> <td>Single cube trusses</td> <td>"single_cube_truss"</td> <td>275</td> <td>11</td> </tr> <tr> <td>Random topologies</td> <td>"three_cube_truss_random"</td> <td>1000</td> <td>50</td> </tr> <tr> <td>Topologically optimized designs</td> <td>"cube_opt_shapes"</td> <td>1500</td> <td>None</td> </tr> </tbody> </table>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.