Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

6,623

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

6,623 results for “generalization”

Learn how ShareScore rates datasets ↗
edi56/100

North Temperate Lakes LTER General Lake Model Parameter Set for Lake Mendota, Summer 2016 Calibration

The General Lake Model (GLM), an open source, one-dimensional hydrodynamic model, was used to simulate various physical, chemical, and biological variables on Lake Mendota between 15 April 2016 and 11 November 2016. GLM (v.2.1.8) was coupled to the Aquatic EcoDynamics (AED) module library via the Framework for Aquatic Biogeochemical Modeling (FABM). GLM-AED requires four major “scripts†to run the model. First, the glm2.nml file configures lake metadata, meteorological driver data, stream inflow and outflow driver data, and physical response variables. Second, the aed2.nml file configures various biogeochemical modules for the simulation of oxygen, carbon, phosphorus, and nitrogen, among others. Third, aed2_phyto_pars.nml configures all parameters pertaining to phytoplankton dynamics. And fourth, aed2_zoop_pars.nml configures all parameters pertaining to zooplankton dynamics. This dataset contains parameter descriptions and values as they were used to simulate organic carbon and greenhouse gas production on Lake Mendota in summer 2016. Meteorological data and stream files used in this calibration are also included in this dataset. Additional methods and model descriptions can be found in J.A. hart’s Masters Thesis, University of Wisconsin-Madison Center for Limnology, May 2017. Readers are referred to the GLM (Hipsey et al. 2014) and AED (Hipsey et al. 2013) science manuals for further details on model configuration.

openCC (other)Dec 2022View details →
OpenNeuro52/100

Modeling an auditory stimulated brain under altered states of consciousness using the generalized ising model

Open the record for dataset details and reuse information.

openCC0Jan 2021View details →
zenodo52/100

Gold-Caps_LMD-Matched_General

<p>This dataset contains captions for the <a href="https://colinraffel.com/projects/lmd/">Lakh MIDI Dataset-matched</a> music dataset (~30,000 tracks with accompanying MIDI files).</p><p>These captions were generated by the <strong>gpt-4-1106-preview</strong> chat endpoint prompted to describe each track based on the track title and artist. The captions have not been filtered or post-processed in any way.</p><p><strong>Prompt used:</strong><br>"Give a general description of the track &lt;title&gt; by &lt;artist_name&gt; in one sentence. Don't mention the title or artist."</p>

opencc-by-4.0Nov 2023View details →
zenodo52/100

Datasets for: Generalizing Monin-Obukhov Similarity Theory (1954) for Complex Atmospheric Turbulence, Stiperski and Calaf 2023, PRL

<p>Scaling variables for the generalized flux-variance scaling relations that include turbulence anisotropy. Dataset is a companion to the manuscript &nbsp;Stiperski, I., Calaf, M., 2023: Generalizing Monin-Obukhov similarity theory (1954) for complex atmospheric turbulence. Physical Review Letters, 130 (12), 124001,&nbsp; &nbsp;https://doi.org/10.1103/PhysRevLett.130.124001</p> <p>The dataset contains the turbulence statistics from 13 datasets:&nbsp; AHATS, Cabauw, CASES-99, METCRAX II campaign (NEAR&nbsp; and RIM towers), T-Rex campaign (Central tower - TRexC, West tower - TRexW) and i-Box measurement network (CCS-VF0 tower - i-Box0, CS-SF1 tower - i-Box1, CS-NF10 tower - i-Box10, CS-NF27 tower - i-Box27, CS-MT21 tower - i-BoxTop, im Hinteren Eis tower - imHint).</p> <p><br>Data are organized in csv files for each datasets and only contain high quality (for applied criteria see the Supplemental Material of the companion paper, https://journals.aps.org/prl/supplemental/10.1103/PhysRevLett.130.124001) data with 30 min averaging for unstable stratification and 1 min for stable stratification. Since the data were used for scaling, there is no reference to time, but the measurement height is provided as an additional variable.&nbsp;</p> <p>Meaning of variables:</p> <p>zeta - z/L where z is height above ground and L is the local Obukhov length</p> <p>SigmaU - $\overline{u'u'}/u_*$ scaled standard deviation of streamwise velocity, where $u_*$ is the local friction velocity</p> <p>SigmaU - $\overline{v'v'}/u_*$ scaled standard deviation of spanwise velocity</p> <p>SigmaU - $\overline{v'v'}/u_*$ scaled standard deviation of surface-normal velocity</p> <p>SigmaT - $\overline{T'T'}/T_*$ scaled standard deviation of sonic temperature, where $T_*$ is the local temperature scale</p> <p>SigmaEpsU - scaled dissipation rate of the streamwise velocity</p> <p>SigmaEpsW - scaled dissipation rate of the surface-normal velocity&nbsp;</p>

opencc-by-4.0Mar 2024View details →
zenodo52/100

Gender codification of 412 national (general) and European Parliament elections in six European countries (2003-2021)

<p>This dataset has been produced by applying the Manifesto Gender Analysis (MGA) codebook to 412 national (general) and European Parliament elections in the six countries participating in the UNTWIST project (Denmark, Germany, Hungary, Spain, Switzerland, and the UK) from 2003 to 2021.</p> <p>&nbsp;The Manifesto Gender Analysis coding procedure, developed by WP4 of the UNTWIST consortium, aims to analyse gender-related content in party manifestos. It relies on existing manifestos collected by MARPOR and EM projects from 2003-2021 in six national contexts: Denmark, Germany, Hungary, Spain, Switzerland, and the United Kingdom. The process involves splitting manifestos into quasi-sentences, coding them based on a scheme inspired by previous projects and feminist typology, and completing an expert survey. This method ensures comprehensive analysis and potential scalability through computational methods.&nbsp;</p> <p>The coding procedure involves a series of essential steps, divided in two main activities: the classification of manifestos&rsquo; quasi-sentences, and the completion of a survey dedicated to more general concepts which can be gauged by evaluating the content of the entire documents. In the latter case, then, the unit of measure of each coder consists in the manifesto document, whereas in the former the units of measure are quasi-sentences - i.e., arguments denoting a verbal expression of a political idea or issue. Coders are instructed to split sentences containing multiple arguments into quasi-sentences and ensure that each quasi-sentence encapsulates a single political idea or issue.&nbsp;</p> <p>Once the manifestos are split into said units, coders classify the arguments following the MGA coding scheme. The coding scheme (MGA) consists of 5 domains and 25 coding categories, covering various aspects of gender-related issues. Each domain includes an "other" category for relevant statements that do not fit precisely into the defined categories. Apart from coding categories related to specific themes, the coding scheme then includes additional dimensions. The classification process consists of seven steps: (1) assessing whether the quasi-sentence addresses gender-related issues, (2) defining both the domain and coding category, (3) determining whether the quasi-sentence refers to a specific recipient or group based on gender and/or sexual orientation, (4) evaluating intersectionality, (5)<strong> </strong>assigning the sentiment or connotation, (6) determining if it's related to a goal, issue, or policy, and (7) characterising the policy if applicable.</p> <p>After completing the classification of the quasi-sentences in a given manifesto, coders fill in a survey for each manifesto document. The surveys provide information that cannot be directly inferred from the quasi-sentences, focusing on the gender ontology of a manifesto, the degree to which a manifesto entails a binary conception of sexes, the extent to which a manifesto promotes a patriarchal conception of the society, and how much a manifesto promotes heterosexuality as the only normal and socially acceptable sexual orientation of individuals. While the last four characteristics are gauged relying on quasi-interval measures (scales ranging from 0 to 10), the first one, gender ontology, consists in a categorical variable which distinguishes between manifestos with an essentialist ontology &ndash; gender and sex are the same and inseparable &ndash;, a constructivist ontology &ndash; biological sex is mediated through social construction of femininity and masculinity &ndash;, and other or undefined ontologies.</p>

opencc-by-sa-4.0Jun 2024View details →
zenodo52/100

Generalized Deception Dataset

<p>We took labeled&nbsp;datasets from five&nbsp;different deception-detection tasks with no licensing issues&nbsp;and converted them to a standard format. We inspected each dataset for quality and generated new cleaned versions.&nbsp;</p> <table> <thead> <tr> <th scope="col">Task</th> <th scope="col"># Deceptive</th> <th scope="col"># Truthful</th> </tr> </thead> <tbody> <tr> <td>Product Reviews</td> <td>10493</td> <td>10481</td> </tr> <tr> <td>Phishing</td> <td>6134</td> <td>9202</td> </tr> <tr> <td>Job Scams</td> <td>608</td> <td>13735</td> </tr> <tr> <td>Political Statements</td> <td>5669</td> <td>7167</td> </tr> <tr> <td>Fake News</td> <td>27486</td> <td>34615</td> </tr> </tbody> </table> <p>Our data is structured as five jsonlines files (one for each task) with a text to classify and a Boolean is_deceptive label.&nbsp;</p> <p>Sample data point:&nbsp;</p> <pre><code>{ "text":"the Annies List political group supports third-trimester abortions on demand.", "is_deceptive":true } </code></pre> <p>&nbsp;</p> <p>Changelog</p> <p><strong>1.1</strong></p> <ul> <li>Fixed flipped labels in the job scams dataset.&nbsp;</li> </ul>

opencc-by-4.0May 2022View details →
OpenNeuro48/100

Model-based fMRI reveals co-existing specific and generalized concept representations

Open the record for dataset details and reuse information.

openCC0Jan 2020View details →
zenodo48/100

Data for The Generally Curious: Thematically Distinct Datasets of 4chan's /pol/ Discussion Forum's 'General Threads'

<p>Over the second half of the 2010s, the /pol/ (&lsquo;politically incorrect&rsquo;) forum on the 4chan image board has emerged as a space within which various extreme political ideologies are discussed and cultivated, occasionally informing off-site acts of political extremism. While previous research has often studied this space as a unified whole, it is relevant to more specifically demarcate different publics within 4chan&rsquo;s /pol/ board, apart from studying it as an &lsquo;amorphous blob&rsquo;. This paper focuses specifically on &lsquo;generals&rsquo; - recurring threads with a specific thematic focus identified by a particular vernacular phrase or tag. By identifying them it is possible to subset the board&rsquo;s archive into multiple distinct datasets comprising discussions about a particular topic, such as Donald Trump, the Syria war, or British politics. We provide a dataset containing 58,841 opening posts and 13,697,738 replies to those, divided over 329 thematically distinct &lsquo;general thread&rsquo; collections. In this paper we outline our data collection and query protocol, the structure of the data and its rationale, as well as a number of suggested research uses for this new data.</p> <p>&nbsp;</p>

opencc-by-4.0Jan 2020View details →
zenodo48/100

Benchmark dataset for preprint: "EDEN: A high-performance, general-purpose, NeuroML-based neural simulator"

<p>The benchmark files and scripts to reproduce the figures of the preprint&nbsp;&nbsp;&quot;EDEN: A high-performance, general-purpose, NeuroML-based neural simulator&quot; ( https://arxiv.org/abs/2106.06752 )</p> <p>The benchmarks require a computer running Linux with Docker installed.</p> <p>Unpack the paper_experiments.zip file and follow the instructions in the README.md file to run the benchmarks and reproduce the figures.</p> <p>&nbsp;</p>

opencc-by-4.0Sep 2021View details →
zenodo48/100

Github commit data for the article "Beyond Zipf's law: Exploring the discrete generalized beta distribution in open-source repositories"

<p><span>This dataframe corresponds to the data used in the Nowak's et al. 2024 article "Beyond Zipf&rsquo;s law: Exploring the discrete generalized beta distribution in open-source repositories" (see reference below).</span></p> <p><span>It consists of the distirbutions of number of commits per user across a number of GitHub repositories.&nbsp;<br><br>There are three columns:</span></p> <ul> <li><span>repository: the repository name</span></li> <li><span># of commits: the number of commits of a given individual</span></li> <li><span>rank: the user rank in the repository (by decreasing number of commits)<br><br></span></li> </ul> <p><strong><span>Reference:</span></strong></p> <p><span>Nowak, P., Santolini, M., Singh, C., Siudem, G., &amp; Tupikina, L. (2024). Beyond Zipf&rsquo;s law: Exploring the discrete generalized beta distribution in open-source repositories.&nbsp;<em>Physica A: Statistical Mechanics and Its Applications</em>, <em>649</em>, 129927. <a href="https://doi.org/10.1016/j.physa.2024.129927">https://doi.org/10.1016/j.physa.2024.129927</a></span></p>

opencc-by-4.0Aug 2024View details →
zenodo48/100

Generalization of a density-dependent ecosystem function in dominant aquatic macroinvertebrates

<div> <div> <p>This Zenodo record contains the supporting data and code for the publication 'Generalization of a density-dependent ecosystem function in dominant aquatic macroinvertebrates', published in Oikos (<a title="DOI to publication" href="https://doi.org/10.1111/oik.10774">https://doi.org/10.1111/oik.10774</a>). The data are described in detail in the corresponding publication. The data archive contains a ReadMe file, two text files with the empirical data, and a corresponding R script for analysis. All required data to reproduce the full analysis from the original publication are provided.</p> <p>In order to reproduce the analysis and figures, run&nbsp;<code>DensityDependenceAnalysis20240319.R</code>. Make sure that your working directory is the actual folder containing the data files&nbsp;<code>Data_Field.txt</code>&nbsp;and&nbsp;<code>Data_Lab.txt</code>. If run in Rstudio, this should happen automatically. Else this is easily achieved by (re)starting R (or R Studio) by double-clicking the R script file from the folder. The script will produce all the figures from the paper, organized in a folder&nbsp;<code>AnalysisYYYYMMDD</code>&nbsp;and two subfolders&nbsp;<code>CheckFigs</code>&nbsp;and&nbsp;<code>SuppFigs</code>. Figures are prepared as pixel graphics (PNG).</p> </div> </div>

opencc-by-4.0Aug 2024View details →
zenodo48/100

One shot generalization in humans revealed through a drawing task - Dataset

<p>Dataset for:</p> <p>One shot generalization in humans revealed through a drawing task. Henning Tiedemann, Yaniv Morgenstern, Filipp Schmidt, Roland W Fleming. bioRxiv 2021.05.31.446461; doi: https://doi.org/10.1101/2021.05.31.446461</p>

opencc-by-4.0Aug 2021View details →
edi48/100

General Lake Model-Aquatic EcoDynamics model parameter set for Falling Creek Reservoir, Vinton, Virginia, USA 2013-2019

The General Lake Model (GLM), an open-source, one-dimensional hydrodynamic model, was used to simulate physical, chemical, and biological variables in Falling Creek Reservoir, Vinton, Virginia, USA between 15 May 2013 and 31 December 2019. GLM (v.3.2.0a3) was coupled to the Aquatic EcoDynamics (AED) module library via the Framework for Aquatic Biogeochemical Modeling (FABM). GLM-AED requires three configuration files to run the model. First, the glm3.nml file configures lake metadata (including hypsometry), meteorological driver data, stream inflow and outflow driver data files, and physical response variables (mixing parameters and sediment heat zones). Second, the aed2_20220111_2DOCpools.nml file configures various biogeochemical modules for the simulation of oxygen, carbon, silica, nitrogen, phosphorus, organic matter, and phytoplankton. Third, the aed2_phyto_pars_4Jan2022.nml file configures all parameters pertaining to phytoplankton dynamics. Meteorological data, two surface stream inflow files, a submerged oxygenation inflow file, and outflow file used in this calibration are also included.

openCC (other)May 2022View details →
edi48/100

West Falmouth Harbor Eelgrass Bed Extent, 2010-2019, generalized

West Falmouth Harbor (West Falmouth, MA, USA) has been experiencing a dramatic increase in nitrogen loading from an upgradient municipal wastewater treatment facility since the early 2000’s. As part of a long-term study into the effects of this nitrogen enrichment, we have been assessing changes in the extent and health of the eelgrass (Zostera marina) community within the harbor. We conducted side scan sonar surveys and analyzed them using SonarWiz and ArcGIS to delineate the extent of continuous seagrass bed within the harbor. The goal of this dataset is to assess changes in eelgrass habitat extent; small bare patches within the larger seagrass bed due to anchor scars and mooring scouring were not digitized. Groundtruthing of the edges was done with a combination of diver surveys, underwater video, and surface surveys at low tide. Surveys were run in every year from 2010 through 2022 and we plan to continue them into future years. Data analysis is ongoing and shapefiles will be added as they are run through quality assessment. Details for the 2010 data collection and analysis are detailed in Hayn 2012: Exchange of nitrogen and phosphorus between a shallow estuary and coastal waters, Appendix 3, permanently archived at https://hdl.handle.net/1813/31380.

openCC (other)Jan 2023View details →
zenodo44/100

Data and code from: Insect biomass decline scaled to species diversity: General patterns derived from a hoverfly community

<p>To study changes in&nbsp;flying insect communities, and hoverflies in particular, malaise trap samples from a German site&nbsp;were compared between two years (Hallmann et al. 2020).&nbsp;The data files deposited here&nbsp;contain&nbsp;data obtained from six malaise traps in the Wahnbachtal (North Rhine-Westphalia, Germany, 50.851944N, 7.320833E) that were deployed in 1989 and again in 2014, at the exact same locations. Traps were situated in wet meadows as well as tall perennial meadows, in close proximity to shrub corridors, to forest&ndash;grassland borders, and to the Wahnbach River and surrounded by agricultural land, essentially a rather heterogeneous habitat. The Wahnbach River and the greater part of the valley&nbsp;are protected for watershed purposes and are subject to nature conservation management by the Wahnbach Talperrenverband. Hence, several restrictions apply to safeguard against water contamination.</p> <p>Total insect biomass collected with these traps was already included in Hallmann et al. (2017), but here we focus on additional information: the abundance and richness of hoverflies (Syrphidae) in each of the collected samples (pots). Methodologies of collection are described in Sorg (1990), Schwan et al. (1993), Sorg et al. (2013), Hallmann et al. (2017), and Ssymank et al. (2018). &nbsp;In brief, malaise traps were deployed throughout the growing season and operated continuously (day and night). Malaise trap construction (e.g., size, material, colouring, and ground sealing) and placing (e.g., positioning, orientation, and slope of the locations) were standardised in all aspects. Insect samples were preserved in 80% ethanol solution. Catches of the six&nbsp;traps investigated in the present study were emptied regularly: On average exposure intervals were 7.0 d (SD = 0.5) in 1989 and 16.7 d (SD = 5.6) in 2014. Across the six traps in 2014 the total exposure time (in number of days) was 42% higher compared to 1989. All collected samples (n = 196) were used in the present analysis with in total 19,604 individual&nbsp;hoverflies counted, distributed over 162 species and 59 genera.</p> <p>To assess how environmental conditions have changed over the 25 year, several additional datasets were assembled. Climatic<br> data were obtained from 169 climatic stations and were used to interpolate daily weather variables to each trap location, using spatiotemporal kriging. These steps are described in detail in Hallmann et al. (2017).</p> <p>Our analysis (see R code)&nbsp;consists of three components. First, we&nbsp;considered total abundance, species richness, and species diversity, at two&nbsp;temporal scales: pooled per year, i.e., across the sampling season, and seasonally&nbsp;(i.e., per day), and we compared these metrics between 1989 and&nbsp;2014. Second, we examined how total flying biomass (i.e., the weight of all&nbsp;trapped insects, of which hoverflies are only a small proportion) related to&nbsp;total abundance as well as species richness of hoverflies. Third, we derived&nbsp;persistence probabilities and population growth rate trends per species, to&nbsp;examine interspecific variation in these parameters.</p> <p>Descriptions of the deposited files:</p> <p><strong>Groups.csv</strong><br> MF_NR&nbsp;= identifier of each of the six malaise trap locations<br> yrf&nbsp;= year of sampling<br> pot&nbsp;= sample identifier<br> dt = number of sampling days<br> from.dnr = day-of-the-year on which a pot was attached to a malaise trap<br> to.dnr = day-of-the-year on which a pot was collected from a malaise trap<br> mean.daynr = mean day-of-the-year of the sampling period<br> Nspec = number of different hoverfly species found in a pot<br> Nind = number of hoverfly individuals found in a pot</p> <p><strong>Counts.csv</strong><br> A matrix of counts of individual hoverflies per pot per species. The 196 rows represent the pots in the same order as in the file &#39;Groups.csv&#39;. The columns represent the 162 different hoverfly species found. The scientific species names are indicated in the column headers.</p> <p><strong>PairedData.csv</strong><br> pot =&nbsp;sample identifier<br> JAHR&nbsp;= year of sampling<br> MF_NR&nbsp;= identifier of each of the six malaise trap locations<br> dt = number of sampling days<br> from.dnr = day-of-the-year on which a pot was attached to a malaise trap<br> to.dnr = day-of-the-year on which a pot was collected from a malaise trap<br> NI&nbsp;= number of hoverfly individuals found in a potbiomass.daily<br> NSP&nbsp;= number of different hoverfly species found in a pot<br> biomass.daily = daily fresh weight [gram]&nbsp;of flying insects: total fresh weight in a&nbsp;pot&nbsp;divided by the number of sampling days.</p> <p><strong>ModelFrame.csv</strong><br> MF_NR&nbsp;= identifier of each of the six malaise trap locations<br> yrf = year of sampling<br> pot =&nbsp;sample identifier<br> dt = number of sampling days<br> from.dnr = day-of-the-year on which a pot was attached to a malaise trap<br> to.dnr = day-of-the-year on which a pot was collected from a malaise trap<br> mean.daynr = mean day-of-the-year of the sampling period<br> plot = identifier of each of the six malaise trap locations<br> date = date for which the weather variables are interpolated<br> daynr = day-of-the-year&nbsp;for which the weather variables are interpolated<br> altitude = altitude [m] of the malaise trap locations<br> year = year of sampling<br> temperature = interpolated temperature [degrees Celsius]<br> precipitation = interpolated precipitation [mm per day]<br> wind.speed = interpolated wind speed [m/s]</p> <p><strong>Data_Rcode.pdf</strong><br> This pdf&nbsp;provides the R-code behind the analysis of&nbsp;the Hoverfly data. Three datasets are provided along with this R-code document, namely &quot;Counts.csv&quot;,&nbsp;&quot;Groups.csv&quot;, &quot;PairedData.csv&quot; and &quot;ModelFrame.csv&quot;. Additionally, the BUGS-code &quot;&quot;syrphidModel.jag&quot;&nbsp;is required for running the daily-activity model in JAGS.</p> <p><strong>syrphidModel.jag</strong><br> This&nbsp;BUGS-code is required for running the daily-activity model in JAGS.</p>

opencc-by-4.0Nov 2020View details →
zenodo44/100

Datos generales sobre revistas de la Universidad Nacional, Costa Rica

<p>Consiste en un mapeo general de caracter&iacute;sticas de las revistas que se publican en la Universidad Nacional. Facilita informaci&oacute;n m&iacute;nima como el ISSN, sitio web, nombre de la revista y a&ntilde;o de inicio.</p>

opencc-by-nc-nd-3.0Dec 2020View details →
zenodo44/100

RDF version of the data from Choi, JS., Ha, M.K., Trinh, T.X. et al. Towards a generalized toxicity prediction model for oxide nanomaterials using integrated data from different sources. Sci Rep 8, 6110 (2018)

<p>RDF version of the data from Choi, JS., Ha, M.K., Trinh, T.X. et al. Towards a generalized toxicity prediction model for oxide nanomaterials using integrated data from different sources. Sci Rep 8, 6110 (2018)</p>

opencc-by-4.0Jan 2021View details →
zenodo44/100

Data release for "OrchID: a Generalized Framework for Taxonomic Classification of Images Using Evolved Artificial Neural Networks"

<p><strong>Abstract</strong></p> <p>Taxonomic expertise for the identification of species is rare and costly. On-going advances in computer vision and machine learning have led to the development of numerous semi- and fully automated species identification systems. However, these systems are rarely agnostic to specific morphology, rarely can perform taxonomic &ldquo;approximation&rdquo; (by which we mean partial identification at least to higher taxonomic level if not to species), and frequently rely on costly scientific imaging technologies.</p> <p>We present a generic, hierarchical identification system for automated taxonomic approximation of organisms from images. We assessed the effectiveness of this system using photographs of slipper orchids (Cypripedioideae), for which we implemented image pre-processing, segmentation, and colour and shape feature extraction algorithms to obtain digital phenotypes for 116 species. The identification system trained on these digital phenotypes uses a nested hierarchy of artificial neural networks for pattern recognition and automated classification that mirrors the Linnean taxonomy, such that user-submitted photos can be assigned a genus, section, and species classification by traversing this hierarchy.</p> <p>Performance of the identification system varied depending on photo quality, number of species included for training, and desired taxonomic level for identification. High quality photos were scarce for some taxa and were under-represented in the training set, resulting in imbalanced network training. The image features used for training were sufficient to reliably identify photos to the correct genus but less so to the correct section and species.</p> <p>The outcomes of this project include a library of feature extraction algorithms called <em>ImgPheno</em>, a collection of scripts for neural network training called <em>NBClassify</em>, a library for evolutionary optimization of artificial neural network construction called <em>AI::FANN::Evolving</em> and a planned web application called <em>OrchID</em> for identification of user-submitted images. All project outcomes are open source and freely available.</p> <p><strong>About this release</strong></p> <p>This release corresponds belongs with our response to the reviewers of PLoS One. At this stage of the review cycle the manuscript is assessed as &#39;minor revision&#39;. Consequently, we don&#39;t anticipate making more releases until publication.</p>

opencc-zeroOct 2015View details →
zenodo44/100

Old Mandu (बूढ़ी मांडू), Dhār district, Madhya Pradesh. General view of the fort wall.

<p>Old Mandu (बूढ़ी मांडू), Dhār district, Madhya Pradesh. General view of the north line of the fort wall. Photo 2/2010.</p>

opencc-by-4.0Mar 2017View details →
zenodo44/100

Old Mandu (बूढ़ी मांडू), Dhār district, Madhya Pradesh. General view of the main tank with remains of ghāṭ

<p>Old Mandu (बूढ़ी मांडू), Dhār district, Madhya Pradesh. General view of the main tank with remains of ghāṭ. Photo 2010.</p>

opencc-by-4.0Mar 2017View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record