Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

5,155

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

5,155 results for “Data Base”

Learn how ShareScore rates datasets ↗
dryad40/100

Data from: Non-invasive age estimation based on fecal DNA using methylation-sensitive high-resolution melting for Indo-Pacific bottlenose dolphins

<p class="MsoNormal"><span>Age is necessary information for the study of life history of wild animals. A general method to estimate the age of odontocetes is counting dental growth layer groups (GLGs). However, this method is highly invasive as it requires the capture and handling of individuals to collect their teeth.</span><span> Recently, the development of DNA-based age </span><span>estimation methods has been actively studied as an alternative to such invasive methods, of which many have used biopsy samples. However, if DNA-based age estimation can be developed from fecal samples, age estimation can be performed without touching or disrupting individuals, thus establishing an entirely non-invasive method. </span><span>We developed an age estimation model using the methylation rate of two gene regions, <em>GRIA2</em> and <em>CDKN2A,</em> measured through methylation-sensitive high-resolution melting (MS-HRM) from fecal samples of wild Indo-Pacific bottlenose dolphins (<em>Tursiops aduncus</em>). The age of individuals was known through conducting longitudinal individual identification surveys underwater. Methylation rates were quantified from 36 samples. Both gene regions showed a significant correlation between age and methylation rate. The age estimation model was constructed based on the methylation rates of both genes which achieved sufficient accuracy (after LOOCV: MAE = 5.08, <em>R<sup>2</sup></em> = 0.34) for the ecological studies of the Indo-Pacific bottlenose dolphins, with a lifespan of 40-50 years. This is the first study to report the use of non-invasive fecal samples to estimate the age of marine mammals.</span></p>

opencc-zeroNov 2023View details →
zenodo40/100

Data Set for the Journal Article "Automated Preparation of Nanoscopic Structures: Graph-Based Sequence Analysis, Mismatch Detection, and pH-Consistent Protonation with Uncertainty Estimates"

<p>This repository containes the data generated by ASAP and discussed in the journal article [Csizi, K.-S. and Reiher, M., 2023, arXiv:2307.16344], including Cartesian coordinates of training and test set molecules, and MD trajectories.&nbsp;</p>

opencc-by-4.0Nov 2023View details →
zenodo40/100

Data from: Methylation-based markers for the estimation of age in African Cheetah, Acinonyx jubatus

<p>This is a dataset for methylation analyses in cheetah done by EpiTYPER mass array.</p>

opencc-by-4.0Nov 2023View details →
zenodo40/100

EEG and eye-tracking data from a go/no-go saccadic task based on facial expression cues

<p>Electroencephalographic (EEG) and eye-tracking data from 20 healthy individuals who performed a&nbsp;go/no-go saccadic task based on facial expression cues aimed at studying error monitoring processes.</p> <p>&nbsp;</p> <p>Version 2 includes only the EEG data, but&nbsp; with triggers of correct and erroneous actions. Version 2 has an error that was corrected for version 3.</p> <p>&nbsp;</p> <p>EEG Triggers</p> <table> <tbody> <tr> <td> <p><span>19</span></p> </td> <td> <p><span>Eye tracker starts recording</span></p> </td> </tr> <tr> <td> <p><span>1</span></p> </td> <td> <p><span>Beginning of each trial</span></p> </td> </tr> <tr> <td> <p><span>200</span></p> </td> <td> <p><span>Gap between neutral and instruction</span></p> </td> </tr> <tr> <td> <p><span>99</span></p> </td> <td> <p><span>Fixation cross between instruction and saccade</span></p> </td> </tr> <tr> <td> <p><span>2</span></p> </td> <td> <p><span>Instruction no-go happy</span></p> </td> </tr> <tr> <td> <p><span>21</span></p> </td> <td> <p><span>Target left no-go happy</span></p> </td> </tr> <tr> <td> <p><span>22</span></p> </td> <td> <p><span>Target right no-go happy</span></p> </td> </tr> <tr> <td> <p><span>121</span></p> </td> <td> <p><span>Response period for no-go happy after target left</span></p> </td> </tr> <tr> <td> <p><span>122</span></p> </td> <td> <p><span>Response period for no-go happy after target right</span></p> </td> </tr> <tr> <td> <p><span>3</span></p> </td> <td> <p><span>Instruction no-go sad</span></p> </td> </tr> <tr> <td> <p><span>31</span></p> </td> <td> <p><span>Target left no-go sad</span></p> </td> </tr> <tr> <td> <p><span>32</span></p> </td> <td> <p><span>Target right no-go sad</span></p> </td> </tr> <tr> <td> <p><span>131</span></p> </td> <td> <p><span>Response period for no-go sad after target left</span></p> </td> </tr> <tr> <td> <p><span>132</span></p> </td> <td> <p><span>Response period for no-go sad after target right</span></p> </td> </tr> <tr> <td> <p><span>4</span></p> </td> <td> <p><span>Instruction pro-right</span></p> </td> </tr> <tr> <td> <p><span>42</span></p> </td> <td> <p><span>Target right pro-right</span></p> </td> </tr> <tr> <td> <p><span>142</span></p> </td> <td> <p><span>Response period for pro-right</span></p> </td> </tr> <tr> <td> <p><span>5</span></p> </td> <td> <p><span>Instruction pro-left</span></p> </td> </tr> <tr> <td> <p><span>51</span></p> </td> <td> <p><span>Target left pro-left</span></p> </td> </tr> <tr> <td> <p><span>151</span></p> </td> <td> <p><span>Response period for pro-left</span></p> </td> </tr> <tr> <td> <p><span>6</span></p> </td> <td> <p><span>Instruction anti-left</span></p> </td> </tr> <tr> <td> <p><span>61</span></p> </td> <td> <p><span>Target left anti-left</span></p> </td> </tr> <tr> <td> <p><span>161</span></p> </td> <td> <p><span>Response period for anti-left</span></p> </td> </tr> <tr> <td> <p><span>7</span></p> </td> <td> <p><span>Instruction anti-right</span></p> </td> </tr> <tr> <td> <p><span>72</span></p> </td> <td> <p><span>Target right anti-right</span></p> </td> </tr> <tr> <td> <p><span>172</span></p> </td> <td> <p><span>Response period for anti-right</span></p> </td> </tr> <tr> <td> <p><span>199</span></p> </td> <td> <p><span>Final fixation period (1.5s black screen after last saccade)</span></p> </td> </tr> <tr> <td> <p><span>190</span></p> </td> <td> <p><span>Eye tracker stops recording</span></p> </td> </tr> <tr> <td> <p><span>proCorr</span></p> </td> <td> <p><span>Beginning of saccade in the correct direction following a pro-saccade intruction</span></p> </td> </tr> <tr> <td> <p><span>proErr</span></p> </td> <td> <p><span>Beginning of saccade in the erroneous direction following a pro-saccade intruction</span></p> </td> </tr> <tr> <td> <p><span>antiCorr</span></p> </td> <td> <p><span>Beginning of saccade in the correct direction following a anti-saccade intruction</span></p> </td> </tr> <tr> <td> <p><span>antiErr</span></p> </td> <td> <p><span>Beginning of saccade in the erroneous direction following a anti-saccade intruction</span></p> </td> </tr> <tr> <td> <p><span>nogoErr</span></p> </td> <td> <p><span>Beginning of saccade in the erroneous direction following a no-go intruction</span></p> </td> </tr> </tbody> </table>

opencc-by-4.0Oct 2021View details →
zenodo40/100

Primary NMR Data Supporting the Article "Synthetic approach to 2-alkyl-4-quinolones and 2-alkyl-4-quinolone-3-carboxamides based on common β-keto amide precursors"

<p>This archive contains raw 1H/13C FIDs and associated data in Bruker-specific format that can be viewed with Bruker&rsquo;s TopSpin or other appropriate NMR processing software. The subfolders are named in accordance with the compound numbering in the associated research paper (Synthetic Approach to 2-Alkyl-4-quinolones and 2-Alkyl-4-quinolone-3-carboxamides Based on Common &beta;-Keto Amide Precursors).</p> <p>Correspondence: angelov@uni-plovdiv.bg</p> <p>&nbsp;</p>

opencc-by-4.0Nov 2023View details →
dryad40/100

Data from: Building trust takes time: Limits to arbitrage for blockchain-based assets

<p><span>The dataset contains all historical order book snapshots and blockchain network information used to generate the results for the paper "Building Trust takes Time". </span></p> <p><span>A blockchain replaces central counterparties with time-consuming consensus pro</span><span>tocols to record the transfer of ownership.</span> <span>This settlement latency slows cross-</span><span>exchange trading, exposing arbitrageurs to price risk. Off-chain settlement, instead, </span><span>exposes arbitrageurs to costly default risk. We show with Bitcoin network and or</span><span>der book data that cross-exchange price differences coincide with periods of high </span><span>settlement latency, asset flows chase arbitrage opportunities, and price differences </span><span>across exchanges with low default risk are smaller. Blockchain-based trading thus </span><span>faces a dilemma: reliable consensus protocols require time-consuming settlement </span><span>latency, leading to arbitrage limits. Circumventing such arbitrage costs is possible </span><span>only</span> <span>by reinstalling trusted intermediation, which mitigates default risk.</span></p>

opencc-zeroNov 2023View details →
dryad40/100

Trait-based sensitivity of large mammals to a catastrophic tropical cyclone: DNA metabarcoding data

<p>Extreme weather events perturb ecosystems and increasingly threaten biodiversity<sup>1</sup>. Ecologists emphasize the need to forecast and mitigate the impacts of these incidents, which requires knowledge of how risk is distributed among species and environments, but the scale and unpredictability of extreme events complicates assessment<sup>1</sup><sup>–4</sup>. These challenges are compounded for large animals ('megafauna'), which play crucial ecological roles but are hard to study<sup>5</sup>. Traits such as body size, dispersal ability, and habitat affiliation are among the hypothesized determinants of animals' vulnerability to natural hazards<sup>1,6,7</sup>. However, it has rarely been possible to test these propositions or, more generally, to link short- and longer-term effects of weather-related disturbance<sup>8,9</sup>. Here, we show how large herbivores and carnivores in Mozambique responded to Intense Tropical Cyclone Idai, the deadliest storm on record in Africa, across scales ranging from individual decisions in the hours after landfall to community-level responses nearly 20 months later. Animals occupying low-elevation habitats exhibited strong spatial responses to rising floodwaters. Body size predicted species' subsequent numerical responses: small-bodied species exhibited the greatest population declines. We trace this sensitivity to limited mobility, which increased likelihood of death during the flood and constrained animals' capacity to withstand food shortages afterward. Our results identify potentially general trait-based mechanisms underlying animal responses to severe weather and may help to inform strategies for wildlife conservation in a volatile climate.</p> <ol> <li><span><em><span>Climate Change 2022: Impacts, Adaptation and Vulnerability. Contribution of Working Group II to the Sixth Assessment Report of the Intergovernmental Panel on Climate Change</span></em><span> [H.-O. Pörtner, D.C. Roberts, M. Tignor, E.S. Poloczanska, K. Mintenbeck, A. Alegría, M. Craig, S. Langsdorf, S. Löschke, V. Möller, A. Okem, B. Rama (eds.)]. Cambridge University Press. Cambridge University Press, Cambridge, UK and New York, NY, USA, (2022).</span></span></li> <li><span><span>Smith, M. An ecological perspective on extreme climatic events: A synthetic definition and framework to guide future research. <em>J. Ecol.</em> <strong>99</strong>, 656-663 (2011).</span></span></li> <li><span><span>Ummenhofer, C. C., &amp; Meehl, G. A. Extreme weather and climate events with ecological relevance: a review, <em>Phil. Trans. R. Soc. B. </em><strong>372</strong>, 20160135 (2017).</span></span></li> <li><span><span>Jentsch, A., Kreyling, J., &amp; Beierkuhnlein, C. A new generation of climate-change experiments: events, not trends. <em>Front. Ecol. Environ. </em><strong>5</strong>, 365-374 (2007).</span></span></li> <li><span><span>Pringle, R. M., et. al. Impacts of large herbivores on terrestrial ecosystems. <em>Current Biology</em> <strong>33</strong>, R584-R610 (2023).</span></span></li> <li><span><span>Spiller, D. A., Losos, J. B., &amp; Schoener, T. W. Impact of a catastrophic hurricane on island populations. <em>Science </em><strong>281</strong>, 695-697 (1998). </span></span></li> <li><span><span>Schoener, T. W., &amp; Spiller, D. A. Nonsynchronous recovery of community characteristics in island spiders after a catastrophic hurricane. <em>PNAS </em><strong>103</strong>, 2220-2225 (2006).</span></span></li> <li><span><span>Pruitt, N., Little, A. G., Majumdar, S. J., Schoener, T. W., &amp; Fisher, D. N. Call-to-Action: A global consortium for tropical cyclone ecology. <em>TREE </em><strong>34</strong>, 588-590 (2019).</span></span></li> <li><span><span>Lin, T. C., Hogan, J. A., &amp; Chang, C. T. Tropical cyclone ecology: a scale-link perspective. <em>TREE </em><strong>35</strong>, 594-604 (2020).</span></span></li> </ol>

opencc-zeroDec 2022View details →
zenodo40/100

phenology data of lake mixing, and stratification based on the CESM2-LE output

<p>The phenology data of lake mixing, and stratification calculated from the daily output of CESM2 large ensemble. For more information about this dataset, please refer to the manuscript entitled 'Projected changes in the phenology of stratification and overturning in ice-covered lakes of the Northern Hemisphere (Lei Huang et al)'. For more information about CESM2 large ensemble, please refer to https://www.cesm.ucar.edu/community-projects/lens2.&nbsp;</p>

opencc-by-4.0Nov 2023View details →
zenodo40/100

Solar flare forecasting based on magnetogram sequences learning with MViT and data augmentation

<p><strong>Source codes and dataset of the research "Solar flare forecasting based on magnetogram sequences learning with MViT and data augmentation".</strong></p><p>Our work employed PyTorch, a framework for training Deep Learning models with GPU support and automatic back-propagation, to load the MViTv2 s models with Kinetics-400 weights. To simplify the code implementation, eliminating the need for an explicit loop to train and the automation of some hyperparameters, we use the PyTorch Lightning module. The inputs were batches of 10 samples with 16 sequenced images in 3-channel resized to 224 × 224 pixels and normalized from 0 to 1.</p><p>Most of the papers in our literature survey split the original dataset chronologically. Some authors also apply k-fold cross-validation to emphasize the evaluation of the model stability. However, we adopt a hybrid split taking the first 50,000 to apply the 5-fold cross-validation between the training and validation sets (known data), with 40,000 samples for training and 10,000 for validation. Thus, we can evaluate performance and stability by analyzing the mean and standard deviation of all trained models in the test set, composed of the last 9,834 samples, preserving the chronological order (simulating unknown data).</p><p>We develop three distinct models to evaluate the impact of oversampling magnetogram sequences through the dataset. The first model, Solar Flare MViT (SF MViT), has trained only with the original data from our base dataset without using oversampling. In the second model, Solar Flare MViT over Train (SF MViT oT), we only apply oversampling on training data, maintaining the original validation dataset. In the third model, Solar Flare MViT over Train and Validation (SF MViT oTV), we apply oversampling in both training and validation sets.</p><p>We also trained a model oversampling the entire dataset. We called it the "SF_MViT_oTV Test" to verify how resampling or adopting a test set with unreal data may bias the results positively.</p><p><strong>GitHub version</strong></p><p>The .zip hosted here contains all files from the project, including the checkpoint and the output files generated by the codes. We have a clean version hosted on GitHub (<a href="https://github.com/lfgrim/SFF_MagSeq_MViTs">https://github.com/lfgrim/SFF_MagSeq_MViTs</a>), without the magnetogram_jpg folder (which can be downloaded directly on <a href="https://tianchi-competition.oss-cn-hangzhou.aliyuncs.com/531804/dataset_ss2sff.zip">https://tianchi-competition.oss-cn-hangzhou.aliyuncs.com/531804/dataset_ss2sff.zip)</a> and the output and checkpoint files. Most code files hosted here also contain comments on the Portuguese language, which are being updated to English in the GitHub version.</p><p><strong>Folders Structure</strong></p><p>In the Root directory of the project, we have two folders:&nbsp;</p><ul><li>magnetogram_jpg: holds the source images provided by Space Environment Artificial Intelligence Early Warning Innovation Workshop through the link <a href="https://tianchi-competition.oss-cn-hangzhou.aliyuncs.com/531804/dataset_ss2sff.zip">https://tianchi-competition.oss-cn-hangzhou.aliyuncs.com/531804/dataset_ss2sff.zip. </a>It comprises 73,810 samples of high-quality magnetograms captured by HMI/SDO from 2010 May 4 to 2019 January 26. The HMI instrument provides these data (stored in hmi.sharp_720s dataset), making new samples available every 12 minutes. However, the images from this dataset were collected every 96 minutes. Each image has an associated magnetogram comprising a ready-made snippet of one or most solar ARs. It is essential to notice that the magnetograms cropped by SHARP can contain one or more solar ARs classified by the National Oceanic and Atmospheric Administration (NOAA).</li><li>Seq_Magnetogram: contains the references for source images with the corresponding labels in the next 24 h. and 48 h. in the respectively M24 and M48 sub-folders.<ul><li>M24/M48: both present the following sub-folders structure:<ul><li>Seqs16;</li><li>SF_MViT;</li><li>SF_MViT_oT;</li><li>SF_MViT_oTV;</li><li>SF_MViT_oTV_Test.</li></ul></li></ul></li></ul><p>There are also two files in root:</p><ul><li>inst_packages.sh: install the packages and dependencies to run the models.</li><li>download_MViTS.py: download the pre-trained MViTv2_S from PyTorch and store it in the cache.</li></ul><p>M24 and M48 folders hold reference text files&nbsp;(flare_Mclass...) linking the images in the magnetogram_jpg folders or the sequences (Seq16_flare_Mclass...)&nbsp; in the Seqs16 folders with their respective labels. They also hold "cria_seqs.py" which was responsible for creating the sequences and "test_pandas.py" to verify head info and check the number of samples categorized by the label of the text files. All the text files with the prefix "Seq16" and inside the Seqs16 folder were created by "criaseqs.py" code based on the correspondent "flare_Mclass" prefixed text files.</p><p>Seqs16 folder holds reference text files, in which each file contains a sequence of images that was pointed to the magnetogram_jpg folders.</p><p>All SF_MViT... folders hold the model training codes itself (SF_MViT...py) and the corresponding job submission (jobMViT...), temporary input (Seq16_flare...),&nbsp;output (saida_MVIT... and MViT_S...), error (err_MViT...) and checkpoint files (sample-FLARE...ckpt). Executed model training codes generate output, error, and checkpoint files. There is also a folder called "lightning_logs" that stores logs of trained models.</p><p><strong>Naming pattern for the files:</strong></p><ul><li>magnetogram_jpg: follows the format<i> </i>"hmi.sharp_720s.&lt;SHARP-ID&gt;.&lt;date&gt;.magnetogram.fits.jpg" and</li><li>Seqs16: follows the format "hmi.sharp_720s.<i>&lt;</i>SHARP-ID<i>&gt;</i>.&lt;init-date&gt;.to.&lt;end-date&gt;", where:<ul><li>hmi: is the instrument that captured the image</li><li>sharp_720s: is the database source of SDO/HMI.</li><li>&lt;SHARP-ID&gt;: is the identification of SHARP region, and can contain one or more solar ARs classified by the (NOAA).</li><li>&lt;date&gt;: is the date-time the instrument captured the image in the format yyyymmdd_hhnnss_TAI (y:year, m:month, d:day, h:hours, n:minutes, s:seconds).</li><li>&lt;init-date&gt;: is the date-time when the sequence starts, and follow the same format of &lt;date&gt;.</li><li>&lt;end-date&gt;: is the date-time when the sequence ends, and follow the same format of &lt;date&gt;.</li></ul></li><li>Reference text files in M24 and M48 or inside SF_MViT... folders follows the format "&lt;prefix&gt;flare_Mclass_&lt;forecasting-horizon&gt;_&lt;dataset&gt;.txt&lt;over&gt;", where:<ul><li>&lt;prefix&gt;: is Seq16 if refers to a sequence, or void if refers direct to images.</li><li>&lt;forecasting-horizon&gt;: "24h" or "48h".</li><li>&lt;dataset&gt;: is "TrainVal&lt;n&gt;" or "Test". The &lt;n&gt; refers to the split of Train/Val.</li><li>&lt;over&gt;: void or "_over" after the extension (...txt_over): means temporary input reference that was over-sampled by a training model.</li></ul></li><li>All SF_MViT...folders:<ul><li>Model training codes: "SF_MViT_&lt;oversampling-type&gt;_M+_&lt;forecasting-horizon&gt;_&lt;split-type&gt;&lt;gpu-type&gt;", where:<ul><li>&lt;oversampling -type&gt;: void or "oT" (over Train) or "oTV" (over Train and Val) or "oTV_Test" (over Train, Val and Test);</li><li>&lt;forecasting-horizon&gt;: "24h" or "48h";</li><li>&lt;split-type&gt;: "oneSplit" for a specific split or "allSplits" if run all splits.</li><li>&lt;gpu-type&gt;: void is default to run 1 GPU or "2gpu" to run into 2 gpus systems;</li></ul></li><li>Job submission files: "jobMViT_&lt;queue&gt;", where:<ul><li>&lt;queue&gt;: point the queue in Lovelace environment hosted on CENAPAD-SP (<a href="https://www.cenapad.unicamp.br/parque/jobsLovelace">https://www.cenapad.unicamp.br/parque/jobsLovelace</a>)</li></ul></li><li>Temporary inputs: "Seq16_flare_Mclass_&lt;forecasting-horizon&gt;_&lt;dataset&gt;.txt&lt;over&gt;:<ul><li>&lt;dataset&gt;: train or val;</li><li>&lt;over&gt;: void or "_over" after the extension (...txt_over): means temporary input reference that was over-sampled by a training model.</li></ul></li><li>Outputs: "saida_MViT_Adam_10-7&lt;split&gt;", where:<ul><li>&lt;split&gt;: k0 to k4, means the correlated split of the output, or void if the output is from all splits.</li></ul></li><li>Error files: "err_MViT_Adam_10-7&lt;split&gt;", where:<ul><li>&lt;split&gt;: k0 to k4, means the correlated split of the error log file, or void if the error file is from all splits.</li></ul></li><li>Checkpoint files: "sample-FLARE_MViT_S_10-7-epoch=&lt;n-epoch&gt;-valid_loss=&lt;loss-value&gt;-Wloss_k=&lt;n-split&gt;.ckpt", where:<ul><li>&lt;n-opoch&gt;: epoch number of the checkpoint;</li><li>&lt;loss-value&gt;: corresponding valid loss;</li><li>&lt;n-split&gt;: 0 to 4.</li></ul></li></ul></li></ul>

opencc-by-4.0Nov 2023View details →
zenodo40/100

What can radar-based measures of subglacial hydrology tell us about basal shear stress? A case study at Thwaites Glacier, West Antarctica (Interpolated Data)

<p>This dataset accompanies the paper 'What can radar-based measures of subglacial hydrology tell us about basal shear stress? A case study at Thwaites Glacier, West Antarctica' in Journal of Glaciology, and can be used alongside the code found on Github (https://github.com/rohaizharis/inversion_radar2022) to reproduce the figures. The dataset consists of ice-penetrating radar data (specularity and relative reflectivity) and basal shear stress inversions that have been linearly interpolated onto radar flight tracks.</p>

opencc-by-4.0Aug 2023View details →
zenodo40/100

Data of "Aerosol-cloud interactions near cloud base deteriorating the haze pollution in East China"

<p><span>The attached data is observations from ground to 1200 m a.g.l. using a tethered airship in Yangtze Rive Delta of China. The data is for analysis and figures in the study of "Aerosol-cloud interactions near cloud base deteriorating the haze pollution in East China".</span></p>

opencc-by-4.0Dec 2023View details →
dryad40/100

Data from: A cost-effective blood DNA methylation-based age estimation method in domestic cats, Tsushima leopard cats (Prionailurus bengalensis euptilurus), and Panthera species, using targeted bisulfite sequencing and machine learning models

<p><span>Knowledge of individual age can help both in-situ and ex-situ conservation programs to design more efficient and suitable management plans for targeted wildlife species. DNA methylation is one of the epigenetic aging markers that has emerged as a promising tool that can estimate age with high accuracy using only a tiny amount of biological material, which can be collected in a minimally invasive way. Here, we sequenced five targeted genetic regions and used </span><span>8–23</span><span> selected CpG sites to build age estimation models with machine learning methods </span><span>with about only $3–7 per sample</span><span>, using blood samples of seven Felidae species—ranging from small to big, and domestic to endangered species: domestic cats (<em>Felis catus</em>, 139 samples), Tsushima leopard cats (<em>Prionailurus bengalensis euptilurus</em>, 84 samples), and five<em> Panthera </em>species (96 samples). </span><span>The models built achieved satisfactory accuracy—the mean absolute error of the best models was 1.966, 1.348, and 1.552 years in domestic cats, Tsushima leopard cats, and <em>Panthera</em> spp., respectively.</span><span> Our models in domestic cats and Tsushima leopard cats were applicable to individuals regardless of health conditions, indicating the high applicability of our models to samples collected from diverse situations, e.g., rescued individuals in the context of conservation. We also showed the possibility of developing universal age estimation models for the five<em> Panthera</em> spp. using two of the five genetic regions, suggesting an even lower cost to use our models for future applications.</span></p>

opencc-zeroJan 2024View details →
zenodo40/100

Data for "Mechanically-Sensitive Fluorochromism by Molecular Domino Transformation in a Schiff Base Crystal"

<p>The dataset contains input and output files of computational chemistry by Quantum ESPRESSO and Gaussian softwares conducted on two polymorphic crystal structures of 4-nitro-N-salicylideneaniline.</p>

opencc-by-4.0Jun 2024View details →
zenodo40/100

From Pixels to Phenotypes: Integrating Image-Based Profiling with Cell Health Data Improves Interpretability

<p>Code: https://github.com/srijitseal/BioMorph_Space<br> <br> Cell Painting assays generate morphological profiles that are versatile descriptors of biological systems and have been used to predict <em>in vitro</em> and <em>in vivo</em> drug effects. However, Cell Painting features are based on image statistics, and are, therefore, often not readily biologically interpretable. In this study, we introduce an approach that maps specific Cell Painting features into the BioMorph space using readouts from comprehensive Cell Health assays. We validated that the resulting BioMorph space effectively connected compounds not only with the morphological features associated with their bioactivity but with deeper insights into phenotypic characteristics and cellular processes associated with the given bioactivity. The BioMorph space revealed the mechanism of action for individual compounds, including dual-acting compounds such as emetine, an inhibitor of both protein synthesis and DNA replication. In summary, BioMorph space offers a more biologically relevant way to interpret cell morphological features from the Cell Painting assays and to generate hypotheses for experimental validation.</p> <p>&nbsp;</p> <p>The following datasets are released:<br> &nbsp;</p> <p>Cell_Health_median_357_profiles_70_labels.csv :<br> The Cell Heath dataset for CRISPR perturbations.&nbsp;Contains&nbsp;median consensus signatures for the 357 consensus profiles (119 CRISPR perturbations &times; 3 cell lines) Ref: Way et al.</p> <p>Cell_Painitng_CRISPR_Perturbations_357_profiles_827_features_scaled.csv:<br> The Cell Painting dataset for CRISPR perturbations.&nbsp;Contains 827 morphology features (and metadata annotation) for 357 consensus profiles (119 CRISPR perturbations &times; 3 cell lines).&nbsp;Ref: Way et al.</p> <p>Cell_Painting_data_658_compounds_827_Features_scaled.csv<br> The Cell Painting dataset for compound perturbations.&nbsp;Contains 658 structurally unique compounds with 827 Cell Painting features. Ref: Bray et al</p> <p>Endpoints_9_Mitotox_biological_activities_658_compounds.csv<br> The biological assay activity labels&nbsp;for compound perturbations.&nbsp;Contains 658 structurally unique compounds with&nbsp;9 biological activity consensus hit calls.&nbsp;Ref: ToxCast/MoleculeNet</p> <p>BioMoprh_pvalue_658_compunds_398_BioMorph_terms.csv:<br> The dataset of standardised BioMorph term p-values. Contains&nbsp;398 BioMorph terms for the 658 compounds in the biological activity dataset.&nbsp;<br> <br> References:&nbsp;<br> Way et al. Predicting cell health phenotypes using image-based morphology profiling. Mol Biol Cell. 2021;32(9):995-1005.<br> Bray et al. A dataset of images and morphological profiles of 30 000 small-molecule treatments using the Cell Painting assay. Gigascience. 2017;6(12):1-5.&nbsp;<br> MoleculeNet: Wu&nbsp;et al. MoleculeNet: A benchmark for molecular machine learning. Chem Sci. 2018;9(2):513-530.&nbsp;<br> ToxCast:&nbsp;Exploring ToxCast Data | US EPA https://www.epa.gov/chemical-research/exploring-toxcast-data (accessed Jul 9, 2023).</p>

opencc-by-4.0Jan 2024View details →
zenodo40/100

Supplementary Data for "Streamlining Vocabulary Conversion to SKOS: A YAML-based Approach to Facilitate Participation in the Semantic Web"

<p>This dataset contains quality assessment results for 26 vocabularies. The assessment was conducted using the <a href="https://skos-play.sparna.fr/skos-testing-tool/">qSKOS vocabulary quality assessment tool</a>.</p> <p>The 26 assessed vocabularies were converted from their original formats into the Simple Knowledge Organization System (SKOS) data model using the approach described in our paper titled <a href="https://doi.org/10.1007/978-3-031-62362-2_9">"Streamlining Vocabulary Conversion to SKOS: A YAML-based Approach to Facilitate Participation in the Semantic Web"</a>, presented at the <a href="https://doi.org/10.1007/978-3-031-62362-2">24th International Conference on Web Engineering (ICWE 2024)</a>.</p> <p>The dataset contains a quality assessment for the following vocabularies:</p> <ol> <li>A Taxonomy of Evaluation Towards Standards</li> <li>Cross-Device Taxonomy</li> <li>What Makes a Data-driven Business Model? A Consolidated Taxonomy</li> <li>DDI Aggregation Method</li> <li>DDI Mode of Collection</li> <li>Building a New Taxonomy for Data Discretization Techniques</li> <li>Demopaedia</li> <li>Data Science Glossary</li> <li>A Taxonomy of Evaluation Approaches in Software Engineering</li> <li>Evaluation Thesaurus</li> <li>The Glossary of Human Computer Interaction</li> <li>Human-Factors Taxonomy</li> <li>A Taxonomy to Structure and Analyze Human&ndash;Robot Interaction</li> <li>A Taxonomy of Interaction for Instructional Multimedia</li> <li>A Taxonomy of Interrogation Methods</li> <li>Design Vocabulary for Human&ndash;IoT Systems Communication</li> <li>Understanding Movement and Interaction: An Ontology for Kinect-Based 3D Depth Sensors</li> <li>Thesaurus Mass Communication</li> <li>Mixed-Initiative Human-Robot Interaction: Definition, Taxonomy, and Survey</li> <li>A Taxonomy of Quality of Service and Quality of Experience of Multimodal Human-Machine Interaction</li> <li>A Human-Centered Taxonomy of Interaction Modalities and Devices</li> <li>A Taxonomy of Spatial Interaction Patterns and Techniques</li> <li>A Taxonomy of Social Errors in Human-Robot Interaction</li> <li>Taxonomy of Digital Research Activities in the Humanities</li> <li>Virtual Reality and the CAVE: Taxonomy, Interaction Challenges and Research Directions&nbsp;</li> <li>Cross-Device Interaction</li> </ol>

opencc-by-4.0Feb 2024View details →
zenodo40/100

Fig. 5. Mean Gonadosomatic Index for C in Color pattern variation in Cichla temensis (Perciformes: Cichlidae): Resolution based on morphological, molecular, and reproductive data

Fig. 5. Mean Gonadosomatic Index for C. temensis variants grouped by CPV grade. a) Females from the Igapó Açú (Region 1). b) Females from the rio Caures (Region 2). c) Males from the Igapó Açú region. d) Males from the rio Caures. A significant correlation between GSI and CPV Grade was found for males and females in both collecting regions, p &lt;.01 for a, b, d, and d.

opencc-by-4.0Dec 2012View details →
zenodo40/100

Fig. 4 in Color pattern variation in Cichla temensis (Perciformes: Cichlidae): Resolution based on morphological, molecular, and reproductive data

Fig. 4. Maximum-likelihood phylogeny of 50 sequences sampled from the paca and açu variants of Cichla temensis (Genbank accession numbers HQ230011 - HQ230016) The phylogeny was rooted a posteriori with Cichla species of the clade A (sensu Willis et al. 2010) (GU295691- GU295704). The scale represents an HKY85 genetic distance.

opencc-by-4.0Dec 2012View details →
zenodo40/100

Fig. 3. a in Color pattern variation in Cichla temensis (Perciformes: Cichlidae): Resolution based on morphological, molecular, and reproductive data

Fig. 3. a) Mean (± SEM) lateral line scale counts for C. temensis, C. monoculus, and C. orinocensis. ANOVA showed no significant differences among the C. temensis variants but revealed significant differences interspecifically. Post hoc t tests (horizontal starred bar) revealed that all species were significantly different, p &lt;0.0001*. b) Mean (± SEM) body depth to Standard Length ratio (adjusted for gonad size differential) for C. temensis, C. monoculus, and C. orinocensis. ANOVA showed no significant differences among the C. temensis variants but revealed significant differences interspecifically. Post hoc t tests (horizontal starred bar) revealed that all C. temensis were significantly different from both sympatric species, p &lt;0.0001*.

opencc-by-4.0Dec 2012View details →
zenodo40/100

Fig. 2 in Color pattern variation in Cichla temensis (Perciformes: Cichlidae): Resolution based on morphological, molecular, and reproductive data

Fig. 2. Collecting regions in two cyclically flooding drainages in the rio Amazon basin. Region 1, the Igapó-Açu region, a blackwater tributary complex of the rio Madeira, provided specimens of C. temensis and C. monoculus. Region 2, the rio Caures, a blackwater tributary of the rio Negro, provided specimens of C. temensis and C. orinocensis.

opencc-by-4.0Dec 2012View details →
dryad40/100

Data from: Response to MHC-based olfactory cues in a mate choice context in two species of darter (Percidae: Etheostoma)

<p>Mate choice is hypothesized to play an important role in maintaining high diversity at major histocompatibility complex (MHC) genes in vertebrates. Many studies have revealed that females across taxa prefer the scent of males with MHC genotypes different to their own. In this study we tested the "opposites-attract" hypothesis in two species of darter with known differences in female criteria used in mate choice: in the fantail darters (a paternal-care species), females prefer males with visual traits related to nest guarding and egg tending, while in rainbow darters (not a paternal-care species) female mate choice criteria are unknown. In dichotomous mate-choice trials, we presented females of both species with the scents of conspecific males with MHC class IIb genotypes that were either similar or dissimilar to that of the focal female. We evaluated the proportion of time each female spent with each male and calculated the average strength of female preference for both species. Female fantail darters demonstrated a preference for the scent of males with similar (rather than dissimilar) MHC genotypes, but this result was not statistically significant. Rainbow darter females showed no preference for the scent of males with similar or dissimilar MHC genotypes. Our results do not support the "opposites-attract" hypothesis in darters.</p>

opencc-zeroFeb 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record