Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
110
datasets available to search
ShareScore release 0.9.0
Dataset results
110 results for “associative learning”
Exome sequence analysis identifies rare coding variants associated with a machine learning-based marker for coronary artery disease.
<p>*.sh and *.R are codes to test rare coding variants for association with ISCAD.</p> <p>Petrazzini_etal_2024_*_level_meta_analysis.txt.gz are summary statistics of variant- and gene-level associations of rare coding variants in the exome sequences of 604,914 individuals with an in-silico score for coronary artery disease (ISCAD).</p> <p>Chromosomal positions are mapped to the GRCh38 (hg38) human genome reference.</p> <p>Directions of effect correspond to associations in the UK Biobank, the All of Us Research Program, the BioMe Biobank sample 1 and the BioMe Biobank sample 2, in that order.</p>
Stereoisomers are not Machine Learning's Best Friends: Experimental results of the prediction of the association constant between a cyclodextrin and a guest with Stereo2vec
<p>This study addresses the challenge of accurately identifying stereoisomers in cheminformatics which originates from our objective to apply machine learning to predict association constant between a cyclodextrin and a guest. Identifying stereoisomers is indeed crucial for machine learning applications. Current tools offer various molecular descriptors, including their textual representation as Isomeric SMILES which can distinguish stereoisomers. But such representation is text-based and does not have a fixed size, so a conversion is needed to make it usable to machine learning approaches. Word embedding techniques can be used to solve this problem. Mol2vec, a word embedding approach for molecules, offers such a conversion. Unfortunately, it cannot distinguish between stereoisomers due to its inability to capture the spatial configuration of molecular structures. This study proposes several approaches that use word embedding techniques to handle molecular discrimination using stereochemical information of molecules or considering Isomeric SMILES notation as a text in Natural Language Processing. Our aim is to generate a distinct vector for each unique molecule, correctly identifying stereoisomer information in cheminformatics. The proposed approaches are then compared on our original machine learning task: predicting the association constant between a cyclodextrin and a guest molecule.</p>
RepliChrom: Interpretable machine learning predicts cancer-associated enhancer-promoter interactions using DNA replication timing
<p>This dataset accompanies the study "RepliChrom: Interpretable machine learning predicts cancer-associated enhancer-promoter interactions using DNA replication timing". The study introduces RepliChrom, a computational framework designed to predict enhancer–promoter interactions (EPIs) by leveraging multi-scale replication timing (RT) signals. This approach addresses the fundamental challenge of distinguishing gene targets regulated by distal enhancers from those activated by proximal transcriptional activity-a key problem in understanding the causal basis of complex diseases.</p> <p>Despite recent advances in high-throughput technologies such as Hi-C, ChIA-PET, and Hi-TrAC that allow genome-wide reconstruction of 3D chromatin architecture, the role of DNA replication timing in mediating these spatial interactions remains underexplored. RepliChrom fills this gap by using cell-type-specific RT profiles as predictive features for chromatin interaction inference.</p> <p>To support model development, training, and evaluation, we provide a comprehensive multi-omics dataset covering six human cell lines (K562, GM12878, HeLaS3, IMR90, NHEK, and HUVEC ), encompassing:</p> <p>Hi-C datasets: Processed chromatin interaction loops used to define positive and negative enhancer–promoter interaction pairs.</p> <p>ChIA-PET datasets: Interaction data anchored around transcription factor binding, including POLR2A and CTCF, used for model validation across different interaction types.</p> <p>Hi-TrAC datasets: Targeted chromatin accessibility-derived interaction data, offering complementary validation of the model on alternate platforms.</p> <p>Replication Timing (RT) data: Processed RT signal profiles for each cell line, used to extract multi-scale temporal features as inputs for RepliChrom.</p> <p>These datasets enable reproducibility of the model training process and serve as benchmark resources for future research into DNA replication–mediated regulation of 3D genome architecture.</p> <p><strong>Included Files</strong></p> <p>Hi-C_datasets.zip (2.23 MB): Processed Hi-C interaction pairs for six cell types.</p> <p>ChIA-PET_datasets.zip (818.18 KB): CTCF and POLR2A ChIA-PET interactions across multiple lines.</p> <p>Hi-TrAC_datasets.zip (196 bytes): Hi-TrAC-based chromatin interaction training datasets across multiple lines.</p> <p>Cellline_RT_data.zip (86.17 MB): Replication timing signal data across multiple human cell types for multi-scale replication timing feature extraction.</p> <p><strong>Usage</strong><br>All datasets are intended for academic, non-commercial use. The provided files can be directly used to reproduce the training and evaluation of RepliChrom, and may also support broader applications in enhancer–promoter modeling, replication-timing analysis, and 3D genomics studies. For detailed usage instructions and code implementation, please refer to the GitHub repository: https://github.com/DaoFuying/RepliChrom</p>
Cell-type specific responses to associative learning in the primary motor cortex
<p>The primary motor cortex (M1) is known to be a critical site for movement initiation and motor learning. Surprisingly, it has also been shown to possess reward-related activity, presumably to facilitate reward-based learning of new movements. However, whether reward-related signals are represented among different cell types in M1, and whether their response properties change after cue-reward conditioning remains unclear. Here, we performed longitudinal <i>in vivo</i> two-photon Ca<sup>2+</sup> imaging to monitor the activity of different neuronal cell types in M1 while mice engaged in a classical conditioning task. Our results demonstrate that most of the major neuronal cell types in M1 showed robust but differential responses to both cue and reward stimuli, and their response properties undergo cell-type specific modifications after associative learning. PV-INs' responses became more reliable to the cue stimulus, while VIP-INs' responses became more reliable to the reward stimulus. PNs only showed robust response to the novel reward stimulus, and they habituated to it after associative learning. Lastly, SOM-IN responses emerged and became more reliable to both conditioned cue and reward stimuli after conditioning. These observations suggest that cue- and reward-related signals are represented among different neuronal cell types in M1, and the distinct modifications they undergo during associative learning could be essential in triggering different aspects of local circuit reorganization in M1 during reward-based motor skill learning.</p>
Air quality data from the article "Typhoon-associated air quality over the Guangdong–Hong Kong–Macao Greater Bay Area, China: machine-learning-based prediction and assessment"
<p>This dataset consists of 26 files. The descriptions of the files are as follows:</p> <ul> <li>aqi_TY.csv, pm25_TY.csv, pm10_TY.csv, so2_TY.csv, no2_TY.csv and o3_TY.csv are the observed values of AQI and concentrations of PM<sub>2.5</sub>, PM<sub>10</sub>, SO<sub>2</sub>, NO<sub>2</sub> and O<sub>3</sub> of 36 monitoring stations used in model establish stage on TY days. The time range is June 2014 to December 2020.</li> <li>aqi_NTY.csv, pm25_NTY.csv, pm10_NTY.csv, so2_NTY.csv, no2_NTY.csv and o3_NTY.csv are the observed values of AQI and concentrations of PM<sub>2.5</sub>, PM<sub>10</sub>, SO<sub>2</sub>, NO<sub>2</sub> and O<sub>3</sub> of 36 monitoring stations used in model establish stage on NTY days. The time range is June 2014 to December 2020.</li> <li>station_info.csv is the detailed information of the 36 monitoring stations used in model establish stage, including station number, city, longitude and latitude.</li> <li>aqi_TY_testing.csv, pm25_TY_testing.csv, pm10_TY_testing.csv, so2_TY_testing.csv, no2_TY_testing.csv and o3_TY_testing.csv are the observed values of AQI and concentrations of PM<sub>2.5</sub>, PM<sub>10</sub>, SO<sub>2</sub>, NO<sub>2</sub> and O<sub>3</sub> of 3 monitoring stations used for testing the model on TY days. The time range is June 2014 to December 2020.</li> <li>aqi_NTY_testing.csv, pm25_NTY_testing.csv, pm10_NTY_testing.csv, so2_NTY_testing.csv, no2_NTY_testing.csv and o3_NTY_testing.csv are the observed values of AQI and concentrations of PM<sub>2.5</sub>, PM<sub>10</sub>, SO<sub>2</sub>, NO<sub>2</sub> and O<sub>3</sub> of 3 monitoring stations used for testing the model on NTY days. The time range is June 2014 to December 2020.</li> <li>sta_testing.csv is the detailed information of the 3 monitoring stations used for testing the model, including station number, city, longitude and latitude.</li> </ul>
Data for: Associative learning in the sea anemone Nematostella vectensis
<p>Supporting information for the manuscript "Associative learning in the sea anemone <em>Nematostella vectensis</em>".</p> <p>The following items can be found in this repository:</p> <p>1. <strong>Fig.1D_raw data.xlsx </strong></p> <p>Raw data for the Figure 1D. Manual counting of animals retracting for each condition.</p> <p><strong>2. Fig2CD_raw data.xlsx</strong></p> <p>Raw data for the Figure 2C and D. Output results of the tracking data analysis for each video file.</p> <p><strong>3. 220427_csv files.zip </strong></p> <p>Raw DLC output tracking files for 1 experiment (220427). To be used as an example to run the R markdown script below.</p> <p><strong>4. Tracking_data_Analysis_code.Rmd </strong></p> <p>This file is the R markdown file used to analyze the tracking data. It has been developed to analyze a batch of DLC output tracking files (.csv) for each experiment at once. The calculations carried out by the function are explained in the <em>SI appendix </em>and in the .pdf knitted version of the file below.</p> <p><strong>5. Tracking_data_Analysis_code.pdf </strong></p> <p>Knitted version of the R markdown file.</p>
Curated Dataset of Association Constants Between a Cyclodextrin and a Guest for Machine Learning: Raw Data and Generation Script
<p>Determining the association constant between a cyclodextrin and a guest molecule is an important task for various applications in various industrial and academical fields. However, such a task is time consuming, tedious and requires samples of both molecules. A significant number of association constants and relevant data is available from the literature. The availability of data makes the use of machine learning techniques to predict association constants possible. However, such data is mainly available from tables in articles or appendices. It is necessary to make them available in a computer friendly format and to curate them. Furthermore, the raw data need to be enriched with physicochemical information about each molecule and when such information does not allow to discriminate molecules, some additional data is needed. We present a dataset built from data gathered from the literature. The dataset contains both the original raw data from the articles and the enriched ones. We also provide the scripts used to curate and enrich the raw data.</p>
Plasticity of neuronal dynamics in the lateral habenula for cue-punishment associative learning.
<p>The brain’s ability to associate threats with external stimuli is vital to execute essential behaviours including avoidance. Disruption of this process contributes instead to the emergence of pathological traits which are common in addiction and depression. However, the mechanisms and neural dynamics at the single-cell resolution underlying the encoding of associative learning remain elusive. Here, employing a Pavlovian discrimination task in mice we investigate how neuronal populations in the lateral habenula (LHb), a subcortical nucleus whose excitation underlies negative affect, encode the association between conditioned stimuli and a punishment (unconditioned stimulus). </p>
Data for: Identifying regulators of associative learning using a protein-labelling approach in <em>C. elegans</em>
Open the record for dataset details and reuse information.
Data from: Machine learning identification of microhabitat features associated with occupancy of artificial nestboxes by hazel dormice (Muscardinus avellanarius) in a UK woodland site
Open the record for dataset details and reuse information.
Subjective executive functioning and skill learning during the COVID-19 pandemic associated with perceived loneliness, depressive symptoms, and well-being
Open the record for dataset details and reuse information.
Cell-type specific responses to associative learning in the primary motor cortex
Open the record for dataset details and reuse information.
Food discovery is associated with different reliance on social learning and lower cognitive flexibility across environments in a food caching bird
Open the record for dataset details and reuse information.
Dissociable control of unconditioned responses and associative fear learning by parabrachial CGRP neurons
<p>Parabrachial CGRP neurons receive diverse threat-related signals and contribute to multiple phases of adaptive threat responses in mice, with their inactivation attenuating both unconditioned behavioral responses to somatic pain and fear-memory formation. Because CGRP<sup>PBN</sup> neurons respond broadly to multi-modal threats, it remains unknown how these distinct adaptive processes are individually engaged. We show that while three partially separable subsets of CGRP<sup>PBN</sup> neurons broadly collateralize to their respective downstream partners, individual projections accomplish distinct functions: hypothalamic and extended amygdalar projections elicit assorted unconditioned threat responses including autonomic arousal, anxiety, and freezing behavior, while thalamic and basal forebrain projections generate freezing behavior and, unexpectedly, contribute to associative fear learning. Moreover, the unconditioned responses generated by individual projections are complementary, with simultaneous activation of multiple sites driving profound freezing behavior and bradycardia that are not elicited by any individual projection. This semi-parallel, scalable connectivity schema likely contributes to flexible control of threat responses in unpredictable environments.</p>
Heliconiini butterflies can learn time-dependent reward associations
For many pollinators, flowers provide predictable temporal schedules of resource availability, meaning an ability to learn time-dependent information could be widely beneficial. However, this ability has only been demonstrated in a handful of species. Observational studies of Heliconius butterflies suggest that they may have an ability to form time-dependent foraging preferences. Heliconius are unique among butterflies in actively collecting pollen, a dietary behaviour linked to spatiotemporally faithful 'trap-line' foraging. Time-dependency of foraging preferences is hypothesised to allow Heliconius to exploit temporal predictability in alternative pollen resources. Here, we provide the first experimental evidence in support of this hypothesis, demonstrating that Heliconius hecale can learn opposing colour preferences in two time periods. This shift in preference is robust to the order of presentation, suggesting that preference is tied to the time of day and not due to ordinal or interval learning. However, this ability is not limited to Heliconius, as previously hypothesised, but is also present in a related genus of non-pollen feeding butterflies. This demonstrates that time learning likely pre-dates the origin of pollen-feeding and may be prevalent across butterflies with less specialized foraging behaviours.
Vicarious reward unblocks associative learning about novel cues in male rats
<p>Many species, including humans, are sensitive to social signals and their valuation is important in social learning. When social cues indicate that a conspecific is experiencing reward, they could convey vicarious reward value and prompt social learning. Here, we introduce a task that investigates if mutual reward delivery in male rats can drive social reinforcement learning in a formal associative learning experiment. Using the blocking/unblocking paradigm, we found that when actor rats have fully learned a stimulus-self reward association, adding a cue that predicted additional reward to a partner unblocked associative learning about this cue. In contrast, additional cues that did not predict partner reward remained blocked from acquiring positive associative value. Importantly, this social unblocking effect was still present when controlling for secondary reinforcement but absent when social information exchange was impeded, when mutual reward outcomes were disadvantageously unequal to the actor or when the added cue predicted reward delivery to an empty chamber. Taken together, these results suggest that mutual rewards can drive associative learning in rats and is dependent on vicariously experienced social and food related cues.</p>
Data from: Developmental changes in hippocampal CA1 single neuron firing and theta activity during associative learning
Hippocampal development is thought to play a crucial role in the emergence of many forms of learning and memory, but ontogenetic changes in hippocampal activity during learning have not been examined thoroughly. We examined the ontogeny of hippocampal function by recording theta and single neuron activity from the dorsal hippocampal CA1 area while rat pups were trained in associative learning. Three different age groups [postnatal days (P)17-19, P21-23, and P24-26] were trained over six sessions using a tone conditioned stimulus (CS) and a periorbital stimulation unconditioned stimulus (US). Learning increased as a function of age, with the P21-23 and P24-26 groups learning faster than the P17-19 group. Age- and learning-related changes in both theta and single neuron activity were observed. CA1 pyramidal cells in the older age groups showed greater task-related activity than the P17-19 group during CS-US paired sessions. The proportion of trials with a significant theta (4–10 Hz) power change, the theta/delta ratio, and theta peak frequency also increased in an age-dependent manner. Finally, spike/theta phase-locking during the CS showed an age-related increase. The findings indicate substantial developmental changes in dorsal hippocampal function that may play a role in the ontogeny of learning and memory.
Data from: Social foraging extends associative odor-food memory expression in an automated learning assay for Drosophila melanogaster
<p>Animals socially interact during foraging and share information about the quality and location of food sources. The mechanisms of social information transfer during foraging have been mostly studied at the behavioral level, and its underlying neural mechanisms are largely unknown. Fruit flies have become a model for studying the neural bases of social information transfer, because they provide a large genetic toolbox to monitor and manipulate neuronal activity, and they show a rich repertoire of social behaviors. Fruit flies aggregate, they use social information for choosing a suitable mating partner and oviposition site, and they show better aversive learning when in groups. However, the effects of social interactions on associative odor–food learning have not yet been investigated. Here, we present an automated learning and memory assay for walking flies that allows the study of the effect of group size on social interactions and on the formation and expression of associative odor–food memories. We found that both inter-fly attraction and the duration of odor–food memory expression increase with group size. This study opens up opportunities to investigate how social interactions during foraging are relayed in the neural circuitry of learning and memory expression.</p>
Dataset associated with "Emulating subglacial hydrology in ice sheet models with deep learning methods" by Verjans and Robel.
<p>See Readme file for descriptions.</p>
Supplementary Information: CHAPTER 3 - Classification of genomic features of plant-associated bacteria using machine learning
<p>Appendix A- List of all bacterial genomes used in orthologous genes clustering in the feature extraction step and in the further steps to build and test classifiers’ models. The list includes the isolation source information and the related category for the genome classification and features selection purposes.</p> <p>Appendix B - Distribution of genomes by phylum, family, and genus among the categories defined according to bacteria lifestyle association.</p> <p>Appendix C - Enriched orthogroups by genus according to each enrichment test (Material and Methods). Values for each test are "Y" (enriched), "N" (not enriched), or "Untested" (clusters were untested when there was insufficient phylogenetic signal, they were too small or were found in all genomes).</p> <p>Appendix D - Classification performance of random forest and logistic regression techniques applied to genus-specific datasets of genomic features (orthogroups) using both matrices from gene count number and presence/absence values. Sensitivity is a measure of how well a test identifies true positives; Specificity: is a measure how well a test or model avoids false positives; Positive Predictive Value (Pos. Pred. Value): The probability that a positive prediction is correct; Negative Predictive Value (Neg. Pred. Value): The probability that a negative prediction is correct; Precision: The accuracy of positive predictions; Recall (Sensitivity): The ability to find all relevant cases; F1 Score: A combined measure of precision and recall; Prevalence: The proportion of positive cases in the total; Detection Rate: The proportion of true positive cases identified; Detection Prevalence: The proportion of positive predictions; Balanced Accuracy: An average of sensitivity and specificity; Area Under the Curve (AUC): The overall performance of the model in distinguishing between positive and negative cases.</p> <p>Appendix E - Orthogroups assigned with predicted COGs as an important feature for classifying plant-associated genomes. COG categories: A - RNA processing and modification; B - Chromatin structure and dynamics; C - Energy production and conversion; D - Cell cycle control, cell division, chromosome partitioning; E - Amino acid transport and metabolism; F - Nucleotide transport and metabolism; G - Carbohydrate transport and metabolism; H - Coenzyme transport and metabolism; I - Lipid transport and metabolism; J - Translation, ribosomal structure and biogenesis; K - Transcription; L - Replication, recombination and repair; M - Cell wall/membrane/envelope biogenesis; N - Cell motility; O - Posttranslational modification, protein turnover, chaperones; P - Inorganic ion transport and metabolism; Q - Secondary metabolites biosynthesis, transport and catabolism; R - General function prediction only; S - Function unknown; T - Signal transduction mechanisms; U - Intracellular trafficking, secretion, and vesicular transport; V - Defense mechanisms; W - Extracellular structures; X - Mobilome: prophages, transposons; Y - Nuclear structure; Z - Cytoskeleton.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.