Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
84
datasets available to search
ShareScore release 0.7.1
Dataset results
84 results for “feedback control”
WhiB6 regulation of ESX-1 gene expression is controlled by a negative feedback loop in Mycobacterium marinum
GEO Series GSE99632. Mycobacterium marinum. 10 samples. Type: Expression profiling by high throughput sequencing.
Chimeric activators and repressors define HY5 activity and a feedback control mechanism in Arabidopsis (HY5_ChIPseq)
GEO Series GSE132860. Arabidopsis thaliana. 11 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.
Robustness in the feedback control of the retinoic acid network response to environmental disturbances (BioMark-qRT-PCR)
GEO Series GSE154407. Xenopus laevis. 72 samples. Type: Expression profiling by RT-PCR.
Asymmetric robustness in the feedback control of the retinoic acid network response to environmental disturbances
GEO Series GSE154408. Xenopus laevis. 168 samples. Type: Expression profiling by RT-PCR; Expression profiling by high throughput sequencing.
A super enhancer-controlled BRD4/ERα-RET-ERα positive feedback loop promotes ERα-positive breast cancer [RNA-seq]
GEO Series GSE186644. Homo sapiens. 2 samples. Type: Expression profiling by high throughput sequencing.
Feedback control of Set1 protein levels is important for proper H3K4 methylation patterns
GEO Series GSE55082. Saccharomyces cerevisiae. 8 samples. Type: Genome binding/occupancy profiling by array.
A super enhancer-controlled BRD4/ERα-RET-ERα positive feedback loop promotes ERα-positive breast cancer [GRO-seq]
GEO Series GSE186643. Homo sapiens. 4 samples. Type: Other.
A super enhancer-controlled BRD4/ERα-RET-ERα positive feedback loop promotes ERα-positive breast cancer
GEO Series GSE186646. Homo sapiens. 8 samples. Type: Other; Expression profiling by high throughput sequencing; Genome binding/occupancy profiling by high throughput sequencing.
Negative feedback control of neuronal activity by microglia
GEO Series GSE149897. Mus musculus. 36 samples. Type: Expression profiling by high throughput sequencing.
An Esrrb and Nanog Cell Fate Regulatory Module Controlled by Feedback Interactions [ChIP-seq]
GEO Series GSE31640. Mus musculus. 14 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.
Translational regulation of specific mRNAs controls feedback inhibition and survival during macrophage activation [array]
GEO Series GSE52449. Mus musculus. 30 samples. Type: Expression profiling by array.
Epigenomic analyses of FOXA1-mediated enhancer reprogramming to control tumor subtype plasticity through HH-BMP-FOXA1 signalling feedback between urothelial tumor and stromal fibroblasts [FOXA1 ChIP-s
GEO Series GSE141354. Homo sapiens. 20 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.
Epigenomic analyses of FOXA1-mediated enhancer reprogramming to control tumor subtype plasticity through HH-BMP-FOXA1 signalling feedback between urothelial tumor and stromal fibroblasts [ATAC-seq]
GEO Series GSE141351. Homo sapiens. 16 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.
An Esrrb and Nanog Cell Fate Regulatory Module Controlled by Feedback Interactions [methylation]
GEO Series GSE33191. Mus musculus. 24 samples. Type: Methylation profiling by genome tiling array.
Epigenomic analyses of FOXA1-mediated enhancer reprogramming to control tumor subtype plasticity through HH-BMP-FOXA1 signalling feedback between urothelial tumor and stromal fibroblasts [H3K27ac ChIP
GEO Series GSE148819. Homo sapiens. 20 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.
A lncRNA identifies Irf8 enhancer element in negative feedback control of dendritic cell differentiation
GEO Series GSE198651. Mus musculus. 25 samples. Type: Expression profiling by high throughput sequencing; Genome binding/occupancy profiling by high throughput sequencing.
CMFeed: A Benchmark Dataset for Controllable Multimodal Feedback Synthesis
<p><strong>Overview</strong><br>The Controllable Multimodal Feedback Synthesis (CMFeed) Dataset is designed to enable the generation of sentiment-controlled feedback from multimodal inputs, including text and images. This dataset can be used to train feedback synthesis models in both uncontrolled and sentiment-controlled manners. Serving a crucial role in advancing research, the CMFeed dataset supports the development of human-like feedback synthesis, a novel task defined by the dataset's authors. Additionally, the corresponding feedback synthesis models and benchmark results are presented in the <a href="https://github.com/MIntelligence-Group/CMFeed/" target="_blank" rel="noopener">associated code and research publication</a>. <br><br><em>Task Uniqueness</em>: The task of controllable multimodal feedback synthesis is unique, distinct from LLMs and tasks like VisDial, and not addressed by multi-modal LLMs. LLMs often exhibit errors and hallucinations, as evidenced by their auto-regressive and black-box nature, which can obscure the influence of different modalities on the generated responses [<a href="https://www.nature.com/articles/s41586-024-07421-0" target="_blank" rel="noopener">Ref1</a>; <a href="https://arxiv.org/abs/2311.05232" target="_blank" rel="noopener">Ref2</a>]. Our approach includes an interpretability mechanism, as detailed in the supplementary material of the corresponding <a href="https://github.com/MIntelligence-Group/CMFeed/" target="_blank" rel="noopener">research publication</a>, demonstrating how metadata and multimodal features shape responses and learn sentiments. This controllability and interpretability aim to inspire new methodologies in related fields.<br><br><strong>Data Collection and Annotation</strong><br>Data was collected by crawling Facebook posts from major news outlets, adhering to ethical and legal standards. The comments were annotated using four sentiment analysis models: FLAIR, SentimentR, RoBERTa, and DistilBERT. Facebook was chosen for dataset construction because of the following factors:<br>• Facebook was chosen for data collection because it uniquely provides metadata such as news article link, post shares, post reaction, comment like, comment rank, comment reaction rank, and relevance scores, not available on other platforms. <br>• Facebook is the most used social media platform, with 3.07 billion monthly users, compared to 550 million Twitter and 500 million Reddit users. [<a href="https://en.wikipedia.org/wiki/List_of_social_platforms_with_at_least_100_million_active_users" target="_blank" rel="noopener">Ref</a>] <br>• Facebook is popular across all age groups (18-29, 30-49, 50-64, 65+), with at least 58% usage, compared to 6% for Twitter and 3% for Reddit. [<a href="https://sproutsocial.com/insights/new-social-media-demographics/" target="_blank" rel="noopener">Ref</a>]. Trends are similar for gender, race, ethnicity, income, education, community, and political affiliation [<a href="https://pewresearch.org/internet/fact-sheet/social-media/" target="_blank" rel="noopener">Ref</a>] <br>• The male-to-female user ratio on Facebook is 56.3% to 43.7%; on Twitter, it's 66.72% to 23.28%; Reddit does not report this data. [<a href="https://khoros.com/resources/social-media-demographics-guide" target="_blank" rel="noopener">Ref</a>]</p> <p><em>Filtering Process</em>: To ensure high-quality and reliable data, the dataset underwent two levels of filtering:<br>a) Model Agreement Filtering: Retained only comments where at least three out of the four models agreed on the sentiment.<br>b) Probability Range Safety Margin: Comments with a sentiment probability between 0.49 and 0.51, indicating low confidence in sentiment classification, were excluded.<br>After filtering, 4,512 samples were marked as XX. Though these samples have been released for the reader's understanding, they were not used in training the feedback synthesis model proposed in the corresponding research paper.<br><br><strong>Dataset Description</strong><br>• Total Samples: 61,734<br>• Total Samples Annotated: 57,222 after filtering.<br>• Total Posts: 3,646<br>• Average Likes per Post: 65.1<br>• Average Likes per Comment: 10.5<br>• Average Length of News Text: 655 words<br>• Average Number of Images per Post: 3.7<br><br><strong>Components of the Dataset</strong><br>The dataset comprises two main components:<br>• <em>CMFeed.csv</em> File: Contains metadata, comment, and reaction details related to each post.<br>• <em>Images</em> Folder: Contains folders with images corresponding to each post.<br><br><strong>Data Format and Fields of the CSV File</strong><br>The dataset is structured in CMFeed.csv file along with corresponding images in related folders. This CSV file includes the following fields:<br>• <em>Id</em>: Unique identifier <br>• <em>Post</em>: The heading of the news article.<br>• <em>News_text</em>: The text of the news article.<br>• <em>News_link</em>: URL link to the original news article.<br>• <em>News_Images</em>: A path to the folder containing images related to the post.<br>• <em>Post_shares</em>: Number of times the post has been shared.<br>• <em>Post_reaction</em>: A JSON object capturing reactions (like, love, etc.) to the post and their counts.<br>• <em>Comment</em>: Text of the user comment.<br>• <em>Comment_like</em>: Number of likes on the comment.<br>• <em>Comment_reaction_rank</em>: A JSON object detailing the type and count of reactions the comment received.<br>• <em>Comment_link</em>: URL link to the original comment on Facebook.<br>• <em>Comment_rank</em>: Rank of the comment based on engagement and relevance.<br>• <em>Score</em>: Sentiment score computed based on the consensus of sentiment analysis models.<br>• <em>Agreement</em>: Indicates the consensus level among the sentiment models, ranging from -4 (all negative) to 4 (all positive). 3 negative and 1 positive will result into -2 and 3 positives and 1 negative will result into +2.<br>• <em>Sentiment_class</em>: Categorizes the sentiment of the comment into 1 (positive) or 0 (negative).<br><br><strong>More Considerations During Dataset Construction<br></strong>We thoroughly considered issues such as the choice of social media platform for data collection, bias and generalizability of the data, selection of news handles/websites, ethical protocols, privacy and potential misuse before beginning data collection. While achieving completely unbiased and fair data is unattainable, we endeavored to minimize biases and ensure as much generalizability as possible. Building on these considerations, we made the following decisions about data sources and handling to ensure the integrity and utility of the dataset:<em><br><br>• Why not merge data from different social media platforms? </em>We chose not to merge data from platforms such as Reddit and Twitter with Facebook due to the lack of comprehensive metadata, clear ethical guidelines, and control mechanisms—such as who can comment and whether users' anonymity is maintained—on these platforms other than Facebook. These factors are critical for our analysis. Our focus on Facebook alone was crucial to ensure consistency in data quality and format.</p> <p><em>• Choice of four news handles</em><strong>:</strong> We selected four news handles—BBC News, Sky News, Fox News, and NY Daily News—to ensure diversity and comprehensive regional coverage. These news outlets were chosen for their distinct regional focuses and editorial perspectives: BBC News is known for its global coverage with a centrist view, Sky News offers geographically targeted and politically varied content learning center/right in the UK/EU/US, Fox News is recognized for its right-leaning content in the US, and NY Daily News provides left-leaning coverage in New York. Many other news handles such as NDTV, The Hindu, Xinhua, and SCMP are also large-scale but may contain information in regional languages such as Indian and Chinese, hence, they have not been selected. This selection ensures a broad spectrum of political discourse and audience engagement.</p> <p><em>• Dataset Generalizability and Bias</em><strong>:</strong> With 3.07 billion of the total 5 billion social media users, the extensive user base of Facebook, reflective of broader social media engagement patterns, ensures that the insights gained are applicable across various platforms, reducing bias and strengthening the generalizability of our findings. Additionally, the geographic and political diversity of these news sources, ranging from local (NY Daily News) to international (BBC News), and spanning political spectra from left (NY Daily News) to right (Fox News), ensures a balanced representation of global and political viewpoints in our dataset. This approach not only mitigates regional and ideological biases but also enriches the dataset with a wide array of perspectives, further solidifying the robustness and applicability of our research.</p> <p><em>• Dataset size and diversity:</em> Facebook prohibits the automatic scraping of its users' personal data. In compliance with this policy, we manually scraped publicly available data. This labor-intensive process requiring around 800 hours of manual effort, limited our data volume but allowed for precise selection. We followed ethical protocols for scraping Facebook data , selecting 1000 posts from each of the four news handles to enhance diversity and reduce bias. Initially, 4000 posts were collected; after preprocessing (detailed in Section 3.1), 3646 posts remained. We then processed all associated comments, resulting in a total of 61734 comments. This manual method ensures adherence to Facebook’s policies and the integrity of our dataset.</p> <p><strong>Ethical considerations, data privacy and misuse prevention<br></strong>The data collection adheres to Facebook’s ethical guidelines [<a href="https://developers.facebook.com/terms/" target="_blank" rel="noopener">Ref</a>]. We manually scraped publicly available data in compliance with Facebook's ethical guidelines prohibiting automatic scraping [<a href="https://www.facebook.com/help/463983701520800" target="_blank" rel="noopener">Ref</a>]. We collected data that is publicly available, specifically corresponding to news articles that are publicly accessible following the protocols for ethically scraping Facebook data [<a href="https://webscraping.blog/how-to-scrape-facebook/" target="_blank" rel="noopener">Ref</a>]. The human-generated comments are included without identifying information. Aiming to prevent potential misuse, we have proactively designed our feedback synthesis system with an integral interpretability module. This feature is crucial as it not only helps in explaining how decisions are made within the system but also in detecting and preventing any misuse, such as the creation of misleading or manipulative content. Our original motivation for integrating this technology was to ensure that it is used responsibly and ethically, enhancing its positive impact while minimizing risks. By focusing on developing robust and transparent systems, we aim to foster trust and encourage the responsible use of technology in line with our ethical commitments.<br><br><strong>Code and Citation</strong><br>• Code Repository: <a href="https://github.com/MIntelligence-Group/CMFeed/" target="_blank" rel="noopener">https://github.com/MIntelligence-Group/CMFeed/</a><br>• Citing the Dataset: Users of the dataset should cite the corresponding paper described at the above GitHub Repository.<br><br><strong>License & Access</strong><br>• This dataset is released for academic research only and is free to researchers from educational or research institutes for non-commercial purposes.<br>• Note that you are downloading this corpus at your own risk. No guarantee is provided, e.g. regarding the goodness of the corpus nor towards any subsequent effects. You may use it free of charge, and modify it as you wish, but clearly specify modifications if you pass modified material on.<br><br><strong>Contact</strong><br>Please send any questions about this dataset to:<br>• Puneet Kumar (puneet.kumar@oulu.fi),<br>• Sarthak Malik (sarthak_m@mt.iitr.ac.in),<br>• Balasubramanian Raman (bala@cs.iitr.ac.in),<br>• Xiaobai Li (xiaobai.li@zju.edu.cn).</p>
Splitting Function enables Dual Feedback Regulation to Control JAK2/STAT5 Signaling for a Wide Ligand Range
GEO Series GSE26151. Mus musculus. 20 samples. Type: Expression profiling by array.
A super enhancer-controlled BRD4/ERα-RET-ERα positive feedback loop promotes ERα-positive breast cancer [ChIP-seq]
GEO Series GSE186645. Homo sapiens. 2 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.
A feedback loop between heterochromatin and the nucleopore complex controls germ-cell to oocyte transition during Drosophila oogenesis
GEO Series GSE186982. Drosophila melanogaster. 22 samples. Type: Expression profiling by high throughput sequencing; Genome binding/occupancy profiling by high throughput sequencing.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.