Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
94
datasets available to search
ShareScore release 0.9.0
Dataset results
94 results for “Generative AI”
Human-AI Collaboration: A tool to enable AI model generation with human-in-the-loop
<p>Human-AI collaboration enables domain experts to contribute their expertise with the goal of enhancing the knowledge learned by the AI models from the patterns in the data. This enables the integration of domain-specific knowledge to enrich the data for further improvement of the models through retraining. The human-AI collaboration is composed of multiple sub-components and interfaces that enables communication with external systems such as data sources, model repositories, machine configurations and decision support systems.</p> <p>Human-AI Collaboration component is developed using Python programming language. The frontend is developed using Streamlit1. The backend is developed using python and the API is implemented using FastAPI2. The choice of the programming language was made because of its wide usage and vast user base. The frameworks Streamlit and FastAPI are chosen because of the rich features for functionality and documentation as well as suitability for data analysis tasks. The applications are packaged as docker images for deployment. The application runs as a web application served by nginx for reverseproxying and users can access it via client applications such as web browsers or REST clients like Postman.</p>
Trypanosoma Epitope Dataset: Valid Epitopes and Randomly Generated Peptides with Biochemical Metrics and AI-Generated Scores
<p>This dataset contains information about valid linear B-Cell epitopes from the Trypanosoma genus, as well as randomly generated peptides. It includes biochemical metrics generated by the EpiBuilder-1.0 tool and scores generated by the BepiPred-3.0 software. The data was originally collected from the IEDB and UniProtKB platforms and has been processed and enhanced with these informations for researchers interested in understanding the molecular interactions between Trypanosoma protozoans and the immune system.</p>
Relationships among the generated content type, task, AI technique, and application domain in the retrieved works
<p>Relationships among the generated content type, task, AI technique, and application domain in the retrieved works. Part of the study "What do we mean by GenAI?"</p>
A collection of AI generated images visualising various RDM aspects
<p>This publication contains images visualising various RDM aspects. These images were generated by the <a href="https://www.forschungsdaten.uni-bonn.de/en" target="_blank" rel="noopener">Research Data Service Center</a> team at the University of Bonn and are used in the workshop "Research Data Management: A Crash Course" conducted since 2021 by the Research Data Service Center. The slide deck is available as a related publication (see the related works section below for details).</p> <p>The images were generated with the help of <a href="https://help.openai.com/en/articles/8932459-creating-images-in-chatgpt">ChatGPT</a>. </p> <p>In this version, due to legal reasons, we changed the images.</p>
Generation of a network slicing dataset: the foundations for AI-based B5G resource management
<p><span>This paper introduces a comprehensive network slicing dataset designed to empower artificial intelligence (AI), and other data-based resource management and network performance prediction applications, in 5G and beyond (B5G) networks. The dataset, generated through a packet-level simulator, captures the complexities of network slicing considering the three main network slice types defined by 3GPP: Enhanced Mobile Broadband (eMBB), Ultra-Reliable Low Latency Communications (URLLC), and Massive Internet of Things (mIoT). It includes a wide range of network scenarios with varying topologies, slice instances, and traffic flows. The included scenarios consist of transport networks, excluding the RAN infrastructure.</span></p> <p><span>Each sample consists of pairs of (network scenario, performance metrics). The network configuration includes network topology, traffic characteristics, routing configurations, while the performance metrics are the delay, jitter, and loss for each flow. The dataset is generated with a custom network slicing admission control module, enabling the simulation of realistic scenarios without violating SLAs.</span></p> <p><span>This network slicing dataset is a valuable asset for the research community, unlocking opportunities for innovations in 5G and B5G networks.</span></p>
TWIGMA: A dataset of AI-Generated Images with Metadata From Twitter
<p><strong>Update May 2024: Fixed a data type issue with "id" column that prevented twitter ids from rendering correctly.</strong></p> <p>Recent progress in generative artificial intelligence (gen-AI) has enabled the generation of photo-realistic and artistically-inspiring photos at a single click, catering to millions of users online. To explore how people use gen-AI models such as DALLE and StableDiffusion, it is critical to understand the themes, contents, and variations present in the AI-generated photos. In this work, we introduce TWIGMA (TWItter Generative-ai images with MetadatA), a comprehensive dataset encompassing 800,000 gen-AI images collected from Jan 2021 to March 2023 on Twitter, with associated metadata (e.g., tweet text, creation date, number of likes).</p> <p>Through a comparative analysis of TWIGMA with natural images and human artwork, we find that gen-AI images possess distinctive characteristics and exhibit, on average, lower variability when compared to their non-gen-AI counterparts. Additionally, we find that the similarity between a gen-AI image and human images (i) is correlated with the number of likes; and (ii) can be used to identify human images that served as inspiration for the gen-AI creations. Finally, we observe a longitudinal shift in the themes of AI-generated images on Twitter, with users increasingly sharing artistically sophisticated content such as intricate human portraits, whereas their interest in simple subjects such as natural scenes and animals has decreased. Our analyses and findings underscore the significance of TWIGMA as a unique data resource for studying AI-generated images.</p> <p>Note that in accordance with the privacy and control policy of Twitter, <strong>NO raw content from Twitter is included</strong> in this dataset and users could and need to retrieve the original Twitter content used for analysis using the Twitter id. In addition, users who want to access Twitter data should consult and follow rules and regulations closely at the official Twitter developer policy at https://developer.twitter.com/en/developer-terms/policy. </p> <p> </p> <p> </p>
RU-AI: A Large Multimodal Dataset for Machine Generated Content Detection
<p>This repository contains all the collected and aligned data for RU-AI dataset. It is constructed based on three large publicly available datasets: Flickr8K, COCO, and Places205, by adding their corresponding machine-generated pairs from five different generative models in each modality. </p>
Nanoparticle Size Estimation by Scanning Transmission Electron Microscopy and Generative AI
<p>The "raw" directories contain unaltered simulated and experimental data. The train and val directories contain normalized data used to train the models of the manuscript. The dataframes directory contains all information about the atomic models. Exp info contains info about the raw experimental data (excluding the gas-cell data). </p>
Exploring New Possibilities of Life with Generative AI
<h2><a name="_Toc171508708"></a>ReLife - Prologue</h2> <p><strong>Abstract:</strong> ReLife is an innovative academic research project that aims to explore the frontiers of artificial intelligence, neuroscience, and virtual reality to offer a new chance at life for individuals in critical situations. The project's objective is to allow users to revisit a crucial point in their past and experience an alternative life generated by AI, providing a new perspective on existence.</p> <p>Advancements in artificial intelligence and virtual reality technologies have opened up new possibilities for applications in areas once thought unimaginable. The ReLife project emerges with the vision of utilizing these technologies to create an alternative life experience for people who, due to critical situations, seek a new opportunity to live.</p> <p>The core concept of ReLife is to allow a user to choose a crucial point in their past and make a different decision without perceiving they are in a simulation. From this new choice, generative AI creates an alternative life path, which the user lives in an immersive and interactive virtual reality environment.</p> <p><strong>Tech Objectives:</strong></p> <ul> <li><strong>Data Collection:</strong> Capture and securely store users' personal data, histories, and memories efficiently.</li> <li><strong>Infrastructure and Support:</strong> Develop the technical infrastructure necessary to support the continuous operation of ReLife.</li> <li><strong>Ethical and Legal Aspects:</strong> Establish ethical guidelines and legal compliance policies for consciousness transfer and data usage.</li> <li><strong>Security and Privacy:</strong> Ensure the protection of users' data against unauthorized access.</li> <li><strong>Practical Applications:</strong> Identify practical use cases and develop case studies to demonstrate the benefits of ReLife.</li> <li><strong>User Interface:</strong> Create intuitive and user-friendly interfaces to facilitate user interaction with the system.</li> <li><strong>Consciousness Transfer:</strong> Research and develop technologies for consciousness transfer and neural mapping.</li> <li><strong>Simulation and Virtual Reality:</strong> Create realistic and immersive virtual environments where users can live their new lives.</li> <li><strong>AI Modeling:</strong> Develop AI algorithms to generate alternative life scenarios based on users' decisions.</li> </ul> <p><strong>Conclusion:</strong> ReLife represents a significant step in exploring how technology can transform our lives, offering new opportunities and hope to those who need it most. Through a multidisciplinary approach, the project aims not only to advance technical knowledge but also to address the complex ethical and social issues associated with these innovations.</p> <h2><a name="_Toc171508709"></a>Introduction</h2> <h3><a name="_Toc171508710"></a>Context and Motivation</h3> <p>The ReLife project emerges at a time when advancements in artificial intelligence and virtual reality are challenging the limits of what is possible. The motivation behind this project is twofold: to stimulate the study and evolution of Artificial Intelligence, and to offer a new chance for those who, due to some unfortunate reason or choice, have had their lives drastically affected.</p> <p>In the era of science fiction, the idea of revisiting the past and making different decisions seemed unattainable. However, with current technological developments, we are approaching this possibility. ReLife aims to extend human consciousness through technology, breaking the physical barriers that limit our existence. This project not only challenges the status quo but also explores new frontiers of human knowledge, making previously unimaginable heights a reality.</p> <p>Through a multidisciplinary approach, encompassing neuroscience, artificial intelligence, and virtual reality, ReLife aims to create an immersive and realistic alternative life experience. This not only offers new perspectives of existence but also contributes to scientific and technological research, promoting significant advancement in the field of AI and its applications.</p> <h2><a name="_Toc171508711"></a>Assumptions of the ReLife Project</h2> <ol> <li><strong>Brain in Good Condition</strong>:</li> <ul> <li><strong>Description</strong>: The user must have a brain in good condition, without severe brain injuries, to allow effective and accurate cognitive mapping.</li> <li><strong>Justification</strong>: Brain injuries can hinder neural mapping and consciousness transfer, compromising the quality of the simulation.</li> </ul> <li><strong>Technological Advancements</strong>:</li> <ul> <li><strong>Description</strong>: Significant advancements in neuroscience, neural mapping technology, and generative AI are assumed.</li> <li><strong>Justification</strong>: Advanced technologies are essential for data collection, AI modeling, and creating realistic simulations.</li> </ul> <li><strong>Support Infrastructure</strong>:</li> <ul> <li><strong>Description</strong>: A robust infrastructure is necessary to support the storage and processing of large volumes of data.</li> <li><strong>Justification</strong>: Ensuring the system functions efficiently and securely.</li> </ul> <li><strong>Ethical and Legal Aspects</strong>:</li> <ul> <li><strong>Description</strong>: Compliance with all ethical and legal standards related to consciousness transfer and the use of personal data.</li> <li><strong>Justification</strong>: Protecting users' rights and ensuring regulatory compliance.</li> </ul> <li><strong>Security and Privacy</strong>:</li> <ul> <li><strong>Description</strong>: Implementation of robust security and privacy measures to protect users' data.</li> <li><strong>Justification</strong>: Ensuring the integrity and confidentiality of users' information.</li> </ul> <li><strong>Immersive Environment and Simulation</strong>:</li> <ul> <li><strong>Description</strong>: Creation of realistic and immersive virtual environments where users can live their new lives without realizing they are in a simulation.</li> <li><strong>Justification</strong>: Ensuring that the simulations are realistic and coherent with users' daily experiences, providing complete immersion.</li> </ul> </ol> <p>For more details, download the attached document</p>
Investigating the Use of AI-Generated Exercises for Beginner and Intermediate Programming Courses: A ChatGPT Case Study
<p>In recent years, artificial intelligence (AI) has been increasingly used in education and supports teachers in creating educational material and students in their learning progress. AI- driven learning support has recently been further strengthened by the release of ChatGPT, in which users can retrieve expla- nations for various concepts in a few minutes through chat. However, to what extent the use of AI models, such as ChatGPT, is suitable for the creation of didactically and content-wise good exercises for programming courses is not yet known. Therefore, in this paper, we investigate the use of AI-generated exercises for beginner and intermediate programming courses in higher education using ChatGPT. We created 12 exercise sheets with ChatGPT for a beginner to intermediate programming course focusing on the objects-first approach. We report our process, prompts, and experience using ChatGPT for this task and outline good practices we identified. The generated exercises are assessed and revised, primarily using ChatGPT, until they met the requirements of the programming course. We assessed the quality of these exercises by using them in our course as external teaching assignment at the University of Education Ludwigsburg and let the students evaluate them. Results indicate the quality of the generated exercises and the time-saving for creating them using ChatGPT. However, our experience showed that while it is fast to generate a good version of an exercise, almost every exercise requires minor manual changes to improve its quality.</p>
Generative AI in University Communication - Survey Data (June 2023)
<p>Der Datensatz mit dem Titel "Generative KI in der Hochschulkommunikation - Umfragedaten (Juni 2023)" erfasst Informationen zur Einführung und Nutzung von generativer künstlicher Intelligenz (KI) im Kontext der Hochschulkommunikation. Die Umfrage, die im Juni 2023 unter 318 deutschen Hochschulen durchgeführt wurde, von denen 101 geantwortet haben, untersucht verschiedene Aspekte, darunter Bekanntheit und Wissen über verschiedene KI-Tools (z.B. ChatGPT), Diskussionen in Gremien, das Vorhandensein von Richtlinien für die Nutzung, das Vorhandensein von Arbeitsgruppen für generative KI, strategische Ziele und Initiativen, Schulungsangebote für generative KI-Tools und die wahrgenommene Bedeutung von generativen KI-Tools in der Hochschulkommunikation. Ziel des Datensatzes ist es, Einblicke in die aktuelle Landschaft und Praxis der Integration generativer KI im universitären Umfeld zu geben.</p><p>The dataset, titled "Generative AI in University Communication - Survey Data (June 2023)," captures information related to the adoption and utilization of generative artificial intelligence (AI) in the context of university communication. This survey, conducted in June 2023 among 318 German universities of which 101 responded, explores various aspects, including awarenes and knowledge of various AI tools (e.g. ChatGPT), discussions in committees, the existence of guidelines for usage, the presence of working groups for generative AI, strategic goals and initiatives, training offerings for generative AI tools, and the perceived importance of generative AI tools in university communication. The dataset aims to provide insights into the current landscape and practices regarding the integration of generative AI within university settings.</p><p> </p><p>More information here: <a href="https://www.hof.uni-halle.de/projekte/hochki/">https://www.hof.uni-halle.de/projekte/hochki/</a></p>
Reproducibility of "Diffusion-based Generative AI for Exploring Transition States from 2D Molecular Graphs"
<p>This file is the source data to ensure reproducibility of the paper "Diffusion-based Generative AI for Exploring Transition States from 2D Molecular Graphs". It contains the logs and results of all DFT calculations associated with transition states generated using the model proposed in the paper. It also includes code to reproduce the core findings of the paper, which can be done by running reproduce.sh. To accurately reproduce the results of the paper, use the v1.0.0 virtual environment from "https://github.com/seonghann/tsdiff".</p>
Generative AI in the Advancement of Viral Therapeutics for Predicting and Targeting Immune-Evasive SARS-CoV-2 Mutations
<p>This dataset <strong>encompasses</strong> and describes the following features:</p> <ul> <li>Mutations in viruses like SARS-CoV-2 can make them escape vaccines and treatments.</li> <li>Accurately predicting these mutations is crucial for developing effective countermeasures.</li> <li>The study uses a type of AI called a Generative Adversarial Network (GAN) to analyze the virus's spike protein, which plays a key role in infection.</li> <li>The GAN generates protein sequences similar to natural ones, but which are also likely to evade immune responses.</li> <li>By analyzing these generated sequences, the researchers improve their AI model's ability to predict real-world escape mutations.</li> <li>This improved prediction could help design better vaccines and treatments, and prepare for future viral threats.</li> </ul>
Generative AI in Crowdwork: A Survey of Workers at Three Platforms
<p>Dataset surveying the use and opinions on Generative AI by crowdworkers over three crowdsourcing platforms and across three continents. </p>
Generative AI in Crowdwork for Web and Social Media Research: A Survey of Workers at Three Platforms
<p>Data generated for the paper "<span>Generative AI in Crowdwork for Web and Social Media Research:</span><br><span>A Survey of Workers at Three Platforms</span>" presented at ICWSM 2024.</p>
Generative AI in Healthcare: Revolutionizing Patient Care and Medical Innovation
<p><a href="https://autorexa.com/transforming-patient-care-the-role-of-generative-ai-in-modern-healthcare/">Generative AI in Healthcare</a> is revolutionizing the medical field by enhancing diagnostics, accelerating drug discovery, enabling personalized treatments, and improving patient engagement. From creating synthetic medical images for training to developing tailored treatment plans, this technology is transforming patient care and medical research. With its ability to analyze vast datasets and automate complex processes, generative AI is driving efficiency, reducing costs, and delivering precise solutions. Learn how generative AI is shaping the future of healthcare with innovation, accuracy, and accessibility.</p>
Generative AI in University Communication, 2nd Wave - Survey Data (May 2024)
<p>Der Datensatz mit dem Titel "Generative KI in der Hochschulkommunikation, 2. Welle - Umfragedaten (Mai 2024)" erfasst Informationen zur Einführung und Nutzung von generativer künstlicher Intelligenz (KI) im Kontext der Hochschulkommunikation. Die Umfrage, die im Mai 2024 unter 318 deutschen Hochschulen durchgeführt wurde, von denen 82 geantwortet haben, untersucht verschiedene Aspekte, darunter Bekanntheit und Wissen über verschiedene KI-Tools (z.B. ChatGPT), Diskussionen in Gremien, das Vorhandensein von Richtlinien für die Nutzung, das Vorhandensein von Arbeitsgruppen für generative KI, strategische Ziele und Initiativen, Schulungsangebote für generative KI-Tools und die wahrgenommene Bedeutung von generativen KI-Tools in der Hochschulkommunikation. Ziel des Datensatzes ist es, Einblicke in die aktuelle Landschaft und Praxis der Integration generativer KI im universitären Umfeld zu geben. Die Daten der ersten Erhebung sind unter https://doi.org/10.5281/zenodo.10254904 zu finden.</p> <p>The dataset, titled "Generative AI in University Communication - Survey Data (May 2024)," captures information related to the adoption and utilization of generative artificial intelligence (AI) in the context of university communication. This survey, conducted in June 2024 among 318 German universities of which 82 responded, explores various aspects, including awarenes and knowledge of various AI tools (e.g. ChatGPT), discussions in committees, the existence of guidelines for usage, the presence of working groups for generative AI, strategic goals and initiatives, training offerings for generative AI tools, and the perceived importance of generative AI tools in university communication. The dataset aims to provide insights into the current landscape and practices regarding the integration of generative AI within university settings. Data of the first wave can be found here: https://doi.org/10.5281/zenodo.10254904<br><br>More information here: <a href="https://www.hof.uni-halle.de/projekte/hochki/">https://www.hof.uni-halle.de/projekte/hochki/</a></p>
Replication Package: Navigating the Complexity of Generative AI Adoption in Software Engineering
<p>This paper explores the adoption of Generative Artificial Intelligence (AI) tools and Large Language Models (LLMs) within the domain of software engineering, focusing on the influencing factors at the individual, technological, and social levels. We applied a convergent mixed-methods approach to offer a comprehensive understanding of AI adoption dynamics. We initially conducted a structured interview study with 100 software engineers, drawing upon the Technology Acceptance Model (TAM), the Diffusion of Innovations theory (DOI), and the Social Cognitive Theory (SCT) as guiding theoretical frameworks. Employing the Gioia Methodology, we derived a preliminary theoretical model of AI adoption in software: the Human-AI Collaboration and Adaptation Framework (HACAF). This model was then validated using Partial Least Squares – Structural Equation Modeling (PLS-SEM) based on data from 183 software professionals. Our research unveils the complex dynamics at play in AI adoption within software engineering. Findings indicate that at this early stage of AI integration, the compatibility of AI tools within existing development workflows predominantly drives their adoption, challenging conventional technology acceptance theories. The impact of perceived usefulness, social factors, and personal innovativeness seems less pronounced than expected. The study provides crucial insights for future AI tool design and offers a framework for developing effective organizational implementation strategies.</p>
The Impact of Generative AI on Student Learning Outcomes: A Statistical Analytical Approach - Dataset
Open the record for dataset details and reuse information.
Geoparsing with Large Language Models: Leveraging the linguistic capabilities of generative AI to improve geographic information extraction
<h2>Geoparsing with Large Language Models</h2> <p>The .zip file included in this repository contains all the code and data required to reproduce the results from our paper. Note, however, that in order to run the OpenAI models, users will required an OpenAI API key and sufficient API credits.</p> <div> <h3>Data</h3> <p>The data used for the paper are in the <code>datasetst</code> and <code>results</code> folders.</p> <ul> <li> <p>**Datasets: **This contains the XML files (LGL and Geovirus) and Json files (News2024) used to benchmark the models. It also contains all the data used to fine-tune the gpt-3.5 model, the prompt templates sent to the LLMs, and other data used for mapping and data creation.</p> </li> <li> <p>**Results: **This contains the results for the models on the three datastes. The folder is separated by dataset, with a single <code>.csv</code> file giving the results for each model on each dataset separately. The <code>.csv</code> file is structured so that each row contains either a predicted toponym and an associated true toponym (along with assigned spatial coordinates), if the model correctly identified a toponym; otherwise the true toponym columns are empty for false positives and the predicted columns are empty for false negatives.</p> </li> </ul> <h3>Code</h3> <p>The code is split into two seperate folders <code>gpt_geoparser</code> and <code>notebooks</code>.</p> <ul> <li>**GPT_Geoparser: **this contains the classes and methods used process the XML and JSON articles (<code>data.py</code>), interact with the Nominatim API for geocoding (<code>gazetteer.py</code>), interact with the OpenAI API (<code>gpt_handler.py</code>), process the outputs from the GPT models (<code>geoparser.py</code>) and analyse the results (<code>analysis.py</code>).</li> <li><strong>Notebooks</strong>: This series of notebooks can be used to reproduce the results given in the paper. The file names a reasonably descriptive of what they do within the context of the paper.</li> </ul> <h3>Code/software</h3> <h3>Requirements</h3> <ul> <li>Numpy</li> <li>Pandas</li> <li>Geopy</li> <li>Scitkit-learn</li> <li>lxml</li> <li>openai</li> <li>matplotlib</li> <li>Contextily</li> <li>Shapely</li> <li>Geopandas</li> <li>tqdm</li> <li>huggingface_hub</li> <li>Gnews</li> </ul> <h3>Access information</h3> <p>Other publicly accessible locations of the data:</p> <ul> <li>The LGL and GeoVirus datasets can also be obtained <a href="https://github.com/milangritta/Pragmatic-Guide-to-Geoparsing-Evaluation" target="_blank" rel="noopener">here<span> (opens in new window)</span></a>.</li> </ul> <h3>Abstract</h3> <div> <p>Geoparsing- the process of associating textual data with geographic locations - is a key challenge in natural language processing. The often ambiguous and complex nature of geospatial language make geoparsing a difficult task, requiring sophisticated language modelling techniques. Recent developments in Large Language Models (LLMs) have demonstrated their impressive capability in natural language modelling, suggesting suitability to a wide range of complex linguistic tasks. In this paper, we evaluate the performance of four LLMs - GPT-3.5, GPT-4o, Llama-3.1-8b and Gemma-2-9b - in geographic information extraction by testing them on three geoparsing benchmark datasets: GeoVirus, LGL, and a novel dataset, News2024, composed of geotagged news articles published outside the models' training window. We demonstrate that, through techniques such as fine-tuning and retrieval-augmented generation, LLMs significantly outperform existing geoparsing models. The best performing models achieve a toponym extraction F1 score of 0.985 and toponym resolution accuracy within 161 km of 0.921. Additionally, we show that the spatial information encoded within the embedding space of these models may explain their strong performance in geographic information extraction. Finally, we discuss the spatial biases inherent in the models' predictions and emphasize the need for caution when applying these techniques in certain contexts.</p> </div> <h3>Methods</h3> <div> <p>This contains the data and codes required to reproduce the results from our paper. The LGL and GeoVirus datasets are pre-existing datasets, with references given in the manuscript. The News2024 dataset was constructed specifically for the paper. </p> <p>To construct the News2024 dataset, we first created a list of 50 cities from around the world which have population greater than 1000000. We then used the GNews python package <a href="https://pypi.org/project/gnews/" target="_blank" rel="noopener">https://pypi.org/project/gnews/<span> (opens in new window)</span></a> to find a news article for each location, published between 2024-05-01 and 2024-06-30 (inclusive). Of these articles, 47 were found to contain toponyms, with the three rejected articles referring to businesses which share a name with a city, and which did not otherwise mention any place names.</p> <p>We used a semi autonmous approach to geotagging the articles. The articles were first processed using a Distil-BERT model, fine tuned for named entity recognicion. This provided a first estimate of the toponyms within the text. A human reviewer then read the articles, and accepted or rejected the machine tags, and added any tags missing from the machine tagging process. We then used OpenStreetMap to obtain geographic coordinates for the location, and to identify the toponym type (e.g. city, town, village, river etc). We also flagged if the toponym was acting as a geo-political entity, as these were reomved from the analysis process. In total, 534 toponyms were identified in the 47 news articles. </p> </div> </div>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.