Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
997
datasets available to search
ShareScore release 0.9.0
Dataset results
997 results for “AWARENESS”
#nowplaying-RS: A New Benchmark Dataset for Building Context-Aware Music Recommender Systems
<p>Music recommender systems can offer users personalized and contextualized recommendation and are therefore important for music information retrieval. An increasing number of datasets have been compiled to facilitate research on different topics, such as content-based, context-based or next-song recommendation. However, these topics are usually addressed separately using different datasets, due to the lack of a unified dataset that contains a large variety of feature types such as item features, user contexts, and timestamps. To address this issue, we propose a large-scale benchmark dataset called #nowplaying-RS, which contains 11.6 million music listening events (LEs) of 139K users and 346K tracks collected from Twitter. The dataset comes with a rich set of item content features and user context features, and the timestamps of the LEs. Moreover, some of the user context features imply the cultural origin of the users, and some others—like hashtags—give clues to the emotional state of a user underlying an LE. In this paper, we provide some statistics to give insight into the dataset, and some directions in which the dataset can be used for making music recommendation. We also provide standardized training and test sets for experimentation, and some baseline results obtained by using factorization machines.</p> <p>The dataset contains three files:</p> <ul> <li>user_track_hashtag_timestamp.csv contains basic information about each listening event. For each listening event, we provide an id, the user_id, track_id, hashtag, created_at </li> <li>context_content_features.csv: contains all context and content features. For each listening event, we provide the id of the event, user_id, track_id, artist_id, content features regarding the track mentioned in the event (instrumentalness, liveness, speechiness, danceability, valence, loudness, tempo, acousticness, energy, mode, key) and context features regarding the listening event (coordinates (as geoJSON), place (as geoJSON), geo (as geoJSON), tweet_language, created_at, user_lang, time_zone, entities contained in the tweet).</li> <li>sentiment_values.csv contains sentiment information for hashtags. It contains the hashtag itself and the sentiment values gathered via four different sentiment dictionaries: AFINN, Opinion Lexicon, Sentistrength Lexicon and vader. For each of these dictionaries we list the minimum, maximum, sum and average of all sentiments of the tokens of the hashtag (if available, else we list empty values). However, as most hashtags only consist of a single token, these values are equal in most cases. Please note that the lexica are rather diverse and therefore, are able to resolve very different terms against a score. Hence, the resulting csv is rather sparse. The file contains the following comma-separated values: <hashtag, vader_min, vader_max, vader_sum,vader_avg, afinn_min, afinn_max, afinn_sum, afinn_avg, ol_min, ol_max, ol_sum, ol_avg, ss_min, ss_max, ss_sum, ss_avg >, where we abbreviate all scores gathered over the Opinion Lexicon with the prefix 'ol'. Similarly, 'ss' stands for SentiStrength. </li> </ul> <p>Please also find the training and test-splits for the dataset in this repo. Also, prototypical implementations of a context-aware recommender system based on the dataset can be found at <a href="https://github.com/asmitapoddar/nowplaying-RS-Music-Reco-FM">https://github.com/asmitapoddar/nowplaying-RS-Music-Reco-FM</a>.</p> <p>If you make use of this dataset, please cite the following paper where we describe and experiment with the dataset:</p> <p>@inproceedings{smc18,<br> title = {#nowplaying-RS: A New Benchmark Dataset for Building Context-Aware Music Recommender Systems},<br> author = {Asmita Poddar and Eva Zangerle and Yi-Hsuan Yang},<br> url = {http://mac.citi.sinica.edu.tw/~yang/pub/poddar18smc.pdf},<br> year = {2018},<br> date = {2018-07-04},<br> booktitle = {Proceedings of the 15th Sound & Music Computing Conference},<br> address = {Limassol, Cyprus},<br> note = {code at https://github.com/asmitapoddar/nowplaying-RS-Music-Reco-FM},<br> tppubtype = {inproceedings}<br> }</p>
Figure 6 in Society´s awareness for protection of soils, its biodiversity and function in 2030 - We need a more intrinsic approach
Figure 6. The mole standing at the entrance of the Senckenberg exhibition is a popular motif for selfies.
Figure 3 in Society´s awareness for protection of soils, its biodiversity and function in 2030 - We need a more intrinsic approach
Figure 3. The card game 'Soil Builder' by Helga ZumkowskiXylander (2017 b) (only a selection of cards is shown).
Figure 2 in Society´s awareness for protection of soils, its biodiversity and function in 2030 - We need a more intrinsic approach
Figure 2. Pupils experiencing soil life in a class using dissecting microscopes. Soil samples for investigation were taken by the pupils themselves.
Figure 7 in Society´s awareness for protection of soils, its biodiversity and function in 2030 - We need a more intrinsic approach
Figure 7. Tardigrade as a soft toy is one of few soil animals which found their way to childrens' rooms.
Figure 5 in Society´s awareness for protection of soils, its biodiversity and function in 2030 - We need a more intrinsic approach
Figure 5. The international touring exhibition 'The thin skin of the earth' displays units of soil biodiversity, research, heterogeneity and destruction. The exhibition had over 250.000 visitors till now.
Figure 4 in Society´s awareness for protection of soils, its biodiversity and function in 2030 - We need a more intrinsic approach
Figure 4. Klara Kugelspringer and her friends driven from their home by man-made erosion (from Zumkowski-Xylander 2017a)
Figure 1 in Society´s awareness for protection of soils, its biodiversity and function in 2030 - We need a more intrinsic approach
Figure 1. Picture from the VR-animation 'Adventure Soil Life' part 'leaf litter'. By SMNG/.hapto modified after Xylander (2019).
Data and models for: Learning Ordering in Crystalline Materials with Symmetry-Aware Graph Neural Networks
<p>Data (ver 1.1) and trained models for our paper "<a href="https://arxiv.org/abs/2409.13851">Learning Ordering in Crystalline Materials with Symmetry-Aware Graph Neural Networks</a>". If you use such data or models, please cite our paper. These three directories need to be downloaded and copied into our source codes in order to reproduce our paper: <a href="https://github.com/learningmatter-mit/PerovskiteOrderingGCNNs">https://github.com/learningmatter-mit/PerovskiteOrderingGCNNs</a></p> <ul> <li>data: All data files for training and evaluating GCNNs, with a copy archived on the Materials Data Facility (<a href="https://doi.org/10.18126/ncqt-rh18">DOI: 10.18126/ncqt-rh18</a>)</li> <li>saved_models: All saved model files for evaluating GCNNs</li> <li>best_models: All best model files for evaluating GCNNs</li> </ul>
Climate Change Awareness in the Arab Barometer Wave 5 Survey
<p>The Arab Barometer Wave V 2018-2019 is based on a nationally representative probability sample of the population aged 18 and above. In most countries, the sample includes 2,400 citizens. The data were conducted in face-to-face public opinion surveys (CAPI and PAPI). See technical reports by country for country-specific information. You can find the data, codebooks and all relevant information on the Arab Barometer website.</p> <p>Our dataset contains country weighted counts of different answer options and the re-weighted values of the answers given to the Arab Barometer Wave 5 question:</p> <p>Q108 : How serious a problem do you think the following issues are: Is climate change a very serious problem, a somewhat serious problem, not a very serious problem, not at all a serious problem?</p> <p>Get the country averages and aggregates from Zenodo</p> <p>Get the plot in jpg or png from figshare.</p> <p> </p> <p>See the detailed PDF documentation.</p>
Exploring Augmented Reality Privacy Icons for Smart Home Devices and their Effect on Users' Privacy Awareness
<p><strong>Exploring Augmented Reality Privacy Icons for Smart Home Devices and their Effect on Users' Privacy Awareness</strong></p> <p><strong>Authors</strong></p> <p>Kathrin Knutzen, Florian Weidner, Wolfgang Broll</p> <p> </p> <p><strong>About</strong></p> <p>This data represents the supplementary material for the conference paper with above title submitted at ISMAR 2021.</p> <p> </p> <p><strong>Contents</strong></p> <p>The supplementary material contains five files:</p> <ol> <li>The abstraction of paraphrases and transcripts after each condition respectively.<br> According to qualitative content analysis procedure, the conducted interviews were transcribed, paraphrased and subsequently abstracted to generate a category system. Every category is described with a definition and some exemplary quotes. Statements of participants are condensed and abstracted. Number of participants who made statements regarding a category, and most prominent valence are taken as basis to generate tree maps in Figure 5 and 6.<br> Please note that the prevalences represent the views or opinions of the participants on the single categories. Also, the mentioned categories have several subcategories and only the most important regarding privacy awareness are mentioned in the article.<br> <br> Transcripts, audio files and paraphrases are available upon request.</li> <li>The experimental task description. It served as exposition for the task that the participants had to complete.</li> <li>The interview guideline. Please note that this study was part of a larger project that also focused on topics such as usability and immersion, however, the article reports only on privacy awareness.</li> <li>The R script file to generate the tree maps in Figures 5 and 6. The dataset is created using data from the abstraction Excel sheet.</li> <li>A demonstration video of the experimental setup.</li> </ol>
Dataset used in "Uncertainty-Aware Learning for Improvements in Image Quality of the Canada-France-Hawaii Telescope" (https://arxiv.org/abs/2107.00048)
<p>'x_train.p', 'y_train.p': pickle files for training split containing 50,757 samples</p> <p>'x_val.p', 'y_val.p': pickle file for validation split containing 5,640 samples</p> <p>'x_test.p', 'y_test.p': pickle file for test split containing 6,267 samples</p> <p>'feature_names.p': pickle file containing names of all 119 features</p>
COVID-19++: A Citation-Aware Covid-19 Dataset for the Analysis of Research Dynamics
<p>COVID-19++ is a citation-aware COVID-19 dataset for the analysis of research dynamics. In addition to primary COVID-19 related articles and preprints from 2020, it includes citations and the metadata of first-order cited work. All publications are annotated with MeSH terms, either from the ground truth, or via ConceptMapper, if no ground truth was available. </p> <p>The data is organized in CSV files</p> <p>- Paper metadata (paper_id, publdate, title, data_source): paper.csv</p> <p>- Annotation data, mapping paper_id to MeSH terms: annotation.csv </p> <p>- Authorship data, mapping paper_id to author, optionally with ORCID: authorship.csv<br> - Paired DOIs of citing and cited papers: references.csv</p> <p>The column data source within the paper metadata has the value KE (for metadata from ZB MED KE), PP (for preprints) or CR (for cited resources from CrossRef)<br> </p> <p>This work was supported by BMBF within the programme ``Quantitative Wissenschaftsforschung'' under grant numbers 01PU17013A, 01PU17013B, 01PU17013C.<br> </p>
Mobile for Mothers: a randomized quasi-controlled mobile health intervention to augment maternal health awareness and behavior of pregnant women in tribal societies.
<p>The files contain anonymized raw baseline and end-line data. Anyone using the data must cite the original studies given in the references</p>
Data for - Tracking one-in-a-million: Large-scale benchmark for microbial single-cell tracking with experiment-aware robustness metrics
<p><strong>Large-scale Corynebacterium glutamicum data set with Segmentation and Tracking Annotation</strong></p> <p>We provide five time-lapse sequences with manually corrected segmentation and tracking annotations of growing <strong><em>C. glutamicum</em></strong> cultivations. The dataset contains more than 1.4 million cell observations in 29k cell tracks and 14k cell divisions. We provide videos of the annotations (videos.zip) and the dataset in <a href="http://celltrackingchallenge.net/datasets/">Cell Tracking Challenge</a> format (ctc_format.zip). In the videos, cell contours are rendered in yellow, cell links between frames are colored red and cell divisions, and their links are colored in blue.</p> <p><strong>Data Acquisition</strong></p> <p><strong><em>Corynebacterium glutamicum</em></strong> ATCC 13032 was cultivated in BHI-medium at 30°C in this study. From and overnight preculture, the main culture was inoculated the next day with a starting OD600 of 0.05 and grown at 120 rpm to a OD600 of 0.25. A chip was fabricated, according to <a href="https://doi.org/10.1039/D0LC00711K">(Täuber et al., 2020)</a>, and fixed to the microscope’s holder. The main culture cells were transferred to monolayer growth chambers (height = 720 nm) on the microfluidic chip. Flow through the microfluidic device was mediated by pressure driven pumps with a pressure of 100 mbar on the medium reservoir.</p> <p>The time-lapse phase contrast images of five monolayer growth chambers were taken every minute using an inverted microscope (Nikon Eclipse Ti2) with a 100x oil emersion objective and a DS-QI2 camera (Nikon) at 15 % relative DIA-illumination intensity and 100 ms exposure time. The spatial image resolution is 0.072 μm/px.</p>
Data of: Cross-sectional survey on Germans' awareness for refugees' information barriers
<p>The present dataset is the result of a cross-sectional online survey, which had been conducted to examine selected predictors of Germans' problem awareness in the form of perceived information barriers that refugees face, placing an emphasis on the role of positive intercultural contact experiences. The survey content is based on an extended version of the Empathy-Attitude-Action model and was carried out with a sample of Germans. </p> <p>The dataset is in xlsx-format and the variable-descriptions can be found in the headings of the spreadsheet. </p>
Decreased community-acquired pneumonia coincided with rising awareness of precautions before governmental containment policy in Japan
<p>Google community mobility data provides regional-level movement trends by location type (retail and recreation, grocery and pharmacy, parks, transit stations, workplaces, and residential) as a relative change from the day-of-the-week-wise average from January 3 through February 6, 2020. We obtained the daily national average of Google mobility in Japan in all location types but residential to measure the contact behavior of individuals outside the home.</p>
Video no. 4 'A short exercise on body awareness - movement technique in working on music choreography' by Barbara Dutkiewicz
<p>The footage complements the publication by Barbara Dutkiewicz (2023) ‚<em>Choreography of music. Process of creation according to the principles of plastique animée on the example of music by Henryk Mikołaj Górecki 'Kleines Requiem für eine Polka op. 66’, choreography created by Barbara Dutkiewicz and Iga Eckert</em>’, (DOI:<a href="https://doi.org/10.5281/zenodo.7789711"> https://doi.org/10.5281/zenodo.7789711</a>), written as a part of the intellectual output of the project EURHYTHMICS IN EDUCATION AND ARTISTIC PRACTICE (ERASMUS+). implemented in 2020-2023. Workshop with students of The Karol Szymanowski Academy of Music in Katowice, Poland and Universität für Musik und darstellende Kunst Wien, Austria took place during LTT, host Katowice 21-25.03.2022. Artistic direction: Associate Professor Barbara Dutkiewicz (PhD.hab) and Iga Eckert, MA.</p> <p>ISBN 978-83-963687-3-7</p> <p>The article is supplemented by photos and four videos:</p> <p>- Video no. 1 ’Motif, phrase in working on music choreography’ by Barbara Dutkiewicz <a href="https://doi.org/10.5281/zenodo.7922029">https://doi.org/10.5281/zenodo.7922029</a></p> <p>- Video no 2. ’Exercises with polymetric structure in working on music choreography’ by Iga Eckert and Barbara Dutkiewicz <a href="https://doi.org/10.5281/zenodo.7916295">https://doi.org/10.5281/zenodo.7916295</a> </p> <p>- Video no. 3 ’Movement stylization and multi planarity in working on music choreography’ by Barbara Dutkiewicz <a href="https://doi.org/10.5281/zenodo.7922104">https://doi.org/10.5281/zenodo.7922104</a></p> <p>- Video no. 4 'A short exercise on body awareness - movement technique in working on music choreography’ by Barbara Dutkiewicz <a href="https://doi.org/10.5281/zenodo.7916598">https://doi.org/10.5281/zenodo.7916598</a></p> <p>This article and videos are published on the digital platform "Atlas of Eurhythmics" <a href="https://www.kmh.se/in-english/atlas-of-eurhythmics/results-from-the-project.html">https://www.kmh.se/in-english/atlas-of-eurhythmics/results-from-the-project.html</a> <em> </em>at modul:<em> <strong>"Plastique animée - tradition and contemporary performing"</strong></em><strong>.</strong> </p> <p> </p>
CausalOrca: An ORCA-based Diagnostic Dataset for Causally-aware Multi-agent Trajectory Prediction
<p>CausalOrca is a synthetic diagnostic dataset created through controlled simulations. It is designed to provide annotations of ground-truth causal effects and fine-grained agent categories for social interactions in multi-agent scenarios. The dataset is constructed using a modified RVO2 simulator and incorporates the ORCA optimization-based collision avoidance algorithm known for crowd simulation. With full control over scene configurations, the dataset enables the collection of motion behaviors in paired scenes before and after agent removal, generating a large set of counterfactual pairs with annotations of ground-truth causal effects. CausalOrca can serve as a valuable resource for studying and developing causally-aware neural representations of social interactions and trajectory prediction models. Please see the <a href="https://github.com/rebuttal-anonymous/causalorca">GitHub repository</a> for a more detailed description of the dataset, including dataset statistics and documentation on how to use, visualize, and generate the data.</p>
SARA - A Collection of Sensitivity-Aware Relevance Assessments
<p><strong>SARA - A Collection of Sensitivity-Aware Relevance Assessments</strong></p> <p>Presented here is a collection of Sensitivity-Aware Relevance Assessments for the UC Berkely labelled subset of the Enron Email Collection. The Hearst [1] labelled version of the Enron Email Collection is a subset of the CMU collection that contains 1702 emails that were annotated as part of a class project at UC Berkley. Students in the Natural Language Processing course were tasked with annotating the emails as relevant or not relevant to 53 different categories. Therefore, the labelled version of the Enron email collection provides a rich taxonomy of labels which can be used for multiple definitions of sensitivity such as the Purely Personal and Personal but in a Professional Context. The categories that the emails are labelled for can be seen in [Table 1](#table-1). The files for the labelled version of the Enron Email Collection are available from the<a href="https://bailando.berkeley.edu/enron_email.html"> UC Berkely website</a>.</p> <p>We deploy a topic modelling approach to identify topical themes in the labelled Enron collection that serve as a basis for our information needs which are in turn used to gather queries and relevance assessments, the notebook for which is available <a href="https://colab.research.google.com/drive/1r_mzjVETN5ytYbdJEKuuTUOopFiWjyWR?usp=sharing">here</a>. Two separate crowdsourcing tasks are carried out in the development of SARA. Firstly, query formulations are crowdsourced to represent the information needs and, secondly, relevance assessments are crowdsourced for a pooled set of documents from the labelled Enron collection for each of the information needs.</p> <p>The SARA Collection of Sensitivity-Aware Relevance Assessments is available through the popular ir_datasets library. More information can be found on the ir_datasets <a href="https://github.com/allenai/ir_datasets">GitHub</a> and <a href="https://ir-datasets.com/">website</a>.</p> <p><strong>Information Needs</strong></p> <p>To create our set of sensitivity-aware relevance assessments for the labelled Enron email collection, we first identify a set of topical subjects that reflect the contents of the emails in the collection. We use a topic modelling approach to identify the information needs. When identifying topics to be used as information needs, we are interested in identifying general themes that relate to the topics of discussion that might likely be covered in the contents (i.e., the body) of the emails in the collection. The topics are chosen to be broad enough to be able to reasonably expect that there would be relevant documents in the collection, and not so specific that it would require specialist knowledge to make a judgement of relevance on the subject. Subsequently, we manually construct short passages of text to serve as descriptions of the information needs that are to be searched for in the collection by the crowdworkers. The information needs that the crowdworkers are available in the <em>information_needs.tsv </em>file.</p> <p><strong>Queries</strong></p> <p>In order to collect relevance assessments for pairs of emails and information needs, different query formulations are first needed to generate pools of documents. Query formulations for each topic are collected from crowdworkers from the Prolific crowdwork platform. Ten information needs are shown to each crowdworker and they are asked to provide a query formulation that they would use to get relevant documents to satisfy the information need they are presented with. Three queries for each of the fifty information needs are released. The resulting queries are available in the <em>repeated_queries.tsv</em> file.</p> <p><strong>Relevance Assesments</strong></p> <p>Crowdworkers are shown an information need and an email and asked to rate the document as being either <em>Highly Relevant</em>, <em>Partially Relevant</em>, or <em>Not Relevant</em> to the information need. Each information need/email pair is judged by three crowdworkers and a majority vote is used to generate a ground truth label. Since each information need / email pair is judged by three crowdworkers and there are three possible labels, it is possible for each of the labels to be selected by one crowdworker. In practice, this only happened for 134 pairs. In such cases, ties are broken by having one of the authors read the document and make an additional judgement. In order to ensure that sensitive documents definitely have relevance labels they were also judged by one of the authors for each of the information needs. The relevance assessments are available in the <em>repeated_qrels.txt</em> file. The relevance assessments are in the format 'query iteration document relevancy'. The iteration column is used for IR_Datasets and can be safely ignored and the document name is the filename used in the labelled Enron collection.</p> <p><em>Table 1</em></p> <table> <thead> <tr> <th scope="col"> <table> <thead> <tr> <th>1) Coarse genre</th> <th>2) Included/forwarded information</th> <th>3) Primary topics (If coarse genre 1.1 is selected)</th> <th>4) Emotional tone (If not neutral)</th> </tr> </thead> <tbody> <tr> <td>1.1 Company Business, Strategy, etc. (See 3)</td> <td>2.1 Includes new text in addition to forwarded material</td> <td>3.1 Regulations and regulators (includes price caps)</td> <td>4.1 Jubilation</td> </tr> <tr> <td>1.2 Purely Personal</td> <td>2.2 Forwarded email(s) including replies</td> <td>3.2 Internal projects -- progress and strategy</td> <td>4.2 Hope / anticipation</td> </tr> <tr> <td>1.3 Personal but in professional context (e.g., it was good working with you)</td> <td>2.3 Business letter(s) / document(s)</td> <td>3.3 Company image -- current</td> <td>4.3 Humor</td> </tr> <tr> <td>1.4 Logistic Arrangements (meeting scheduling, technical support, etc.)</td> <td>2.4 News article(s)</td> <td>3.4 Company image -- changing / influencing</td> <td>4.4 Camaraderie</td> </tr> <tr> <td>1.5 Employment arrangements (job seeking, hiring, recommendations, etc.)</td> <td>2.5 Government / academic report(s)</td> <td>3.5 Political influence / contributions / contacts</td> <td>4.5 Admiration</td> </tr> <tr> <td>1.6 Document editing/checking (collaboration)</td> <td>2.6 Government action(s) (such as results of a hearing, etc.)</td> <td>3.6 California energy crisis / California politics</td> <td>4.6 Gratitude</td> </tr> <tr> <td>1.7 Empty message (due to missing attachment)</td> <td>2.7 Press release(s)</td> <td>3.7 Internal company policy</td> <td>4.7 Friendship / affection</td> </tr> <tr> <td>1.8 Empty message</td> <td>2.8 Legal documents (complaints, lawsuits, advice)</td> <td>3.8 Internal company operations</td> <td>4.8 Sympathy / support</td> </tr> <tr> <td> </td> <td>2.9 Pointers to url(s)</td> <td>3.9 Alliances / partnerships</td> <td>4.9 Sarcasm</td> </tr> <tr> <td> </td> <td>2.10 Newsletters</td> <td>3.10 Legal advice</td> <td>4.10 Secrecy / confidentiality</td> </tr> <tr> <td> </td> <td>2.11 Jokes, humor (related to business)</td> <td>3.11 Talking points</td> <td>4.11 Worry / anxiety</td> </tr> <tr> <td> </td> <td>2.12 Jokes, humor (unrelated to business)</td> <td>3.12 Meeting minutes</td> <td>4.12 Concern</td> </tr> <tr> <td> </td> <td>2.13 Attachment(s) (assumed missing)</td> <td>3.13 Trip reports</td> <td>4.13 Competitiveness / aggressiveness</td> </tr> <tr> <td> </td> <td> </td> <td> </td> <td>4.14 Triumph / gloating</td> </tr> <tr> <td> </td> <td> </td> <td> </td> <td>4.15 Pride</td> </tr> <tr> <td> </td> <td> </td> <td> </td> <td>4.16 Anger / agitation</td> </tr> <tr> <td> </td> <td> </td> <td> </td> <td>4.17 Sadness / despair</td> </tr> <tr> <td> </td> <td> </td> <td> </td> <td>4.18 Shame</td> </tr> <tr> <td> </td> <td> </td> <td> </td> <td>4.19 Dislike / scorn</td> </tr> </tbody> </table> </th> </tr> </thead> <tbody> </tbody> </table> <p> </p><p> </p><p>The Sensitivity-Aware Relevance Assessments dataset is held under an Attribution-NonCommercial 4.0 International (CC BY-NC 4.0) licence which allows for it to be adapted, transformed and built upon.</p> <p></p> <p></p> <p>Questions and comments are welcomed via <a href="http://mailto:j.mckechnie.1@research.gla.ac.uk">email</a>.</p> <p><strong>References</strong></p> <p>[1] Marti A Hearst. 2005. Teaching applied natural language processing: Triumphs and tribulations. In Proc. of Workshop on Effective Tools and Methodologies for Teaching NLP and CL.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.