Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
31
datasets available to search
ShareScore release 0.9.0
Dataset results
31 results for “Blog”
Határátkelő blog
<p>This object has been created as a part of the web harvesting project of the Eötvös Loránd University Department of Digital Humanities <a href="https://elte-dh.hu/en">ELTE DH</a>. Learn more about the workflow <a href="https://www.aclweb.org/anthology/2020.wac-1.5/">HERE</a> about the software used <a href="https://github.com/ELTE-DH/WebArticleCurator/">HERE</a>.The aim of the project is to make online news articles and their metadata suitable for research purposes. The archiving workflow is designed to prevent modification or manipulation of the downloaded content. The current version of the curated content with normalized formatting in standard <a href="https://tei-c.org/">TEI XML</a> format with Schema.org encoded metadata is available <a href="https://doi.org/10.5281/zenodo.5831354">HERE</a>. The detailed description of the raw content is the following:</p><ul><li>The portal's archived content (from 2012-07-19 to 2021-08-09) in WARC format available <a href="https://doi.org/10.5281/zenodo.5536989">HERE</a> (crawled: 2021-08-09T18:47:30.917190 - 2021-08-09T19:18:44.861392).</li></ul><br>Please fill in the following form before requesting access to this dataset:<a href="https://forms.office.com/r/VnP05QYTGV">ACCES FORM</a>
The relationship between linguistic expression and symptoms of depression, anxiety, and suicidal thoughts: A longitudinal study of blog content
<p>To investigate the associations between linguistic features and symptoms of depression, generalised anxiety, and suicidal ideation, we extracted linguistic features from individuals’ blog content and correlated it with validated mental health data in a longitudinal study (n=38). Depressive symptoms were assessed using the self-report Patient Health Questionnaire (PHQ-9), anxiety symptoms using the self-report Generalised Anxiety Disorder Scale (GAD-7), and social media data was analysed using the Linguistic Inquiry and Word Count (LIWC) tool for linguistic features. Bivariate and multivariate analyses were performed to investigate the correlations between the linguistic features and mental health scores between subjects. We then used the multivariate regression model to predict longitudinal changes in mood within subjects.</p>
Ports, Past and Present Blog (perma.cc link tabular data)
<p>A list of blog and calendar posts from the Ports, Past and Present project written on Wordpress. This tabular data is a modified output from the <a href="https://perma.cc/">perma.cc</a> folder containing a series of WARC records captured on the 26th of June, 2023.</p>
Entradas del blog "Lo Inquieto" y su relevancia.
<p><strong>Conjunto de datos sobre cada entrada escrita en el blog en el apartado de literatura, información sobre el autor, la fecha, titulo y la entrada del blog.</strong></p>
Wordpress blog export, posts from 2006--18 July, 2015.
<p>An XML Wordpress export of the 440 blog posts and associated comments by Henry Rzepa up to July 18, 2015.</p>
BRAIN Journal-Participative Teaching with Mobile Devices and Social Networks for K-12 Children-Figure 7. The projects's Twitter page for educational and research micro-blogging
<p>The fourth stage was the use of social media to build a social learning network based on content and experience sharing and creation. The following web 2.0 services were used as a distributed platform able to support our experimental learning system: a) Panoramio (http://www.panoramio.com/user/7606828) (Figure 6) as a geo-referenced photo-sharing service over Google Maps and Google Earth for sharing project’s essential results; b) Twitter (https://twitter.com/maps_of_time) (Figure 7), as a social network and micro-blogging service for short announcements and comments; c) Google+ (https://plus.google.com/114705936110835992130?hl=en#114705936110835992130/posts?hl=en) (Figure 8), as a platform for sharing and tagging multiple content (photo, video), blogging and video chatting service with the recent Google Hangout, for sharing educational content; d) Google Drive (Levin, 2013) for cloud storage and collaborative document editing. We also created a YouTube channel for public distribution of video content (https://www.youtube.com/TimemapsNet), (Rusu et al., 2013) and a Facebook page of Vădastra School (https://www.facebook.com/scoalaVădastra). </p>
BRAIN Journal-Participative Teaching with Mobile Devices and Social Networks for K-12 Children-Figure 9.The educational blog on Google+ Time Maps page–the weaving techniques
<p>For this subject two video films were posted on Google+ (a performance and a 3D reconstruction), slightly different from those available on the Time Maps web site, but containing the same information. The children had to make a little effort to relate this information with the one presented on the site, to make a connection between the questions, the fragments from videos at which the answers referred to and the information from the site. The set of questionnaires lead the school children through the majority of data offered by the web site regarding to the two historical periods (Figures 9, 10).</p>
BRAIN Journal-Participative Teaching with Mobile Devices and Social Networks for K-12 Children-Figure 10. The educational blog on Google+ Time Maps page – the glass manufacturing techniques
<p>For this subject two video films were posted on Google+ (a performance and a 3D<br> reconstruction), slightly different from those available on the Time Maps web site, but containing<br> the same information. The children had to make a little effort to relate this information with the one<br> presented on the site, to make a connection between the questions, the fragments from videos at<br> which the answers referred to and the information from the site.<br> The set of questionnaires lead the school children through the majority of data offered by the<br> web site regarding to the two historical periods (Figures 9, 10).</p>
Material related to the blog that reports on the parrot LUT
<p><strong>Material related to the blog that reports on the parrot LUT</strong></p> <p>The blog was published at the Node: <a href="http://thenode.biologists.com/parrot-lut/research/">http://thenode.biologists.com/parrot-lut/research/</a></p> <p> </p> <p><strong>-Source</strong></p> <p>The ‘morgenstemning’ LUT was originally described in:</p> <p>M. Geissbuehler and T. Lasser - "How to display data by color schemes compatible with red-green color perception deficiencies”, Optics Express, 2013</p> <p>The ‘inferno’ LUT was originally created by Stéfan van der Walt and Nathaniel Smith (<a href="http://bids.github.io/colormap/">http://bids.github.io/colormap/</a>).</p> <p>The ‘pseudocolorMM’ LUT was derived from MetaMorph software (version 7.6).</p> <p>The ‘royal’ and ‘Fire’ LUT are available in ImageJ (version 1.49j)</p> <p>The ‘parrot’ LUT was designed by Joachim Goedhart and first described here:<br> <a href="http://thenode.biologists.com/parrot-lut/research/">http://thenode.biologists.com/parrot-lut/research/</a></p> <p><br> <strong>-Distribution</strong></p> <p>The colormaps Magma, Inferno, Plasma and Viridis are available under a CC0 "no rights reserved" license (<a href="https://creativecommons.org/share-your-work/public-domain/cc0">https://creativecommons.org/share-your-work/public-domain/cc0</a>) and are present in FIJI.</p> <p><br> The colormaps Mongenstemning & Parrot are free software: you can redistribute them and/or modify<br> it under the terms of the GNU General Public License as published by<br> the Free Software Foundation, either version 3 of the License, or<br> (at your option) any later version.</p> <p>These colormaps are distributed in the hope that they will be useful,<br> but WITHOUT ANY WARRANTY; without even the implied warranty of<br> MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the<br> GNU General Public License for more details: <<a href="http://www.gnu.org/licenses/">http://www.gnu.org/licenses/</a>>.</p>
Data for blog post on dfm.io: "An experiment in open science: exoplanet population inference"
<p>The data set used be the blog post "An experiment in open science: exoplanet population inference" published at https://dfm.io/posts/exopop/</p>
Post Blog Sketchfab
Ecco il nostro articolo sul blog di Sketchfab nella categoria Cultural Heritage https://sketchfab.com/blogs/community/teramomusiva-creating-3d-models-of-archaeological-finds-to-highlight-heritage Seguendo le annotazioni Potrete leggere l'articolo. #eurotech #universitàdichieti Source: Objaverse 1.0 / Sketchfab
WSSSPE 5.1 - Data for speed blog analysis
<p>Data and coding for the thematic/framework analysis of speed blogs produced at WSSSPE 5.1.</p> <p>WSSSPE5.1_blogs.docx contains the original text of the speed blogs, obtained from the Software Sustainability Institute website. Authors of the individual speed blogs are credited in the text of each post.</p> <p>WSSSPE5.1_blogpost_coding.xlxs contains the results of the thematic analysis.</p>
BLOQUES Spanish learners blog corpus (Japanese L1)
<p><strong>Description</strong></p> <p>List of 2,669 URLs corresponding to blog posts written by Japanese speakers who are learning Spanish.</p> <p>Posts belong to 46 blogs (from 41 users) mainly from the Blogger and WordPress domains.</p> <p><strong>Related publications:</strong></p> <p>Valverde, P. (2020). "Parece más KAWAII ¿no?": los blogs como práctica de la escritura extensiva fuera del aula. In Journal of Inquiry and Research, n. 112, Kansai Gaidai University, pp. 261-277. <a href="http://doi.org/10.18956/00007940">http://doi.org/10.18956/00007940</a></p> <p>Valverde, P. (2018). Un corpus de blogs de aprendices japoneses de español (A blog corpus of Japanese learners of Spanish). In Bargalló Escrivá, M., Forgas Berdet, E. & Nomdedeu Rull, A. (Eds.), Léxico y cultura en LE/L2: corpus y diccionarios. Asociación para la Enseñanza del Español como Lengua Extranjera, pp. 845-857, ISBN 978-84-09-04375-0. <a href="https://cvc.cervantes.es/ensenanza/biblioteca_ele/asele/pdf/28/28_0078.pdf">https://cvc.cervantes.es/ensenanza/biblioteca_ele/asele/pdf/28/28_0078.pdf</a></p> <p>Valverde, P. (2016). Japanese L1 speakers blogging in Spanish: motivations, topics and linguistic properties. In Moreno Ortiz, A. & Pérez-Hernández, C (Eds.) EPiC Series in Language and Linguistics. 8th International Conference on Corpus Linguistics, vol.1, pp. 424-437, ISSN 2398-5283. <a href="https://doi.org/10.29007/k2jk">https://doi.org/10.29007/k2jk</a></p>
Update to figure 1 of the blog "The growth of interdisciplinarity"
<p>Figure 1 in the blog "<a href="https://ui.adsabs.harvard.edu/blog/citations-journals">The growth of interdisciplinarity</a>" shows the growth in the number of papers over the 20 years of the study, for all refereed articles, those in the 50 journal sample, and the rest. The journals in the main sample (the In Sample set) are all major international journals; all astronomy journals which have published at least fifty papers with 100 or more citations in the last 20 years are included in this sample. The journals outside the sample (the Out of Sample set) are of two basic types: smaller, mostly national, astronomy journals, such as ARep, AstL, MmSAI, BASI, ChA&A, BaltA, etc., and journals which are not normally considered “astronomy” journals but which sometimes publish astronomy-related articles. There are more than 350 different journals which have published one or more astronomy-related articles which have received 100 or more citations in the last 20 years.</p> <p>This entry provides an update to this figure, with data up to and including 2023.</p>
Blog-1K
<p>The Blog-1K corpus is a redistributable authorship identification testbed for contemporary English prose. It has 1,000 candidate authors, 16K+ posts, and a pre-defined data split (train/dev/test proportional to ca. 8:1:1). It is a subset of the <a href="http://www.kaggle.com/datasets/rtatman/blog-authorship-corpus">Blog Authorship Corpus</a> from Kaggle. The MD5 for Blog-1K is '0a9e38740af9f921b6316b7f400acf06'.</p> <p>1. Preprocessing</p> <p>We first filter out texts shorter than 1,000 characters. Then we select one thousand authors whose writings meet the following criteria:<br> - accumulatively at least 10,000 characters,<br> - accumulatively at most 49,410 characters,<br> - accumulatively at least 16 posts,<br> - accumulatively at most 40 posts, and <br> - each text has at least 50 function words found in the Koppel512 list (to filter out non-English prose).</p> <p>Blog-1K has three columns: 'id', 'text', and 'split', where 'id' corresponds to its parent corpus.</p> <p>2. Statistics</p> <p>Its creation and statistics can be found in <a href="https://codeberg.org/haining/Blog-1K/src/branch/main/blog-1k_generation_and_stats.ipynb">the Jupyter Notebook</a>.</p> <table> <tbody> <tr> <td>Split</td> <td># Authors</td> <td># Posts</td> <td># Characters</td> <td>Avg. Characters Per Author (Std.)</td> <td>Avg. Characters Per Post (Std.)</td> </tr> <tr> <td>Train</td> <td>1,000</td> <td>16,132</td> <td>30,092,057</td> <td>30,092 (5,884)</td> <td>1,865 (1,007)</td> </tr> <tr> <td>Validation</td> <td>935</td> <td>2,017</td> <td>3,755,362</td> <td>4,016 (2,269)</td> <td>1,862 (999)</td> </tr> <tr> <td>Test</td> <td>924</td> <td>2,017</td> <td>3,732,448</td> <td>4,039 (2,188)</td> <td>1,850 (936)</td> </tr> </tbody> </table> <p><br> 3. Usage</p> <pre><code class="language-python">import pandas as pd df = pd.read_csv('blog1000.csv.gz', compression='infer') # read in training data train_text, train_label = zip(*df.loc[df.split=='train'][['text', 'id']].itertuples(index=False))</code></pre> <p> </p> <p>4. License<br> All the materials is licensed under the ISC License.</p> <p><br> 5. Contact<br> Please contact <a href="mailto:hw56@indiana.edu">its maintainer</a> for questions.</p>
Swedish Diachronic Corpus - User-generated Data - Blog text
<p>The Swedish Diachronic Corpus (https://www2.lingfil.uu.se/person/pettersson/svediakorp/) is a project funded by Swe-Clarin (<a href="https://sweclarin.se/eng">https://sweclarin.se/eng</a>). The purpose of the project is to provide a corpus of texts covering the time period from Old Swedish to present day, with a wide variety of text types and freely available for download and search. The texts are provided in a plain text format and in a uniform CoNLL format, with placeholders for linguistic annotation. </p> <p>The dataset provided here is the blog section of the User-generated Data in the Swedish Diachronic Corpus. For other datasets within the corpus, see further the corpus website: https://www2.lingfil.uu.se/person/pettersson/svediakorp/. </p> <p>The project members are Eva Pettersson (Uppsala University) and Lars Borin (University of Gothenburg). For questions or comments, or if you are aware of any corpus resource that could be included in the Swedish Diachronic Corpus, don't hesitate to contact us!<br> <br> Eva Pettersson Department of Linguistics and Philology, Uppsala University eva.pettersson@lingfil.uu.se<br> Lars Borin Department of Swedish, University of Gothenburg lars.borin@svenska.gu.se</p>
Wordpress blog export, posts from 2006--7 March, 2020..
<p>An XML archive of the blogs posted at https://www.ch.ic.ac.uk/rzepa/blog during the period 2008-2020.</p>
Romanian micro-blogging named entity recognition (MicroBloggingNERo)
<p>MicroBloggingNERo is a manually annotated corpus for named entity recognition in Romanian micro-blogging texts. <br> It provides gold annotations for organizations, locations, persons, time expressions, legal references, medical devices,<br> chemicals, anatomical parts and disorders found in micro-blogging texts. The text was anonymized, by replacing all<br> URLs with <url>, user references with <user>, person names, specific locations and organizations with new randomized names. <br> Anonymization was realized in the same way, regardless of the micro-blogging platform specific format.</p> <p>Since names were replaced with new random ones, any resemblance to real individuals is by pure chance of the random names<br> generator. No real person is depicted in the included messages.</p> <p><br> DATA</p> <p>The MicroBloggingNERo corpus is available in different formats: text, span-based, and token-based. </p> <p>Text files are in the folder "text" with .txt extension, in UTF-8 encoding.</p> <p>Span-based annotations are given in BRAT (https://brat.nlplab.org/) ann format. These annotations can be found in folders starting with "ann_".</p> <p>Token-based annotations are given in CONLLUP files, following the CoNLL-U Plus format https://universaldependencies.org/ext-format.html .<br> Part-of-speech tagging was realized using UDPIPE. <br> Named entity annotations are placed in the column "RELATE:NE" (the 11th column) as defined in the "global.columns" metadata field.<br> Automatic processing was performed through the RELATE platform (https://relate.racai.ro).</p> <p>The archive contains: </p> <p>- ann_EVERYTHING <br> Folder in which all the files are in .ann format and contains annotations of: legal references, persons, locations, organizations, time, chemicals, medical devices, anatomical parts and disorders. <br> Overlapping annotations of organizations and time entities inside legal references were allowed. </p> <p>- ann_EVERYTHING_LARGEST_SPAN <br> Folder in which all the files are in .ann format and contains annotations of: legal references, persons, locations, organizations, time, chemicals, medical devices, anatomical parts and disorders. <br> Overlapping annotations were not allowed and only the longest named entities were annotated. This affects primarily the legal references class.</p> <p>- ann_LEGAL_PER_LOC_ORG_TIME <br> Folder in which all the files are in .ann format and contains annotations of: legal references, persons, locations, organizations and time. <br> There are no overlapping annotations. </p> <p>- ann_PER_LOC_ORG_TIME <br> Folder in which all the files are in .ann format and contains annotations of: persons, locations, organizations and time. <br> There are no overlapping annotations. </p> <p>- ann_BIOMEDICAL<br> Folder in which all the files are in .ann format and contains annotations of: medical devices, chemicals, anatomical parts and disorders. <br> There are no overlapping annotations. </p> <p>- conllup_EVERYTHING_LARGEST_SPAN<br> Folder in which all the files are in .ann format and contains annotations of: legal references, persons, locations, organizations, time, chemicals, medical devices, anatomical parts and disorders. <br> There are no overlapping annotations. </p> <p>- conllup_LEGAL_PER_LOC_ORG_TIME <br> Folder in which all the files are in .conllup format and contains annotations of: legal references, persons, locations, organizations and time. <br> Overlapping annotations were not allowed and only the longest named entities were annotated. </p> <p>- conllup_PER_LOC_ORG_TIME <br> Folder in which all the files are in .conllup format and contains annotations of: persons, locations, organizations and time. <br> Overlapping annotations were not allowed and only the longest named entities were annotated. </p> <p>- conllup_BIOMEDICAL<br> Folder in which all the files are in .conllup format and contains annotations of: medical devices, chemicals, anatomical parts and disorders. <br> Overlapping annotations were not allowed and only the longest named entities were annotated. </p> <p>- text <br> Folder containing the raw texts.</p> <p>- splits.tsv<br> Proposed splits into train,test,valid following a distribution of 70-15-15% for each entity class, based on the ann_EVERYTHING_LARGEST_SPAN folder</p> <p>LICENSING</p> <p>This work is provided under the license CC BY-NC-ND 4.0 (Attribution-NonCommercial-NoDerivatives 4.0 International).<br> The license can be viewed online here: https://creativecommons.org/licenses/by-nc-nd/4.0/ <br> and the full text here: https://creativecommons.org/licenses/by-nc-nd/4.0/legalcode . </p> <p><br> CONTACT</p> <p>Research Institute for Artificial Intelligence "Mihai Draganescu", Romanian Academy<br> Web: http://www.racai.ro <br> Contact emails: vasile@racai.ro , maria@racai.ro , vergi@racai.ro , elena@racai.ro</p>
Supplementary Material for "Is Personalization Worth It? Notifying Blogs about a Privacy Issue Resulting from Poorly Implemented Consent Banners"
<p>This record contains the supplementary material for the paper "Is Personalization Worth It? Notifying Blogs about a Privacy Issue Resulting from Poorly Implemented Consent Banners" which was published at ARES 2024. In detail, we present all of our message texts (initial notification, reminder, debriefing and responses to individual replies), the translated questions of our short survey, and the anonymized data of our study outcome.</p>
Data from: Bringing ecology blogging into the scientific fold: measuring reach and impact of science community blogs
The popularity of science blogging has increased in recent years, but the number of academic scientists who maintain regular blogs is limited. The role and impact of science communication blogs aimed at general audiences is often discussed, but the value of science community blogs aimed at the academic community has largely been overlooked. Here, we focus on our own experiences as bloggers to argue that science community blogs are valuable to the academic community. We use data from our own blogs (n = 7) to illustrate some of the factors influencing reach and impact of science community blogs. We then discuss the value of blogs as a standalone medium, where rapid communication of scholarly ideas, opinions, and short observational notes can enhance scientific discourse, and discussion of personal experiences can provide indirect mentorship for junior researchers and scientists from underrepresented groups. Finally, we argue that science community blogs can be treated as a primary source and provide some key points to consider when citing blogs in peer-reviewed literature.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.