Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

17

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

17 results for “authorship analysis”

Learn how ShareScore rates datasets ↗
zenodo40/100

The Authorship of Stephen King's Books Written Under the Pseudonym "Richard Bachman": A Stylometric Analysis (data)

<p>This data accompanies a paper for the 2nd Annual Conference for Computational Literary Studies: &quot;The Authorship of Stephen King&rsquo;s Books Written Under the Pseudonym &#39;Richard Bachman&#39;: A Stylometric Analysis&quot;.</p> <p><strong>Abstract</strong>:</p> <p>Between 1977 and 1984, Stephen King published five novels under the pseudonym &ldquo;Richard Bachman&rdquo;. Reviewers noted similarities between King&rsquo;s and Bachman&rsquo;s writing styles when <em>Thinner&nbsp;</em>(1984) was published, ultimately leading to King&rsquo;s unmasking. We investigate, using the Juola protocol, whether computational techniques can correctly identify King as the author of the Bachman books out of a selection of contemporary candidate authors &ndash; Dean Koontz, Peter Straub, and Thomas Harris. We also perform a post-hoc analysis of the use of pop-culture references and brand names in Bachman, King, Koontz, Straub, and Harris novels, based on comments in reviews of Bachman and King novels. The references extracted from the Bachman books occurred significantly more often in King&rsquo;s texts than in the others&rsquo;, showing that attentive readers could have &ldquo;heard King&rsquo;s voice&rdquo; in the Bachman books through what a reviewer denigratingly called King&rsquo;s &ldquo;compulsion to list brand-name products and his affinity for pop-cult teenage junk&rdquo;. These results contribute to the vexed issue of explainability, which is a recurrent challenge in author identification for literary texts.</p> <p>&nbsp;</p> <p>Below is a description of each file in this repository:</p> <p><strong>bachman_segments_features_array_1000token_segments.csv</strong>,&nbsp;<strong>bachman_segments_features_array_5000token_segments.csv</strong>, and <strong>bachman_segments_features_array_10000token_segments.csv</strong> contain the feature spaces created by vectorizing 1,000-, 5,000-, and 10,000-token segments of Bachman, King, Koontz, and Straub books. Each row of the csv files contains the vectorized segment, the segment&#39;s author, the book the segment was drawn from, the book&#39;s publication date, and&nbsp;the segment number.&nbsp;</p> <p>&nbsp;</p> <p><strong>bachman_segments_author_candidate_cosine_distances_1000token_segments.csv</strong>,&nbsp;<strong>bachman_segments_author_candidate_cosine_distances_5000token_segments.csv</strong>,and&nbsp;<strong>bachman_segments_author_candidate_cosine_distances_10000token_segments.csv&nbsp;</strong>contain the Bachman segment number, bootstrap iteration number (from 0 and 9,999), the distractor&nbsp;author of the randomly-sampled segment, and the cosine distance between the Bachman segment vector and the distractor author&#39;s randomly-sampled segment vector (calculated using the data stored in the&nbsp;bachman_segments_features_array_1000token_segments.csv,&nbsp;bachman_segments_features_array_5000token_segments.csv, and bachman_segments_features_array_10000token_segments.csv files).</p> <p>&nbsp;</p> <p><strong>bachman_segments_author_candidate_ranks_1000token_segments.csv</strong>,&nbsp;<strong>bachman_segments_author_candidate_ranks_5000token_segments.csv</strong>,and&nbsp;<strong>bachman_segments_author_candidate_ranks_10000token_segments.csv </strong>contain the same columns as the 3 files described in the previous paragraph, but the cosine distance between Bachman segment and distractor author segment is converted to a ranking. For each bootstrap iteration there are 4 (one for each candidate author)&nbsp;rows containing the distance ranking between the Bachman segment and a candidate author segment. In a particular bootstrap iteration, if a King segment had the smallest cosine distance to a Bachman segment, King has the ranking &quot;1&quot;, and if a Koontz segment had the second smallest distance to a Bachman segment, Koontz has the ranking &quot;2&quot;, and so on.&nbsp;</p> <p>&nbsp;</p> <p><strong>predicted_author_candidate_raw_counts_1000token_segments.csv</strong>,<strong>&nbsp;predicted_author_candidate_raw_counts_5000token_segments.csv</strong>, and<strong> predicted_author_candidate_raw_counts_10000token_segments.csv&nbsp;</strong>contain the total number of times King, Straub, Harris, and Koontz segments received a certain distance ranking in the files described in the previous paragraph.&nbsp;</p> <p>&nbsp;</p> <p><strong>predicted_author_candidate_proportions_1000token_segments.csv</strong>, <strong>predicted_author_candidate_proportions_5000token_segments.csv</strong>, and and&nbsp;<strong>predicted_author_candidate_proportions_10000token_segments.csv</strong>&nbsp;contain a Bachman book title, and percentage of that book&#39;s&nbsp;segments that received the distance rankings 1-4 of each author. For example, in&nbsp;<strong>predicted_author_candidate_proportions_10000token_segments.csv, </strong><em>The Long Walk</em>&#39;s segments were ranked as most similar&nbsp;(rank= &quot;1&quot;) to King segments in 73.3% of bootstrap iterations.&nbsp;</p> <p>&nbsp;</p> <p><strong>pop_culture_refs_counts_books_10000token_segments.csv&nbsp;</strong>contains the author and book title of a randomly-sampled 10,000-token segment from the aforementioned book, the iteration (from 0 to 99), and the number of pop culture references found in the segment that match those extracted from Bachman books.</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0May 2023View details →
zenodo36/100

Network visualisation of co-authorship analysis of countries

<p>VOSviewer mapping shows network visualisation of co-authorship analysis of countries for&nbsp;<span>focusing on MICP research in the context of hydrodynamics (1999-2024).</span></p>

opencc-by-4.0Apr 2024View details →
dryad36/100

International comparison of cross-disciplinary integration in industry 4.0: A co-authorship analysis using academic literature databases

<p>In innovation strategy, a type of Schumpeterian competitive strategy in business administration, "intra-individual diversity" has attracted attention as one factor for creating innovation. In this study, we redefine "framework for identifying researchers' areas of expertise" as "a framework for quantifying intra-individual diversity among researchers. Note that diversity here refers to authorship of articles in multiple research fields. The application of this framework then made it possible to visualize organizational diversity by accumulating the intra-individual diversity of researchers and to discuss the innovation strategy of the organization. The analysis in this study discusses how countries are promoting research on the topics of artificial intelligence (AI), big data, and Internet of Things (IoT) technologies, which are at the core of Industry 4.0, from an innovation perspective. Note that Industry 4.0 is a technological framework that aims to "improve the efficiency of all social systems," "create new industries," and "increase intellectual productivity."  For the analysis, we used 19-year bibliographic data (2000–2018) from the top 20 countries in terms of the number of papers in AI, big data, and IoT technologies. As the results, this study classified the styles of cross-disciplinary fusion into four patterns in AI and three patterns in big data. This study did not consider the results in IoT because of only small differences between countries. Furthermore, regional differences in the style of cross-disciplinary fusion were also observed, and the global innovation patterns in Industry 4.0 were classified into seven categories. In Europe and North America, the cross-disciplinary integration style was similar to that between the United States, Germany, the Netherlands, Spain, England, Italy, Canada, and France. In Asia, the cross-disciplinary fusion style was similar between China, Japan, and South Korea.</p>

opencc-zeroSep 2022View details →
zenodo36/100

Dataset of "Gender analysis and co-authorship networks in the scientific production of oncology in Spain (2011–2021)"

Open the record for dataset details and reuse information.

opencc-by-4.0Sep 2024View details →
dryad36/100

International comparison of cross-disciplinary integration in industry 4.0: A co-authorship analysis using academic literature databases

Open the record for dataset details and reuse information.

publicSep 2022View details →
zenodo32/100

PAN19 Authorship Analysis: Cross-Domain Authorship Attribution

<p>Authorship attribution is an important problem in information retrieval and computational linguistics but also in applied areas such as law and journalism where knowing the author of a document (such as a ransom note) may enable e.g. law enforcement to save lives. The most common framework for testing candidate algorithms is the closed-set attribution task: given a sample of reference documents from a restricted and finite set of candidate authors, the task is to determine the most likely author of a previously unseen document of unknown authorship. This task may be quite challenging in&nbsp;<strong>cross-domain conditions</strong>, when documents of known and unknown authorship come from different domains (e.g., thematic area, genre). In addition, it is often more realistic to assume that the true author of a disputed document is not necessarily included in the list of candidates.</p> <p><strong>Fanfiction</strong>&nbsp;refers to fictional forms of literature which are nowadays produced by admirers (&#39;fans&#39;) of a certain author (e.g. J.K. Rowling), novel (&#39;Pride and Prejudice&#39;), TV series (Sherlock Holmes), etc. The fans heavily borrow from the original work&#39;s theme, atmosphere, style, characters, story world etc. to produce new fictional literature, i.e. the so-called&nbsp;<strong>fanfics</strong>. This is why fanfiction is also known as transformative literature and has generated a number of controversies in recent years related to the intellectual rights property of the original authors (cf. plagiarism). Fanfiction, however, is typically produced by fans without any explicit commercial goals. The publication of fanfics typically happens online, on informal community platforms that are dedicated to making such literature accessible to a wider audience (e.g.&nbsp;<a href="https://www.fanfiction.net/">fanfiction.net</a>). The original work of art or genre is typically refered to as a&nbsp;<strong>fandom</strong>.</p> <p>This edition of PAN focuses on cross-domain attribution in fanfiction, a task that can be more accurately described as&nbsp;<strong>cross-fandom attribution in fanfiction</strong>. In more detail, all documents of unknown authorship are fanfics of the same fandom (target fandom) while the documents of known authorship by the candidate authors are fanfics of several fandoms (other than the target-fandom). In contrast to the PAN-2018 edition of this task, we focus on&nbsp;<strong>open-set attribution</strong>&nbsp;conditions, namely the true author of a text in the target domain is not necessarily included in the list of candidate authors.</p> <p>Each problem consists of a set of known fanfics by each candidate author and a set of unknown fanfics located in separate folders. The file&nbsp;<code>problem-info.json</code>&nbsp;that can be found in the main folder of each problem, shows the name of folder of unknown documents and the list of names of candidate author folders.</p> <p>The fanfics of known authorship belong to several fandoms (excluding the target fandom). The file&nbsp;<code>fandom-info.json</code>&nbsp;(it can be found in the main folder of each problem) provides information about the fandom of all fanfics of known authorsihp, as follows.</p> <p>The true author of each unknown document can be seen in the file&nbsp;<code>ground-truth.json</code>, also found in the main folder of each problem. Note that all unknown documents that are not written by any of the candidate authors belong to the&nbsp;<code>&lt;UNK&gt;</code>&nbsp;class.</p> <p>In addition, to handle a collection of such problems, the file&nbsp;<code>collection-info.json</code>&nbsp;includes all relevant information. In more detail, for each problem it lists its main folder, the language (either&nbsp;<code>&quot;en&quot;</code>,&nbsp;<code>&quot;fr&quot;</code>,&nbsp;<code>&quot;it&quot;</code>, or&nbsp;<code>&quot;sp&quot;</code>), and the encoding (always&nbsp;<code>UTF-8</code>) of documents.</p>

openDec 2018View details →
zenodo32/100

PAN20 Authorship Analysis: Celebrity Profiling

<p><strong>Synopsis</strong></p> <ul> <li>Task: Given the Twitter feeds of the followers, determine the occupation, age, and gender of a celebrity.</li> <li>Evaluation: [<a href="https://github.com/pan-webis-de/pan-code/tree/master/clef20/celebrity-profiling">code</a>]</li> <li>Baselines: [code]</li> <li>See the full Shared Task [<a href="https://pan.webis.de/clef20/pan20-web/celebrity-profiling.html">here</a>]</li> </ul> <p>The datasets contain three files: a <code>follower-feeds.ndjson</code> as input, a <code>labels.ndjson</code> as output, and a <code>celebrity-feeds.ndjson</code> for additional study. Each file lists all celebrities as JSON objects, one per line and identified by the <code>id</code> key. The training dataset contains 1,920 celebrities and is balanced towards gender and occupation. The supplement dataset contains the remaining 8,265 celebrities but is not balanced in any way.</p> <p>&nbsp;</p> <p>The <code>follower-feeds.ndjson</code> contains the English tweets of at least 10 followers for each celebrity, with at least 50 tweets each excluding retweets.</p> <pre><code class="language-json">{"id": 1234, "text": [["a tweet of follower 1", "another tweet of follower 1", ...], ["a tweet of follower 2", ...], ...]} {"id": 5678, "text": [["a tweet of follower 1", "another tweet of follower 1", ...], ["a tweet of follower 2", ...], ...]}</code></pre> <p>&nbsp;</p> <p>The <code>celebrity-feeds.ndjson</code> contains the Twitter timelines of the original celebrities, formatted as:</p> <pre><code class="language-json">{"id": 1234, "text": ["a tweet of celebrity 1", "another tweet of celebrity 1", ...]} {"id": 5678, "text": ["a tweet of celebrity 2", "another tweet", ...]}</code></pre> <p>&nbsp;</p> <p>The <code>labels.ndjson</code> contains the classes that should be predicted. A valid submission has to produce a <code>labels.ndjson</code> given the <code>follower-feeds.ndjson</code> and contain an entry for each <code>id</code> given in the input.</p> <pre><code class="language-json">{"id": 1234, "occupation": "sports", "gender": "female", "birthyear": 2002} {"id": 5678, "occupation": "professional", "gender": "male", "birthyear": 1990}</code></pre> <p>The following values are possible for each of the traits:</p> <pre><code>occupation := {sports, performer, creator, politics} birthyear := {1940, ..., 1999} gender := {male, female}</code></pre> <p>&nbsp;</p>

openFeb 2020View details →
zenodo32/100

PAN20 Authorship Analysis: Authorship Verification

<p><strong>Task</strong></p> <p>Authorship verification is the task of deciding whether two texts have been written by the same author based on comparing the texts&#39; writing styles.</p> <p>In the coming three years at PAN&nbsp;2020 to PAN&nbsp;2022, we develop a new experimental setup that addresses three key questions in authorship verification that have not been studied at scale to date:</p> <ul> <li> <p>Year 1 (PAN 2020): Closed-set verficiation.<br> Given a large training dataset comprising of known authors who have written about a given set of topics, the test dataset contains verification cases from a subset of the authors and topics found in the training data.</p> </li> <li> <p>Year 2 (PAN 2021): Open-set verification.<br> Given the training dataset of Year&nbsp;1, the test dataset contains verification cases from previously unseen authors and topics.</p> </li> <li> <p>Year 3 (PAN 2022):&nbsp;<em>Suprise task</em>.<br> The task of the last year of this evaluation cycle (to be announced at a later time) will be designed with an eye on realism and practical application.</p> </li> </ul> <p>This evaluation cycle on authorship verification provides for a renewed challenge of increasing difficulty within a large-scale evaluation. We invite you to plan ahead and participate in all three of these tasks.</p> <p>More information at:&nbsp;<a href="https://pan.webis.de/clef20/pan20-web/author-identification.html">PAN @ CLEF 2020 - Authorship Verification</a></p> <p>&nbsp;</p> <p><strong>Citing the Dataset</strong></p> <p>If you use this dataset for your research, please be sure to cite the following paper:<br> <br> Sebastian Bischoff, Niklas Deckers, Marcel Schliebs, Ben Thies, Matthias Hagen, Efstathios Stamatatos, Benno Stein, and Martin Potthast.&nbsp;The Importance of Suppressing Domain Style in Authorship Analysis.&nbsp;CoRR,&nbsp;abs/2005.14714,&nbsp;May&nbsp;2020.</p> <p>Bibtex:</p> <pre><code>@Article{stein:2020k, author = {Sebastian Bischoff and Niklas Deckers and Marcel Schliebs and Ben Thies and Matthias Hagen and Efstathios Stamatatos and Benno Stein and Martin Potthast}, journal = {CoRR}, month = may, title = {{The Importance of Suppressing Domain Style in Authorship Analysis}}, url = {https://arxiv.org/abs/2005.14714}, volume = {abs/2005.14714}, year = 2020 }</code></pre> <p>&nbsp;</p>

openDec 2019View details →
zenodo28/100

Research collaboration patterns in sustainable mining – a co-authorship analysis of publications

<p>This dataset includes data related to 4220 articles on sustainable mining published from 1983 to 2018. The Scopus database was selected as a data source. Detailed data applies to co-authored articles. The number of authors and affiliations (country, institution, sector) were taken into account. Data has been cleaned in terms of names of institutions and countries.<br> In given sets the following data were included:<br> - Distribution of articles in sustainable mining from 1983 to 2018<br> - Distribution of joint articles and the types of joint articles from 1983 to 2018<br> - Team size in terms of the number of authors of articles in sustainable mining from 1983 to 2018<br> - Team size in terms of the number of authors&#39; institutions in articles in sustainable mining from 1983 to 2018<br> - Team size in terms of the number of authors&#39; countries in articles in sustainable mining from 1983 to 2018</p> <p><strong>Please note, the field separator used in these files is a semicolon, while the decimal separator is a comma. Each of the files has two header lines.</strong></p>

opencc-by-4.0May 2020View details →
zenodo28/100

PAN22 Authorship Analysis: Style Change Detection

<p>This is the dataset for the <a href="https://pan.webis.de/clef22/pan22-web/style-change-detection.html">Style Change Detection</a> task of PAN 2022.</p> <p><strong>Task</strong></p> <p>The goal of the style change detection task is to identify text positions within a given multi-author document at which the author switches. Hence, a fundamental question is the following: If multiple authors have written a text together, can we find evidence for this fact; i.e., do we have a means to detect variations in the writing style? Answering this question belongs to the most difficult and most interesting challenges in author identification: Style change detection is the only means to detect plagiarism in a document if no comparison texts are given; likewise, style change detection can help to uncover gift authorships, to verify a claimed authorship, or to develop new technology for writing support.</p> <p>Previous editions of the Style Change Detection task aim at e.g., detecting whether a document is single- or multi-authored (<a href="https://pan.webis.de/clef18/pan18-web/style-change-detection.html">2018</a>), the actual number of authors within a document (<a href="https://pan.webis.de/clef19/pan19-web/style-change-detection.html">2019</a>), whether there was a style change between two consecutive paragraphs (<a href="https://pan.webis.de/clef20/pan20-web/style-change-detection.html">2020</a>,&nbsp;<a href="https://pan.webis.de/clef21/pan21-web/style-change-detection.html">2021</a>) and where the actual style changes were located (<a href="https://pan.webis.de/clef21/pan21-web/style-change-detection.html">2021</a>). Based on the progress made towards this goal in previous years, we again extend the set of challenges to likewise entice novices and experts:</p> <p>Given a document, we ask participants to solve the following three tasks:</p> <ul> <li><strong>[Task1] Style Change Basic:</strong>&nbsp;for a text written by two authors that contains a single style change only, find the position of this change (i.e., cut the text into the two authors&rsquo; texts on the paragraph-level),</li> <li><strong>[Task2] Style Change Advanced:</strong>&nbsp;for a text written by two or more authors, find all positions of writing style change (i.e., assign all paragraphs of the text uniquely to some author out of the number of authors assumed for the multi-author document)</li> <li><strong>[Task3] Style Change Real-World:</strong>&nbsp;for a text written by two or more authors, find all positions of writing style change, where style changes now not only occur between paragraphs, but at the sentence level.</li> </ul> <p>All documents are provided in English and may contain an arbitrary number of style changes, resulting from at most five different authors.</p> <p><strong>Data</strong></p> <p>To develop and then test your algorithms, three datasets including ground truth information are provided (<em>dataset1</em>&nbsp;for task 1,&nbsp;<em>dataset2</em>&nbsp;for task 2, and&nbsp;<em>dataset3</em>&nbsp;for task 3).</p> <p>Each dataset is split into three parts:</p> <ol> <li><em>training set:</em>&nbsp;Contains 70% of the whole dataset and includes ground truth data. Use this set to develop and train your models.</li> <li><em>validation set:</em>&nbsp;Contains 15% of the whole dataset and includes ground truth data. Use this set to evaluate and optimize your models.</li> <li><em>test set:</em>&nbsp;Contains 15% of the whole dataset, no ground truth data is given. This set is used for evaluation (see later).</li> </ol> <p>You are free to use additional external data for training your models. However, we ask you to make the additional data utilized freely available under a suitable license.</p> <p><strong>Input Format</strong></p> <p>The datasets are based on user posts from various sites of the StackExchange network, covering different topics. We refer to each input problem (i.e., the document for which to detect style changes) by an ID, which is subsequently also used to identify the submitted solution to this input problem. We provide one folder for train, validation, and test data for each dataset, respectively.</p> <p>For each problem instance&nbsp;<code>X</code>&nbsp;(i.e., each input document), two files are provided:</p> <ol> <li><code>problem-X.txt</code>&nbsp;contains the actual text, where paragraphs are denoted by&nbsp;<code>\n</code>&nbsp;for tasks 1 and 2. For task 3, we provide one sentence per paragraph (again, split by&nbsp;<code>\n</code>).</li> <li><code>truth-problem-X.json</code>&nbsp;contains the ground truth, i.e., the correct solution in JSON format. An example file is listed in the following (note that we list keys for the three tasks here): <pre><code>{ "authors": NUMBER_OF_AUTHORS, "site": SOURCE_SITE, "changes": RESULT_ARRAY_TASK1 or RESULT_ARRAY_TASK3, "paragraph-authors": RESULT_ARRAY_TASK2 }</code></pre> <p>The result for task 1 (key &quot;changes&quot;) is represented as an array, holding a binary for each pair of consecutive paragraphs within the document (0 if there was no style change, 1 if there was a style change). For task 2 (key &quot;paragraph-authors&quot;), the result is the order of authors contained in the document (e.g.,&nbsp;<code>[1, 2, 1]</code>&nbsp;for a two-author document), where the first author is &quot;1&quot;, the second author appearing in the document is referred to as &quot;2&quot;, etc. Furthermore, we provide the total number of authors and the Stackoverflow site the texts were extracted from (i.e., topic). The result for task 3 (key &quot;changes&quot;) is similarly structured as the results array for task 1. However, for task 3, the&nbsp;<code>changes</code>&nbsp;array holds a binary for each pair of consecutive&nbsp;<em>sentences</em>&nbsp;and they may be multiple style changes in the document.</p> <p>An example of a multi-author document with a style change between the third and fourth paragraph (or sentence for task 3) could be described as follows (we only list the relevant key/value pairs here):</p> <pre><code>{ "changes": [0,0,1,...], "paragraph-authors": [1,1,1,2,...] }</code></pre> <p>&nbsp;</p> </li> </ol> <p><strong>Output Format</strong></p> <p>To evaluate the solutions for the tasks, the results have to be stored in a single file for each of the input documents and each of the datasets. Please note that we require a solution file to be generated for each input problem for each dataset. The data structure during the evaluation phase will be similar to that in the training phase, with the exception that the ground truth files are missing.</p> <p>For each given problem&nbsp;<code>problem-X.txt</code>, your software should output the missing solution file&nbsp;<code>solution-problem-X.json</code>, containing a JSON object holding the solution to the respective task. The solution for tasks 1 and 3 is an array containing a binary value for each pair of consecutive paragraphs (task 1) or sentences (task 3). For task 2, the solution is an array containing the order of authors contained in the document (as in the truth files).</p> <p>An example solution file for tasks 1 and 3 is featured in the following (note again that for task 1, changes are captured on the paragraph level, whereas for task 3, changes are captured on the sentence level):</p> <pre><code>{ "changes": [0,0,1,0,0,...] }</code></pre> <p>For task 2, the solution file looks as follows:</p> <pre><code>{ "paragraph-authors": [1,1,2,2,3,2,...] }</code></pre> <p>&nbsp;</p>

openMar 2022View details →
dryad28/100

Trends in authorship demographics for manuscripts published in Endocrine journals: A 70-year analysis

<p><em><span>Background</span> </em></p> <p><span>Over the previous few decades, demographics, gender, and the amount of papers published have all changed considerably. One of the fields of medicine that has yet to be extensively investigated is endocrinology.</span> </p> <p><em><span>Material and Methods</span> </em></p> <p><span>Journal of Endocrinology and General &amp; Comparative Endocrinology are two landmark journals that publish articles from around the world. We examined each decade during the 70-year period from 1961 to 2021. Funding source, first author – last author gender, their demographics and proportion of papers with at least one female author were the parameters considered while studying each publication. We predicted that the number of female authors per paper would increase with time, as would the range of degrees held by the authors, demographical variations in authorship, and the funding source. Our goal was also to determine the distribution of female first authors and senior authors in endocrinology journals over a 70-year period, as well as to check the gender combinations using the Punnett square. </span> </p> <p><em><span>Results</span> </em></p> <p><span>Female initial authors rose from 7% to 29.6% (p&lt;0.0006) between 1961 and 2021, whereas female senior authors rose from 15.6% to 22.2%. Despite women's small contributions to first and senior authors, female participation rose from 17.48% (25/143) to 70% (170/250) between 1961 and 2021. Male-Female and Female-Male combinations rose with Chi-Square = 124.6, (p&lt;0.0001). Europe and the Americas had the most female academic medical contributors (p&lt;0.0001) Regardless of author status, female participation rose from 17.48% in 1961 to 68% in 2021.</span> </p> <p><em><span>Conclusion </span> </em></p> <p><span>In papers published in endocrinology journals, there was a rising trend in female contributions to academic medicine. Even with the large growth of female endocrinologists, there is still a disparity in why the increase in female authors is comparably fewer.</span> </p>

opencc-zeroMay 2022View details →
dryad28/100

Trends in authorship demographics for manuscripts published in Endocrine journals: A 70-year analysis

Open the record for dataset details and reuse information.

publicMay 2022View details →
zenodo24/100

PAN19 Authorship Analysis: Celebrity Profiling

<p><strong>Paper:</strong> https://webis.de/publications.html?q=wiegmann_2019a</p><p><strong>Source Dataset:</strong> https://files.webis.de/data-in-progress/data-research/social-media-analysis/acl19-celebrity-profiling/</p><p>&nbsp;</p><p>Celebrities are among the most prolific users of social media, promoting their personas and rallying followers. This activity is closely tied to genuine writing samples, rendering them worthy research subjects in many respects, not least author profiling.</p><p>The Celebrity Profiling task this year is to predict four traits of a celebrity from their social media communication. The traits are the degree of fame, occupation, age, and gender. The social media communication is given as the teaser messages from past tweets. The goal is to develop a piece of software which predicts celebrity traits from the teaser history.</p><p>The training dataset contains&nbsp;two files: a&nbsp;feeds.ndjson&nbsp;as input and a&nbsp;labels.ndjson&nbsp;as output. Each file lists all celebrities as JSON objects, one per line and identified by the&nbsp;id&nbsp;key.</p><p>The input file contains the cid and a list of all teaser messages for each celebrity.</p><p>{"id":&nbsp;1234,&nbsp;"text":&nbsp;["a tweet",&nbsp;"another tweet",&nbsp;...]}</p><p>The output file contains the cid and&nbsp;a value for each trait for each celebrity from the input file.</p><p>{"id":&nbsp;1234,&nbsp;"fame":&nbsp;"star",&nbsp;"occupation":&nbsp;"sports",&nbsp;"gender":&nbsp;"female",&nbsp;"birthyear":&nbsp;2002}</p><p>The following values are possible for each of the traits:</p><p>fame&nbsp;:=&nbsp;{rising,&nbsp;star,&nbsp;superstar}&nbsp; occupation&nbsp;:=&nbsp;{sports,&nbsp;performer,&nbsp;creator,&nbsp;politics,&nbsp;manager,&nbsp;science,&nbsp;professional,&nbsp;religious}&nbsp; birthyear&nbsp;:=&nbsp;{1940,&nbsp;...,&nbsp;2012}&nbsp; gender&nbsp;:=&nbsp;{male,&nbsp;female,&nbsp;nonbinary}</p><p>&nbsp;</p>

openJan 2019View details →
zenodo20/100

PAN19 Authorship Analysis: Style Change Detection

<p>This is the data set for the <a href="http://pan.webis.de/clef19/pan19-web/style-change-detection.html">Style Change Detection task</a>&nbsp;of <a href="http://pan.webis.de/clef19/pan19-web/">PAN@CLEF 2019</a>.</p> <p>The goal of the style change detection task is to identify text positions within a given multi-author document at which the author switches. Detecting these positions is a crucial part of the authorship identification process, and for multi-author document analysis in general. Note that, for this task, we make the assumption that a change in writing style always signifies a change in author.</p> <p><strong>Tasks</strong></p> <p>Given a document, we ask participants to answer the following two questions:</p> <ul> <li>Was the given document written by multiple authors? (task 1)</li> <li>For each pair of consecutive paragraphs in the given document: is there a style change between these paragraphs? (task 2)</li> </ul> <p>In other words, the goal is to determine whether the given document contains style changes and if it indeed does, we aim to find the position of the change in the document (between paragraphs).</p> <p>All documents are provided in English and may contain zero up to ten style changes, resulting from at most three different authors. However, style changes may only occur between paragraphs (i.e., a single paragraph is always authored by a single author and does not contain any style changes).</p> <p><strong>Data</strong></p> <p>To develop and then test your algorithms, two data sets including ground truth information are provided. Those data sets differ in their topical breadth (i.e., the number of different topics that are covered in the documents contained).&nbsp;dataset-narrow&nbsp;contains texts from a relatively narrow set of subjects matters (all related to technology), whereas&nbsp;dataset-wide&nbsp;adds additional subject areas to that (travel, philosophy, economics, history, etc.).</p> <p>Both of those data sets are split into three parts:</p> <ul> <li><em>training set:</em>&nbsp;Contains 50% of the whole data set and includes ground truth data. Use this set to develop and train your models.</li> <li><em>validation set:</em>&nbsp;Contains 25% of the whole data set and includes ground truth data. Use this set to evaluate and optimize your models.</li> <li><em>test set:</em>&nbsp;Contains 25% of the whole data set. For the documents on the test set, you are not given ground truth data. This set is used for evaluation.</li> </ul> <p><br> <strong>Input Format</strong></p> <p>Both dataset-narrow and dataset-wide are based on user posts from various sites of the StackExchange network, covering different topics. We refer to each input problem (i.e., the document for which to detect style changes) by an ID, which is subsequently also used to identify the submitted solution to this input problem.</p> <p>The structure of the provided datasets is as follows:</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp;</p> <pre><code class="language-json">train/     dataset-narrow/     dataset-wide/ validation/     dataset-narrow/     dataset-wide/ test/     dataset-narrow/     dataset-wide/</code></pre> <p>&nbsp;</p> <p>For each problem instance&nbsp;X&nbsp;(i.e., each input document), two files are provided:</p> <p>problem-X.txt&nbsp;contains the actual text, where paragraphs are denoted by&nbsp;\n\n.<br> truth-problem-X.json&nbsp;contains the ground truth, i.e., the correct solution in JSON format:</p> <pre><code class="language-json">{     "authors": NUMBER_OF_AUTHORS,     "structure": ORDER_OF_AUTHORS,     "site": SOURCE_SITE,     "multi-author": RESULT_TASK1,     "changes": RESULT_ARRAY_TASK2 }</code></pre> <p>The result for task 1 (key &quot;multi-author&quot;) is a binary value (1 if the document is multi-authored, 0 if the document is single-authored). The result for task 2 (key &quot;changes&quot;) is represented as an array, holding a binary for each pair of consecutive paragraphs within the document (0 if there was no style change, 1 if there was a style change). If the document is single-authored, the solution to task 2 is an array filled with 0s. Furthermore, we provide the order of authors contained in the document (e.g.,&nbsp;[A1, A2, A1]&nbsp;for a two-author document), the total number of authors and the Stackoverflow site the texts were extracted from (i.e., topic).</p> <p>An example of a multi-author document, where there was a style change between the third and fourth paragraph could look as follows (we only list the two relevant key/value pairs here):&nbsp;</p> <pre><code class="language-json">{     "multi-author": 1,     "changes": [0,0,1,...] }</code></pre> <p>A single-author document would have the following form (again, only listing the two relevant key/value pairs):</p> <pre><code class="language-json">{     "multi-author": 0,     "changes": [0,0,0,...] }</code></pre> <p>&nbsp;</p>

restrictedJan 2019View details →
zenodo16/100

PAN22 Authorship Analysis: Authorship Verification

<p><strong>Download</strong></p> <p>Access to our corpus can be requested via the Aston Institute for Forensic Linguistics Databank:&nbsp;<a href="https://fold.aston.ac.uk/handle/123456789/17">https://fold.aston.ac.uk/handle/123456789/17</a></p> <p><strong>Task</strong></p> <p>Authorship verification is the task of deciding whether two texts have been written by the same author based on comparing the texts&#39; writing styles. In previous editions of PAN, we explored the effectiveness of authorship verification technology in several languages and text genres. In the two most recent editions, cross-domain authorship verification using fanfiction texts was examined. Despite certain differences between fandoms, the task of cross-fandom authorship verification has proved to be relatively feasible. In the current edition, we focus on more challenging scenarios where each author verification case considers two texts that belong to different DTs (cross-DT authorship verification). This will allow us to study the ability of stylometric approaches to capture authorial characteristics that remain stable across DTs even when very different forms of expression are imposed by the DT norms.</p> <p>Based on a new corpus in English, we provide cross-DT authorship verification cases using the following DTs:</p> <ul> <li>Essays</li> <li>Emails</li> <li>Text messages</li> <li>Business memos</li> </ul> <p>The corpus comprises texts of around 100 individuals. All individuals have similar age (18-22) and are native English speakers. The topic of text samples is not restricted while the level of formality can vary within a certain DT (e.g., text messages may be addressed to family members or non-familial acquaintances).</p> <p>More information at:&nbsp;<a href="https://pan.webis.de/clef22/pan22-web/author-identification.html">Authorship Verification 2022</a></p>

restrictedMar 2022View details →
zenodo16/100

PAN21 Authorship Analysis: Style Change Detection

<p>This is the dataset for the <a href="https://pan.webis.de/clef21/pan21-web/style-change-detection.html">Style Change Detection</a> task&nbsp;of PAN 2021.</p> <p>The goal of the style change detection task is to identify text positions within a given multi-author document at which the author switches.&nbsp;</p> <p><strong>Tasks</strong></p> <p>Given a document, we ask participants to answer the following three questions:</p> <ul> <li><em>Single vs. Multiple.</em>&nbsp;Given a text, find out whether the text is written by a single author or by multiple authors (task 1).</li> <li><em>Style Change Basic.</em>&nbsp;Given a text written by two or more authors and that contains a number of style changes, find the position of the changes (task 2).</li> <li><em>Style Change Real-World.</em>&nbsp;Given a text written by two or more authors, find all positions of writing style change, i.e., assign all paragraphs of the text uniquely to some author out of the number of authors you assume for the multi-author document (task 3).</li> </ul> <p>All documents are provided in English and may contain an arbitrary number of style changes, resulting from at most five different authors. However, style changes may only occur between paragraphs (i.e., a single paragraph is always authored by a single author and does not contain any style changes).</p> <p><strong>Data</strong></p> <p>The dataset is split into three parts:</p> <ol> <li><em>training set:</em>&nbsp;Contains 70% of the whole data set and includes ground truth data. Use this set to develop and train your models.</li> <li><em>validation set:</em>&nbsp;Contains 15% of the whole data set and includes ground truth data. Use this set to evaluate and optimize your models.</li> <li><em>test set:</em>&nbsp;Contains 15% of the whole data set. For the documents on the test set, you are not given ground truth data. This set is used for evaluation.</li> </ol> <p>The dataset is based on user posts from various sites of the StackExchange network, covering different topics. We refer to each input problem (i.e., the document for which to detect style changes) by an ID, which is subsequently also used to identify the submitted solution to this input problem. We provide one folder for train, validation, and test data.</p> <p>For each problem instance&nbsp;<code>X</code>&nbsp;(i.e., each input document), two files are provided:</p> <ol> <li><code>problem-X.txt</code>&nbsp;contains the actual text, where paragraphs are denoted by&nbsp;<code>\n\n</code>.</li> <li><code>truth-problem-X.json</code>&nbsp;contains the ground truth, i.e., the correct solution in JSON format: <pre><code class="language-json">{ "authors": NUMBER_OF_AUTHORS, "site": SOURCE_SITE, "multi-author": RESULT_TASK1, "changes": RESULT_ARRAY_TASK2, "paragraph-authors": RESULT_ARRAY_TASK3 }</code></pre> The result for task 1 (key &quot;multi-author&quot;) is a binary value (1 if the document is multi-authored, 0 if the document is single-authored). The result for task 2 (key &quot;changes&quot;) is represented as an array, holding a binary for each pair of consecutive paragraphs within the document (0 if there was no style change, 1 if there was a style change). If the document is single-authored, the solution to task 2 is an array filled with 0s. For task 3 (key &quot;paragraph-authors&quot;), the result is the order of authors contained in the document (e.g.,&nbsp;<code>[1, 2, 1]</code>&nbsp;for a two-author document), where the first author is &quot;1&quot;, the second author appearing in the document is referred to as &quot;2&quot;, etc. Furthermore, we provide the total number of authors and the Stackoverflow site the texts were extracted from (i.e., topic).<br> <br> An example of a multi-author document, where there was a style change between the third and fourth paragraph could look as follows (we only list the relevant key/value pairs here): <pre><code class="language-json">{ "multi-author": 1, "changes": [0,0,1,...], "paragraph-authors": [1,1,1,2,...] }</code></pre> A single-author document would have the following form (again, only listing the relevant key/value pairs): <pre><code class="language-json">{ "multi-author": 0, "changes": [0,0,0,...], "paragraph-authors": [1,1,1,...] }</code></pre> <p>&nbsp;</p> </li> </ol>

restrictedMar 2021View details →
zenodo16/100

PAN20 Authorship Analysis: Style Change Detection

<p>This is the data set for the <a href="https://pan.webis.de/clef20/pan20-web/style-change-detection.html">Style Change Detection task</a>&nbsp;of PAN 2020.</p> <p>The goal of the style change detection task is to identify text positions within a given multi-author document at which the author switches. Detecting these positions is a crucial part of the authorship identification process, and for multi-author document analysis in general. Note that, for this task, we make the assumption that a change in writing style always signifies a change in author.</p> <p><strong>Tasks</strong></p> <p>Given a document, we ask participants to answer the following two questions:</p> <ul> <li>Was the given document written by multiple authors? (task 1)</li> <li>For each pair of consecutive paragraphs in the given document: is there a style change between these paragraphs? (task 2)</li> </ul> <p>In other words, the goal is to determine whether the given document contains style changes and if it indeed does, we aim to find the position of the change in the document (between paragraphs).</p> <p>All documents are provided in English and may contain zero up to ten style changes, resulting from at most three different authors. However, style changes may only occur between paragraphs (i.e., a single paragraph is always authored by a single author and does not contain any style changes).</p> <p><strong>Data</strong></p> <p>To develop and then test your algorithms, two data sets including ground truth information are provided. Those data sets differ in their topical breadth (i.e., the number of different topics that are covered in the documents contained).&nbsp;<em>dataset-narrow</em>&nbsp;contains texts from a relatively narrow set of subjects matters (all related to technology), whereas&nbsp;<em>dataset-wide</em>&nbsp;adds additional subject areas to that (travel, philosophy, economics, history, etc.).</p> <p>Both of those data sets are split into three parts:</p> <ol> <li><em>training set:</em>&nbsp;Contains 50% of the whole data set and includes ground truth data. Use this set to develop and train your models.</li> <li><em>validation set:</em>&nbsp;Contains 25% of the whole data set and includes ground truth data. Use this set to evaluate and optimize your models.</li> <li><em>test set:</em>&nbsp;Contains 25% of the whole data set. For the documents on the test set, you are not given ground truth data. This set is used for evaluation (see later).</li> </ol> <p>&nbsp;</p> <p><strong>Input Format</strong></p> <p>Both dataset-narrow and dataset-wide are based on user posts from various sites of the StackExchange network, covering different topics. We refer to each input problem (i.e., the document for which to detect style changes) by an ID, which is subsequently also used to identify the submitted solution to this input problem.</p> <p>The structure of the provided datasets is as follows:</p> <pre> <code> train/ dataset-narrow/ dataset-wide/ validation/ dataset-narrow/ dataset-wide/ test/ dataset-narrow/ dataset-wide/ </code></pre> <p>For each problem instance&nbsp;<code>X</code>&nbsp;(i.e., each input document), two files are provided:</p> <ol> <li><code>problem-X.txt</code>&nbsp;contains the actual text, where paragraphs are denoted by&nbsp;<code>\n\n</code>.</li> <li><code>truth-problem-X.json</code>&nbsp;contains the ground truth, i.e., the correct solution in JSON format: <pre><code>{ "authors": NUMBER_OF_AUTHORS, "structure": ORDER_OF_AUTHORS, "site": SOURCE_SITE, "multi-author": RESULT_TASK1, "changes": RESULT_ARRAY_TASK2 }</code></pre> <p>The result for task 1 (key &quot;multi-author&quot;) is a binary value (1 if the document is multi-authored, 0 if the document is single-authored). The result for task 2 (key &quot;changes&quot;) is represented as an array, holding a binary for each pair of consecutive paragraphs within the document (0 if there was no style change, 1 if there was a style change). If the document is single-authored, the solution to task 2 is an array filled with 0s. Furthermore, we provide the order of authors contained in the document (e.g.,&nbsp;<code>[A1, A2, A1]</code>&nbsp;for a two-author document), the total number of authors and the Stackoverflow site the texts were extracted from (i.e., topic).</p> <p>An example of a multi-author document, where there was a style change between the third and fourth paragraph could look as follows (we only list the two relevant key/value pairs here):&nbsp;</p> <pre><code>{ "multi-author": 1, "changes": [0,0,1,...] }</code></pre> <p>A single-author document would have the following form (again, only listing the two relevant key/value pairs):</p> <pre><code>{ "multi-author": 0, "changes": [0,0,0,...] }</code></pre> </li> </ol> <p>&nbsp;</p>

restrictedFeb 2020View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record