Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
52
datasets available to search
ShareScore release 0.9.0
Dataset results
52 results for “fiction”
"Is Heidi really happier in the mountains? A mixed-methods investigation of spatial affect in fiction." - Data
<p>This repository provides access to the data used in Grisot, G & Herrmann, J. B. (2024) "Is Heidi really happier in the mountains? A mixed-methods investigation of spatial affect in fiction"</p> <p>It contains the following datasets:</p> <ul> <li><a href="https://zenodo.org/api/records/14235844/draft/files/all_entities.csv/content" target="_blank" rel="noopener noreferrer">all_entities.csv</a>: the spatial entities lists used in the paper (see Grisot, G & Herrmann, J. B., 2023)</li> <li><a href="https://zenodo.org/api/records/14235844/draft/files/corpus_books_aggr_sent_norm.csv/content" target="_blank" rel="noopener noreferrer">corpus_books_aggr_sent_norm.csv</a>: a corpus of N=184 Swiss literary narrative texts written in German between 1822 and 1940 by 69 Swiss authors, with sentiment values and spatial entities identified in each sentence.</li> <li><a href="https://zenodo.org/api/records/14235844/draft/files/heidi_clean_aggr_sent.csv/content" target="_blank" rel="noopener noreferrer">heidi_clean_aggr_sent.csv</a>: the 1880 digitised edition of the novel <em>Heidi</em>, as available from E-Rara, with sentiment values and spatial entities identified in each sentence.</li> <li><a href="https://zenodo.org/api/records/14235844/draft/files/sentiart.csv/content" target="_blank" rel="noopener noreferrer">sentiart.csv</a>: the sentiment lexcon SentiArt (Jacobs, 2019)</li> </ul>
ERC Locus Ludi. Play and Games in Antiquity. Sorting Fun From Fiction: Were 'Tesserae' Gaming Pieces ?
<p>Clare Rowan, Warwick, is the PI of the ERC project Token Communities in the ancient Mediterranean. She explores for the first in depth the central roles that tokens played in cultural, religious, political and economic life in antiquity. Her talk: Sorting Fun From Fiction: Were “tesserae” Gaming Pieces?, was given at the International Conference, Play and Games in Antiquity. Definition, Transmission, Reception, September 17-19, 2018, Swiss Museum of Games.</p> <p>Movie/Music: <a href="https://www.edwanmusic.com/">https://www.edwanmusic.com/</a></p> <p>More about: www.locusludi.ch</p>
EMNLP-23-Bootstrapping-a-Violence-Detector-for-Fan-Fiction
<p>Data for the paper `Trigger Warnings: Bootstrapping a Violence Detector for Fan Fiction`. </p><p><strong>Code</strong>: https://github.com/webis-de/emnlp23-bootstrapping-a-violence-detector-for-fan-fiction</p><p><strong>Publication</strong>: tbd. </p><p><strong>Citation</strong>: https://webis.de/publications.html?q=wolska_2023</p><p> </p>
French Fiction of the 16-18th century
<p>A corpus containing all digitized French novels from the beginning of print (the first entry is from 1473) to the 18th century.</p> <p>French novels of the period have been identified using the Y2 quote of the French National Library Catalog that has served to classify past and present collections of novels in France from 1730 to 1996. Combined use of digitized sources from Gallica, Google Books, Archive.org and other digital library made it possible to attain a high representativeness: 78% of the novels of the 1450-1600 and 68% of the novels of the 1600-1700 have been retrieved.</p> <p>The corpus is part of a planned collection of French Fiction (1050-1920) that will also integrate <a href="https://github.com/Jean-Baptiste-Camps/Geste/">Geste</a> (a medieval corpus curated by Jean-Baptiste Camps) and <a href="https://zenodo.org/record/4751204#.YJv2Xes6_y9">Fictions littéraires de Gallica</a> (a 1600-1950 corpus extracted from Gallica with Pierre-Carl Langlais, with a strong focus on the 19th century). While it aims to bridge the two pre-existing part of the collection, it is also a more ambitious experiment of systematic collection of existing digital sources.</p> <p>The project remains very much a work-in-progress at this stage. Occasional errors in the metadata and the identification of the unique work are still possible. Besides, the identification of multi-volumes remain challenging in digital sources beyond Gallica.</p> <p>The repository includes the following files:</p> <ul> <li>The metadata of available and unavailable file for all novels identified in the 16th century (corpus_roman_metadata_16.tsv) and the 17th century (corpus_roman_metadata_17.tsv). All the editions have been temptatively assigned to a unique work (work_id) based on theo title, the author and additional metadata. This dataset includes both information on a specific digitized volume (volume_file, volume_title, volume_date, volume_edition_id) and on the earliest edition of the work recorded by the French national library (first_edition, first_edition_titre, first_edition_date), as well as the identification of the author (prenom_auteur, nom_auteur) and the complete list of all available edition (list_edition_bnf). When digitized files are not available for a given work, the information on the volume is replaced with a missing data mark (NA).An edition-based dataset was initially contemplated, but it turned out to be much harder than expected: the French National Catalog do not record all the available editions and runs of the period and it would have been necessary to check and create unique edition IDs for numerous Google Books volume.</li> <li>The complete text of the novels when available (corpus_roman_16_text.tsv and corpus_roman_17_text.tsv). The use of contemporary OCR software on early modern text have long yielded poor results as words, typographies, and even letters were markedly different than the corpus theses software were trained on. Consequently, numerous volumes from Gallica have simply no OCR, as the results were below the quality requirement of the digital library. New historical OCR models will mmake it possible to create a reliable OCR on the entire corpus. The dataset includes all the text at the page-level whenever there is some text on the page. Page numbering is based on the absolute numbering of the file, not on the original numbering of the edition.</li> <li>A classified dataset of 159 novels from the 17th century in four major genres of the period: chilvaric novel, love novel, historical novel and comic novel. The classification is based on an exceptional source of 1731, the catalog of novels from Nicolas Lenglet du Fresnoy (published as the second volume of <em>De l'usage des romans</em>). The classified dataset include both the text (as in corpus_roman_17_text.tsv) at the page-level and the lemmatization realized with a trained syntaxic model on 17th century French (<a href="https://github.com/e-ditiones/LEM17">https://github.com/e-ditiones/LEM17</a>)</li> <li>A classification model created with the classified dataset. This "Fresnoy" model has a high accuracy (93%) which can be parrtly attributed to overfitting (as there is a limited amount of novels per genre). The model can be reused with <a href="https://github.com/Numapresse/TidySupervise">Tidysupervise</a>, a small R extension to create supervised text models.</li> </ul>
Extraction and Analysis of Fictional Character Networks: A Survey
<p><strong>Description. </strong>Resources used in our survey of character network extraction and analysis methods. This version of the resources does not correspond to the published paper, but to a later update used in the (longer) arXiv/HAL version of the paper. The resources used in the original published paper are provided in v1.0.0.</p> <p><strong>Updates. </strong>The latest version of the data (especially the bibliographic table) is available on a GitHub page, which is easier to update: <a href="https://compnet.github.io/CharNetReview/">https://compnet.github.io/CharNetReview/</a>. This Zenodo repository is no longer updated.</p> <p><strong>Citation. </strong>If you use these data or figures, please cite the following paper:</p> <ul> <li>V. Labatut and X. Bost, “<em>Extraction and Analysis of Fictional Character Networks: A Survey</em>,” ACM Computing Surveys 52(5):89, 2019. ⟨<a href="https://hal.archives-ouvertes.fr/hal-02173918">hal-02173918</a>⟩ DOI: <a href="http://doi.org/10.1145/3344548">10.1145/3344548</a></li> </ul> <p><br><code>@Article{Labatut2019,</code><br><code> author = {Labatut, Vincent and Bost, Xavier},</code><br><code> title = {Extraction and Analysis of Fictional Character Networks: A Survey},</code><br><code> journal = {ACM Computing Surveys},</code><br><code> year = {2019},</code><br><code> volume = {52},</code><br><code> number = {5},</code><br><code> pages = {89},</code><br><code> doi = {10.1145/3344548},</code><br><code>}</code></p>
Blank predetermination in the Iberian Acheulean: fact or fiction? Insight from the cleaver on flake assemblage from Casal do Azemel site (Leiria, Portugal) by a Geometric Morphometric approach
<p>The increase of data available for the study of the Middle Pleistocene in the Iberian Peninsula has favoured the understanding of the technological trends of the regional Acheulean techno-complex. This has features of Large Flake Acheulean -LFA- with a significant presence of cleavers on flake, a specific tool type that is of great cultural and technological value. Moreover, these tools are privileged to discuss the importance of predetermination in Acheulean assemblages. Following this reason, besides the traditional techno-typological approach, we perform 2D Geometric Morphometric Analysis (GMA) to explore this topic on the cleaver on flake assemblage from Casal do Azemel (Leiria, Portugal), a paradigmatic Iberian Acheulean site with a large number of cleavers on flake (more than 100 pieces), one of the largest collections of this type of tools in Western Europe.</p> <p>The results obtained note a strong degree of shape homogeneity and suggest that the different technological solutions underlying the definition of the distal cutting edge, or the intensity of secondary reshaping, do not produce major differences in the overall shape of the tool. Therefore, this homogeneity is a consequence of the existence of a specific pattern of support selection and/or is the outcome of the already predetermined nature of the chosen blank. These observations allowed to discuss the significance of blank predetermination in the Acheulean, highlighting the existence of highly structured technological and cognitive prerequisites.</p>
Figure 17 in The dark side of birds: melanism-facts and fiction
Figure 17. Chukar Alectoris chukar (NHMUK 1939.12.9.3715), same mutation as in Fig. 16. (Harry Taylor, © Natural History Museum, London)
Figure 22 in The dark side of birds: melanism-facts and fiction
Figure 22. Forms of melanism (C and D) in Asian Blue Quail Synoicus chinensis that affect female plumage much more than male plumage (Pieter van den Hooven)
Figure 24 in The dark side of birds: melanism-facts and fiction
Figure 24 (facing page). Pl. 15 in A. B. Meyer (1887) Unser Auer-, Rackel- und Birkwild und seine Abarten showing what Meyer believed to be Black Grouse Lyrurus tetrix × ptarmigan Lagopus sp. hybrids; however, these birds are aberrant-coloured female Black Grouse (cf. Figs. 25–26) (Hein van Grouw, © Natural History Museum, London)
Figure 18 in The dark side of birds: melanism-facts and fiction
Figure 18. Hazel Grouse Tetrastes bonasia (NHMUK 1987.24.122) with a form of melanism that has altered the typical pattern and markings, resulting in a paler appearance (Harry Taylor, © Natural History Museum, London)
Figure 16 in The dark side of birds: melanism-facts and fiction
Figure 16. Specimen of Red-legged Partridge Alectoris rufa with a form of melanism that has changed the usual pattern and markings (left, NHMUK 1923.1.29.1) and a normal-coloured individual (right, NHMUK 1907.12.20.7); due to a mutation, the black head and throat markings are altered and the solid-coloured back, shoulders and wing feathers (based mainly on eumelanin) now exhibit distinctive patterns based on both melanin types, and some parts resemble the flanks plumage (Harry Taylor, © Natural History Museum, London)
Figure 5 in The dark side of birds: melanism-facts and fiction
Figure 5. 'Yellow' is a dominant mutation of the agouti gene in Common Quail Coturnix coturnix (NHMUK 1996.41.441, England) resulting in a coloration based on mainly phaeomelanin (Harry Taylor, © Natural History Museum, London) Figure 6.. Hooded Crow Corvus corone cornix (left, NHMUK 1965.M.19466), Carrion Crow C. c. corone (right, NHMUK 2013.5.13) and their hybrid offspring, with in centre a first-generation hybrid (NHMUK 1965.M.19463), left of it a backcross Hooded Crow (NHMUK 1879.3.7.1) and right a backcross Carrion Crow (NHMUK 1925.2.12.1); both the Hooded Crow and all of the hybrids display many patterned feathers with black centres and grey fringes (Harry Taylor, © Natural History Museum, London)
Figure 15 in The dark side of birds: melanism-facts and fiction
Figure 15. Feral Pigeons Columba livia with phaeomelanised plumage (ash-red) but different wing patterns: (A) Barred (the white primaries and head feathers are the product of leucism), (B) Chequer and (C) T-pattern chequer, Leiden, the Netherlands, 19 August 2007 (Hein van Grouw)
Figure 20 in The dark side of birds: melanism-facts and fiction
Figure 20. Normal-coloured and black-shouldered female Indian Peafowls Pavo cristatus; the 'black-shoulder' mutation has a major effect on female plumage (Hein van Grouw)
Figure 3 in The dark side of birds: melanism-facts and fiction
Figure 3. Common Woodpigeon Columba palumbus specimens (NHMUK 1930.8.14.1, NHMUK 2000.11.1, NHMUK 1923.3.8.1); the phaeomelanised plumage is probably the product of a similar mutation as 'recessive red' in Feral Pigeon C. livia (cf. Fig. 2D) (Harry Taylor, © Natural History Museum, London) Figure 4. 'Perdix montana', the phaeomelanistic variety of Grey Partridge P. perdix, NHMUK 1939.12.9.3717, Nasavad, Slovakia, 10 November 1932 (at left) and NHMUK 1939.12.9.3716, Norfolk, England, October 1911; the gene for this recessive mutation is present throughout the range of the species and is therefore a frequently recurring variety on the borderline of being recognised as a morph (Harry Taylor, © Natural History Museum, London)
Figure 14A in The dark side of birds: melanism-facts and fiction
Figure 14A. Melanistic Tawny Owl Strix aluco, Switzerland, April 2015 (Bertrand Ducret); (B) melanistic morph of Ural Owl Strix uralensis, southern Poland, May 2008 (Chris van Rijswijk)
Figure 10. Melanistic White Wagtail Motacilla a in The dark side of birds: melanism-facts and fiction
Figure 10. Melanistic White Wagtail Motacilla a. alba, Ardivachar, South Uist, Outer Hebrides, Scotland, September 2015; the normal black head markings have overrun their boundaries, whilst the rest of the plumage is darker as well (John Kemp) Figure 11. Melanistic Great Tit Parus major, Rotterdam, the Netherlands, November 2008; the normal black head and breast markings have overrun their boundaries, whilst the rest of the plumage is darker as well (Harvey van Diek)
Figure 19 in The dark side of birds: melanism-facts and fiction
Figure 19. Normal-coloured and black-shouldered male Indian Peafowls Pavo cristatus; apart from parts of the wing and shoulders, this mutation does not affect the rest of male plumage (Hein van Grouw)
Figure 21. A in The dark side of birds: melanism-facts and fiction
Figure 21. A form of melanism in Mallard Anas platyrhynchos that affects female plumage much more than male plumage (Hein van Grouw)
Figure 9 in The dark side of birds: melanism-facts and fiction
Figure 9. Common Chaffinch Fringilla coelebs, Utö Island, Finland, 13 April 2013, with increased phaeomelanin resulting in predominantly reddish-brown plumage (Jorma Tenovuo)
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.