Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

32

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

32 results for “documentary”

Learn how ShareScore rates datasets ↗
zenodo48/100

Documentary sources of case studies on the issues a data protection officer faces on a daily basis

<p>The dataset contains the text of the documents that are sources of evidence used in [1] and [2] to distill our reference scenarios according to the methodology suggested by Yin in [3].</p> <p>The dataset is composed of 95 unique document texts spanning the period 2005-2022. This dataset makes available a corpus of documentary sources useful for outlining case studies related to scenarios in which the DPO finds himself operating in the performance of his daily activities.</p> <p>The language used in the corpus is mainly Italian, but some documents are in English and French. For the reader&#39;s benefit, we provide an English translation of the title of each document.</p> <p>The documentary sources are of many types (for example, court decisions, supervisory authorities&#39; decisions, job advertisements, and newspaper articles), provided by different bodies (such as supervisor authorities,&nbsp; data controllers, European Union institutions, private companies, courts, public authorities, research organizations, newspapers, and public administrations),&nbsp; and redacted from distinct professional roles (for example, data protection officers, general managers, university rectors, collegiate bodies, judges, and journalists).</p> <p>The documentary sources were collected from 31 different bodies. Most of the documents in the corpus (a total of 83 documents) have been transformed into Rich Text Format (RTF), while the other documents (a total of 12) are in PDF format. All the documents have been manually read and verified.<br> The dataset is helpful as a starting point for a case studies analysis on the daily issues a data protection officer face. Details on the methodology can be found in the accompanying papers.</p> <p>The available files are as follows:</p> <ul> <li><strong>documents-texts.zip</strong>&nbsp;--&gt;&nbsp;contain a directory of .rtf files (in some cases .pdf files) with the text of documents used as sources for the case studies. Each file has been renamed with its SHA1 hash so that it can be easily recognized.</li> <li><strong>documents-metadata.csv</strong>&nbsp;--&gt;&nbsp;Contains a CSV file&nbsp;with the metadata&nbsp;for each document used as a source for the case studies.</li> </ul> <p>This dataset is the original one used in the publication [1] and the preprint containing the additional material [2].</p> <p>[1] F. Ciclosi and F. Massacci, &quot;The Data Protection Officer: A Ubiquitous Role That No One Really Knows&quot; in IEEE Security &amp; Privacy, vol. 21, no. 01, pp. 66-77, 2023, doi: 10.1109/MSEC.2022.3222115, url: https://doi.ieeecomputersociety.org/10.1109/MSEC.2022.3222115.</p> <p>[2] F. Ciclosi and F. Massacci, &quot;The Data Protection Officer, an ubiquitous role nobody really knows.&quot; arXiv preprint arXiv:2212.07712, 2022.</p> <p>[3] R. K. Yin, Case study research and applications. Sage, 2018.</p>

opencc-by-4.0Apr 2023View details →
zenodo44/100

Documentary practices and uses in rural context of north-western Iberia

<p>Abstract, slides, datasets, and draft transcript of the presentation given at the XXIIe Colloque de pal&eacute;ographie latine, Prague (14-16 September).</p> <p><strong>Brief abstract: </strong></p> <p>Our proposal intends to examine the documentary practices and uses found in the context of rural, secular communities in the north-western Iberian Peninsula from the late eleventh to the thirteen centuries. These practices will be considered in contrast with those from urban, mostly ecclesiastical, communities. We will assess the form and function of secular diplomas made by rural scribes, before considering the graphic systems and styles used in communities of professional scribes with different levels of literacy. Finally, the mobility of practices and their authors will be evaluated, taking the parish church and the rural monastery as meeting points.</p> <p>Ultimately, this will provide a clearer picture of the connections between rural and urban scribes and their communities, mutual exchanges, reception of new palaeographical and codicological developments and cultural interaction. This innovative assessment of the evolution of documentary practices in north-western Iberia will lead to a better understanding of writing practices across all levels of society.</p> <p><strong>Proposal:</strong></p> <p>Within the context of the transition from Visigothic to Caroline minuscule from the end of the eleventh century until well into the twelfth century in north-western Iberia, and alongside the progressive increase in the use of Romance that will result in the emergence of vernacular languages by the thirteenth century, our proposal intends to address documentary practices and uses in the context of secular rural communities in contrast with those communities culturally located at the centre of change and which are mostly ecclesiastical. With that end in mind, this paper combines three of the perspectives proposed for this congress. We will depart from the analysis of the form and function of secular diplomas made by scribes from rural contexts, before considering the graphic systems and styles used in communities of professional scribes with different levels of literacy. Finally, we will assess the mobility of practices and their authors, taking the parish church and the rural monastery as meeting points.</p> <p>Form and function: We will examine the interaction between the visual and material aspects typical of framed charters, their production and content, in the context of rural secular communities (what distinguishes the secular documentary corpus in terms of its materiality and use?) and in opposition to the charters associated with central ecclesiastical communities. A comparison will be made between the processes of documentary creation in both environments, their raison d&rsquo;&ecirc;tre and usefulness, and the effect on the resulting product as a material object to be read but also seen.</p> <p>The scripts and styles of writing communities: We will study the graphic typologies and documentary practices of professional scribes working for central ecclesiastical institutions (cathedrals and monasteries) in contrast to those in centres of a lower cultural level (rural parishes and minor monasteries), examining their possible interaction with regards to the process of absorption and internalization of new palaeographic and diplomatic developments. In this sense, we propose to delve into the life and development of graphic activity of ecclesiastical and lay agents with limited literacy and to evaluate their role in the configuration of identities.</p> <p>Individual mobility and encounters: The scope of action of the rural scribes will be evaluated alongside their relationship with the secular communities they served. Similarly, we will consider their role as integrating agents as well as their possible itinerancy within (certainly) parochial and, possibly, monastic contexts.</p> <p>As a result of the study of these three perspectives, we will provide a clearer picture of the encounters in the written culture between rural and central scribes and their respective communities, pondering mutual exchanges, reception of new developments and how cultural interaction could affect the evolution of documentary practices in north-western Iberia. This will bring about a better understanding of the practice of writing across all levels of society.</p>

opencc-by-4.0Aug 2023View details →
zenodo40/100

Freedom from infection - Documentary

<p>Short film documentary about Freedom from Infection project: The Peruvian Amazon is characterized by low&nbsp;access to health services due to the fact that most of the communities are in remote locations, and their only means of transportation are boats. We have identified multiple challenges in our current reporting system to detect Malaria cases occurring in Amazon based communities. &quot;Freedom from Infection&quot; seeks to provide decision makers with a tool to calculate the probability that transmission has been interrupted in these communities in order to have a Malaria-free Amazon.&nbsp;</p> <p>Freedom from Infection is a collaboration between Universidad Peruana Cayetano Heredia (UPCH) and the London School of Hygiene &amp; Tropical Medicine (LSHTM), financed by the Bill and Melinda Gates Foundation (BMGF).&nbsp;</p>

opencc-by-4.0Feb 2022View details →
zenodo40/100

dataset of documentaries on genocide and atrocity

<p>This is a tab-separated file listing documentaries on genocides. The main genocides focused on are: The Herero and Nama genocide (1904-08), the Armenian genocide, the Holocaust, the Indonesian genocide (1965-66), the Cambodian genocide, the Rwandan genocide, the Bosnian genocide. Other atrocities are also recorded but more sporadically</p> <p>Variables/columns and [dtype]:</p> <p>'FILM' [string], 'DATE' [integer], 'DIRECTOR' [string], 'ATROCITY, GENOCIDE' [string, one of 'Herero &amp; Nama', 'Armenian g.', 'Holodomor', 'Holocaust', 'Indonesian g.', 'Cambodian g.', 'Rwandan g.', 'Bosnian g.'; documentaries on further atrocities are also recorded in my dataset but not pursued as systematically], 'DATA STATE' [float, internally used variable], 'LANGUAGE' [string], 'DURATION' [integer], 'PRODUCER' [string], 'COUNTRY' [string], 'ACQUIRED' [string], 'LINKS' [string], 'COMMENTS' [string, plot synopsis and interviewed persons may be recorded here as well as salient features of sound and cinematography], 'PERP REPR' [string, one of &lsquo;direct&rsquo;, &lsquo;archival&rsquo;, &lsquo;reenactment&rsquo;, &lsquo;animation&rsquo;], 'PERP GROUPS' [string], 'PERP GENDER' [string, one of &lsquo;M&rsquo;, &lsquo;F&rsquo;, &lsquo;F, M&rsquo;], 'VIOLENCE RATIONALE/CAUSES' [string], 'COLLABORATOR GROUPS' [string], 'VICTIM REPR' [string], 'VICTIM GROUPS' [string], 'VICTIM GENDER INTERVIEWS' [string, one of &lsquo;M&rsquo;, &lsquo;F&rsquo;, &lsquo;F, M&rsquo;], 'SEXUAL VIOLENCE' [boolean], 'DOC TYPE' [string], 'RATINGS' [float, if available], 'AWARD' [string]</p> <p>&nbsp;</p> <p>A related and continuously updated website that allows for visualizing this data (as well as scraped Yad Vashem and Cinematography of the Holocaust filmographic data) is available here: https://jackewiebohne.shinyapps.io/shiny/</p>

opencc-by-4.0Jun 2024View details →
zenodo40/100

Full interview of Christine Chinkin for the documentary "The legacy of the Tokyo Women's Tribunal"

<p>Full interview of Christine Chinkin for the documentary &quot;The legacy of the Tokyo Women&#39;s Tribunal&quot;</p> <p>https://www.lse.ac.uk/women-peace-security/research/Gendered-Peace-the-legacy-of-the-Tokyo-Womens-Tribunal&nbsp;</p> <p>Video editing: Sophie Mallett</p> <p>Captions: Marie Toseland</p> <p>&nbsp;</p>

opencc-by-4.0Dec 2022View details →
zenodo40/100

Dataset and evaluation for HTR models for Latin and French Medieval Documentary Manuscripts

<p><strong>1. Dataset presentation.</strong></p> <p>This is the dataset used to produce the HTR models applied to documentary Latin and French manuscripts presented in the paper: Sergio Torres Aguilar, Vincent Jolivet. <strong>Handwritten Text Recognition for Documentary Medieval<br> Manuscripts. </strong>2022. https://hal.science/hal-03892163</p> <p>The dataset contains mostly charters and registers from the Late-medieval period (12th-15th). The training and evaluation, entailing 1855 pages, 120k lines of text and almost 1M tokens, were conducted using three freely available ground-truth corpora :</p> <p><strong>The Alcar-HOME database </strong>: https://zenodo.org/record/5600884</p> <p><strong>The e-NDP corpus </strong>: https://github.com/chartes/e-NDP_HTR</p> <p><strong>The Himanis project </strong>: https://zenodo.org/record/5535306</p> <p>The final model operates in a multilingual environment (Latin and French) and it is able to recognize several Latin script families (mostly <em>Textualis</em> and <em>Cursiva</em>) in documents produced in ca. 12th - 15th centuries. During the evaluation the models shows an accuracy of <strong>94.01%</strong> on the validation set and a CER (character error ratio) of about <strong>0.12</strong> to <strong>0.17</strong> on four external unseen datasets. A fine-tuning exercise using 10 ground-truth pages can raise these results to a CER between <strong>0.06</strong> to <strong>0.10</strong> respectively.</p> <p>&nbsp;</p> <p><strong>2. Dataset contents .</strong></p> <p>a) <em>GT_list : </em>List containing the GT file names which constitute the training, evaluation and test sets. The images and transcriptions can be downloaded from their original repositories.</p> <p>b) <em>Training :</em> Contains the training and testing results (evaluation and prediction files) presented in the original paper for the two training phases: Regular (Textualis and Cursiva separated training) and Quartiles (mixed training by quartiles).</p> <p>c) <em>Useful_scripts :</em> Scripts to produce the HTR metrics (CER, WER, SER) and plot the model&#39;s accuracy.</p> <p>d) <em>Best_model :</em> Contains the best multilingual and multi-script model.</p>

opencc-by-4.0Jan 2023View details →
zenodo40/100

Interwoven Sound Spaces Documentary

<p>This is a documentary shot and edited by Tim Nowitzki about the Interwoven Sound Spaces project (2022).</p> <p>Interwoven Sound Spaces is an interdisciplinary project which brought bringing together telematic music performance, interactive textiles, interaction design, and artistic research. A team of researchers collaborated with two professional contemporary music ensembles based in Berlin, Germany, and Pite&aring;, Sweden, and four composers, with the aim of creating a telematic distributed concert taking place simultaneously in two concert halls and online. Central to the project was the development of interactive textiles capable of sensing the musicians&rsquo; movements while playing acoustic instruments, and generating data the composers used in their works. Musicians, instruments, textiles, sounds, halls, and data formed a network of entities and agencies that was is reconfigured for each piece, showing how networked music practice enables distinctive musicking techniques.&nbsp;</p>

opencc-by-4.0Sep 2023View details →
zenodo36/100

Independent Documentary Filmmakers from China, Hong Kong, and Taiwan Web Archive collection derivatives

<p>Web archive derivatives of the <a href="https://archive-it.org/collections/12172">Literary Authors from Europe and Eurasia Web Archive</a> collection from the <a href="https://archive-it.org/home/IvyPlus">Ivy Plus Libraries Confederation</a>. The derivatives were created with the <a href="https://github.com/archivesunleashed/aut/">Archives Unleashed Toolkit</a> and <a href="https://cloud.archivesunleashed.org/">Archives Unleashed Cloud</a>.</p> <p>The <strong>ivy-12126-parquet.tar.gz</strong> derivatives&nbsp;are&nbsp;in&nbsp;the <a href="https://parquet.apache.org/">Apache&nbsp;Parquet format</a>,&nbsp;which&nbsp;is&nbsp;a <a href="http://en.wikipedia.org/wiki/Column-oriented_DBMS">columnar&nbsp;storage</a> format. These derivatives are generally small enough to work with on your local machine, and can be easily converted to Pandas DataFrames. See <a href="https://github.com/archivesunleashed/notebooks/blob/master/datathon-nyc/parquet_pandas_stonewall.ipynb">this</a> notebook for examples.</p> <p><strong>Domains</strong></p> <pre><code class="language-java">.webpages().groupBy(ExtractDomainDF($"url").alias("url")).count().sort($"count".desc)</code></pre> <p>Produces&nbsp;a&nbsp;DataFrame&nbsp;with&nbsp;the&nbsp;following&nbsp;columns:</p> <ul> <li>domain</li> <li>count</li> </ul> <p><strong>Web&nbsp;Pages</strong></p> <pre><code class="language-java">.webpages().select($"crawl_date", $"url", $"mime_type_web_server", $"mime_type_tika", RemoveHTMLDF(RemoveHTTPHeaderDF(($"content"))).alias("content"))</code></pre> <p>Produces&nbsp;a&nbsp;DataFrame&nbsp;with&nbsp;the&nbsp;following&nbsp;columns:</p> <ul> <li>crawl_date</li> <li>url</li> <li>mime_type_web_server</li> <li>mime_type_tika</li> <li>content</li> </ul> <p><strong>Web&nbsp;Graph</strong></p> <pre><code class="language-java">.webgraph()</code></pre> <p>Produces&nbsp;a&nbsp;DataFrame&nbsp;with&nbsp;the&nbsp;following&nbsp;columns:</p> <ul> <li>crawl_date</li> <li>src</li> <li>dest</li> <li>anchor</li> </ul> <p><strong>Image&nbsp;Links</strong></p> <pre><code class="language-java">.imageLinks()</code></pre> <p>Produces&nbsp;a&nbsp;DataFrame&nbsp;with&nbsp;the&nbsp;following&nbsp;columns:</p> <ul> <li>src</li> <li>image_url</li> </ul> <p><a href="https://github.com/archivesunleashed/aut-docs/blob/master/current/binary-analysis.md#binary-analysis"><strong>Binary&nbsp;Analysis</strong></a></p> <ul> <li>Audio</li> <li>Images</li> <li>PDFs</li> <li>Presentation&nbsp;program&nbsp;files</li> <li>Spreadsheets</li> <li>Text&nbsp;files</li> <li>Word&nbsp;processor&nbsp;files<br> &nbsp;</li> </ul> <p>The <strong>ivy-12126-auk.tar.gz </strong>derivatives<strong> </strong>are the <a href="https://cloud.archivesunleashed.org/derivatives">standard set of web archive derivatives</a> produced by the Archives Unleashed Cloud.</p> <ul> <li><strong>Gephi </strong>file, which can be loaded into <a href="https://gephi.org/">Gephi</a>. It will have basic characteristics already computed and a basic layout.</li> <li><strong>Raw Network</strong> file, which can also be loaded into <a href="https://gephi.org/">Gephi</a>. You will have to use that network program to lay it out yourself.</li> <li><strong>Full text</strong> file. In it, each website within the web archive collection will have its full text presented on one line, along with information around when it was crawled, the name of the domain, and the full URL of the content.</li> <li><strong>Domains count</strong> file. A text file containing the frequency count of domains captured within your web archive.</li> </ul>

opencc-by-4.0Jan 2020View details →
zenodo36/100

WasteScapes: Augmented Documentary as Critical Place-Based Learning

<p>WasteScapes: Augmented Documentary as Critical Place-Based Learning by Liz Miller</p><p>24th April 2023 held in the context of the virtual lecture series, hosted by the University of Bayreuth as a part of the research project Digital Documentary Practices</p><p>Wastescapes is an augmented documentary project, an App that engages creatively and critically with places, stories, and concepts related to waste in the city of Montréal, Canada. The WasteScapes App guides individuals to sites with entangled layers of history such as former dumps, underground waste infrastructures, or repurposed waste lands and challenges users to consider how our patterns of consumption have impacted land, water and other species over time. In WasteScapes being in place matters: short audio narrations and photos are unlocked only when someone visits a site. The project permitted Miller and her collaborators to think beyond fixed representation towards sensory awareness, embodied pedagogies, and expanded documentary practices. In this lecture, Miller will discuss WasteScapes and the potential of augmented documentary to shift a user's focus from character to place, from individual to ecosystem, and in doing so address the more than human as well as the power dynamics of place. Elizabeth (Liz) Miller, MFA, is a filmmaker and a Full Professor in Communication Studies at Concordia University. Her multi-platform collaborative documentary projects on timely issues such as water privatization, wetlands, environmental justice, and climate change have won awards and influenced decision makers. She is the co-author of Going Public: The Art of Participatory Practice (2017) and has written book chapters and articles on co-creation, environmental media, and place-based pedagogies. Her most recent project, WasteScapes, incorporates augmented documentary, cycling tours, installations, and educational resources.&nbsp;</p><p>Elizabeth Miller elizabeth.miller@concordia.ca&nbsp;</p><p>http://coms.concordia.ca/faculty/miller.html&nbsp;</p><p>http://redlizardmedia.com&nbsp;</p><p>http://theshorelineproject.org/</p>

opencc-by-4.0Nov 2023View details →
zenodo36/100

The VR Documentary – History, Theory, Examples

<p>While VR has now almost become established on the games market for home use by several large commercial manufacturers of displays and content providers, the use in the educational context is not yet very large. This applies in particular to artistic projects that, like documentary film, want to combine public education and creative expression. But especially at festivals that take place regularly (such as the DOK Neuland section of the Leipzig Documentary Film Festival), there are spaces to present innovative non-fictional projects – and the demand is not small, there are usually even waiting lists for use. The lecture would like to approach this new form between creative practice, world building and interactive user integration in three parts: First, it will deal with the (recent) history of VR and ask about the essential factors in the implementation of current forms. Second, methods and theories are presented that emphasize the potential for innovation – also compared to linear documentary film or other digital forms like web docs – and I will outline the special nature of the form between experience, storytelling and participation. Third and lastly, of course, some examples should also be presented briefly, which have received a lot of encouragement from users and practitioners in recent years. These are selected in such a way that they describe as many different modalities of non-fictional VR as possible. Florian Mundhenke, PD Dr. phil. habil., since 2020 Associate Professor for Cultural and Media Studies (DAAD) at the Institute for Modern Languages and Cultural Studies at University of Alberta in Edmonton/Canada. From 2018-2020 Temporary Professor (W3) for Media Studies and Media Culture at University of Leipzig. Before he was Senior Lecturer at the Institute for Media and Communication (IMK) at University of Hamburg, Associate Professor for Media Hybridity ("Juniorprofessor für Mediale Hybride") at the University of Leipzig, and Research Assistant at University of Marburg. He was speaker of the DFG-funded research network "Cinema as an experience space": www.erfahrungsraum-kino.de. PhD dissertation on the phenomenon of chance in film in 2008 (Marburg: Schueren). Habilitation (Lecture qualification thesis) on hybrid forms between documentary and fictional film in 2016 (Wiesbaden: Springer VS). Fields of research include intersections of arts and media in history and practice (intermedial, transcultural), non-fictional media (documentary film, VR films, i-docs, AR), methods and concepts of Digital Humanities in media studies, theories of media genres and genre development, intersectional media theory (gender, race, class, ideology), and world cinema with a focus on the Far East (Japan, Korea, Taiwan).&nbsp;</p><p>Florian Mundhenke https://apps.ualberta.ca/directory/person/mundhenk&nbsp;</p><p>Email: mundhenk@ualberta.ca</p>

opencc-by-4.0Dec 2023View details →
zenodo36/100

Managing competition between legacy television services and video streaming platforms in Hungary in the early 2020s – A case study [Secondary documentary sources]

<p>Secondary documentary sources used in the paper entitled &quot;Managing competition between legacy television services and video streaming platforms in Hungary in the early 2020s &ndash; A case study&quot;</p>

opencc-by-4.0Feb 2024View details →
zenodo36/100

TRIDIS: HTR model for Multilingual Medieval and Early Modern Documentary Manuscripts (11th-16th)

<p><strong>TRIDIS (Tria Digita Scribunt)</strong> is a Handwriting Text Recognition model trained on semi-diplomatic transcriptions from medieval and Early Modern Manuscripts. It is suitable for work on documentary manuscripts, that is, manuscripts arising from legal, administrative, and memorial practices more commonly from the Late Middle Ages (13th century and onwards). It can also show good performance on documents from other domains, such as literature books, scholarly treatises and cartularies providing a versatile tool for historians and philologists in transforming and analyzing historical texts.</p> <p>A paper presenting the first version of the model is available here: Sergio Torres Aguilar, Vincent Jolivet. <strong>Handwritten Text Recognition for Documentary Medieval Manuscripts. </strong>Journal of Data Mining and Digital Humanities.<strong> </strong>2023. https://hal.science/hal-03892163</p> <p>&nbsp;</p> <h3>Transcriptions rules :</h3> <p>Since the majority of the training documents come from diplomatic editions, the transcriptions were <strong>normalized</strong> to contemporary reading standards, and <strong>abbreviations were expanded</strong> with the aim of facilitating a more fluid reading of the document.</p> <p>The following rules were applied:</p> <ul> <li>The abbreviations have been expanded, both those by suspension (<code>facimꝰ</code> ---&gt; <code>facimus</code>) and by contraction (<code>d&ntilde;i</code> --&gt; <code>domini</code>). Likewise, those using conventional signs (<code>⁊</code> --&gt; <code>et</code> ; <code>ꝓ</code> --&gt; <code>pro</code>) have been resolved.&nbsp;</li> <li>The named entities (names of persons, places and institutions) have been <code>capitalized</code>. The beginning of a block of text as well as the original capitals used by the scribe are also capitalized.</li> <li>The consonantal <code>i</code> and <code>u</code> characters have been transcribed as <code>j</code> and <code>v</code> in both French and Latin.</li> <li>The punctuation marks used in the manuscript like:&nbsp;<code>.</code> or <code>/</code> or <code>|</code> have not been systematically transcribed as the transcription has been standardized with modern punctuation.</li> <li>Corrections and words that appear cancelled in the manuscript have been transcribed surrounded by the sign <code>$</code> at the beginning and at the end.</li> </ul> <p>&nbsp;</p> <h3>Versions :</h3> <p><strong>Version 1 </strong>of the model was trained on charters and registers dataset from the Late Medieval period (12th-15th centuries). The training and evaluation involved 1855 pages, 120k lines of text, and almost 1M tokens, conducted using three freely available ground-truth corpora:</p> <ul> <li>The Alcar-HOME database: <a href="../record/5600884" target="_new">https://zenodo.org/record/5600884</a></li> <li>The e-NDP corpus: <a href="../record/7575693" target="_new">https://zenodo.org/record/7575693</a></li> <li>The Himanis project: <a href="../record/5535306" target="_new">https://zenodo.org/record/5535306</a></li> </ul> <p><strong>Version 2</strong> of the model has added new datasets from feudal books and legal proceedings (14th-16th centuries), incorporating an additional 115k lines and more than 1.2M tokens to the previous version using other corpora like:</p> <ul> <li>K&ouml;nigsfelden Abbey corpus: <a href="../record/5179361" target="_new">https://zenodo.org/record/5179361</a></li> <li>Monumenta Luxemburgensia.</li> </ul> <p>&nbsp;</p> <h3>Accuracy</h3> <p>TRIDIS was trained using a CNN+RNN+CTC architecture within the Kraken suite (https://kraken.re/). This final model operates in a multilingual environment (Latin, Old French, and Old Spanish) and is capable of recognizing several Latin script families (mostly Textualis and Cursiva) in documents produced circa 11th - 16th centuries. During evaluation, the model showed an accuracy of 93.1% on the validation set and a CER (Character Error Ratio) of about 0.11 to 0.15 on four external unseen datasets. Fine-tuning the model with 10 ground-truth pages can improve these results to a CER of between 0.06 to 0.10, respectively.</p> <h3>Other formats</h3> <p>The ground truth used for version 2 was also employed to train a Transformer HTR model that combines TrOCR as the encoder with a RoBERTa medieval model as the decoder. This model exhibits a slighly better performance in terms of CER metrics to the current TRIDIS version and shows an improved WER by about 25%. The model is available on the Hugging Face Hub: <a href="https://huggingface.co/magistermilitum/tridis_HTR">magistermilitum/tridis_HTR</a></p>

opencc-by-4.0Mar 2024View details →
zenodo36/100

Building a Ngalawa Double Outrigger Logboat in Bagamoyo, Tanzania: A Craftsman at his Work. 3D Model and Documentary Film Files.

<p>The 3D model and&nbsp;documentary film detailing&nbsp;the building of the&nbsp;<em>Bahari Yetu, Urithi Wetu</em>&nbsp;<em>ngalawa</em>&nbsp;accompany&nbsp;an article on building a Ngalawa, a double outrigger logboat. The&nbsp;article documents master logboat-builder Alalae Mohamed&rsquo;s construction of a&nbsp;<em>ngalawa</em>&nbsp;fishing vessel in Bagamoyo, Tanzania, in 2019. The&nbsp;<em>ngalawa</em>&nbsp;is an extended logboat with double outrigger and lateen sail: used by low-income, artisanal fishers. It is the most common marine vessel type of the East African coast. This article follows the construction process from Alalae&rsquo;s selection and the felling of the tree(s) to the launching of the vessel. It outlines the tools and materials used, details the sequence he followed, and presents his choices and considerations made along the way. It is accompanied by a documentary film recording the construction process, a 3D digital model of the vessel and detailed construction drawings.&nbsp;</p>

opencc-by-4.0Feb 2022View details →
zenodo36/100

Hong Kong Marketplace Film Documentary

<p>Hong Kong Marketplace Documentary&nbsp;https://www.youtube.com/watch?v=dIiwJX4tgmY</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-2.0Aug 2022View details →
zenodo36/100

CATCH-EyoU: Meanings and Practices of Youth Participation and Cases of Successful Participation: Cross-national Documentary Evidence of Youth Civic and Political Participation Initiatives

<p>The data set integrates documentary evidence of national youth civic and political participation initiatives in Italy, Sweden, Germany, Greece, Portugal, Czech Republic, UK, and Estonia. The evidence are represented as pamphlets, brochures, leaflets, newsletters, social media and websites, you tube videos, meme and vine campaigns. The participation initiatives are selected with a focus on initiatives used by national youth citizenship organizations for generating, supporting, communicating and engaging young people&rsquo;s active participation in diverse and varied causes (the environment, voting at 16, anti-fees and cuts campaigns, rights around civic spaces and youth centers, employment and jobs-related campaigns, issues around race and religion, volunteering, refugees, local housing and education).</p> <p>&nbsp;</p> <p>The data set consists of: (1) a spreadsheet integrating data from all partner countries, and containing qualitative descriptions of activities and cataloguing associated materials of youth civic and political organizations; (2) a textual report containing exemplar images of publicly available online newsletters, minutes, reports and campaign materials from across the consortium (where feasible), copyright cleared and creative commons screenshots of websites, invitation letters, records of personal conversations with members of youth civic teams, quotes from researcher-conversations on social media or f-2-f with members of teams with names redacted.</p> <p>Personal data and all non-public data has been anonymized in order to protect participants&rsquo; privacy.</p> <p>The data can be re-used by researchers who want to compare our cross- European data with similar data collected in different countries, to perform textual analysis (content analysis and/or data mining) on our data. Also other stakeholders may be interested in reanalyzing our data for comparative aims.</p>

opencc-by-nc-4.0Dec 2017View details →
zenodo36/100

Data from : Geochemical and Documentary Topography of a Medieval Silver Valley

<p>These data are the source of an interdisciplinary investigation (archaeology, geochemistry, history) of a medieval silver and lead production site located in southern France, in the Minier valley (Occitanie, Aveyron, Le-Viala-du-Tarn). In order to identify the production sites, in situ geochemical surveys were carried out using a portable X-ray fluorescence spectrometer and differential GPS, guided on the analysis of medieval archival sources. The cartographic representation of the metal concentrations in the surface horizons shows significant enrichment of zinc and lead in the vicinity of the mines. This first type of enrichment makes it possible to highlight the activities of separation of sphalerite and silver-bearing galena. The galena thus isolated on the hillsides is then transported to the vicinity of watercourses, where it is crushed, washed, and smelted. These secondary activities result in a last type of enrichment in which only lead is found in large quantities. The cross-referencing of the information made it possible to overcome the challenges related to the location of the mineral processing workshops, which were often invisible on the surface. The medieval workshops have been located and a function suggested, outlining the first trends in the spatial and social division of labour and providing a solid corpus for future archaeological excavations. Finally, this study highlights the persistence of significant metal contamination in the soils of a rural valley and encourages the consideration of former mining areas when examining the environmental impact of metal production.</p>

opencc-by-sa-4.0Oct 2024View details →
dryad36/100

Wildlife documentaries present a diverse, but biased, portrayal of the natural world

<p>1. Wildlife-documentary production has expanded over recent decades, while studies report reduced direct contact with nature. The role of documentaries and other electronic content in educating people about biodiversity is therefore likely to be growing increasingly important. This study investigated whether the content of wildlife documentaries is an accurate reflection of the natural world and whether conservation messaging in documentaries has changed over time.</p> <p>2. We sampled an online film database (n = 105) to quantify the representation of taxa and habitats over time, and compared this with actual taxonomic diversity in the natural world. We assessed whether the precision with which an organism could be identified from the way it was mentioned varied between taxa or across time, and whether mentions of conservation and anthropogenic impacts on the natural world changed over time.</p> <p>3. Mentions of organisms (n = 374) were very biased towards vertebrates (81.1% of mentions) relative to invertebrates (17.9% of mentions), despite vertebrates representing only 3.4% of described species, compared to 74.9% for invertebrates. Mentions were highly variable across groups and between time periods, particularly for insects, fish and reptiles. Plants had a consistently low representation across time periods.</p> <p>4. A range of habitats was represented, the most common being tropical forest and the least common being deep ocean, but there was no change over time.</p> <p>5. Mentions identifiable to species were significantly different between taxa, with 41.8% of mentions of vertebrates identifiable to species compared with just 7.5% of invertebrate mentions and 10% of plant mentions. This did not change over time.</p> <p>6. Conservation was mentioned in 16.2% of documentaries overall but in almost 50% of documentaries in the current decade. Anthropogenic impacts were mentioned in 22.1% of documentaries and never before the 1970s.</p> <p>7. Our results show that documentaries provide a diverse picture of nature with an increasing focus on conservation, with likely benefits for public awareness. However, they overrepresent vertebrate species, potentially directing public attention towards these taxa. We suggest widening the range of taxa featured to redress this and call for a greater focus on threats to biodiversity to improve public awareness.</p>

opencc-zeroDec 2022View details →
zenodo36/100

Entangled Ecologies. Interactive Documentary between Living Archive, Responsible Witnessing and Relational Co-Creation

<p>The videoessay is also accessible:&nbsp;<a href="https://trametrami.avinus.org/publikationen/3-2023">https://trametrami.avinus.org/publikationen/3-2023</a></p> <p>Recent re-conceptualizations of participation as well as theories of digital transformation culture have shifted the notion of the archive seeing not so much as a static, institutional body but rather as a dynamic, living epistemic environments. Building on these approaches, this presentation discusses emerging phenomena in interactive documentary focusing on ecological emergency &ndash; seeing this crisis itself as a complex ecology of issues where images and conceptualization past, present and future meet. Taking paradigmatic projects addressing the issue of climate change &ndash; <em>The Shore Line</em> (2017) and <em>Climate Witness Project </em>(2019) &ndash; I suggest tentative answers to the question in which way i-docs can contribute to tackle complex change. Are there images from the past which help us to better cope with the present and to envision a more sustainable future? How can one raise awareness of immanent ecological risks when menace is almost invisible, unprecedented and un-imaginable? Is there a way to negotiate multifaceted entanglements through documentaries which are neither paralyzing ecodystopian narratives nor naive ecotopian success-stories? The hypothesis underlying my approach is that one possible solution resides in documentaries which are built on principles of pluri-perspectivity and polyphony to transcend dualisms and to build bridges leading from history to the forthcoming. Discourses from various traditions are brought into dialogue: theories of documentary film meet ecocriticism; network theory encounters reflections on the epistemic dimension of non-fiction; and concepts of co-creation and intervention are related to processes of witnessing, doing documentary and responsible action-taking.</p> <p>&nbsp;</p> <p>more information: <a href="https://did.avinus.org/">https://did.avinus.org/</a>&nbsp;and&nbsp;<a href="https://trametrami.avinus.org/publikationen">https://trametrami.avinus.org/publikationen&nbsp;</a></p>

opencc-by-4.0Jun 2023View details →
zenodo36/100

As We May Remember. The Future of Remembrance from the Perspective of Documentary Archives

<p>Can we already discern the structures of future memory cultures? From this fundamental question, I will examine the mediated forms of current memory culture, with a particular focus on documentary films. Digital technologies are currently leading to significant changes in our knowledge culture, which will inevitably impact the shaping of memory cultures and the structures of media memory in the near future. These changes manifest in two opposing processes: on the one hand, in the archival situation, characterized by non-accessibility, poor archiving, or even the physical decay of documentary material (such as analog video), and on the other hand, attempts at digital preservation or even reconstitution of archives through various media transformation processes - particularly re-mediatizations, which are shaped by new forms of media expression (i-docs, VR/AR-technologies etc.). The latter is not only a challenge for media historiography, but opens up new possibilities for the memory work of GLAM and memorial sites. In this contribution, I will explore the tension between disappearing archive material, using the example of the archival situation of German documentary films, and selected new digital forms of re-mediatization, focusing on themes such as the Holocaust and the Nazi era.</p> <p>&nbsp;</p> <p>See also:&nbsp;Weber, Thomas. (2023, June 30). As We May Remember. The Future of Remembrance from the Perspective of Documentary Archives. In TraMeTraMi: Vols. 4-2023.&nbsp;<a href="https://trametrami.avinus.org/publikationen/4-2023">https://trametrami.avinus.org/publikationen/4-2023</a>&nbsp;</p> <p>&nbsp;</p> <p><strong>Thomas</strong> <strong>Weber</strong> is Professor for media studies at the University of Hamburg. He was one of the leaders of the DFG-project &ldquo;History of the german documentary film after 1945&rdquo; and leads several other projects in the field of documentary film (see<a href="http://www.dokartlabor.avinus.de/"> www.dokartlabor.avinus.de</a>)&nbsp;His books include: <em>Webdokumentationen</em> 2021; <em>Medienkulturen des Dokumentarischen</em> 2017 (ed. with Carsten Heinze); <em>Mediale Transformationen des Holocausts</em> 2013 (ed. with Ursula von Keitz); &ldquo;Documentary Film in Media Transformation&rdquo;, InterDisciplines &ndash; Journal of History and Sociology. Vol 4, No 1 (2013). Further information see<a href="http://www.thomas-weber.avinus.de/"> www.thomas-weber.avinus.de</a></p> <p>&nbsp;</p>

opencc-by-4.0Jun 2023View details →
zenodo36/100

Revisions. NS-Regime, WW2 and Holocaust in West-German Documentaries in the Early Federal Republic of Germany

<p>This contribution presents an English summary of some of the important ideas of the project &quot;Revisionen. Nationalsozialismus, Holocaust und Zweiter Weltkrieg im dokumentarischen Film der fr&uuml;hen Bundesrepublik Deutschland&quot; (Revisions. National Socialism, Holocaust and World War II in Documentary Film in the Early Federal Republic of Germany) by the two authors G&ouml;tz Lachwitz and Thomas Weber, which was published as a book. It deals with the hitherto little-noticed examination of the Nazi past, the Holocaust and the Second World War by the West German documentary films made up to 1961, which sought a new social approach to the past and thus sought to contribute to a resolution of the conflicts in their time. The project does not deal solely with well-known films such as <em>Nuit et Brouillard</em> by Alain Resnais or <em>Mein Kampf</em> by Erwin Leiser, but takes a look at a broader selection. Building on in-depth archival work, it also discusses films that have rarely or not at all been taken up in the specialized literature. Even if the discourses discernible in most of the films tie in with familiar interpretations, they still surprise us with a diversity of formats, distribution channels, and performance contexts, for example in documentary television series or political education work, which points to a more differentiated approach to the Nazi past in the early Federal Republic than has generally been assumed to date.</p> <p>Published in TraMeTraMi, 7-2023:&nbsp;https://trametrami.avinus.org/publikationen/7-2023&nbsp;</p> <p>Further Informations: https://produkte.avinus.de/produkt/lachwitz-weber-revisionen</p>

opencc-by-4.0Jul 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record