Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

21,320

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

21,320 results for “Transcription”

Learn how ShareScore rates datasets ↗
zenodo40/100

List of manuscripts containing John Chrysostom's Homilies and the relevant manual transcriptions

<p>This dataset consists of a <strong>list of all manuscripts</strong> (in the form of a .csv file) used as data in experiments with HTR training via Transkribus. The manuscripts are dated between the 10th-14th&nbsp;centuries and transmit John Chrysostom&rsquo;s <em>Homilies on St. Paul&rsquo;s Epistles to Titus</em>. Homilies 1 and 5 were exploited for the training process.&nbsp;In addition, <strong>19 XML source files</strong> are provided in the <strong>TEI standards</strong> format, which contains a sample of the manual transcription used as ground truth data for training HTR models.</p> <p>Specifically, the&nbsp;<strong>sample_dataset_chrysostomus_ad-titum.csv</strong><strong>&nbsp;</strong>file includes the following columns:</p> <ul> <li><strong>Sigla:</strong> a capital letter used in critical editions to refer to a specific manuscript in an abbreviated form.</li> <li><strong>Manuscripts:&nbsp;</strong>the name of each manuscript, containing the library and the catalogue number assigned to it.</li> <li><strong>Folia:&nbsp;</strong>the folia (i.e., pages) of each manuscript used in the experiments. A different sequence of folia from the same manuscript is recorded in a separate&nbsp;row of this file.</li> <li><strong>Ground truth data sample [file_name]:&nbsp;</strong>the file name of the TEI/XML files that&nbsp;correspond&nbsp;to each manuscript.</li> <li><strong>Image files:&nbsp;</strong>most digital reproductions of manuscripts are under some degree of copyright protection. So, instead of the image files, in this column, one can find a link to the relevant library&#39;s digital archive (if applicable).</li> </ul> <p><strong>**IMPORTANT NOTE:</strong> Version 1 contained an erroneous file. Please use only version 1.2.</p>

opencc-by-4.0Feb 2023View details →
zenodo40/100

Interview Transcripts on Ocean Modeling and Biogeochemical Modeling

<p>This data package contains interview transcripts of interviews that we conducted to understand the domain of ocean modeling with a focus on human interaction, processes and tooling to support our effort to create domain-specific languages to support the work of scientists in this domain.</p> <p>This data package contains the anonymized transcripts of interviews from two phases. The <strong>first phase</strong> was dedicated to collect any theme regarding the domain including processes to develop, maintain and use scientific models (in<br> particular in ocean sciences), as well as, human interaction, technical aspects, and other work environment related themes.</p> <p>The <strong>second phase</strong> focus on the development of biogeochemical models. However, we also included questions regarding the domain.</p>

opencc-by-4.0Jul 2023View details →
zenodo40/100

Interviews transcriptions

<p>This is a full set of transcriptions of all the interviews conducted to complete a research dedicated towards evaluating emotive research content, identifying areas of interest that may have not been the core areas in some of the original research papers, and creating a digital archaeo game or the purpose of testing the extracted theory in terms of emotive triggers in nostalgic and negative emotive archaeological environments. The transcriptions have been handwritten and reviewed several times but may contain some grammatical and spelling mistakes due to some issues with unstable connections during the video and audio recording or due to the inability to understand some of the spoken languages at times. Due to the nature of the research the interviewed subjects have&nbsp; agreed in writing to be named and therefore it contains the real names of the people involved in the experiment.</p> <p>Overall this raw data set has been used to further investigate the area of emotive studies but may in the future be useful for many other applications as there were a wide variety of topics being discussed and the open-ended questions structure made this a very fruitful resource.</p>

opencc-by-4.0Jul 2023View details →
zenodo40/100

Dataset supporting the paper "Expanding the coverage of regulons from high-confidence prior knowledge for accurate estimation of transcription factor activities"

<p>Datasets involved in the construction and benchmarking of the CollecTRI-derived regulons&nbsp;as presented in the paper &quot;Expanding the coverage of regulons from high-confidence prior knowledge for accurate estimation of transcription factor activities&quot;.</p>

opencc-by-4.0Mar 2023View details →
zenodo40/100

Analysis Products: Transcription factor stoichiometry, motif affinity and syntax regulate single cell chromatin dynamics during fibroblast reprogramming to pluripotency

<p>This record contains analysis products for the paper &quot;Transcription factor stoichiometry, motif affinity and syntax regulate single cell chromatin dynamics during fibroblast reprogramming to pluripotency&quot; by Nair, Ameen&nbsp;<em>et al</em>.&nbsp;Please refer to the READMEs in the directories, which are summarized below.</p> <p>The record&nbsp;contains the following files:<br> <br> `clusters.tsv`:&nbsp;<strong>&nbsp;</strong>contains the cluster id, name and colour of clusters&nbsp;in the paper</p> <p><strong>scATAC.zip</strong></p> <p>Analysis products for the single-cell ATAC-seq data. Contains:</p> <p>- `cells.tsv`: list of barcodes that pass QC. Columns include:<br> &nbsp;&nbsp; &nbsp;- `barcode`<br> &nbsp;&nbsp; &nbsp;- `sample`: (time point)<br> &nbsp;&nbsp; &nbsp;- `umap1`<br> &nbsp;&nbsp; &nbsp;- `umap2`<br> &nbsp;&nbsp; &nbsp;- `cluster`<br> &nbsp;&nbsp; &nbsp;- `dpt_pseudotime_fibr_root`: pseudotime values treating a fibroblast cell as root<br> &nbsp;&nbsp; &nbsp;- `dpt_pseudotime_xOSK_root`: pseudotime values treating xOSK cell as root<br> - `peaks.bed`: list of peaks of 500bp across all cell states. 4th column contains the peak set label. Note that ~5000 peaks are not assigned to any peak set and are marked as NA.<br> - `features.tsv`: 50 dimensional representation of each cell&nbsp;<br> - `cell_x_peak.mtx.gz`: sparse matrix of fragment counts within peaks. Load using scipy.io.mmread in python or readMM in R. Columns correspond to cells from `cells.tsv` (combine sample + barcode). Rows correspond to peaks in `peaks.bed`&nbsp;</p> <p><strong>scATAC_clusters.zip</strong></p> <p>Analysis products corresponding to cluster pseudo-bulks of the single-cell ATAC-seq data.&nbsp;</p> <p>- `clusters.tsv`: contains the cluster id, name and colour used in the paper<br> - `peaks`: contains `overlap_reproducibilty/overlap.optimal_peak` peaks called using ENCODE bulk ATAC-seq pipeline in the narrowPeak format.<br> - `fragments`: contains per cluster fragment files&nbsp;</p> <p><strong>scATAC_scRNA_integration.zip</strong></p> <p>Analysis products from the integration of scATAC with scRNA. Contains:</p> <p>- `peak_gene_links_fdr1e-4.tsv`: file with peak gene links passing FDR 1e-4. For analyses in the paper, we filter to peaks with absolute correlation &gt;0.45.<br> - `harmony.cca.30.feat.tsv`: 30 dimensional co-embedding for scATAC and scRNA cells obtained by CCA followed by applying Harmony over assay type.<br> - `harmony.cca.metadata.tsv`: UMAP coordinates for scATAC and scRNA cells derived from the Harmony CCA embedding. First column contains barcode.</p> <p><strong>scRNA.zip</strong></p> <p>Analysis products for the single-cell RNA-seq data. Contains:</p> <p>- `seurat.rds`: seurat object that contains expression data (raw counts, normalized, and scaled), reductions (umap, pca), knn graphs, all associated metadata. Note that barcode suffix (1-9 corresponds to samples D0, D2, ..., D14, iPSC)<br> - `genes.txt`: list of all genes<br> - `cells.tsv`: list of barcodes that pass QC across samples. Contains:<br> &nbsp;&nbsp; &nbsp;- `barcode_sample`: barcode with index of sample (1-9 corresponding to D0, D2, ..., D14, iPSC)&nbsp;<br> &nbsp;&nbsp; &nbsp;- `sample`: sample name (D0, D2, .., D14, iPSC)<br> &nbsp;&nbsp; &nbsp;- `umap1`<br> &nbsp;&nbsp; &nbsp;- `umap2`<br> &nbsp;&nbsp; &nbsp;- `nCount_RNA`<br> &nbsp;&nbsp; &nbsp;- `nFeature_RNA`<br> &nbsp;&nbsp; &nbsp;- `cluster`<br> &nbsp;&nbsp; &nbsp;- `percent.mt`: percent of mitochondrial transcripts in cell<br> &nbsp;&nbsp; &nbsp;- `percent.oskm`: percent of OSKM transcripts in cell<br> - `gene_x_cell.mtx.gz`: sparse matrix of gene counts. Load using scipy.io.mmread in python or readMM in R. Columns correspond to cells from `cells.tsv` (barcode suffix contains sample information). Rows correspond to genes in `genes.txt`&nbsp;<br> - `pca.tsv`: first 50 PC of each cell<br> - `oskm_endo_sendai.tsv`: estimated raw counts (cts, may not be integers) and log(1+ tp10k) normalized expression (norm) for endogenous and exogenous (Sendai derived) counts of POU5F1 (OCT4), SOX2, KLF4 and MYC genes. Rows are consistent with `seurat.rds` and `cells.tsv`</p> <p><strong>multiome.zip</strong></p> <p><em>multiome/snATAC:</em></p> <p>These files are derived from the integration of nuclei from multiome (D1M and D2M), with cells from day 2 of scATAC-seq (labeled D2).&nbsp;</p> <p>- `cells.tsv`: This is the list of nuclei barcodes that pass QC from multiome AND also cell barcodes from D2 of scATAC-seq. Includes:<br> &nbsp;&nbsp; &nbsp;- `barcode`<br> &nbsp;&nbsp; &nbsp;- `umap1`: These are the coordinates used for the figures involving multiome in the paper.<br> &nbsp;&nbsp; &nbsp;- `umap2`: ^^^&nbsp;<br> &nbsp;&nbsp; &nbsp;- `sample`: D1M and D2M correspond to multiome, D2 corresponds to day 2 of scATAC-seq<br> &nbsp;&nbsp; &nbsp;- `cluster`: For multiome barcodes, these are labels transfered from scATAC-seq. For D2 scATAC-seq, it is the original cluster labels.&nbsp;<br> - `peaks.bed`: This is the same file as scATAC/peaks.bed. List of peaks of 500bp. 4th column contains the peak set label. Note that ~5000 peaks are not assigned to any peak set and are marked as NA.<br> - `cell_x_peak.mtx.gz`: sparse matrix of fragment counts within peaks. Load using scipy.io.mmread in python or readMM in R. Columns correspond to cells from `cells.tsv` (combine sample + barcode). Rows correspond to peaks in `peaks.bed`.<br> - `features.no.harmony.50d.tsv`: 50 dimensional representation of each cell prior to running Harmony (to correct for batch effect between D2 scATAC and D1M,D2M snMultiome). Rows correspond to cells from `cells.tsv`.<br> - `features.harmony.10d.tsv`: 10 dimensional representation of each cell after running Harmony. Rows correspond to cells from `cells.tsv`.</p> <p><em>multiome/snRNA:</em></p> <p>- `seurat.rds`: seurat object that contains expression data (raw counts, normalized, and scaled), reductions (umap, pca),associated metadata. Note that barcode suffix (1,2 corresponds to samples D1M, D2M). Please use the UMAP/features from snATAC/ for consistency.<br> - `genes.txt`: list of all genes (this is different from the list in scRNA analysis)<br> - `cells.tsv`: list of barcodes that pass QC across samples. Contains:<br> &nbsp;&nbsp; &nbsp;- `barcode_sample`: barcode with index of sample (1,2 corresponding to D1M, D2M respectively)&nbsp;<br> &nbsp;&nbsp; &nbsp;- `sample`: sample name (D1M, D2M)<br> &nbsp;&nbsp; &nbsp;- `nCount_RNA`<br> &nbsp;&nbsp; &nbsp;- `nFeature_RNA`<br> &nbsp;&nbsp; &nbsp;- `percent.oskm`: percent of OSKM genes in cell<br> - `gene_x_cell.mtx.gz`: sparse matrix of gene counts. Load using scipy.io.mmread in python or readMM in R. Columns correspond to cells from `cells.tsv` (barcode suffix contains sample information). Rows correspond to genes in `genes.txt`&nbsp;</p>

opencc-by-4.0Aug 2023View details →
zenodo40/100

Interruption Audio & Transcript: Derived from Group Affect and Performance Dataset

<p><strong>Licensing</strong></p> <p>This dataset is adapted from the Group Affect and Performance dataset which is released under the Creative Commons Attribution-NonCommercial 4.0 International (CC BY-NC 4.0) license.&nbsp;<a href="https://creativecommons.org/licenses/by-nc/4.0/">https://creativecommons.org/licenses/by-nc/4.0/</a></p> <p>&nbsp;</p> <p><strong>Description</strong></p> <p>This dataset contains the audio files containing manually annotated cases of overlapped utterances, classified into True Interruptions and False Interruptions. It is derived from the <a href="https://github.com/gmfraser/gap-corpus/tree/master">Group Affect and Performance</a> dataset created by the University of the Fraser Valley, Canada. Original conversation transcripts and audio files have been supplied for context. The Group Affect and Performance dataset provides a rich source of interruptions and overlapped utterances in general, yielding 200 True Interruptions from 355 instances of overlapped utterances in the 14 Group meetings which were annotated.</p> <p>&nbsp;</p> <p><strong>Structure</strong></p> <p>This dataset is structured into three parts:</p> <p>1. data.json contains a list of all instances of overlapped utterances, classified into &lsquo;interruption&rsquo; and &lsquo;non-interruption&rsquo; corresponding to True and False Interruptions respectively. Each instance is uniquely identified by the Group in which it occurred, the speaker and the starting time of the utterance.</p> <p>2. The 'audio' directory contains the audio of each instance of overlapped utterances corresponding to those found in data.json. The naming convention of the files is as such: &lsquo;Group [group number]: [utterance start time] - [utterance end time].wav&rsquo;.</p> <p>3. Also included is a copy of the original dataset which includes the full audio and transcript. This allows the full meeting to be heard and any context for interruptions to be evaluated.</p> <p>Note that directories 2. and 3. can be accessed by unzipping audio-and-transcripts.zip.</p> <p>&nbsp;</p> <p><strong>Data Collection Protocol</strong></p> <p>Of paramount importance to our process are the definitions of an overlapped utterance and a True Interruption. A False Interruption is simply an overlapped utterance which is not a True Interruption. These definitions directly impact the dataset; for overlapped utterance it informs which data points are included in our dataset and for True Interruption it informs the classes assigned to each sample.</p> <p>In defining an overlapped utterance, our primary aim is to create an overarching class encompassing interruptions and all instances that could be deemed a True Interruption. For this reason, we omit cases where the timing misplaced speech and early-onset responses.</p> <p>An overlapped utterance is defined as an instance where one interlocutor provides speech or noise during another interlocutor&rsquo;s speech, creating an overlap that may be deemed a possible interruption when considering its timing alone. For this reason we omit cases of where the timing indicates misplaced speech or early-onset responses.</p> <p>Our definition of True Interruption&nbsp;is an instance where an interrupting party intentionally attempts to take over a turn of the conversation from an interruptee and, in doing so, creates an overlap in speech.</p> <p>As previously mentioned, due to the &lsquo;intent&rsquo; part of this definition, we avoid cases of misplaced speech and early-onset responses. The former is enforced by not considering cases of overlapped speech which begin within 300ms of each other since this is an estimate for the average human reaction time of articulating a vowel in response to a speech stimuli. The latter is enforced by not considering speech starting within the last 10% of first utterance in the overlapped speech. Note that this approach fails to filter out all cases of misplaced speech, so we manually remove the remaining instances.</p> <p>&nbsp;</p> <p><strong>Methodology</strong></p> <p>Three main steps were taken to produce this dataset:</p> <p>1.&nbsp;Parsing the transcripts for cases of overlapping speech</p> <p>2.&nbsp;Manually annotating these cases per our protocol and adding them to data.json</p> <p>3.&nbsp;Extracting audio samples from data.json and adding them to the audio folder</p> <p>&nbsp;</p> <p>If you use this dataset, please cite the following paper:</p> <p>&nbsp;</p> <p>Doyle, D.; Şerban, O. Interruption Audio &amp; Transcript: Derived from Group Affect and Performance Dataset.&nbsp;<em>Data</em>&nbsp;<strong>2024</strong>,&nbsp;<em>9</em>, 104. https://doi.org/10.3390/data9090104</p> <p>&nbsp;</p> <p>@article{data9090104,</p> <p>AUTHOR = {Doyle, Daniel and Şerban, Ovidiu},</p> <p>TITLE = {Interruption Audio &amp; Transcript: Derived from Group Affect and Performance Dataset},</p> <p>JOURNAL = {Data},</p> <p>VOLUME = {9},</p> <p>YEAR = {2024},</p> <p>NUMBER = {9},</p> <p>ARTICLE-NUMBER = {104},</p> <p>URL = {https://www.mdpi.com/2306-5729/9/9/104},</p> <p>ISSN = {2306-5729},</p> <p>DOI = {10.3390/data9090104}</p> <p>}</p>

openother-ncSep 2023View details →
zenodo40/100

Forschungsprojekt: Digitalisierung, Klassifikationen und Gesundheits-Apps. Dataset A - Transcript of interviews with health app users and developers

<p><strong>Allgemeine Hinweise Data Set A</strong></p> <p>1. Titel des Forschungsprojekts in dessen Rahmen die Daten entstanden sind:</p> <p>&quot;Digitale Gesundheitsklassifikationen in Apps - Praktiken und Probleme ihrer Entwicklung und situativen Anwendung. Projektleitung: Prof. Dr. Rainer Diaz-Bone. Bearbeitung: Valeska Cappel Dipl. Soz., Miriam Kutt (Hilfsassistenz). Laufzeit: 2019-2023. Finanzierung: Schweizer Nationalfonds.&quot;</p> <p>2. Prim&auml;rforscherInnen:</p> <ul> <li>Rainer Diaz-Bone</li> <li>Valeska Cappel</li> <li>Miriam Kutt</li> </ul> <p>3. Publikationsjahr:</p> <ul> <li>2023</li> </ul> <p>4. Hinweise zur Verf&uuml;gbarkeit</p> <p>Die Daten werden &uuml;ber LORY (Lucerne Open Repository) dauerhaft in Form von Transkripten zug&auml;nglich gemacht.</p> <p>5. Fachgebiet</p> <ul> <li>Soziologie</li> </ul> <p>6. Kategorie und Schlagw&ouml;rter</p> <ul> <li>Gesundheitswesen</li> <li>Selbstvermessung</li> <li>Gesundheits-Apps</li> <li>Klassifikationen</li> <li>Pragmatismus</li> <li>Economics of convention</li> <li>Soziologie der Konventionen</li> <li>Digitalisierung</li> </ul> <p>7. Abstract, wozu die Daten erhoben wurden</p> <p>Ziel der Datenerhebung war es die Bewusstseinsstrukturen, Meinungen, Einstellungen und Aushandlungs- und Probleml&ouml;sungsprozesse der Versuchsteilnehmer (Experten aus dem Gesundheitswesen und der App-Nutzenden) zu pr&auml;ventiven Gesundheits-Apps zu erheben. F&uuml;r ein kontrolliertes Erhebungsverfahren wurde dazu die qualitative Methode des Interviews durchgef&uuml;hrt. Mit dieser Methode wurden die Daten als Teil-Aspekte der Realit&auml;t der Befragten verstanden und erhoben. Die Daten beinhalten Interpretations- und Deutungs-, und Bewertungsschemata, Motivationsstrukturen und Aushandlungsprozesse sowie Sachinformationen zu dem Umgang und der Entwicklung von Gesundheits-Apps.</p> <p>8. Untersuchungsgebiet</p> <p>Das Untersuchungsgebiet liegt im Bereich der digitalen Gesundheit und beschr&auml;nkt sich im Speziellen auf die Prozesse der Entwicklung von Gesundheits-Apps, sowie die Nutzung der Gesundheits-Apps. Untersuchungsgebiet waren damit Personen, die in einem Unternehmen im Gesundheitsfeld arbeiten und an der App-Entwicklung beteiligt sind, sowie Personen die Gesundheits-Apps in ihrem Alltag aktiv benutzen.</p> <p>9. Gesamtheit auf die generalisiert werden k&ouml;nnte (&bdquo;Grundgesamtheit&ldquo;)</p> <p>Im Rahmen der qualitativen Interviews wurden 20 Personen befragt.</p> <p>10. Auswahlverfahren und Stichproben</p> <p>Die Daten im Projekt wurden anhand qualitativer Methoden gewonnen. Die F&auml;lle und Daten wurden &uuml;ber die Methode der &bdquo;Theoretical-Sampling-Technik&ldquo; ausgew&auml;hlt. Die genauen Begr&uuml;ndungen zur Auswahl der F&auml;lle wurden im Verlauf des Projektes theoretisch erarbeitet. Dabei wurde sich dem Forschungsfeld der pr&auml;ventiven Gesundheits-Apps mit heuristischen Vermutungen angen&auml;hert, die anhand der Konzepte der Theorie der Konventionen und einer machttheoretischen Perspektive Foucaults entwickelt wurden. Gesundheit und Gesundheitshandlungen wurden dabei aus einer pragmatischen Perspektive als ein Ergebnis von Koordinationsbem&uuml;hungen zwischen Akteuren, Gegenst&auml;nden, Technologien und Machtstrukturen verstanden. In den Analyseschritten der Codierung und Auswertung der Interviews und Dokumente, wurden zur Qualit&auml;tssicherung Memos angelegt. Diese Memos bildeten eine theoretische Grundlage f&uuml;r weitere Auswahlprozesse des Materials. Diese Methode des Theoretical-Samplings wurde zudem durch die Methode des Schneeball-Systems erg&auml;nzt, indem interviewte Personen immer nach weiteren Kontakten befragt wurden und neu genannte Kontakte vor dem Hintergrund des Theoretical-Samplings ausgew&auml;hlt wurden.&nbsp;&nbsp;</p> <p>11. Erhebungszeitraum</p> <ul> <li>2019-2022</li> </ul> <p>12. Sprache</p> <ul> <li>19 Interviews deutsch</li> <li>1 Interview englisch</li> </ul> <p>13. Gr&ouml;&szlig;e des Datensatzes</p> <ul> <li>1,07 MB</li> </ul> <p>14. Verwendete Dateiformate und notwendige Software</p> <ul> <li>Dateiformat: RTF</li> </ul> <p>Software: RTF Standard-Textprogrammen auf unterschiedlichen Betriebssystemen (Bspw. Word, Wordpad, LibreOffice, OpenOffice)</p>

opencc-by-4.0Sep 2023View details →
zenodo40/100

Development and validation of a novel plasmid chassis system for screening of metabolite-responsive transcription factors

<p>This dataset contains the raw data that lie at the basis of the results discussed in <strong>Chapter 3: Development and validation of a novel plasmid chassis system for screening of metabolite-responsive transcription factors&nbsp;</strong>of the PhD thesis of Amber Bernauw.&nbsp;The README.txt file provides more information on the&nbsp;different data files.</p>

opencc-by-4.0Sep 2023View details →
zenodo40/100

In vivo screening of Lrp-type transcription factors in Escherichia coli

<p>This dataset contains the raw data that lie at the basis of the results discussed in&nbsp;<strong>Chapter 4: <em>In vivo</em> screening of Lrp-type&nbsp;transcription factors in <em>Escherichia coli</em></strong><strong>&nbsp;</strong>of the PhD thesis of Amber Bernauw.&nbsp;The README.txt file provides more information on the&nbsp;different data files.</p>

opencc-by-4.0Sep 2023View details →
zenodo40/100

Image stacks for full-body transcription factor expression atlas with completely resolved cell identities in C. elegans

<p>Each image stack presented as&nbsp;zip file. Once decompressed, each folder&nbsp;contain &#39;.ano&#39;&nbsp;linker file,&nbsp;straightening C. elegans L1 images file, the segmentation mask image file and&nbsp;the cell annotation file.&nbsp;The image files are stored in Peng Hanchuan RAW/TIFF format, and the cell annotation file is stored in simple comma separated values format. To&nbsp;visualize the image stack data, drag the &#39;.ano&#39; linker file to VANO interface.&nbsp;</p> <p>vano_win32_1.741.zip contains VANO for worm visualization.</p>

opencc-by-4.0Feb 2023View details →
zenodo40/100

TumourMethData Transcript Counts

<p>RNA-Seq Kallisto transcript counts for TumourMethData</p>

opencc-by-4.0Oct 2023View details →
zenodo40/100

Greetings From! Historical Postcards Address Transcription Dataset

<p>This dataset provides both Ground Truth (GT) and Handwritten Text Recognition (HTR) transcriptions of historical postcard addresses, stemming from a project to extract address information from historical picture postcards from Belgium, France, Germany, Luxembourg, the Netherlands, and the UK. The dataset encapsulates the back of 500 historically significant postcards.</p><p>The research associated with this dataset will be presented at <strong>Computational Humanities Research Conference, December 6--8, 2023, Paris, France.</strong></p><p><strong>Scope and Content:</strong></p><ul><li><strong>HTR Material</strong>: Handwritten Text Recognition outputs for 500 postcards.</li><li><strong>GT Material</strong>: Ground Truth transcriptions created by human transcribers for the same set of 500 postcards.</li></ul><p><strong>File Structure and Formats:</strong></p><p><i>For both HTR and GT Material, the following files are provided:</i></p><ul><li><strong>JPEG Images</strong>: Scanned or digitized images of the postcards.</li><li><strong>.txt</strong>: Plain text transcriptions of the postcards.</li><li><strong>_tei.xml</strong>: Transcriptions rendered in the TEI XML format.</li><li><strong>.pdf</strong>: PDF presentation of the postcards along with their transcriptions.</li><li><strong>mets.xml</strong>: METS (Metadata Encoding and Transmission Standard) schema for the data.</li><li><strong>page</strong> folder: XML files for individual images, offering metadata and structural information.</li><li><strong>metadata.xml</strong>: metadata concerning the dataset.</li><li><strong>GT_addresses_GPT4.json &amp; HTR_addresses_GPT4.json</strong>: JSON files detailing individual address data for each postcard in structured format.</li></ul><p><strong>Annotation and Transcription:</strong></p><ul><li><strong>GT</strong>: Ground Truth data was annotated by human transcribers who examined both the images of the postcards and the outputs of the HTR system. Transcribers made corrections according to predefined conventions: using <strong>#</strong> for illegible characters, <strong>*</strong> at the start of lines without address information (e.g., person's name), and starting a line with <strong>@</strong> for irrelevant lines.</li><li><strong>HTR</strong>: The HTR versions emerged from state-of-the-art HTR systems (<a href="https://readcoop.eu/introducing-transkribus-super-models-get-access-to-the-text-titan-i/">Transkribus Text Titan I</a>). The <strong>.json</strong> files hold precise address details derived from the main data, which were processed using OpenAI's GPT-4 Large Language Model.</li></ul>

opencc-by-sa-4.0Oct 2023View details →
zenodo40/100

Supplementary data accompanying HOCOMOCO v12 collection of transcription factor binding motifs

<p><strong>*** SUMMARY ***</strong></p><p>This dataset contains supplementary data accompanying HOCOMOCO v12 collection</p><p>of DNA binding motifs for human and mouse transcription factors, https://hocomoco.autosome.org</p><p>&nbsp;</p><p>The contents include:</p><p>- the complete initial set of motifs discovered from ChIP-Seq and HT-SELEX data;</p><p>- the curated subset of motifs associated with distinct motif subtypes that were used in benchmarking;</p><p>- the benchmarking results and the resulting final motif collections, including motif logos;</p><p>- accompanying metadata.</p><p>&nbsp;</p><p>Please refer to the README and the HOCOMOCO website for further details.</p><p>&nbsp;</p>

opencc-by-4.0Oct 2023View details →
zenodo40/100

target_rhabdoid_wgbs_hg19 and transcript counts

<p>A HDF5-backed RangedSummarizedExperiment for WGBS Data (hg19&nbsp;CpG sites)&nbsp;for 69 rhabdoid&nbsp;tumours and RNA-seq transcript counts for 65 of these samples from the paper 'Chun, Hye-Jung E., et al. "Genome-wide profiles of extra-cranial malignant rhabdoid tumors reveal heterogeneity and dysregulated developmental pathways." <i>Cancer cell</i> 29.3 (2016): 394-406.'.&nbsp;</p>

opencc-by-4.0Oct 2023View details →
dryad40/100

Transcript- and annotation-guided genome assembly of the European starling

Open the record for dataset details and reuse information.

publicJul 2022View details →
dryad40/100

Data from: Suppression of Huntington’s disease somatic instability by transcriptional repression and direct CAG repeat binding

Open the record for dataset details and reuse information.

publicSep 2025View details →
dryad40/100

Data from: Experimentally induced active and quiet sleep engage non-overlapping transcriptional programs in Drosophila

Open the record for dataset details and reuse information.

publicOct 2023View details →
dryad40/100

Data from: Pleomorphic effects of three small-molecule inhibitors on transcription elongation by <em>Mycobacterium tuberculosis</em> RNA polymerase

Open the record for dataset details and reuse information.

publicOct 2025View details →
dryad40/100

Supplementary data from: Inherent single-cell heterogeneity of the transcriptional response to hypoxia in cancer cells

Open the record for dataset details and reuse information.

publicSep 2025View details →
dryad40/100

Data from: Nascent transcription reveals regulatory changes in extremophile fishes inhabiting hydrogen sulfide-rich environments

Open the record for dataset details and reuse information.

publicMay 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record