Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
394
datasets available to search
ShareScore release 0.9.0
Dataset results
394 results for “thesis”
Epiphytic macrolichens in relation to forest management and topography in a western Oregon watershed, 1997-1999 (Berryman thesis)
Epiphytic macrolichen communities were sampled in 117 coniferous stands in Blue River watershed of western Oregon. Stands were sampled across various stand types defined by stand structure, according to age classes of the younger tree cohort and remnant tree retention. Remnant trees were those in an older cohort that remained following a stand disturbance that initiated tree regeneration, such as a timber harvest or natural forest fire. Stands were located in upland and riparian forests of two vascular plant series (western hemlock and true fir). Presence and abundance of all epiphytic macrolichen species were sampled in a 0.4 ha circular Forest Health Monitoring (FHM) plot in the 117 stands. Epiphytic lichen biomass (oven-dried, kg/ha) was estimated for three functional groups: nitrogen-fixing cyanolichens, forage lichens, and matrix lichens in 63 of the 117 stands.
Genome assemblies of four MDR B. fragilis isolates using PacBio data - supporting the PhD Thesis
<p>Unicycler and Canu assemblies using PacBio data of four MDR B. fragilis isolates.</p> <p>Data supporting the PhD Thesis <em>Epidemiology and Genomics of antimicrobial resistance in the Bacteroides fragilis group </em>by Thomas Vognbjerg Sydenham, The research unit of Clinical Microbiology, Department of Clinical Research, Faculty of Health Sciences, University of Southern Denmark September 2019.</p> <p> </p>
Datasets from Approximate equality of character strings and its application to record linkage in metadata of scientific publications thesis
<p>The datasets were produced in my thesis project. The thesis (in Czech language) explores the application of approximate string matching in scientific publication record linkage process. An introduction to record matching along with five commonly used metrics for string distance (Levenshtein, Jaro, Jaro-Winkler, Cosine distances and Jaccard coefficient) are provided. These metrics are applied on publication metadata from V3S current research information system of the Czech Technical University in Prague. Based on the findings, optimal thresholds in the F1, F2 and F3-measures are determined for each metric.</p> <p>Thesis citation:<br> DOBIÁŠOVSKÝ, Jan. <em>Approximate equality of character strings and its application to record linkage in metadata of scientific publications</em> [online]. Praha, 2020 [cit. 2020-05-04]. Masters thesis. Charles University. Faculty of Arts. Institute of Information Studies and Librarianship.</p> <p> </p>
Audio data from thesis Perception and Production of Nanning Mandarin Fourth Tone
<p>Recordings of 4 female speakers (S1, S2, S3, and S4) of Nanning-accented Mandarin Chinese reading preconstructed sentences. Recordings of those 4 female speakers telling a story based on 5 pages from Mercer Mayer's wordless picture book <em>Frog on His Own</em> (FOHO). Recording of perception test and perception test warm-up given to 26 Chinese living in Nanning, Guangxi. List of corpus sentences read, test prompts, and test sheet.</p>
Datasets used in the thesis "Contributions to High-Dimensional Pattern Recognition"
<p>Datasets used in the thesis "Contributions to High-Dimensional Pattern Recognition" and related publications.</p>
Data for H Wierstorf, Perceptual Assessment of Sound Field Synthesis, PhD thesis, TU Berlin, (2014)
<p>This publication contains data that was used to generate figures in the thesis:</p> <p>H. Wierstorf, Perceptual Assessment of Sound Field Synthesis, PhD thesis, TU Berlin, (2014), http://dx.doi.org/10.14279/depositonce-4310</p> <p>The code to generate all the figures is available at https://github.com/hagenw/phd-thesis excluding the data provided by this data set. The data contains no measured data, but the result of numerical simulations for which the actual code to generate it, is also part of the repository at github.</p>
Supplementary Materials for SE Carter PhD Thesis
<p>Included:</p> <p>- Additional Phase I: Value and Privacy Preference Survey graphs (see Version 1)</p> <p>- VcPA Profile Design Graphs and Statistics for Phase II: Mock App Store Study (v2 includes correct cluster numbers)</p> <p>- Code List for Phase III: Semi-Structured Interviews (see Version 1)</p>
Data supporting the Master thesis "Monitoring von Open Data Praktiken - Herausforderungen beim Auffinden von Datenpublikationen am Beispiel der Publikationen von Forschenden der TU Dresden"
<p>Data supporting the Master thesis "Monitoring von Open Data Praktiken - Herausforderungen beim Auffinden von Datenpublikationen am Beispiel der Publikationen von Forschenden der TU Dresden" (Monitoring open data practices - challenges in finding data publications using the example of publications by researchers at TU Dresden) - Katharina Zinke, Institut für Bibliotheks- und Informationswissenschaften, Humboldt-Universität Berlin, 2023</p> <p>This ZIP-File contains the data the thesis is based on, interim exports of the results and the R script with all pre-processing, data merging and analyses carried out. The documentation of the additional, explorative analysis is also available. The actual PDFs and text files of the scientific papers used are not included as they are published open access.</p> <p>The folder structure is shown below with the file names and a brief description of the contents of each file. For details concerning the analyses approach, please refer to the master's thesis (publication following soon).</p> <p>## Data sources </p> <p>Folder 01_SourceData/</p> <p>- PLOS-Dataset_v2_Mar23.csv (PLOS-OSI dataset)</p> <p>- ScopusSearch_ExportResults.csv (export of Scopus search results from Scopus)</p> <p>- ScopusSearch_ExportResults.ris (export of Scopus search results from Scopus)</p> <p>- Zotero_Export_ScopusSearch.csv (export of the file names and DOIs of the Scopus search results from Zotero)</p> <p>## Automatic classification </p> <p>Folder 02_AutomaticClassification/</p> <p>- (NOT INCLUDED) PDFs folder (Folder for PDFs of all publications identified by the Scopus search, named AuthorLastName_Year_PublicationTitle_Title) </p> <p>- (NOT INCLUDED) PDFs_to_text folder (Folder for all texts extracted from the PDFs by ODDPub, named AuthorLastName_Year_PublicationTitle_Title)</p> <p>- PLOS_ScopusSearch_matched.csv (merge of the Scopus search results with the PLOS_OSI dataset for the files contained in both)</p> <p>- oddpub_results_wDOIs.csv (results file of the ODDPub classification)</p> <p>- PLOS_ODDPub.csv (merge of the results file of the ODDPub classification with the PLOS-OSI dataset for the publications contained in both)</p> <p>## Manual coding </p> <p>Folder 03_ManualCheck/</p> <p>- CodeSheet_ManualCheck.txt (Code sheet with descriptions of the variables for manual coding)</p> <p>- ManualCheck_2023-06-08.csv (Manual coding results file)</p> <p>- PLOS_ODDPub_Manual.csv (Merge of the results file of the ODDPub and PLOS-OSI classification with the results file of the manual coding)</p> <p>## Explorative analysis for the discoverability of open data<br> <br>Folder04_FurtherAnalyses </p> <p>Proof_of_of_Concept_Open_Data_Monitoring.pdf (Description of the explorative analysis of the discoverability of open data publications using the example of a researcher) - in German</p> <p>## R-Script </p> <p>Analyses_MA_OpenDataMonitoring.R (R-Script for preparing, merging and analyzing the data and for performing the ODDPub algorithm)</p>
Glyco Atlas Supplemental Tables (Leo Alexander Dworkin) PhD Thesis
<p>The complex multi-step process of glycosylation occurs in a single cell, yet current analytics generally cannot measure the output (the glycome) of a single cell. Here, we addressed this discordance by testing usage of single cell transcriptomics to break this barrier. We investigated how single cell RNA-seq data can be used to characterise the state of the glycosylation machinery and metabolic network in single cell. The metabolic network involves 214 glycosylation and modification enzymes with their contributions to the glycome outlined in our previously built atlas of cellular glycosylation pathways. We studied differential mRNA regulation of enzymes at the organ and single cell level, finding that most of the general protein and lipid oligosaccharide scaffolds are produced by enzymes exhibiting limited transcriptional regulation among cells. We predict key enzymes within different glycosylation pathways to be highly transcriptionally regulated as regulatable hotspots of the cellular glycome. We designed the Glycopacity software that enables investigators to extract and interpret glycosylation information from transcriptome data and define hotspots of regulation.</p>
Karst diagram from PhD thesis of Laia Comas-Bru
<p>Schematic illustration showing the formation of speleothems</p> <p>Formats: editable pdf and jpeg.</p> <p>Versions: one with the parameters modifying stable oxygen isotopes and one without.</p> <p>Original image: Fig 1.3 of "Schematic illustration showing the formation of speleothems" PhD thesis of Laia Comas-Bru. University College Dublin, Ireland.2015.</p>
Life tables and graphs for Bahry (2022) - Equilibrium conditions in the evolution of senescence [MSc thesis, Carleton Univeristy]
<p>Life table data, and derived quantities, for <em>Equilibrium Conditions in the Evolution of Senescence</em> (Bahry, 2022, MSc thesis); adapted from the supplementary data of (Jones et al., 2014). Life table data for human (Japan 2009), human (Aché hunter-gatherer), fruit fly, Soay sheep, freshwater hydra, and desert tortoise.</p> <p>Basic life table quantities: age interval <span class="math-tex">\((X)\)</span>; survival function <span class="math-tex">\((l_X)\)</span>; and age-specific interval fecundity <span class="math-tex">\((m_X)\)</span>. Derived quantities include interval average force of mortality; reproductive value; residual reproductive value; Hamilton's indicators of the age-specific forces of selection; and actual age-specific mortality vs. predicted age-specific mortality based on models treated in (Bahry, 2022).</p> <p>In the original life tables of Jones et al. (2014), desert tortoises negatively senesce over the range of observed ages, but had a final observed cut-off age of 74; this causes reproductive value to artifactually fall to 0 as age-approached the cutoff. To get around this, I also used an extrapolated desert tortoise life table, assuming the age-74 mortality and fecundity rates remained constant until age 1000, then using the extrapolated life table to calculate reproductive value (and Hamilton's indicators) up to the cutoff age 74.</p> <p><strong>References</strong></p> <p>Bahry, D. (2022). <em>Equilibrium Conditions in the Evolution of Senescence</em> [Master's thesis, Carleton University].</p> <p>Jones, O. R. et al. (2014). Diversity of ageing across the tree of life. <em>Nature</em> 505: 169–174. https://doi.org/10.1038/nature12789</p>
Text-fig. 6. Correlation of the Cheringoma and Mazamba formations on the basis of benthic foraminiferans and mammals respectively. Identifications of foraminiferans are from Newton (1924) and Abrard (1928), and the ranges of foraminiferans are from Sella-Kiel et al. (1998). The time scale is from Gradstein et al. (2020). The distribution of Nummulites atacicus is included, but it is not known whether it is reworked from older deposits. If the identification is valid, it would support the thesis that there was a period of Ypresian deposition in the vicinity during which remains of the species were fossilised. in Stratigraphy, Chronology And Palaeontology Of The Tertiary Rocks Of The Cheringoma Plateau, Mozambique
Text-fig. 6. Correlation of the Cheringoma and Mazamba formations on the basis of benthic foraminiferans and mammals respectively. Identifications of foraminiferans are from Newton (1924) and Abrard (1928), and the ranges of foraminiferans are from Sella-Kiel et al. (1998). The time scale is from Gradstein et al. (2020). The distribution of Nummulites atacicus is included, but it is not known whether it is reworked from older deposits. If the identification is valid, it would support the thesis that there was a period of Ypresian deposition in the vicinity during which remains of the species were fossilised.
Dan Richmond PhD thesis depository
<p>Data depository for PDB file of 5,6-Tryp, and HTML files of taxonomic plots.</p>
Data from the MA Thesis "Does She Talk Differently?: Exploring Implications of Gender in US Presidential and Vice Presidential Debates" and Coded Transcriptions of the 7 Analyzed Debates
<p>The raw data collected for the master's thesis "Does She Talk Differently?: Exploring Implications of Gender in US Presidential and Vice Presidential Debates", the tables and graphs created based on the data as well as the transcriptions for the seven debates analyzed for the research paper can be found in the files.</p>
Diagramms and Figures for a Bachelor Thesis: "Morphological Diversity of Bryophytes: Methodological Approaches and Ecological Implications"
<p>This depository holds all diagramms and figures that were created with the collected data used in the Bachelor's Thesis using R Version 4.4.0 or Python Version 3.12.4. <br><br></p>
Dataset Thesis of Diskriminasi Gender Berita Kriminalitas Online Di Indonesia
<p>Dataset pada file ZIP berisi data ringkasan berita kriminalitas di indonesia pada 1 Januari - 31 Desember 2023, data hasil pra-pemrosesan, dan data ekstraksi fitur vektor menggunakan word embedding sebelum dan sesudah debiasing.</p>
AQL queries and benchmark results from PhD thesis "ANNIS: A graph-based query system for deeply annotated text corpora"
<p>These are the queries, the benchmark results and the evaluation scripts of the thesis "ANNIS: A graph-based query system for deeply annotated text corpora" (Thomas Krause 2018, Humboldt-Universität zu Berlin)</p> <p><strong>diss_2018-01-12_v0.5.0.csv </strong><br> Results of all configurations of executed benchmarks for graphANNIS and also the baseline times of relANNIS.</p> <p><strong>queries.zip</strong><br> Contains folders for each corpus containing all queries used for the benchmark. Each file-name begins with the ID of the query. The extension denotes the type, which can be one of the following:</p> <ul> <li><em>".</em>aql" contains the original AQL (ANNIS query language) query which was collected</li> <li>".json" is JSON representation of the parsed AQL query</li> <li>".count" is the number of matches a query should have</li> <li>".time" is the average time in milliseconds that was needed to execute the query in relANNIS on the benchmark system</li> <li>".corpora" contains the name of the corpus the query belongs to (should be only one corpus and the same as the folder name in the selection of queries in this data set)</li> <li>".relplan" contains the PostgreSQL plan for the query</li> <li>".graphplan" contains the graphANNIS plan for the query</li> </ul> <p><strong>evaluation-scripts.py/evaluation-scripts.ipynb</strong><br> Python scripts to perform the evaluation and generate the images. This are both a Python-file and the original notebook file that can be used with the Jupyter Notebook application.</p> <p><strong>relannis_benchmark_scripts.zip </strong><br> The files in this zip-file can be used to execute the benchmarks in the relANNIS system by piping the into the "annis.sh" command line tool of relANNIS</p>
Data of the PhD thesis "Merge-and-Shrink Abstractions for Classical Planning: Theory, Strategies, and Implementation"
<p>This data set contains raw data and parsed data of all experiments [1] run for the PhD thesis. They were generated using lab (see https://doi.org/10.5281/zenodo.399255).</p> <p>The raw data files (sievers-phd2017-raw-data-part*.tar.gz) contain a subdirectory for each experiment, each containing a subdirectory for each planner run of the experiment, distributed over the directories runs-*. For each run, there are the input PDDL files, domain.pddl and problem.pddl, the compressed output as generated by the translator component of Fast Downward (output.sas.xz), the run log file "run.log" (stdout), possibly also a run error file "run.err" (stderr), and the run script "run" used to start the experiment. The latter cannot be directly used, however, because the directory containing source code and build (compiled object files) have been removed for space reasons. The code is publicly available under https://doi.org/10.5281/zenodo.1163381. The (lab) scripts for parsing run.log are also available in the main directory of each experiment. All other scripts and a corresponding lab version are available on request.</p> <p>For each raw data experiment, the parsed data file (sievers-phd2017-parsed-data.tar.gz) also contains a directory of the same name, with "-eval" appended. It contains a single file called "properties" that combines all of the experiment's parsed data (which can be and was generated from the raw data using lab and the parser scripts). They are in the json format and can be used for easy manipulation of the data. The directories with the prefix "paper-" and "talk-" are combinations of other directories (using the "fetch" mechanism of lab). It is recommended to use these, because due to technical errors, the original eval directories do not contain all runs of all planners (to be more precise: they contain all runs, but a subset of the planner have not been started in these experiments for technical errors and thus considered not solving the task). The missing ones have been run separately, see the directories with "missing-runs" in their name. This is also the reason some of these directories ("paper-", "talk-") contain files named "old-properties" and "fixed-properties" besides the actual "properties". "old-properties" are those with missing/faulty runs, "fixed-properties" are as "old-properties", however with the data of faulty runs removed, and "properties" are as "fixed-properties", however with the addition of the fixed missing runs (in fact, these always contain *all* fixed missing runs of all experiments, for technical reasons).</p> <p>The file sievers-phd2017-parsed-data-all-and-random-merge-strategies.tar.gz contains parsed data of earlier experiments (see [1]), for which no raw data has been archived. The directories contain properties files in the json format.</p> <p>[1] except raw data for the parsed data "sota-symba-spmas-eval" (which in the meantime was added to a separate data set available under https://doi.org/10.5281/zenodo.1189912) and all re-used experiments from the paper "An Analysis of Merge Strategies for Merge-and-Shrink Heuristics" (Silvan Sievers, Martin Wehrle and Malte Helmert, ICAPS 2016), for which the raw data was too large to be archived.</p>
Research data, sources and documents for thesis on Exploring Complexity Metrics for Artifact-Centric Business Process Models
<p>Research data, sources and documents for thesis on Exploring Complexity Metrics for Artifact-Centric Business Process Models This repository contains the supplemental material for the <a href="https://pqdtopen.proquest.com/pubnum/10759956.html">thesis "Exploring Complexity Metrics for Artifact-Centric Business Process Models" by Marin, Mike A., Ph.D., University of South Africa (South Africa), 2017.</a></p>
Data and Code For PhD Thesis- Making an Impression: An Assessment of the Role of Print Surfaces Within the Technological, Commercial, Intellectual and Cultural Trajectory of Book Illustration c.1780-c.1860.
<p><strong>Overview</strong></p> <p>These datasets support research found in the PhD Thesis entitled: Making an Impression: An Assessment of the Role of Print Surfaces Within the Technological, Commercial, Intellectual and Cultural Trajectory of Book Illustration c.1780-c.1860, submitted by William Finley. The data was provided by the British Library as part of a wider ambition to digitise millions of book illustrations (further details can be found here: <a href="https://github.com/BL-Labs/imagedirectory">https://github.com/BL-Labs/imagedirectory</a>). The main dataset contains 6063 rows and 108,784 illustrations. All of the data is held in comma-separated values (csv) files. The collection has been subdivided in order to interrogate the dataset further. All of the datasets have been interrogated using Rscript (R 3.4.2 El Capitan Build), the codes of which have been included in the dataset repository. </p> <p><strong>Guidance to Datasets</strong></p> <p>'Counts of Illustration by Size over Time'</p> <ul> <li>Title: Title of Book</li> <li>Author: Author of Book </li> <li>Pub_Place: Location the book was published</li> <li>Book_Id: British Library image identifier </li> <li>YEAR: Year the book was published</li> <li>Image_Count: Number of illustrations belonging to that book</li> <li>N: Number of illustrations according to the size of illustration</li> <li>IMAGE_TYPE: Size of Illustration </li> </ul> <p>'Density Graphs of the Position of Illustrations on the Page'</p> <ul> <li>X1: The X position in pixels of the top left of the image</li> <li>Y1: The Y position in pixels of the top left of the image</li> <li>X2: Width in pixels of the Image</li> <li>Y2: Height in pixels of the Image</li> <li>PERCENT_PAGE: Percentage of the page taken up by illustrations</li> <li>TYPE: Percentage Range of the Page taken up by illustrations </li> <li>VOL: Volume</li> </ul> <p>'Frequency of Illustrations Across First 100 Pages of the Book 1800-1850'</p> <ul> <li>Page: Page number </li> <li>X1: The X position in pixels of the top left of the image</li> <li>Y1: The Y position in pixels of the top left of the image</li> <li>X2: Width in pixels of the image</li> <li>Y2: Height in pixels of the image</li> <li>PERCENTPAGE: Percentage of the page occupied by illustration</li> <li>SIZE: Size category each illustration belongs to</li> </ul> <p>'Illustration Arrangement in Single Books and Editions'</p> <ul> <li>Book_ID: British Library image identifier</li> <li>Page_NO: Page Number Illustration is found on </li> <li>X: The X position in pixels of the top left of the image</li> <li>Y: The Y position in pixels of the top left of the image </li> <li>WIDTH: Width in pixels of the image</li> <li>HEIGHT: Height in pixels of the image</li> <li>IMAGE_SIZE: Percentage of the page occupied by illustration</li> </ul> <p>'Relative Frequency of Printing Methods Over Time'</p> <ul> <li>Publisher: Publisher of the book</li> <li>Title: Title of the book</li> <li>first_author: Author</li> <li>pub_place: Location the book was published</li> <li>book_identifier: British Library image identifier</li> <li>Year: Year the book was published </li> <li>COUNT_IMAGE: Number of illustrations found in a given book</li> <li>YEAR_BOOK_COUNT: Number of book published in a given year</li> <li>YEAR_IMAGE: Number of Illustrations found in books published in a given year</li> <li>AVERAGE_IMAGE: Average number of illustrations found in books published in a given year</li> <li>IMAGE_METHOD: Print method used to produce the illustration</li> <li>IMAGE_SIZE_COUNT: Number of illustrations printed in books in a given year according to its size</li> <li>IMAGE_SIZE: Categories of image sizes</li> </ul> <p><strong>Contact</strong></p> <p>Will Finley can be contacted via the following email. Information listed is accurate at the time of publication</p> <p>Email: wafinley1@sheffield.ac.uk</p> <p> </p> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.