Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,298
datasets available to search
ShareScore release 0.9.0
Dataset results
1,298 results for “Archive”
Interviews on Current Practices for Describing and Providing Access to UK Public Sector Web Archives
<p>This dataset contains qualitative interview data which investigated current practice for describing and providing access to UK Public Sector Web Archives. Participants included staff responsible for the management and curation of the following web archives:</p><ul><li>UK Web Archive (four of the six Legal Deposit libraries: the British Library, Bodleian Libraries, Cambridge University Library, and the National Library of Scotland)</li><li>UK Government Web Archive (The National Archives)</li><li>UK Parliament Web Archive (Parliamentary Archives)</li><li>NRS Web Archive (National Records of Scotland) and</li><li>PRONI Web Archive (Public Record Office Northern Ireland).</li></ul><p>Available to the public are the University of Dundee (UoD) ethics application for this study, including the research data management plan and information provided to organisations before participating in the study. The report of interview codes and code groups (the 'Codebook') demonstrates the connections made across responses. This is supplemented by a redacted report of quotations by code, organised by code group and document.</p><p>This qualitative interview data, and subsequent analysis, forms the basis of the Masters thesis 'Web Archives for All? Towards Equitable Access to UK Public Sector Web Archives' submitted as part of the MLitt Archives and Records Management at the University of Dundee. </p>
Data archive for: Exploring the use of machine learning to improve vertical profiles of temperature and moisture
<p>Vertical profiles of temperature and dewpoint are useful in predicting deep convection that leads to severe weather that threatens property and lives. Currently, forecasters rely on observations from radiosonde launches and numerical weather prediction (NWP) models. Radiosonde observations are, however, temporally and spatially sparse, and NWP models contain inherent errors that influence short-term predictions of high-impact events. This work explores using machine learning (ML) to postprocess NWP model forecasts, combining them with satellite data to improve vertical profiles of temperature and dewpoint. We focus on different ML architectures, loss functions, and input features to optimize predictions. Because we are predicting vertical profiles at 256 levels in the atmosphere, this work provides a unique perspective at using ML for 1-D tasks. Compared to baseline profiles from the Rapid Refresh (RAP), ML predictions offer the largest improvement for dewpoint, particularly in the mid- and upper-atmosphere. emperature improvements are modest, but CAPE values are improved by up to 40%. Feature importance analyses indicate that the ML models are primarily improving incoming RAP biases. While additional model and satellite data offer some improvement to the predictions, architecture choice is more important than feature selection in fine-tuning the results. Our proposed deep residual UNet performs the best by leveraging spatial context from the input RAP profiles; however, the results are remarkably robust across model architecture. Further, uncertainty estimates for every level are well-calibrated and can provide useful information to forecasters.</p>
Lost in translation: A historical-comparative reconstruction of Proto-Khoe-Kwadi based on archival data - Supplementary Material
<p>Supplementary Materials 1-4 for the following article: </p><p>Fehn, Anne-Maria & Jorge Rocha. 2023. Lost in translation: A historical-comparative reconstruction of Proto-Khoe-Kwadi based on archival data. Diachronica. https://doi.org/10.1075/dia.23022.feh.</p>
Bridging the Divide: Connecting Language Activist Efforts and Language Archives
<p>Bridging the Divide: Connecting Language Activist Efforts and Language Archives</p> <p>Subhashish Panigrahi, Mandana Seyfeddinipur and Susan Kung at the Language Documentation and Archiving conference in the Berlin-Brandenburg Academy of Sciences and Humanities on October 6. 2022</p> <p>Language documentation, revitalization, reclamation, and activism efforts take place all over the world. At the local, grassroots, community and international levels, participants have taken agency and self-organised to engage in these activities to create a documentary record of their own languages, to preserve cultural and linguistic richness, and to reclaim ownership of and control over their languages and cultures, ensuring data sovereignty. In academia, linguists have developed theoretical methods for linguistic language documentation and have created language archives housed at universities. Language activists have created language documentation training materials, organised projects in the Wikimedia ecosystem, formed nonprofits and NGOs, and used social media platforms to self-organise and share their materials. However, many of these grassroots efforts lack access to stable archives that can provide long-term digital preservation of these unique and invaluable materials. Simultaneously, language archives based at universities could provide long-term preservation but lack the connection to activists. In this presentation, we showcase some of these community-based efforts, and we argue for the need to bridge the divide between academically based archives and the "real world" in order to ensure that all language documentation efforts will be preserved for the long-term and accessible and available to all peoples well into the future. We also share examples demonstrating how different kinds of archives fit into the needs and expertise levels of different local activist groups. While taking into account some of the existing practices of community-led efforts for sharing materials online that are more convenient and have better visibility among the viewers, we illustrate the skill development and resource allocation that would be required to migrate to long-term archives. We also discuss the current entry-level barriers of archives that need mitigation for forging activism-academic collaborations and paving the path for robust archives while ensuring the agency of speakers.</p>
Music Data, Archiving for Community Use and Future Directions Through the Decade of Indigenous Languages
<p>Music Data, Archiving for Community Use and Future Directions Through the Decade of Indigenous Languages</p> <p>Linda Barwick</p> <p>Presented 5 October 2022 at the international conference "Where Do We Need to Go From Here?" Language Documentation and Archiving in the International Decade of Indigenous Languages</p>
The best of both worlds: Bringing together community knowledge and design with institutional archives
<p>The best of both worlds: Bringing together community knowledge and design with institutional archives</p> <p>Vera Ferreira, Buachut Watyam, Siripen Ungsitpoonporn, & Mandana Seyfeddinipur</p> <p>Presented 7 October 2022 at the Berlin-Brandenburg Academy of Sciences and Humanities Where Do We Need to Go From Here? Language Documentation and Archiving in the International Decade of Indigenous Languages</p>
The National Archives - Richmond, UK
A videogrammetry scan of The National Archives in Richmond, UK. The scan was made in Agisoft software from 158 frames from this video: https://www.youtube.com/watch?v=ncNVLQoKTo8 Source: Objaverse 1.0 / Sketchfab
Internet Archive Building
The outside of the Internet Archive HQ located in San Francisco California Source: Objaverse 1.0 / Sketchfab
Data from: Evaluating genotyping-in-thousands by sequencing as a genetic monitoring tool for a climate sentinel mammal using non-invasive and archival samples
<p>Genetic tools for wildlife monitoring can provide valuable information on spatiotemporal population trends and connectivity, particularly in systems experiencing rapid environmental change. Though many DNA sequencing approaches still require high quality and quantity of DNA obtained from traditional sources (e.g. blood and tissue), rapid genotyping tools such as Genotyping-in-Thousands by sequencing (GT-seq) have improved our ability to make use of degraded and less concentrated DNA commonly obtained from non-invasive and archival samples. Here, we developed a multi-purpose GT-seq panel (307 single nucleotide polymorphisms) for a climate sentinel mammal (the American pika, <em>Ochotona princeps</em>) for use as a genetic tool for monitoring populations in the Canadian Rocky Mountains. We optimized the panel using contemporary tissue samples (n = 77) and subsequently applied it to archival tissue (n = 17) and contemporary fecal pellet samples (n = 129) to evaluate its effectiveness at identifying individuals and sex, estimating relatedness, and inferring population structure. The panel demonstrated high efficacy with contemporary and archival tissue samples (94.7% and 90.5% genotyping success, respectively) and negligible genotyping error (0.001% and 0.0%, respectively). Despite relatively high genotyping success for fecal pellet samples (79.7%), high genotyping error (28.4%) limited its power as a monitoring tool to assess genetic variation using non-invasive samples and highlighted the need for further optimization around sample and data collection.</p>
Chimpanzee food processing video archive (Waibira community, Budongo Forest, Uganda)
<p>This is a video archive containing data (n = 474 videos) on food processing techniques and other behaviours used by the chimpanzees in the Waibira community, Budongo Forest, Uganda. Each video is labelled with data on date, location, individuals and behaviours present (including non-processing behaviours). </p> <p>Also included is the ethogram used to categorise processing behaviours as they are labelled in the archive, and 19 exemplar clips demonstrating some of these behaviours. This archive is intended to be used as a data ark for future interdisciplinary research. While the videos themselves cannot be made publicly accessible, those interested in collaborations may browse the available videos and contact the researchers to request access to specific videos. </p>
Data archive for the peer-reviewed journal article "Online measurements during simulated atmospheric aging track the strongly increasing oxidative potential of complex combustion aerosols relative to their primary emissions"
<p>This data archive accompanies the article "Online measurements during simulated atmospheric aging track the strongly increasing oxidative potential of complex combustion aerosols relative to their primary emissions", which was accepted in November 2024 in the peer-reviewed journal Environmental Science and Technology Letters. The data archive contains the processed OP_DTT, PM loading, oxidant level, and elemental ratio measurements presented in this journal article. </p>
Archived Model Output and Code for "Marine Boundary Layer Cloud Condensation Nuclei Bias over the Southern Ocean: Comparisons between the Community Atmosphere Model 6 and Field Observations "
<div> <p>This is an archive of CAM6 simulation output used in the paper Marine Boundary Layer Cloud Condensation Nuclei Bias over the Southern Ocean: Comparisons between the Community Atmosphere Model 6 and Field Observations, submitted to the AGU Journal. Codes used to read the nc file is also attached.</p> </div>
Hardware Performance Archive
<h3>SHA256</h3> <h3>MD5</h3>
Archival bundle of the data used for "Predictive Auto-scaling with OpenStack Monasca" (UCC 2021)
<p>This archive contains the data used for the paper</p> <p><strong>Predictive Auto-scaling with OpenStack Monasca</strong><br> <a href="mailto:giacomo.lanciano@sns.it">Giacomo Lanciano</a>*, Filippo Galli, Tommaso Cucinotta, Davide Bacciu, Andrea Passarella<br> 2021 IEEE/ACM 14th International Conference on Utility and Cloud Computing (UCC)<br> <a href="https://doi.org/10.1145/3468737.3494104">10.1145/3468737.3494104</a></p> <p>Follow the instructions provided in the <a href="https://github.com/giacomolanciano/UCC2021-predictive-auto-scaling-openstack">companion repo</a> to automatically download and decompress the archive. The following files are included:</p> <table> <tbody> <tr> <td><strong>File</strong></td> <td><strong>Description</strong></td> </tr> <tr> <td> <p>amphora-x64-haproxy.qcow2</p> </td> <td> <p>Image used to create Octavia amphorae</p> </td> </tr> <tr> <td> <p>distwalk-{lin,mlp,rnn,stc}-<INCREMENTAL-ID>.log</p> </td> <td> <p>distwalk run log</p> </td> </tr> <tr> <td> <p>distwalk-{lin,mlp,rnn,stc}-<INCREMENTAL-ID>-pred.json</p> </td> <td> <p>Predictive metric data exported from Monasca DB</p> </td> </tr> <tr> <td> <p>distwalk-{lin,mlp,rnn,stc}-<INCREMENTAL-ID>-real.json</p> </td> <td> <p>Actual metric data exported from Monasca DB</p> </td> </tr> <tr> <td> <p>distwalk-{lin,mlp,rnn,stc}-<INCREMENTAL-ID>-times.csv</p> </td> <td> <p>Client-side response time for each request sent during a run</p> </td> </tr> <tr> <td> <p>model_dumps/*</p> </td> <td> <p>Dumps of the models and data scalers used for the validation</p> </td> </tr> <tr> <td> <p>predictor.log</p> </td> <td> <p>monasca-predictor log</p> </td> </tr> <tr> <td> <p>predictor-times.log</p> </td> <td> <p>monasca-predictor` log (timing info only)</p> </td> </tr> <tr> <td> <p>predictor-times-{lin,mlp,rnn}.{csv,log}</p> </td> <td> <p>monasca-predictor log (timing info only, group by predictor)</p> </td> </tr> <tr> <td> <p>super_steep_behavior.csv</p> </td> <td> <p>Dataset used to train MLP and RNN models</p> </td> </tr> <tr> <td> <p>test_behavior_02_distwalk-6t_last100.dat</p> </td> <td> <p>distwalk load trace</p> </td> </tr> <tr> <td> <p>ubuntu-20.04-min-distwalk.img</p> </td> <td> <p>Image used to create Nova instances for the scaling group</p> </td> </tr> </tbody> </table> <p>* <em>contact author</em></p>
Replication Archive for Telework, Childcare, and Mothers' Labor Supply
<p>This package contains a main Stata .do file for the results presented in the paper titled, "Telework, Childcare, and Mothers' Labor Supply." It also includes the main dataset and two auxiliary datasets used to generate the estimates produced in the paper. The most updated version of the working paper is available here: <a href="https://doi.org/10.21034/iwp.52">https://doi.org/10.21034/iwp.52</a>. </p>
globalbioticinteractions/AEC-DBCNet: Collaborative databasing of North American bee collections within a global informatics network project archive
<p>Data in this archive are from the <em>Collaborative databasing of North American bee collections within a global informatics network project</em>. Data was originally captured using Arthropod Easy Capture software developed at the American Museum of Natural History (AMNH), New York. Project lead investigators are John Ascher (Principal Investigator) and Jerome Rozen (Co-Principal Investigator) at the AMNH, and Douglas Yanega (Principal Investigator), University of California Riverside.</p> <p><strong>Please use this citation for this archive: </strong>John Ascher, Digital Bee Collections Network data archive from the C<em>ollaborative databasing of North American bee collections within a global informatics network project</em>. Version: 08 Mar 2016. https://doi.org/10.5281/zenodo.1436853</p> <p>This project was supported by the National Science Foundation grant <a href="https://nsf.gov/awardsearch/showAward?AWD_ID=0956388">DBI 0956388</a> and <a href="https://nsf.gov/awardsearch/showAward?AWD_ID=0956340">DBI 0956340</a></p> <p><strong>ABSTRACT</strong> Natural history collections contain millions of bee specimens documenting the geographic ranges, temporal occurrence patterns, and floral associations of the 20,000 described bee species. This project will digitize and consolidate specimen records from 10 bee collections across the United States. The investigators will make or verify species identifications, capture full label data, georeference and error-check localities, and upload this information to publicly accessible databases. Web-based tools will be used to capture data across collections efficiently, validate bee and plant names through automated comparison with taxonomic authority files, and synthesize data on species pages with images, digitized literature records, and other information about bees and their host plants. Data will be uploaded to the Global Biodiversity Information Facility and to Discover Life (www.discoverlife.org), a website that features customizable global maps for all global bee species and dynamic identification keys for North American species. To obtain information needed to conserve and manage pollinators, the investigators will work with ecologists to model geographic and temporal trends in bee populations in relation to environmental variables. Bees are the most important pollinators of the approximately 1/3 of crops that require animal pollination. Recent declines in honey bee populations highlight the need to understand better the roles of native bees in agricultural and natural systems. This project will help predict risks to bees and their pollination services from climate change, habitat loss, and other factors. The outreach program Bee Hunt (www.discoverlife.org/bee) will educate the public, including students in underserved communities, about bee diversity and the importance of pollination services. Using digital photography and rigorous research protocols, Bee Hunt will empower people at biological field stations, nature centers, parks, schools, and other sites to collect high-quality data to augment information from specimen records.</p>
Data Archive for "Nonequilibrium Statistical Thermodynamics of Multicomponent Interfaces"
<p>Selected data, including certain simulation output, analysis scripts, and processed data files used for figures.</p>
Replication Archive for "Chaos Before Order: Productivity Patterns in U.S. Manufacturing"
<p>This archive includes the public-use data and STATA code to replicate the analyses based on the public-use data in the referenced paper.</p>
Data archive for paper "Machine Learning Emulation of Urban Land Surface Processes"
<p>This archive contains models, data* (Overview), as well as the Singularity image to optionally rerun experiments described in "<a href="https://doi.org/10.1029/2021MS002744">Machine Learning Emulation of Urban Land Surface Processes</a>".</p> <p><strong>Prerequisites</strong></p> <ul> <li>Linux or macOS with Bash shell.</li> <li><a href="https://sylabs.io/">Singularity</a> (tested with version 3.6.3-1.el8)</li> </ul> <p>Please note that all steps require <a href="https://sylabs.io/">Singularity</a> to be installed on your system. If you are looking for information on how to install or use Singularity, please refer to the <a href="https://sylabs.io/docs">Singularity documentation</a>.</p> <p><strong>Overview</strong></p> <p>A general overview of the repository structure is given below. Due to licensing restrictions analysis and forcing data (*) cannot be included and need to be requested separately (see Initialization). Data derivatives (**) from either analysis or forcing, as well as intermediary data (***), are not included as they can be generated by rerunning experiments (see Usage).</p> <pre><code>. ├── data │ ├── analysis* │ ├── forcing* │ ├── teb │ ├── utils │ ├── wps │ └── wrf ├── hpc ├── models │ ├── teb │ ├── unn │ ├── wps │ └── wrf-unn ├── notebooks ├── outputs │ ├── analysis** │ ├── benchmark*** │ ├── forcing** │ ├── kerastuner*** │ ├── notebooks │ ├── tabular │ ├── teb** │ ├── unn** │ ├── wps*** │ └── wrf ├── paper │ └── figures ├── singularity └── tools </code></pre> <p><strong>Initialization</strong></p> <p>Forcing and analysis data need to be requested separately. The following directories should map to their respective data archives:</p> <ul> <li><code>./data/analysis</code> -> <a href="http://doi.org/10.5281/zenodo.4678387">Grimmond et al. (2013)</a></li> <li><code>./data/forcing</code> -> <a href="http://doi.org/10.5281/zenodo.4679279">Grimmond et al. (2021)</a></li> </ul> <p><strong>Usage</strong></p> <p>To rerun all experiments and reproduce results, run <code>tools/run_all.sh</code> from your command prompt. After completion, all results are saved in the <code>outputs</code> directory. Note that WRF simulations require high CPU time and may take hours or days to complete.</p> <p>Alternatively, if <a href="https://en.wikipedia.org/wiki/Portable_Batch_System">Portable Batch System (PBS)</a> is available on your system, the following helpers may be used instead:</p> <pre><code>qsub hpc/submit_init.pbs qsub hpc/submit_tuner.pbs qsub hpc/submit_unn.pbs qsub hpc/submit_find_median_unn.pbs qsub hpc/submit_wrf.pbs qsub hpc/submit_postprocess.pbs qsub hpc/submit_benchmark.pbs </code></pre> <p>Note that you may need to modify PBS helper scripts to suit your specific environment.</p> <p><strong>Development notes</strong></p> <p>See DEVELOP.md.</p> <p><strong>License</strong></p> <p>The source code developed for this work is licensed under MIT (<code>LICENSE_CODE.txt</code>). For licensing information of third-party software see licenses under the <code>models</code> directory. Data files in this archive, including the initial and boundary condition data from the European Centre for Medium-Range Weather Forecasts (<code>data/wps/ungrib</code>), are licensed under CC BY-NC 4.0 (<code>LICENSE_DATA.txt</code>).</p>
Archiv des Onlinelabors für Digitale Kulturelle Bildung
<p>A collection of qualitative case vignettes illustrating personal experiences in the everyday use of social media. Case vignettes are clustered thematically and were generated between spring 2018 and summer 2020. The archive is provided as a static html website. The dataset as well as the accompanying documentation is in German.</p> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.