Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
867
datasets available to search
ShareScore release 0.9.0
Dataset results
867 results for “repositories”
A dataset of metadata for UK academic institutional repositories, including a census of research software contained.
<p>A dataset of metadata for UK academic institutional repositories, including a census of research software contained.</p> <table> <tbody> <tr> <td><strong>URL</strong></td> <td>The OAI url</td> </tr> <tr> <td><strong>id</strong></td> <td>CORE Identifier</td> </tr> <tr> <td><strong>openDoarId</strong></td> <td>Open DOAR identifier</td> </tr> <tr> <td><strong>name</strong></td> <td>Name of repository</td> </tr> <tr> <td><strong>Russell_member</strong></td> <td>If the university is a member of the Russell Group of research intensive universities</td> </tr> <tr> <td><strong>RSE_group</strong></td> <td>If an RSE group is present (based on Soc of RSE data)</td> </tr> <tr> <td><strong>email</strong></td> <td>Redacted</td> </tr> <tr> <td><strong>uri</strong></td> <td>Not used</td> </tr> <tr> <td><strong>uni_sld</strong></td> <td>Second level domain (the part of the url between . And .ac.uk</td> </tr> <tr> <td><strong>homepageUrl</strong></td> <td>University website</td> </tr> <tr> <td><strong>source</strong></td> <td>Not used</td> </tr> <tr> <td><strong>ris_software</strong></td> <td>the Research Information System software used</td> </tr> <tr> <td><strong>ris_software_enum</strong></td> <td>Resolve ris_software into similar types (e.g. Eprints 3, EPrints3.3.16 both equal eprints)</td> </tr> <tr> <td><strong>metadataFormat</strong></td> <td>the protocol used for metadata</td> </tr> <tr> <td><strong>createdDate</strong></td> <td>Repository creation date</td> </tr> <tr> <td><strong>location</strong></td> <td>location of university</td> </tr> <tr> <td><strong>logo</strong></td> <td>University logo (resolves in error)</td> </tr> <tr> <td><strong>type</strong></td> <td>Only = Repository for this dataset. Can be = journal etc.</td> </tr> <tr> <td><strong>stats</strong></td> <td>Not used</td> </tr> <tr> <td><strong>contains_software_set</strong></td> <td>Whether the OAI-PMH software set is present in the repository.</td> </tr> <tr> <td><strong>Num_sw_records</strong></td> <td>The response of the OAI-PMH query for software (erroneous as discussed in paper)</td> </tr> <tr> <td><strong>Error</strong></td> <td>The category of error returned by the experiment’s OAI-PMH queries (see paper)</td> </tr> <tr> <td><strong>Manual_Num_sw_records</strong></td> <td>The true amount of software contained in the repository as found by a manual exhaustive search of each university website</td> </tr> <tr> <td><strong>Category</strong></td> <td>Whether the repository (a) contains software; (b) can contain software, but doesn’t yet; (c) has no separate type of research output called software or similar</td> </tr> </tbody> </table> <p> </p>
X-COBOL: A Dataset of COBOL Repositories
<p>Despite being proposed as early as 1959, COBOL (Common Business-Oriented Language) still predominantly acts as an integral part of the majority of operations of several financial, banking, and governmental organizations. To support the inevitable modernization and maintenance of legacy systems written in COBOL, it is essential for organizations, researchers, and developers to understand the nature and source code of COBOL programs. However, to the best of our knowledge, we are unaware of any dataset that provides data on COBOL software projects, motivating the need for the dataset. Thus, to aid empirical research on comprehending COBOL in open-source repositories, we constructed a dataset of 84 COBOL repositories mined from GitHub, containing rich metadata on the development cycle of the projects. We envision that researchers can utilize our dataset to study COBOL projects' evolution, code properties and develop tools to support their development. Our dataset also provides 1255 COBOL files present inside the mined repositories.</p>
Project's repository for: Co-immersion in Audio Augmented Virtuality: the Case Study of a Static and Approximated Late Reverberation Algorithm
<p>Repository of the VR scene and the audio data used for the experiment reported in the publication <a href="https://ieeexplore.ieee.org/document/10269056" target="_blank" rel="noopener">available in Open Access</a>:</p> <blockquote> <p>Davide Fantini, Giorgio Presti, Michele Geronazzo, Riccardo Bona, Alessandro Giuseppe Privitera and Federico Avanzini (2023) "Co-immersion in Audio Augmented Virtuality: the Case Study of a Static and Approximated Late Reverberation Algorithm" in <em>IEEE Transactions on Visualization and Computer Graphics (ISMAR special issue)</em></p> </blockquote> <p>The file <a href="../api/files/06c374e2-c54d-40f1-ae23-c4c7afbfba5b/README.md">README.md</a> includes some instructions to use the data in this repository.</p> <p> </p> <p><strong>AUDIO</strong></p> <p>The file <a href="../api/files/06c374e2-c54d-40f1-ae23-c4c7afbfba5b/audio.zip">audio.zip</a> includes the Reaper's projects and audio files used in the experiment to provide the auditory stimuli (simultaneous reverberated speeches) to the participants. Each subfolder corresponds to a different Virtual Acoustics Environment (VAE):</p> <ul> <li><<em>LivingRoom</em>|<em>MARCo</em>|<em>METU</em>> <ul> <li><<em>Living Room</em>|<em>MARCo</em>|<em>METU</em>><em>.rpp</em>: Reaper's project for the VAE</li> <li><em>Bin</em>: folder including the speech data convolved with the late reverberation part of the reverb condition \(B\) for each source position in the VAE</li> <li><em>Freeverb</em>: folder including the speech data convolved with the late reverberation part of the reverb condition \(F_\text{d}\) for each source position in the VAE</li> <li><em>HOA</em>: <ul> <li><em>ER</em>: folder including the speech data convolved with the early reflections part (HOA in A-format) of the reference reverb condition \(H\) for each source position in the VAE</li> <li><em>Ref</em>: folder including the speech data convolved with the entire reference reverb condition \(H\) (HOA in A-format) for each source position in the VAE</li> </ul> </li> </ul> </li> </ul> <p>The reverberated speech data in the <a href="../api/files/06c374e2-c54d-40f1-ae23-c4c7afbfba5b/audio.zip">audio.zip</a> file are obtained using third-party datasets:</p> <ul> <li>The anechoic speech data are retrieved from four speakers (F2, F5, M3, M6) of the <a href="https://doi.org/10.5281/zenodo.6257551">ACE challenge corpus</a></li> <li>The Room Impulse Responses (RIR) in High-Order Ambisonics (HOA) format used to reverberate the speeches are retrieved from: <ul> <li><a href="https://doi.org/10.5281/zenodo.5747753">Living Room</a></li> <li><a href="https://doi.org/10.5281/zenodo.3477602">Concert hall (MARCo)</a></li> <li><a href="https://doi.org/10.5281/zenodo.2635758">Classroom (METU)</a></li> </ul> </li> </ul> <p> </p> <p><strong>VR SCENE</strong></p> <p>The file <a href="../api/files/06c374e2-c54d-40f1-ae23-c4c7afbfba5b/VRscene.zip">VRscene.zip</a> includes the Virtual Reality (VR) scene provided to the participants during the experiment via an Oculus Quest 2. This file includes two subfolders:</p> <ul> <li><em>UDPServer</em>: C# code for the UDP server used for sending the OSC messages for head tracking <ul> <li><em>external/SharpOSC.dll</em>: external library (<a href="https://github.com/ValdemarOrn/SharpOSC">SharpOSC</a>) used to interact with the OSC protocol</li> </ul> </li> <li><em>VR_Headtracking</em>: folder including the Unity project with the VR scene</li> </ul> <p> </p>
Dependency networks of PyPI, npm, CRAN and Bioconductor repositories
<p>A set of datasets with a list of dependencies from the software repositories PyPI, npm, CRAN and Bioconductor.</p>
"Canvas BM" Digital Business Model Template Repository
<p>The Digital Business Model Template Repository consists of 265 unique one-page, diverse business model compositions, constructed as variants or adaptations based on the reference Business Model Canvas (BMC) created by A. Osterwalder. Business Model Canvases are templates composed of key blocks (elements) of the business model, which are interconnected and ready to be filled with content. The process of acquiring templates through quantitative, and then qualitative research, was conducted from November 2020 to October 2022. Each identified template was verified for compliance with licensing rights. The thematic repository should be treated as a business guide, aiding in the selection of suitable tools for designing business models for specific organizations, as well as in the process of creating, analyzing, and modifying business model templates. The Digital Business Model Template Repository can be useful in both a scientific, research, didactic, and individual context, and can also be beneficial in business practice.</p>
Data Repository Accompanying "Controllable single Cooper pair splitting in hybrid quantum dot systems"
<p>Code and datasets associated with the manuscript " Controllable single Cooper pair splitting in hybrid quantum dot systems". With the code and data included here, all necessary fits and analysis can be conducted to produce the figures given in the manuscript and its supplementary material. The only exception is that we include the results of the quantum dot stability diagram simulation, however this simulation involves no new physics and the procedure is described in detail in the manuscript's supplementary information.</p>
List of research data repositories that were shut down
<p>This dataset aggregates information about 191 research data repositories that were shut down. The data collection was based on the registry of research data repositories re3data and a comprehensive content analysis of repository websites and related materials. Documented in the dataset are the period in which a repository was active, the risks resulting in its shutdown, and the repositories taking over custody of the data after.</p>
"Outgassing Composition of the Murchison Meteorite: Implications for Volatile Depletion of Planetesimals and Interior-Atmosphere Connections for Terrestrial Exoplanets" Data Repository
<p>This repository contains the data files, analysis Jupyter notebooks and figures from Thompson et al. 2023 "Outgassing Composition of the Murchison Meteorite: Implications for Volatile Depletion of Planetesimals and Interior-Atmosphere Connections for Terrestrial Exoplanets"</p>
Repository of IVD Patient-Specific FE Models
<p>Free repository of 169 PP FE models of the IVD. Resulting cohort from a morphing process as a free-access repository to further empower the scientific community. This initiative underlines our commitment to promoting standardization and facilitating a more comprehensive understanding of the mechanisms underlying IVD degeneration.</p>
Return on Investment Metrics for Data Repositories in Earth and Environmental Sciences
Despite a growing recognition of the importance of data to the economy and to science, investment in repositories to manage and disseminate that data in easily accessible and understandable ways is scarce. Keeping repository services active and up-to-date for a long time period is difficult due to this funding situation. As a result, repositories must continually provide proof of their value, their Return on Investment (ROI) to their sponsors; yet doing so has always been difficult, problematic and not always successful. In this work, an analysis of approaches for assessing the ROI of several scientific data repositories has identified various techniques that repositories use to report on the impact and value of their data products and services. A survey of selected repositories rated the set of metrics identified and rated each by its importance as well as the ease with which the metric could be measured. The discussion is broken down into considerations for calculating costs, perceived value of repositories and suggested metrics that would allow a repository to calculate an ROI. The authors, representatives of environmental data repositories, concluded that easily obtainable data use metrics, such as data downloads, etc., have limited value while more informative analyses would require additional resources.
Luquillo Critical Zone Observatory (LCZO) Data repository on HydroShare
Data archive for the Luquillo Critical Zone Observatory (LCZO), Puerto Rico. Active from 2009 to 2020. The archive is here: https://www.hydroshare.org/group/144 LCZO focuses on how Critical Zone processes and water balances differ in tropical landscapes with contrasting bedrock but similar climatic and environmental histories. Our infrastructure, sampling strategy, and data management system include watersheds underlain by granodiorite (GD) and volcaniclastic (VC) bedrock in the natural laboratory of the Luquillo Mountains, Puerto Rico. LCZO is one of ten NSF-supported critical zone observatories. The archive is here: https://www.hydroshare.org/group/144 An additional dataset on Forest and Ground Cover classification, DEM, and Beryllium-10 data for the Luquillo Experimental Forest has been published here: https://doi.org/10.4211/hs.0181f4621d184a89a625ee16dd9858a6 An additional datasets for LCZO -- Stream Water Chemistry, Meteorology -- Environmental Monitoring -- Luquillo Mountains -- (2014-Ongoing) https://www.hydroshare.org/resource/b05e1645887f4122a284719bb6cb70dc/ Support for this work was provided by grants BSR-8811902, DEB-9411973, DEB-9705814 , DEB-0080538, DEB-0218039 , DEB-0620910 , DEB-1239764, DEB-1546686, and DEB-1831952 from the National Science Foundation to the University of Puerto Rico as part of the Luquillo Long-Term Ecological Research Program. Additional support provided by the University of Puerto Rico and the International Institute of Tropical Forestry, USDA Forest Service.
Wikimedia Links to UK Repositories (snapshot)
<p>This is a snapshot (08/12/2019) of links from Wikipedia to UK repositories using the "insource" parameter on this special page to search across Wikipedia <a href="https://en.wikipedia.org/w/index.php?sort=relevance&search=">https://en.wikipedia.org/w/index.php?sort=relevance&search=</a></p> <p>It covers institutional open access and data repositories. It also includes non-institutional and subject specific data repositories including figshare, zenodo, datadryad</p> <p>N.B. There are over 150 HEIs in the UK and for practical reasons it is only complete across the Russell Group. Other repositories were partially crowdsourced via the UKCoRR mailing list which can be updated via this Google sheet for future iterations of the dataset <a href="https://docs.google.com/spreadsheets/d/1rJ3Se0y6dJJW43SLeFxaKlSsdrQN5p4lQLK_rMhxAZQ/edit#gid=0">https://docs.google.com/spreadsheets/d/1rJ3Se0y6dJJW43SLeFxaKlSsdrQN5p4lQLK_rMhxAZQ/edit#gid=0</a></p>
Repository: Quantifying environmental impacts of primary aluminum ingot production and consumption: A trade-linked multilevel life cycle assessment
<p>This repository contains the input data, codes and results of the model developed in the paper "Quantifying environmental impacts of primary aluminum ingot production and consumption: A trade-linked multilevel life cycle assessment" published in the Journal of Industrial Ecology (2020) by Alexandre Milovanoff, I. Daniel Posen, Heather L. MacLean.</p>
Dataset of merge conflicts collected from GitHub repositories
<p>Within each nested folder of the archive you will find files A,O,B and M. They each represent a conflict where file O was altered in two different ways, resulting in A and B. Finally, a developer solved the merge conflict committing M as the solution.</p> <p>We have selected these by manually searching for a programming language on GitHub and selecting those repositories that had a large number of forks, commits and contributors.</p>
Data repository for the paper "Tectonics and seismicity in the Northern Apennines driven by slab retreat and lithospheric delamination"
<p>Output data from a numerical modeling study analyzing the Tectonics and seismicity of the Northern Apennines in relation to the geodynamic mechanism (slab retreat and crustal delamination) suggested to be driving the orogenic system.</p> <p>Understanding how long-term subduction dynamics relates to short-term seismicity and crustal tectonics is a challenging but crucial topic in seismotectonics. We attempt to address this issue in the context of the Northern Apennines orogenic belt, which displays characteristic tectonic and seismogenic behaviors on a wide range of spatiotemporal scales. We use a visco-elasto-plastic seismo-thermo-mechanical (STM) modeling approach with a realistic 2D setup based on available geological and geophysical data. In accordance with regional geodynamics, subduction dynamics and seismicity are simulated together, driven solely by slab pull. Our numerical experiments suggest that lower crustal rheology and lithospheric mantle temperatures modulate the crustal tectonics of the Northern Apennines. Results indicate that the observed spatial distribution of the upper crustal tectonic regimes requires buoyant and highly ductile material beneath the suture zone. This allows protrusion of the asthenosphere in the lower crust, lithospheric delamination, and slab retreat. The resulting horizontal velocities and principal stress axis orientations agree with observations, suggesting that slab delamination and retreat are compatible with regional deformation. Our simulations successfully reproduce the presence of seismicity in the thrust front and on normal faults in the interior of the range. Slab temperatures and lithospheric mantle stiffness distinctly affect the cumulative seismic moment release and the spatial distribution of upper crustal earthquakes. The properties of deep, sub-crustal material are thus shown to influence model shallow seismicity, even though the upper crust is largely mechanically decoupled from the lithospheric mantle. Our simulations therefore highlight the important effect of deep crustal rheologies and self-driven subduction dynamics in controlling the shallow, brittle deformation and related seismicity during an ongoing orogeny.</p> <p>The repository consists of the following: 1) the executable code for running the model (i2_istm and in2_istm, the latter of which is used to initialise the model); 2) the setting files for which timesteps to output (mode.t3c and mode_istm.t3c, the latter of which is for the short-term phase of the model), the model setting files (init_istm.t3c), rock type and temperature setup images (prf_app.tif and tfin.tif, respectively); and 3) the output quantities in the model for the last timestep in HDF5 format (app400.gzip.h5), the list of ruptured markers (pick_events_app.txt) and GPS-station-like markers at the surface (eachdt_gpsmarker_app.txt), and the time limits used for computing average velocities from the GPS marker positions (timelims.mat).</p> <p>The files used for the figures in the paper relate to the reference model and 9 other models: 2 models with different rheology for the Adriatic lower crust, 2 models with different temperatures in the mantle, and 5 models with different shear modulus in the Adriatic lithospheric mantle. The two models with different lower crust rheology (granulite and plagioclase) were not run in short-term mode and therefore no GPS-like or ruptured markers logs are available for them. Descriptive prefixes are used to identify which model each file refers to. The rock type setup is common to all models included here. The reference temperature setup is also used for the models with different shear modulus in the slab and the model with granulite lower crust rheology. The model with plagioclase lower crust rheology has a different temperature setup with a hotter lower crust, as mentioned in the paper; it is not a simple exploration of the effect of rheology, but an attempt to get the lower crust to be very ductile through a combination of a ductile rheology (but less so than in the reference model) and high temperatures.</p> <p>For information about the modeling code, setup, results, and interpretation, please refer to the paper. This repository will be updated with the final paper information after publication.</p>
How repositories can contribute their FAIR share
<p>Findable, accessible, interoperable and reusable (FAIR) data are an increasingly important aspect of open scholarship. Increasing the production and use of FAIR data requires a wide range of stakeholders across the research ecosystem to actively play their parts. FAIRsFAIR – Fostering Fair Data Practices in Europe – aims to supply practical solutions for applying the FAIR data principles throughout the research data life cycle, which can be included in the strategies that research organisations and implemented to enable a FAIR data culture.</p> <p>This 90 minute virtual workshop focuses on the work FAIRsFAIR carries out in collaboration with repositories to enable them to play their role in helping to make and keep data FAIR over time. It covers an early draft of a transition support programme for repositories wishing to improve their capacity to support FAIR data production and use. Following an overview of the draft programme, attendees review and discuss the draft support programme, consider how it might be applied within their own repositories, and how they can support and promote relevant aspects of the programme within their institutions and the wider community.</p> <p>This record is a virtual workshop recording.</p> <p>Slides: Herterich, Patricia, & Davidson, Joy. (2020, June). How repositories can contribute their FAIR share. Zenodo. <a href="http://doi.org/10.5281/zenodo.3871523">http://doi.org/10.5281/zenodo.3871523</a></p>
Open data repository, Boehm-Sturm et al., Phenotyping placental oxygenation in Lgals1 deficient mice using 19F MRI
<p>Open data repository of journal article "Phenotyping placental oxygenation in Lgals1 deficient mice using <sup>19</sup>F MRI"</p>
The Collection Management System Collection - Crowd-sourcing a list of digital repository options
<p><strong>The Collection Management System Collection - Crowd-sourcing a list of digital repository options</strong></p> <p>This dataset contains a list of digital repository options for collection management systems. It has been started and complited by Ashley Blewer.<br> The data set contains:</p> <ul> <li>a PDF capture of the blog describing motivation and background, columns of the spreadsheet and further resources; originally published at https://bits.ashleyblewer.com/blog/2017/08/09/collection-management-system-collection/</li> <li>The dataset / spreadsheet of The Collection Management System Collection, originally published at https://docs.google.com/spreadsheets/d/1cXOug3qM0pNNeD_wssiVEv9c0W1Y5I1VDTnSPTk7fb4/<br> The data was exported from the google spreadsheet on November 14th 2020 into the following formats: <ul> <li>PDF</li> <li>XLSX</li> <li>CSV</li> <li>TSV</li> </ul> </li> </ul> <p>The list contains basic information, administration considerations, interface considerations, technical considerations and social considerations for 70 different repository systems.</p>
Piburgersee core meta data repository for the publication "Seismic control of large prehistoric rockslides in the Eastern Alps"
<p>This dataset comprises the core meta data of Plansee, which is the basis for the publication Oswald et al. "Seismic control of large prehistoric rockslides in the Eastern Alps".</p> <p>The core meta data belongs to a 8m long sediment core composed of 12 individual core sections (see Plansee_core_data.xlsx). For each individual core section the core image (_coreimage.jpg), the CT data (_CT.rar), XRF data, (_XRF.txt) and multi-sensor core logging data (_MSCL.csv) are provided.</p>
Plansee seismic and core meta data repository for the publication "Seismic control of large prehistoric rockslides in the Eastern Alps"
<p>This dataset comprises the raw seismic data and core meta data of Plansee, which is the basis for the publication Oswald et al. "Seismic control of large prehistoric rockslides in the Eastern Alps".</p> <p>Seismic profiles are provided as .SGY files (Plansee_seismics_SGYfiles.rar)</p> <p>The core meta data belongs to a 7m long sediment core composed of 10 individual core sections (see Plansee_core_data.xlsx). For each individual core section the core image (_coreimage.jpg), the CT data (_CT.rar) and multi-sensor core logging data (_MSCL.csv) are provided.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.