Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
867
datasets available to search
ShareScore release 0.9.0
Dataset results
867 results for “Repositories”
Repository of Thoracolumbar Osteo-Ligamentous Spine Patient-Personalised FE Meshes
<p><em><strong>Repository of Thoracolumbar Osteo-Ligamentous Spine Patient-Personalised FE Meshes:</strong></em></p> <ul> <li>42 FE input files (.inp extension, Abaqus software, Simulia) representing patient-personalised thoracolumbar spine hexahedral meshes, including point coordinates, mesh connectivity IDs and element sets. Each input file is almost 99MB (totally 3.86GB), and it includes vertebras and IVDs hexahedral meshes; pelvis, sacrum, and the femoral head triangulated meshes; and ligaments.</li> <li>An excel file: "<a href="../records/10994164/files/Descriptive_List%20(42_P-S_FE_Models).xlsx?download=1">Descriptive_List (42_P-S_FE_Models).xlsx</a>" reporting measured spinopelvic parameters for 42 patient-personalised FE hexahedral meshes. The Excel file includes measured spinopelvic parameters (PI, PT, SS, LL, LL-PI, GT, RPV, RLL, LDI, RSA, TPA, and scoliosis cobb angle), GAP and IVD centric thickness for FE virtual cohort and patient-personalised FE meshes. Model ID in the excel file is correspondent to the model’s name.</li> <li>One png file: "<a href="../records/10994164/files/42_P-S.png?download=1">42_P-S.png</a>" representing the first 10 patient-personalised thoracolumbar spine hexahedral models out of 42 FE models.</li> </ul> <p><em><strong>Notes:</strong></em></p> <p>1- Model number in "<a href="../records/10994164/files/Descriptive_List%20(42_P-S_FE_Models).xlsx?download=1">Descriptive_List (42_P-S_FE_Models).xlsx</a>" is correspondent to the same model number in the 42 stereolithography (stl) files (.stl extension) representing the thoracolumbar spine triangulated meshes (DOI: <a href="https://doi.org/10.5281/zenodo.8108354">10.5281/zenodo.8108354</a>; <a href="../records/8108354/files/42%20Patient-Specific%20stl%20files.rar?download=1">42 Patient-Specific stl files.rar</a>).</p> <p>2-Models are reconstructed thanks to the Statistical Shape Modelling (SSM) and mesh morphing techniques, i.e., spine sagittal geometrical parameters could be measured by a clinical software like sterEOS or Surgimap, then, different shape modes of morphed-mesh SSM tool could be activated in order to obtain spine deformity. Then, pros and cons analyses between patient geometrical data and synthesized geometrical information could exploit FE patient-personalised thoracolumbar osteo-ligamentous spine models.</p> <p>3- FE inp files can be opened by Abaqus 2019 and later. Any other FE software which supports .inp extension also can open the files.</p> <p><em><strong>Developed by: </strong></em>Morteza Rasouligandomani (Ph.D. in biomedical engineering, Pompeu Fabra university, BCN Med-Tech group, DTIC department, Barcelona, Spain).</p> <p>Email contact: jerome.noailly@upf.edu</p>
Open data repository, Knab et al., Prediction of stroke outcome in mice based on non-invasive MRI and behavioral testing
<p><strong>Open data repository, Knab et al., Prediction of stroke outcome in mice based on non-invasive MRI and behavioral testing</strong></p> <p><strong>Latest version of files: repository_v2.0.zip, Behavior Data_v2.0.xlsx and MRI IDs Testing&Replication Cohort.xlsx (please ignore repository.zip)</strong></p> <p>Open data repository Knab et al. Prediction of stroke outcome in mice based on non-invasvive MRI and behavioral testing</p> <p>Open code and documentation of prediction models available via <a href="https://github.com/major-s/mouse-mcao-outcome-predictor">https://github.com/major-s/mouse-mcao-outcome-predictor</a></p> <p><strong>Content:</strong></p> <p>README.txt</p> <p>This information</p> <p><strong>dat</strong></p> <p>Contains MRI data in NIFTI format and secondary data from atlas registration. For documentation of atlas registration files see https://pubmed.ncbi.nlm.nih.gov/28829217/<br>Files used for the manuscript:<br>t2.nii: t2 weighted image acquired 24 h post stroke<br>masklesion.nii: manually delineated lesion<br>x_masklesion.nii: lesion in atlas space<br>ix_ANO.nii: Allen brain atlas in native space (i.e. matching t2.nii)<br>Lesion volume was calculated by volume of voxels unequal 0 in x_masklesion.nii<br>Overlap of regions defined by ix_ANO.nii with masklesion.nii were used for calculating percent damage in each atlas region</p> <p><strong>prediction_models</strong></p> <p>Contains separated training and test data as xlsx and csv files with lesion volumes in cubic mm of the Allen brain atlas space, percent damage per atlas region and behavioral data. The training data was used as input for training prediction models in MATLAB, the results were created using the test data.<br>The files have following sturcture:<br>Column 1: animal ID<br>Columns 2-537: MRI regions (column title corresponds to the region number as used in the Allen common coordinate framework)<br>Column 538: lesion volume<br>Column 539: initial performance (subacute deficit) = mean performance/deficit on days 2-6<br>Column 540: mean performance/deficit on days 2-6 = initial performance (subacute deficit) - this column equals column 539 but has different header which was used to train the residual from initial deficit<br>Column 541: residual performance/deficit<br>Column 542: test or training group<br>Consecutive rows contain data for each animal specified by the animal id</p> <p>The repository also contains all trained models, prediction results for the test data and tables with resulting median absolute error (MedAE) and 5th, 25th, 75th and 95 absolute error quantiles for each model.<br>The model files end with '_models.mat' and contain 50 independently trained models each. Each model version is specified by number 1-50.<br>The result files end with '_test_results.mat' or '_test_results.xlsx', files with MedAE and quantiles end with '_test_errors.xlsx' or '_test_errors.csv. The common part of filenames specifies the used paradigm<br>Folder 'subacute deficit prediction' contains:<br> - initial_performance_from_lesion_volume: prediction of subacute deficit using lesion volume<br> - initial_performance_from_segmented_mri: prediction of subacute deficit using segmented mri<br>Folder 'long-term outcome prediction' contains:<br> - lesion_volume: prediction of residual deficit using lesion volume<br> - segmented_mri: prediction of residual deficit using segmented_mri<br> - initial_performance: prediction of residual deficit using subacute deficit<br>Folder 'mri_inc_oob_imp' contains models trained using increasing number of mri segments sorted according to the out-of-bag importance. The number of used segments is given in the file name. The models, results and errors are separated in subfolders.</p> <p>Files with equal file name and different extension always contain the same data</p> <p><strong>templates</strong><br>Allen atlas, template, brain mask, hemisphere masks, tissue probability masks in NIFTI format including annotations of region IDs and parameter.m file for use in MATLAB toolbox ANTx2<br> </p>
The Minimum Information about a Biosynthetic Gene Cluster (MIBiG) data repository
<p>This dataset was originally published alongside the Minimum Information about a Biosynthetic Gene Cluster (MIBiG) data standard publication(s).</p> <p>It contains JSON files following the MIBiG data standard. Additional information on proteins/genes associated to biosynthetic gene clusters described by MIBiG can be found in the GenBank (gbk) and fasta files.</p> <p>This dataset was uploaded with permission from the corresponding author(s).</p> <p>For more information, see https://mibig.secondarymetabolites.org/.</p>
GAPs Data Repository on Return: Guideline, Data Samples and Codebook
<p><span>The GAPs Data Repository provides a comprehensive overview of available qualitative and quantitative data on national return regimes, now accessible through an advanced web interface at <a href="https://data.returnmigration.eu/" target="_new"><span>https://data.returnmigration.eu/</span></a><span>. </span></span></p> <p><span>This updated guideline outlines the complete process, starting from the initial data collection for the return migration data repository to the development of a comprehensive web-based platform. Through iterative development, participatory approaches, and rigorous quality checks, we have ensured a systematic representation of return migration data at both national and comparative levels.</span></p> <p><span>The Repository organizes data into five main categories, covering diverse aspects and offering a holistic view of return regimes: country profiles, legislation, infrastructure, international cooperation, and descriptive statistics. These categories, further divided into subcategories, are based on insights from a literature review, existing datasets, and empirical data collection from 14 countries. The selection of categories prioritizes relevance for understanding return and readmission policies and practices, data accessibility, reliability, clarity, and comparability. Raw data is meticulously collected by the national experts. </span></p> <p><span>The transition to a web-based interface builds upon the Repository’s original structure, which was initially developed using REDCap </span><span>(Research Electronic Data Capture). It <span> </span>is a secure web application for building and managing online surveys and databases.</span><span>The REDCAP ensures systematic data entries and store them on Uppsala University’s servers while significantly improving accessibility and usability as well as data security. It also enables users to export any or all data from the Project when granted full data export privileges. Data can be exported in various ways and formats, including Microsoft Excel, SAS, Stata, R, or SPSS for analysis. At this stage, the Data Repository design team also converted tailored records of available data into public reports accessible to anyone with a unique URL, without the need to log in to REDCap or obtain permission to access the GAPs Project Data Repository. Public reports can be used to share information with stakeholders or external partners without granting them access to the Project or requiring them to set up a personal account. Currently, all public report links inserted in this report are also available on the Repository’s webpage, allowing users to export original data.<span> </span></span></p> <p><span>This report also includes a detailed codebook to help users understand the structure, variables, and methodologies used in data collection and organization. This addition ensures transparency and provides a comprehensive framework for researchers and practitioners to effectively interpret the data.</span></p> <p><span>The GAPs Data Repository is committed to providing accessible, well-organized, and reliable data by moving to a centralized web platform and incorporating advanced visuals. This Repository aims to contribute inputs for research, policy analysis, and evidence-based decision-making in the return and readmission field.</span></p> <p><span>Explore the GAPs Data Repository at <a href="https://data.returnmigration.eu/" target="_new">https://data.returnmigration.eu/</a>.</span></p>
European Building Vulnerability Data Repository
<p>A repository for the European vulnerability database developed as part of the European Seismic Risk Model 2020 (ESRM20).</p> <p>More information available in the following paper: Crowley et al. (2021) “Open models and software for assessing the vulnerability of the European building stock,” COMPDYN 2021, 8th ECCOMAS Thematic Conference on Computational Methods in Structural Dynamics and Earthquake Engineering, Greece.</p>
European Exposure Model Data Repository
<p>A repository of the exposure data used to develop the ESRM20 exposure models.</p> <p>More information available here: <a href="https://eu-risk.eucentre.it/exposure/">https://eu-risk.eucentre.it/exposure/</a></p>
A Repository of 100+ Years of Measured Soil Freezing Characteristic Curves
<p>The temperature of the soil can be used as a proxy to represent the soil ice content through a soil freezing characteristic curve (SFCC). This mathematical construct relates the soil ice content to a specific temperature for a particular soil. SFCCs depend on many factors including soil properties (e.g., porosity, composition, etc.), soil pore water pressure, dissolved salts, (hysteresis in) freezing/thawing point depression, and degree of saturation, all of which can be site-specific and time varying. SFCCs have been measured using various methods for diverse soils since 1921, and to date this data has not been broadly compared, in part because it has not previously been compiled in a single data set. The dataset presented in this publication includes SFCC data digitized or received from authors, and includes both historic and modern studies.</p>
Materials for Design Open Repository. High Entropy Alloys
<p>The current dataset is composed of a collection of High Entropy Alloys (HEAs). It contains the alloy composition, the number of chemical elements (No), the phase in a simple form (S_Phase), where 4 classes of phases were considered, namely amorphous (AM), intermetallic (IM), solid solution (SS), and solid solution + intermetallic (SS+IM). It contains also a second phase column (Phase), where we added the type of phase present in alloys with SS and repeated the S_Phase entry for the other cases. We have calculated 13 design parameters (see their definition below) used to design HEAs, known as the parametric approach. Finally, a set of columns containing the chemical elements and their corresponding fraction in the alloy is included. This dataset was developed in the framework of the European project ACHIEF for the discovery of novel materials to be used in industrial processes.</p> <ol> <li>Mean atomic radius <em>a</em> (Å) <ul> <li><span class="math-tex">\(a = \displaystyle\sum_{i=1}^{n} c_i r_i\)</span></li> </ul> </li> <li>Atomic size difference δ <ul> <li><span class="math-tex">\(\delta = \sqrt{\displaystyle\sum_{i=1}^{n} c_i \bigg(1 - \dfrac{r_i}{a} \bigg)^2}\)</span></li> </ul> </li> <li>Average melting temperature <em>T<sub>m</sub></em> (K) <ul> <li><span class="math-tex">\(T_m = \displaystyle\sum_{i=1}^{n} c_i T_{mi}\)</span></li> </ul> </li> <li>Average melting temperature standard deviation (K) <ul> <li><span class="math-tex">\(\sigma_{T_m} = \sqrt{\displaystyle\sum_{i=1}^{n} c_i \bigg(1 - \dfrac{T_{mi}}{T_m} \bigg)^2}\)</span></li> </ul> </li> <li>Mixing enthalpy Δ<em>H<sub>mix</sub></em> (kJ/mol) <ul> <li><span class="math-tex">\(\Delta H_{mix} = 4 \displaystyle\sum_{i \neq j} c_i c_j H_{ij}\)</span></li> </ul> </li> <li>Mixing enthalpy standard deviation (kJ/mol) <ul> <li><span class="math-tex">\(\sigma_{\Delta H_{mix}} = \sqrt{\displaystyle\sum_{i \neq j} c_i c_j (H_{ij} - \Delta H_{mix})^2}\)</span></li> </ul> </li> <li>Ideal mixing entropy <em>S<sub>id</sub></em> (<em>R</em>)<strong>*</strong> <ul> <li><span class="math-tex">\(S_{id} = \Delta S_{mix} = -R \displaystyle\sum_{i=1}^{n} c_i \ln c_i\)</span></li> </ul> </li> <li>Electronegativity <em>χ</em> <ul> <li><span class="math-tex">\(\chi = \displaystyle\sum_{i=1}^{n} c_i \chi_i\)</span></li> </ul> </li> <li>Electronegativity difference in a multi-component alloy system <ul> <li><span class="math-tex">\(\Delta\chi = \displaystyle\sqrt{\sum_{i=1}^{n} c_i(\chi_i - \chi)^2}\)</span></li> </ul> </li> <li>Valence electron concentration <em>VEC</em> <ul> <li><span class="math-tex">\(VEC = \displaystyle\sum_{i=1}^{n} c_i \cdot VEC_i\)</span></li> </ul> </li> <li>Valence electron concentration standard deviation <ul> <li><span class="math-tex">\(\sigma_{VEC} = \sqrt{\displaystyle\sum_{i=1}^{n} c_i (VEC_i - VEC)^2}\)</span></li> </ul> </li> <li>Mean bulk modulus <em>K </em>(GPa) <ul> <li><span class="math-tex">\(K = \displaystyle\sum_{i=1}^{n} c_i K_i\)</span></li> </ul> </li> <li>Bulk modulus standard deviation (GPa) <ul> <li><span class="math-tex">\(\sigma_{K} = \sqrt{\displaystyle\sum_{i=1}^{n} c_i (K_i - K)^2}\)</span></li> </ul> </li> <li>Young's modulus <em>E</em> (GPa) <ul> <li><span class="math-tex">\(E = \displaystyle\sum_{i=1}^{n} c_i E_i\)</span></li> </ul> </li> <li>Shear modulus <em>G</em> (GPa) <ul> <li><span class="math-tex">\(G = \displaystyle\sum_{i=1}^{n} c_i G_i\)</span></li> </ul> </li> </ol> <p>where <em>n</em> is the number of components in the alloy system, <em>c<sub>i</sub></em> is the stoichiometric ratio, <em>r<sub>i</sub></em> is the atomic radius, <em>T<sub>mi</sub></em> is the melting temperature, <em>χ<sub>i</sub></em> is the Pauli electronegativity, <em>VEC<sub>i</sub></em> is the valence electron concentration, and <em>K<sub>i</sub></em> is the bulk modulus, <em>E<sub>i</sub></em> is the Young's modulus, and <em>G<sub>i</sub></em> is shear modulus for the <em>i</em>-th component of the alloy. <em>H<sub>ij</sub></em> is the binary mixing enthalpy in the liquid phase, and <em>R</em> is the gas constant.</p> <p><strong>*Note:</strong> the ideal mixing entropy <em>S<sub>id</sub></em> units in the first version of the dataset appear as kJ/mol, but they should be written in terms of the gas constant <em>R</em>, e.g., the compound Ag<sub>2</sub>Al has <em>S<sub>id</sub></em> = 0.636 <em>R</em>, where <em>R</em> = 8.314 J · K<sup>−1</sup> · mol<sup>−1</sup>. The second version the <em>S<sub>id</sub></em> units are corrected and two new features are included.</p>
Data repository for "Interface rotation in Cu/Nb accumulative roll bonded (ARB) nanolaminates"
<p>Data repository for "Interface rotation in Cu/Nb accumulative roll bonded (ARB) nanolaminates"</p> <p>This repository contains raw experimental data for "Interface rotation in Cu/Nb accumulative roll bonded (ARB) nanolaminates" manuscript. Please refer to the manuscript for the data interpretation.</p> <p>The repository structure:</p> <ul> <li><code>CuNbARB-sample-photo.jpg</code> shows a photograph of as-received Cu(63nm)/Nb(63nm) accumulative roll bonded (ARB) nanolaminate sample</li> <li><code>ARB_63nm_DRX</code> contains X-ray diffraction measurements</li> <li><code>RD/TDXFIBmilling</code> folders contain focused ion beam images captured during pillar milling <ul> <li>In <code>RDX/TDX</code>, <code>X</code> refers to pillar number (see Supplementary information.org for the full pillar list). The numbers in the file names inside refer to the corresponding pillars.</li> </ul> </li> <li><code>RD/TDXSEMbefore</code> folders contain scanning electron (SEM) images of the as-fabricated pillars</li> <li><code>RD/TDX-compression</code> folders contain in situ pillar compression data, including some of the SEM images captured before/after the compression, raw load-displacement data (in <code>.hys</code> native Hysitron piconindenter format), load-displacement data exported to raw text (see Supplementary information for examples how to plot load-displacement using the raw text files), SEM videos, SEM videos combined with the load-displacement data, and accelerated videos</li> <li><code>RD/TDX-SEMafter</code> folders contain SEM images of the compressed pillars</li> <li>Supplementary-info folder contains supplementary information</li> </ul> <p>Author: I. Radchenko, W. Zhu, L. Qing, E. Navarro, R. Sahay, P.S. Lee, N. Raghavan, O. Thomas, A.S. Budiman, K. Chen</p>
Rapid Landslide Risk Zoning toward Multi-Slope Units of the Neikuihui Tribe for Preliminary Disaster Management repository
<p> Taiwan features steep terrain and a fragile geology environment accompanied by frequent earthquakes and typhoons annually. Meanwhile, with the booming economy and rapid population growth, activities pivot from metropolises to the Taiwan's suburban and mountain areas. However, for example, the Neikuihui tribe in northern Taiwan evolves landslide disasters during extreme rainfall events. To rapidly examine landslide risk in the tribe area for preliminary disaster management, the well-known principle of Risk, which comprises Hazard, Exposure, and Vulnerability, was carefully adapted to scrutinize 14 slope units around the Neikuihui tribe region. The framework of risk zoning is improved based on the previous quantified findings regarding the inventory of the deep-seated landslides in southern Taiwan. Moreover, the proposed procedures comprehensively assess susceptibility, activity, exposure, and vulnerability of each slope unit. The rapid risk zoning analysis of multi-slope units delivers a sloping unit with a high level of landslide risk, and this slope unit did suffer from landslide disasters in the 2016 typhoon event. This study preliminarily proves that the proposed framework and details of rapid risk zoning can help identify a relatively high-risk slope unit around a tribal region and address pre-countermeasures for disaster management.</p>
Rosalia: An experimental research site to study hydrological processes in a forest catchment - data repository
<p>This repository is a supplement to the paper <strong>Fürst, J., et al. (2021). “Rosalia: an experimental research site to study hydrological processes in a forest catchment.” Earth Syst. Sci. Data 13(8): 4019-4034.</strong></p> <p>Experimental watersheds have a long tradition as research sites in hydrology and have been used as far back as the late 19<sup>th</sup> and early 20<sup>th</sup> century. The University of Natural Resources and Life Sciences Vienna (BOKU) has been operating the experimental research forest site called “Rosalia” with an area of 950 ha since 1875 to support and facilitate research and education. Recently, BOKU researchers from various disciplines extended the “Rosalia” instrumentation towards a full ecological-hydrological experimental watershed. The overall objective is to implement a multi-scale, multi-disciplinary observation system that facilitates the study of water, energy and solute transport processes in the soil-plant-atmosphere continuum.</p> <p>This repository contains the datasets collected by a monitoring network of 4 discharge gauging stations, 7 rain-gauges, together with observations of air and water temperature, relative humidity and conductivity. In four profiles, soil water content and temperature are recorded in different depths. In 2019, additionally a program to collect isotopic data in precipitation and discharge was started. On one site, also Nitrate, TOC and turbidity are monitored. All data collected since 2015, including in total 56 high resolution time series data (10 min sampling interval), are provided to the scientific community.</p>
largeEELproject_computational_repository
<p>Data and code repository for the thesis project: High-throughput spatial transcriptome<br> profling with EEL-FISH: Enhancing sensitivity, spatial coverage, and computational<br> bottlenecks</p>
Microbes go to school - Output repository
<p>This dataset is an output repository for the final report of the Agora project "Microbes go to school" funded by SNF from 2020 to 2022. This project aims at using service-learning to bridge the gap between university and school, and disseminate knowledge in microbiology and biodiversity in the classroom by engaging students as communicators. In this repository, you'll find general content about the project (gallery, course descriptions, and the article we published), pedagogical content (protocols of the activities edited by us, original content produced by the students that was evaluated, and the feedback form that we sent to the teachers to evaluate the students), and outreach content (guide for trainers, newsletters and recipes).</p>
Public WhatsApp groups from the Brazilian online repositories
<p>This repository contains two gzip files with public WhatsApp groups collected from two Brazilian repositories, they are:</p> <ul> <li><a href="https://gruposdezap.com/">Grupos de Zap</a> - db_grupos_whats.json.gz</li> <li><a href="https://gruposwhats.app/">Grupo de Whats</a> - db_grupos_zap.json.gz</li> </ul> <p>The groups were collected on 01/2022.</p> <p>The files contain five properties: </p> <ul> <li><strong>title</strong>: <em>group title</em></li> <li><strong>description</strong>: <em>description of the group informed by the administrator</em></li> <li><strong>created_date</strong>: <em>creation date when the group was registered in the repository</em></li> <li><strong>num_vizualization</strong><em>: times the group was seen on the site - only for the zap groups repository</em></li> <li><strong>category</strong>: <em>group category</em></li> </ul> <p>If you use this dataset cite your paper, please:</p> <ul> <li><a href="https://doi.org/10.1145/3539637.3557056"><em>"Click Here to Join": A Large-Scale Analysis of Topics Discussed by Brazilian Public Groups on WhatsApp</em> </a></li> </ul> <p><em>Daniel Kansaon, Philipe Melo, and Fabrício Benevenuto. 2022. “Click Here to Join”: A Large-Scale Analysis of Topics Discussed by Brazilian Public Groups on WhatsApp. In Brazilian Symposium on Multimedia and Web (WebMedia ’22), November 7–11, 2022, Curitiba, Brazil. ACM, New York, NY, USA, 11 pages. https://doi.org/10.1145/3539637.3557056</em></p>
Data repository for the publication "Economic Interests Cloud Hazard Reductions in the European Regulation of Substances of Very High Concern"
<p>This repository contains the data and scripts associated with the article “Economic Interests Cloud Hazard Reductions in the European Regulation of Substances of Very High Concern“, written by Jessica Coria, Erik Kristiansson and Mikael Gustavsson.</p>
Extracted MSR GitHub Repository URLs
<p>This dataset contains text files of <a href="https://github.com">GitHub</a> URLs pointing to hosted git repositories.</p> <p>These URLs come from mining software repository (MSR) datasets. URLs are built by taking the repository owner's name (OWNER) and it's name (REPO) and appending them to https://github.com/. There is one URL per line. <em>URLs have not been tested for their current availibility</em>. An example URL format is provided below:</p> <pre><code>https://github.com/OWNER/REPO</code></pre> <p> Current URLs are from the following datasets:</p> <ul> <li>libraies.io January 12th, 2020 dataset <ul> <li>Jeremy Katz, "Libraries.io Open Source Repository and Dependency Metadata". Zenodo, Jan. 12, 2020. doi: 10.5281/zenodo.3626071.</li> </ul> </li> <li>RepoReapers/reaper dataset <ul> <li>Munaiah, N., Kroh, S., Cabrey, C. et al. Curating GitHub for engineered software projects. Empir Software Eng 22, 3219–3253 (2017). https://doi.org/10.1007/s10664-017-9512-6</li> </ul> </li> <li>GH Torrent dataset <ul> <li>G. Gousios, “The GHTorent dataset and tool suite,” in <em>Proceedings of the 10th Working Conference on Mining Software Repositories</em>, San Francisco, CA, USA, May 2013, pp. 233–236.</li> </ul> </li> </ul>
3D models (NXS): Towards a spatial data repository for archaeological research in the Romanian Mostiștea Basin and Danube Valley
<p><span>Spatial data are crucial in archaeological research, where orthophotos, digital elevation models, and 3D models are widely used for mapping, documenting, and monitoring archaeological sites. The introduction of affordable and compact unmanned aerial vehicles (UAVs) has significantly advanced the use of UAV-based photogrammetry in the past 20 years. Recently, compact airborne systems have also enabled the capture of thermal, multispectral, and aerial laser scanning data. This study presents the data acquired with different platforms and sensors at Chalcolithic archaeological sites in Romania's Mostiștea Basin and Danube Valley. Since laser scanning and photogrammetry generate large data volumes, data storage and dissemination must also be carefully considered. Based on a thorough study of system performance, data acquisition and processing methods, and data outputs, a workflow for the systematic mapping and documentation of sites has been proposed. Given the experience obtained in the last 5 summer campaigns (2018-2023), 19 sites have been accurately mapped, of which 5 sites are mapped using airborne laser scanning. 18 sites are documented using multispectral photogrammetry, and for 17 sites, interactive image-based 3D models are acquired using true-color photogrammetry. All data are stored on a publicly accessible website for visualization, as well as on an open-data platform for data exchange. For the multispectral data, a raster tile service has been implemented, allowing the use of the data in a GIS environment.</span></p>
3D models (true color, TIF): Towards a spatial data repository for archaeological research in the Romanian Mostiștea Basin and Danube Valley
<p>Spatial data are crucial in archaeological research, where orthophotos, digital elevation models, and 3D models are widely used for mapping, documenting, and monitoring archaeological sites. The introduction of affordable and compact unmanned aerial vehicles (UAVs) has significantly advanced the use of UAV-based photogrammetry in the past 20 years. Recently, compact airborne systems have also enabled the capture of thermal, multispectral, and aerial laser scanning data. This study presents the data acquired with different platforms and sensors at Chalcolithic archaeological sites in Romania's Mostiștea Basin and Danube Valley. Since laser scanning and photogrammetry generate large data volumes, data storage and dissemination must also be carefully considered. Based on a thorough study of system performance, data acquisition and processing methods, and data outputs, a workflow for the systematic mapping and documentation of sites has been proposed. Given the experience obtained in the last 5 summer campaigns (2018-2023), 19 sites have been accurately mapped, of which 5 sites are mapped using airborne laser scanning. 18 sites are documented using multispectral photogrammetry, and for 17 sites, interactive image-based 3D models are acquired using true-color photogrammetry. All data are stored on a publicly accessible website for visualization, as well as on an open-data platform for data exchange. For the multispectral data, a raster tile service has been implemented, allowing the use of the data in a GIS environment.</p>
LLM Research Repository
<p><strong>Overview</strong></p> <p>Welcome to the Large Language Models (LLM) Repository, a curated collection aimed at researchers, practitioners, and enthusiasts in the field of Natural Language Processing (NLP). This repository offers resources related to Large Language Models, including research papers, theses, tools, datasets, courses, open-source models, and benchmarks. </p> <p><strong>1. Research Papers</strong></p> <p>A compilation of seminal and cutting-edge research papers that shape the field of Large Language Models. This section includes:</p> <ul> <li>Foundational Papers: Groundbreaking papers that laid the framework for LLM research.</li> <li>Recent Advances: Latest research examining novel architectures, training techniques, and applications.</li> <li>Survey and Review Articles: Comprehensive surveys and reviews that aggregate findings and offer insightful analysis on various aspects of LLMs.</li> </ul> <p><strong>2. Theses</strong></p> <p>A collection of master's and doctoral theses that focus on various facets of LLMs, providing in-depth explorations of core concepts, novel methodologies, and empirical studies from all over the world.</p> <p><strong>3. Tools</strong></p> <p>A list of tools and libraries essential for working with Large Language Models. This section encompasses:</p> <ul> <li>Development Frameworks: Popular libraries and frameworks.</li> <li>Utilities: Tools for data preprocessing, model deployment, and inference acceleration.</li> </ul> <p><strong>4. Datasets</strong></p> <p>A collection of datasets tailored for training and evaluating Large Language Models. This section includes:</p> <ul> <li>Text Corpora: Large-scale text datasets from diverse domains such as news articles, books, and social media.</li> <li>Annotated Datasets: Datasets with human annotations for tasks such as named entity recognition, sentiment analysis, and machine translation.</li> </ul> <p><strong>5. Courses</strong></p> <p>A list of university courses related to Large Language Models.</p> <p><strong>6. Open Source Models</strong></p> <p>Access to state-of-the-art open-source Large Language Models, allowing you to leverage pre-trained models for various applications. This section includes:</p> <ul> <li>Model Repositories: Links to GitHub repositories and model zoos hosting popular LLMs such as GPT, BERT, T5, and their variants.</li> <li>Pre-trained Models: Ready-to-use models available through platforms like the Hugging Face Model Hub, including detailed usage instructions and licensing information.</li> <li>Customized Implementations: Specialized versions and fine-tuned models tailored for specific tasks or domains.</li> </ul> <p><strong>7. Benchmarks</strong></p> <p>A suite of benchmarks designed to evaluate the performance and robustness of Large Language Models. This section features:</p> <ul> <li>Standard Benchmarks: Widely-accepted benchmarks like GLUE, SuperGLUE, and the LAMBADA dataset.</li> <li>Challenge Sets: Curated datasets that test specific capabilities of models, such as commonsense reasoning, multilingual understanding, or adversarial robustness.</li> </ul>
Exploiting the Greenland volcanic ash repository to date caldera-forming eruptions and widespread isochrons during the Holocene
<p>Polar ice-cores have long been recognised as unrivalled repositories of past volcanic events. Although tephra products from local eruptions tend to dominate these records, improvements in micro-sampling and analytical techniques are uncovering a growing number of cryptotephras erupted from exceptionally distant volcanoes. We present a series of nine Middle Holocene cryptotephra deposits detected within the NGRIP ice-core that originate from five different volcanic regions across the Northern Hemisphere (Alaska, Cascades, Iceland, Japan, Kamchatka). Unique compositional signatures are employed to identify ash from three large caldera-forming events in Kamchatka (KS<sub>2 </sub>from Ksudach), the Cascades (Mazama) and North East Japan (Mashu), along with ash from the Hekla 4 eruption in Iceland. High-precision ice-core ages (adopting a 1950 datum for the GICC05 timescale assigned to the Greenland ice cores) are derived for each eruption: Hekla 4 (4325 ± 8 a b1.95k), KS<sub>2</sub> (7089 ± 26 a b1.95k), Mashu (i-f) (7473 ± 33 a b1.95k) and Mazama (7562 ± 35 a b1.95k), all of which can be employed as chronological fix-points in other proxy records where these deposits are also preserved. Four further cryptotephra deposits and one macro-deposit (in the GRIP ice core) are also identified and traced to sources in Iceland and Alaska. The cryptotephra originating from Alaska is correlated to a deposit identified in lake records from the Kenai Peninsula, thought to originate from Redoubt Volcano. The remaining four deposits are typical of the products of Katla, Grímsvötn and Veiðivötn in Iceland. This ensemble of mid-Holocene tephra deposits highlights the pivotal position of the Greenland ice-sheet and its ice-cores to capture deposition from the convergence of several far-travelled ash clouds. Precise age estimates derived from the annually resolved ice-core record greatly enhances the value of these tephra isochrons.</p> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.