Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
3,118
datasets available to search
ShareScore release 0.7.1
Dataset results
3,118 results for “resources”
Resources for The Fundamental Limit of Jet Tagging
<p>Resources related to The Fundamental Limit of Jet Tagging (arxiv:2411.02628)</p> <p>Includes:</p> <p>-In total 11,200,0000 qcd and 11,200,0000 top jets generated from corresponding transformer-based models (Tmodels) trained using the JetClass Dataset.</p> <p>-The trained Tmodels.</p> <p>-The LLR predictions from the Tmodels, for different number of constituents.</p> <p>-The predictions corresponding to the classifier-based jet taggers.</p>
GeoERA RESOURCE H3O-PLUS data set which contains hydraulic properties of prime aquifers and aquitards in the Dutch-Flemish-German cross-border area
<p>Dataset which contains information about hydraulic properties of harmonized hydrogeological units in the Dutch-Flemish-German cross-border region which was compiled in the GeoERA RESOURCE project under WP3 H3O-PLUS. The harmonization of the 3D geometry of the cross-border hydrogeological units in the H3O projects constituted a major step towards a common hydrogeological dataset of the Roer Valley Graben and thus the harmonization of groundwater flow models. The database that was compiled provides the characterization of these hydrogeological units with respect to their hydraulic properties, primarily their hydraulic conductivity.<br> The associated report and appendices describe the database of hydraulic properties of aquifers and aquitards based on common criteria. Attention is also given to the characterization of hydraulic properties of faults.</p>
GeoERA RESOURCE CHAKA data set which contains time series of precipitation and discharge of springs in the CHAKA pilot areas (D5.5)
<p>Dataset which contains time series of precipitation and discharge of springs in the pilot areas of the CHAKA work package of the GeoERA RESOURCE project. The file contains precipitation and spring discharge data of 16 pilot areas in the Karst & Chalk work package. A description of the application of the dataset for the characterisation of the typology of karst systems in given in the D5.3 deliverable of GeoERA RESOURCE of which the pdf is provided. Further information about the CHAKA results can be assessed though the webservices of the European Geological Data Infrastructure (EGDI). </p>
Online Resources for Strullu-Derrien et al - The 330–320 Million-Year-Old Tranchée des Malécots (Chaudefonds-sur-Layon, South of the Armorican Massif, France): a Rare Geoheritage Site Containing In Situ Palaeobotanical Remains
<p>This repository contains the following files associate with "The 330–320 Million-Year-Old Tranchée des Malécots (Chaudefonds-sur-Layon, South of the Armorican Massif, France): a Rare Geoheritage Site Containing In Situ Palaeobotanical Remains" by Christine Strullu-Derrien, Alan RT Spencer, Christopher J Cleal and Victor O. Leshyk.</p> <p><strong>Online Resource 1</strong> Model data as a .zip archive (301.5MB) containing .obj/.mtl and texture files for each 3D reconstruction (Models #1-4, whole site reconstruction, detailed reconstruction of the trench, and model of the mine site).</p> <p><strong>Online Resource 2</strong> Video animation showing whole site 3D model (.mp4 | 37.7MB), with quick fly-through of the Tranchée des Malécots showing exposed rock and bedding of the SW wall.</p> <p><strong>Online Resource 3</strong> Video animation showing 3D Model #1 (.mp4 | 35.1MB).</p> <p><strong>Online Resource 4</strong> Video animation showing 3D Model #2 (.mp4 | 83.5MB).</p> <p><strong>Online Resource 5</strong> Video animation showing 3D Model #3 (.mp4 | 45.3MB).</p> <p><strong>Online Resource 6</strong> Video animation showing 3D Model #4 (.mp4 | 65.8MB).</p> <p><strong>Online Resource 7</strong> Video animation showing 3D model of the Malécots mine headframe (.mp4 | 14.9.0MB).</p>
SARS-CoV-2 genomics resources for Galaxy
<p>Reference and custom annotation data expected as input by Galaxy SARS-CoV-2 variation analysis workflows developed by covid19.galaxyproject.org</p>
Aphidinae comparative genomics resource
<p>Here we provide early access to 18 new genome assemblies, including 8 assembled to chromosome-scale, for aphids from the subfamily Aphidinae. For consistency and to aid comparative analysis, all genomes have been annotated using the same repeat masking and RNA-seq-based gene prediction pipeline. Using this pipeline we also provide new annotations for three previously published genome assemblies.</p> <p>The genome assemblies and annotations are made freely available without restriction, we only request that this Zenodo resource is cited when using the data. Raw sequence data upload to NCBI is underway and full details of all accessions will be given in an updated version of this resource. Manuscripts are in preparation describing the individual genome assemblies in detail and larger comparative genome analyses and we will update this resource with additional citation information as papers are published.</p> <p>Full details of all genome assemblies and annotations included in this release are given in the attached "Data_Description.pdf" document. </p> <p><strong>Aphid species included in this release (bold type = chromosome-scale assembly):</strong></p> <p><em><strong>Aphis fabae</strong><br> Aphis glycines </em>(updated annotation)<br> <em><strong>Aphis gossypii</strong><br> Aphis thalictri<br> Aphis rumicis<br> Brachycaudus cardui<br> Brachycaudus helichrysi<br> Brachycaudus klugkisti<br> <strong>Brevicoryne brassicae</strong><br> Diuraphis noxia<br> <strong>Macrosiphum albifrons</strong><br> Metopolophium dirhodum<br> Myzus cerasi </em>(updated annotation)<br> <em>Myzus ligustri<br> Myzus lythri<br> Myzus varians<br> Pentalonia nigronervosa </em>(updated annotation)<br> <em><strong>Phorodon humuli</strong><br> <strong>Rhopalosiphum padi<br> Sitobion avenae<br> Sitobion miscanthi</strong></em></p>
Bayesian Online Learning for Energy-Aware Resource Orchestration in Virtualized RANs - Dataset
<p>Dataset providing a set of measurement of performance and power consumpetion of a virtualized Base Station (srseNB).</p>
SSHOC - National Gallery - Raphael Research Resource CIDOC CRM Mapped Dataset
<p>In 2007 the <a href="https://cima.ng-london.org.uk/documentation">Raphael Research Resource</a> project began to examine how complex conservation, scientific and art historical research could be combined in a flexible digital form. Exploring the presentation of interrelated high resolution images and text, along with how the data could be stored in relation to an event driven ontology in the form of <a href="http://www.w3.org/TR/rdf-concepts/">RDF triples</a>. The original <a href="https://cima.ng-london.org.uk/documentation">main user interface</a> is still live, In 2021/21 as part of the <a href="https://www.sshopencloud.eu/">SSHOC Project</a> the raw data stored within the system was mapped to the <a href="https://www.cidoc-crm.org/">CIDOC CRM</a> using a custom set of Python scripts (<a href="https://doi.org/10.5281/zenodo.6461654">https://doi.org/10.5281/zenodo.6461654</a>). The SSHOC work aimed to make this data more <a href="https://www.go-fair.org/fair-principles/">FAIR</a> so in addition to mapping it to a standard ontology, to increase Interoperability, it has also been made available in the form of <a href="http://en.wikipedia.org/wiki/Linked_Data">open linkable data</a> combined with a <a href="http://en.wikipedia.org/wiki/SPARQL">SPARQL</a> end-point. This live data presentation can been found <a href="https://rdf.ng-london.org.uk/sshoc/">Here</a>.</p> <p>This deposit contains the CIDOC-CRM mapped data formatted in XML and an example model diagram representing some of the key relationships covered in the data-set.</p>
Inventory of tools and resources for crop diversification available for stakeholders
<p>The aim of the database is to give an overview of existing resources, tools and methods to promote crop diversification strategies (rotation, multiple cropping, intercropping) at different levels (including the value chain and territory levels). This version contains 143 resources.</p> <p>Each resource is described with a set of criteria: strategies used / described in the resource, purpose of the resource (what is an end-user doing with the resource), expected performances, area of validity, context of use, but also characteristics for use (cost, training, required time to collect data…).</p> <p>A toolbox was also designed to support end-users to navigate among this database and aims to help different type of end-users to identify interesting and adapted resources to foster crop diversification.</p> <p></p>
Map of Tigray's mineral resources (north Ethiopia) - Canadian and other international mining licences
<p>In Tigray, artisanal mining of gold in the low-lying areas with outcropping Precambrian rocks is one of the major off-farm income sources. The 17<sup>th</sup> C. Portuguese traveller Barradas had already mentioned gold production in Tembien. Rural youth seasonally migrate to inhospitable lowlands and gorges such as the largely uninhabited Weri’i River valley, to search for placer gold, washed out from weathered gold-containing quartz veins within the meta-sediments and meta-volcanics. In recent decades, large-scale gold exploration and mining of gold deposits has been carried out in various parts of Tigray by local (such as the Ezana Mining Development P.L.C.) and several foreign exploration companies particularly from Canada. Recently, The Ethiopia Cable exposed links between big Canadian mining interests and a renewed PR campaign (involving Canadian professor and lobbyist Ann Fitz-Gerald) to whitewash the Ethiopian and Eritrean governments. Earlier on, it had already been suggested that one of the reasons for the Canadian government being very late in officially addressing the atrocities in the ongoing Tigray war, might be related to the country’s mining interests in Tigray.</p> <p>Here we contextualise Tigray’s gold and base metal resources, and present a map of active and applied mineral exploration and mining licenses in Tigray. The largest exploration license areas are concessions of Canadian companies, followed by the U.S. and the United Kingdom.</p>
WaterGAP2.2d model derived Potential evapotranspiration and Renewable water resources variables with standard and modified PET calculation methods
<p>This data set is produced as a part of the ''Improving the quantification of climate change hazards by hydrological models: A simple ensemble approach for considering the uncertain effect of vegetation response to climate change on potential evapotranspiration" journal publication (in preparation). WaterGAP2.2d global hydrological model with two different settings; 1) with standard PET method Priestley-Taylor (PT) and 2) with modified approach (PT-MA) (please refer to the publication for more details on the method) used to derive the data set. The bias-adjusted GCM-derived (GFDL-ESM2M, HadGEM2-ES, IPSL-CM5A-LR, and MIROC5) climate data under RCP2.6 and RCP8.5 emission scenarios were used as the input. The model-derived potential evapotranspiration and the renewable water resources variables are available from 1981 to 2099 on the monthly scale for each land grid cell (spatial resolution: 0.5 degrees x 0.5 degrees). The data files are in the netCDF format (.nc4). </p>
Resource status collected during AWOPS tests
<p>The dataset is about resource status collected during the validation of the Automated Work Planning Services (AWOPS). The dataset includes:</p> <ul> <li>5 csv files regarding crews composition involved at JEA renovation works from March 7<sup>th</sup> to April 11<sup>th</sup>;</li> <li>9 json files regarding crews efforts in the same period.</li> </ul>
CafeteriaSA corpus: Scientific abstracts annotated across different food semantic resources
<p>In the last decades, a great amount of work has been done in predictive modeling of issues related to human and environmental health. Resolution of issues related to healthcare is made possible by the existence of several biomedical vocabularies and standards, which play a crucial role in understanding health information, together with a large amount of health-related data. However, despite the large number of available resources and work done in the health and environmental domains, there is a lack of semantic resources that can be utilized in the food and nutrition domain, as well as their interconnections. For this purpose, in an European Food Safety Authority-funded project CAFETERIA, we have developed the first annotated corpus of 500 scientific abstracts that consists of 6,407 annotated food entities with regard to Hansard taxonomy, 4,299 for FoodOn, and 3,623 for SNOMED-CT. The CafeteriaSA corpus will enable further development of natural language processing methods for food information extraction from textual data that will allow extracting of food information from scientific textual data.</p>
CafeteriaFCD corpus: Food consumption data annotated with regard to different food semantic resources
<p>The FoodBase curated version which contains 1,000 manually evaluated recipes, annotated with the appropriate semantic tags from the Hansard Taxonomy, FoodON and SNOMED-CT.</p>
A chromosome-level genome resource for studying virulence mechanisms and evolution of the coffee rust pathogen Hemileia vastatrix
<p>Recurrent epidemics of coffee leaf rust, caused by the fungal pathogen <em>Hemileia vastatrix,</em> have constrained the sustainable production of Arabica coffee for over 150 years. The ability of <em>H. vastatrix </em>to overcome resistance in coffee cultivars and evolve new races is inexplicable for a pathogen that supposedly only utilizes clonal reproduction. Understanding the evolutionary complexity between <em>H. vastatrix</em> and its only known host, including determining how the pathogen evolves virulence so rapidly is crucial for disease management. Achieving such goals relies on the availability of a comprehensive and high-quality genome reference assembly. To date, two reference genomes have been assembled and published for <em>H. vastatrix</em> that, while useful, remain fragmented and do not represent chromosomal scaffolds. Here, we present a complete scaffolded pseudochromosome-level genome resource for <em>H. vastatrix </em>strain 178a (Hv178a). Our initial assembly revealed an unusually high degree of gene duplication (over 50% BUSCO basidiomycota_odb10 genes). Upon inspection, this was predominantly due to a single scaffold that itself showed 91.9% BUSCO Completeness. Taxonomic analysis of predicted BUSCO genes placed this scaffold in Exobasidiomycetes and suggests it is a distinct genome, which we have named Hv178a associated fungal genome (Hv178a AFG). The high depth of coverage and close association with Hv178a raises the prospect of symbiosis, although we cannot completely rule out contamination at this time. The main Ca. 546 Mbp Hv178a genome was primarily (97.7%) localised to 11 pseudochromosomes (51.5 Mb N50), building the foundation for future advanced studies of genome structure and organization. Citation: https://doi.org/10.1101/2022.07.29.502101</p>
TIME4CS WP4 Mapping of citizen science training resources
<p>This dataset was compiled as part of the TIME4CS project, WP4, and lists identified citizen science training resources, as of July 2022.</p> <p>The <a href="https://eu-citizen.science/">EU-citizen.science</a> platform provided the basis for mapping CS training in Europe, as the team behind the platform has put considerable effort into compiling, and encouraging the CS community to contribute, CS training resources. Additionally, training courses were identified based on the case studies in WP1, as most universities do not list their courses on the EU-citizen.science platform.</p>
BY-COVID Work Package 2 List of Resources
<p>Work Package 2: <em>Accessing heterogeneous data across domains and jurisdictions for enabling the downstream processing of COVID-19 and future pandemic episodes data </em>has gathered a relevant list of resources for the following areas: Non-patient related, Human-patient biomolecular, Human-patient clinical and health and socio-economics. </p>
Four Reference Quality Genome Assemblies of Pyrenophora teres f. maculata: A Resource for Studying the Barley Spot Form Net Blotch Interaction
<p>Updated draft genome assembly (FASTA) and annotation (GFF) for the <em>P. teres </em>f.<em> maculata </em>isolate FGOB10Ptm-1. </p>
Four Reference Quality Genome Assemblies of Pyrenophora teres f. maculata: A Resource for Studying the Barley Spot Form Net Blotch Interaction
<p>Updated draft genome assembly (FASTA) and annotation (GFF) for the <em>P. teres </em>f.<em> maculata </em>isolate P-A14. </p>
DISTANT-CTO: A Zero Cost, Distantly Supervised Approach to Improve Low-Resource Entity Extraction Using Clinical Trials Literature
<p><strong>Datasets</strong></p> <ol> <li>DISTANT-CTO is a weakly-labelled dataset of 'Intervention' and 'Comparator' entity annotated sentences. The dataset was obtained using candidate generation the approach described in "DISTANT-CTO: A Zero Cost, Distantly Supervised Approach to Improve Low Resource Entity Extraction Using Clinical Trials Literature". <ol> <li>distantcto_high_conf.txt - ds conf 1.0 (full dataset)</li> <li>extraction1_pos_posnegtrail_conf09.txt - ds conf 0.9 (partial dataset)</li> </ol> </li> <li>The physio test set is a dataset comprising 153 PICO annotated randomized controlled trial abstracts from Physiotherapy and Rehabilitation. This dataset was used as an additional benchmark to evaluate the generalization power of the weakly annotated dataset and NER model for this sub-domain.</li> </ol> <p> </p> <p><strong>Utility</strong></p> <p>The dataset could be used as an input for training 'Intervention' named-entity recognition (NER) models.</p> <p> </p> <p><strong>Availability</strong></p> <p>This directory includes extraction1_pos_posnegtrail_conf09.txt - This text data file contains all the weak annotations (source intervention terms mapped onto target sentences) from clinicaltrials.org (CTO) with a confidence score of 0.9 and above.</p> <p>The directory also includes ‘physio_sent_annot2POS_posnegtrail.txt’ – This data file contains manually annotated (Intervention entity) data from the physiotherapy and rehabilitation domain. It follows a roughly similar structure as described in the ‘Description for long targets’ section. (‘Participant’ and ‘Outcome’ annotations are removed from this file)</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.