Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
355
datasets available to search
ShareScore release 0.9.0
Dataset results
355 results for “data extraction”
Example raw quantification outputs used for extracting view data
This data contains the raw quantification outputs from platforms FragPipe, Maxquant, DIA-NN and Spectronaut. We use them as inputs for extracting view data serving as inputs to our newly designed multi-view proteomics framework.
Extract of the project data from the LIFE KPI webtool. Deliverable 2.5 of the LIFE NatuReef project: Nature-based reef solution for coastal protection and marine biodiversity enhancement. LIFE22-NAT-IT-LIFE-NatuReef/101113742
<p>Key performance indicator, a quantifiable measure of performance over time for a specific objective.</p>
Weekly supervised Multilingual Data Set to train Named Entity Recognition for Symptom Extraction
<p>Data Sets were generated using the Weakly Supervised NER pipeline (https://github.com/HUMADEX/Weekly-Supervised-NER-pipline) to train the symptom extraction NER models. </p> <p><strong>Supported Languages and dataset locations for the specific language:</strong></p> <p> English (base language): https://huggingface.co/HUMADEX/english_medical_ner<br> German: https://huggingface.co/HUMADEX/german_medical_ner<br> Italian: https://huggingface.co/HUMADEX/italian_medical_ner<br> Spanish: https://huggingface.co/HUMADEX/spanish_medical_ner<br> Greek: https://huggingface.co/HUMADEX/german_medical_ner<br> Slovenian: https://huggingface.co/HUMADEX/slovenian_medical_ner<br> Polish: https://huggingface.co/HUMADEX/polish_medical_ner<br> Portuguese: https://huggingface.co/HUMADEX/portugese_medical_ner</p> <p> </p> <p><strong>Dataset Building </strong></p> <ul> <li>Data Integration and Preprocessing</li> <li>Data Cleaning</li> <li>Annotation with Stanza's i2b2 Clinical Model </li> <li>Translation into the targeted language</li> <li>Word Alignment </li> <li>Data Augmentation </li> </ul> <p><strong>Acknowledgement</strong><br>This dataset had been created as part of joint research of HUMADEX research group (https://www.linkedin.com/company/101563689/) and has received funding by the European Union Horizon Europe Research and Innovation Program project SMILE (grant number 101080923) and Marie Skłodowska-Curie Actions (MSCA) Doctoral Networks, project BosomShield ((rant number 101073222). Responsibility for the information and views expressed herein lies entirely with the authors.</p> <p><strong>Authors:</strong><br>dr. Izidor Mlakar, Rigona Sallauka, dr. Umut Arioz, dr. Matej Rojc</p> <p><strong>Please cite as:</strong></p> <p><span>Article title: Weakly-Supervised Multilingual Medical NER For Symptom Extraction For Low-Resource Languages</span><br><span>Doi: 10.20944/preprints202504.1356.v1</span><br><span>Website: </span><a title="https://www.preprints.org/manuscript/202504.1356/v1" href="https://www.preprints.org/manuscript/202504.1356/v1">https://www.preprints.org/manuscript/202504.1356/v1</a></p>
Monthly averaged lightning and trace gases data extracted from EMAC simulations (2007, T42L90MA resolution)
<pre>About Dataset Monthly averaged lightning and trace gases data extracted from EMAC simulations (2007, T42L90MA resolution) Authors: Francisco J. Pérez-Invernon, Francisco J. Gordillo-Vázquez, Heidi Huntrieser, Patrick Jöckel and Eric J. Bucsela Description of the data CTR simulations: CTR_*.nc files LNOfs simulation: LNOfs_*.nc files *tr_*.nc: Monthly averaged trace gases *lnox*.nc: Monthly averaged lightning data<br>*ECHAM5*.nc: Monthly averaged dynamical variables<br>*grid_def*.nc: Monthly averaged grid variables<br>*tropop*.nc: Monthly averaged tropospheric variables </pre> <pre>File format: netcdf</pre> <p> </p>
Assessment of animal diseases caused by bacteria resistant to antimicrobials: Dogs and cats - Appendix B: Excel file with all data extracted
<p>Information on all the full-text studies that were assessed, including the reason for exclusion for those that were excluded at the full-text screening and the data extracted from the included studies, can be consulted here. </p> <p>The extensive literature review was carried out by the University of Copenhagen under the contract OC/EFSA/ALPHA/2020/02 – LOT 1 (https://ted.europa.eu/udl?uri=TED:NOTICE:457654-2020:TEXT:EN:HTML)</p>
Raw data extracted from ChEMBL
<p>Raw data files extracted from ChEMBL for the MELLODDY project.</p>
IFC Building Models for Automated Extraction of Data from Balconies
<p>This is a set of example Industry Foundation Classes (IFC) building models to extract domain specific construction information. More speficially, to extract locations of potential placement sites for thermal bridges between balconies and their neighbouring floors.</p>
Extracted data from primary literature examining impacts of recreational activities on freshwater ecosystems
<p>Aquatic ecosystems are attractive sites for recreation. However, human presence at or on aquatic ecosystems can have a range of ecological impacts, creating trade-offs between recreation as ecosystem service and biodiversity conservation. There is currently no synthesis of evidence regarding the ecological impacts associated with various forms of aquatic recreation, to compare the magnitude of effects between types of recreation. Therefore, conservation conflicts surrounding water-based recreation are difficult to manage. We conducted a global meta-analysis, differentiating various recreational impacts and the type of recreational uses in four categories: shore use, shoreline angling, swimming and boating; and studied ecological impacts directed at three levels of biological organization: individuals, populations, and communities. We screened over 13,000 articles and identified 94 suitable studies providing 701 effect sizes for inclusion in the meta-analysis. Aggregated across all animal and plant taxa, impacts of boating and shore use resulted in highly significant effects on almost all levels of biological organization. Regarding taxonomic groups, the most negative effects of water-based recreation were observed in invertebrates, whereas effects on birds were most pronounced at individual levels and not significant at community levels. From a conservation perspective, fostering water-based recreation and the ecological services they provide must be balanced with ecological impacts associated with the activities. Although generalizations are challenging, local scale effects of activity-specific constraints seem unlikely to be effective if other forms of water-based recreation continue.</p>
Data extracted from APSIM simulations for Sorghum
<p>The data are supplemental material for paper entitled APSIM-powered framework for effective post-rainy sorghum agri-system design in India.</p> <p>The APSIM sorghum model was run using environmental data spanning 30 years with a total of 13,824 Genotype × Management combinations per grid. To cover the Indian rabi sorghum production tract 311 grid items were used resulting in a total of 4299264 simulations. The computation took approximately 14 days including several downtime periods and generated 14.6 TB of output data. APSIM natively generates the data in raw text format, therefore for follow-up processing it was necessary to conduct extraction and parsing of relevant pieces of information. For the purposes of extraction and transformation into .csv file format a program was written in C# language to select only the relevant data.</p>
Assessment of animal diseases caused by bacteria resistant to antimicrobials: Horses - Appendix B: Excel file with all data extracted
<p>Information on all the full-text studies that were assessed, including the reason for exclusion for those that were excluded at the full-text screening and the data extracted from the included studies, can be consulted here. </p> <p>The extensive literature review was carried out by the University of Copenhagen under the contract OC/EFSA/ALPHA/2020/02 – LOT 1 (https://ted.europa.eu/udl?uri=TED:NOTICE:457654-2020:TEXT:EN:HTML)</p>
FlyTracker Video Analysis and Data Extraction
<p>Protocol video showing how to use the MATLAB package "FlyTracker" to analyze locomotor behavior in a video and extract the data in a format compatible with any worksheet software.</p> <p>Full uncompressed video made with FinalCut Pro (higher quality than version available at STAR Protocols)</p> <p>MATLAB package: <a href="https://github.com/kristinbranson/FlyTracker/archive/refs/heads/main.zip">FlyTracker</a></p> <p>MATLAB Script: <a href="https://github.com/LaurentSeroude/FlyTrackerExtraction">FlyTracker Extraction</a></p> <p>Peer-reviewed publications:</p> <p>Genome 64,139,2021 <a href="https://github.com/LaurentSeroude/FlyTrackerExtraction/blob/main/Genome%2064%2C139%2C2021.pdf">PDF</a></p> <p><a href="https://star-protocols.cell.com/protocols/2193">STAR Protocols 3,101888,2022</a></p> <p>Peer-reviewed protocol: STAR Protocols in press</p>
Data Extraction Sheet for the Review
<p>Approximately 30 variables will be extracted from the publications that are included in the review. This will include information on the:</p> <ol> <li>Study characteristics <ol> <li>Publication (title, year of publication, author(s) and their affiliation, journal, type of document)</li> </ol> </li> <li>Variables of interest <ol> <li>Concepts that are used for health research on racialised minority groups and how they are operationalised</li> <li>Research methodology and methods used</li> <li>The data used, and how this is collected, and applied</li> </ol> </li> </ol> <p>A full overview of the variables to be extracted can be found in this Data Extraction sheet.</p>
A STP-HSI index method for urban built-up area extraction based on multi-source remote sensing data
<p>The changes of urban built-up areas can reflect the process of urbanization, and it can reflect the population, economy, and cultural development of the city. Therefore, accurate and timely extraction of urban built-up areas plays an important role in the dynamic management of the city. In the existing research, single-source remote sensing data is used to extract urban built-up areas, and there is a problem that the spectrum of urban areas and non-urban areas is easily confused. Multi-source remote sensing data, including luojia-1 remote sensing data, Landsat 8 OLI remote sensing data, etc., can make up for the spectrum confusing issues.</p> <p>We fuse the time series information of night light remote sensing data, neighborhood information and point of interest (POI) data in spatial dimension, and propose a built-up area extraction method that integrates night light time and space information and POI information.</p>
Data Extraction table for the study Machine-based Stereotypes: How Machine Learning Algorithms Evaluate Ethnicity from Face Data
<p>This table contains the data extraction results for the study Machine-based Stereotypes: How Machine Learning Algorithms Evaluate Ethnicity from Face Data. It contains 24 columns and 74 rows.</p>
Data extraction Diabetes cochrane review
<p>This is the data that was extracted using Covidence. please note that some changes to risk of bias assessments were made later on review of the ratings by the lead author, upon discussion and agreement with the raters. These changes are noted directly in the Cochrane review. </p>
Data Extraction Sheet for the Review - 2
<p>Approximately 30 variables will be extracted from the publications that are included in the review. This will include information on the:</p> <ol> <li>Study characteristics <ol> <li>Publication (title, year of publication, author(s) and their affiliation, journal, type of document)</li> </ol> </li> <li>Variables of interest <ol> <li>Concepts that are used for health research on racialised minority groups and how they are operationalised</li> <li>Research methodology and methods used</li> <li>The data used, and how this is collected, and applied</li> </ol> </li> </ol> <p>A full overview of the variables to be extracted can be found in this Data Extraction sheet.</p>
Data Extraction Form
<p>This Excel sheet contains the data extraction and quality assessment for our Systematic Literature Review (SLR) titled "Challenges in Constructing Deep Learning Software - A Systematic Literature Review."</p>
Data extraction format
<p>data extraction format used to determine the survival status of patients admitted to the intensive care unit.</p>
Data from: Drop it all: Extraction-free detection of non-indigenous marine species through optimized direct-droplet digital PCR
<p>Molecular biosecurity surveillance programs increasingly use environmental DNA (eDNA) for detecting marine non-indigenous species (NIS). However, the current molecular detection workflow is cumbersome, prone to errors and delays, and is limited in providing knowledge about eDNA beyond the spatial and temporal extent of the sampling. These limitations can hinder management efforts and restrict the "opportunity window" for a rapid response to new marine NIS incursions. Emerging innovative field-deployable digital droplet PCR (ddPCR) systems offer improved workflow efficiency by autonomously analyzing targeted free-floating extra-cellular eDNA (free-eDNA) signals. Despite their potential, these systems have not been tested in marine environments. Thus, an aquarium study was conducted with three distinct marine NIS: <span>the Mediterranean fanworm <em>Sabella spallanzanii</em>, the ascidian clubbed tunicate <em>Styela clava</em>, and the brown bryozoan <em>Bugula neritina</em></span> to evaluate the detectability of free-eDNA in seawater. The detectability of targeted free-eDNA was assessed by directly analyzing aquarium water samples using an optimized species-specific ddPCR assay, without filtration or DNA extraction, so-called, "direct-ddPCR". The results demonstrated the consistent detection of <em>Sabella spallanzanii</em> and <em>Bugula neritina</em> free-eDNA when these organisms were present in high abundance. Once organisms were removed, the free-eDNA signal exponentially declined, noting that free-eDNA persisted between 24-72 hours. Results indicate that organism biomass, specimen characteristics (e.g., stress and viability), and species-specific biological differences may influence free-eDNA detectability. These results are critical for implementing <em>in-situ</em> nucleic acid automated continuous sensing systems for marine biosurveillance, enabling point-of-need detection and <span>rapid management response to biosecurity threats. </span></p>
Sample extraction and SNP sequencing data for: Identification of sex-linked SNP markers in wild populations of monomorphic birds
<p><span>Single-nucleotide polymorphism (SNP) analyses are a powerful tool for population genetics, pedigree reconstruction and phenotypic trait mapping. However, the untapped potential of SNP markers to discriminate the sex of individuals in species with reduced sexual dimorphism or of individuals during immature stages remains a largely unexplored avenue. Here, we develop a novel protocol for molecular sexing of birds based on the detection of unique Z- and W-linked SNP markers. Our method is based on the identification of two unique loci, one in each sexual chromosome. Individuals are considered males when they show no calls for the W-linked SNP and are heterozygotic or homozygotic for the Z-linked SNP, while females show both Z- and W-linked SNP calls. We validated the method in the Jackdaw (<em>Corvus</em> <em>monedula</em>). The reduced sexual dimorphism in this species makes it difficult to sex individuals in the wild. We assessed the reliability of the method using 36 individuals of known sex and found that their sex was correctly assigned in 100% of cases. The sex-linked markers also proved to be widely applicable to discriminate males and females from a sample of 927 genotyped individuals of different maturity stages with an accuracy of 99.5%. Given that SNP markers are increasingly used in quantitative genetic analyses of wild populations, the approach we propose has a great potential to be integrated into broader genetic research programmes without the need for additional sexing techniques.</span></p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.