Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
773
datasets available to search
ShareScore release 0.7.1
Dataset results
773 results for “data science”
Open Source Software in Data Science
<p>This upload includes an anonymized data set of a survey first launched in 2022. The survey has been revised since. The data set. however, contains answers of the first launch.</p>
Data coverage, biases, and trends in a global citizen-science resource for monitoring avian diversity
<p><strong>Aim:</strong> Understanding and addressing the global biodiversity crisis requires ecological information compiled continuously from across the globe. Data from citizen science initiatives are useful for quantifying species' ecological niches and geographical distributions but can be difficult to apply towards biodiversity monitoring. The presence of fixed geographical locations reduces the opportunistic nature of citizen science data, allowing for more reliable and nuanced trend estimation. The eBird citizen-science programs contains predefined locations whose bird assemblages are sampled across years ('hotspots'). For hotspots to function as a biodiversity monitoring resource, issues related to data coverage, biases, and trends need to be addressed.</p> <p><strong>Location:</strong> Global.</p> <p><strong>Methods:</strong> We estimated the survey completeness of species richness at 300,500 eBird hotspots during the years 2002 to 2022. We documented sampling biases at eBird hotspot and non-hotspot locations during 2022 based on protection status, temperature, precipitation, and landcover.</p> <p><strong>Results:</strong> A total of 10,410 bird species (<em>ca</em>. 96.9% of total) were recorded at hotspots. The number hotspots and the quantity of data and unique participants and quality of species richness estimates has increased worldwide with the Nearctic containing the strongest and most consistent trends. Compared to non-hotspots, hotspots over sampled areas with higher protection status. Hotspots and non-hotspots over sampled warmer and wetter locations in the Antarctic, Nearctic, and Palearctic, and cooler locations in the Afrotropics, Australasia, and the Neotropics. Hotspots and especially non-hotspots over sampled urban areas. Hotspots and non-hotspots under sampled shrublands in Australasia. Hotspots and especially non-hotspots under sampled forests in the Afrotropics, Indomalaya, Neotropics, and Oceania.</p> <p><strong>Main conclusions:</strong> Hotspots have captured a large component of the world's avian diversity but have done so inconsistently across space and time. Data quantity and quality are increasing in many regions, but the presence of sampling biases and spatial uncertainty needs to be addressed when applying the data.</p>
Dataset: Problem-centred interviews results for Matching Data Life Cycle and Research Processes in Engineering Sciences
<p>The authors would like to thank the Federal Government and the Heads of Government of the Länder, as well as the Joint Science Conference (GWK), for their funding and support within the framework of the NFDI4Ing consortium. Funded by the German Research Foundation (DFG) - project number 442146713.</p>
MaTableGPT: GPT-based Table Data Extractor from Materials Science Literature
<p>Using the MaTableGPT tool, which extracts table data from material literature based on GPT, performance data of catalysts were extracted from 11,077 documents on OER catalysts. GPT-3.5 fine-tuning was conducted using 126 tables, achieving an accuracy of 96.8%. The data consists of DOI, catalyst, performance, and property(reaction_type, value, current_density, electrolyte, overpotential, potential, substrate, versus, condition). For more details, please refer to the DOI.</p>
Combined data file for Jokinen et al. "Terrestrial organic matter input drives sedimentary trace metal sequestration in a human-impacted boreal estuary", Science of the Total Environment 717, 2020
<p>The datafile contains all the new raw data presented in the figures in the publication.</p>
Video delle lezioni - Open Science come e perché / Corso Universitario Aggiornamento Professionale per Data Steward, Università di Torino marzo 2024
<p>Video delle lezioni su Open Science, FAIR data management, politiche europee per CUAP data steward Università di Torino</p> <p>Contenuto;</p> <ol> <li>[DA AGGIUNGERE]</li> </ol>
Data for "Citizen science as a valuable tool for environmental review" - Frontiers in Ecology and the Environment - Callaghan et al.
<p>This dataset is the dataset used in Callaghan et al. Citizen science as a valuable tool for environmental review. Frontiers in Ecology and the Environment. The data are Environmental Impact Statement titles, and our coding of those documents. See paper for details.</p>
FIGURE 10 in An analysis of fossil identification guides to improve data reporting in citizen science programs
FIGURE 10. An example of a †Cosmopolitodus hastalis photo enhanced by illustration.
Austrian Science Fund (FWF) Publication Cost Data 2017
<p><br> Following the approach for the datasets in 2013 (<a href="http://dx.doi.org/10.6084/m9.figshare.988754">http://dx.doi.org/10.6084/m9.figshare.988754</a>), 2014 (<a href="https://dx.doi.org/10.6084/m9.figshare.1378610.v14">https://dx.doi.org/10.6084/m9.figshare.1378610.v14</a>), 2015 (<a href="https://doi.org/10.6084/m9.figshare.3180166">https://doi.org/10.6084/m9.figshare.3180166</a>) and 2016 (<a href="https://doi.org/10.5281/zenodo.810596">https://doi.org/10.5281/zenodo.810596</a>), the Austrian Science Fund (FWF) is making the publication costs spent in 2017 (esp. for Open Access) publically available.</p> <p>The dataset includes payments for publications of authors funded by the Austrian Science Fund (FWF) via following programmes:</p> <p>Peer-Reviewed Publications: <a href="https://www.fwf.ac.at/en/research-funding/fwf-programmes/peer-reviewed-publications/">https://www.fwf.ac.at/en/research-funding/fwf-programmes/peer-reviewed-publications/</a></p> <p>Stand-Alone Publications: <a href="https://www.fwf.ac.at/en/research-funding/fwf-programmes/stand-alone-publications/">https://www.fwf.ac.at/en/research-funding/fwf-programmes/stand-alone-publications/</a></p>
[DATA_SCIENCE] Interviews: The Secure Anonymised Information Linkage Databank (SAIL)
<p>This is a set of interview transcripts executed by Niccolò Tempini between September 2015 and October 2016, as part of the ERC project "The Epistemology of Data-Intensive Science", and in the context of a case study of SAIL. Please read the "Notes on transcript editing" document for further information.</p> <p>Two papers that specifically make use of these interviews have been published as of 2018:</p> <ul> <li>Tempini, N., Leonelli, S., 2018. Concealment and discovery: The role of information security in biomedical data re-use. Soc Stud Sci 48, 663–690. <a href="https://doi.org/10.1177/0306312718804875">https://doi.org/10.1177/0306312718804875</a></li> <li>Tempini, N., 2016. Science Through the “Golden Security Triangle”: Information Security and Data Journeys in Data-intensive Biomedicine, in: Proceedings of the 37th International Conference on Information Systems (ICIS 2016). <a href="https://aisel.aisnet.org/icis2016/ISHealthcare/Presentations/20/">https://aisel.aisnet.org/icis2016/ISHealthcare/Presentations/20/</a></li> </ul> <p>The transcripts document SAIL researchers' experience of infrastructure development and data curation and re-use practices. Researchers have consented to have these transcripts made available as Open Data. Other interviewees did not give consent, so those transcripts are held securely by the research team in Exeter.<br> You also find the information sheet provided to interviewees, which gives you the context for this project. Further information and related publications can be found at www.datastudies.eu.</p>
Data sharing in the life sciences : a study of researchers at the Norwegian University of Life Sciences
<p>Survey data collected in 2012 as basis for master thesis in library and information sciences titled "Data sharing in the life sciences : a study of researchers at the Norwegian University of Life Sciences" </p>
Illustrative Darwin core archive to output data from a citizen science platform to a collection management system
<p>Illustrative DwC archive to send data back to a collection management system from a citizen sciences platform. This illustrative archive displays the specimens used for the trans-institutional and trans-platform pilot project held in the frame of ICEDIG.</p> <p>Further description of its content in the milestone28 document, worpackage 5.2 of the ICEDIG project.</p>
Data for "Corruption Risk in Contracting Markets: A Network Science Perspective"
<p>EU public procurement data, scored for corruption risk. Includes deduped issuers and winners. Includes data dictionary. For more information contact: johanneswachs@gmail.com</p>
Data set used in "Diving into science and conservation recreational divers can monitor reef assemblages"
<p>This data set was used in the manuscript "Diving into science and conservation: recreational divers can monitor reef assemblages", accepted for publication in Perspectives in Ecology and Conservation. It is composed by reef species (fishes and turtles) counts in different size classes performed by recreational and experience scientific divers following a sampling protocol. Additionally there is information about recreational divers diving experience, underwater waste register and the scores for their evaluation about conducting scientific data collection.</p>
Langmark: annotations for scenes with semantic inconsistencies connecting distributional semantic models to vision science – data and code
<p>Data (including object annotations) and code from the following manuscript:</p> <p><em>Langmark: annotations for scenes with semantic inconsistencies connecting distributional semantic models to vision science</em>.</p>
Detecting the effect of intensive agriculture on Odonata diversity using citizen science data
<p>These datasets are used in the following article :</p> <p><strong><span>Detecting the effect of intensive agriculture on Odonata diversity using citizen science data</span></strong></p> <p><strong><span><span>By </span></span></strong><span>Renaud Baeta<sup>1</sup>, Justine Léauté<sup>1,2</sup>, Éric Sansault<sup>1</sup> and Sylvain Pincebourde<sup>2*</sup></span></p> <p><span><strong>In </strong></span><em><span>Ecological Applications </span></em><span>(Research article)</span></p> <p><strong><span>Abstract</span></strong></p> <p><span><span><span>Agricultural areas represent one of the major ecosystems of the world. Intensification of agricultural practices produced openfields characterized by low biological diversity. Nevertheless, the distance up to which intensive agricultural fields alter surrounding natural systems is rarely quantified. We determined the spatial scale at which agricultural landscapes alter the diversity of Odonates, a key taxon in wetland ponds, and we tested to what extent citizen-science data can be used reliably for this purpose. We compiled 7,731 observations made in a portion of the region Centre-Val-de-Loire (France) over 10 years by naturalists on 729 water bodies to analyze the effect of agricultural landscapes (mainly wheat, rapeseed, sunflower) on the species richness of both damselflies and dragonflies in lentic systems. Sixty species were reported over the 10-years period. For dragonflies, intensive agricultural landscapes best explained their richness at the scales of 800m and 1,600m for overall and autochthonous species, respectively, when using the full dataset. The spatial scale was smaller for damselflies, at 200m for both overall and autochthonous species. These distances were not severely impacted when constraining the data to consider several biases. Multi-model averaging showed that the proportion of intensive agriculture decreased species richness, despite the potential biases inherent to an imperfect database acquired by citizens. This imperfect citizen dataset allows to infer the lowest effect size of agriculture on species richness. Quantitatively, this effect was more important for autochthonous species. Interestingly, both relatively rare taxa and common or generalist species can be under threat in intensive agricultural landscapes, calling for more ecotoxicological studies. The influence of agricultural practices from a distance implies that conservation and management plans of wetland ponds should consider the landscape ecological characteristics and not only the pond features. Conservation efforts focusing too locally on a site may be undermined because intensive agriculture from a distance limits the potential for the site to recover highly diverse communities. These distant effects should be integrated by policy-makers when deciding which wetland pond should benefit from a conservation plan or which conservation action may be planned, implementing for instance buffer zones and/or ecological corridors composed of natural vegetation. </span></span></span></p>
D1.4 - Consolidated OA data set for responsible research practices and trust in science: POIESIS National and Regional Database
<p>Deliverable 1.4 of the POIESIS project is a consolidated open access data set for responsible research practices and trust in science. D1.4 consists of two databases: the POIESIS National Database and the POEISIS Regional Database. The method and rationale behind the construction of all indicators in both databases are laid out in Work Package 1’s previous deliverable, D1.3 (Bauer et al., 2024). The two database are accompanied by a Guidance Sheet which provides details and relevant notes regarding the structure of the two databases. </p> <p> </p>
Data Science Tasks used in "AI Support for Data Scientists: An Empirical Study on Workflow and Alternative Code Recommendation"
<p>This entry contains the supplementary files for a scientific article. </p> <p>The dataset contains the necessary files for the two data science tasks used in the experiment study from scientific article "AI Support for Data Scientists: An Empirical Study on Workflow and Alternative Code Recommendation"</p> <p> </p>
Dataset and Model for Data Science Projects
<p>This zenodo is for the data science projects hosted in this GitHub repository: https://github.com/oathaha/data-science-portfolio/tree/main</p>
Dataset from TableLabler: Scalable Labeling of Data Tables with Language Models for Tabular Dataset Creation [Scalable Data Science]
<p>TableLabler: Scalable Labeling of Data Tables with Language Models for Tabular Dataset Creation [Scalable Data Science]</p> <p>Pre-publicatoin upload for VLDB review.</p> <p>Code is at https://github.com/RelationalAI/annotated-tables/. The Github repository name is from a previous draft version and cannot be changed. It is code for TableLabler.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.