Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
6,025
datasets available to search
ShareScore release 0.9.0
Dataset results
6,025 results for “Science of science”
Langmark: annotations for scenes with semantic inconsistencies connecting distributional semantic models to vision science – data and code
<p>Data (including object annotations) and code from the following manuscript:</p> <p><em>Langmark: annotations for scenes with semantic inconsistencies connecting distributional semantic models to vision science</em>.</p>
Detecting the effect of intensive agriculture on Odonata diversity using citizen science data
<p>These datasets are used in the following article :</p> <p><strong><span>Detecting the effect of intensive agriculture on Odonata diversity using citizen science data</span></strong></p> <p><strong><span><span>By </span></span></strong><span>Renaud Baeta<sup>1</sup>, Justine Léauté<sup>1,2</sup>, Éric Sansault<sup>1</sup> and Sylvain Pincebourde<sup>2*</sup></span></p> <p><span><strong>In </strong></span><em><span>Ecological Applications </span></em><span>(Research article)</span></p> <p><strong><span>Abstract</span></strong></p> <p><span><span><span>Agricultural areas represent one of the major ecosystems of the world. Intensification of agricultural practices produced openfields characterized by low biological diversity. Nevertheless, the distance up to which intensive agricultural fields alter surrounding natural systems is rarely quantified. We determined the spatial scale at which agricultural landscapes alter the diversity of Odonates, a key taxon in wetland ponds, and we tested to what extent citizen-science data can be used reliably for this purpose. We compiled 7,731 observations made in a portion of the region Centre-Val-de-Loire (France) over 10 years by naturalists on 729 water bodies to analyze the effect of agricultural landscapes (mainly wheat, rapeseed, sunflower) on the species richness of both damselflies and dragonflies in lentic systems. Sixty species were reported over the 10-years period. For dragonflies, intensive agricultural landscapes best explained their richness at the scales of 800m and 1,600m for overall and autochthonous species, respectively, when using the full dataset. The spatial scale was smaller for damselflies, at 200m for both overall and autochthonous species. These distances were not severely impacted when constraining the data to consider several biases. Multi-model averaging showed that the proportion of intensive agriculture decreased species richness, despite the potential biases inherent to an imperfect database acquired by citizens. This imperfect citizen dataset allows to infer the lowest effect size of agriculture on species richness. Quantitatively, this effect was more important for autochthonous species. Interestingly, both relatively rare taxa and common or generalist species can be under threat in intensive agricultural landscapes, calling for more ecotoxicological studies. The influence of agricultural practices from a distance implies that conservation and management plans of wetland ponds should consider the landscape ecological characteristics and not only the pond features. Conservation efforts focusing too locally on a site may be undermined because intensive agriculture from a distance limits the potential for the site to recover highly diverse communities. These distant effects should be integrated by policy-makers when deciding which wetland pond should benefit from a conservation plan or which conservation action may be planned, implementing for instance buffer zones and/or ecological corridors composed of natural vegetation. </span></span></span></p>
D1.4 - Consolidated OA data set for responsible research practices and trust in science: POIESIS National and Regional Database
<p>Deliverable 1.4 of the POIESIS project is a consolidated open access data set for responsible research practices and trust in science. D1.4 consists of two databases: the POIESIS National Database and the POEISIS Regional Database. The method and rationale behind the construction of all indicators in both databases are laid out in Work Package 1’s previous deliverable, D1.3 (Bauer et al., 2024). The two database are accompanied by a Guidance Sheet which provides details and relevant notes regarding the structure of the two databases. </p> <p> </p>
Citizen Science Scan 2023 Belgium - dataset
Open the record for dataset details and reuse information.
Data Science Tasks used in "AI Support for Data Scientists: An Empirical Study on Workflow and Alternative Code Recommendation"
<p>This entry contains the supplementary files for a scientific article. </p> <p>The dataset contains the necessary files for the two data science tasks used in the experiment study from scientific article "AI Support for Data Scientists: An Empirical Study on Workflow and Alternative Code Recommendation"</p> <p> </p>
OA-Concepts and Wikipedia-Links for "The different AI of Science and Wikipedia"
<p>The files contain the data for the VosViewer analyses in "The Different Artificial Intelligences of Science and Wikipedia" (Korte et al. 2024).</p> <p>OpenAlex:</p> <p>As described in the paper, all works from OpenAlex from 2001 and 2022 with the concept "Artificial Intelligence" and a concept score > 0.3 were downloaded. In August 23 using the OpenAlex API. The files contain for each work all concepts with a concept score > 0.3 in one line each separeted by dots. This allows co-occurrence analyses of concepts in VosViewer. For the analyses in the paper whitespaces and "(" in the concepts were removed.</p> <p>First line for 2001: <br>Random forest. Mathematics. AdaBoost. Statistics. Tree (set theory). Generalization. Support vector machine. Generalization error. Measure (data warehouse). Artificial intelligence. Pattern recognition (psychology). </p> <p>Wikipedia:</p> <p>As described in the paper, all Wikipedia pages of the category "Artificial Intelligence" and of all direct sibling categories were downloaded. In August 23 with the Wikipedia Periodic Revisions tool of El Baff and Hecking (https://github.com/DLR-SC/wikipedia-periodic-revisions): . The files contain for each page all hyperlinks to other Wikipedia pages for the dates of 12.31.2005 and 12.31.2021 in one line each separeted by dots. This allows co-occurrence analyses of links in Wikipedia.</p> <p>First line of 2006: <br>vehicleregistrationplate. corporation. googlesearch. googleplatform. googol. google(disambiguation). menlopark,california. erice.schmidt. sergeybrin. lawrencee.page. georgereyes. internet. [...]</p>
Dataset and Model for Data Science Projects
<p>This zenodo is for the data science projects hosted in this GitHub repository: https://github.com/oathaha/data-science-portfolio/tree/main</p>
Ilustração Frase Open Science
<p>Título: Open Science: Just Science Done Right!</p> <p>Descrição: Esta ilustração está relacionada à Ciência Aberta, destacando a frase “Open Science: Just Science Done Right!”, de Jon Tennant, um defensor conhecido da ciência aberta e do acesso aberto. A arte feita à mão utiliza um lettering fluido e criativo para reforçar a mensagem sobre a importância de tornar a ciência acessível, transparente e colaborativa. Os toques de laranja vibrante trazem dinamismo à obra, simbolizando a energia e o impacto transformador da ciência aberta na sociedade global.</p> <p>Autor(a): Malu de Carvalho (Instagram: @mapadebiblio)</p> <p>Licença: Este trabalho está licenciado sob CC BY-NC-ND 4.0.</p> <p>Assunto: Ciência Aberta, Acesso Aberto, Comunicação Científica, Ilustração, Lettering</p> <p>DOI: 10.5281/zenodo.13853213</p>
Dataset from TableLabler: Scalable Labeling of Data Tables with Language Models for Tabular Dataset Creation [Scalable Data Science]
<p>TableLabler: Scalable Labeling of Data Tables with Language Models for Tabular Dataset Creation [Scalable Data Science]</p> <p>Pre-publicatoin upload for VLDB review.</p> <p>Code is at https://github.com/RelationalAI/annotated-tables/. The Github repository name is from a previous draft version and cannot be changed. It is code for TableLabler.</p>
The European Language Social Science Thesaurus (ELSST)
<p>The European Language Social Science Thesaurus (ELSST) is a broad-based, multilingual thesaurus for the social sciences. It is owned and published by the Consortium of European Social Science Data Archives (CESSDA) and its national Service Providers. The thesaurus consists of over 3,400 concepts and covers the core social science disciplines: politics, sociology, economics, education, law, crime, demography, health, employment, information and communication technology, and environmental science.</p> <p>ELSST is used for data discovery within CESSDA and facilitates access to data resources across Europe, independent of domain, resource, language or vocabulary.</p> <p><strong><em>Recommended Citation</em></strong>: CESSDA and Service Providers (2024) The European Language Social Science Thesaurus (ELSST) (Version 5), <a href="https://elsst.cessda.eu">https://elsst.cessda.eu</a>. DOI: 10.5281/zenodo.13843400</p>
Venus Express IMA solar wind data for VeRa radio science observations
<p>The Venus Express (VEX) Analyser of Space Plasmas and Energetic Atoms (ASPERA-4)<br>ion mass spectrometer (IMA) observed the the pristine solar wind parameters<br>at Venus between 2006 and 2014. Details can be found in<br>Barabash et al. (2007) Planetary and Space Science, Vol. 55 (12), pp. 1772-1792<br>The Analyser of Space Plasmas and Energetic Atoms (ASPERA-4) for the <br>Venus Express mission.</p> <p>This file is based on the VEX_IMA_SW_NVP_20060101.cef data file provided<br>by M. Fränz (MPI for Solar System Research, Goettingen, Germany) and contains only those ASPERA-4<br>solar wind observations used in the publication Peter et al. (2024) Icarus<br>The variability of the topside ionospheres of Venus and Mars<br>as seen by recent radio science observations (Appendix A4 Fig. 2).</p> <p>Data are based on hourly average IMA spectra obtained when VEX was at larger distance than<br>1 Venus radius from nominal bow shock (Martinecz 2009, PhD). Density and total velocity<br>are obtained by integration over the IMA_EXTRA proton spectrum. Dynamic pressure<br>by product of density and velocity. This file contains only orbits used for VERA analyis.</p>
Fault Traces Dataset for Zou and Fialko Earth and Space Sciences Manuscript
<p>The 'xx_fault.dat' contains the linked fault traces of different regions (nz: Northern New Zealand; nv: Basin and Range Province; ca: Ventura County, California; np: Pennsylvania and Northern New Jersey); The 1st and 2nd columns are UTM coordinates; The 3rd column is the random number assigned for distinguishing each fault traces.</p> <p>The 'xx_len.dat' contains the length of each fault trace in the corresponding regions, in km.</p> <p>The other '.dat' files and the 'SunData.xls' contain the length of fractures from outcrop and lab data. All of them are frequency distribution, except the 'LaHouve_Villemin.dat' which is already in cumulative distribution. The unit of 'SunData.xls' is mm; for the two 'Bahat' datasets is cm; for the rest of the outcrop data is m.</p> <p>The .m files are the codes for calculating cumulative length distribution, frequency density distribution, and fault connection.</p> <p> </p> <p> </p> <p>For any questions please contact Xiaoyu Zou via x3zou@ucsd.edu</p>
MULTIPLIERS_WP5_Science learning project on AMR_UAB_Public data_v1
<p><span>The datasets contain the following data related to the science learning project on ARM: </span></p> <ul> <li><span>Summary of transcripts from interviews with OSC members, and student groups (Pseudo-/Anonymised).</span></li> </ul>
PROOF OF THE CLASSICAL MECHANICAL COSMOS IN FIVE ADVERTISEMENT POSTERS THIS CURES SINISTER SCIENCE AND PROVIDES AN EASY WAY OUT OF HOMO SAPIENS' 5000-YEAR WARPATH TOWARD EXTINCTION
<p>Gerhard Ris former DA magistrate and lawyer with thirty years experience in courts of law.</p> <p>The previous DOI publications are in the links below. My last publication on Zenodo ten days ago has 37 views and 43 downloads.</p> <p>This five-page poster article further improves the Elementary Scientist Exam and the six other Zenodo-published articles, including the elementary list with the Unambiguous 1000 Elementary Terms, Definitions, and Descriptions in Elementary Science, Courtrooms, and Schools. (Proven to have a very high view-to-download ratio with 317 views and 243 downloads in six months after publication. A glitch (?) at CERN probably causes a 0 view 0 download error.) The undisputed reason for the improvement follows from the indisputably best-practice correct definition of science.</p> <p>A decent systematic trial and error search for all the laws of everything and an all-inclusive Bildung in Education Permanente.</p> <p>This has been solved and published in mentioned publications as One Law of Nature makes One Law of Human Nature. These two laws elevate all sciences to the wordiness of an exact science. Mind, that this definition can be reduced by redefining certain terms. It’s a never-ending process akin to growing wise oak of the tree-like tower of science continuously broadening and improving its base. This has been successfully done for the first time in 5000 years.</p> <p>Mathematics must be made to fit Nature and not the other way around.</p> <p>Science proper per definition demands all-inclusive Bildung. Bildung requires marketing, advertisement, and sales of science proper. To that end please note this selection of thumbnails as posters advertising the classes explaining the model in chunks of nigh 4-minutes.</p>
Exploring Data at Scale with Arkouda: A Practical Introduction to Scalable Data Science
<p><span>Data scientists can be thought of as modern-day explorers, venturing into the vast unknown of information. However, this exciting journey is not without its hurdles. One of the biggest challenges they face is the sheer immensity of data they encounter. Modern datasets cannot fit in laptop memory, containing terabytes or even petabytes of information. Working with such massive data requires specialized tools and techniques to extract meaningful insights. As data sets are growing ever larger, data science demands interactivity, where scientists can learn while working with the data. At the same time, data science demands scalability, where scientists are able to work with data sets in their entirety. Data scientists have naturally been drawn to Python as it provides interactivity through its read, evaluate, print loop and performance through its utilization of libraries written in other languages, like C and Fortran. These libraries typically are not designed for HPC and run into problems when attempting to scale. The gap that Arkouda fills in the data science landscape is a library that is both interactive, providing a familiar Python API, and scalable, leveraging a scalable Chapel server in the backend. Arkouda is a framework for scalable Python packages for interactive data science and has applications ranging from oceanography to net flow analysis.</span></p>
Data Matrix Theme-Specific Analysis of the Recommendation on Science and Scientific Researchers (RSSR): Open Access, Open Data, and Open Science
<p>This Table sets out findings from the mapping exercise conducted as part of the objectives of subtask 6.1 of the RRING project.</p> <p>Aim: Alignment of RRI to advance the UN SDGs.</p> <p>Objectives:</p> <ul> <li>Mapping the RSSR to the SDGs </li> </ul> <p>Mapping the RSSR to the SDGs is aimed at providing new perspectives, ideas and approaches that can help to improve the operationalization and implementation of each SDG, <em>by facilitating the integration of RRI (or RRI-like) practices in the SDGs, to make them more achievable.</em> The impact of the new perspectives, ideas and approaches in SDG operationalization and implementation will be aimed at the level of <em>national and international policy (making); future research and innovation projects (in industry and academia); as well as education and training of researchers, policy makers and other stakeholders.</em></p> <p>Two documents were used for this task:</p> <ul> <li>2017 Recommendation on Science and Scientific Researchers ([RSSR], UNESCO), and</li> <li>the United Nations 2030 Agenda for Sustainable Development with the 17 Sustainable Development Goals (SDGs).</li> </ul>
comparison_sciences Dataset
<p>The <code>comparison_sciences</code> dataset contains word-frequency and other non-consumptive-use data about 553,699 unique English-language news documents (no duplicate or close-variant documents) that contain the words "science" or "sciences." The documents came from U.S. mainstream and student news sources published during 1977-2019 (though mostly from 1985-2019). WE1S researchers use this data to understand how public discourse about the humanities compares to public discourse about science.</p> <p>We gathered this data using keyword searches for "science," which found articles containing either (or both) the words "science" and "sciences." We took data from the top 10 circulating newspapers in the U.S. and from University Wire sources (student newspapers). Documents in this dataset may also contain the word "humanities," just as documents in the <a href="https://zenodo.org/record/5068311"><code>humanities_keyword</code></a> dataset may contain the words "science" or "sciences."</p> <p><em>(See <a href="https://we1s.ucsb.edu/research/we1s-materials/">WE1S Research Materials Overview</a> for the relation between the project's "datasets" and "collections.")</em></p>
Table 1: Publications found in the Scopus and Web of Science database from 2000 to 2019 (excluding duplicate articles) on the topic of biorefineries.
<p>Table referring to the study of systematic literature review on the social impact of small and medium-sized biorefineries.</p>
Data from: Processing citizen science- and machine-annotated time-lapse imagery for biologically meaningful metrics
Time-lapse cameras facilitate remote and high-resolution monitoring of wild animal and plant communities, but the image data produced require further processing to be useful. Here we publish pipelines to process raw time-lapse imagery, resulting in count data (number of penguins per image) and 'nearest neighbour distance' measurements. The latter provide useful summaries of colony spatial structure (which can indicate phenological stage) and can be used to detect movement – metrics which could be valuable for a number of different monitoring scenarios, including image capture during aerial surveys. We present two alternative pathways for producing counts: 1) via the Zooniverse citizen science project Penguin Watch and 2) via a computer vision algorithm (Pengbot), and share a comparison of citizen science-, machine learning-, and expert- derived counts. We provide example files for 14 Penguin Watch cameras, generated from 63,070 raw images annotated by 50,445 volunteers. We encourage the use of this large open-source dataset, and the associated processing methodologies, for both ecological studies and continued machine learning and computer vision development.
GLOBE Mosquito Habitat Mapper Citizen Science Data 2017-2020
<p>Three Cases: Metadata and Procedures</p> <p>The data sets described here were used in an article submitted to the journal GeoHealth in 2021. The data files and further supplemental links (including general information about GLOBE data) can be accessed at https://observer.globe.gov/get-data/mosquito-habitat-data.</p> <p>Case 1: Removal of records with suspect geolocation data. A Python script was applied to remove records where the measured position (in decimal degrees) was identical to the GLOBE MGRS site position. GPS-obtained latitude and longitude coordinates are reported in decimal degrees, so records identified by whole numbers were also removed. This procedure removed 5704 (23%) of the 24983 records in the Mosquito Habitat Mapper database, with 19,279 records remaining. The secondary data sets cleaned only for geolocation anomalies were labeled Case 1.</p> <p>Case 2: Identifying suspected training events. For this test, we sought to identify groups of data that exceeded 10 records sharing these characteristics. Another Python script was employed to extract the photos for ease of visual inspection. Because we needed to manually review the photo records, we set the threshold for groups at >10, so that the analysis could be completed in the time allotted. Groups identified thought this procedure were outputted as case 2: groups. The resulting data set cleaned of groups >10 was labeled Case 2. The resulting data set included 20,006 records and identified 2,447 records found in clusters we postulated were training events.</p> <p>Case 3: The Case 3 secondary dataset result from applying the Python scripts used to create Cases 1 and 2. We used the Case 3 data sets, with improved geolocation and large groups eliminated, in the following analysis.</p> <p>Acknowledgments: These data were obtained from NASA and the GLOBE Program and are freely available for use in research, publications and commercial applications. When data from GLOBE are used in a publication, we request this acknowledgment be included: "These data were obtained from the GLOBE Program." Please include such statements, either where the use of the data or other resource is described, or within the acknowledgments section of the publication.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.