Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
6,025
datasets available to search
ShareScore release 0.9.0
Dataset results
6,025 results for “Science of science”
Dataset supplementing Marx, S., Gruenhage, G., Walper, D., Rutishauser, U., Einhäuser, W. (2015). Competition with and without priority control: linking rivalry to attention through winner-take-all networks with memory. Annals of the New York Academy of Sciences. 1339, 138-153.
<p>Data supplementing the paper Marx, S., Gruenhage, G., Walper, D., Rutishauser, U., Einhäuser, W. (2015). Competition with and without priority control: linking rivalry to attention through winner-take-all networks with memory. <em>Annals of the New York Academy of Sciences. 1339, </em>138-153. doi: 10.1111/nyas.12575 The files can be freely used for scientific purposes, provided this reference is appropriately cited.</p> <p>Files contain the behavioral data, the model can be found at https://doi.org/10.5281/zenodo.573026</p> <p> </p> <p>The following files are contained in this folder:</p> <p>dataExp1.mat contains the data of experiment 1</p> <p>The variables durationLeft and durationRight contain 5 x 6 x 6 cell arrays with the dominance durations for the left and right grating, respectively. Dimensions are subject x contrast level left x contrast level right.</p> <p><br> dataExp2.mat contains the data of experiment 2</p> <p>Variables buttonStart, buttonEnd and whichButton contain 3x4x5 (contrast levels x blank duration levels x subjects) cell arrays that contain the start time and end time of each button press, and which button (1/2) was pressed, respectively.</p> <p>Variables presStart and presEnd contain 3x4x5 (contrast levels x blank duration levels x subjects) cell arrays that contain start and end of each blank period. All time stamps refer to the onset of the first blanking trial (end of continuous presentation)</p> <p>Variable prevPerz contains the percept (button) that was pressed at the end of the continuous presentation period.</p> <p><br> figure3_human.m, figure4_human.m and figure6_human.m exemplify the usage of the data by re-plotting the figures containing human data of the aforementioned paper</p>
Supplementary material for the publication: "Efficient Surrogate Models for Materials Science Simulations: Machine Learning-based Prediction of Microstructure Properties"
<p><span><span><span>This dataset contains supplementary code, images and models for the publication „Efficient Surrogate Models for Materials Science Simulations: Machine Learning-based Prediction of Microstructure Properties“.</span></span></span></p> <p> </p> <p><span><span><span>The content will be updated and additionally linked to the corresponding git repositories.</span></span></span></p>
Package and Dependency Metadata for CZI Hackathon: Mapping the Impact of Research Software in Science
<p>A collection of useful datasets extracted from <a href="https://packages.ecosyste.ms">https://packages.ecosyste.ms</a> and <a href="https://repos.ecosyste.ms/">https://repos.ecosyste.ms</a> for use at the CZI Hackathon: Mapping the Impact of Research Software in Science.</p><p>All data is provided as NDJSON (new line delimited JSON), each line represents a valid JSON object, and they are separated by newline characters. There are <a href="https://pypi.org/project/ndjson/">python</a> and <a href="https://www.rdocumentation.org/packages/ndjson/versions/0.9.0/topics/stream_in">R</a> libraries for reading these files, or you can maually read each line and parse each line as a single JSON object.</p><p>Each ndjson file has been compressed with gzip (actual command: `tar -czvf`) to reduce download size, they expand to significantly bigger files after extraction.</p><h4>Package Data</h4><p>Package names from cran, bioconductor and pypi that have been parsed by the <a href="https://github.com/chanzuckerberg/software-mentions">software-mentions</a> project (data: <a href="https://datadryad.org/stash/dataset/doi:10.5061/dryad.6wwpzgn2c">https://datadryad.org/stash/dataset/doi:10.5061/dryad.6wwpzgn2c</a>) are collected together with their latest release at time of publishing along with the names of their dependencies, those dependency names have then also been recursively fetched with latest release and dependencies until the full list of transitive dependencies is included. </p><p>Note: This approach uses a simplified method of dependency resolution, always picking the latest version of each package rather than taking into account each dependencies specific version range requirements, this is primarily due to time constraints and allows all software ecosystems to be processed in the same way. A future improvement would be to use each package ecosystem's specific dependency resolution algorithm to compute the full transitive dependency tree for each mentioned software package.</p><h4>GitHub Data</h4><p>Two different approaches were taken for collecting data for referenced GitHub mentions:</p><p>1. `github.ndjson` is metadata for each repository from GitHub, including "manifest" files which are known files that contain dependency information for a project such as requirements.txt, DESCRIPTION and package.json, parsed using <a href="https://github.com/ecosyste-ms/bibliothecary">https://github.com/ecosyste-ms/bibliothecary</a>, which may include transitive dependencies that have been discovered in a `lockfile` within the repository.</p><p>2. `github_packages.ndjson` is metadata for each package that was found on any package manager that references the GitHub url as it's repository url/source/homepage, these packages, like the cran and pypi data above, include the latest release and their direct dependencies. There may be more than one package for each GitHub URL as it is a one to many relationship. `github_packages_with_transitive.ndjson` follows the same format but also includes the extra resolved transitive dependencies of all packages using the same approach as with cran and pypi data above with the same caveats. </p><p>There are also many more ecosystems referenced in these files than just cran, bioconductor and pypi, https://packages.ecosyste.ms provides a standardized metadata format for all of them to enable comparison and simplification of automation.</p><h4>Contact</h4><p>If you would like any help, support or more data from Ecosyste.ms please do get in touch via email: hello@ecosyste.ms or open an issue on GitHub: https://github.com/ecosyste-ms/packages/issues</p>
Dataset - Understanding the software and data used in the social sciences
<p>This is a repository for a UKRI Economic and Social Research Council (ESRC) funded project to understand the software used to analyse social sciences data.</p><p>Any software produced has been made available under a BSD 2-Clause license and any data and other non-software derivative is made available under a CC-BY 4.0 International License. Note that the software that analysed the survey is provided for illustrative purposes - it will not work on the decoupled anonymised data set.</p><p>Exceptions to this are:</p><ul><li>Data from the UKRI ESRC is mostly made available under a <a href="https://creativecommons.org/licenses/by-nc-sa/4.0/">CC BY-NC-SA 4.0</a> Licence.</li><li>Data from Gateway to Research is made available under an <a href="https://www.nationalarchives.gov.uk/doc/open-government-licence/version/3/">Open Government Licence</a> (Version 3.0).</li></ul><h2>Contents</h2><ul><li>Survey data & analysis: esrc_data-survey-analysis-data.zip</li><li>Other data: esrc_data-other-data.zip</li><li>Transcripts: esrc_data-transcripts.zip</li><li>Data Management Plan: esrc_data-dmp.zip</li></ul><h3>Survey data & analysis</h3><p>The survey ran from 3rd February 2022 to 6th March 2023 during which 168 responses were received. Of these responses, three were removed because they were supplied by people from outside the UK without a clear indication of involvement with the UK or associated infrastructure. A fourth response was removed as both came from the same person which leaves us with 164 responses in the data.</p><p>The survey responses, Question (Q) Q1-Q16, have been decoupled from the demographic data, Q17-Q23. Questions Q24-Q28 are for follow-up and have been removed from the data. The institutions (Q17) and funding sources (Q18) have been provided in a separate file as this could be used to identify respondents. Q17, Q18 and Q19-Q23 have all been independently shuffled.</p><p>The data has been made available as Comma Separated Values (CSV) with the question number as the header of each column and the encoded responses in the column below. To see what the question and the responses correspond to you will have to consult the survey-results-key.csv which decodes the question and responses accordingly. </p><p><strong>A pdf copy of the survey questions is </strong><a href="https://github.com/softwaresaved/esrc-software-study/blob/main/Docs/esrc-survey.pdf"><strong>available</strong></a><strong> on GitHub.</strong></p><p>The survey data has been decoupled into:</p><ul><li>survey-results-key.csv - maps a question number and the responses to the actual question values.</li><li>q1-16-survey-results.csv- the non-demographic component of the survey responses (Q1-Q16).</li><li>q19-23-demographics.csv - the demographic part of the survey (Q19-Q21, Q23).</li><li>q17-institutions.csv - the institution/location of the respondent (Q17).</li><li>q18-funding.csv - funding sources within the last 5 years (Q18).</li></ul><p>Please note the code that has been used to do the analysis will not run with the decoupled survey data. </p><h3>Other data files included</h3><ul><li>CleanedLocations.csv - normalised version of the institutions that the survey respondents volunteered.</li><li>DTPs.csv - information on the UKRI Doctoral Training Partnerships (DTPs) scaped from the UKRI <a href="https://esrc.ukri.org/skills-and-careers/doctoral-training/doctoral-training-partnerships/doctoral-training-partnership-dtp-contacts/">DTP contacts</a> web page in October 2021.</li><li>projectsearch-1646403729132.csv.gz - data snapshot from the <a href="https://gtr.ukri.org/">UKRI Gateway to Research</a> released on the 24th February 2022 made available under an <a href="https://www.nationalarchives.gov.uk/doc/open-government-licence/version/3/">Open Government Licence</a>.</li><li>locations.csv - latitude and longitude for the institutions in the cleaned locations.</li><li>subjects.csv - research classifications for the ESRC projects for the 24th February data snapshot.</li><li>topics.csv - topic classification for the ESRC projects for the 24th February data snapshot.</li></ul><h3>Interview transcripts</h3><p>The interview transcripts have been anonymised and converted to markdown so that it's easier to process in general. List of interview transcripts:</p><ul><li>1269794877.md</li><li>1578450175.md</li><li>1792505583.md</li><li>2964377624.md</li><li>3270614512.md</li><li>40983347262.md</li><li>4288358080.md</li><li>4561769548.md</li><li>4938919540.md</li><li>5037840428.md</li><li>5766299900.md</li><li>5996360861.md</li><li>6422621713.md</li><li>6776362537.md</li><li>7183719943.md</li><li>7227322280.md</li><li>7336263536.md</li><li>75909371872.md</li><li>7869268779.md</li><li>8031500357.md</li><li>9253010492.md</li></ul><h3>Data Management Plan</h3><p>The study's Data Management Plan is provided in PDF format and shows the different data sets used throughout the duration of the study and where they have been deposited, as well as how long the SSI will keep these records. </p>
Reassessing science communication for effective farmland biodiversity conservation.
<p>Data set and code for analysing a communication case study about biodiversity conservation and farming in the European decision-making environment. It includes: (1) a literature corpus, consisting of 5988 digital press releases and news texts, covering the period of 2015-2020, from 40 different organizations involved in European farming and food decision-making processes; (2) R code scripts for analysis and graphic representation.</p>
Figures 10-12 in A new genus and species of armored scale insect (Hemiptera: Diaspididae) from Australia found in the historic Koebele Collection of the California Academy of Sciences John W. Dooley III
Figures 10-12. Diaspididae spp., habitus, detail of venter of L1 lobes, and pygidium ventral (left) and dorsal (right). 10) Dichosoma convexa (after Brimblecombe 1957). 11) Duplaspidiotus claviger (after Ferris 1937). 12) Eulaingia stenophyllae (after Borchsenius and Williams 1963).
Figure 3 in A new genus and species of armored scale insect (Hemiptera: Diaspididae) from Australia found in the historic Koebele Collection of the California Academy of Sciences John W. Dooley III
Figure 3. Protomorgania koebelei adult female (pygidium). A) L1 lobes fused ventrally, appressed dorsally; B) single simple plate between L1 and position of L2 seta; C) anal pore; D) sclerotized arch; E) chitinized and finely stippled cuticle around vulva.
Figure 2 in A new genus and species of armored scale insect (Hemiptera: Diaspididae) from Australia found in the historic Koebele Collection of the California Academy of Sciences John W. Dooley III
Figure 2. Protomorgania koebelei adult female (thorax and Abdomen). A) anterior perispiracular pores; B) dorsal microducts; C) dorsal microducts, magnified; D) roughened cuticle.
Figure 19 in A new genus and species of armored scale insect (Hemiptera: Diaspididae) from Australia found in the historic Koebele Collection of the California Academy of Sciences John W. Dooley III
Figure 19. Pseudotargionia glandulosa, habitus, detail of venter of L1 lobes, and pygidium ventral (left) and dorsal (right) (after Ferris 1937).
Figure 1. Protomorgania koebelei adult female. A in A new genus and species of armored scale insect (Hemiptera: Diaspididae) from Australia found in the historic Koebele Collection of the California Academy of Sciences John W. Dooley III
Figure 1. Protomorgania koebelei adult female. A) habitus; B) tubercle; C) anterior spiracle; D) posterior spiracle; E) pygidial lobes; F) slide mounted female habitus; G) habitus on host; H) close-up of habitus on host
Figures 7-9 in A new genus and species of armored scale insect (Hemiptera: Diaspididae) from Australia found in the historic Koebele Collection of the California Academy of Sciences John W. Dooley III
Figures 7-9. Diaspididae spp., habitus, detail of venter of L1 lobes, and pygidium ventral (left) and dorsal (right) (after Brimblecombe 1957). 7) Diaphoraspis orbata. 8) Diaspidopus distinctus. 9) Diastolaspis novata.
Figures 4-6 in A new genus and species of armored scale insect (Hemiptera: Diaspididae) from Australia found in the historic Koebele Collection of the California Academy of Sciences John W. Dooley III
Figures 4-6. Diaspididae spp., habitus, detail of venter of L1 lobes, and pygidium ventral (left) and dorsal (right) (after Brimblecombe 1957). 4) Achorphora obliqua. 5) Acontonidia triangulari. 6) Aspidonymus woodwardi.
Figures 16-18 in A new genus and species of armored scale insect (Hemiptera: Diaspididae) from Australia found in the historic Koebele Collection of the California Academy of Sciences John W. Dooley III
Figures 16-18. Diaspididae spp., habitus, detail of venter of L1 lobes, and pygidium ventral (left) and dorsal (right). 16) Neoleonardia extensa (after Ferris 1938). 17) Neomorgania eucalypti (after Ferris 1937). 18) Pseudaonidia duplex (after Ferris 1937).
HoneyBee: Progressive Instruction Finetuning of Large Language Models for Materials Science
<p>We propose an instruction-based process for trustworthy data curation in materials science (MatSci-Instruct), which we then apply to finetune a LLaMa-based language model targeted for materials science (HoneyBee). MatSci-Instruct helps alleviate the scarcity of relevant, high-quality materials science textual data available in the open literature, and HoneyBee is the first billion-parameter language model specialized to materials science. In MatSci-Instruct we improve the trustworthiness of generated data by prompting multiple commercially available large language models for generation with an Instructor module (e.g. Chat-GPT) and verification from an independent Verifier module (e.g. Claude). Using MatSci-Instruct, we construct a dataset of multiple tasks and measure the quality of our dataset along multiple dimensions, including accuracy against known facts, relevance to materials science, as well as completeness and reasonableness of the data. Moreover, we iteratively generate more targeted instructions and instruction-data in a finetuning-evaluation-feedback loop leading to progressively better performance for our finetuned HoneyBee models. Our evaluation on the MatSci-NLP benchmark shows HoneyBee's outperformance of existing language models on materials science tasks and iterative improvement in successive stages of instruction-data refinement. We study the quality of HoneyBee's language modeling through automatic evaluation and analyze case studies to further understand the model's capabilities and limitations. Our code and relevant datasets are publicly available at https://github.com/BangLab-UdeM-Mila/NLP4MatSci-HoneyBee.</p>
SeaPaCS graphic elaboration of the Protocol for marine micro-plastic collection and monitoring in citizen science and for building a L.A.D.I. trawling tool
<p>This is a graphic elaboration (in Italian) of the protocol "SeaPaCS deliverable - protocol for plastic monitoring in citizen science" in English and Italian is a deliverable of the SeaPaCS project (Participatory Citizen Science Against Marine Pollution), funded by IMPETUS (project ID 101058677). The protocol and the visual elaboration has been freely adapted from "<i>LADI and the Trawl</i>" by Coco Coyle with Melissa Novaceski, Emily Wells and Max Liboiron, as published by the Civic Laboratory for Environmental Action Research, August 2016. The graphic elaboration (as the protocol) in both languages, consists of three parts: 1) how to build a DIY low cost manta trawl device (LADI - Low-Tech Aquatic Detection Debris Instrument) to monitor plastic pollution, adjusted to materials availability and costs in Italy; 2) how to monitor (the sampling itself and towing procedure); and 3) how to categorize plastic debris back on land. </p>
science-gallery
<h4><strong>Source of original data</strong></h4><ul><li><a href="http://www.sociopatterns.org/datasets/infectious-sociopatterns-dynamic-contact-networks/">Infectious SocioPatterns dynamic contact networks</a></li></ul><h4><strong>References</strong></h4><p>If you use this data, please cite the following:</p><ul><li><a href="https://doi.org/10.1016/j.jtbi.2010.11.033">What's in a crowd? Analysis of face-to-face behavioral networks</a>. Isella et al., Journal of Theoretical Biology (2011).</li><li><a href="http://www.sociopatterns.org/">The SocioPatterns collaboration</a></li></ul>
SCOPING REVIEW OF EMPIRICAL LITERATURE ON EVALUATION OF EDITORIAL POLICIES SUPPORTING OPEN SCIENCE PRACTICES IN SCHOLARLY JOURNALS
<p>Excel file containing the description of methods & materials used in the scoping review, the whole dataset of included studies, and respective descriptive metadata.</p>
Reporting Database of the INCENTIVE project's monitoring and evaluation activities of the Citizen Science Hubs pilot operation (M16-M31)
<p>The current excel file constitutes the INCENTIVE Reporting Database, meaning a repository of the data that were accumulated in WP4 (described within Deliverable 4.1). The results presented here are the main input for the elaboration of Deliverable 4.2 and Deliverable 4.3.</p>
Mixed population trends inside a California protected area: Evidence from long-term community science monitoring
<p><span>Protected areas are one of the most widespread and accepted conservation interventions, yet their population trends are rarely compared to regional trends to gain insight into their effectiveness. Here, we leverage two long-term community science datasets to demonstrate mixed effects of protected areas on long-term bird population trends. We analyzed 31 years of bird transect data recorded by community volunteers across all major habitats of Stanford University's Jasper Ridge Biological Preserve to determine the population trends for a sample of 66 species. We found that nearly a third of species experienced long-term declines, and on average, all species declined by 12%. Further, we averaged species trends by conservation status and key life history attributes to identify correlates and possible drivers of these trends. Observed increases in some cavity-nesters and declines of scrub-associated species suggest that long-term fire suppression may be a key driver, reshaping bird communities through changes in forest and chaparral structure and composition. Additionally, we compared our results to those of the North American Breeding Bird Survey's Central California Coast region (n = 55 species) to place Jasper Ridge in a broader context. Most species experienced similar directional population trends inside vs. outside of the preserve, and only eight species (14.5%) did better inside this small, protected area. Therefore, we must identify relevant management strategies for declining populations and explicitly consider how existing protected areas target and manage each species. Further, this analysis underscores the importance of local and national community science for revealing nuanced long-term bird population trends.</span></p>
Simulations used in the Ocean Science Journal submission titled "Internal and forced ocean variability in the Mediterranean Sea " by Benincasa et al., 2024
<p>Temperature (votemper) and current speed datasets from the EAS5 (Clementi et al., <em>Mediterranean Sea Analysis and Forecast (CMEMS MED-Currents, EAS5 system),</em> 2019; Coppini et al., <em>The Mediterranean forecasting system. Part I: evolution and performance</em>, EGUsphere, pp. 1–50, 2023) simulations used in the manuscript titled "<em>Internal and forced ocean variability in the Mediterranean Sea"</em> and submitted to the journal Ocean Science by Benincasa et al. </p> <p>The daily fields are at 2 depth levels ( 0 = 0 m, 2 = 30 m) and in the 2 seasons (JFMA = winter, JASO = summer) for the entire Mediterranean Sea. The vertical profile of the temperature field up to about 950 m depth is available at 8 locations distributed over the basin. The depth levels are found in <em>depth.pkl</em>: the first column represents the depth of the various levels, whereas the second is the increments between 2 consecutive depth levels. </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.