Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
593
datasets available to search
ShareScore release 0.9.0
Dataset results
593 results for “workshop”
IML workshop challenge on jet mass regression
<p>This dataset is associated with the LPCC IML (Lhc Physics Center at Cern Inter-experimental Machine Learning) working group. It was produced for the second IML annual workshop (April 2018).</p> <p>This dataset is part of a machine learning "challenge" on jet mass regression at future circular collider (FCC) conditions. Further details can be found on the challenge page, here:</p> <p><a href="https://gitlab.cern.ch/IML-WG/IML_challenge_2018/wikis/home">https://gitlab.cern.ch/IML-WG/IML_challenge_2018/wikis/home</a></p>
meshography workshop
<p>Leveraging the meshing of frog gastrocnemius muscle from digital photography using artificial neural networks.</p> <p>This data collection supports the finding in our paper submitted to the INTERNATIONAL JOURNAL OF NUMERICAL METHODS IN BIOMEDICAL ENGINEERING", on Oct.20,2018.</p> <p>There are two folders (left and right) containing 15 images each, processed for edge detection of the left and right object contours.</p> <p>There are also some Matlab scripts and functions, for data preparation and training of a Bayesian backpropagation neural network.</p> <p>The script (kod2.m) is associated with the polynomial fitting procedure to the contours, and it utilized the function (fiterror.m) in the process.</p> <p>The file (polinomtablo4th.mat) contains the output of the fitting procedure, and is used by the neural networking script (plotneur4th.m).</p> <p>Finally, the unit cylinder mesh are processed through the network, to produce the 1767 node, 1440 hexahedral-type element FEAP mesh (wIplant2), of the muscle shown in (mesh_view2.png).</p> <p>The codes are based in the MATLAB R15b, image processing and neural network toolboxes.</p>
AVANT Workshop "Antibiotic-free pork production: realiy or chimera?"
<p>The aim of the workshop organized by the EU Innovation Action AVANT is to facilitate dialogue and consensus on specific themes pertinent to the future of antibiotic-free pork production. The involvement of significant partners such as SEGES, COOPERL and FVE, alongside other relevant national and international stakeholders, will offer a diverse array of perspectives and expertise crucial for addressing the multifaceted challenges and opportunities of this type of production in the years to come. One significant outcome could be the development of a position paper encompassing the collective insights, agreements, and recommendations arising from the discussions held during the workshop.</p>
crowdsourced body parameters of workshop attendants at the Helmholtz MT ARD ST3 meeting
<p>This data set was crowdsourced at the 2021 Helmholtz MT ARD ST3 meeting from attendants of the Machine Learning Tutorial on Sep 30, 2021. For more details on the event, see<br> https://indico.desy.de/event/28823/</p> <p>The CSV contains 4 columns:</p> <p>- is_female : fill with 1 if participant is female, filled with 0 if not female</p> <p>- shoesize_europe : your european shoesize</p> <p>- weight_kg : your weight in kilograms</p> <p>- height_cm : your height in centimeters</p> <p>Participants were encouraged to +1 or -1 to individual body properties in case they do not feel confident providing their true numbers.</p>
BIP! NDR (NoDoiRefs): a dataset of citations from papers without DOIs in computer science conferences and workshops
<h2>Overview</h2> <p>In the field of Computer Science, conference and workshop papers serve as important contributions, carrying substantial weight in research assessment processes, compared to other disciplines. However, a considerable number of these papers are not assigned a Digital Object Identifier (DOI), hence their citations are not reported in widely used citation datasets like OpenCitations and Crossref, raising limitations to citation analysis. While the Microsoft Academic Graph (MAG) previously addressed this issue by providing substantial coverage, its discontinuation has created a void in available data.</p> <p>BIP! NDR aims to alleviate this issue and enhance the research assessment processes within the field of Computer Science. To accomplish this, it leverages a workflow that identifies and retrieves Open Science papers lacking DOIs from the DBLP Corpus, and by performing text analysis, it extracts citation information directly from their full text.</p> <p>The current version of the dataset contains <em>~4.3M citations</em> made by approximately <em>211K open access Computer Science conference or workshop papers</em> that, according to DBLP, do not have a DOI. The DBLP snapshot used for this version was the one released on <em>September 2025</em>. </p> <h2>Dataset files</h2> <h3>1. Core Non-DOI Citation Dataset - bip_ndr_{version}.tar.gz</h3> <p>The dataset is formatted as a JSON Lines (JSONL) file (one JSON Object per line) to facilitate file splitting and streaming. </p> <p>Each JSON object has three main fields:</p> <ul> <li> <p>“_id”: a unique identifier,</p> </li> <li> <p>“citing_paper”, the “dblp_id” of the citing paper,</p> </li> <li> <p>“cited_papers”: array containing the objects that correspond to each reference found in the text of the “citing_paper”; each object may contain the following fields:</p> <ul> <li> <p>“dblp_id”: the “dblp_id” of the cited paper. Optional - this field is required if a “doi” is not present.</p> </li> <li> <p>“doi”: the doi of the cited paper. Optional - this field is required if a “dblp_id” is not present.</p> </li> <li> <p>“bibliographic_reference”: the raw citation string as it appears in the citing paper.</p> </li> </ul> </li> </ul> <p>Changes from previous version:</p> <ul> <li>Added more papers from DBLP.</li> </ul> <h3>2. Citation Intents Dataset - bip_ndr_ci_{version}.tar.gz</h3> <p>This file enriches the BIP! NDR dataset with citation-level intent classification.<br>It preserves the same base structure of the previous file, while adding a nested array of "citations" with each element of "cited_papers".</p> <p>Each "citation" provides the local textual context, section, and intent of the citation in the following format:</p> <ul> <li>"citation_id": Unique identifier in the format {citing_id}>{cited_id}_CIT{index} linking the citing and cited entities.</li> <li>"section": The section of the citing paper where the citation occurs (e.g., Introduction, Methods, Results).</li> <li>"intent": Inferred purpose of the citation based on textual context (see classification schema below).</li> </ul> <p>The "intent" field follows the SciCite classification schema, which categorizes citations into three high-level functional types:</p> <ol> <li>background information: The citation states, mentions, or points to the background information giving more context about a problem, concept, approach, topic, or importance of the problem in the field.</li> <li>method: Making use of a method, tool, approach or dataset.</li> <li>results comparison: Comparison of the paper's results/findings with the results/findings of other work.</li> </ol> <p>The classification is done with the <a href="https://huggingface.co/sknow-lab/Qwen2.5-14B-CIC-SciCite">Qwen2.5-14B-CIC-SciCite fine-tuned Large Language Model, published by Athena RC</a>. </p> <p>Changes from previous version: </p> <ul> <li>Added more papers with intent</li> </ul>
TOPS Graphic for the 2022 PACE Applications Workshop
<p>This graphic was made to be displayed on a TV screen during the NASA 2022 PACE (Plankton, Aerosol, Cloud, ocean Ecosystem) Applications Workshop held on September 14-15, 2022. </p>
Thoth's GraphQL API Workshop
<p>As Thoth continues to enable presses to manage metadata for their open access books and export it in a number of different formats to various platforms, catalogues and other dissemination channels, we want to make sure all its users are fully aware of Thoth’s open API capabilities. Led by Thoth’s software engineer Javier Arias, the team of COPIM’s Work Package 5 will guide users through its GraphQL API and demo how to query the API, write queries, and export the resulting data into a spreadsheet. Ultimately, this workshop aims to familiarize non-technical users who would like to learn how to directly retrieve data from Thoth’s API with Thoth’s versatile capabilities.</p> <p>Details of the workshop available at: <a href="https://www.copim.ac.uk/outputs/events/220819-thoth-api-workshop/">https://www.copim.ac.uk/outputs/events/220819-thoth-api-workshop/</a></p> <p>Learn more about Thoth at <a href="https://thoth.pub/">https://thoth.pub/</a></p>
US National Native Bee Monitoring RCN Data Management Workshop: Public Domain Videos
<p>The US National Native Bee Monitoring Research Coordination Network (RCN) held a two-day workshop on data management best practices for native bee inventory, survey, and monitoring data on March 28 and 30, 2023. Videos in this data set were played at the workshop. These videos are released into the public domain. This data set includes the following videos:</p> <ul> <li>Ecological Metadata Standards to Enable Data Reuse by Julien Brun</li> <li>Useful Photo Management for Bee Species by Sam Droege</li> <li>Trait Data Models and Vocabulary by Jen Hammock</li> <li>Symbiota: open-source community portals for insect data management by Andrew Johnston</li> <li>Moving data from the field to the world by Jonathan Koch</li> <li>Exploring data using Discover Life by Clare Maffei</li> <li>USDA Data Sharing Policies and Opportunities by Cynthia Sims Parr</li> <li>Why Share Species Interaction Data? by Jorrit H. Poelen</li> <li>Big-Bee: Sharing Bee Interactions & Traits by Katja C. Seltmann</li> <li>Let’s talk about data by Katja C. Seltmann</li> <li>Responsible use of museum specimens & their data by Erika M. Tucker</li> </ul>
US National Native Bee Monitoring RCN Data Management Workshop: CC BY Videos
<p>The US National Native Bee Monitoring Research Coordination Network (RCN) held a two-day workshop on data management best practices for native bee inventory, survey, and monitoring data on March 28 and 30, 2023. Videos in this data set were played at the workshop. These videos are released with a CC BY license. Please cite the presenter(s) of the video(s) you use. This data set includes the following videos:</p> <ul> <li>The ABeeCs of Data Attribution: Please use your magic words by David Bloom</li> <li>Darwin Core Geography: How to make your locality data complete and accurate by David Bloom</li> <li>Best Practices for Managing Native Bee Molecular Data by Michael G. Branstetter</li> <li>A Trait Database for Bees by Elizabeth A. Crisfield</li> <li>Biotic interaction data and invasive species assessment by Quentin Groom</li> <li>OpenTraits Network (OTN) & TRY Plant Trait Database by Jens Kattge</li> <li>Semantics modeling of phenotypic trait data with ontologies by Diego S. Porto</li> <li>WorldFAIR: towards making plant-pollinator data FAIR by Maarten Trekels</li> </ul>
Research Data Management and data protection in the Social Sciences [Workshop recording]
<p>This online workshop organized by The Austrian Social Science Data Archive (AUSSDA) focused on the Research Data Management basics, Data Management Plans and common data protection issues in the Social Sciences.</p> <p>The first part of the workshop was dedicated to RDM basics and Data Management Plans (DMPs). In many projects, DMPs are mandatory deliverables that need to be submitted at the beginning of a project and are updated throughout the project life cycle. During the workshop, it was explained which aspect funders expect to be part of DMPs in Social Sciences and how researchers can benefit from (writing) these documents.</p> <p>In the second part of the workshop, data protection issues that are common in Social Sciences were addressed and how they can be handled. In particular, differences in the curation of quantitative and qualitative data need in order to comply with data protection regulations in general and AUSSDA deposit guidelines in particular. Presentation on how AUSSDA scans quantitative data for potential data protection violations using STATA and gives participants the opportunity to test the code on their own data and devices.</p> <p>The video is available on<a href="https://www.youtube.com/watch?v=DhiL9J-Iwqg"> the CESSDA Training YouTube channel</a>.</p> <p> </p>
Anonymisation for data sharing in practice [Online Workshop. Recording]
<p>The goal of this event was to show trainers the tools they need to teach the fundamentals of data anonymisation and disclosure control in training sessions while also giving them hands-on experience with current open source technologies (sdcMicro). Some of the concepts and techniques presented, included k-anonymity, top/bottom coding and aggregation with practical examples and recommendations on incorporating anonymisation into research designs.</p> <p> </p> <p>The video is available on<a href="https://www.youtube.com/watch?v=JeJ6OOxXZwo&t=328s"> the CESSDA Training YouTube channel</a>.</p> <p> </p> <p> </p>
How to Ensure Researchers Share Their FAIR Data: Practical Tips and Tools [Online Workshop, Recording]
<p>The online hands-on workshop was aimed at trainers and support staff covering critical elements of data sharing and available tools and resources for supporting Open Science including:<br> • Open Science resources and Data Management Planning<br> • Consent and Ethical considerations<br> • Legislation and Licence frameworks<br> The objectives of the workshop were i) to raise awareness of key tools and resources available for Open Science training ii) to enable a platform to exchange ideas regarding key training topics and iii)n to provide training materials and worksheets for future reuse.<br> The workshop consisted of presentations, demos, a roundtable discussion on ethical considerations, a showcase of licence frameworks at different European archives and an exercise with all participants fostering an exchange of experiences focused on learnt lessons.</p> <p>The video is available on<a href="https://www.youtube.com/watch?v=uztTCRFRZHg"> the CESSDA Training YouTube channel</a>.</p>
FEDORA.Transcription and photos from the Intensive Creative Workshop I on Acceleration, Complexity and Interdisciplinarity
<p>The dataset contains transcription and photos from the Intensive Creative Workshop I on Acceleration, Complexity, and Interdisciplinarity realised within the FEDORA project. The workshop was held in Bologna on 9th July 2021 and involved 16 participants with diverse backgrounds: science communicators, artists, graphic designers, novelists, photographers, researchers, school teachers, and video producers. They worked in two teams on the barriers, challenges and problems associated with introducing a transdisciplinary approach in science education and also with facing the issues of complexity and futurization by using and owning new languages.</p> <p>The results are detailed in the deliverable D2.2: First draft of recommendations on “new languages” (Confidential Deliverable) for the design of materials; in the D2.3<a href="https://doi.org/10.5281/zenodo.7518940"> </a>Multimedia report of Intensive workshop I on acceleration, complexity and interdisciplinarity <a href="https://doi.org/10.5281/zenodo.7518940">https://doi.org/10.5281/zenodo.7518940</a> and they are part of the framework developed for deliverable D2.5: Framework for aligning science education with society: the search for new languages and narratives to enhance imagination and the capacity to talk about contemporary challenges:<a href="https://doi.org/10.5281/zenodo.7519100"> https://doi.org/10.5281/zenodo.7519100</a></p> <p> </p>
Why, what and how do European healthcare managers use performance data? Results of a survey and workshop among members of the European Hospital and Healthcare Federation (Data set; anonymised)
<p>The dataset presents results of a descriptive cross-sectional study based on a survey, delivered through an online self-reported questionnaire. The questionnaire was distributed to managers of hospitals and other health care organisations in a purposive sample of participants to the Exchange Programmes of the European Hospital and Health Care Federation (HOPE) eliciting information on the actual use of performance data in hospitals and other healthcare organisations in Europe in 2019.<br> Data collected through the online questionnaire was analysed using univariate descriptive statistics. Analyses were conducted using the R statistical program version 3.6.1. Respondents were, for certain parts of the analysis, sub-grouped by their reported managerial position and experience, as well as the type of organisation they work for. Analysis was done on a full sample of respondents, including the primary, 2019 HOPE Exchange Programme participants, and the secondary study population, 2015-2018 Exchange Programme alumni and local hosts.</p>
Teaser SUNRISE Stakeholder Workshop
<p>Promotion video: The next big event for the SUNRISE community will take place on 17-18 June, 2019 in Brussels.</p>
EI Single-Cell RNA-Seq Workshop 2020
<p>Datasets to be used for the "Single-Cell RNA-Seq Workshop 2020" at the Earlham Institute, Norwich, UK.</p>
The Software Sustainability Institute's Collaborations Workshop 2015 (CW15) attendees computational tools dataset
<p>Contains the question, raw data, and cleaned data for producing the most used software word cloud for those who attended the Software Sustainability Institute's Collaborations Workshop 2015 (CW15) held at the Oxford e-Research Institute, Oxford, UK from 25-27 March 2015</p>
CIS OCR Workshop v1.0: OCR and postcorrection of early printings for digital humanities
<p>The 2-day CIS OCR Workshop on "OCR and postcorrection of early printings for digital humanities" originally held at LMU, Munich 14/15 September 2015 (see http://www.cis.lmu.de/ocrworkshop).</p> <p>Release date: 2016-02-25</p> <p><br /> CIS OCR Workshop by Uwe Springmann, Florian Fink is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.</p>
Figure 9b. from: Defining the Scholarly Commons - Reimagining Research Communication. Report of Force11 SCWG Workshop, Madrid, Spain, February 25-27, 2016 - Research Ideas and Outcomes 2: e9340 (26 May 2016) https://doi.org/10.3897/rio.2.e9340
<p>Figure 9b. - Visualization showing one example of development of a group's vision and progress to group principlesFigure 9a.One group's vision as a collection of ideas (session 7)Figure 9b.One group's vision as a interconnected elements (triples, derived after session 7)Figure 9c.One group's suggested principles (session 11)</p>
Figure 7c. from: Defining the Scholarly Commons - Reimagining Research Communication. Report of Force11 SCWG Workshop, Madrid, Spain, February 25-27, 2016 - Research Ideas and Outcomes 2: e9340 (26 May 2016) https://doi.org/10.3897/rio.2.e9340
Figure 7c. - Use of Trello during the workshopFigure 7a.Example of Trello card with tags. "using the public domain" - the idea name; #G2 - the idea came from the Group 2; #viz - the idea is ready to be included in the visualization; #triple - the idea has a link to another idea in the group's vision. Figure 7b.Example of Trello card comment with link to another idea. #dep_on - idea depends on another idea; "https://trello.com/c/RaNcbKSe" - is the short link to the Trello card with the connected idea. In the visualization it would be represented as a line connecting two ideas.Figure 7c.Screenshot of public Trello board <br>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.