Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

6,766

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

6,766 results for “project”

Learn how ShareScore rates datasets ↗
zenodo44/100

Dron view of the 2021 field trial for BRESOV project in SERIDA

<p>This video shows the field trial with 311 snap bean lines in organic conditions.</p> <p>The field trial had three plots per line with ten plants distributed in 1 m.</p> <p>This wide diversity is being evaluated considering morpho-agronomic characters.</p> <p>This work is part of the BRESOV project funded by the EU (Grant agreement ID:&nbsp;774244 ).</p> <p>This video was recorded at the SERIDA facilities, Villaviciosa, Asturias, Spain in July 2021.</p> <p>DOI&nbsp;&nbsp;10.5281/zenodo.5573917</p> <p><strong>Also available in the link:&nbsp;</strong> https://www.youtube.com/watch?v=dJa9UDFQnOE&amp;t=5s</p>

opencc-by-4.0Jul 2021View details →
zenodo44/100

Public Utility Data Liberation Project (PUDL) Data Release

<h2><strong>v2025.10.0 (2025-10-14)</strong></h2> <p>This is a regular monthly data release, primarily intended to ensure that PUDL has the most up-to-date EIA-860M data. It also happens to include final EIA-860 data for 2024, and some newly integrated EIA-923 financial data and PHMSA natural gas data. See below for details.</p> <h3>Expanded Data Coverage</h3> <h4>EIA-860</h4> <ul> <li> <p>Updated EIA-860 with final release data from 2024. See issue <a href="https://github.com/catalyst-cooperative/pudl/issues/4616">#4616</a> and PR <a href="https://github.com/catalyst-cooperative/pudl/pull/4617">#4617</a>.</p> </li> </ul> <h4>EIA-860M</h4> <ul> <li> <p>Updated EIA-860M monthly generator report with newly published data for August of 2025. See issue <a href="https://github.com/catalyst-cooperative/pudl/issues/4639">#4639</a> and PR <a href="https://github.com/catalyst-cooperative/pudl/pull/4638">#4638</a>.</p> </li> </ul> <h3>New Data</h3> <h4>PHMSA</h4> <ul> <li> <p>Added eight transformed table containing annual data from PHMSA natural gas distributors from 1970 to the present. Note that these containing mostly numeric values are named as <code><span>_core</span></code> - indicating that these tables have not been fully cleaned and validated. We&rsquo;ve published these tables to make the 50+ years of PHMSA data we&rsquo;ve extracted and mapped available for others to use and for contributors to more easily improve incrementally. See <a href="https://github.com/catalyst-cooperative/pudl/issues/3770">#3770</a> and <a href="https://github.com/catalyst-cooperative/pudl/pull/4005">#4005</a>.</p> </li> <li> <p>The first cleaned table, <code><span>core_phmsagas__distribution_operators</span></code> has been added to our PUDL database. Thanks to <a href="https://github.com/sponsors/seeess1">@seeess1</a> for all of your work on this!</p> </li> </ul> <h4>EIA 923</h4> <ul> <li> <p>Thanks to contributions from <a href="https://github.com/sponsors/alexclippinger">@alexclippinger</a>, we&rsquo;ve added cleaned EIA923 Schedule 8B Financial Information to the PUDL database as <a href="https://catalystcoop-pudl.readthedocs.io/en/v2025.10.0/data_dictionaries/pudl_db.html#i-core-eia923-yearly-byproduct-expenses-and-revenues"><span>_core_eia923__yearly_byproduct_expenses_and_revenues</span></a>. Once harvested, this table will be replaced with a well-normalized version of the same data, but it is being published in this form until then. See <a href="https://github.com/catalyst-cooperative/pudl/issues/4099">#4099</a> and <a href="https://github.com/catalyst-cooperative/pudl/issues/2448">#2448</a>, and <a href="https://github.com/catalyst-cooperative/pudl/pull/4636">#4636</a>.</p> </li> </ul> <h3>Documentation</h3> <ul> <li> <p>Added data source pages for:</p> <ul> <li> <p><a href="https://catalystcoop-pudl.readthedocs.io/en/v2025.10.0/data_sources/censuspep.html"><span>Population Estimates Program's (PEP) Federal Information Processing Series (FIPS) Codes</span></a>; see issue <a href="https://github.com/catalyst-cooperative/pudl/issues/4375">#4375</a> and PR <a href="https://github.com/catalyst-cooperative/pudl/pull/4622">#4622</a>.</p> </li> <li> <p><a href="https://catalystcoop-pudl.readthedocs.io/en/v2025.10.0/data_sources/sec10k.html"><span>U.S. Securities and Exchange Commission (SEC) Form 10-K</span></a>; see issue <a href="https://github.com/catalyst-cooperative/pudl/issues/4329">#4329</a>, <a href="https://github.com/catalyst-cooperative/pudl/issues/4347">#4347</a> and PR <a href="https://github.com/catalyst-cooperative/pudl/pull/4562">#4562</a>.</p> </li> </ul> </li> </ul> <h3>New Data Tests &amp; Data Validations</h3> <ul> <li> <p>After investigating some modest discrepancies between our imputed hourly electricity demand and prior work by <a href="https://github.com/sponsors/truggles">@truggles</a> &amp; <a href="https://github.com/sponsors/awongel">@awongel</a>, we&rsquo;re removing the &ldquo;EXPERIMENTAL&rdquo; warning label that we had on those tables. See <a href="https://github.com/catalyst-cooperative/pudl-examples/pull/10">our discussion about the imputation results in the PUDL Examples repo</a>. The <a href="https://www.kaggle.com/code/catalystcooperative/06-pudl-imputed-electricity-demand">associated notebook is available on Kaggle</a></p> <p>This relates to the PUDL imputed demand values in following tables:</p> <ul> <li> <p><a href="https://catalystcoop-pudl.readthedocs.io/en/v2025.10.0/data_dictionaries/pudl_db.html#out-eia930-hourly-operations"><span>out_eia930__hourly_operations</span></a></p> </li> <li> <p><a href="https://catalystcoop-pudl.readthedocs.io/en/v2025.10.0/data_dictionaries/pudl_db.html#out-eia930-hourly-subregion-demand"><span>out_eia930__hourly_subregion_demand</span></a></p> </li> <li> <p><a href="https://catalystcoop-pudl.readthedocs.io/en/v2025.10.0/data_dictionaries/pudl_db.html#out-eia930-hourly-aggregated-demand"><span>out_eia930__hourly_aggregated_demand</span></a></p> </li> </ul> </li> </ul> <h3>Deprecations</h3> <ul> <li> <p>We have finally shut down our long-suffering <a href="https://datasette.io">Datasette</a> deployment, but are still working on achieiving feature parity in the new <a href="https://viewer.catalyst.coop">PUDL Data Viewer</a>. We have <a href="https://github.com/catalyst-cooperative/eel-hole/issues/36">an epic tracking our progress</a>. See issue <a href="https://github.com/catalyst-cooperative/pudl/issues/4481">#4481</a> and PR <a href="https://github.com/catalyst-cooperative/pudl/pull/4605">#4605</a> for the removal of Datasette references within the main PUDL repo.</p> </li> </ul> <h2><strong>Other PUDL v2025.10.0 Resources</strong></h2> <ul> <li><a href="https://catalystcoop-pudl.readthedocs.io/en/v2025.10.0/data_dictionaries/pudl_db.html">PUDL v2025.10.0 Data Dictionary</a></li> <li><a href="https://catalystcoop-pudl.readthedocs.io/en/v2025.10.0/">PUDL v2025.10.0 Documentation</a></li> <li><a href="https://registry.opendata.aws/catalyst-cooperative-pudl/">PUDL in the AWS Open Data Registry</a></li> <li>PUDL v2025.9.1 in a free, public AWS S3 bucket: s3://pudl.catalyst.coop/v2025.10.0/</li> <li>PUDL v2025.9.1 in a requester-pays GCS bucket: gs://pudl.catalyst.coop/v2025.10.0/</li> <li><a href="https://doi.org/10.5281/zenodo.17352325">Zenodo archive of the PUDL GitHub repo for this release</a></li> <li><a href="https://github.com/catalyst-cooperative/pudl/releases/tag/v2025.10.0">PUDL v2025.10.0 release on GitHub</a></li> <li><a href="https://pypi.org/project/catalystcoop.pudl/2025.10.0">PUDL v2025.10.0 package in the Python Package Index (PyPI)</a></li> </ul> <h2><strong>Contact Us</strong></h2> <p><strong>If you're using PUDL, we would love to hear from you!</strong> Even if it's just a note to let us know that you exist, and how you're using the software or data. Here's a bunch of different ways to get in touch:</p> <ul> <li><a href="https://github.com/catalyst-cooperative">Follow us on GitHub</a></li> <li>Use the <a href="https://github.com/catalyst-cooperative/pudl/issues">PUDL Github issue tracker</a> to let us know about any bugs or data issues you encounter</li> <li><a href="https://github.com/orgs/catalyst-cooperative/discussions">GitHub Discussions</a> is where we provide user support.</li> <li>Watch our <a href="https://github.com/orgs/catalyst-cooperative/projects/9">GitHub Project</a> to see what we're working on.</li> <li>Email us at <a href="mailto:hello@catalyst.coop">hello@catalyst.coop</a> for private communications.</li> <li>On Mastodon: <a href="https://mastodon.energy/@catalystcoop">@CatalystCoop@mastodon.energy</a></li> <li>On BlueSky: <a href="https://bsky.app/profile/catalyst.coop">@catalyst.coop</a></li> <li>On Twitter: <a href="https://twitter.com/CatalystCoop">@CatalystCoop</a></li> <li>Connect with us <a href="https://www.linkedin.com/company/catalyst-cooperative/">on LinkedIn</a></li> <li>Play with our data and notebooks <a href="https://www.kaggle.com/catalystcooperative">on Kaggle</a></li> <li>Combine our data with ML models <a href="https://huggingface.co/catalystcooperative">on HuggingFace</a></li> <li>Learn more about us on our website: <a href="https://catalyst.coop">https://catalyst.coop</a></li> <li>Subscribe to our announcements list for <a href="https://catalyst.coop/updates">email updates</a>.</li> </ul>

opencc-by-4.0Nov 2024View details →
zenodo44/100

Vectors for Goode's Homolosine projection

<p>This dataset contains useful vector maps to work with with Goode&#39;s Homolosine projection. The list of files included are:</p> <ul> <li>&nbsp;<strong>CounterDomain.geojson</strong> - a polygonal approximation of the Homolosine projection counter-domain. This can be used to fix vectors wrongly&nbsp;&nbsp;projected by programmes that consider the counter-domain to be infinite. It&nbsp;can also be used to represent the seas in global mapping.</li> <li>&nbsp;<strong>ParallelsMeridians.geojson</strong> -&nbsp;a set of meridians and parallels to be used in the creation of global maps.</li> <li><strong>Homolosine.crs</strong> - the PROJ string defining the Homolosine projection (referenced by the GeoJSON slides)</li> <li><strong>LICENCE</strong> - full text of the licence (EUPL-1.2)</li> </ul> <p>These datasets were generated with the open souce programme homolosine-vectors, available at:&nbsp;<a href="https://gitlab.com/ldesousa/homolosine-vectors">https://gitlab.com/ldesousa/homolosine-vectors</a></p>

opencc-by-sa-4.0Oct 2018View details →
zenodo44/100

Online Real-Time Delphi Survey for the research project "MENARA" - Compilation of all Comments to Closed and Open Questions

<p><strong>Looking into the Futures: Delphi Survey about the MENA region</strong></p> <p>In order to get a more realistic overview of the situation and trends, of the potentials, problems and potentials of the countries of the MENA region a Real Time Delphi survey was conducted. This is an important tool of modern future research. It was managed by the IZT- Institute for Future Studies in Berlin. A group of 139 experts and researchers from different institutes and organizations were invited to participate at the Online Real-Time Delphi Survey (RTD) about possible and likely futures of the MENA region. The experts were asked to answer questions and provide their opinions on twelve topics such as social unrest, youth unemployment, urbanization, gender equality, security etc. In this dataset all comments to the closed and the open questions are compiled.</p> <p>The output was one of the basic material used for the creation of future regional scenarios for mid-term (2025) and long-term (2050) time horizons. Focus scenarios were produced in order to exemplify selected characteristic and important future options, in terms of chances and risks (e.g. energy futures).</p>

opencc-by-4.0Mar 2019View details →
zenodo44/100

Survey for coordinators of agroecology research projects

<p>This dataset (Project_coordinators_survey.csv) contains data related to the survey launched within the Task 1.3 of AE4EU project for the coordinators of agroecology research projects funded by European, transnational, and national programmes identified in the mapping activities. Together with the dataset, the structure of the questionnaire related to this survey is also provided (Project_coordinators_questionnaire_structure.pdf).</p> <p>Answers from respondents were anonymised before the publication. Informed consent was obtained from participants to the survey to use their answers and quotations for research and publication</p>

opencc-by-4.0Oct 2022View details →
zenodo44/100

List of validated primers of gilthead sea bream (Sparus aurata) and European seabass (DIcentrarchus labrax) developed in PerformFISH project (D2.3)

<p>The document contains all the primers identified for the screening of genes tested for their potential as biomarkers&nbsp;&nbsp;to predict quality performance in gilthead sea bream and European sea bass larvae and juveniles in the context of PERFORMFISH (WP2). The spreadsheet has the following information: Pathway, phisiologic process in which the gene is involved; name of protein that &nbsp;gene produces; gene code; acession n&ordm;, code given in the consulted databases and the sequence extracted for primer design; FW and RV primer, forward and reverse primer sequence specific for target gene; melt temperature, &nbsp;optimized temperature that primers work at ; amplicon size, size in base pairs of the product produced &nbsp;with the &nbsp;primers; eff%, efficency of primers; r2; source, the origin of the primers, &quot;in house&quot; or &quot;literature&quot; (including available DOI. &nbsp;Each pair of primers are classified using a &quot;traffic light&quot; system indicating their validation status.</p>

opencc-by-4.0Oct 2022View details →
zenodo44/100

WoS and Scopus records for the bibliometric analysis in the output D2.2 Digital transformation of research and innovation roadmap of the reSEArch-EU project

<p>These files represent the exported WoS and Scopus records, used in the output&nbsp;D2.2 Digital transformation of research and innovation roadmap &nbsp;of the Horizont project reSEArch-EU, implemented by the SEA-EU university alliance.</p>

opencc-by-4.0Nov 2022View details →
zenodo44/100

Report on a Survey among Organisers of Citizen Science Projects - Dataset and Report

<p>In this publication you will find the report on a survey among organisers of citizen science projects developed by&nbsp;Michael Str&auml;hle &amp; Christine Urban (alphabetical order), Wissenschaftsladen Wien - Science Shop Vienna. You will also find, the datasets which contain&nbsp;the responses obtained from the very short questionnaire &quot;VSQ&quot; and the responses used for the report.</p>

opencc-by-4.0Nov 2022View details →
zenodo44/100

[Dataset] Does Volunteer Engagement Pay Off? An Analysis of User Participation in Online Citizen Science Projects

<p>Corresponding dataset for the publication &quot;Does Volunteer Engagement Pay Off? An Analysis of User Participation in Online Citizen Science Projects&quot;, a conference paper for the conference&nbsp;CollabTech 2022:&nbsp;<a href="https://link.springer.com/book/10.1007/978-3-031-20218-6">Collaboration Technologies and Social Computing</a>&nbsp;and&nbsp;published as part of the&nbsp;<a href="https://link.springer.com/bookseries/558">Lecture Notes in Computer Science</a>&nbsp;book series (LNCS,volume 13632) <a href="https://link.springer.com/chapter/10.1007/978-3-031-20218-6_5">here</a>. Usernames have been anonymised.</p> <p>The structure of the&nbsp;dataset is as follows:</p> <p><strong>Annotations</strong>&nbsp;</p> <p><em>List of annotations made per day for each of the analysed projects.</em></p> <p><code>annotations.csv&nbsp;</code></p> <p><strong>Comments&nbsp;</strong></p> <p><em>Total list of comments with several data fields (i.e., comment id, text, reply_user_id)</em></p> <p><code>comments.csv&nbsp;</code></p> <p><strong>Rolechanges</strong>&nbsp;</p> <p><em>List of roles per user to determine number of role changes&nbsp;</em></p> <p><code>478_rolechanges.csv</code></p> <p><code>1104_rolechanges.csv</code></p> <p><code>...</code></p> <p><strong>Totalnetworkdata</strong>&nbsp;</p> <p><em>Network data (edge and node sets) for the given projects (without time slices).</em></p> <p>Edges&nbsp;</p> <ul> <li> <p><code>478_edges.csv</code></p> </li> <li> <p><code>1104_edges.csv</code></p> </li> </ul> <p>Nodes&nbsp;</p> <ul> <li> <p><code>478_nodes.csv</code>&nbsp;</p> </li> <li> <p><code>1104_nodes.csv</code>&nbsp;</p> </li> </ul> <p><strong>Trajectories</strong>&nbsp;</p> <p><em>Network data (edge and node sets) for the given projects and all time slices (Q1&nbsp;2016 - Q4 2021)</em></p> <p>478&nbsp;</p> <ul> <li>Edges&nbsp; <ul> <li> <p><code>edges_4782016_q1.csv</code></p> </li> <li> <p><code>edges_4782016_q2.csv</code></p> </li> <li> <p><code>edges_4782016_q3.csv</code></p> </li> <li> <p><code>edges_4782016_q4.csv</code></p> </li> </ul> </li> <li> <p>...</p> </li> <li>Nodes&nbsp; <ul> <li><code>nodes_4782016_q1.csv</code></li> <li> <p><code>nodes_4782016_q4.csv</code></p> </li> <li> <p><code>nodes_4782016_q3.csv</code></p> </li> <li> <p><code>nodes_4782016_q2.csv</code></p> </li> <li> <p><code>...</code></p> </li> </ul> </li> </ul> <p>&nbsp;</p> <p>1104&nbsp;</p> <ul> <li> <p>Edges&nbsp;</p> <ul> <li> <p><code>...</code></p> </li> </ul> </li> <li> <p>Nodes&nbsp;</p> <ul> <li> <p><code>...</code></p> </li> </ul> </li> <li> <p><code>...</code></p> </li> </ul> <p>&nbsp;</p>

opencc-by-4.0Nov 2022View details →
zenodo44/100

Icons of the tipping points in the Earth System from the H2020 COMFORT project (820989)

<p>Triple threat processes and/or other forcings can lead to changes in the ocean happening fast and abruptly. These changes, referred to as &ldquo;tipping points&rdquo;, are critical thresholds in a marine system that, when exceeded, can lead to a significant change in the state of the system, which often can be irreversible. This product has been prepared with the financial support of Norges forskningsr&aring;d (Research Council of Norway) (309382) and the European Union&rsquo;s Horizon 2020 research and innovation programme under grant agreement No 820989 (project COMFORT, Our common future ocean in the Earth system &ndash; quantifying coupled cycles of carbon, oxygen, and nutrients for determining and achieving safe operating spaces with respect to tipping points). The work reflects only the author&rsquo;s/authors&rsquo; view; the European Commission and their executive agency are not responsible for any use that may be made of the information the work contains.</p>

opencc-by-4.0Nov 2022View details →
zenodo44/100

Availability of information on citizen science activities, checked against the Activities & Dimensions Grid of Citizen Science on the basis of some projects

<p>The research resulting in this report aimed at answering the following questions:</p> <ul> <li> <p>Which information on citizen science activities is online available that matches the Activity &amp; Dimension Grid of Citizen Science or goes beyond it?&nbsp;&nbsp;</p> </li> <li> <p>Is there any contradictory information?</p> </li> <li> <p>What can be the reason for the availability or non-availability of information about citizen science activities?</p> </li> <li> <p>How does/could this impact on the CS Track&rsquo;s recommendations?</p> </li> </ul> <p>The corresponding dataset consists of the results of a keyword-based search in the WP2 project database. The information retrieval resulted in 3318 projects on which information is available in German or English.</p> <p>More information on this research can be found in D2.2 section 3.2.</p>

opencc-by-4.0Nov 2022View details →
zenodo44/100

Demographic factors and the environmental Kuznets curve: global plastic pollution by 2050 could be 2 to 4 times worse than projected

<p>These data are made of two files. One file provides the observed data we collected and cleaned from the World Bank database. The second file provides the simulation results from the STIRPAT model we designed based on the&nbsp;observed data abovementioned. Our results can be summarised as follows:</p> <p>Since 2015, the detrimental effects of plastic pollution have attracted media, public, and governmental attention. Considering economic growth is inevitable and a key driver of plastic contamination, it is worthwhile to analyze the environmental Kuznets curve (EKC) relationship between economic development and plastic pollution. To this end, we contribute by being the first to (i) use the Stochastic Impacts by Regression on Population, Affluence, and technology model (STIRPAT model) to investigate this EKC relationship; (ii) provide a comprehensive analysis of how demographic factors affect plastic pollution; and (iii) use panel model techniques to examine the drivers of plastic pollution. Our empirical results support an inverted U-shaped relationship between plastic pollution and income. They show that at current trends, global plastic pollution (that is, annual discard of inadequately managed plastic waste) is expected to grow from 52 million tons per year in 2020 to 257 million tons per year in 2050.</p>

opencc-by-4.0Dec 2022View details →
zenodo44/100

Analysis of the content of the H2020 project websites related to LEAs and IA.

<p>This is the dataset used in the article entitled &quot;The disconnect between the goals of trustworthy AI for law enforcement and the EU research agenda&quot;. You can find more information about the results obtained, as well as the methodology used in the paper.</p>

opencc-by-4.0Dec 2022View details →
zenodo44/100

Codes to replicate statistical analysis in UPLIFT project Deliverable 2.4 Synthesis report, Chapter 6

<p>Policies attempting to mitigate the effects of urban inequality, often disregard affected citizens&rsquo; experiences, and thus fail to achieve&nbsp;maximum impact. By incorporating these perspectives into the policy design process, the project &quot;Urban PoLicy Innovation to address inequality with and for Future generaTions&quot;&nbsp;(UPLIFT), funded under the EU Horizon 2020 program&nbsp;aims to find innovative interventions in a bottom-up approach. The aims of UPLIFT project are to understand patterns and trends of inequality across Europe and to understand how individuals experience and adapt to inequality through participatory research. Moreover the project will together with the communities in four locations, co-design a policy tool aimed at addressing and reducing inequality and socio-economic divisions. The activity and results of the project can be followed at&nbsp;<a href="https://www.uplift-youth.eu/">https://www.uplift-youth.eu/</a>.</p> <p>Deliverable 2.4 (Synthesis report:&nbsp;socioeconomic inequalities in different urban contexts) is the final deliverable of work package 2 of the UPLIFT project, which aims to synthetize the main outcomes of the urban reports that described the policy environment around vulnerable individuals in the fields of education, employment and housing in 16 functional urban areas of the EU. In addition section 6.2 &quot;Statistical analysis of linkages between economic development of cities, their public policy performance and inequality outcomes&quot; of the report provides a statistical analysis of&nbsp;how local economic competitiveness&nbsp;and the local policy context affect urban deprivation and inequality among the young in European cities. The analysis is based on data from 2006, 2009, 2012, 2015 and 2019 of the Quality of Life in European Cities survey.</p>

opencc-by-4.0Nov 2022View details →
zenodo44/100

Data for Project 'Feasibility, Usability and Acceptance of a Newly Developed Exergame-Based Training Concept for Older Adults with Mild Neurocognitive Disorder - A Pilot Randomized Controlled Trial'

<p>Data for Project &#39;Feasibility, Usability and Acceptance of a Newly Developed Exergame-Based Training Concept for Older Adults with Mild Neurocognitive Disorder - A Pilot Randomized Controlled Trial&#39; (trial&nbsp;registered at clinicaltrials.gov (<a href="https://clinicaltrials.gov/ct2/show/NCT04996654">NCT04996654</a>; date of registration: 11 July 2021), consisting&nbsp;of:</p> <p>(1) the&nbsp;original and complete data set for all primary outcomes (&#39;Data_Primary-Outcomes_Brain-IT-Pilot-Feasibility-RCT_for-publication.xlsx&#39;);</p> <p>(2) the original and complete data set for all secondary outcomes (&#39;Data_Secondary-Outcomes_Brain-IT-Pilot-Feasibility-RCT_for-publication.xlsx&#39;);</p> <p>(3) the&nbsp;original and complete data set for all other outcomes (i.e. baseline factors (demographic data, type of usual care interventions) and training heart rate; &#39;Data_Other-Outcomes_Brain-IT-Pilot-Feasibility-RCT_for-publication.xlsx&#39;);</p> <p>(4) folder including the raw and processed heart rate variability (HRV) and electroencephalography (EEG)&nbsp;data for all participants and measurements (HRV-and-EEG_raw-and-processed-data.zip);</p> <p>(5)&nbsp;a corresponding README file including (a) general information, (b) data and file overview, (c) sharing and access information, (d) methodological information, and (e) data-specific information.</p>

opencc-by-4.0Dec 2022View details →
zenodo44/100

OpenChart-SE: A corpus of artificial Swedish electronic health records for imagined emergency care patients written by physicians in a crowd-sourcing project

<p>Electronic health records (EHRs) are a rich source of information for medical research and public health monitoring. Information systems based on EHR data could also assist in patient care and hospital management. However, much of the data in EHRs is in the form of unstructured text, which is difficult to process for analysis. Natural language processing (NLP), a form of artificial intelligence, has the potential to enable automatic extraction of information from EHRs and several NLP tools adapted to the style of clinical writing have been developed for English and other major languages. In contrast, the development of NLP tools for less widely spoken languages such as Swedish has lagged behind. A major bottleneck in the development of NLP tools is the restricted access to EHRs due to legitimate patient privacy concerns. To overcome this issue we have generated a citizen science platform for collecting artificial Swedish EHRs with the help of Swedish physicians and medical students. These artificial EHRs describe imagined but plausible emergency care patients in a style that closely resembles EHRs used in emergency departments in Sweden. In the pilot phase, we collected a first batch of 50 artificial EHRs, which has passed review by an experienced Swedish emergency care physician. We make this dataset publicly available as OpenChart-SE corpus (version 1) under an open-source license for the NLP research community. The project is now open for general participation and Swedish physicians and medical students are invited to submit EHRs on the project website (<a href="https://github.com/Aitslab/openchart-se">https://github.com/Aitslab/openchart-se</a>), where additional batches of quality-controlled EHRs will be released periodically. &nbsp;</p> <p>&nbsp;</p> <p><strong>Dataset content</strong></p> <p><em>OpenChart-SE, version 1 corpus (txt files and and dataset.csv)</em></p> <p>The OpenChart-SE corpus, version 1, contains 50 artificial EHRs (note that the numbering starts with 5 as 1-4 were test cases that were not suitable for publication). The EHRs are available in two formats, structured as a .csv file and as separate textfiles for annotation. Note that flaws in the data were not cleaned up so that it simulates what could be encountered when working with data from different EHR systems. All charts have been checked for medical validity by a resident in Emergency Medicine at a Swedish hospital before publication.</p> <p>&nbsp;</p> <p><em>Codebook.xlsx</em></p> <p>The codebook contain information about each variable used. It is in XLSForm-format, which can be re-used in several different applications for data collection.</p> <p>&nbsp;</p> <p><em>suppl_data_1_openchart-se_form.pdf</em></p> <p>OpenChart-SE mock emergency care EHR form.</p> <p>&nbsp;</p> <p><em>suppl_data_3_openchart-se_dataexploration.ipynb</em></p> <p>This jupyter notebook contains the code and results from the analysis of the OpenChart-SE corpus.</p> <p>&nbsp;</p> <p>More details about the project and information on the upcoming preprint accompanying the dataset can be found on the project website (<a href="https://github.com/Aitslab/openchart-se">https://github.com/Aitslab/openchart-se</a>).</p>

opencc-by-4.0Dec 2022View details →
zenodo44/100

Model agreement and trend analysis data associated to the publication: "Impact of climate change on site characteristics of eight major astronomical observatories using high-resolution global climate projections until 2050"

<p>This dataset is associated with the following&nbsp;publication:</p> <p>Haslebacher, C., Demory, M.-E., Demory, B.-O., Sarazin, M., and Vidale, P. L., &ldquo;Impact of climate change on site characteristics of eight major astronomical observatories using high-resolution global climate projections until 2050. Projected increase in temperature and humidity leads to poorer astronomical observing conditions&rdquo;, <em>Astronomy and Astrophysics</em>, vol. 665, 2022. doi:10.1051/0004-6361/202142493.</p> <p>In the folder &#39;model_agreement&#39;, there are pickle files from which a python dictionary can be extracted with:</p> <pre><code>with open('mypklfile.pkl', 'rb') as myfile: dload = pickle.load(myfile)</code></pre> <p>Pickle files ending with &#39;_d_obs_ERA5.pkl&#39; contain in situ data and ERA5 data. Pickle files ending with &#39;d_model.pkl&#39; contain PRIMAVERA model data. A few explanations:<br> - &#39;ds_sel&#39;: contains monthly timeseries of selected intersecting data<br> - &#39;ds_taylor&#39;: contains data used for the Taylor diagram&nbsp;(Figs. 4-10)<br> - &#39;ds_mean_month&#39;: contains seasonal cycle&nbsp;for plotting (Figs. 4-10)<br> -&nbsp;&#39;ds_mean_year&#39;: contains yearly timeseries for plotting (Figs. 4-10)&nbsp;</p> <p>The subfolder &#39;median_nc_u_v_t&#39; contains NETCDF files with the median and interquartile range of the wind speed in u and v direction, the temperature and geopotential height. This was used for Figs. G1-G8 and to calculate the refractive index structure constant Cn2.</p> <p>The subfolder &#39;skill_score_classification&#39; contains csv files with the sorted skill score classifications. The column headers are: model_name, skill score, correlation coefficient, standard deviation, centred root mean square error.</p> <p>The folder &#39;trend_analysis&#39; contains for each variable csv files of ERA5 and PRIMAVERA monthly time series used for&nbsp;trend analysis, pdf files of analysis summaries, csv files of Bayesian analysis results and png files of longitude-latitude maps of trends (analysed with linear regression). Additionally, there is a csv file of&nbsp;averaged in situ pressures.</p> <p>Code that generated and used this data&nbsp;is available on github:&nbsp;<a href="https://github.com/CarolineHaslebacher/Astroclimate-future-project">https://github.com/CarolineHaslebacher/Astroclimate-future-project</a>&nbsp;&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Jan 2023View details →
zenodo44/100

IMMERSE Horizon 2020 Project Downstream User Toolbox – data for tutorial on impact of wave coupling on surface particle dispersion simulations

<p>Exemplary data for tutorial on impact of wave coupling on surface particle dispersal simulations<br> <a href="https://github.com/immerse-project/Downstream-Users-Toolbox/tree/main/T8.3_WaveCoupling_ParticleTransport_UniU">https://github.com/immerse-project/Downstream-Users-Toolbox/tree/main/T8.3_WaveCoupling_ParticleTransport_UniU</a><br> created as part of the downstream user toolbox of the IMMERSE Horizon 2020 project (<a href="https://immerse-ocean.eu/">https://immerse-ocean.eu/</a>).</p> <p>In the tutorial the impact of new options for the representation of wave-current interactions in the NEMO ocean model (<a href="https://www.nemo-ocean.eu/">https://www.nemo-ocean.eu/</a>) on surface particle simulations are tested in a case study for the Mediterranean Sea. The tutorial consists of two jupyter notebooks: Parcels_CalcTraj.ipynb and CompTraj_uncoupledVScoupled.ipynb. Parcels_CalcTraj.ipynb calculates Lagrangian particle trajectories based on velocity output &nbsp;from ocean only as well as coupled ocean-wave model simulation by making use of the OceanParcels software (<a href="https://oceanparcels.org/">https://oceanparcels.org/</a>). CompTraj_uncoupledVScoupled.ipynb compares dispersal statistics of Lagrangian particle trajectories calculated from ocean-only vs coupled ocean-wave model simulations.</p> <p>This repository contains the surface velocity and ocean model grid data needed to run Parcels_CalcTraj.ipynb, as well as the trajectory data produced by Parcels_CalcTraj.ipynb, which is needed to run CompTraj_uncoupledVScoupled.ipynb. The surface velocity data stems from two simulations with a regional high-resolution (1/24&deg; horizontal resolution) model configuration for the Mediterranean Sea: a coupled ocean-wave model simulation and a complimentary ocean-only simulation. These model simulations make use of the NEMO v4.2-RC ocean model, the Wave Watch 3 v.6.07 wave model, the OASIS3-MCT coupler, and ECMWF atmospheric fields; they are described in detail in IMMERSE deliverable D5.7 &ldquo;Assessment of wave-current effects on the circulation in theMed-MFC system&rdquo;<strong>.</strong></p>

opencc-by-4.0Jan 2023View details →
zenodo44/100

The e-NDP project : collaborative digital edition of the Chapter registers of Notre-Dame of Paris (1326-1504). Ground-truth for handwriting text recognition (HTR) on late medieval manuscripts.

<p>The <a href="https://endp.hypotheses.org/">e-NDP project</a>, funded by the ANR, is led by the <a href="https://lamop.hypotheses.org/6870">LaMOP</a> (Julie Claustre and Darwin Smith).</p> <p>The project&#39;s partners are the Archives nationales, the&nbsp;Biblioth&egrave;que nationale de France (Department of Manuscripts, Biblioth&egrave;que de l&#39;Arsenal), the &Eacute;cole nationale des chartes and the Biblioth&egrave;que Mazarine.</p> <p>The e-NDP project aims at renewing our knowledge on <strong>Notre-Dame de Paris cathedral</strong> through the creation of a collaborative digital edition of the registers of its Chapter (1326-1504, <em>AN LL 105-128</em>), the community of 51 canons meeting three times a week on set days to take all administrative, financial and practical decisions pertaining to the cathedral, its estate and the society living in its cloister. This corpus has never been the object of a comprehensive study to understand the workings and history of this urban enclave and powerful community. The collaborative digital edition is based on a process of<strong> handwriting text recognition (HTR)</strong>, tested and supervised by scholars, researchers and engineers combining expertise in Medieval history, paleography, philology and digital humanities. The edition shall allow a better insight into the Chapter&rsquo;s administration, into its economical and political power within Paris, and the relationships it maintained with other institutions in the city.</p> <p>&nbsp;</p> <p><strong>Section 1 : The e-NDP ground-truth dataset for Handwriting text recognition.</strong></p> <p>The full e-NDP corpus kept today in the French National Archives and was entirely digitized and described in its&nbsp;<a href="https://www.siv.archives-nationales.culture.gouv.fr/siv/rechercheconsultation/consultation/ir/consultationIR.action?formCaller=GENERALISTE&amp;irId=FRAN_IR_059635">catalog</a>&nbsp;in 2022.</p> <p>The first major goal of the&nbsp;e-NDP projet is to propose a first automatic transcription of the 14k pages composing the 26 chapter registers. To achieve this goal representative samples from&nbsp;each one of the volumes were selected and transcribed in order to train a specialized HTR model able to propose a high quality automatic transcription. The collected ground-truth released on this repository currently has <strong>512 pages from the 26 registers</strong> of the cathedral chapter preserved in the National Archives (LL105 - LL128, <strong>1326-1504</strong>). The transcriptions were manually completed in <strong>two rounds</strong> by a group of 12 contributors, historians and paleographers, over the course of 2021-2022 using <a href="https://escriptorium.paris.inria.fr/">eScriptorium </a>as annotation environment.&nbsp;&nbsp;</p> <p>&nbsp;</p> <p><strong>Ground-truth features :</strong></p> <p><br> <em>Number of hands </em>: according to our estimates no fewer than 18&nbsp;main hands were involved in the writing of the registers during the medieval period.&nbsp;</p> <p><em>Language</em> : More than 98% of the content of the registers was written in Latin, the rest in French. The exact percentage is hard to estimate because the vernacular language is often used in formulae, notes and comments. It is rare to find entire pages or blocks written in French.&nbsp;</p> <p><em>Script family</em> : The registers were written using a Cursive script (ca. late XIIIe - XVIe).</p> <p><em>Documental typology</em> : The volumes containing the chapter conclusions were conceived to serve&nbsp;as memorial&nbsp;records, but above all as documents for regular use and consultation in the daily practice of administration and management. In diplomatics the notion of &quot;documentary manuscripts&quot; is used to describe this kind of sources&nbsp;also by opposition to books and litterary or&nbsp;normative&nbsp;manuscripts.</p> <table align="center"> <caption><strong>Ground truth statistics</strong></caption> <tbody> <tr> <th>Text units</th> <th>Count</th> </tr> <tr> <td>Pages</td> <td>512</td> </tr> <tr> <td>Annotated regions (see section 2)</td> <td>2448</td> </tr> <tr> <td>Lines of text</td> <td>34231</td> </tr> <tr> <td>Tokens</td> <td>205083</td> </tr> <tr> <td>Characters</td> <td>3320407</td> </tr> </tbody> </table> <p>&nbsp;</p> <p><strong>Rules of transcription :</strong></p> <ul> <li>The abbreviations have been resolved, both those by suspension (<code>facimꝰ</code> ---&gt; <code>facimus</code>) and by contraction (<code>d&ntilde;i</code> --&gt; <code>domini</code>). Likewise, those using conventional signs (<code>⁊</code> --&gt; <code>et</code> ; <code>ꝓ</code> --&gt; <code>pro</code>) have been resolved.&nbsp;</li> <li>The named entities (names of persons, places and institutions) have been <code>capitalized</code>. The beginning of a block of text as well as the original capitals used by the notary are also capitalized.</li> <li>The consonantal <code>i</code> and <code>u</code> characters have been transcribed as <code>j</code> and <code>v</code> in both French and Latin.</li> <li>The punctuation marks used in the text: <code>.</code> and <code>/</code> have been transcribed, but the transcription has not been standardized with modern punctuation.</li> <li>Corrections and words that appear cancelled in the manuscript have been transcribed surrounded by the sign <code>$</code> at the beginning and at the end.</li> <li>More specific transcription rules can be found into the file <code>transcription_guidelines.pdf</code></li> </ul> <p>&nbsp;</p> <p><strong>Section 2. e-NDP Layout Segmentation.</strong></p> <p>Layout segmentation is a compulsory step before HTR recognition in order to distinguish sections and regions inside a document. This process intend to separate interdependant page zones to produce a recognition in a section-sequence order and not in a line-sequence order which mix textual and peri-textual content.</p> <p>The regions of 364&nbsp;pages (see <code>GT-layout_list</code>) of the e-NDP corpus were annotated using a 5 sections vocabulary (see <code>endp_layout_regions</code>) in order to describe&nbsp;the page distribution in all the 26 volumes :</p> <ol> <li><em>Block</em>&nbsp;: All the central text blocks, that normally corresponds to the main content called &quot;conclusions&quot; in registers.</li> <li><em>Liste</em>&nbsp;: List of names of the canons who were present during the meeting. Normally located before the <em>conclusions</em>.</li> <li><em>Entr&eacute;e</em>&nbsp;: Marginal notes or entries to inform about the content of <em>conclusions</em>.</li> <li><em>Date</em>&nbsp;: Paragraph contending the date. Normally at the head of a <em>conclusion</em>, but separate of the main body.</li> <li><em>Num&eacute;rotation</em>&nbsp;: Page numbers in roman or arabic. Usually appear in the top corners of the pages.</li> </ol> <table align="center"> <caption><strong>Layout GT statistics</strong></caption> <tbody> <tr> <th>Region</th> <th>Count</th> </tr> <tr> <td>block</td> <td>833</td> </tr> <tr> <td>liste</td> <td>431</td> </tr> <tr> <td>date</td> <td>448</td> </tr> <tr> <td>entr&eacute;e</td> <td>205</td> </tr> <tr> <td>num&eacute;rotation</td> <td>531</td> </tr> </tbody> </table> <p>&nbsp;</p> <p><strong>Section 3. The e-NDP HTR modeling.</strong></p> <p>The e-NDP project has progressively trained several HTR models adapted to work on late medieval cursive in order to accelerate the production of ground truth. Currently the best model delivers an average&nbsp;<strong>CER (Character error ratio) of 9.7%</strong> in handwriting recognition on&nbsp;the 26 registers (see <code>endp_learning_curve</code>) and can serve as generalist model&nbsp;for other manuscripts of the same period and similar script family. These models and their training implementation details can be found in the project&#39;s github <a href="https://github.com/chartes/e-NDP_HTR">repository</a>.&nbsp;</p> <p>Additionally, the automatic HTR transcriptions of the 26 registers (14k pages, 4.5M tokens) enriched with lexical and semantical information has been the subject of a first <a href="https://nosketch-engine.lamop.fr/#dashboard?corpname=endp">online publication</a> using the NoSketch engine that allows advanced data mining based on the combination of data, metadata and NLP features.&nbsp;</p> <p>&nbsp;</p> <p><strong>Section 4. Dataset content.</strong></p> <p>This zip dataset contains :</p> <p>- <code>HTR_ground_truth</code> : Two folders containing the jpg / jpeg images and their curated transcriptions in PAGE XML format.</p> <p>- <code>images_docs</code> : 4 files illustrating the different phases of the project (list of GT for layout segmentation, layout ontologie, transcription guideline and HTR evaluation curves)</p>

opencc-by-4.0Feb 2023View details →
zenodo44/100

Datasets generated by rurAllure project - promotion of rural museums and heritage sites in the vicinity of European pilgrimage routes

<p>These datasets have been generated as part of rurAllure project (funded by the European Union&rsquo;s Horizon 2020 Research and Innovation programme under grant agreement no 101004887). Main goal of rurAllure is the promotion of rural museums and heritage sites in the vicinity of European pilgrimage routes: https://rurallure.eu/project/about/</p>

opencc-by-4.0Feb 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record