Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

6,766

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

6,766 results for “project”

Learn how ShareScore rates datasets ↗
zenodo44/100

Excel data collection template on descriptive political representation in national parliaments of the projects Pathways to Power and InclusiveParl adapted for the ActEU project

<p>This file contains the empty data collection template and variable and value labels to code biographical data on legislators for WP4 in the ActEU project. It is an abbreviated version of the codebooks produced by the Pathways to Power project and by the InclusiveParl project.</p>

opencc-by-nc-4.0Sep 2024View details →
zenodo44/100

Zonal Statistics of Weather Indicators for Brazilian Municipalities from the BR-DWGD Project

<p>This dataset presents daily weather indicators for Brazilian municipalities computed with zonal statistics using the data from the&nbsp;<a href="https://sites.google.com/site/alexandrecandidoxavierufes/brazilian-daily-weather-gridded-data" target="_blank" rel="noopener">BR-DWGD project</a> (version 3.2.3), from 1961-01-01 to 2024-03-20.</p> <p>&nbsp;</p> <table> <tbody> <tr> <td>File</td> <td>Indicator</td> <td>Unit</td> </tr> <tr> <td>pr_3.2.3.parquet</td> <td>Precipitation</td> <td>mm</td> </tr> <tr> <td>ETo_3.2.3.parquet</td> <td>Evapotranspiration</td> <td>mm</td> </tr> <tr> <td>Tmax_3.2.3.parquet</td> <td>Maximum temperature</td> <td>&deg;C</td> </tr> <tr> <td>Tmin_3.2.3.parquet</td> <td>Minimum temperature</td> <td>&deg;C</td> </tr> <tr> <td>Rs_3.2.3.parquet</td> <td>Solar radiation</td> <td>MJm-2</td> </tr> <tr> <td>u2_3.2.3.parquet</td> <td>Wind speed at 2 m height</td> <td>m/s</td> </tr> <tr> <td>RH_3.2.3.parquet</td> <td>Relative humidity</td> <td>%</td> </tr> </tbody> </table> <p>The methodology to compute the zonal statistics follows <a href="https://doi.org/10.1017/eds.2024.3" target="_blank" rel="noopener">https://doi.org/10.1017/eds.2024.3</a> .</p>

opencc-by-4.0Apr 2023View details →
zenodo44/100

Data set for "Cell class-specific long-range axonal projections of neurons in mouse whisker-related somatosensory cortices"

<p>Data set for: Liu Y, Bech P, Tamura K, D&eacute;lez LT, Crochet S, Petersen CCH (2024) Cell class-specific long-range axonal projections of neurons in mouse whisker-related somatosensory cortices. eLife 13: RP97602. https://doi.org/10.7554/eLife.97602</p> <p>There are 3 files in this upload:</p> <p>1. The file named "2024_Liu_eLife.pdf" is the Open Access pdf of the online publication in eLife.</p> <p>2. The file named "Liu_anatomy_data_code.zip" (~35 GB) is a zipped version of a folder "Liu_anatomy_data_code" (~111 GB), which contains the anatomical data analysed in the study along with the Python codes used to generate the published figures 1-7 and their associated figure supplements.&nbsp;</p> <p>3. The file named "Liu_function_data_code.zip" (~10 GB) is a zipped version of a folder "Liu_function_data_code" (~35 GB), which contains the functional data analysed in the study along with the Python codes used to generate the published figure 8 and its associated figure supplement.&nbsp;</p> <p>After unzipping, the Python codes should run as a Jupyter notebook (anatomy .ipynb code) or Python code (function .py code) in Anaconda.</p>

opencc-by-4.0Oct 2024View details →
zenodo44/100

Deliverable 1.1.1.1 BEL-Float project | Dataset containing the results of numerical simulations (motions, forces) of the operational performance analysis - Input files

<p>This dataset contains the parent input used to generate the simulation files of the DeepCwind OC4 semi-submersible combined with the 5MW NREL turbine for various wind and wave conditions. The basis of the OpenFAST input files are taken from&nbsp;<a href="https://github.com/OpenFAST/r-test/tree/main/glue-codes/openfast/5MW_OC4Semi_WSt_WavesWN">OpenFAST r-test GitHub repository (5MW_OC4Semi_WSt_WavesWN)</a>&nbsp;and adapted to simulate various wind and wave conditions. The turbulent wind field as the input to the InflowWind module is generated using&nbsp;<a href="https://www.nrel.gov/wind/nwtc/turbsim.html">TurbSim</a>. The simulations are performed on a modified version of OpenFAST v3.5.3 to which adaptation to the code is made to extract additional Morison drag output up to 16 cylindrical members. This adapted code is&nbsp;<a href="https://github.com/abkpribadi/openfast/tree/Morison_additional_output">uploaded on GitHub as a branch from a forked OpenFAST repository</a>. In total there are 1152 simulation results consists of 768 irregular waves and 384 regular waves cases. The complete dataset is divided into 9 sub-datasets, see "Related work" section. A report describing this dataset will be made available on BEL-Float project website by November 2024: https://www.owi-lab.be/bel-float.</p>

opencc-by-4.0Oct 2024View details →
zenodo44/100

C2SMARTER Year 1 Project "Enhancing Transit Access and Safety Through Equitable Micromobility Solution"

<p>These 6 PDF files are the maps produced from Task 1 of the C2SMARTER Year 1 project Enhancing Transit Access and Safety Through Equitable Micromobility Solution.</p> <p>Site A and Site B are transit underserved areas (census tracts in El Paso, Texas) identified in Task 1 of this project.</p> <p>The first 2 maps shows the underserved areas overlaid with bus stops (taken from the General Transit Feed Specification or GTFS database).</p> <p>The next 2 maps shows the underserved areas overlaid with locations of crashes involving pedestrians and bicycles from 1/1/2024 to 7/30/2024..</p> <p>The last 2 maps color coded the streets in Site A and Site B with bicycle level of traffic stress (LTS).</p>

opencc-by-4.0Aug 2024View details →
zenodo44/100

Song capturing lived-experiences of flooding and climate resilience with St. Eugenes Choir Newtownstewart (BluePrint project)

<p>This audio piece represents one of the creative risk communication outputs co-created within the BluePrint project. Between March and October 2024, socially engaged artist Sara Walmsley worked creatively with flood-affected community representatives in Newtownstewart, Co. Tyrone and Eglinton, Co. Derry-Londonderry exploring their lived-experiences of flooding and need for climate adaptation and resilience.&nbsp;</p> <p>In the audio piece, you will hear the melodic, polyphonic harmonies of St. Eugene&rsquo;s Church choir (Newtownstewart) as they give music to the words of members of their community whose homes were destroyed and lives endangered by flood water. The piece captures the voices of those striving to adapt to our changing climate, those who are responding to the urgency by finding solace, hope, strength and courage in the unending and unsurprising resilience and creativity of our communities.&nbsp;</p> <p>The BluePrint project is led by the MaREI Centre, University College Cork, with partners the Playhouse, Derry City and Strabane District Council, and Mayo County Council. The BluePrint project is a recipient of the&nbsp;Creative Climate Action fund, an initiative from the Creative Ireland Programme. It is funded by the Department of Tourism, Culture, Arts, Gaeltacht, Sport and Media in collaboration with the Department of the Environment, Climate and Communications.&nbsp;</p> <p>Find out more: <a href="https://www.marei.ie/project/blueprint/">https://www.marei.ie/project/blueprint/</a></p>

opencc-by-sa-4.0Nov 2024View details →
zenodo44/100

Supplemental_Data_S1 for "Kmer Manifold Approximation and Projection for visualizing DNA sequences"

<p>This dataset includes the results generated by KMAP software applied to the htselexdata dataset. Each folder within the dataset contains outputs from multiple dimensionality reduction techniques, including KMAP, UMAP, t-SNE, and MDS. Additionally, motifs and logos have been derived using both KMAP and MEME methods. This data provides insights into motif patterns and structures, which can be beneficial for further bioinformatics and computational biology analyses.</p>

opencc-by-4.0Nov 2024View details →
zenodo44/100

Twiter Dataset on climate change discussions: COP27, IPCC, climate refugees and Doñana - Clint project

<p><strong>CLINT Data</strong></p> <p>This repository contains the date used in the project CLINT and the paper &nbsp;"<a href="https://arxiv.org/abs/2410.21187">A cross-platform analysis of polarization and echo chambers in climate change discussions</a>"&nbsp;&nbsp;</p> <p><strong>Open Twitter Data</strong></p> <p>We used the Twitter&rsquo;s search to gather historical tweets and the streaming API to follow specified accounts and also collect in real-time tweets that mention specific keywords. To comply with <a href="https://developer.twitter.com/en/developer-terms/agreement-and-policy">Twitter&rsquo;s Terms of Service</a>, we are only publicly releasing the tweet IDs of the collected tweets. The data is released for non-commercial research use.&nbsp;</p> <p><strong>With Twitter's changes to its Academic API policies, it&rsquo;s no longer possible to collect or rehydrate tweets </strong><strong>as we usually did, however we open data in case at some point it will become feasible to do it.</strong></p> <table> <tbody> <tr> <td>&nbsp;</td> <td><strong>IPCC</strong></td> <td><strong>Do&ntilde;ana</strong></td> <td><strong>Climate Refugees</strong></td> <td><strong>COP27</strong></td> </tr> <tr> <td><strong>Number of tweets</strong></td> <td>352,723&nbsp;</td> <td>1,487,425</td> <td>1,938,932</td> <td>6,225,508&nbsp;</td> </tr> <tr> <td><strong>Number of authors</strong></td> <td>157,056</td> <td>290,782</td> <td>841,454&nbsp;</td> <td>1,351,903&nbsp;</td> </tr> <tr> <td><strong>First tweet date</strong></td> <td>2023-03-18</td> <td>2019-01-01</td> <td>2008-03-10&nbsp;</td> <td>2022-09-01&nbsp;</td> </tr> <tr> <td><strong>Last tweet date</strong></td> <td>2023-03-26</td> <td>2023-04-30</td> <td>2022-12-31&nbsp;</td> <td>2022-11-27</td> </tr> </tbody> </table> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp;</p>

opencc-by-4.0Nov 2024View details →
zenodo44/100

Van Dijk et al. (2021), A meta-analysis of projected global food demand and population at risk of hunger for the period 2010–2050, data and scripts

<p>This repository contains all data and R scripts to reproduce the figures in Van Dijk et al. (2021), A meta-analysis of global food demand and population at risk of hunger projections for the period 2010-2050, Nature Food. More specifically, it includes two databases: (1) A database with standardized information to describe the characteristics of 57&nbsp;studies that were identified by the systematic literature review and (2) The&nbsp;Global Food Security Projections Database v1.0.1&nbsp;with harmonized projections for three&nbsp;global food security indicators: food consumption in kcal per capita and total kcal, and population at risk of hunger. The database also includes projections for total global population that are required to derive the global food security indicators.</p> <p>The two scripts (nf_figures.r and nf_meta_regression.r) can be used to reproduce the figures and tables in the main paper and the supplementary information. Please start with the first script, which sources the second script.&nbsp;</p> <p>This is the first version of the Global Food Projections Database. We expect to update the data, including additional studies and variables in the future. For issues and suggestions, please contact michiel.vandijk@wur.nl.</p> <p>&nbsp;</p>

opencc-by-4.0Dec 2020View details →
zenodo44/100

Ensemble projections (+ uncertainties) of contemporary (2012-2031) and future (2081-2100) mean annual plankton/phytoplankton/zooplankton species diversity (and species turn-over in time) for the global surface open ocean.

<p><em><strong>Gridded spatial fields (raster objects) containing the species distribution models (SDMs) projections of mean annual plankton total plankton, phytoplankton and zooplankton species diversity from Benedetti et al. (2021). </strong></em></p> <p>The present .grd file (&#39;rasterStack&#39; object in R) contain the fields of mean annual surface plankton/phytoplankton/zooplankton species diversity for the contemporary (2012-2031) and future (2081-2100) conditions of the global open ocean (i.e., data underlying those maps in Figure 1 and Figure 3 of Benedetti et al., 2021). Layers quantifying the uncertainty (i.e., the variablity across models projections estimated through the standard deviation) in ensemble projections were also added (i.e., data underlying the maps in Supplementary Figure 4). See the Methods section of Benedetti et al. (2021) for a full description of the methodology and the ensemble SDMs forecasting framework. The raster layers follow the 1&deg;x1&deg; cell grid of the World Ocean Atlas (https://www.ncei.noaa.gov/).</p> <p>In short, we empirically modelled the monthly and mean annual diversity patterns stemming from the distribution of 860 plankton species (336 phytoplankton, 524 zooplankton) spanning 13 phyla, 71 orders and 324 genera through an ensemble approach based on SDMs. The considered species cover a wide range of traits and functions, representing 10 major plankton functional groups (PFGs; three phytoplankton and seven zooplankton groups). We compiled the species occurrence records from various data sources (available here: https://zenodo.org/record/5101349#.YO7Dqm469lM) and aggregated them onto a monthly-resolved 1&deg;x1&deg; grid, excluding observations from regions where the seafloor is shallower than 200 m. We matched these binned open ocean records with observation-based climatologies of environmental predictors (temperature, dissolved oxygen concentration, solar irradiance, macronutrients concentration, chlorophyll a concentration) that reflect the climatic and biogeochemical conditions of the surface open ocean. Four types of SDMs (generalized linear models, generalized additive models, artificial neural networks, and random forests) were fitted to model the species&rsquo; current environmental habitat suitability patterns. For each SDMs, we used four alternative pools of predictors. Assuming niche conservatism, we projected each of the 16 resulting species-level habitat suitability models into the future using outputs from five ESMs belonging to the Coupled Model Intercomparison Project 5 (CMIP5) that were forced by the Representative Concentration Pathway 8.5 (RCP8.5) scenario of high greenhouse gas concentrations. To this end, we first computed the modelled monthly climatologies of the selected predictors for the 2012-2031 and 2081-2100 periods, and derive the future monthly anomalies from the differences between these two time periods. These anomalies were added to the observation-based monthly climatologies (i.e., those used to train the SDMs) to estimate the future environmental conditions of the ocean, and projected the SDMs in these future conditions. Finally, we estimated the mean annual present and future alpha diversity (species richness; SR) and beta diversity (species turnover through time) patterns for both trophic levels, for each cell, from the ensemble of SDMs. SR ensembles are estimated as the sum of all species&rsquo; habitat suitability patterns averaged across all 80 possible combinations (i.e., &quot;ensemble members&quot;) of SDMs (n = 4), ESMs (n = 5) and predictor pools (n = 4). To assess the uncertainties of our diversity projections based on the ensemble members, we compute the interquartile range of the 80 ensemble members SR projections. We calculate species turnover as the change in mean annual species composition between present and future time based on Jaccard&rsquo;s dissimilarity index and by decomposing this total turnover into the true species turnover (ST, also known as species replacement) and the nestedness (SR change) components. Numerous tests are conducted to ensure the robustness of the results with regard to the spatially and temporally highly uneven sampling effort as well as with regard to the relative role of different predictors.</p> <p><strong>This project has received funding from the European Union&rsquo;s Horizon 2020 research and innovation programme under grant agreement No 862923. This output reflects only the author&rsquo;s view, and the European Union cannot be held responsible for any use that may be made of the information contained therein.</strong></p>

opencc-by-4.0Jul 2021View details →
zenodo44/100

RE-Lab-Projects/TRY_DE_2015_2045: Test Reference Years (TRY) for 15 typical regions in germany with special regards on realisitc radiation data on a 1min timescale

<p>Test Reference Years (TRY) for 15 typical regions in germany with special regards on realisitc radiation data on a 1min timescale</p> <p><strong>Summary:</strong></p> <p>The data set contains the updated test reference years (TRY) of the German Weather Service (DWD). By subdividing into 15 TRY regions, each postcode area can be assigned a representative weather data set. It should be emphasized that in addition to a mean, current test reference year for a region, there is also a year with extreme summer and extreme winter weather. To take climate change into account, there is then a time series for the year 2045 for each test reference year based on the IPCC climate models. This means that a total of 90 weather data sets are available with a one-hour time resolution.</p> <p>In order to use the data in simulations with a temporal resolution of 1min or 15min, the data set was extended by linear interpolation. While this approach is justifiable for air pressure and temperature, for example, it does not depict high fluctuations in solar radiation. Therefore, based on the one-minute open data measurement data set of the Baseline Surface Radiation Network, with an algorithm by Hofmann et. al. the time series of global radiation are newly generated for all test reference years. Another algorithm by Hofmann et. al. was used to calculate the corresponding diffuse radiation times series.</p> <p><strong>Sources:</strong></p> <ul> <li>Raw data from DWD: <a href="https://kunden.dwd.de/obt/">https://kunden.dwd.de/obt/</a> -&gt; <code>1_raw-data</code></li> <li>Synthetic 1min radiation data: <a href="http://pvmodelling.org/">http://pvmodelling.org/</a> -&gt; <code>2_synthetic-radiation</code></li> </ul> <p><strong>How to use or recreate the final dataset:</strong></p> <ol> <li>clone/download this repository</li> <li>unzip the files from the data.zip file <ol> <li><a href="https://github.com/RE-Lab-Projects/TRY_DE_2015_2045/releases/download/v1.4.0/data.zip">https://github.com/RE-Lab-Projects/TRY_DE_2015_2045/releases/download/v1.4.0/data.zip</a></li> </ol> </li> <li>Use or recreate the final dataset <ol> <li>use: Final datasets are then located in -&gt; <code>3_processed-data</code></li> <li>recreate: run the <code>process-data.py</code></li> </ol> </li> </ol> <p><strong>Test reference stations / regions</strong></p> <p>No. | lon | lat | station | region<br> 1 | 53.5591 | 8.5872 | Bremerhaven | Nordseek&uuml;ste<br> 2 | 54.0878 | 12.1088 | Rostock | Ostseek&uuml;ste<br> 3 | 53.5299 | 10.0078 | Hamburg | Nordwestdeutsches Tiefland<br> 4 | 52.3938 | 13.0651 | Potsdam | Nordostdeutsches Tiefland<br> 5 | 51.4562 | 7.0568 | Essen | Niederrheinisch-westf&auml;lische Bucht und Emsland<br> 6 | 550.6461 | 7.9426 | Bad Marienburg | N&ouml;rdliche und westliche Mittelgebirge, Randgebiete<br> 7 | 51.3334 | 9.4725 | Kassel | N&ouml;rdliche und westliche Mittelgebirge, zentrale Bereiche<br> 8 | 51.7239 | 10.6069 | Braunlage | Oberharz und Schwarzwald (mittlere Lagen)<br> 9 | 50.8233 | 12.9181 | Chemnitz | Th&uuml;ringer Becken und S&auml;chsisches H&uuml;gelland<br> 10 | 50.3226 | 11.9124 | Hof | S&uuml;d&ouml;stliche Mittelgebirge bis 1000 m<br> 11 | 50.4312 | 12.9522 | Fichtelberg | Erzgebirge, B&ouml;hmer- und Schwarzwald oberhalb 1000 m<br> 12 | 49.4902 | 8.4637 | Mannheim | Oberrheingraben und unteres Neckartal<br> 13 | 48.2432 | 12.5286 | M&uuml;hldorf | Schw&auml;bisch-fr&auml;nkisches Stufenland und Alpenvorland<br> 14 | 48.6536 | 9.8666 | St&ouml;tten | Schw&auml;bische Alb und Baar<br> 15 | 47.4945 | 11.1046 | Garmisch Partenkirchen | Alpenrand und -t&auml;ler</p> <p><strong>Content</strong></p> <ul> <li><strong>files</strong>: 90 test reference years (TRY) <pre><code>15 test reference regions x 3 reference conditions (average year, extreme summer, extreme winter) x 2 reference projections (year 2015 and year 2045) </code></pre> </li> <li><strong>columns per file</strong>: <pre><code>datetime [yyyy-MM-dd hh:mm:ss+01:00/02:00] temperature [degC] pressure [hPa] wind direction [deg] wind speed [m/s] cloud coverage [1/8] humidity [%] direct irradiance [W/m^2] diffuse irradiance [W/m^2] synthetic global irradiance [W/m^2] synthetic diffuse irradiance [W/m^2] clear sky irradiance [W/m^2] </code></pre> </li> <li><strong>length</strong>: 1 year</li> <li><strong>time increment</strong>: 60s / 900s / 3600s</li> </ul> <p><strong>Important hints</strong>:</p> <ul> <li>all files in <code>3_processed-data</code> were calculated with the skript <code>process-data.py</code></li> <li><em>A value with, for example, a timestamp 12:00:00 represents the mean value from this timestamp until the following timestamp.</em></li> <li><em>datetime column is in CET / CEST</em></li> </ul>

opencc-by-4.0Jul 2021View details →
zenodo44/100

Model projection of the effect of climate change and fishing pressure on key species of the South East Asia Seas

<p>The dataset contain Projection from the Size-Spectra Bioclimatic Envelop Model (SS-DBEM), this work was part of the GCRF Blue communities Programme (www.blue-communities.org). The model provides distribution and abundance and/or biomass of fish and other species of commercial interest under climate change and fishing pressure. The model outputs are yearly abundance/biomass on a 0.5-by-0.5 degree grid, covering the period from 2000 to 2098. Further description of the model and relevant references are listed in the following file: Guide-fish-model-output-use.docx</p> <p>The model was run under two climate scenario: RCP4.5 and RCP8.5, with different combinations of fishing pressure expressed as the Maximum Sustainable Yield (MSY) for the following values: 0 (no fishing, climate change alone will cause variation in fish biomass), 1 (sustainable fishing), 2, 3 (overfishing), and, 4 (overfishing with destructive practice). The intent is not to reproduce current fishing level but to provide a range of scenarios with which the future of fisheries can be explored.</p> <p>We projected fish species that were identified as key in the South East Asia seas region by our regional partners.The full list is provided in document: Fish-list-modelguide.xlsx</p> <p>There are 4 zip files that contain the model outputs of in either abundance (number of fish) or biomass grams of fish) for the two climate scenario. For example Biomass-RCP45.zip will contain model outputs in biomass for projections under RCP4.5 and all MSY. within the zip files are .csv files of the outputs for each species under the 5 MSY (0 to 4), the individual file names identify the species (identified by a 6digit code), the output provided (abundance or biomass), the RCP (8.5 or 4.5), and the MSY (0, 1, 2, 3, or 4). For example the file labelled 600107-Abundance-rcp85-msy4.csv contains the outputs for species 600107 (Skipjack tuna, <em>Katsuwonnus pelamis</em>), as abundance, under RCP8.5 with MSY4. Headers indicate what is in each column (latitude, longitude and year).</p> <p>&nbsp;</p> <p>Note: some knowledge of Python, R, or a similar software is recommended to ensure easy of use.</p>

opencc-by-4.0Jul 2021View details →
zenodo44/100

Introducing the DHARPA Project: An interdisciplinary lab to enable critical DH practice

<p>Traditional humanities research has been leveraged but also destabilised by the increasing accessibility of digitized sources and computational tools for analysis. As traditional close reading methods alone are no longer sufficient to analyse such an unprecedented mass of digital data, a plethora of platforms have appeared in order to help researchers query and visualise the networks and patterns latent in these sources. However, this flood of data and push-button technologies has also threatened to obscure through abundance and instill in scholars a false sense of mastery. Some scholars have criticized DH for its naive and starry-eyed application of computational techniques, citing not only how the uncritical adoption of black-box technologies affects substantive research but also how it reproduces a positivism at odds with the purpose of humanities inquiry (e.g., Liu 2012; Chun 2013; Jagoda 2013; Raley 2013; Allington et al 2016; Brennan 2017; Grimshaw 2018). Both uncritical DH practice and the software it employs can be compared to a &ldquo;Mechanical Turk,&rdquo; with the decisions and interventions made by the researcher hidden from view and only the well-oiled and seemingly autonomous product on display. These trends have constituted a crisis for humanities scholarship but also an extraordinary opportunity to transform the field.</p> <p>In this presentation, we introduce the Digital History Advanced Research Projects Accelerator (DHARPA), a diverse and interdisciplinary research and development laboratory based in the University of Luxembourg&rsquo;s Centre for Contemporary and Digital History (C&sup2;DH). While software development is central to our project, our aim is not merely to build more tools but to encourage methodologies that self-reflexively examine the interaction of technology and historical practice. We want to show how the application of expertise works in tandem with technology to produce knowledge, how digitally enabled research is not a product but rather a process, reliant on the critical engagement of the scholar. We want scholars to open the black box and to be empowered to tinker with what&rsquo;s inside.</p> <p>DHARPA is multifaceted:</p> <p><strong>Software and infrastructure development:</strong></p> <p>Our software will be free and open-source and backed by a long-term sustainability plan and training opportunities to encourage widespread and confident adoption. Users will be able to run the software remotely or to download the software for local use, and to rely on a generalizable hosting infrastructure that ensures privacy, portability, and sustainability. Built for both humanities scholars and social scientists, we plan to include modules for: data ingestion, data standardisation, textual analysis, network analysis, geographical analysis and bring them within a seamless environment, where work can flow between tasks from experimentation and modelling to presentation and dissemination.</p> <p><strong>Developing best practice through design:</strong></p> <p>Central to DHARPA is the creation of a Virtual Research Environment that will enable scholars to engage with their data while promoting critical historical practice. Our interactive software will cultivate the holistic practice of interweaving data, code, computational functionality, and metadata with a narrative of researcher&rsquo;s choices and actions. Documenting academic labour makes its value evident, while also making it reproducible and keeping it honest: it allows scholars to take ownership of their interventions.</p> <p><strong>Collaborations:</strong></p> <p>DHARPA staffs a responsive laboratory that will rapidly prototype and launch applied solutions. We work with other academics at the C&sup2;DH and within a wider, international and interdisciplinary scholarly community.</p> <p><strong>Training and outreach:</strong></p> <p>As the project progresses, we will be introducing our software to the academic community through workshops at the University of Luxembourg and international conferences as well as through publications. Further to our ethos that scholarship is an iterative process, we welcome feedback to improve our software and specifications.</p>

opencc-by-4.0Jul 2021View details →
zenodo44/100

ghtorrent-projects Dataset

<p>A hypergraph dataset mined from the <a href="http://ghtorrent.org">GHTorrent</a> project is presented. The dataset contains two files</p> <p>1. project_members.txt: Contains GitHub projects with at least 2 contributors and the corresponding contributors (as a hyperedge). The format of the data is: &lt;repo_id&gt; &lt;simplex_size&gt; &lt;simplex&gt; &lt;name&gt; &lt;created_at&gt;</p> <p>2. num_followers.txt: Contains all GitHub users and their number of followers.</p> <p>The artifact also contains the SQL queries used to obtain the data from GHTorrent (<a href="https://ghtorrent.org/files/schema.pdf">schema</a>).</p>

opencc-by-4.0Jul 2021View details →
zenodo44/100

Visualization and perception of data gaps in the context of Citizen Science projects: Gradation of Reporting Activity

<p>Online experiment about the influence of different numbers of levels of representation of reporting activity&nbsp; (total number of reports for all birds in the given time span and region) on proportion of correct responses and subjective evaluation of the task (NASA-TLX). Effects of representation with three (3) levels and effects of representation with five (5) levels are investigated. Two groups of members of ornitho.de were tested: experts - persons with access to database (more than 10 reports per month in average) and novices - persons without access to database (less than 10 reports per month in average). Two different tasks were given. The evaluation of statements on a map and the selection of grid fields that met a given requirement.</p>

opencc-by-4.0Aug 2021View details →
zenodo44/100

Visualization and perception of data gaps in the context of Citizen Science projects: Video tutorial support

<p>Online experiment about the influence of the availability of a video tutorial on proportion of correct responses and subjective evaluation of the task (NASA-TLX). Two different tasks were given. The evaluation of statements on a map and the selection of grid fields that met a given requirement.</p>

opencc-by-4.0Aug 2021View details →
zenodo44/100

A unified genealogy of modern and ancient genomes: Unified, inferred tree sequences of 1000 Genomes, Human Genome Diversity, and Simons Genome Diversity Projects

<p>Unified, inferred tree sequences built from&nbsp;the 1000 Genomes phase 3, Human Genome Diversity, and Simons Genome Diversity Projects. Each tree sequence is the arm of an autosome (the short arm of acrocentric chromosomes are not included).&nbsp;Tree sequences were inferred using&nbsp;<a href="https://tsinfer.readthedocs.io/">tsinfer</a>&nbsp;version 0.2.1,&nbsp;dated using&nbsp;<a href="https://tsdate.readthedocs.io/en/latest/">tsdate</a> version 0.1.4&nbsp;and compressed using&nbsp;<a href="https://tszip.readthedocs.io/en/stable/">tszip</a>. All data is in GRCh38.</p> <p>The full data pipeline used to generate these tree sequences and associated metadata is available on&nbsp;<a href="https://github.com/awohns/unified_genealogy_paper">GitHub</a>. A description can be found in the Supplementary Material of <a href="https://www.biorxiv.org/content/10.1101/2021.02.16.431497v2">Wohns et al. (2021)</a>.</p> <p>Tree sequences can&nbsp; be decompressed as follows:</p> <pre><code>$ tsunzip hgdp_tgp_sgdp_chr1_p.dated.trees.tsz</code></pre> <p>Once decompressed, trees files can be loaded and processed in Python using&nbsp;<a href="https://tskit.readthedocs.io/">tskit</a>.&nbsp;</p> <pre><code>import tskit ts = tskit.load("hgdp_tgp_sgdp_chr1_p.dated.trees") # ts is an instance of tskit.TreeSequence print("The short arm of chromosome 1 contains {} trees".format(ts.num_trees))</code></pre> <p>Metadata associated with nodes contain&nbsp;the mean and variance of tsdate&#39;s posterior distribution on node time. To access these values, we can use:</p> <pre><code>import json node = ts.node(10000) metadata_dict = json.loads(node.metadata) print("The mean of the posterior distribution on the age of node 10000 is {} generations".format(metadata_dict["mn"])) print("The variance of the posterior distribution on the age of node 10000 is {} generations".format(metadata_dict["vr"]))</code></pre> <p>Age estimates for&nbsp;each variant site can be derived from the mean of the age estimates of the&nbsp;upper and lower bounding nodes of the oldest mutation associated with a site. tsdate includes <a href="https://tsdate.readthedocs.io/en/latest/python-api.html?highlight=sites_time_from_ts#tsdate.sites_time_from_ts">a function to find the age estimates of all sites in the tree sequence</a>:</p> <pre><code>import tsdate site_times = tsdate.sites_time_from_ts(ts, node_selection='arithmetic')</code></pre> <p>This returns a numpy array which has a length equal to the number of sites.</p> <p>Accessing variant sites in the tree sequence provides&nbsp;the position and id of variants:</p> <pre><code>site = ts.site(1000) site_metadata = json.loads(site.metadata) print("The position of site 1000 is {} and its ID is {}.".format(site.position, site_metadata["ID"]))</code></pre> <p>Metadata associated with individuals and populations was derived from the original sources (<a href="http://ftp.1000genomes.ebi.ac.uk/vol1/ftp/technical/working/20130606_sample_info/20130606_g1k.ped">TGP</a>, <a>HGDP</a>, and <a href="https://sharehost.hms.harvard.edu/genetics/reich_lab/sgdp/SGDP_metadata.279public.21signedLetter.samples.txt">SGDP</a>)&nbsp;and converted to JSON form. For example, to access individual metadata we can use:</p> <pre><code>ind = ts.individual(0) metadata_dict = json.loads(ind.metadata)</code></pre> <p>The metadata_dict variable will now contain&nbsp;all the metadata for the individual with ID 0 as a dictionary. Metadata associated with populations can be found in a similar way. Population IDs are associated with individuals via their constituent nodes. For example,</p> <pre><code>pop_metadata = [json.loads(pop.metadata) for pop in ts.populations()] ind_node = ts.node(ind.nodes[0]) ind_pop_metadata = pop_metadata[ind_node.population]</code></pre> <p>After this, the&nbsp;ind_pop_metadata variable will contain the population level metadata for individual ID 0.</p>

opencc-by-4.0Sep 2021View details →
zenodo44/100

Water analysis dataset > Water Sentinels pilot project - ACTION project

<p><a href="http://www.ocean-alive.org/en/water-sentinels">WATER SENTINELS</a> is a citizen-science pilot by Ocean Alive to promote water quality at Sado estuary.</p> <p>This dataset presents&nbsp;data and metadata for water quality assessment on&nbsp;water samples collected at Sado estuary by citizens during the 6-month pilot. The samples were analyzed by students and researchers at Set&uacute;bal School of Technology (Polytechnic Institute of Set&uacute;bal).&nbsp;</p> <p>The data collected by citizens at the field was location (GPS coordinates), temperature and water transparency (using the secchi disc method). In the laboratory researchers evaluated pH, salinity, conductivity, and concentration of ammonia, nitrate, phosphate, suspended solids&nbsp;(SST) and organic matter&nbsp;(SOM and DOM).</p> <p>WATER SENTINELS is one of the pilots supported by <a href="https://actionproject.eu/water-sentinels/"><strong>ACTION</strong></a> project. This project has received funding from the European Union&rsquo;s Horizon 2020 research and innovation programme under grant agreement number 824603.</p>

opencc-by-4.0Sep 2021View details →
zenodo44/100

Backyard Beetles and Pollinators Dataset - EREN/NEON Flexible Learning Project

<p>This dataset comes from the EREN-NEON flexible learning project &#39;Backyard Beetles and Pollinators.&#39; It can be used for teaching field, computational, or hybrid courses.&nbsp;The dataset is standardized, visual observations of insect plant visitors, identified to standard functional groups, to indirectly assess pollination and construct plant-pollinator interaction networks. Insects were identified to morphospecies in the field using reference images. We also collected information about the flowers the insects were observed on - including functional type information about the color, size, and type of flower - as well as cover.&nbsp; This information is part of an ongoing course-based undergraduate research project to both teach about plants, insects, and functional biodiversity in a flexible and inclusive way - while collaboratively assessing interaction networks across landscapes and time.&nbsp;</p> <p>To use the flexible lesson materials or join the collaboration, get more information here:&nbsp;https://erenweb.org/eren-neon-flexible-learning-projects/&nbsp; &nbsp;Or contact the project lead,&nbsp;Dr. Stack Whitney, directly at kxwsbi [at] RIT [dot] edu.&nbsp;</p>

opencc-by-4.0Aug 2021View details →
zenodo44/100

ESG Hound Published Findings on Starship Project Environmental Assessment

<p>Starship Project Environmental Assessment Critique and Findings.&nbsp;The Boca Chica launch site&#39;s environmental impact statement draft is full of errors and missing important details - ESG Hound is on the case.</p> <p>&nbsp;</p> <p>Latest version of dataset at <a href="https://www.esghound.com">https://www.esghound.com</a>.</p>

opencc-by-4.0Oct 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record