Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
747
datasets available to search
ShareScore release 0.7.1
Dataset results
747 results for “Open Data”
Data for publication "Benefits of open access to researchers from lower-income countries: A global analysis of reference patterns in 1980–2020"
<p>Data to reproduce figures for the publication "Benefits of open access to researchers from lower-income countries: A global analysis of reference patterns in 1980–2020" (DOI: 10.1177/01655515241245952). Each file contains the data underlying the figure corresponding to the file name.</p>
BeauAMP : processing and consolidation of open data on public procurement in France (2015-2023)
<p>This accurate and comprehensive dataset encapsulates the main information published on the BOAMP website (the official journal for public procurement notices in France) from 2015 to 2023, enriched with the individual characteristics of contracting authorities and holders of public contracts. After converting the notices into a processed table, we use a machine learning algorithm to estimate the SIRETs (i.e. national identifiers) of the contracting parties, so that we can merge the open data on public procurement with individual information on public and private agents (size, legal status, main activity, geolocation...). Finally, we estimate the geolocation of foreign firms. The dataset contains about 300,000 public contracts and describes more than 1,000,000 interactions between approximately 16,000 public entities and 130,000 companies. It covers over 100 variables on the contract features, the outcome of the award procedure, the characteristics of contracting authorities and the characteristics of awarded firms.</p> <p> </p> <p>See similar data from 2024 : https://zenodo.org/records/17187786</p>
Open data repository, Knab et al., Prediction of stroke outcome in mice based on non-invasive MRI and behavioral testing
<p><strong>Open data repository, Knab et al., Prediction of stroke outcome in mice based on non-invasive MRI and behavioral testing</strong></p> <p><strong>Latest version of files: repository_v2.0.zip, Behavior Data_v2.0.xlsx and MRI IDs Testing&Replication Cohort.xlsx (please ignore repository.zip)</strong></p> <p>Open data repository Knab et al. Prediction of stroke outcome in mice based on non-invasvive MRI and behavioral testing</p> <p>Open code and documentation of prediction models available via <a href="https://github.com/major-s/mouse-mcao-outcome-predictor">https://github.com/major-s/mouse-mcao-outcome-predictor</a></p> <p><strong>Content:</strong></p> <p>README.txt</p> <p>This information</p> <p><strong>dat</strong></p> <p>Contains MRI data in NIFTI format and secondary data from atlas registration. For documentation of atlas registration files see https://pubmed.ncbi.nlm.nih.gov/28829217/<br>Files used for the manuscript:<br>t2.nii: t2 weighted image acquired 24 h post stroke<br>masklesion.nii: manually delineated lesion<br>x_masklesion.nii: lesion in atlas space<br>ix_ANO.nii: Allen brain atlas in native space (i.e. matching t2.nii)<br>Lesion volume was calculated by volume of voxels unequal 0 in x_masklesion.nii<br>Overlap of regions defined by ix_ANO.nii with masklesion.nii were used for calculating percent damage in each atlas region</p> <p><strong>prediction_models</strong></p> <p>Contains separated training and test data as xlsx and csv files with lesion volumes in cubic mm of the Allen brain atlas space, percent damage per atlas region and behavioral data. The training data was used as input for training prediction models in MATLAB, the results were created using the test data.<br>The files have following sturcture:<br>Column 1: animal ID<br>Columns 2-537: MRI regions (column title corresponds to the region number as used in the Allen common coordinate framework)<br>Column 538: lesion volume<br>Column 539: initial performance (subacute deficit) = mean performance/deficit on days 2-6<br>Column 540: mean performance/deficit on days 2-6 = initial performance (subacute deficit) - this column equals column 539 but has different header which was used to train the residual from initial deficit<br>Column 541: residual performance/deficit<br>Column 542: test or training group<br>Consecutive rows contain data for each animal specified by the animal id</p> <p>The repository also contains all trained models, prediction results for the test data and tables with resulting median absolute error (MedAE) and 5th, 25th, 75th and 95 absolute error quantiles for each model.<br>The model files end with '_models.mat' and contain 50 independently trained models each. Each model version is specified by number 1-50.<br>The result files end with '_test_results.mat' or '_test_results.xlsx', files with MedAE and quantiles end with '_test_errors.xlsx' or '_test_errors.csv. The common part of filenames specifies the used paradigm<br>Folder 'subacute deficit prediction' contains:<br> - initial_performance_from_lesion_volume: prediction of subacute deficit using lesion volume<br> - initial_performance_from_segmented_mri: prediction of subacute deficit using segmented mri<br>Folder 'long-term outcome prediction' contains:<br> - lesion_volume: prediction of residual deficit using lesion volume<br> - segmented_mri: prediction of residual deficit using segmented_mri<br> - initial_performance: prediction of residual deficit using subacute deficit<br>Folder 'mri_inc_oob_imp' contains models trained using increasing number of mri segments sorted according to the out-of-bag importance. The number of used segments is given in the file name. The models, results and errors are separated in subfolders.</p> <p>Files with equal file name and different extension always contain the same data</p> <p><strong>templates</strong><br>Allen atlas, template, brain mask, hemisphere masks, tissue probability masks in NIFTI format including annotations of region IDs and parameter.m file for use in MATLAB toolbox ANTx2<br> </p>
Local Geohistory Project: Open Data
<p>The Local Geohistory Project aims to educate users and disseminate information concerning the geographic history and structure of political subdivisions and local government. This repository contains the data used to populate the <a href="https://www.localgeohistory.pro/en/">project website</a>. The tab-separated values (TSV) files containing the data are available in the <strong>data</strong> folder, and metadata is available in the <strong>metadata</strong> folder.</p> <p>Currently, the open dataset only contains information related to New Jersey and Pennsylvania, with several scattered events concerning neighboring jurisdictions, mostly that currently border either state.</p> <p>This repository does not contain the application code, which can be found in the <a href="https://github.com/localgeohistoryproject/application">Application repository</a>, nor does it contain the table data for the bundled <strong>calendar</strong> extension.</p>
SARS-CoV-2 wastewater surveillance data and metadata in the Open Data Model format. Part 1: Québec City
<p>SARS-CoV-2 wastewater surveillance data and metadata in the Open Data Model format. Part 1: Québec City Authors</p> <ul> <li>Therrien, J-D<sup>1</sup></li> <li>Maere, T.<sup>1</sup></li> <li>Sanchez-Quete, F.<sup>2</sup></li> <li>Tsitouras, A.<sup>2</sup></li> <li>Goitom, E.<sup>3</sup></li> <li>Cloutier, F.<sup>4</sup></li> <li>Dufour, D.<sup>4</sup></li> <li>Proulx, F. <sup>4</sup></li> <li>Nicolaï, N.<sup>1</sup></li> <li>Philippe, R.<sup>1</sup></li> <li>Tohidi, M.<sup>1</sup></li> <li>Dorner, S.<sup>3</sup></li> <li>Frigon, D.<sup>2</sup></li> <li>Vanrolleghem, P.A.<sup>1</sup></li> </ul> <p>Affiliations</p> <ul> <li><sup>1</sup> model<em>EAU</em>, Département de génie civil et de génie des eaux, Université Laval</li> <li><sup>2</sup> Microbial Community Engineering Lab (MiCEL), Department of Civil Engineering, McGill University</li> <li><sup>3</sup> Polytechnique Montréal</li> <li><sup>4</sup> Ville de Québec</li> </ul> <p>General Remarks</p> <p>Wastewater-based surveillance of SARS-CoV-2 virus can detect between 1 and 30 infected individuals per 100,000 (including asymptomatic ones) by analyzing the population's sewage. As such, this method is very attractive since it costs only a fraction of clinical testing (as low as 1%). Human faeces may contain the virus a few days before a person becomes ill. Thus, this approach allows for detection of outbreaks 2-7 days before the increase in reported cases stemming from clinical screening tests (Bibby et al., 2021). Wastewater-based surveillance complements clinical testing by geolocating outbreaks, which may help targeting intensive screening programs. Moreover, it provides a quick indication of whether new public health measures (e.g., masks, social distancing, confinement, and curfew) are effective.</p> <p>Sampling</p> <p>The reported dataset contains open data collected in the province of Québec as part of the SARS-CoV-2 wastewater-based surveillance program <a href="https://www.centreau.ulaval.ca/en/covid/">CentrEau</a>-COVID. Four of the largest cities in the province (Montréal, Laval, Québec City, and Trois-Rivières), as well as the municipalities of four rural regions (Mauricie, Centre-du-Québec, Bas-St-Laurent, and Gaspésie) participated in the program. The entire dataset includes 31 sampling sites covering approximately half the population of the province of Québec (population size of 8.5 million). The timeframe covered by the dataset varies for each site. The earliest surveillance program was launched in March 2020, others followed soon after. Samples were collected using various methods, such as 24h composite samples, grab samples, and passive sampling using variations on the Moore swab method (Schang et al., 2020)</p> <p>Analysis</p> <p>Prior to the analysis of the samples for SARS-CoV-2, physiochemical parameters such as total suspended solids (TSS), turbidity, conductivity, ammonium concentration, and pH were measured. The samples were subsequently concentred by filtration using a MEC filter (0.45 um), followed by total RNA extraction using the Qiagen AllPrep PowerViral DNA/RNA Kit (Qiagen, USA) with some modifications (beta-mercaptoethanol concentration raised to 10% and lysis performed at 55 °C for 30 minutes) (Ahmed et al., 2020). SARS-CoV-2 viral RNA was detected by a one-step RT-qPCR. To assess the RNA recovery rate of the procedure, samples were spiked before extraction with a known concentration of Bovine Respiratory Syncytial Virus (BRSV) using the Zoetis INFORCE 3 vaccine (Zoetis, USA). In addition to SARS-CoV-2, samples were assessed for Pepper Mild Mottle Virus (PMMoV), the daily load of which is hypothesized to represent the fecal load contributions to the samples at a given site and time. PCR conditions and primer used to collect viral data are described in the files <code>primers.md</code> and <code>PCR conditions.md</code>.</p> <p>Compilation</p> <p>The measurements on wastewater samples carried out by the participating laboratories of this study are found in the <code>WWMeasure</code> table. The values provided by municipalities come from laboratories accredited by the Centre d'expertise en analyse environnementale du Québec (CEAEQ), in compliance with the latter's quality assurance protocols. The COVID-19-related public health data found in the <code>CPHD</code> table were collected from the Institut National de Santé Publique du Québec (INSPQ)'s public reports. Wastewater data taken in-situ at the sampling sites (e.g., the flow at pumping stations or water resource recovery facilities (WRRFs)) are found in the <code>SiteMeasure</code> table and were taken by the institutions responsible for managing the sites. All of the data, stemming from multiple sources, were combined into the <a href="https://github.com/Big-Life-Lab/ODM">Open Data Model (ODM)</a> standard format using the <a href="https://github.com/modelEAU/ODM-Import">ODM-Import python package</a> (see also Structure).</p> <p>Validation</p> <p>Wastewater and sample data were manually assessed for quality by our research collaborators. Data points for which the quality appeared to be uncertain were tagged with the value <code>True</code> in the <code>qualityFlag</code> column. Conversely, data deemed of good quality have a quality flag of <code>False</code>. Data that were not checked have a quality flag of <code>NA</code>. Textual comments describing the issues with the data points in more detail are also included in the dataset using the <code>notes</code> column of the relevant tables. Note that data validation was carried out by the data custodians responsible for each city in the dataset according to available resources. As the project continues and data validation is undertaken on more sections of the dataset, data may be re-analyzed, flagged, or commented as needed. Revisions to the dataset will be reported to the best of our ability.</p> <p>Structure</p> <p>The data contained in this dataset has been structured according to the <a href="https://github.com/Big-Life-Lab/ODM">Open Data Model (ODM) for Wastewater-Based Surveillance</a>. This model provides a standardized dictionary to collect and share data and metadata stemming from wastewater-based surveillance programs. By convention, it splits all data into 10+ thematic tables with each record representing a unique measurement, i.e., long format. For convenience, the <code>wide</code> folder presents the data found in all the other tables in a wide format, i.e., multiple measurements are aligned by <code>timestamp</code>, with each column representing a different parameter.</p> <p>Acknowledgements</p> <p>The authors would like to acknowledge that this dataset was collected thanks to the financial support of the Fonds de Recherche du Québec, the Molson Foundation, the Trottier Family Foundation, CentrEau and NSERC. The authors would also like to acknowledge the efforts of Douglas Manuel (Ottawa Hospital) and Howard Swerdfeger (Public Health Agency of Canada) for their original idea for the Open Data Model and continued development.</p> <p>References</p> <ol> <li> <p>Ahmed, W., Bertsch, P.M., Bivins, A., Bibby, K., Farkas, K., Gathercole, A., Haramoto, E., Gyawali, P., Korajkic, A., McMinn, B.R., Mueller, J.F., Simpson, S.L., Smith, W.J.M., Symonds, E.M., Thomas, K. v., Verhagen, R., Kitajima, M., 2020. Comparison of virus concentration methods for the RT-qPCR-based recovery of murine hepatitis virus, a surrogate for SARS-CoV-2 from untreated wastewater. Science of the Total Environment 739. <a href="https://doi.org/10.1016/j.scitotenv.2020.139960">https://doi.org/10.1016/j.scitotenv.2020.139960</a></p> </li> <li> <p>Bibby, K., Bivins, A., Wu, Z., North, D., 2021. Making waves: Plausible lead time for wastewater based epidemiology as an early warning system for COVID-19. Water Research 202, 117438. <a href="https://doi.org/10.1016/j.watres.2021.117438">https://doi.org/10.1016/j.watres.2021.117438</a></p> </li> <li> <p>Schang, C., Crosbie, N., Nolan, M., Poon, R., Wang, M., Jex, A., Scales, P., Schmidt, J., Thorley, B.R., Henry, R., Kolotelo, P., Langeveld, J., Schilperoort, R., Shi, B., Einsiedel, S., Thomas, M., Black, J., Wilson, S., McCarthy, D.T., 2020. Passive sampling of viruses for wastewater-based epidemiology: a case-study of SARS-CoV-2 [WWW Document]. URL <a href="https://www.researchgate.net/publication/347103410\_Passive\_sampling\_of\_viruses\_for\_wastewater-based\_epidemiology\_a\_case-study\_of\_SARS-CoV-2?channel=doi&linkId=5fd800f392851c13fe892393&showFulltext=true">https://www.researchgate.net/publication/347103410\_Passive\_sampling\_of\_viruses\_for\_wastewater-based\_epidemiology\_a\_case-study\_of\_SARS-CoV-2?channel=doi&linkId=5fd800f392851c13fe892393&showFulltext=true</a> (accessed 1.18.21).</p> </li> </ol>
Fatiando a Terra data v1.0.0: A curated collection of open geophysics data for tutorials and documentation
<p>This repository holds curated sample datasets that can be used in the documentation and tutorials of the <a href="https://www.fatiando.org/">Fatiando a Terra</a> project. All datasets are cleaned and formatted versions of openly available data under permissive licenses or in the public domain.</p> <p>More information about datasets and the code for cleaning, formatting, and preprocessing the data can be found at: <a href="https://github.com/fatiando/data">https://github.com/fatiando/data</a></p> <p>See the README.md file for information on data sources and their original licenses.</p> <p><strong>NOTE:</strong> This collection uses <a href="https://semver.org/">semantic versioning</a> (i.e., MAJOR.MINOR.BUGFIX). Major releases mean that backwards incompatible changes were made to the data. Minor releases add new data without changing existing files. Bug fix releases fix errors in a previous release that makes the data unusable. Changes to the current data files will always be published as a major release unless the file(s) in the previous release was unusable/corrupted.</p>
Global Environmental and Weather data for PyPSA-Earth: An Open Optimisation Model of the Earth Energy System.
<p><strong>PyPSA-Earth </strong>is an open model dataset of the global power system at different network levels that cover our Earth. The African model can be built using the code provided at <a href="https://github.com/pypsa-meets-africa/pypsa-africa">https://github.com/pypsa-meets-africa/pypsa-africa</a>. Other regions follow soon under the same code base.</p> <p>Since the GitHub codebase is not suited for handling large changing files, we provide here separate <strong>data bundles and cutouts</strong> to be downloaded and extracted as noted in the <a href="https://pypsa-meets-africa.readthedocs.io/en/latest/index.html">documentation</a></p> <p>The below-provided <strong>cutouts </strong>are spatiotemporal subsets of the Earth weather data from the <a href="https://software.ecmwf.int/wiki/display/CKB/ERA5+data+documentation">ECMWF ERA5</a> reanalysis dataset and the <a href="https://wui.cmsaf.eu/safira/action/viewDoiDetails?acronym=SARAH_V002">CMSAF SARAH-2</a> solar surface radiation dataset for the <strong>year 2013</strong>. They have been prepared by and are for use with the <a href="https://github.com/PyPSA/atlite">atlite</a> tool (<a href="https://atlite.readthedocs.io/">https://atlite.readthedocs.io/</a>). They can be reproduced or extended for other weather years (approx. 40-50 years) around the world by using the <a href="https://github.com/pypsa-meets-africa/pypsa-africa/blob/main/scripts/build_cutout.py">build.cutout.py</a></p> <p><strong>ECMWF ERA5</strong></p> <ul> <li><strong>Source: </strong><a href="https://cds.climate.copernicus.eu/cdsapp#!/dataset/reanalysis-era5-single-levels?tab=overview">https://cds.climate.copernicus.eu/cdsapp#!/dataset/reanalysis-era5-single-levels?tab=overview</a></li> <li><strong>Terms of Use: </strong><a href="https://cds.climate.copernicus.eu/api/v2/terms/static/20180314_Copernicus_License_V1.1.pdf">https://cds.climate.copernicus.eu/api/v2/terms/static/20180314_Copernicus_License_V1.1.pdf</a></li> </ul>
First Three-dimensional Quantification of Planktic Food Chain lower levels (Copepods) for the Ross Sea region Marine Protected Area (RSRMPA), Antarctica: Using FAIR-inspired legacy data with Machine Learning, and Open Source GIS
<p>This dataset is relative to the paper entitled: "First Three-dimensional Quantification of Planktic Food Chain lower levels (Copepods) for the Ross Sea region Marine Protected Area (RSRMPA), Antarctica: Using FAIR-inspired legacy data with Machine Learning, and Open Source GIS" publishing in journal Diversity (MPDI).</p> <p>Abstract:</p> <p>Zooplankton is a fundamental group in all aquatic ecosystems located the base of the food chain. It forms a link between the lower trophic levels with secondary consumers and shows marked fluctuations of populations with environmental change, especially reacting to heating and water acidification. At sea copepod crustaceans account for app. 70% in abundance of zooplankton and are a target of monitoring activities in key areas such as the Southern Ocean. In this study we have used FAIR-inspired legacy data (dating back to the ‘80s) collected in the Ross Sea by the Italian National Antarctic Program in GBIF.org. Together with other open-access GIS data sources and tools it allows generating, for the first time, three-dimensional predictive distribution maps for twenty-six copepod species. These predictive maps were obtained by applying machine learning techniques to grey literature data, which were visualized in open-source GIS platforms. In a Species Distribution Modeling (SDM) framework we used machine learning with three types of algorithms (TreeNet, RandomForest and Ensemble) to analyze the presence and absence of copepods at different areas and depth classes in function of environmental descriptors obtained from the Polar Macroscope Layers present in Quantartica. The models allow for the first time to map-predict the food chain in quantitative terms showing the relative index of occurrence (RIO) and identified the presence for each copepod species analyzed in the Ross Sea. Our results show marked geographical preferences that vary with species and trophic strategy. This study demonstrates that machine learning is a successful method in accurately predicting Antarctic copepod presence, also providing useful data to orient future sampling and management of wildlife and conservation.</p>
Dataset: maturity of transparency of open data ecosystems in 22 smart cities
<p>This dataset contains data collected during a study "<a href="https://www.sciencedirect.com/science/article/pii/S2210670722002281?casa_token=8xHhtKug0xEAAAAA:POKIQswXhPdbwqgi5A8q98xitcUju_VS8T7oSP6YujXdABZlc5bNn4vEHzHGoxoW16mT6hA-HZ4#!">Transparency of open data ecosystems in smart cities: Definition and assessment of the maturity of transparency in 22 smart cities</a>" (Sustainable Cities and Society (SCS), vol.82, 103906) conducted by Martin Lnenicka (University of Pardubice), Anastasija Nikiforova (University of Tartu), Mariusz Luterek (University of Warsaw), Otmane Azeroual (German Centre for Higher Education Research and Science Studies), Dandison Ukpabi (University of Jyväskylä), Visvaldis Valtenbergs (University of Latvia), Renata Machova (University of Pardubice).</p> <p>This study inspects smart cities’ data portals and assesses their compliance with transparency requirements for open (government) data by means of the expert assessment of 34 portals representing 22 smart cities, with 36 features.</p> <p>It being made public both to act as supplementary data for the paper and in order for other researchers to use these data in their own work potentially contributing to the improvement of current data ecosystems and build sustainable, transparent, citizen-centered, and socially resilient open data-driven smart cities.</p> <p>***Purpose of the expert assessment***<br> The data in this dataset were collected in the result of the applying the developed benchmarking framework for assessing the compliance of open (government) data portals with the principles of transparency-by-design proposed by Lněnička and Nikiforova (2021)* to 34 portals that can be considered to be part of open data ecosystems in smart cities, thereby carrying out their assessment by experts in 36 features context, which allows to rank them and discuss their maturity levels and (4) based on the results of the assessment, defining the components and unique models that form the open data ecosystem in the smart city context.</p> <p>***Methodology***<br> Sample selection: the capitals of the Member States of the European Union and countries of the European Economic Area were selected to ensure a more coherent political and legal framework. They were mapped/cross-referenced with their rank in 5 smart city rankings: IESE Cities in Motion Index, Top 50 smart city governments (SCG), IMD smart city index (SCI), global cities index (GCI), and sustainable cities index (SCI). A purposive sampling method and systematic search for portals was then carried out to identify relevant websites for each city using two complementary techniques: browsing and searching.<br> To evaluate the transparency maturity of data ecosystems in smart cities, we have used the transparency-by-design framework (<a href="https://www.sciencedirect.com/science/article/pii/S0736585321000447?casa_token=7K8YGcYWbQcAAAAA:_HnV50rvxwQmDYyTjLYCmUkhDM2Qpsu8TPPBgOxajkV6ammJ1BBwgtQEnMMZdVk5ONxrGNY8hOw">Lněnička & Nikiforova, 2021</a>)*.<br> The benchmarking supposes the collection of quantitative data, which makes this task an acceptability task. A six-point Likert scale was applied for evaluating the portals. Each sub-dimension was supplied with its description to ensure the common understanding, a drop-down list to select the level at which the respondent (dis)agree, and a comment to be provided, which has not been mandatory. This formed a protocol to be fulfilled on every portal. Each sub-dimension/feature was assessed using a six-point Likert scale, where strong agreement is assessed with 6 points, while strong disagreement is represented by 1 point.<br> Each website (portal) was evaluated by experts, where a person is considered to be an expert if a person works with open (government) data and data portals daily, i.e., it is the key part of their job, which can be public officials, researchers, and independent organizations. In other words, compliance with the expert profile according to the International Certification of Digital Literacy (ICDL) and its derivation proposed in <a href="https://www.emerald.com/insight/content/doi/10.1108/OIR-05-2020-0204/full/html?casa_token=6Yd7zSiQMg0AAAAA:yT8d_thrh84stDSVbax8eXLm5vP9LkrwZZFMzC_vql9vZNoQP_iYBHCZ0NOndkvusIx9TZAvJLWBp6lqe9bymm-xHaZ93k2mYfoHXKdVq52A0a7MlwGa">Lněnička et al. (2021)</a>* is expected to be met.<br> When all individual protocols were collected, mean values and standard deviations (SD) were calculated, and if statistical contradictions/inconsistencies were found, reassessment took place to ensure individual consistency and interrater reliability among experts’ answers.<br> *<a href="https://www.sciencedirect.com/science/article/pii/S0736585321000447?casa_token=7K8YGcYWbQcAAAAA:_HnV50rvxwQmDYyTjLYCmUkhDM2Qpsu8TPPBgOxajkV6ammJ1BBwgtQEnMMZdVk5ONxrGNY8hOw">Lnenicka, M., & Nikiforova, A. (2021). Transparency-by-design: What is the role of open data portals?. Telematics and Informatics, 61, 101605</a><br> *<a href="https://www.emerald.com/insight/content/doi/10.1108/OIR-05-2020-0204/full/html?casa_token=6Yd7zSiQMg0AAAAA:yT8d_thrh84stDSVbax8eXLm5vP9LkrwZZFMzC_vql9vZNoQP_iYBHCZ0NOndkvusIx9TZAvJLWBp6lqe9bymm-xHaZ93k2mYfoHXKdVq52A0a7MlwGa">Lněnička, M., Machova, R., Volejníková, J., Linhartová, V., Knezackova, R., & Hub, M. (2021). Enhancing transparency through open government data: the case of data portals and their features and capabilities. Online Information Review.</a></p> <p>***Test procedure***<br> (1) perform an assessment of each dimension using sub-dimensions, mapping out the achievement of each indicator<br> (2) all sub-dimensions in one dimension are aggregated, and then the average value is calculated based on the number of sub-dimensions – the resulting average stands for a dimension value - eight values per portal<br> (3) the average value from all dimensions are calculated and then mapped to the maturity level – this value of each portal is also used to rank the portals.</p> <p>***Description of the data in this data set***<br> Sheet#1 "comparison_overall" provides results by portal<br> Sheet#2 "comparison_category" provides results by portal and category<br> Sheet#3 "category_subcategory" provides list of categories and its elements<br> </p> <p>***Format of the file***<br> .xls</p> <p>***Licenses or restrictions***<br> CC-BY</p> <p>For more info, see README.txt</p>
OmniFold Weights | CMS 2011A Open Data | Jet Primary Dataset | pT 375-700 GeV
<p>Unfolding weights corresponding to a selection of jets from the <a href="https://doi.org/10.5281/zenodo.3340205">Jet Primary Dataset of the CMS 2011A Open Data in MOD HDF5 format</a> and associated simulated datasets. The unfolding is performed in a high-dimensional manner using the <a href="https://arxiv.org/abs/1911.09107">OmniFold</a> method, which can unfold all observables simultaneously. <a href="https://arxiv.org/abs/1810.05165">Particle Flow Networks</a> are used in Step 1 and Step 2 of the OmniFold method to process the full phase space information. The datasets and neural networks were accessed/built via the <a href="https://energyflow.network/">EnergyFlow Python package</a>. An upcoming version of the package will contain an example/demo demonstrating how to use these weights.</p> <p>The phase space selections for the data, sim, and gen datasets (using the terminology of the OmniFold paper) are:</p> <ul> <li>data: <span class="math-tex">\(p_T^{\rm jet}\in [375, 700]\)</span> GeV, <span class="math-tex">\(|\eta^{\rm jet}|<2.4\)</span>, jet quality <span class="math-tex">\(\ge\)</span> 2</li> <li>sim: <span class="math-tex">\(p_{T,\text{corr}}^{\rm jet} \in [375, 700]\)</span> GeV, <span class="math-tex">\(|\eta^{\rm jet}| < 2.4\)</span>, gen jet matched ('gen_jet_pts != -1' in EnergyFlow), jets from the <a href="https://doi.org/10.5281/zenodo.3341500">170</a> and <a href="https://doi.org/10.5281/zenodo.3341772">1800</a> MC datasets are excluded</li> <li>gen: Matched to sim jet</li> </ul> <p>The omnifold_weights.npz file contains two arrays, 'wssim' corresopnding to the Step 1 weights <span class="math-tex">\(\omega_n\)</span>, and 'wsgen' corresponding to the Step 2 weights <span class="math-tex">\(\nu_n\)</span>, for iteration <span class="math-tex">\(n\)</span>. The shape of each of these arrays is (6, 16489054), with the first axis being the iteration axis and the second axis being the event axis. There are 5 iterations, but 6 sets of weights in each array, with the 0th entry being the starting weights.</p>
CMS Open Data 2012 datasets for dimuon exercises
<p>These datasets are a subset of the CMS Open data with 2021 data-taking conditions for education purposes.</p> <p>In this version, the data and simulation files are compressed into one big file for easy access. They are stored in two different formats (CSV and PKL) with the same content, therefore just use one of them.</p> <p>Once unzipped:</p> <p>- Data files, starting with output_data_CMS_Run2012B, correspond to 4429.37 /pb of data collected by the CMS Experiment. They are a subset of the dataset on reference [1].</p> <p>- Simulation files, starting with output_sim_CMS_MonteCarlo2012, are a subset of the dataset referenced on [2]. The number of generated events in this case is 30458871, and the cross section is 3503.71.</p> <p>All the files were processed with a modified version of the AOD2NanoAODOutreachTool [3]. The small modifications are related to the number of triggers stored, and some objects like taus were removed.</p> <p> </p> <p>--------------------------------------------------------</p> <p>[1] CMS collaboration (2017). DoubleMuParked primary dataset in AOD format from Run of 2012 (/DoubleMuParked/Run2012B-22Jan2013-v1/AOD). CERN Open Data Portal. DOI:<a href="http://doi.org/10.7483/OPENDATA.CMS.YLIC.86ZZ">10.7483/OPENDATA.CMS.YLIC.86ZZ</a></p> <p>[2] Wunsch, Stefan; (2019). DYJetsToLL dataset in reduced NanoAOD format for education and outreach. CERN Open Data Portal. DOI:<a href="http://doi.org/10.7483/OPENDATA.CMS.SRRA.2GON">10.7483/OPENDATA.CMS.SRRA.2GON</a></p> <p>[3] https://github.com/cms-opendata-analyses/AOD2NanoAODOutreachTool</p>
Modular control of human movement during running: an open access data set
<p>The human body is an outstandingly complex machine including around 1000 muscles and joints acting synergistically. Yet, the coordination of the enormous amount of degrees of freedom needed for movement is mastered by our one brain and spinal cord. The idea that some synergistic neural components of movement exist was already suggested at the beginning of the XX century. Since then, it has been widely accepted that the central nervous system might simplify the production of movement by avoiding the control of each muscle individually. Instead, it might be controlling muscles in common patterns that have been called muscle synergies. Only with the advent of modern computational methods and hardware it has been possible to numerically extract synergies from electromyography (EMG) signals. However, typical experimental setups do not include a big number of individuals, with common sample sizes of five to 20 participants. With this study, we make publicly available a set of EMG activities recorded during treadmill running from the right lower limb of 135 healthy and young adults (78 males, 57 females). Moreover, we include in this open access data set the code used to extract synergies from EMG data using non-negative matrix factorization and the relative outcomes. Muscle synergies, containing the time-invariant muscle weightings (motor modules) and the time-dependent activation coefficients (motor primitives), were extracted from 13 ipsilateral EMG activities using non-negative matrix factorization. Four synergies were enough to describe as many gait cycle phases during running: weight acceptance, propulsion, early swing and late swing. We foresee many possible applications of our data, that we can summarize in three key points. First, it can be a prime source for broadening the representation of human motor control due to the big sample size. Second, it could serve as a benchmark for scientists from multiple disciplines such as musculoskeletal modelling, robotics, clinical neuroscience, sport science, etc. Third, the data set could be used both to train students or to support established scientists in the perfection of current muscle synergies extraction methods.</p> <p>The "RAW_DATA.RData" R list consists of elements of S3 class "EMG", each of which is a human locomotion trial containing cycle segmentation timings and raw electromyographic (EMG) data from 13 muscles of the right-side leg. Cycle times are structured as data frames containing two columns that correspond to touchdown (first column) and lift-off (second column). Raw EMG data sets are also structured as data frames with one row for each recorded data point and 14 columns. The first column contains the incremental time in seconds. The remaining 13 columns contain the raw EMG data, named with the following muscle abbreviations: ME = gluteus medius, MA = gluteus maximus, FL = tensor fasciæ latæ, RF = rectus femoris, VM = vastus medialis, VL = vastus lateralis, ST = semitendinosus, BF = biceps femoris, TA = tibialis anterior, PL = peroneus longus, GM = gastrocnemius medialis, GL = gastrocnemius lateralis, SO = soleus.</p> <p>The file "dataset.rar" contains data in older format, not compatible with the R package <a href="https://CRAN.R-project.org/package=musclesyneRgies">musclesyneRgies</a>.</p>
Basic data visualisations for Figshare State of Open Data 2021 survey
<p>R markdown files for:</p> <ul> <li>Downloading and cleaning data from the State of Open Data survey 2021</li> <li>Basic visualisations of responses to questions in the State of Open Data survey 2021</li> <li>HTML file of those visualisations.</li> </ul> <p>Free text fields are included in the markdown but have been turned off for knitting and in the HTML file.</p>
Data and codes: Changing the Academic Gender Narrative through Open Access
<p>This Zenodo entry includes data and R codes used to generate the figures included in the manuscript "Changing the Academic Gender Narrative through Open Access", authored by members of the Curtin Open Knowledge Initiative (COKI). These include data that are either publicly available or derived through the COKI data infrastructure.</p> <p>The R file includes codes used to generate Figures 1, 2, 3, 4, 1A, 2A and 3A. It uses data contained in the files "au_data_all.csv", "au_groupings.csv", "uk_data_all.csv" and "uk_groupings.csv".</p> <p>This entry also includes the full data files (.csv and .xlsx) for Figures 5 and 6 included in the manuscript:</p> <ul> <li>Figure 5: ‘Percentages of women academic staff (headcount) compared to the total number of academics in the institution for 43 Australian universities by grouping, 2020’. The analysis is of publicly available data sourced from the Australian Department of Education, Skills and Employment.</li> <li>Figure 6: ‘Percentages of women academic staff (headcount) compared to the total number of academics in the institution for a subset of 165 United Kingdom higher education institutions by grouping, 2020’. The analysis is of publicly available data sourced from the United Kingdom Higher Education Statistics Agency (HESA).</li> </ul>
D4.1. YouCount open data sample from the evaluation – current stand
<p>The H2020 YouCount project runs from February 2021 to January 2024. The overarching objectives are to generate new knowledge and innovations to increase the social inclusion of youth through co-creative youth citizen social science (Y-CSS) and to provide evidence of the actual outcomes of Y-CSS. Multiple case studies—consisting of 10 co-creative Y-CSS projects with young citizen scientists (YCS) aged between about 13-29 years old across nine countries in Europe—will provide knowledge about the positive drivers of social inclusion in general. The cases will further produce knowledge as well as innovations in relation to social participation, social belonging, and citizenship.</p> <p>The YouCount evaluation design for process and outcome evaluation of Y-CSS is a multi-method approach that spans across the whole duration of the project and is carried out by the WP4 of UNIVIE. It therefore is to be classified as current work in progress, as some methods only just have been implemented and will be analyzed in the future, to estimated cross-case comparisons. The deliverable aims at making the research design, as well as the current stand of the evaluative studies, transparent and publicly available. This happens in the spirit of open science, with the goal of doing “Science for and with Society”. Hereby outlined is the theoretical design, the way of carrying it out, and the current stand of each study implementation in the overall project.</p> <p>Moreover, the D4.1 includes open data regarding the outcome methodology (pre-survey questionnaire) and a sav.file with a sample of open data collected from the current pre-survey data. See more details in the report. The attached sav.-file can provide knowledge about the used variables, to estimate occurring answering patterns very roughly, ad to familiarize with the implementation of such a pre-post-survey. However, it is to be noted that this data set is exemplary, anonymized and potentially also not complete yet and must be handled and used accordingly. Due to the relatively low number of participants (yet), this research is to be characterized as early stage research and only depicts a moment in time. This being said, the YouCount project is designed to gather a huge quantity of data that promises a variety of concrete research outputs, so future data samples will be richer for in-depth analyses. At this point, more quantitative as well as qualitative data is needed to estimate real impacts of Y-CSS in all its facettes.</p>
D2.2 Open data concerning social inclusion provided on the project homepage - Emerging findings
<p>Authors to the case posters and contributors from consortium partners are described in the deliverable. </p> <p>The H2020 YouCount project runs from February 2021 to January 2024 and the consortium consists of 11 partners from nine European countries. Multiple case studies—consisting of 10 co-creative Y-CSS projects with young citizen scientists (YCS) aged between about 13-29 years old across nine countries in Europe—will provide knowledge about the positive drivers of social inclusion in general. The cases will further produce knowledge as well as innovations in relation to social participation, social belonging, and citizenship.</p> <p>In line with YouCount’s commitment to Open Science and Data Management based on the FAIR Principles, D2.2 provides a sample of open data concerning social inclusion from the research and innovation activities during the implementation period. The open data is based on informed consent and includes the following files included in the report:</p> <p>1. File 1 Homepage 30-06-22, Case descriptions.</p> <p>2. File 2 Case posters 08-06-2022, Experiences with inclusive co-creative Y-CSS in multiple case study. </p> <p>3. File 3 Narrative text 26-06-22, Experiences with developing the YouCount app toolkit, methodology.</p> <p>4. File 4 Quotes 25-06-22, Views and experiences with social inclusion of youths, YouCount/ECSA WG EIE webinars 2021 and YouCount newsletters 2022.</p> <p>5. File 5 Links to YouCount app toolkit, 28-06-22, Youths’ views and experiences with social inclusion opportunities, observations.</p> <p>Notably, the open data are based on a co-creative and flexible research design and comes in an early phase of the case studies. They can thus only be used as emerging data and preliminary findings. Still, the data contain valuable information of the research experiences and voices from young people found in the early phase of conducting hands on co-creative Y-CSS. More systematic open social inclusion data will be provided later in the project. </p> <p>The open data can also be found at the project website <a href="https://www.youcountproject.eu/">Home - YouCount - Social Citizen Science (youcountproject.eu)</a>.</p> <p>Note! They are shared under CC-BY (text) and CC-BY-ND (images case posters) due to confidentiality issues.</p>
Monitoring open access publishing of NWO funded research (2015-2021) data set
<p>This is the dataset underlying the report "Monitoring open access publishing of NWO funded research" (<a href="https://doi.org/10.5281/zenodo.7041897">https://doi.org/10.5281/zenodo.7041897</a>)</p> <p>The report presents statistics on the extent to which publications from the period 2015–2021 funded by NWO are available in Open Access. The analyses presented in this report also cover publications funded by the Netherlands Organisation for Health Research and Development ZonMw. This report builds on two earlier reports, published in <a href="https://zenodo.org/record/4446042">2020</a> and <a href="https://zenodo.org/record/5056043">2021</a>, covering publications from the period 2015–2018 and 2015-2020, respectively.</p>
Supplementary data to `Do science maps from open access literature capture the overall topic structure of an academic field?`
<p>The dataset contains the 8,528 academic articles records related to Sustainable Food research sourced with the query `TS=("sustainab*" NEAR/2 "food*")` .</p> <p>They are the records present in the largest component of the citation network, as specified in the manuscript. </p> <p>The dataset was sourced from OpenAlex based on the original data used in the manuscript and it is composed of the following columns:</p> <table> <tbody> <tr> <td><em><strong>Column</strong></em></td> <td><em><strong>Description</strong></em></td> </tr> <tr> <td>Id</td> <td>OpenAlex ID</td> </tr> <tr> <td>DOI</td> <td>Document Object Identifier</td> </tr> <tr> <td>display_name</td> <td>The article title</td> </tr> <tr> <td>publication_year</td> <td>The publication year of the article</td> </tr> <tr> <td>open_access</td> <td>An object with details of the open access status of the article</td> </tr> </tbody> </table> <p>We choose the `.rdata` format for easy loading in R. Use the function `load()` to add the data frame to the enviroment. </p>
Usability of Open Data Portals
<p><span>This dataset reports on the usability of Open Data Portals as reported by the European public. Knowledge of open science and data concepts is also reported. </span></p>
LLODIA (Linguistic Linked Open Data for Diachronic Analysis)
<p>LLODIA (Linguistic Linked Open Data for Diachronic Analysis) model developed within the Nexus Linguarum WG4 UC4.2.1 use case in humanities.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.