Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
4,694
datasets available to search
ShareScore release 0.7.1
Dataset results
4,694 results for “data analysis”
Data physicalization papers analysis (2019) V1
<p>Dataset containing a "report" of the reading of articles listed on <a href="http://dataphys.org/wiki/Bibliography">http://dataphys.org/wiki/Bibliography</a> with classification by human sense used in prototype/proposal.</p>
Processed data and analysis results for 104 RBPs
<p>This repository makes available the processed data and the results of our SURF paper. </p> <p>The paper presents the <strong>S</strong>tatistical <strong>U</strong>tility for <strong>R</strong>BP <strong>F</strong>unctions (SURF) for integrative analysis of RNA-seq and CLIP-seq data. The goal of SURF is to identify alternative splicing (AS), alternative transcription initiation (ATI), and alternative polyadenylation (APA) events regulated by individual RBPs and elucidate protein-RNA interactions governing these events. We applied the SURF pipeline to analyze 104 RBP data sets (from <a href="https://www.encodeproject.org">ENCODE</a>) and performed downstream analysis. Check out the browsable results from this <a href="http://www.statlab.wisc.edu/shiny/surf/">shiny</a> app!</p> <p>The current repository includes:</p> <ul> <li>meme.326.input.zip -- input of 326 MEME runs on SURF-inferred location features</li> <li>meme.326.output.zip -- output of 326 MEME runs on SURF-inferred location features</li> <li>surf_inferred_feature.gtf -- SURF-inferred location features for 52 RBPs</li> <li>gencode.v24.annotation.filtered.gtf -- filtered genome annotation used for ENCODE data analysis</li> <li>Homo_sapiens.GRCh37.71.primary_assembly.protein_coding.gtf -- genome annotation used for simulation study</li> <li>simulation_truth.txt -- truth parameters used for RNA-seq simulation</li> <li>[RBP].results.rds -- SURF output (each an R object) for 104 RNA-binding proteins ([RBP] the protein name). </li> </ul> <p>For reproducing the results, the source code is available at DOI: <a href="https://doi.org/10.5281/zenodo.3779853">10.5281/zenodo.3779853</a>.</p>
Experimental data for PanDDA analysis of the bromodomain of human FALZ
<p>The repository contains processed data from the entire crystallographic fragment screen of the bromodomain of human nucleosome-remodeling factor subunit BPTF (FALZ). Crystals of FALZ were screened against the DSPL and 3D-Fragment Consortium Libraries by X-ray Crystallography at the XChem facility of Diamond Light Source beamline I04-1 (FALZ_XChem_screen.tar.bz2). Additionally, metadata about the experiment can be found in the <em>mainTable</em> of the corresponding SQLite database file (FALZ_XChem_screen.sqlite). All identified ligand-bound structures were deposited in the Protein Data Bank under Group ID <strong><a href="https://www.rcsb.org/search/structure?q=pdbx_deposit_group.group_id:G_1002123">G_1002123</a></strong>. The individual PDB codes are:</p> <ul> <li>FALZA-x0079 5R4G</li> <li>FALZA-x0085 5R4H</li> <li>FALZA-x0172 5R4I</li> <li>FALZA-x0177 5R4J</li> <li>FALZA-x0271 5R4K</li> <li>FALZA-x0309 5R4L</li> <li>FALZA-x0402 5R4M</li> <li>FALZA-x0438 5R4N</li> </ul> <p>All structures necessary to reproduce the deposited PanDDA event maps which were used for ligand identification were deposited in the Protein Data Bank under PDB ID <a href="https://www.rcsb.org/structure/5R4O">5R4O</a> (group ID <strong><a href="https://www.rcsb.org/search/structure?q=pdbx_deposit_group.group_id:G_1002124">G_1002124</a></strong>).</p> <p> </p> <p><strong><em>Usage:</em></strong></p> <p>download <em>FALZ_XChem_screen.tar.bz2</em> and save into the desired project directory, e.g.</p> <pre><strong>/home/me/FALZ</strong></pre> <p>unpack the tar archive:</p> <pre><strong>tar –xvjf FALZ_XChem_screen.tar.bz2</strong></pre> <p>run pandda, e.g.</p> <pre><strong>pandda.analyse data_dirs="/home/me/FALZ/*" out_dir="/home/me/FALZ_pandda" pdb_style=dimple.pdb mtz_style=dimple.mtz</strong></pre> <p>For more information about PanDDA, please check the <a href="http://www.ccp4.ac.uk/html/pandda.html">PanDDA CCP4 program documentation</a>.</p> <p> </p> <p><em><strong>Reference:</strong></em></p> <p>Pearce, N. M. <em>et al.</em> A multi-crystal method for extracting obscured crystallographic states from conventionally uninterpretable electron density. <em>Nature Communications</em> <strong>8</strong>, ncomms15123 (2017).</p> <p> </p>
Experimental data for PanDDA analysis of human JMJD1B
<p>The repository contains processed data from the entire crystallographic fragment screen of human JMJD1B at the XChem facility of Diamond Light Source beamline I04-1 (JMJD1BA_XChem_screen.tar.bz2). Additionally, metadata about the experiment can be found in the <em>mainTable</em> of the corresponding SQLite database file (JMJD1BA_XChem_screen.sqlite). All identified ligand-bound structures were deposited in the Protein Data Bank under Group ID <strong><a href="https://www.rcsb.org/search/structure?q=pdbx_deposit_group.group_id:G_1002146">G_1002146</a>. </strong>All structures necessary to reproduce the deposited PanDDA event maps which were used for ligand identification were deposited in the Protein Data Bank under PDB ID <a href="https://www.rcsb.org/structure/5R7X">5R7X</a> (group ID <strong><a href="https://www.rcsb.org/search/structure?q=pdbx_deposit_group.group_id:G_1002141">G_1002141</a></strong>).</p> <p> </p> <p><strong><em>Usage:</em></strong></p> <p>download <em>JMJD1BA_XChem_screen.tar.bz2</em> and save into the desired project directory, e.g.</p> <p><strong>/home/me/JMJD1B</strong></p> <p>unpack the tar archive:</p> <p><strong>tar –xvjf JMJD1BA_XChem_screen.tar.bz2</strong></p> <p>run pandda, e.g.</p> <p><strong>pandda.analyse data_dirs="/home/me/JMJD1B/*" out_dir="/home/me/JMJD1B_pandda" pdb_style=dimple.pdb mtz_style=dimple.mtz</strong></p> <p>For more information about PanDDA, please check the <a href="http://www.ccp4.ac.uk/html/pandda.html">PanDDA CCP4 program documentation</a>.</p> <p> </p> <p><strong><em>Reference:</em></strong></p> <p>Pearce, N. M. <em>et al.</em> A multi-crystal method for extracting obscured crystallographic states from conventionally uninterpretable electron density. <em>Nature Communications</em> <strong>8</strong>, ncomms15123 (2017).</p> <p> </p>
Data and code for the analysis in "Assessing the impact of non-pharmaceutical interventions on SARS-CoV-2 transmission in Switzerland"
<p>Data and code used for the analysis in <em>Assessing the impact of non-pharmaceutical interventions on SARS-CoV-2 transmission in Switzerland</em> (Lemaitre et al., Swiss Medial Weekly 2020).</p>
Data Analysis for "Laser Cooling of a Nanomechanical Oscillator to Its Zero-Point Energy"
<p>Data Analysis for the paper "Laser Cooling of a Nanomechanical Oscillator to Its Zero-Point Energy". All the original data and analysis codes in Matlab are provided. In addition, we provide a python notebook with detailed description of the data analysis.</p>
Data cleaning and analysis for the Master's thesis: DIFFERENCES IN CONSUMER PREFERENCES FOR UNWEATHERED AND WEATHERED WOOD
<p>The data and analytical support the Master's thesis submitted by Hana Remesova at the University of Primorska<br> Faculty of Mathematics, Natural Sciences, and Information Technologies. The .csv files are data files, the .Rmd file is an R markdown which can be run. The product of knitting the .Rmd file is the .html.</p>
Systematic Data Analysis and Diagnostic Machine Learning Reveal Differences between Compounds with Single- and Multitarget Activity
<p>The deposited files contain balanced data sets of multi-target (MT) and single-target (ST) compounds (CPDs) used for machine learning studies (https://dx.doi.org/10.1021/acs.molpharmaceut.0c00901). The first file (st_mt_data.tsv) contains 15,142 MT- and 15,081 ST-CPDs and the second (st_dt_data.tsv) 1828 DT- and 1776 ST-CPDs. For each CPD, a nonstereo_aromatic_SMILES representation, the original ChEMBL_cid, UniProt (target) IDs, and CPD category (CPD_CAT) (i.e. DT/MT/ST) is provided. DT stands for 'diverse-target' and denotes a subset of MT-CPDs (as detailed in the publication). In addition, a CPD is tagged “Y” if it continued to be present in the data set after removal of 50% randomly selected CPDs or 50% CPD nearest neighbors (NN), respectively.</p>
Extreme to phenomenal storm wave impacts on a steep rocky coast, north Mayo, Ireland: video data, image analysis, runup and flow velocity calculations for waves of storms Fionn and Gareth.
<p>The primary data are video (.mp4) files of extreme storm wave impacts on the sites of high elevation (>=20m above high water mark) coastal boulder deposits, recorded during storms Fionn (16/01/2018) and Gareth (12/03/2019), at (54.320355, -9.569633) on the north Mayo coast of Ireland, while the significant wave height was in the range [11m,14m]. There are also .png and .jpg files derived from frames of some of the videos, relating to the analysis of the impacting wave kinematics (runup/landward propagation and flow velocities), together with physical measurements for scale determination and runup/velocity/measurement uncertainty calculations in Excel. The files EventX.mp4 are the primary data for the wave impacts EventX. The files EventX_Frame_Y.jpg are frames sampled from EventX.mp4 at constant time intervals in the temporal vicinity of the impact. The files EventX_Edges_Y.png are the edges derived from the frames with the Canny edge detector. The files EventX_Registration_Y.jpg are the impacting wavefront edges with topographical edges registered on the file ReferenceImage.jpg The files EventX.jpg are the stacked registrations for all Y, from which the impact kinematics are derived. The file Scale_Registration_Position_Velocity_Measurements_AndUncertainty.xlsx contains physical measurements for scale determination, measurements of registration error, and the calculations of impact runup/landward displacement and flow velocities, with their uncertainties. The files JetX_Leacht_a_Chúil.mp4/g are videos of large jet-producing impacts at another site.</p> <p>The files DSCN0066.MP4-DSC0085.MP4 are the raw video observations of Storm Gareth, recorded from 15:35-18:41 UT on 12 March 2019 with a Nikon Coolpix W100, while the significant wave height increased from 12m to in excess of 14m (the timestamp of these videos in Properties->Details->Media Created is one hour later than the UT of creation, because the camera's clock was set to Irish Summer Time). The file GPO15366.MP4 is an example of the GoPro (Hero 5) videos recorded simultaneously.</p>
Data from the parametric analysis of masonry panels with limit analysis
<p>For each one of the simulations performed from the parametric analysis of masonry panels with limit analysis, this database contains a .txt, a .vtk and a .png file. In the .txt file the elapse time and the collapse multiplier of each simulation can be found. The .vtk file contains all the geometry and displacement values of every masonry panel. Finally, the .png file presents the collapse mechanism obtained. </p>
Long-term live imaging and multiscale analysis identify heterogeneity and core principles of epithelial organoid morphogenesis - Image data
<p>The dataset contains raw imaging data from the work:</p> <p>"Long-term live imaging and multiscale analysis identify heterogeneity and core principles of epithelial organoid morphogenesis"</p> <p>The dataset is organized as the following: the "FigureX_" or SupplementaryFigure_X" suffix in the filename refers to the figure in the paper in which the raw data is analyzed and/or visualized. The data is "raw", i.e. not processed. However, in many cases, maximum projections of the original 3D image stacks have been uploaded due to size limitations. The total size of the image stacks approaches 0.5TB. To access the full 3D image stacks please contact the corresponding author (Francesco Pampaloni, fpampalo@bio.uni-frankfurt.de).</p> <p><strong>Authors</strong></p> <p>Lotta Hof<sup>1</sup>*, Till Moreth<sup>1</sup>*, Michael Koch<sup>1</sup>, Tim Liebisch<sup>2</sup>, Marina Kurtz<sup>3</sup>, Julia Tarnick<sup>4</sup>, Susanna M. Lissek<sup>5</sup>, Monique M.A. Verstegen<sup>6</sup>, Luc J.W. van der Laan<sup>6</sup>, Meritxell Huch<sup>7</sup>, Franziska Matthäus<sup>2</sup>, Ernst H.K. Stelzer<sup>1</sup>, Francesco Pampaloni<sup>1§</sup></p> <p><sup>1</sup>Physical Biology Group, Buchmann Institute for Molecular Life Sciences (BMLS), Goethe-Universität Frankfurt am Main, Frankfurt am Main, Germany</p> <p><sup>2</sup>Faculty of Biological Sciences, Goethe-Universität Frankfurt am Main, Frankfurt am Main, Germany</p> <p><sup>3</sup>Department of Physics, Goethe-Universität Frankfurt am Main, Frankfurt am Main, Germany</p> <p><sup>4</sup>Deanery of Biomedical Science, University of Edinburgh, Edinburgh, United Kingdom</p> <p><sup>5</sup>Experimental Medicine and Therapy Research, University of Regensburg, Regensburg, Germany</p> <p><sup>6</sup>Department of Surgery, Erasmus MC – University Medical Center, Rotterdam, The Netherlands</p> <p><sup>7</sup>The Wellcome Trust/CRUK Gurdon Institute, University of Cambridge, Cambridge, United Kingdom. Present address: Max Planck Institute of Molecular Cell Biology and Genetics, Dresden, Germany</p> <p>*contributed equally</p> <p><sup>§</sup>corresponding author: fpampalo@bio.uni-frankfurt.de</p> <p><strong>Abstract</strong></p> <p><em>Background</em></p> <p>Organoids are morphologically heterogeneous three-dimensional cell culture systems and serve as an ideal model for understanding the principles of collective cell behaviour in mammalian organs during development, homeostasis, regeneration and pathogenesis. To investigate the underlying cell organisation principles of organoids, we imaged hundreds of pancreas and cholangio carcinoma organoids in parallel using light sheet and bright field microscopy for up to seven days.</p> <p><em>Results</em></p> <p>We quantified organoid behaviour at single-cell (microscale), individual-organoid (mesoscale), and entire-culture (macroscale) levels. At single-cell resolution, we monitored formation, monolayer polarisation and degeneration, and identified diverse behaviours, including lumen expansion and decline (size oscillation), migration, rotation and multi-organoid fusion. Detailed individual organoid quantifications lead to a mechanical 3D agent-based model. A derived scaling law and simulations support the hypotheses that size oscillations depend on organoid properties and cell division dynamics, which is confirmed by bright field microscopy analysis of entire cultures.</p> <p><em>Conclusion</em></p> <p>Our multiscale analysis provides a systematic picture of the diversity of cell organisation in organoids by identifying and quantifying the core regulatory principles of organoid morphogenesis.</p>
DATA ANALYSIS - SARS-COV-2 ( Del69-70 VARIANT ) – NEW UK MUTANTS
<p>The data for S - genome sequence analysis known as Del69-70 is under variant of concern ( VOC ) . It is also termed as variant of investigation ( VUI ) . The data for VUI is statistically analysed by datewise and regionwise . The software used for data analysis is CURVE FINDER V.1.4 . The reproducibility of correlation and standard error is reported here for analysis of scattered data an attempt to study the Rational Fit and Harris Fit .</p>
Detection of Functionally Similar Code Clones: Data, Analysis Software, Benchmark
<p>We analysed 2,800 programs in Java and C for which we knew they are functionally similar. We checked if existing clone detection tools are able to find these functional similarities and classified the non-detected differences. We make all used data, the analysis software as well as the resulting benchmark available here.</p>
Flow diagram for analysis of high-throughput sequencing data
<p>Tex code and resulting pdf image, summarising the data processing pipeline of high-throughput sequencing data (fastq format files), through mapping the data to a reference genome, and then discovery and genotyping of sequence variants. The latter stage uses both 'GATK Haplotype Caller' for smaller variants, such as single-nucleotide polymorphisms and insertion-deletion polymorphisms, and Genomestrip for variants such as deletions and duplications greater than 1000 nucelotide bases in length. Note that the flow diagram is intended to represent what steps were take in the study, and does not necessarily represent the current optimum methods.</p> <p>The manuscript for which this image is a part of can be found open-access at F1000 Research "Whole genome resequencing of a laboratory-adapted <em>Drosophila melanogaster </em>population sample" https://f1000research.com/articles/5-2644/v1 doi: 10.12688/f1000research.9912.1</p> <p> </p>
The Landscape of Research Data Repositories in 2015. A re3data Analysis
<p>The attached data sets provides an overview of the landscape of research data repositories in 2015. They are based on an analysis of the re3data - registry of research data repositories from December 2015.</p>
Quantitative Content Analysis Data for Hand Labeling Road Surface Conditions in New York State Department of Transportation Camera Images
<p><strong>Foundational Codebook and Data: </strong></p> <p>Traffic camera images from the New York State Department of Transportation (511ny.org) are used to create a hand-labeled dataset of images classified into to one of six road surface conditions: 1) severe snow, 2) snow, 3) wet, 4) dry, 5) poor visibility, or 6) obstructed. Six labelers (authors Sutter, Wirz, Przybylo, Cains, Radford, and Evans) went through a series of four labeling trials where reliability across all six labelers were assessed using the Krippendorff’s alpha (KA) metric (Krippendorff, 2007). The online tool by Dr. Freelon (Freelon, 2013; Freelon, 2010) was used to calculate reliability metrics after each trial, and the group achieved inter-coder reliability with KA of 0.888 on the 4th trial. This process is known as quantitative content analysis, and three pieces of data used in this process are shared, including: 1) a PDF of the codebook which serves as a set of rules for labeling images, 2) images from each of the four labeling trials, including the use of New York State Mesonet weather observation data (Brotzge et al., 2020), and 3) an Excel spreadsheet including the calculated inter-coder reliability (ICR) metrics and other summaries used to asses reliability after each trial. The data are included in NYSDOT_quantitative_content_analysis.zip.</p> <p>The broader purpose of this work is that the six human labelers, after achieving inter-coder reliability, can then label large sets of images independently, each contributing to the creation of larger labeled dataset used for training supervised machine learning models to predict road surface conditions from camera images. The xCITE lab (xCITE, 2023) is used to store camera images from 511ny.org, and the lab provides computing resources for training machine learning models.</p> <p><strong>Obstructed Class Variation: </strong></p> <p>There are many applications for labeling roadside camera images, and as a variation of the foundational codebook, an addendum codebook provides another version of labeling the obstructed class. Specifically, this variation prioritizes labeling an image as “obstructed” only in extreme circumstances where there is a camera- or image- specific problem that prevents the assessment of any road surfaces. For labelers who want to use this version of the obstructed class (in this document) and also the other five weather-related classes (in the foundational codebook), the guidance is to use both documents in tandem, making sure to use the obstructed rules/definitions in this document while disregarding the obstructed rules/definitions in the foundational codebook. Alternatively, this codebook may be used alone in applications where the goal is to solely classify obstructed vs not obstructed. To ensure reliability and quality of this variation, quantitative content analysis was conducted on this addendum codebook, just as it was for the foundational codebook. Two labelers were tested with a sample of 30 images and achieved inter-coder reliability with Krippendorff's Alpha of 0.934 after one trial. The data, including the addendum codebook and labeling trial data (images and results) are included in ObstructedVariation_quantitative_content_analysis.zip.</p> <p>This material is based upon work supported by the U.S. National Science Foundation under Grant No. RISE-2019758.</p>
Data associated to "The Direct Cost of Contaminated Brownfield Sites on Real Estate in France: A Quasi-Exhaustive Hedonic Price Analysis"
<p>Data for replication of main results in "The Direct Cost of Contaminated Brownfield Sites on Real Estate in France: A Quasi-Exhaustive Hedonic Price Analysis". The folder "data_estim" contains all necessary data to replicate all estimations in the article (see the R code "codes_cbs-cost") with three .csv files: dvf_estim.csv, dvfbasol_estim.csv and cell200_simulation.csv. The variable names in these files are as follow:</p><p> </p><p>Identifier Variables:</p><p>- IDMUTATION: identifier for each transacted property</p><p>- comm_code: identifier for each commune defined in 2021</p><p>- admin_code: identifier for urban areas defined in 2021</p><p>- iris2014_code: identifier for each neighborhood defined in 2014</p><p>- cell200_code: identifier for each 200-meters gredded cells</p><p>- dvf_x: longitude of each transacted property (EPSG: 2154, Lambert-93, RGF93)</p><p>- dvf_y: latitude of each transacted property (EPSG: 2154, Lambert-93, RGF93)</p><p>- basol_code: identifier for each CBS (only reported in dvfbasol_estim.csv)</p><p>- anneemut: year of transaction for each property</p><p> </p><p>Dependent Variable:</p><p>- pm2: price in euro per square meter of transacted properties</p><p> </p><p>Interest Variables:</p><p>- areaha_basol250: area in hectare of CBS between 0 and 250 meters from transacted property</p><p>- areaha_basol500: area in hectare of CBS between 250 and 500 meters from transacted property</p><p>- areaha_basol1000: area in hectare of CBS between 500 and 1000 meters from transacted property</p><p>- areaha_basol2000: area in hectare of CBS between 1000 and 2000 meters from transacted property</p><p>- areaha_basol3000: area in hectare of CBS between 2000 and 3000 meters from transacted property</p><p>- area250_indpro: area in hectare of CBS with industrial manufacturing activities between 0 and 250 meters from transacted property</p><p>- area500_indpro: area in hectare of CBS with industrial manufacturing activities between 250 and 500 meters from transacted property</p><p>- area1000_indpro: area in hectare of CBS with industrial manufacturing activities between 500 and 1000 meters from transacted property</p><p>- area2000_indpro: area in hectare of CBS with industrial manufacturing activities between 1000 and 2000 meters from transacted property</p><p>- area3000_indpro: area in hectare of CBS with industrial manufacturing activities between 2000 and 3000 meters from transacted property</p><p>- area250_indoth: area in hectare of CBS with industrial non-manufacturing activities (extractive) between 0 and 250 meters from transacted property</p><p>- area500_indoth: area in hectare of CBS with industrial non-manufacturing activities (extractive) between 250 and 500 meters from transacted property</p><p>- area1000_indoth: area in hectare of CBS with industrial non-manufacturing activities (extractive) between 500 and 1000 meters from transacted property</p><p>- area2000_indoth: area in hectare of CBS with industrial non-manufacturing activities (extractive) between 1000 and 2000 meters from transacted property</p><p>- area3000_indoth: area in hectare of CBS with industrial non-manufacturing activities (extractive) between 2000 and 3000 meters from transacted property</p><p>- area250_othact: area in hectare of CBS with other or unknown activities between 0 and 250 meters from transacted property</p><p>- area500_othact: area in hectare of CBS with other or unknown activities between 250 and 500 meters from transacted property</p><p>- area1000_othact: area in hectare of CBS with other or unknown activities between 500 and 1000 meters from transacted property</p><p>- area2000_othact: area in hectare of CBS with other or unknown activities between 1000 and 2000 meters from transacted property</p><p>- area3000_othact: area in hectare of CBS with other or unknown activities between 2000 and 3000 meters from transacted property</p><p>- areaha_specific250: area in hectare of CBS specific to a unique CBS between 0 and 250 meters from transacted property (only reported in dvfbasol_estim.csv)</p><p>- areaha_specific500: area in hectare of CBS specific to a unique CBS between 250 and 500 meters from transacted property (only reported in dvfbasol_estim.csv)</p><p>- areaha_specific1000: area in hectare of CBS specific to a unique CBS between 500 and 1000 meters from transacted property (only reported in dvfbasol_estim.csv)</p><p>- areaha_specific2000: area in hectare of CBS specific to a unique CBS between 1000 and 2000 meters from transacted property (only reported in dvfbasol_estim.csv)</p><p> </p><p>Robustness Variables:</p><p>- pm2mean_iris: average transaction price per square meter of neighborhood IRIS</p><p>- shpoorhouse: share in percentage of poor households </p><p>- dvfschool_nb250: number of schools within 250 meters of property</p><p>- dvfschool_nb500: number of schools within 500 meters of property</p><p>- dvfschool_nb1000: number of schools within 1000 meters of property</p><p>- dvfschool_nb2000: number of schools within 2000 meters of property</p><p>- dvfschool_nb3000: number of schools within 3000 meters of property</p><p>- dvfroad_nb250: number of road connections within 250 meters of property</p><p>- dvfroad_nb500: number of road connections within 500 meters of property</p><p>- dvfroad_nb1000: number of road connections within 1000 meters of property</p><p>- dvfroad_nb2000: number of road connections within 2000 meters of property</p><p>- dvfroad_nb3000: number of road connections within 30000 meters of property</p><p>- dvfrail_nb250: number of railway stations within 250 meters of property</p><p>- dvfrail_nb500: number of railway stations within 500 meters of property</p><p>- dvfrail_nb1000: number of railway stations within 1000 meters of property</p><p>- dvfrail_nb2000: number of railway stations within 2000 meters of property</p><p>- dvfrail_nb3000: number of railway stations within 3000 meters of property</p><p> </p><p>Control Variables:</p><p>- center_dist: distance in kilometers of transacted property from urban area center</p><p>- sterr: surface area in square meter of parcel of each property</p><p>- sbati: surface area in square meter of building surfaces</p><p>- vente_cla: transaction through a classical process (binary variable)</p><p>- vente_adj: transaction through adjudicated process (binary variable)</p><p>- vente_ech: transaction through special exchange process (binary variable)</p><p>- vente_exp: transaction through expropriation process (binary variable)</p><p>- vente_efa: transaction before completion (binary variable)</p><p>- nblocmai: number of houses in each transaction</p><p>- nblocapt: number of apartments in each transaction</p><p>- nblocdep: number of building dependencies in each transaction</p><p>- nblocact: number of properties for commercial purpose in each transaction</p><p>- nbapt1pp: number of apartment with 1 room in each transaction</p><p>- nbapt2pp: number of apartment with 2 rooms in each transaction</p><p>- nbapt3pp: number of apartment with 3 rooms in each transaction</p><p>- nbapt4pp: number of apartment with 4 rooms in each transaction</p><p>- nbapt5pp: number of apartment with 5 and more rooms in each transaction</p><p>- nbmai1pp: number of house with 1 room in each transaction</p><p>- nbmai2pp: number of house with 2 rooms in each transaction</p><p>- nbmai3pp: number of house with 3 rooms in each transaction</p><p>- nbmai4pp: number of house with 4 rooms in each transaction</p><p>- nbmai5pp: number of house with 5 and more rooms in each transaction</p><p>- pm2mean_comm: average transaction price in euro per square meter of commune</p><p>- dvfmonument_nb500: number of historical monuments between 0 and 500 meters from transacted property</p><p>- dvfmonument_nb1000: number of historical monuments between 500 and 1000 meters from transacted property</p><p>- dvfmonument_nb2000: number of historical monuments between 1000 and 2000 meters from transacted property</p><p>- dvfindus_nb500: number of active industrial sites between 0 and 500 meters from transacted property</p><p>- dvfindus_nb1000: number of active industrial sites between 500 and 1000 meters from transacted property</p><p>- dvfindus_nb2000: number of active industrial sites between 1000 and 2000 meters from transacted property</p><p>- sh_apt: share of apartments in neighborhood IRIS</p><p>- sh_1945: share in percentage of properties with a building age before 1945</p><p>- sh_1970: share in percentage of properties with a building age before 1970</p><p>- sh_1990: share in percentage of properties with a building age before 1990</p><p>- sh_ap90: share in percentage of properties with a building age between 1990 and 2015</p><p>- sh_2015: share in percentage of properties with a building age after 2015</p><p>- clc1000_urbanhousing: share in percentage of land within 1000 meters of transacted properties with housing</p><p>- clc1000_urbanpark: share in percentage of land within 1000 meters of transacted properties with urban parks</p><p>- clc1000_recreation: share in percentage of land within 1000 meters of transacted properties with recreative activities</p><p>- clc1000_industrial: share in percentage of land within 1000 meters of transacted properties with industrial activities</p><p>- clc1000_transport: share in percentage of land within 1000 meters of transacted properties with transport infrastructures</p><p>- clc1000_nature: share in percentage of land within 1000 meters of transacted properties with natural land use</p><p>- clc1000_agr: share in percentage of land within 1000 meters of transacted properties with agricultural land use</p><p>- clc1000_forest: share in percentage of land within 1000 meters of transacted properties with forest</p><p>- clc1000_water: share in percentage of land within 1000 meters of transacted properties with water</p><p> </p><p> </p>
Dataset and data analysis of activities of Paris between 1829 and 1907
<p>Dataset construction and data analysis of 'A typology of activities over a century of urban growth', <em>Nature Cities</em>, DOI: <a href="https://www.doi.org/10.1038/s44284-024-00108-7">10.1038/s44284-024-00108-7</a></p>
Data for: Bayesian Analysis for Remote Biosignature Identification on exoEarths (BARBIE) 2: Using Grid-Based Nested Sampling in Coronagraphy Observation Simulations for O2 and O3
<p>We present all of the data across our SNR and abundance study for the molecules O2 and O3 for an exoEarth twin. The wavelength range is from 0.515-1 micron, with 25 evenly spaced 20% bandpasses in this range. The SNR ranges from 3-20, and the abundance values range in log space in steps of 0.5 and/or 0.25 (all presented in VMR in the associated table). We present the lower and upper wavelength per bandpass, the input O2 and O3 values (abundance case), the retrieved O2 and O3 values (presented as the log10(VMR)), the lower and upper limits of the 68% credible region (presented as the log10(VMR)), and the log-Bayes factor for O2 and O3. For more information about how these were calculated, please see Bayesian Analysis for Remote Biosignature Identification on exoEarths (BARBIE) 2: Using Grid-Based Nested Sampling in Coronagraphy Observation Simulations for O2 and O3, accepted and currently available on arXiv. </p> <p>To open this csv as a Pandas dataframe, use the following command:</p> <p>your_dataframe_name = pd.read_csv(f'zenodo_table.csv', dtype={'Input O2': str, {'Input O3': str}})</p>
Identifying patterns and recommendations of and for sustainable open data initiatives: a benchmarking-driven analysis of open government data initiatives among European countries
<p>This dataset contains data collected during a study <a href="https://www.sciencedirect.com/science/article/pii/S0740624X23000989"><em><strong>"Identifying patterns and recommendations of and for sustainable open data initiatives: a benchmarking-driven analysis of open government data initiatives among European countries"</strong></em></a> conducted by <em>Martin Lnenicka (University of Pardubice, Pardubice, Czech Republic), Anastasija Nikiforova (University of Tartu, Tartu, Estonia), Mariusz Luterek (University of Warsaw, Warsaw, Poland), Petar Milic (University of Pristina - Kosovska Mitrovica, Kosovska Mitrovica, Serbia), Daniel Rudmark (University of Gothenburg and RISE Research Institutes of Sweden, Gothenburg, Sweden), Sebastian Neumaier (St. Pölten University of Applied Sciences, Austria), Caterina Santoro (KU Leuven, Leuven, Belgium), Cesar Casiano Flores (University of Twente, Twente, the Netherlands), Marijn Janssen (Delft University of Technology, Delft, the Netherlands), Manuel Pedro Rodríguez Bolívar (University of Granada, Granada, Spain).</em></p> <p>It is being made public both to act as supplementary data for "<em>Identifying patterns and recommendations of and for sustainable open data initiatives: a benchmarking-driven analysis of open government data initiatives among European countries</em>", Government Information Quarterly*, and in order for other researchers to use these data in their own work. </p> <p>***Methodology***</p> <p>The paper focuses on benchmarking of open data initiatives over the years and attempts to identify patterns observed among European countries that could lead to disparities in the development, growth, and sustainability of open data ecosystems. </p> <p>This study examines existing benchmarks, indices, and rankings of open (government) data initiatives to find the contexts by which these initiatives are shaped, both of which then outline a protocol to determine the patterns. The composite benchmarks-driven analytical protocol is used as an instrument to examine the understanding, effects, and expert opinions concerning the development patterns and current state of open data ecosystems implemented in eight European countries - Austria, Belgium, Czech Republic, Italy, Latvia, Poland, Serbia, Sweden. 3-round Delphi method is applied to identify, reach a consensus, and validate the observed development patterns and their effects that could lead to disparities and divides. Specifically, this study conducts a comparative analysis of different patterns of open (government) data initiatives and their effects in the eight selected countries using six open data benchmarks, two e-government reports (57 editions in total), and other relevant resources, covering the period of 2013–2022.</p> <p>***Description of the data in this data set***</p> <p>The file "OpenDataIndex_<em>2013_</em>2022" collects an overview of 27 editions of 6 open data indices - for all countries they cover, providing respective ranks and values for these countries. These indices are:</p> <p>1) Global Open Data Index (GODI) (4 editions)</p> <p>2) Open Data Maturity Report (ODMR) (8 editions)</p> <p>3) Open Data Inventory (ODIN) (6 editions)</p> <p>4) Open Data Barometer (ODB) (5 editions)</p> <p>5) Open, Useful and Re-usable data (OURdata) Index (3 editions)</p> <p>6) Open Government Development Index (OGDI) (2 editions)</p> <p>These data shapes the third context - open data indices and rankings. The second sheet of this file covers countries covered by this study, namely, Austria, Belgium, Czech Republic, Italy, Latvia, Poland, Serbia, Sweden. It serves the basis for Section 4.2 of the paper.</p> <p>Based on the analysis of selected countries, incl. the analysis of their specifics and performance over the years in the indices and benchmarks, covering 57 editions of OGD-oriented reports and indices and e-government-related reports (2013-2022) that shaped a protocol (see paper, Annex 1), 102 patterns that may lead to disparities and divides in the development and benchmarking of ODEs were identified, which after the assessment by expert panel were reduced to a final number of 94 patterns representing four contexts, from which the recommendations defined in the paper were obtained. These patterns are available in the file "OGDdevelopmentPatterns". The first sheet contains the list of patterns, while the second sheet - the list of patterns and their effect as assessed by expert panel.</p> <p>***Format of the file***<br>.xls, .csv (for the first spreadsheet only)</p> <p>***Licenses or restrictions***<br>CC-BY</p> <p> </p> <p>For more info, see README.txt<br> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.