Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,721

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,721 results for “network data”

Learn how ShareScore rates datasets ↗
edi48/100

Supplemental soil moisture and temperature data from the saddle catchment sensor network, 2019 - 2021.

Hand-held soil moisture measurements were taken at 8 of 16 soil moisture sensors within the sensor network at Niwot Ridge to supplement the continuous measurement system at these locations. The hand-held measurements occur much less frequently than the 10 min sensor data, but they are important in determining the spatial variability of soil moisture in the alpine.

openCC (other)Feb 2023View details →
zenodo44/100

BTC-Blockchain-Network-Data

<p>File GKvolatility-2011-2017.txt contains the Garmann-Klass dailiy volatility.</p> <p>File BTC-BlockchainData_2017.txt contains the processed BTC blockchain transaction network.</p> <p>Column 1 = source node id<br> Column 2 = target node id<br> Column 3 = transaction time - unix timestamp of the transaction (seconds since<br> % 1970-01-01)<br> Column 4 = amount of bitcoins trasferred from source to node at transaction time</p> <p>For details about the data and citation:<br> ino Antulov-Fantulin, Dijana Tolic, Matija Piskorec, Zhang Ce, Irena Vodenska, Inferring short-term volatility indicators from Bitcoin blockchain, Complex Networks and Their Applications VII. COMPLEX NETWORKS 2018. Studies in Computational Intelligence, vol 813. Springer, https://doi.org/10.1007/978-3-030-05414-4_41</p> <p>Link on arxiv: https://arxiv.org/abs/1809.07856</p>

opencc-by-4.0Dec 2019View details →
zenodo44/100

alexandru-uta/cloud_network_variability_data: Cloud Network Variability Data

<p># Dataset on Cloud Network Variability</p> <p>This is the dataset obtained while benchmarking public and private clouds for the following article:</p> <p>### Alexandru Uta, Alexandru Custura, Dmitry Duplyakin, Ivo Jimenez, Jan Rellermeyer, Carlos Maltzahn, Robert Ricci, Alexandru Iosup. Is Big Data Performance Reproducible in Modern Cloud Networks?. In proceedings of 17th USENIX Symposium on Networked Systems Design and Implementation (NSDI), 2020, February 25-27, Santa Clara, USA.</p> <p>The dataset contains bandwidth variability data for:<br> - Amazon EC2<br> - Google Compute Engine<br> - Microsoft Azure (limited data)<br> - Scaleway (limited data)<br> - SURFsara HPCCloud</p> <p>The dataset contains TCP latency (RTT) data for:<br> - Amazon EC2<br> - Google Compute Engine</p> <p>The dataset contains data regarding token bucket sizes (explained in depth in the aforementioned article) for Amazon EC2.</p> <p>Please note that the full archive is over 3.5 GB of data, made up of hundreds of thousands of small files. This is why we decided to archive the data, which we then split into smaller files, to accommodate for the Github max file size of 100 MB.</p> <p>===== DETAILS ON THE ARCHIVE FORMAT =====</p> <p><strong>1. The Bandwidth Variability data:</strong><br> - archived in the bandwidth_variability_data.tar.bz2.parta, .partb, .partc<br> - after unpacking the archive, the data is split in directories per machine type, and experiment type, for example:<br> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; --- perfvar-aws-m5xlarge-fullspeed: contains iperf3 output files for continuous communication between 2 m5.xlarge VMs&nbsp;&nbsp;&nbsp; in Amazon EC2.<br> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; --- perfvar-google-4cpu-bursty-5s30s: contains iperf3 output files for bursty communication (5 seconds communication, 30 seconds break; repeat)</p> <p><strong>2. The Latency Variability data:</strong><br> - archived in the file latency_study.tar.bz2<br> - after unpacking the archive, the directories contain TCP dump RTT data and iperf3 outputs</p> <p><strong>3. The Token Bucket AWS study:</strong><br> - archived in the file token_bucket_study.tar.bz2<br> - contains files of form INSTANCE_TYPE-REGION-TIMESTAMP.{bw, raw, tb}<br> - files with extension .raw contain raw iperf3 client output<br> - files with extension .bw contain bandwidth samples taken at 1 second intervals from the iperf3 utility<br> - files with extension .tb contain a triple of form &lt;time_to_token_bucket_depletion, high_token_bucket_bandwidth, low_token_bucket_bandwidth&gt;<br> - example of a file name: c5.2xlarge-us-west-1-1567733720.raw</p> <p>&nbsp;</p>

openapache2.0Dec 2019View details →
zenodo44/100

Data and R Code from "A novel approach to sustainability assessment of food supply chains using networks of ecosystem services"

<p>Data and R code from this paper applying network analysis (iGraph) to&nbsp;two case studies pre and post agroecological transitions in Central America and Tanzania, Africa from the IPES-Food report. Further descriptions of this data and code can be found within the extended manuscript. R Code relies on the data from the scenarios (e.g., Nodes and Relations CSVs) and creates the output network metrics (e.g., Node Metric CSVs).&nbsp;</p>

opencc-by-4.0Feb 2020View details →
zenodo44/100

Internet use: participating in social networks [percentage of individuals] processed Eurostat data [CEEMID indicator]

<p>The indicator&nbsp;&#39;<strong>Internet use: participating in social networks (creating user profile, posting messages or other contributions to facebook, twitter, etc.) [percentage of individuals]</strong>&#39; from the Eurostat statistical product&nbsp;<em>Individuals who used the internet, frequency of use and activities.</em></p> <p>- NUTS2013 regional codes are recoded to NUTS2016<br> - missing data is handled with last observation carry forward, next observation carry back, linear interpolation<br> -NUTS2 areas are imputed when only NUTS1 level data is available.&nbsp;<br> <br> The original dataset is available here:<br> <a href="https://appsso.eurostat.ec.europa.eu/nui/show.do?dataset=isoc_r_iuse_i&amp;lang=en">https://appsso.eurostat.ec.europa.eu/nui/show.do?dataset=isoc_r_iuse_i&amp;lang=en</a></p> <p>More about CEEMID: <a href="http://ceemid.eu">www.ceemid.eu</a><br> Get in touch: <a href="http://danielantal.eu/#contact">danielantal.eu/#contact</a></p>

opencc-by-4.0Apr 2020View details →
zenodo44/100

Data from: The impact of human mobility networks on the global spread of COVID-19

<p>This is&nbsp;empirical dataset from the paper &quot;The impact of human mobility networks on the global spread of COVID-19&quot;. Specifically, the dataset includes several files: (a) the COVID-19 network - an origin/destination matrix (i.e., &quot;covid_network.csv&quot;); (b) the common language network - edgelist format (i.e. &quot;edge_list_comlang.csv&quot;); (c) the same continent network - edgelist format (i.e., &quot;edge_list_continent.csv&quot;; (d) the contiguity network (i.e., &quot;edge_list_contig.csv&quot;);&nbsp; (e) the migration network - edgelist format (i.e., &quot;edge_list_migration_in.csv&quot;; (f) the tourism network - edgelist format (i.e., edge_list_tourism_in.csv&quot;); (g) the list of nodes (countries) corresponding to files (b)-(e) (i.e., &quot;nodes.csv&quot;).&nbsp;Additionally, we uploaded the Rcode used in the paper (i.e. &quot;code&quot;), as a .pdf file format,&nbsp;the&nbsp;data source for the figures included in the paper (i.e., &quot;covid_network_matrix.csv&quot;, &quot;matrix_migration_out.csv&quot;, &quot;matrix_tourism.csv&quot; - Figure 1; &quot;Fig_2_a_matrix_comlang.csv&quot;, Fig_2_b_matrix_contig.csv&quot;, &quot;Fig_2_c_matrix_continent.csv&quot; - Figure 2; &quot;Fig_3.graphmlz - Figure 3; Fig_4.graphmlz - Figure 4)&nbsp;and the &quot;global network of COVID-19 onset&quot; (an individual-level data) (i.e., &quot;global_covid_network.csv&quot;).&nbsp;</p> <p>For details, please, see the Methods section of the paper:&nbsp;The impact of human mobility networks on the global spread of COVID-19&nbsp;(Hancean, M.-G., Slavinec, M., Perc, M).&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Oct 2020View details →
zenodo44/100

Network properties data and code used in "Ecological plasticity governs ecosystem services in multilayer networks".

<p>Code and network properties data used in the analyses presented in &quot;Ecological plasticity governs ecosystem services in multilayer networks&quot;. Further information can be requested of the author David A. Bohan (David.Bohan@inrae.fr).</p>

opencc-by-4.0Nov 2020View details →
zenodo44/100

Least cost network data for ancient camel transportation in the Eastern desert of Egypt - Desert Networks HiSoMA CNRS

<p>This repository contains the data necessary for the realization of a least cost network for camel transport during antiquity (Ptolemaic and Roman period) in the Egyptian eastern desert. The details of the network construction and data processing can be found in the the associated paper and datapaper.</p> <p>Study paper:<br> Mani&egrave;re, L., Cr&eacute;py, M., Redon, B. (2020) Building a Model to reconstruct the Hellenistic and Roman Road Networks of the Eastern desert of Egypt, a Semi-Empirical Approach Based on Modern Travelers&rsquo; Itineraries. DOI : <a href="http://doi.org/10.5334/jcaa.67">http://doi.org/10.5334/jcaa.67</a></p> <p>Datapaper:<br> Mani&egrave;re, L., Cr&eacute;py, M., Redon, B. (2020) Geospatial data from the &ldquo;Modelling the Hellenistic and Roman Road Networks of the Eastern desert of Egypt, a Semi-Empirical Approach Based on Modern Travelers&rsquo; Itineraries&rdquo; paper. DOI : <a href="http://doi.org/10.5334/joad.71">http://doi.org/10.5334/joad.71</a></p>

opencc-by-4.0Oct 2020View details →
zenodo44/100

SiLeNe data -- a Sinitic Lexical Network

<p>This dataset is the semi-raw data used to build the Sinitic Lexical Network, browsable at https://silene.magistry.fr/</p> <p>It is a compilation of various sources of lexical description of different languages for which sinograms are a traditional script. These sources were released as Open Data by&nbsp; different third parties and are combined here for convenience in cross-lingual linguistic studies.</p> <p>This work is a graph-based extention of a work started with Fabienne Marc on Mandarin Chinese in the early 2000&#39;s. It has since then shifted to a multilingual perspective on the Chinese script, with a larger team of young researchers.</p> <p>For the sake of simplicity, the format is a single CSV file where each row describes the reading of a sinogram in a specific word in a specific language. Its design was dictated by our need for the work on the book &laquo;&Agrave; L&#39;&Eacute;coute des Sinogrammes&raquo; (ALES), but it can be converted to build the SiLeNe graph.</p>

opencc-by-4.0Dec 2020View details →
zenodo44/100

Social networks predict the life and death of honey bees - Data

<p><strong>Interaction matrices and metadata used in &quot;Social networks predict the life and death of honey bees&quot;</strong></p> <p><a href="https://www.biorxiv.org/content/10.1101/2020.05.06.076943v2">Preprint: Social networks predict the life and death of honey bees</a></p> <p>See the README file in <a href="https://doi.org/10.5281/zenodo.4435058">bb_network_decomposition</a> for example code.</p> <p><strong>The following files are included:</strong></p> <p><strong>interaction_networks_20160729to20160827.h5</strong></p> <p>The social interaction networks as a dense tensor and metadata.</p> <p>Keys:</p> <ul> <li>interactions: Tensor of shape (29, 2010, 2010, 9) (days x individuals x individuals x interaction_types). I_{d,i,j,t} = log(1 + x), where x is the number of interactions of type t between individuals i and j at recording day d. See the methods section of paper of the interaction types.</li> <li>labels: Names of the 9 interaction types in the order they are stored in the interactions tensor.</li> <li>bee_ids: List of length 2010, mapping from sequential index used in the interaction tensor to the original BeesBook tag ID of the individual</li> </ul> <p><strong>alive_bees_bayesian.csv </strong></p> <p>This file contains the results of the bayesian lifetime model with one row for each bee.</p> <p>Columns:</p> <ul> <li>bee_id: Numerical unique identifier for each individual.</li> <li>days_alive: Number of bees the bees was determined to be alive. If the individual was still alive at the end of the recording, the number of days from the day she hatched until the end of the recording.</li> <li>death_observed: Boolean indicator whether the death occurred during the recording period.</li> <li>annotated_tagged_date: Hatch date of the individual, i.e. the date she was tagged.</li> <li>inferred_death_date: The death date as determined by the model.</li> </ul> <p><strong>bee_daily_data.csv</strong></p> <p>This file contains one row per bee per day that she was alive for the focal period.</p> <p>Columns:</p> <ul> <li>bee_id: Numerical unique identifier for each individual.</li> <li>date: Date in year-month-day format.</li> <li>age: Age in days. Can be NaN if the bee has no associated death_date.</li> <li>network_age, network_age_1, network_age_2: The first three dimensions of network age.</li> <li>dance_floor, honey_storage, near_exit, brood_area_total: Normalized (sum to 1). Can be NaN if a bee had no high confidence detections (&gt;0.9) for a given day. Can be 0 if a bee was only seen outside of the annotated areas.</li> <li>location_descriptor_count: The number of minutes the bee was seen in one of the location labels during that day. I.e., dance_floor * location_descriptor_count calculates the number of minutes, the bee was seen on the dance floor on the given day.</li> <li>death_date: Date the bee was last seen in the colony in year-month-day format. Can be NaN for individuals that did not die until the end of the recording period.</li> <li>circadian_rhythm: R&sup2; value of a sine with a period of one day fitted to the velocity data of the individual over three days. Can be NaN if the fit did not converge due to a lack of data points.</li> <li>velocity_peak_time: Phase of the circadian sine fit in hours as an offset to 12:00 UTC. Can be NaN if circadian_rhythm is NaN.</li> <li>velocity_day, velocity_night: Mean velocity of the individual between 09:00-18:00 UTC and 21:00-06:00 UTC, respectively. Can be NaN if no velocity data was available for that interval.</li> <li>days_left: Difference in days between date and death_date. Can be NaN if death_date is NaN.</li> </ul> <p><strong>location_data.csv</strong></p> <p>This file contains subsampled position information for all bees during the focal period. The data contains one row for every individual for every minute of the recording if that individual was seen at least once during that minute with a tag confidence of at least 0.9. The first matching detection for each individual is used.</p> <p>Columns:</p> <p>In addition to the bee_id and date columns as in the bee_daily_data.csv, the file contains these additional columns:</p> <ul> <li>cam_id, cams: The cam_id is a numerical identifier from {0, 1, 2, 3}. Each side of the hive is filmed by two cameras where {0, 1} and {2, 3} record the same side respectively. The cams column contains values either &ldquo;(0, 1)&rdquo; or &ldquo;(2, 3)&rdquo; and indicates to which sides of the hive this detection belongs.</li> <li>x_pos_hive, y_pos_hive: The spatial positions in millimeters on the hive. The two cameras from one side share a common coordinate system.</li> <li>location: The label that was assigned to the comb at (x_pos_hive, y_pos_hive) on the given date. The label &ldquo;other&rdquo; indicates detections that were outside of any annotated region. The label &ldquo;not_comb&rdquo; indicates the wooden frame or empty space around the comb.</li> <li>timestamp, date: The timestamp indicates the beginning of each one-minute sampling interval and is given in UTC, as indicated (example: &ldquo;2016-08-13 00:00:00+00:00&rdquo;). The date part of the timestamp is repeated in the &ldquo;date&rdquo; column. Both are given in year-month-day format.</li> </ul> <p><strong>Software used to acquire and analyze the data:</strong></p> <ul> <li><a href="https://doi.org/10.5281/zenodo.4435058">bb_network_decomposition: Network age calculation and regression analyses</a></li> <li><a href="https://github.com/BioroboticsLab/bb_pipeline/releases/tag/2016">bb_pipeline: Tag localization and decoding pipeline</a></li> <li><a href="https://github.com/BioroboticsLab/bb_pipeline_models/releases/tag/2016">bb_pipeline_models: Pretrained localizer and decoder models for bb_pipeline</a></li> <li><a href="https://github.com/BioroboticsLab/bb_binary/releases/tag/2016">bb_binary: Raw detection data storage format</a></li> <li><a href="https://doi.org/10.5281/zenodo.4436419">bb_irflash: IR flash system schematics and arduino code</a></li> <li><a href="https://github.com/BioroboticsLab/bb_imgacquisition/releases/tag/2016">bb_imgacquisition: Recording and network storage </a></li> <li><a href="https://github.com/BioroboticsLab/bb_behavior/releases/tag/2016">bb_behavior: Database interaction and data (pre)processing, velocity calculation</a></li> <li><a href="https://github.com/BioroboticsLab/bb_circadian/releases/tag/2016">bb_circadian: Circadian rhythm calculations</a></li> <li><a href="https://github.com/BioroboticsLab/bb_tracking_2016/releases/tag/2016">bb_tracking: Tracking of bee detections over time</a></li> <li><a href="https://github.com/BioroboticsLab/bb_wdd/releases/tag/2016">bb_wdd: Automatic detection and decoding of honey bee waggle dances</a></li> <li><a href="https://github.com/BioroboticsLab/bb_interval_determination/releases/tag/2016">bb_interval_determination: Homography calculation</a></li> <li><a href="https://github.com/BioroboticsLab/bb_stitcher/releases/tag/2016">bb_stitcher: Image stitching</a></li> </ul> <p>&nbsp;</p>

opencc-by-4.0Jan 2021View details →
zenodo44/100

NoSyms: A neural network approach to detecting data structures in raw memory

<p>This data was used for a experiments with graph convolutional neural networks for memory forensics as part of a bachelor thesis (included as pdf).<br> <br> Abstract:<br> <br> This work presents a neural network based approach for data structure detection in raw memory that does not require an entirely matching description of the target data structure. Instead, it&rsquo;s merely necessary to provide multiple descriptions of data structures similar to the target as training data in the form of debugging symbols. The core contribution of this work is a formal description and implementation of encoding data structure definitions as well as raw memory contents such that they can be processed by graph convolutional neural networks. A description and implementation of a neural network meant to detect data structures in the memory contents of a Linux Kernel demonstrates the practical applicability of the described approach.<br> <br> The Code is available on GitHub <a href="https://github.com/NiklasBeierl/nosyms">https://github.com/NiklasBeierl/nosyms</a>.<br> <br> nokaslr_dump is the qemu memory snapshot used to test&nbsp;the model.<br> nokaslr.raw is the &quot;raw&quot; form of the snapshot as produced by Volatility 3&#39;s layerwriter plugin.<br> symbols-training-data contains the Volatility symbol JSON files from which training data was derived.<br> nokaslr_pointers.csv lists the kernel space pointers in the snapshot and<br> nokaslr_tasks.csv lists task structs in the snapshot. Both were&nbsp;extracted via a Volatility plugins that are included in the GitHub Repo.<br> vmlinux-5.4.0-58-generic.json is the symbol file for the kernel the snapshot was taken from.<br> other-symbols.zip contains symbol files I generated vor various other kernels but did not end up using, use at your own discretion.</p>

opencc-by-4.0Jun 2021View details →
zenodo44/100

Data release for "OrchID: a Generalized Framework for Taxonomic Classification of Images Using Evolved Artificial Neural Networks"

<p><strong>Abstract</strong></p> <p>Taxonomic expertise for the identification of species is rare and costly. On-going advances in computer vision and machine learning have led to the development of numerous semi- and fully automated species identification systems. However, these systems are rarely agnostic to specific morphology, rarely can perform taxonomic &ldquo;approximation&rdquo; (by which we mean partial identification at least to higher taxonomic level if not to species), and frequently rely on costly scientific imaging technologies.</p> <p>We present a generic, hierarchical identification system for automated taxonomic approximation of organisms from images. We assessed the effectiveness of this system using photographs of slipper orchids (Cypripedioideae), for which we implemented image pre-processing, segmentation, and colour and shape feature extraction algorithms to obtain digital phenotypes for 116 species. The identification system trained on these digital phenotypes uses a nested hierarchy of artificial neural networks for pattern recognition and automated classification that mirrors the Linnean taxonomy, such that user-submitted photos can be assigned a genus, section, and species classification by traversing this hierarchy.</p> <p>Performance of the identification system varied depending on photo quality, number of species included for training, and desired taxonomic level for identification. High quality photos were scarce for some taxa and were under-represented in the training set, resulting in imbalanced network training. The image features used for training were sufficient to reliably identify photos to the correct genus but less so to the correct section and species.</p> <p>The outcomes of this project include a library of feature extraction algorithms called <em>ImgPheno</em>, a collection of scripts for neural network training called <em>NBClassify</em>, a library for evolutionary optimization of artificial neural network construction called <em>AI::FANN::Evolving</em> and a planned web application called <em>OrchID</em> for identification of user-submitted images. All project outcomes are open source and freely available.</p> <p><strong>About this release</strong></p> <p>This release corresponds belongs with our response to the reviewers of PLoS One. At this stage of the review cycle the manuscript is assessed as &#39;minor revision&#39;. Consequently, we don&#39;t anticipate making more releases until publication.</p>

opencc-zeroOct 2015View details →
zenodo44/100

Data archive: CICT for single cell RNA-seq network inference

<p>This archive contains benchmarking input data and results for using single cell gene expression data to infer gene regulatory networks (GRN) by the Causal Inference with Composition of Transactions (CICT) method and a selected set of published methods. This accompanies the manuscript "Robust discovery of gene regulatory networks from single-cell gene expression data by Causal Inference Using Composition of Transactions" (Shojaee and Huang, Brief in Bioinform 2023. DOI: 10.1093/bib/bbad370). The CICT code is available at the GitHub repo (https://github.com/hlab1/scRNAseqWithCICT/).</p><p>The original CICT algorithm was described in Shojaee et al. (arXiv:1608.02658, 2016). The benchmarked methods were included in the BEELINE benchmarking pipeline (Pratapa et al., Nat Methods 2020), to which we added DEEPDRIM (Chen et al., Brief Bioinform 2021), SCENIC (Aibar et al., Nat Methods 2017), Inferelator 3.0 (Gibbs et al., Bioinformatics 2022), and CellOracle (Kamimoto et al., Nature 2023). The output directory names are (subdirectories within each dataset):</p><p>* CICT_ewMIshrink_RFmaxdepth10_RFntrees20/: CICT for simulated data<br>* CICT_v2/: CICT for experimental data<br>* CELLORACLEDB/: CellOracle for experimental data<br>* DEEPDRIM72_ewMIshrink_RFmaxdepth10_RFntrees20/: DEEPDRIM for simulated data<br>* DEEPDRIM72_v2/: DEEPDRIM for experimental data<br>* INFERELATOR38_ewMIshrink_RFmaxdepth10_RFntrees20/: Inferelator-Prior for simulated data<br>* INFERELATOR38_v2/: Inferelator-Prior for experimental data<br>* INFERELATOR34_ewMIshrink_RFmaxdepth10_RFntrees20/: Inferelator-NoPrior for experimental data<br>* INFERELATOR34_v2/: Inferelator-NoPrior for experimental data<br>* GENIE3/: GENIE3<br>* GRNBOOST2/: GRNBOST2<br>* LEAP/: LEAP<br>* PIDC/: PIDC<br>* PPCOR/: PPCOR<br>* SCENICDB/: SCENIC for experimental data<br>* SCNS/: SCNS<br>* SCODE/: SCODE<br>* SCRIBE/: SCRIBE<br>* SINCERITIES/: SINCERITIES<br>* SINGE/: SINGE<br>* RANDOM/: RANDOM</p><p>The methods were benchmarked against two kinds of scRNA-seq datasets:<br>* Simulated datasets produced by the SERGIO simulator from a synthetic network (Dibaeinia et al., Cell Systems 2020), including complete datasets and datasets with dropouts with shape parameter k=6.5 and rate parameter q=10, 30, 50, 70, 80.&nbsp;<br>* Experimental datasets compiled by the BEELINE pipeline, evaluated at three different levels L0, L1 and L2, with three types of ground truth networks.<br>&nbsp; &nbsp; * Evaluation levels:<br>&nbsp;&nbsp; &nbsp; &nbsp; &nbsp;* L0: 500 highly varying genes plus TFs<br>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;* L1: 1000 highly varying genes plus TFs<br>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;* L2: 500 highly varying genes, TFs and 500 genes randomly selected that excluded the 1000 highly varying genes from L1.<br>&nbsp; &nbsp; * Types of ground truths:<br>&nbsp;&nbsp; &nbsp; &nbsp; &nbsp;* Cell-type-specific ChIP-seq ground truth (L0, L1, L2)<br>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;* Non-specific ChIP-seq ground truth (L0_ns, L1_ns, L2_ns)<br>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;* Loss-of-function/gain-of-function ground truth (L0_lofgof, L1_lofgof, L2_lofgof)</p><p>The directory structure is organized in accordance with the BEELINE benchmarking pipeline. For complete details please please see the BEELINE documentation (https://murali-group.github.io/Beeline/) and Github repo (https://github.com/Murali-group/Beeline).</p><p>&nbsp;</p>

opencc-by-nc-sa-4.0Jun 2023View details →
zenodo44/100

Data to "The olfactory network of larval Xenopus laevis regenerates accurately after olfactory nerve transection"

<p>This record contains analysis scripts (written in Matlab) as well as raw and processed data to reproduce the results shown in:</p> <p>Hawkins S. J., G&auml;rtner Y., Offner T., Weiss L., Maiello G., Hassenkl&ouml;ver T., and Manzini I.&nbsp; (under review) The olfactory network of larval <em>Xenopus laevis</em> regenerates accurately after olfactory nerve transection.&nbsp;</p>

opencc-by-4.0Mar 2024View details →
zenodo44/100

Prebuilt Electricity Network for PyPSA-Eur based on OpenStreetMap Data

<p>This dataset contains a<strong> topologically connected representation of the European high-voltage grid (220 kV to 750 kV)</strong> <strong>constructed using OpenStreetMap data</strong>. Input data was retrieved using the Overpass turbo API (<a title="Overpass turbo" href="https://overpass-turbo.eu" target="_blank" rel="noopener">https://overpass-turbo.eu</a>). A heurisitic cleaning process was used to for lines and links where electrical parameters are incomplete, missing, or ambiguous. Close substations within a radius of <strong>500 m</strong> are aggregated to single buses, exact locations of underlying substations is preserved. Unique identifiers for lines and links are preserved, e.g. an AC line/cable with the ID <em>way/83742802-1</em> can be viewed on OpenStreetMap using the query <a title="OpenStreetMap example (AC)" href="https://www.openstreetmap.org/way/83742802" target="_blank" rel="noopener">https://www.openstreetmap.org/way/83742802</a>. A DC line/cable with the ID <em>relation/15781671</em> can be accessed using the query <a title="OpenStreetMap example (DC)" href="https://www.openstreetmap.org/relation/15781671" target="_blank" rel="noopener">https://www.openstreetmap.org/relation/15781671</a></p> <p>A detailed explanation on the <strong>background, methodology, and validation </strong>can be found in the article published in <a href="https://www.nature.com/articles/s41597-025-04550-7"><strong>Nature Scientific Data</strong></a>:</p> <blockquote> <p><em>Xiong, B., Fioriti, D., Neumann, F., Riepin, I., Brown, T.</em> Modelling the high-voltage grid using open data for Europe and beyond. <em>Sci Data</em> <strong>12</strong>, 277 (2025). <a href="https://doi.org/10.1038/s41597-025-04550-7" target="_blank" rel="noopener">https://doi.org/10.1038/s41597-025-04550-7</a></p> </blockquote> <p><strong>Countries</strong> included in the dataset:</p> <blockquote> <p>Albania (AL), Austria (AT), Belgium (BE), Bosnia and Herzegovina (BA), Bulgaria (BG), Croatia (HR), Czech Republic (CZ), Denmark (DK), Estonia (EE), Finland (FI), France (FR), Germany (DE), Greece (GR), Hungary (HU), Ireland (IE), Italy (IT), Kosovo (XK), Latvia (LV), Lithuania (LT), Luxembourg (LU), Moldova (MD), Montenegro (ME), Netherlands (NL), North Macedonia (MK), Norway (NO), Poland (PL), Portugal (PT), Romania (RO), Serbia (RS), Slovakia (SK), Slovenia (SI), Spain (ES), Sweden (SE), Switzerland (CH), Ukraine (UA), United Kingdom (GB)</p> </blockquote> <p>The dataset was constructed as part of the workflow within the open-source, sector-coupling model PyPSA-Eur and will be updated continuously as data and/or the cleaning process improves.&nbsp;</p> <p><strong>PyPSA-Eur</strong> is an open model dataset of the European power system at the transmission network level that covers the full ENTSO-E area. It can be built using the code provided at <a href="https://github.com/PyPSA/PyPSA-eur">https://github.com/PyPSA/PyPSA-eur</a>.</p> <p><strong>Not all data dependencies</strong> are shipped with the <a href="https://github.com/PyPSA/PyPSA-eur">code repository</a>, since git is not suited for handling large changing files. Instead we provide separate <strong>data bundles</strong> to be downloaded and extracted as noted in the <a href="https://pypsa-eur.readthedocs.io/en/latest/installation.html">documentation</a>.</p> <p>While the <a href="https://github.com/PyPSA/PyPSA-eur">code</a> and provided dataset in PyPSA-Eur is released as free software under the MIT,&nbsp;<strong>different licenses and terms of use</strong> apply to the underlying input data.</p> <p><strong>Extract from OpenStreetMap Terms of Use</strong></p> <blockquote> <p>OpenStreetMap<sup><a href="https://www.openstreetmap.org/copyright#trademarks">&reg;</a></sup> is <em>open data</em>, licensed under the <a href="https://opendatacommons.org/licenses/odbl/">Open Data Commons Open Database License</a> (ODbL) by the <a href="https://osmfoundation.org/">OpenStreetMap Foundation</a> (OSMF).</p> <p>You are free to copy, distribute, transmit and adapt our data, as long as you credit OpenStreetMap and its contributors. If you alter or build upon our data, you may distribute the result only under the same licence. The full <a href="https://opendatacommons.org/licenses/odbl/1.0/">legal code</a> explains your rights and responsibilities.</p> <p>Our documentation is licensed under the <a href="https://creativecommons.org/licenses/by-sa/2.0/">Creative Commons Attribution-ShareAlike 2.0</a> license (CC BY-SA 2.0).</p> </blockquote> <p>This processed dataset is provided under the Open Data Commons Open Database License (ODbL 1.0) license.</p> <p><strong>Changelog from version 0.5 to 0.6:<br></strong></p> <ul> <li>Added electric parameters to lines (e.g. nominal current, resistance r, reactance x, susceptance b). This allows the dataset to be used outside of PyPSA/PyPSA-Eur.</li> <li>Interactive map.html now bundled with the dataset.</li> <li>Tags columns include what the element contains (e.g. merged lines contain lines that were aggregated together).</li> </ul> <p><strong>Changelog from version 0.4 to 0.5:<br></strong></p> <ul> <li>Exact locations of original substations and converter stations (interior point/Pole of Inaccessibility) are preserved.</li> <li>Clustering resolution improved from 5000 to 500 meters.</li> <li>Lines of same electric parameters are merged, if they cross a virtual bus (that is not a real substation).</li> <li>Information from OSM relations are used, wherever applicable. To avoid doubling, members (ways) of the relation are dropped in the set of lines, accordingly.</li> <li>There are now unique transformers for each voltage level in each station. Transformers now have a nominal capacity, representing the maximum of line capacities connected to either side/bus of the transformer (n-0, nominal capacity).</li> <li>Wherever applicable, OSM IDs are preserved and used in the index of the network components.</li> </ul>

openodc-odblNov 2024View details →
zenodo44/100

scGraph2Vec: a deep generative model for gene embedding augmented by Graph Neural Network and single-cell omics data

<p>This repository contains the training data and source code to reproduce the results of our paper:<br>scGraph2Vec: a deep generative model for gene embedding augmented by Graph Neural Network and single-cell omics data</p> <p>More description can be also found in GitHub (https://github.com/LPH-BIG/scGraph2Vec).</p>

opencc-zeroJun 2024View details →
zenodo44/100

Data set associated to the manuscript entitled Carbon emissions from inland waters may be underestimated: evidence from European river networks fragmented by drying by López-Rojo et. al

<p>CO2 and CH4 emissions and several associated environmental variables &nbsp;were taken in 6 European drying river networks, in 20 river reaches per river network. The field work was carried across 3 sampling campaigns in 2021, coinciding with 3 hydrological seasons (pre-dry, dry and post-rewetting) to encompass most of the hydrological variability. Each time, measures were taken in the habitats available (flowing water, dry riverbeds, isolated pools).</p>

opencc-by-4.0Jan 2024View details →
zenodo44/100

Data for "Unfolding the structural stability of nanoalloys via symmetry-constrained genetic algorithm and neural network potential"

<p><strong>PtNi_alloy_eam.db</strong> is the dataset (ase.db object) consisting of 55982 intially sampled Pt-Ni alloy structures with EAM energies and forces.</p> <p><strong>PtNi_alloy_dft.db</strong>&nbsp;is the dataset (ase.db object) consisting of the final 6828 resampled&nbsp;Pt-Ni alloy structures&nbsp;with DFT energies and forces calculated by VASP. This is the&nbsp;training set for the NNP, and could be very useful for fitting other machine learning models.</p> <p><strong>PtNi_nanoalloy_vertices_nnp.db</strong> is the dataset (ase.db object) consisting of all the vertices (stable structures) on the convex hulls obtained from NNP-based SCGA runs on 36 Pt-Ni nanoalloy systems. The energies are given by the NNP. Additional information such as mixing energy, motif and&nbsp;symmetry axis are also saved in the dataset and can be queried by the &#39;data&#39;&nbsp;keyword. An&nbsp;xyz format trajectory of these stable structures&nbsp;is also uploaded.</p> <p>All the input files and scripts for hybrid MC-MD&nbsp;simulations, QBC resampling, DFT&nbsp;calculations, NNP training, NNP-based SCGA runs&nbsp;and convex hull analysis are provided in&nbsp;<strong>inputs_and_scripts.zip</strong>.</p>

opencc-by-4.0Aug 2021View details →
zenodo44/100

Dataset of behavioral and neurophysiological data of a virtual sailing task published in: "Providing task instructions during motor training enhances performance and modulates attentional brain networks"

<p>Dataset belonging to the behavioral and neurophysiological data of the publication: &quot;Providing task instructions during motor training enhances performance and modulates attentional brain networks&quot;. The two uploaded Zip files contain kinematic and electroencephalographic data of 36 participants for the Obstacle and HorizonTask.</p>

opencc-by-4.0Jul 2021View details →
zenodo44/100

Data for "Comparing ultrastable lasers at 7×10-17 fractional frequency instability through a 2220 km optical fibre network"

<p>Here we share the relevant data of the manuscript &ldquo;Comparing ultrastable lasers at 7&times;10<sup>-17</sup> fractional frequency instability through a 2,220 km optical fibre network&rdquo;.</p> <p>Raw data was acquired using multiple synchronised, dead-time free frequency counters in Lambda-mode [1]. The integration time for each data point was 1 s. The data provided here have been processed to reflect the fractional frequency difference between the ultrastable lasers at NPL and PTB, scaled to 1542 nm. Specifically,<br> <span class="math-tex">\(y=(\nu_{\text{NPL(ULE)}}\frac{777327}{1126090}-\frac{767233}{767235}\nu_{\text{PTB(Si)}})/194.4 \ \text{THz}\)</span></p> <p>where <span class="math-tex">\(y\)</span>&nbsp;is the value recorded in the data files,&nbsp;<span class="math-tex">\(\nu_{\text{NPL(ULE)}}\)</span>&nbsp;and <span class="math-tex">\(\nu_{\text{PTB(Si)}}\)</span>&nbsp;are the optical frequencies of the ultrastable lasers at NPL (referenced to a ULE cavity) and PTB (referenced to Si cavity), respectively. The numerators and the denominators of the scaling factors correspond to mode numbers of the optical frequency comb at NPL and PTB, respectively. The expression for <span class="math-tex">\(y\)</span>&nbsp;corresponds to the fractional transfer beat [2] between the NPL and PTB ultrastable lasers.</p> <p>The file</p> <ul> <li>&ldquo;833000_s_874000_s_data_for_fig_2.txt&rdquo;</li> </ul> <p>&nbsp;contains the timeseries data used to compute the modified Allan deviation reported in <strong>Fig. 2a</strong>. The &ldquo;0&rdquo; values correspond to invalid data due to glitches in the operation of the optical fibre link. A linear drift of 40 mHz s<sup>-1</sup> has been removed in these data.</p> <p>The file</p> <ul> <li>&ldquo;432000_s_912077_s_data_for_fig_3.txt&rdquo;</li> </ul> <p>contains the timeseries data used in <strong>Fig. 3.</strong> The &ldquo;0&rdquo; values correspond to invalid data due to glitches in the operation of the optical fibre link. These data have additionally been high pass filtered with a cut off frequency of 1 mHz to decouple the short-term instability of the optical fibre link from the drift of the ultrastable lasers (with a characteristic time &gt;1000 s), as described in the manuscript.</p> <p>The files</p> <ul> <li>&ldquo;222000_s_232000_s_data_for_supp_fig_1.txt&rdquo;,</li> <li>&ldquo;270000_s_288000_s_data_for_supp_fig_1.txt&rdquo;,</li> <li>&ldquo;754000_s_765000_s_data_for_supp_fig_1.txt&rdquo;,</li> <li>&ldquo;832000_s_890000_s_data_for_supp_fig_1.txt&rdquo;,</li> </ul> <p>contain the timeseries data used to compute the modified Allan deviation reported in <strong>Supplementary Fig. 1</strong>. The &ldquo;0&rdquo; values correspond to invalid data due to glitches in the operation of the optical fibre link. A linear drift of&nbsp;40 mHz s<sup>-1</sup> has been removed in these data.</p> <p>The temporal starting point is displayed in seconds in the title of the files relative to 00:00 UTC of 2019/07/06.</p> <p>&nbsp;</p> <p><strong>References</strong></p> <p>[1]&nbsp;Dawkins, S. T., McFerran, J. J. &amp; Luiten, A. N. Considerations on the Measurement of the Stability of Oscillators with Frequency Counters.&nbsp;<em>IEEE Transactions on ultrasonics, ferroelectrics, and frequency control</em>&nbsp;<strong>54</strong>, 918-925 (2007).</p> <p>[2]&nbsp;Telle, H.R., Lipphardt, B. &amp; Stenger, J. Kerr-lens, mode-locked lasers as transfer oscillators for optical frequency measurements.&nbsp;<em>Appl. Phys. B</em>&nbsp;<strong>74</strong>, 1-6 (2002).</p> <p>&nbsp;</p>

opencc-by-4.0Dec 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record