Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,721
datasets available to search
ShareScore release 0.9.0
Dataset results
1,721 results for “network data”
Network Modelling using Transcription Sequence Data Reveals Novel MAPK Interactions Important for Drug Resistance
GEO Series GSE55743. Aspergillus fumigatus. 53 samples. Type: Expression profiling by high throughput sequencing.
Intergrated expression profiling and ChIP-SEQ analyses of ChREBP-mediated glucose response network (expression data only)
GEO Series GSE22074. Homo sapiens. 12 samples. Type: Expression profiling by array.
Using transcriptomics data and Adverse Outcome Pathway networks to explore endocrine disrupting properties of Cadmium and PCB-126
GEO Series GSE283372. Danio rerio. 28 samples. Type: Expression profiling by high throughput sequencing.
Complex Disease Subtypes Identified by Network-Based Clustering of Gene Expression Data: Application to COPD
GEO Series GSE76705. Homo sapiens. 229 samples. Type: Expression profiling by array; Third-party reanalysis.
Multi-Omic Integrated networks connect DNA methylation and miRNA with skeletal muscle plasticity to chronic exercise in type 2 diabetic obesity [mRNA data]
GEO Series GSE58249. Homo sapiens. 34 samples. Type: Expression profiling by array.
Newly constructed network models of different WNT signaling cascades applied to breast cancer expression data
GEO Series GSE73857. Homo sapiens. 5 samples. Type: Expression profiling by high throughput sequencing.
Laser microdissection transcriptome data derived gene regulatory networks of developing rice endosperm revealed tissue- and stage-specific regulators modulating starch metabolism
GEO Series GSE181762. Oryza sativa. 24 samples. Type: Expression profiling by array.
Integrating Large-Scale Functional Genomic Data to Dissect the Complexity of Yeast Regulatory Networks
GEO Series GSE11111. Saccharomyces cerevisiae. 9 samples. Type: Expression profiling by array.
REST and Neural Gene Network Dysregulation in iPS Cell Models of Alzheimer’s Disease (RNA-seq data set)
GEO Series GSE117588. Homo sapiens. 6 samples. Type: Expression profiling by high throughput sequencing.
Supporting data for 2019SW002408: Detection of VLF attenuation in the Earth-ionosphere waveguide caused by X-class solar flares using a global lightning location network
<p>This folder contains data from the World Wide Lightning Location Network (WWLLN) used in the paper "<strong>Detection of VLF attenuation in the Earth-ionosphere waveguide caused by X-class solar flares using a global lightning location network</strong>". </p> <p>Contents:</p> <p>1. readme_2019SW002408.pdf: A detailed explanation of the data files in this folder</p> <p>2. station_list_20170906-10.csv: A table of WWLLN stations used in this analysis</p> <p>3. stroke-station_20170906.csv: A list of lightning strokes detected by WWLLN, and the stations detecting each stroke, for the day September 6, 2017.</p> <p>4. stroke-station_20170910.csv: A list of lightning strokes detected by WWLLN, and the stations detecting each stroke, for the day September 10, 2017.</p>
Input data files for RSS-NET analysis of simulated GWAS summary statistics and B cell regulatory network
<p>Details of these data files are provided in https://suwonglab.github.io/rss-net/wtccc_bcell.</p> <p>Contact:<code> xiangzhu[at]psu.edu </code></p>
Data to Replicate paper Improving Bug Detection via Context-based Code Representation Learning and Attention-based Neural Networks part 2
<p>Data to Replicate paper "Improving Bug Detection via Context-based Code Representation Learning and Attention-based Neural Networks" part 2.</p> <p>The author of the paper uploaded dataset to Google Drive. These are the same files, uploaded to Zenodo. Since detection_data.tar.gz exceeded zenodo limits, I split the data into 2 parts <em>detection_data.tar.gz</em> and <em>detection_data.tar.gz</em>. This is the first part. Splitting was achieved on OS X with:</p> <pre><code>split -b 31000m "detection_data.tar.gz" "detection_data.tar.gz."</code></pre> <p>To get original file back, run</p> <pre><code>cat detection_data.tar.gz.* > detection_data.tar.gz</code></pre> <p>GitHub link to the project: <a href="https://github.com/OOPSLA-2019-BugDetection/OOPSLA-2019-BugDetection">https://github.com/OOPSLA-2019-BugDetection/OOPSLA-2019-BugDetection</a></p>
Data to Replicate paper Improving Bug Detection via Context-based Code Representation Learning and Attention-based Neural Networks part 1
<p>Data to Replicate paper "Improving Bug Detection via Context-based Code Representation Learning and Attention-based Neural Networks" part 1. Part 2 accessible here: <a href="https://doi.org/10.5281/zenodo.3719225">https://doi.org/10.5281/zenodo.3719225</a></p> <p>The author of the paper uploaded dataset to Google Drive. These are the same files, uploaded to Zenodo. Since detection_data.tar.gz exceeded zenodo limits, I split the data into 2 parts <em>detection_data.tar.gz</em> and <em>detection_data.tar.gz</em>. This is the first part. Splitting was achieved on OS X with:</p> <pre><code>split -b 31000m "detection_data.tar.gz" "detection_data.tar.gz."</code></pre> <p>To get original file back, run</p> <pre><code>cat detection_data.tar.gz.* > detection_data.tar.gz</code></pre> <p>GitHub link to the project: <a href="https://github.com/OOPSLA-2019-BugDetection/OOPSLA-2019-BugDetection">https://github.com/OOPSLA-2019-BugDetection/OOPSLA-2019-BugDetection</a></p>
Data for manuscript: Sediment Routing and Floodplain Exchange (SeRFE): A spatially explicit model of sediment balance and connectivity through river networks
<p>Data used to calibrate and run the SeRFE model in the applications presented in the manuscript "Sediment Routing and Floodplain Exchange (SeRFE): A spatially explicit model of sediment balance and connectivity through river networks."</p>
Water Distribution--Transportation Interface Network Data
<p>Water distribution network layouts (as .inp files) used in the water distribution-- transportation interface network study submitted for publication. </p>
Dataset for "Comparing Neural Network Based Segmentation of Cardiomyocytes from Different Histology Data"
<p>This dataset contains the trained UNet TensorFlow network, the raw (tiled) histology images, and input and output inference data. The data was used for the work presented at the ISMRM Workshop on Machine Learning in 2018.</p>
IoT-deNAT: Outbound flow-based network traffic data of IoT and non-IoT devices behind a home NAT
<p>This dataset is comprised of NetFlow records, which capture the outbound network traffic of 8 commercial IoT devices and 5 non-IoT devices, collected during a period of 37 days in a lab at Ben-Gurion University of The Negev. The dataset was collected in order to develop a method for telecommunication providers to detect vulnerable IoT models behind home NATs. Each NetFlow record is labeled with the device model which produced it; for research reproducibilty, each NetFlow is also allocated to either the "training" or "test" set, in accordance with the partitioning described in:</p> <p>Y. Meidan, V. Sachidananda, H. Peng, R. Sagron, Y. Elovici, and A. Shabtai, A novel approach for detecting vulnerable IoT devices connected behind a home NAT, Computers & Security, Volume 97, 2020, 101968, ISSN 0167-4048, https://doi.org/10.1016/j.cose.2020.101968. (http://www.sciencedirect.com/science/article/pii/S0167404820302418)</p> <p> </p> <p>Please note:</p> <ul> <li>The dataset itself is free to use, however users are requested to cite the above-mentioned paper, which describes in detail the research objectives as well as the data collection, preparation and analysis.</li> <li>Following is a brief description of the features used in this dataset.</li> </ul> <p> </p> <p># NetFlow features, used in the related paper for analysis</p> <p>'FIRST_SWITCHED': System uptime at which the first packet of this flow was switched<br> 'IN_BYTES': Incoming counter for the number of bytes associated with an IP Flow<br> 'IN_PKTS': Incoming counter for the number of packets associated with an IP Flow<br> 'IPV4_DST_ADDR': IPv4 destination address<br> 'L4_DST_PORT': TCP/UDP destination port number<br> 'L4_SRC_PORT': TCP/UDP source port number<br> 'LAST_SWITCHED': System uptime at which the last packet of this flow was switched<br> 'PROTOCOL': IP protocol byte (6: TCP, 17: UDP)<br> 'SRC_TOS': Type of Service byte setting when there is an incoming interface<br> 'TCP_FLAGS': Cumulative of all the TCP flags seen for this flow</p> <p> </p> <p># Features added by the authors</p> <p>'IP': Prefix of the destination IP address, representing the network (without the host)<br> 'DURATION': Time (seconds) between first/last packet switching</p> <p> </p> <p># Label<br> 'device_model': <type>.<manufacturer>.<model number></p> <p> </p> <p># Partition<br> 'partition': Training or test</p> <p> </p> <p># Additional NetFlow features (mostly zero-variance)<br> 'SRC_AS': Source BGP autonomous system number<br> 'DST_AS': Destination BGP autonomous system number<br> 'INPUT_SNMP': Input interface index<br> 'OUTPUT_SNMP': Output interface index<br> 'IPV4_SRC_ADDR': IPv4 source address<br> 'MAC': MAC address of the source</p> <p> </p> <p># Additional data<br> 'category': IoT or non-IoT<br> 'type': IoT, access_point, smartphone, laptop<br> 'date': Datepart of FIRST_SWITCHED<br> 'inter_arrival_time': Time (seconds) between successive flows of the same device (identified by its MAC address)</p>
Squamous cell carcinoma of the lung: gene expression and network analysis during carcinogenesis (additional data)
<p>Additional files regarding the results of the research.</p>
Theater questionnair data and BP neural network prediction network
<p>More than 1500 questionnaires were distributed to the three theaters, and a total of 1382 valid questionnaires were collected for this study: 416 in small-scale theater, 478 in medium-scale theater, 488 in large-scale theater.</p>
Data from: Management and outcome of primary CNS lymphoma patients in the modern era. a LOC network study
Objective: Studies on primary CNS lymphoma (PCNSL) patients in the "real life" are scarce. Our objective was to analyze, in a nationwide population-based study, the current medical practice in the management of PCNSL. Methods: The LOC database prospectively records all newly diagnosed PCNSL from 32 French centers. Data of patients diagnosed between 2011-2016 were retrospectively analyzed. Results: We identified 1002 immunocompetent patients (43% aged>70 years, median KPS 60). First-line treatment was high-dose methotrexate-based chemotherapy in 92% of cases, with an increasing use of rituximab over time (66%). Patients <60 years old received consolidation treatment in 77% of cases, consisting of whole-brain radiotherapy (WBRT) (54%) or high-dose chemotherapy with autologous stem cell transplantation (HCT-ASCT) (23%). Among the patients > 60 years old, WBRT and HCT-ASCT consolidation were administered in only 9% and 2%, respectively. The complete response rate to initial chemotherapy was 50%. Median progression-free survival was 10.5 months. For relapse, second-line chemotherapy, HCT-ASCT, WBRT and palliative care were offered to 55%, 17%, 10%, and 18% of patients, respectively. The median, 2-year and 5-year overall survival was 25.3 months, 51% and 38%, respectively (< 60 years: not reached (NR), 70% and 61%; > 60 years: 15.4 months, 44% and 28%). Age, KPS, sex and response to induction CT were independent prognostic factors in multivariate analysis. Conclusions: Our study confirms the increasing proportion of elderly within PCNSL and shows comparable outcome in this population-based study with those reported by clinical trials, reflecting a notable application of recent PCNSL advances in the real life.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.