Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,045

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,045 results for “Generated Data”

Learn how ShareScore rates datasets ↗
zenodo44/100

Generation of transcriptional novelty by transposable element insertions in Arabidopsis, Genome Sequencing and eccDNA Data

<p><strong>Raw Illumina sequencing data from the Manuscript entitled &quot;Generation of transcriptional novelty by transposable element insertions in Arabidopsis&quot;</strong></p> <p><strong>A. Illumina genome sequencing reads of Arabidopsis control and hcLines that contain novel transposable element insertions.</strong></p> <p>To identify the genomic position of the new <em>ONSEN</em> insertions, the extracted DNA of the 11 selected lines (nine lines with new insertions and two control lines) was sent to BGI, Hong-Kong for Illumina paired-end 150 bp sequencing, aiming for a minimum of 20X sequencing coverage. Quality control of the raw reads was done using FastQC (Andrews S. (2010). FastQC: a quality control tool for high throughput sequence data. Available online at: <a href="http://www.bioinformatics.babraham.ac.uk/projects/fastqc">http://www.bioinformatics.babraham.ac.uk/projects/fastqc</a>) and trimming/clipping was done using Trimmomatic with parameters ILLUMINACLIP: TruSeq3:2:30:10 LEADING:20 TRAILING:20 SLIDINGWINDOW:4:20 and MINLEN:36. Quality of the reads was deemed excellent and no further actions were taken.</p> <p>Samples identifications: genome_hcLineX with &quot;_1&quot; indicating the forward and &quot;_2&quot; the reverse reads.</p> <p><strong>B. Illumina eccDNA sequencing&nbsp;of Arabidopsis control and hcLines following stress treatments</strong></p> <p>Extrachromosomal circular DNA was prepared and sequenced as follows:&nbsp;twenty plants from each petri dish were pooled separately and DNA was extracted using the CTAB method (<a href="https://dx.doi.org/10.17504/protocols.io.quidwue">dx.doi.org/10.17504/protocols.io.quidwue</a>). Following the mobilome-seq method described in (Lanciano et al., 2017), for all samples, we digested linear DNA from 2 &micro;g of total DNA for 17 hours at 37<sup>o</sup>C using 10 U of PlasmidSafe (<em>LubioScience cat# E3101K</em>), followed by enzyme denaturation (30 mins at 70<sup>o</sup>C). Digested DNA was precipitated with isopropanol supplemented with 1 &micro;g of GlycoBlue coprecipitant (<em>Fisher Scientific cat# 10391565</em>). Circular DNA was then amplified through rolling circle amplification (RCA) with the Illustra TempliPhi kit (<em>GE Healthcare cat# 25-6400-10</em>), following the manufacturer recommendation and leaving the reaction for 16h at 30<sup>o</sup>C. DNA was once again precipitated with isopropanol and sent for Illumina paired end 150 bp sequencing at BGI, Hong Kong.&nbsp;</p> <p>Samples identification:&nbsp;</p> <p>eccDNA_A.thaliana_ctrl:&nbsp;control reads</p> <p>eccDNA_A.thaliana_HS: heat stressed plants reads</p> <p>eccDNA_A.thaliana_AZ_HS: reads of&nbsp;alpha-amanitin, zebularine and heat-stressed plants</p> <p>&quot;R1&quot; indicates forward and &quot;R2&quot; reverse reads.</p> <p>&nbsp;</p>

opencc-by-4.0Jul 2021View details →
zenodo44/100

Electricity demand data and solar generation data from Plymouth. UK

<p>This dataset was used in the Western Power Distribution Presumed Open Data competition in 2021.</p> <p>The data is provided under the&nbsp;Western Power Distribution Open Data Licence.</p> <p>There are five files:</p> <p>pv_train_set4.csv - contains solar panel data - an irradiance, power output and solar panel temperature for each half-hour from 3rd November 2017 through 2nd July 2020.</p> <p>weather_train_set4.csv - contains hourly reanalysis temperature and solar radiation at 6 weather stations near the solar panels (near Plymouth, UK).</p> <p>demand_train_set4.csv - contains half-hourly electricity demand data from a substation near Plymouth, UK, running from&nbsp;3rd November 2017 through 2nd July 2020.</p> <p>pv_test_set4.csv - contains an extra week of data to&nbsp;pv_train_set4.csv, running from 3rd July 2020 through 9th July 2020.</p> <p>demand_test_set4.csv - contains an extra week of data to&nbsp;demand_train_set4.csv,&nbsp;running from 3rd July 2020 through 9th July 2020.</p> <p>&nbsp;</p>

opencc-by-4.0Sep 2021View details →
zenodo44/100

Precipitation oxygen isoscape for mainland China from 1870 to 2017 generated based on data fusion and bias correction of iGCMs simulations

<p>The dataset includes the stable oxygen isotope of precipitation for the mainland of China over the 1870-2017 period, at a spatial resolution of 50-60 km and a monthly temporal resolution. In order to make&nbsp;full use of observations to integrate the advantages of various iGCMs, the combination of data fusion and bias correction methods are used.&nbsp;Some physical-based ancillary data are introduced in the fusion methods, including elevation and meteorological data, to enrich the climate and terrain information in the process of data fusion.&nbsp;Specifically,</p><p>(1) for the 1979-2001 period, nine simulations from six iGCMs (CAM2, GISS E, HadAM3, IsoGSM2, LMDZ4, and MIROC32) and ancillary data are fused with observations by using the CNN fusion method;</p><p>(2) for the 2002-2007 period, seven simulations from four iGCMs (GISS E, IsoGSM2, LMDZ4, and MIROC32) and ancillary data are fused by using the CNN fusion method;</p><p>(3) for the 1969-1978 period, four simulations from three iGCMs (CAM2, GISS E, and HadAM3) and ancillary data are fused by using the CNN fusion method;</p><p>(4) for the 1958-1968 and 2008-2017 periods, two iGCM simulations (CAM2 and HadAM3 for 1958-1968 and IsoGSM2 and LMDZ4 zoomed for 2008-2017) are corrected by using two BCMs, and ensemble mean (mean of four simulations) is then calculated;</p><p>(5) for the 1870-1957 period, one iGCM simulation (HadAM3) is corrected by using two BCMs, and the ensemble mean (mean of two simulations) is then calculated.</p><p>Compared with the existing iGCMs, the isoscape has high quality and stability for a large region in China at the monthly scale.&nbsp;However, it should be noted that the isoscape may be more reliable for the common periods of most iGCMs (1969-2007), but mediocre for other periods.&nbsp;</p>

opencc-by-4.0Dec 2021View details →
zenodo44/100

Dataset of 20 energy prosumers with flexibility data, distributed generation and energy storage

<p>The dataset has 20 prosumers, each with three&nbsp;appliances to provide flexibility for DR events, two PV generation resources, and an energy storage system.&nbsp;The values represent a day using 15 minutes reading periods. All the values are expressed in W, and the matrixes were created as [&nbsp;time_period x info].</p> <p>&nbsp;</p> <p>We would be grateful if you could acknowledge the use of this dataset in your publications. Please use the Zenodo publication to cite this work.</p>

opencc-by-4.0Nov 2022View details →
zenodo44/100

Global gross primary production (GPP) product generated by data fusion based on random forest

<p>Improving the ability of gross primary production (GPP) estimates to capture extreme climate perturbations and reduce the uncertainty of GPP response processes to extreme climate is a new challenge. Based on the random forest algorithm, we integrated the multimodel GPP simulation results published by the Multiscale Synthesis and Terrestrial Model Intercomparison Project, the FLUXNET flux-site-observed GPP, the standardized precipitation index (SPI) and the standardized temperature index (STI) to generate a set of global GPP time-series data products from 2001 to 2010. The new GPP product was named DFRF-GPP, referring to the GPP generated by data fusion based on random forest. DFRF-GPP is highly reliable and can be used as a valuable data source for various applications, especially in high-temperature and drought-related studies.</p>

opencc-by-4.0Nov 2022View details →
zenodo44/100

Data generated for the publication of Xie,S., Valente,L., &Etienne,R.S

<p>This repository shows the data for the publication of Xie,S., Valente,L., &amp;Etienne,R.S.&nbsp; Can we ignore trait-dependent colonization and diversification in island biogeography?</p> <p>All files were obtained via computation at University of Groningen Peregrine High Performance Computing Cluster (HPCC).<br> We use the R package DAISIErobustness and the R package DAISIE to generate the data which was analyzed in the paper. The code for these packages is version controlled on GitHub and is freely available in open-source repositories. See the Related Identifiers section for links to relevant archived versions of both these packages.</p>

opencc-by-4.0Dec 2022View details →
zenodo44/100

Diversity-Driven Unit Test Generation (Data Set)

<p>The goal of automated unit test generation tools is to create a set of test cases for the software under test that achieve the highest possible coverage for the selected test quality criteria. The most&nbsp;effective approaches for achieving this goal at the present time use meta-heuristic optimization&nbsp;algorithms to search for new test cases using fitness functions defined on existing sets of test<br> cases and the system under test. Regardless of how their search algorithms are controlled, however, all existing approaches focus on the analysis of exactly one implementation, the software&nbsp;under test, to drive their search processes, which is a limitation on the information they have&nbsp;available. In this paper we investigate whether the practical effectiveness of white box unit test&nbsp;generation tools can be increased by giving them access to multiple, diverse implementations&nbsp;of the functionality under test harvested from widely available Open Source software repositories. After presenting a basic implementation of such an approach, DivGen (Diversity-driven&nbsp;Generation), on top of the leading test generation tool for Java (EvoSuite), we assess the performance of DivGen compared to EvoSuite when applied in its traditional, mono-implementation&nbsp;oriented mode (MonoGen). The results show that while DivGen outperforms MonoGen in 33%&nbsp;of the sampled classes for mutation coverage (+16% higher on average), MonoGen outperforms<br> DivGen in 12.4% of the classes for branch coverage (+10% higher average).</p>

opencc-by-4.0Aug 2022View details →
zenodo44/100

Two metabolomics data sets (mouse kidney, mouse plasma), generated for the publication Bignon et al., 2023: "Multiomics reveals multilevel control of renal and systemic metabolism by the renal tubular circadian clock".

<p><strong>Publication: </strong>Bignon Y, Wigger L, Ansermet C, Weger BD, Lagarrigue S, Centeno G, Durussel F, G&ouml;tz L, Ibberson M, Pradervand S, Quadroni M, Weger M, Amati F, Gachon F, Firsov D. Multiomics reveals multilevel control of renal and systemic metabolism by the renal tubular circadian clock. J Clin Invest. 2023 Mar 2:e167133. doi: 10.1172/JCI167133. Epub ahead of print. PMID: 36862511.</p> <p>&nbsp;</p> <p><strong>Abstract: </strong> Circadian rhythmicity in renal function suggests rhythmic adaptations in renal metabolism. To decipher the role of the circadian clock in renal metabolism, we studied diurnal changes in renal metabolic pathways using integrated transcriptomic, proteomic, and metabolomic analysis performed on control mice and mice with inducible deletion of the circadian clock regulator Bmal1 in the renal tubule (cKOt). With this unique resource, we demonstrated that ~30% RNAs, ~20% proteins and ~20% metabolites are rhythmic in kidneys of control mice. Several key metabolic pathways including NAD+ biosynthesis, fatty acid transport, carnitine shuttle,and b-oxidation displayed impairments in kidneys of cKOt, resulting in a perturbed mitochondrial activity. Carnitine reabsorption from the primary urine was one of the most impacted processes with a ~50% reduction in plasma carnitine levels and a parallel systemic decrease in tissues carnitine content. This suggests that the circadian clock in the renal tubule controls both kidney and systemic physiology.</p> <p>&nbsp;</p> <p><strong>This record contains two separate mass-spectrometry metabolomics data sets associated with this study:</strong></p> <ol> <li>Metabolic profile of renal tubules, MS/MS data, Metabolon, Morrisville, NC (N=60)</li> <li>Metabolic profile of blood plasma, MS/MS data, Biocrates, Innsbruck, Austria (N=60)</li> </ol> <p>For each data set, original data as received from the platforms and processed data as used in the data analysis are provided. Preprocessing of kidney data included removal of metabolites with more than 80% missing data values, median normalization, imputation and glog2 transformation. Preprocessing of plasma data included filtering of metabolites with any missing data and log2 transformation. Details of data processing are available in the STAR*methods of the publication.</p> <p>&nbsp;</p> <p><strong>Data sets in other repositories associated with the same study:</strong></p> <p>Additional data sets (transcriptomics, proteomics) pertaining to the same&nbsp;study have been deposited in public repositories:</p> <ul> <li>Gene Expression Omnibus (NCBI GEO), GSE216252</li> <li>PRIDE Archive (EMBL-EBI), PXD036803</li> </ul> <p>&nbsp;</p>

opencc-by-4.0Mar 2023View details →
zenodo44/100

Air mass trajectory and connectivity data generated with tropolink (Richard et al., 2023)

<p>Archive containing trajectory and connectivity data generated with tropolink for the preparation of the manuscript Richard et al. (2023, <a href="https://doi.org/10.1029/2023GH000885">https://doi.org/10.1029/2023GH000885</a>), as well as the corresponding specifications (node coordinates, dates and other tropolink&nbsp;options). The archive contains specifications, trajectories and connectivities for the three applications presented in the manuscript:</p><p>- the study of airborne connectivity between areas of production of sugar beet, with starting altitude equal to 250m, 500m and 750m above ground level;</p><p>- the study of airborne connectivity between potyvirus populations;</p><p>- the study of invasion risk of Spodoptera frugiperda in Europe, North Africa and western Asia;</p><p>&nbsp;</p><p>Web application tropolink:&nbsp;https://tropolink.fr/</p><p>Associated gitlab: https://forgemia.inra.fr/tropo-group</p><p>Accompanying wiki: https://forgemia.inra.fr/tropo-group/tropolink/-/wikis</p><p>R code for analyzing tropolink output:&nbsp;https://forgemia.inra.fr/tropo-group/tropolink/-/wikis/Examples</p><p>Richard H., Martinetti D., Lercier D., Fouillat Y., Hadi B., Elkahky M., Ding J., Michel L., Morris C.E., Berthier K., Maupas F.,&nbsp;<br>Soubeyrand S. (2023). Computing geographical networks generated by air-mass movement. GeoHealth 7:e2023GH000885. <a href="https://doi.org/10.1029/2023GH000885">https://doi.org/10.1029/2023GH000885</a>.</p>

opencc-by-4.0May 2023View details →
zenodo44/100

Research generated data supporting the article manuscript "Setting Grounds for Data Literacy in the Sector of Agriculture: Learning About and with Open Data"

<p>In the research 345 MS courses and 216 MS courses data from the ECTS catalogue (2019) of University of Zagreb Faculty of Agriculture were mapped onto the data literacy competence areas (theme) and DL competence areas sub-themes adapted ODI Data Skills Framework (2020) expanding the term &ldquo;skill&rdquo; to &ldquo;competence&rdquo; to include knowledge and attitudes.&nbsp;Teaching staff was interviewed in semi-structured interviews on the data literacy competences covered in their courses and open data use and teaching in their courses as well as their perceived importance for the sector of the course.</p> <p>The upload consists of the following&nbsp;.csv files:</p> <table> <tbody> <tr> <td>readme_DL_OD_Salamonetal.csv</td> </tr> <tr> <td>01DL_OD_Salamonetal.csv</td> </tr> <tr> <td>02DL_OD_Salamonetal.csv</td> </tr> <tr> <td>03DL_OD_Salamonetal.csv</td> </tr> <tr> <td>04DL_OD_Salamonetal.csv</td> </tr> <tr> <td>05DL_OD_Salamonetal.csv</td> </tr> <tr> <td>06DL_OD_Salamonetal.csv</td> </tr> <tr> <td>07DL_OD_Salamonetal.csv</td> </tr> </tbody> </table> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Aug 2023View details →
zenodo40/100

GC-MS data set for Generation of a chromosome-scale genome assembly of the insect-repellant terpenoid-producing Lamiaceae species, Callicarpa americana

<p>RAW GC/MS data set for characterization of class II terpene synthases from <em>Callicarpa americana&nbsp;</em></p>

opencc-by-4.0Feb 2020View details →
zenodo40/100

Constraining the dense matter equation of state with joint analysis of NICERand LIGO/Virgo measurements: Data for generating plots

<p>In this repository you will find a Jupyter&nbsp;notebook with code to generate the plots from the paper <em>Constraining the dense matter equation of state with joint analysis of NICER and LIGO/Virgo measurements</em>&nbsp;by&nbsp;Raaijmakers et al. (2020).</p>

opencc-by-4.0Mar 2020View details →
zenodo40/100

MAOn: A Data-driven Methodology to Generate Living Ontologies

<p>This repository presents the MAnto Lite ontology created with our MAOn methodology in the context of transport and public and the accessibility it provides. Besides, a set of annotated data with the ontology as a validation method is presented. The MAOn methodology is characterized by being data-based, by creating live ontologies and by a thorough evaluation process of the created ontology.</p>

opencc-by-4.0Jun 2020View details →
zenodo40/100

Synthetic COVID-19 Case Reporting Data Generated from an Agent-Based Simulation Model

<p>This is a synthetic case reporting data set for the SARS-CoV-2 epidemic in Austria. The data set statistically reproduces and synthetically augments data on reported cases and was generated with an agent-based simulation model. References to descriptions of the model and the parameterization used to generate the data set is included in the attached PDF file. The data format is described in the README file.</p>

opencc-by-4.0Sep 2020View details →
zenodo40/100

Raman spectroscopic data derived from Calluna vulgaris charcoals, experimentally generated across a range of natural wildfire temperatures

<p>This data has been derived from deconvolved Raman spectra, utilising two first order bands - D (Disordered) and G (Graphitic). Spectra were collected from experimentally pyrolysed charcoals, made from Calluna vulgaris (Ling Heather) separated into three main components; stem, root and flower. For each component at 250, 400, 600 and 800 degrees centigrade respectively, 5 charcoal samples (A, B, C, D, E) were analysed. Following deconvolution, median values for each spectra were produced. These correspond to parameters derived from the Raman data, including D- and G-band width (FWHM), intensity (ID/IG or &#39;R1&#39;) and area (AD/AG) ratios, band separation (G-D or &#39;RBS&#39;), and band width ratios (D-FWHM/G-FWHM). All parameters have been compiled for each component material, and displayed graphically within this dataset.</p>

opencc-by-4.0Nov 2020View details →
zenodo40/100

Onshore & offshore WRF generated wind data

<p>These data sets provide the WRF [1] calculated wind data for Pritzwalk (onshore) and FINO3 (offshore) as Python dictionaries.&nbsp; Additionally, the files contain k-means cluster objects derived from these profiles. These data sets were used for power assessment and design exploration of Airborne Wind Energy Systems using the awebox [2] optimization toolbox.</p> <p>&nbsp;</p> <p>WRF setups are described in detail and used in publication [3,4,5].</p> <p>Wind data are interpolated to fixed heights of: [10, &nbsp; 28, &nbsp; 50, &nbsp; 70, &nbsp; 90,&nbsp; 100,&nbsp; 150,&nbsp; 200,&nbsp; 250,&nbsp; 300,&nbsp; 350, 400,&nbsp; 450,&nbsp; 500,&nbsp; 550,&nbsp; 600, &nbsp; 700,&nbsp; 800,&nbsp; 1000, 1200] meters above ground.</p> <p>&nbsp;</p> <p>Onshore wind data:&nbsp;</p> <ul> <li> <p>Location lat: 53&deg; 10.78&#39; N; long: 12&deg; 11.35&#39; E</p> </li> <li> <p>Time: 1 September 2015 - 31 August 2016</p> </li> <li> <p>Timestep: 10 min</p> </li> </ul> <p>Offshore wind data:&nbsp;</p> <ul> <li> <p>Location lat: 55&deg; 11.7&#39; N, long: 7&deg; 9.5&#39; E</p> </li> <li> <p>Time: 1 September 2013 - 31 August 2014</p> </li> <li> <p>Timestep: 10 min</p> </li> </ul> <p>&nbsp;</p> <p>The clusters are derived from both horizontal wind velocity components using the scikit-learn&rsquo;s k-means clustering algorithm [6]. For our purposes, wind vectors were rotated such that the main wind speed always points in the same direction (u_main,u_deviation).</p> <p>[1]: <a href="https://www.mmm.ucar.edu/weather-research-and-forecasting-model"> Weather Research and Forecasting Model </a></p> <p>[2]: <a href="https://github.com/awebox/awebox">awebox</a></p> <p>[3]: <a href="https://doi.org/10.5194/wes-4-563-2019">Improving mesoscale wind speed forecasts using lidar-based observation nudging for airborne wind energy systems</a></p> <p>[4]: <a href="https://doi.org/10.5194/wes-2020-120">Offshore and onshore ground-generation airborne wind energy power curve characterization </a></p> <p>[5]:<a href="https://doi.org/10.5194/wes-2020-123">Ground-generation airborne wind energy design space exploration </a></p> <p>[6]: <a href="https://scikit-learn.org/stable/modules/generated/sklearn.cluster.KMeans.html">sklearn.cluster.KMeans</a></p>

opencc-by-4.0Nov 2020View details →
zenodo40/100

Training dataset: Generation of a spectral library from HEK-Ecoli Spike-in mass spectrometry data

<p>The five raw files serve as a concise but meaningful training data set in the Galaxy training network (https://galaxyproject.github.io/training-material/).</p> <p>HEK and E.coli cell pellets were lysed with 5 % SDS, 50 mM triethylammonium bicarbonate (TEAB), pH 7.55. The obtained protein extracts were reduced by adding f.c. 5 mM TCEP and alkylated by the addition of f.c. 10 mM iodacetamide. Protein digestion and purification was performed on S-Trap columns. To ensure protein binding to the S-Trap columns, samples were acidified to a final concentration of 1.2 % phosphoric acid (~ pH 2). Six times the sample volume S-Trap buffer (90% aqueous methanol containing a final concentration of 100 mM TEAB, pH 7.1) was added to the samples which were then loaded on the columns and washed with S-Trap buffer. Protein digestion was performed with trypsin and LysC for one hour at 47 &deg;C. Peptides were eluted in three steps with (1) 50 mM TEAB, (2) 0.2 % aqueous formic acid and (3) 50 % acetonitrile containing 0.2 % formic acid. Eluted peptides of HEK and E.coli were mixed in the following ratios (amount in &micro;g):</p> <p>Sample&nbsp;&nbsp; &nbsp;HEK&nbsp;&nbsp; &nbsp;E.coli&nbsp;&nbsp; &nbsp;MS method<br> Sample1&nbsp;&nbsp; &nbsp;2.5&nbsp; &nbsp; &nbsp; 0.00&nbsp; &nbsp; &nbsp; &nbsp; DDA<br> Sample2&nbsp;&nbsp; &nbsp;2.5&nbsp; &nbsp; &nbsp; 0.05&nbsp; &nbsp; &nbsp; &nbsp; DDA<br> Sample3&nbsp;&nbsp; &nbsp;2.5&nbsp; &nbsp; &nbsp; 0.15&nbsp; &nbsp; &nbsp; &nbsp; DDA<br> Sample4&nbsp;&nbsp; &nbsp;2.5&nbsp; &nbsp; &nbsp; 0.40&nbsp; &nbsp; &nbsp; &nbsp; DDA<br> Sample5&nbsp;&nbsp; &nbsp;2.5&nbsp; &nbsp; &nbsp; 0.80&nbsp; &nbsp; &nbsp; &nbsp; DDA</p> <p>Additionally, iRT peptides were added and 1&micro;g of each samples&nbsp;was measured with a Q-Exactive Plus mass spectrometer. Besides the five&nbsp;raw files, we uploaded two&nbsp;fasta files that serve&nbsp;as human and ecoli protein sequence databases, an transition list for the iRT peptides as well as an experimental design for the MaxQuant search.<br> Additionally, we uploaded&nbsp;the Galaxy MaxQuant training result files: protein groups, peptides, mqpar, msms, evidence&nbsp;and PTXQC.</p>

opencc-by-4.0Nov 2020View details →
zenodo40/100

A set of generated Instagram Data Download Packages (DDPs) to investigate their structure and content

<p><strong>Instagram data-download example dataset</strong></p> <p>In this repository you can find a data-set consisting of 11 personal Instagram archives, or Data-Download Packages (DDPs).</p> <p>&nbsp;</p> <p><strong>How the data was generated</strong></p> <p>These Instagram accounts were all new and generated by a group of researchers who were interested to figure out in detail<br> the structure and variety in structure of these Instagram DDPs. The participants user the Instagram account extensively for approximately a week. The participants also intensively communicated with each other so that the data can be used as an example of a network.&nbsp;</p> <p>The data was primarily generated to evaluate the performance of de-identification software. Therefore, the text in the DDPs particularly contain many randomly chosen (Dutch) first names, phone numbers, e-mail addresses and URLS. In addition, the images in the DDPs contain many faces and text as well. The DDPs contain faces and text (usernames) of third parties. However, only content of so-called `professional accounts&#39; are shared, such as accounts of famous individuals or institutions who self-consciously and actively seek publicity, and these sources are easily publicly available. Furthermore, the DDPs do not contain sensitive personal data of these individuals.&nbsp;</p> <p><br> <strong>Obtaining your Instagram DDP</strong></p> <p>After using the Instagram accounts intensively for approximately a week, the participants requested their personal Instagram DDPs by using the following steps. You can follow these steps yourself if you are interested in your personal Instagram DDP.&nbsp;</p> <p>1. Go to www.instagram.com and log in<br> 2. Click on your profile picture, go to *Settings* and *Privacy and Security*<br> 3. Scroll to *Data download* and click *Request download*<br> 4. Enter your email adress and click *Next*<br> 5. Enter your password and click *Request download*</p> <p>Instagram then delivered the data in a compressed zip folder with the format **username_YYYYMMDD.zip** (i.e., Instagram handle and date of download) to the participant, and the participants shared these DDPs with us.</p> <p>&nbsp;</p> <p><strong>Data cleaning</strong></p> <p>To comply with the Instagram user agreement, participants shared their full name, phone number and e-mail address. In addition, Instagram logged the i.p. addresses the participant used during their active period on Instagram. After colleting the DDPs, we manually replaced such information with random replacements such that the DDps shared here do not contain any personal data of the participants.</p> <p>&nbsp;</p> <p><strong>How this data-set can be used</strong></p> <p>This data-set was generated with the intention to evaluate the performance of the de-identification software. We invite other researchers to use this data-set for example to investigate what type of data can be found in Instagram DDPs or to investigate the structure of Instagram DDPs. The packages can also be used for example data-analyses, although no substantive research questions can be answered using this data as the data does not reflect how research subjects behave `in the wild&#39;.&nbsp;</p> <p><br> <strong>Authors</strong></p> <p>The data collection is executed by Laura Boeschoten, Ruben van den Goorbergh and Daniel Oberski of Utrecht University. For questions, please contact l.boeschoten@uu.nl.&nbsp;</p> <p>&nbsp;</p> <p><strong>Acknowledgments</strong></p> <p>The researchers would like to thank everyone who participated in this data-generation project.</p>

opencc-by-4.0Jan 2021View details →
zenodo40/100

Raw NGS data for the study 'Spouse-to-spouse Transmission and Evolution of Hypervariable Region 1 and 5’ Untraslated Region of Hepatitis C Virus Analyzed by Next-generation Sequencing'

<p>This file contains  the original next-generation sequencing data (raw sequences in fastq format) which were analyzed in the study titled: "Spouse-to-spouse Transmission and Evolution of Hypervariable Region 1  and 5’ Untraslated Region of Hepatitis C Virus Analyzed by Next-generation Sequencing".</p> <p> </p> <p> </p>

opencc-zeroJan 2016View details →
zenodo40/100

Graphing and tabulating next-generation sequencing and genotyping data

<p>Making figures and tables for publication. Each zip archive contains input data, shell script to initiate and log R script, one R script for generating several graphs and tables, and the output graphs and tables themselves.</p> <p>Data was generated by whole-genome resequencing of 22 individual D.melanogaster from Sussex-LHM population and 2 from the Sussex RG line, followed by read-mapping, then genotyping with Haplotype Caller and Genomestrip.</p> <p>Locations for raw data, code, logs, extended QC data:</p> <p>Sequence reads NCBI SRA268956</p> <p>NCBI dbSNP https://www.ncbi.nlm.nih.gov/projects/SNP/snp_viewBatch.cgi?sbid=1062461</p> <p>NCBI dbVar accession number pre-release nstd134</p> <p> </p> <p>The pre-print manuscript for this data is available on biorxiv: "Whole genome resequencing of a laboratory-adapted Drosophila melanogaster population sample" http://biorxiv.org/content/early/2016/10/17/081554 doi: http://dx.doi.org/10.1101/081554</p> <p> </p>

opencc-by-4.0Oct 2016View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record