Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,037
datasets available to search
ShareScore release 0.9.0
Dataset results
1,037 results for “large-scale”
The Secret Life of Software Vulnerabilities: A Large-Scale Empirical Study
<p>Online appendix of the paper entitled: "The Secret Life of Software Vulnerabilities: A Large-Scale Empirical Study". It contains all scripts and data required to replicate the four research questions of the study.</p> <p>Abstract: Software vulnerabilities are weaknesses in source code that can be potentially exploited to cause loss or harm. While researchers have been devising a number of methods to deal with vulnerabilities, there is still a noticeable lack of knowledge on their software engineering life cycle, for example how vulnerabilities are introduced and removed by developers. This information can be exploited to design more effective methods for vulnerability prevention and detection, as well as to understand the granularity that these methods should aim at. To investigate the life cycle of software vulnerabilities, we focus on how, when, and under which circumstances vulnerabilities are introduced in software projects, as well as whether, after how long, and how they are removed. We consider 4,097 vulnerabilities with public patches from the National Vulnerability Database—pertaining to 1,163 open-source software projects on GITHUB—and define a six-step process that involves both automated parts (e.g., using the SZZ algorithm to find the vulnerability-inducing commits) and manual analyses (e.g., how vulnerabilities were fixed). The investigated vulnerabilities can be classified in 148 categories, take on average 4.19 commits before being introduced, and remain unfixed for a median of 1,506.50 commits and 691.50 days. Most of them are introduced by developers with high workload, often when doing maintenance activities, and removed with mostly with the addition of new source code aiming at implementing further checks on inputs. We conclude by distilling practical implications on when and how vulnerability detectors should work to better assist developers in early detecting these issues.</p>
Vector sequences in early WIV SRA sequencing data of SARS-CoV-2 inform on a potential large-scale security breach at the beginning of the COVID-19 pandemic
<p>DESCRIPTION</p> <p>Sequences identified as Influenza A virus, Spodoptera frugiperda rhabdovirus and Nipah henipavirus have been previously identified within the early HiSeq 1000 and HiSeq 3000 sequencing data of SARS-CoV-2, SRR11092059,SRR11092060,SRR11092061 and SRR11092062, and were being used to support the hypothesis that a "simultaneous outbreak of multiple zoonotic viruses" have happened in the Huanan Seafood market. https://doi.org/10.31219/osf.io/s4td6</p> <p>However, a closer examination of these sequences revealed that they were not sequences of actual wild viruses, but were in stead fragments left behind from PCR products and cloning vectors harboring both cDNA clones and infectious clones of such viruses, with evidence of viral sequences being joined directly to DNA sequences of vector and non-human origin within the same short reads.</p> <p>Here are the vector sequences and PCR product-like sequences recovered from the earliest WIV SRA sequencing data of Human SARS-CoV-2 from dataset SRR11092059,SRR11092060,SRR11092061,SRR11092062.</p> <p>Sequences associated with Vectors and PCR products from 3 distinct viral species have been obtained: The 3'-end of a Nipah Henipahvirus with fusion to a Hepatitis D virus Ribozyme, a T7 terminator and a Tetracycline resistance gene, The 5'-end of the same Nipah Henipahvirus with fusion to sequences found in diverse vectors, A complete vector genome encoding the HA gene of Influenza A virus subtype H7N9 under a CMV promoter and a bgH polyA terminator, and 221 Contiguous sequences corresponding to the Spodoptera frugiperda rhabdovirus reference genome fused to sequences that were homologous to multiple Plastid sequences and Notably Mitochondrial sequences of Rodents.</p> <p>As sequences corresponding to a rescued infectious clone of a BSL-4 organism (Nipah Henipahvirus) were found in sample sequences that supposedy represents patient samples that were obtained from Hospital ICU and sequenced in a pathogen diagnosis laboratory (which is separate from the Virology Research laboratory which is implied by the context of an Infectious Clone of such an organism, evident by the 3'-HDV ribozyme and T7 terminator fused directly to the 3'-terminus of the Nipah Henipahvirus reads), The discovery of artifact-containing sequences of at least 3 different pathogen species that are phylogenetically and methodologically distinct from each other in samples that were supposedly submitted by a laboratory that is Separate from the virological research laboratories that could have hosted such clone sequences imply extensive crosstalk and cross-contamination between the various laboratories within the Wuhan Institute of Virology, which includes at least one BSL-4 laboratory with evidence of containment breach of a BSL-4 organism and it's subsequent introduction into RNA-seq samples that were processed by a laboratory of distinct and separate purposes than the basic virological research evidenced by the Infectious Clone of the Hipah Henipahvirus.</p> <p>Such a discovery therefore likely imply a major security breach happening within the Wuhan institute of Virology at the time when the first sequences of SARS-CoV-2 was sampled and sequenced, which have important implications on the origins of the SARS-CoV-2 virus itself.</p> <p>METHODS</p> <p>The metagenomic sequencing datasets, SRR11092059,SRR11092060,SRR11092061 and SRR11092062 were first analyzed using the NCBI phylogenetic analysis tool, which identified viral sequences that is not related to SARS-CoV-2 itself. These include Influenza A virus (IAV, subtype H7N9), Spodoptera frugiperda rhabdovirus and Nipah Henipahvirus.</p> <p>The datasets were then subjected to BLAST search using MEGABLAST against the reference sequences of such viruses to verify the existence of the viral sequences and determine the exact sybtype of such viruses and the closest sequences on GenBank that corresponds to the reads. There seuqences are MH926031.1 for the Spodoptera frugiperda rhabdovirus, KY199425.1 for the Influenza A virus and AY988601.1 for the Nipah Henipahvirus.</p> <p>A second round BLAST analysis with these identified sequences were then performed, which unexpectedly revealed numerous reads corresponding to Cloning vectors and non-human Mitochondrial and Plastid sequences being fused directly to the sequences of the identified viral species. Reads were then downloaded and subjected to assembly using the CAP3 sequence assembly program and the EGASSEMBLER tool. Contig sequences were then queried against the NCBI nr/nt database which unanimously identified the original sample sequences as viral sequences inserted into cloning vectors.</p> <p>The complete sequence of the Influenza A virus Haemagluttinin (HA) gene clone was obtained from SRR11092061,SRR11092062 using multiple rounds of BLAST search and sequence assembly expansion on the existing vector-virus junction contigs, and a partial sequence corresponding the 3'-end of Nipah Henipahvirus AY988601.1 fused to a 3'-HDV ribozyme, T7 terminator and a Tet resistance gene was obtained from SRR11092059. In addition, 221 Contig sequences corresponding to the Rhabdovirus MH926031.1 fused to Chloroplast sequence MN524635.1 and Rodent Mitochondrial sequence MT241668.1 have been recovered from SRR11092061.</p> <p>We then performed a BLAST search using the identified vector sequences on SRR11092059,SRR11092060,SRR11092061 and SRR11092062, which confirms the existence of these two vetor sequences in all 4 datasets.</p>
Downscaled surface mass balance in Antarctica: impacts of subsurface processes and large-scale atmospheric circulation
<p>Here is the surface mass balance calculated from a offline subsurface model, that is used in the paper Downscaled surface mass balance in Antarctica: impacts of subsurface processes and large-scale atmospheric circulation.<br> More data are available by contacting nichsen@space.dtu.dk</p>
PDMX: A Large-Scale Public Domain MusicXML Dataset for Symbolic Music Processing
<p>We introduce <strong>PDMX</strong>: a <strong>P</strong>ublic <strong>D</strong>omain <strong>M</strong>usic<strong>X</strong>ML dataset for symbolic music processing. Refer to our <a title="PDMX Paper" href="https://arxiv.org/abs/2409.10831" target="_blank" rel="noopener">paper</a> for more information, and our <a title="PDMX GitHub Repository" href="https://github.com/pnlong/PDMX/" target="_blank" rel="noopener">GitHub repository</a> for any code-related details. Please cite both our paper and <a href="https://arxiv.org/abs/2410.02084" target="_blank" rel="noopener">our collaborators' paper</a> if you use this dataset (see our GitHub for more information).</p> <p>Upon further use of the PDMX dataset, we discovered a discrepancy between the public-facing copyright metadata on the <a href="https://musescore.com/">MuseScore website</a> and the internal copyright data of the MuseScore files themselves, which affected 31,221 (12.29% of) songs. We have decided to proceed with the former given its public visibility on Musescore (i.e. this is what the MuseScore website presents its users with). We have noted files with conflicting internal licenses in the <em><strong>license_conflict</strong></em> column of PDMX. We recommend using the <em><strong>no_license_conflict</strong></em> subset of PDMX (which still includes 222,856 songs) moving forward.</p> <p>Additionally, for each song in PDMX, we not only provide the <em>MusicRender</em> and metadata JSON files, but we also try to include the associated compressed MusicXML (MXL), sheet music (PDF), and MIDI (MID) files when available. Due to the corruption of 42 of the original MuseScore files, these songs lack those associated files (since they could not be converted to those formats) and only include the <em>MusicRender</em> and metadata JSON files. The <em><strong>all_valid</strong></em> subset of PDMX describes the songs where all associated files are valid.</p>
Supplemental dataset to "Cloud sync in response to wave-like large-scale forcings"
<p>This dataset deposits the supplemental materials for the manuscript "Cloud sync in response to wave-like large-scale forcings".</p> <p><a href="https://zenodo.org/api/records/14164798/draft/files/math_note.pdf/content" target="_blank" rel="noopener noreferrer">math_note.pdf</a> A hand-written math derivation note for equations in the appendices. </p> <p><a href="https://zenodo.org/uploads/15304862" target="_blank" rel="noopener noreferrer">movie_wL_006_T24hours.avi</a> A movie of near-surface (z=25m) water vapor mixing ratio for the wL=0.006m/s and T=24 hours experiment (the reference CM1 simulation).</p> <p><a href="https://zenodo.org/api/records/15304862/draft/files/movie_wL_006_T18hours.avi/content" target="_blank" rel="noopener noreferrer">movie_wL_006_T18hours.avi</a> A movie of near-surface (z=25m) water vapor mixing ratio for the wL=0.006m/s and T=18 hours experiment.</p> <p><a href="https://zenodo.org/api/records/15304862/draft/files/movie_wL_006_T12hours.avi/content" target="_blank" rel="noopener noreferrer">movie_wL_006_T12hours.avi</a> A movie of near-surface (z=25m) water vapor mixing ratio for the wL=0.006m/s and T=12 hours experiment.</p> <p><a href="https://zenodo.org/api/records/15304862/draft/files/movie_wL_000.avi/content" target="_blank" rel="noopener noreferrer">movie_wL_000.avi</a> A movie of near-surface (z=25m) water vapor mixing ratio, without large-scale wave-like forcing.</p> <p><a href="https://zenodo.org/api/records/15304862/draft/files/input_sounding/content" target="_blank" rel="noopener noreferrer">input_sounding</a> The initial sounding for all CM1 simulations. </p> <p><a href="https://zenodo.org/api/records/15304862/draft/files/namelist.input/content" target="_blank" rel="noopener noreferrer">namelist.input</a> The namelist file for launching all CM1 simulations. </p> <p><a href="https://zenodo.org/api/records/15304862/draft/files/postprocessing_CM1.zip/content" target="_blank" rel="noopener noreferrer">postprocessing_CM1.zip</a> The postprocessing code of the CM1 simulations, including intermediate output files (.mat) in data postprocessing. </p> <p><a href="https://zenodo.org/api/records/14164798/draft/files/microscopic_model.zip/content" target="_blank" rel="noopener noreferrer">microscopic_model.zip</a> The MATLAB code for the microscopic model. </p> <p><a href="https://zenodo.org/api/records/15304862/draft/files/cm1.F/content" target="_blank" rel="noopener noreferrer">cm1.F</a> The CM1 script where the large-scale vertical velocity is programmed. You can copy it directly to your CM1/src/ path. </p> <p> </p> <p>Feel free to send an email to Dr. Hao Fu (haofu@nju.edu.cn) if you have any questions!</p> <p> </p>
StemGMD: A Large-Scale Audio Dataset of Isolated Drum Stems for Deep Drums Demixing - part 1
<p>We introduce StemGMD, a new large-scale dataset of isolated drum stems that builds upon the extensive MIDI collection found in <a href="https://magenta.tensorflow.org/datasets/groove">Magenta's Groove MIDI Dataset (GMD)</a>.</p> <p>GMD is a 13.6-hour corpus of expressive drum performances executed by ten drummers on a Roland TD-11 electronic drum kit. It contains 1150 MIDI files along with the corresponding full-kit audio mixtures.</p> <p>As a first step in creating StemGMD, we mapped the 22 different MIDI pitches found in the original files onto nine canonical instruments through the reduction scheme proposed in J. Gillick, A. Roberts, J. Engel, D. Eck, and D. Bamman, "Learning to groove with inverse sequence transformations," in International Conference on Machine Learning (ICML), vol. 97, 2019, pp. 2269–2279.</p> <p>Each of the nine resulting MIDI channels was manually synthesized as a 16-bit/44.1 kHz stereo WAV file using ten realistic-sounding acoustic drum kits sourced from the <a href="https://support.apple.com/en-me/guide/logicpro/lgsi2fb2509e/mac">Logic Pro X sample libraries</a>, i.e., Bluebird, Brooklyn, Detroit Garage, East Bay, Heavy, Motown Revisited, Portland, Retro Rock, Roots, and SoCal.</p> <p>As a result, StemGMD contains 1224 hours of audio, which correspond to more than 136 hours of full-kit mixtures. Moreover, StemGMD also contains single hits for each of the drum pieces at ten different velocities ranging from 30 to 127.</p> <p>To the best of our knowledge, StemGMD is the largest publicly available dataset of drums to date. Moreover, it is the first collection of single-instrument clips from all nine pieces in a canonical drum kit, making it well-suited for training deep drums demixing models.</p> <p> </p> <p><strong>*** THIS IS PART 1 OF 2 ***</strong></p> <p><strong>Download part 2 here:</strong> <a href="../records/7882857">https://zenodo.org/records/7882857 </a> (now available!) </p> <p>After downloading both parts, run <strong>unzip_StemGMD.sh</strong> to build the dataset from the split archive files.<br>Once unzipped, StemGMD will take just over <strong>1.13 TB </strong>of memory.</p> <p> </p> <p>The dataset is made available under a <a href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International (CC BY 4.0) License</a>.</p> <p>____________________________</p> <p>We employed the dataset in our paper titled "Toward Deep Drum Source Separation," published in Pattern Recognition Letters. </p> <div> <div> <div> <div> <p>Please, cite this work as: A. I. Mezza, R. Giampiccolo, A. Bernardini, and A. Sarti, "Toward Deep Drum Source Separation," Pattern Recognition Letters, vol. 183, pp. 86-91, 2024, doi: 10.1016/j.patrec.2024.04.026.</p> <pre>@article{mezza2024, title = {Toward deep drum source separation}, author = {Alessandro Ilic Mezza and Riccardo Giampiccolo and Alberto Bernardini and Augusto Sarti}, journal = {Pattern Recognition Letters}, volume = {183}, pages = {86-91}, year = {2024}, issn = {0167-8655}, doi = {https://doi.org/10.1016/j.patrec.2024.04.026} }</pre> </div> </div> </div> </div>
StemGMD: A Large-Scale Audio Dataset of Isolated Drum Stems for Deep Drums Demixing - part 2
<p>We introduce StemGMD, a new large-scale dataset of isolated drum stems that builds upon the extensive MIDI collection found in <a href="https://magenta.tensorflow.org/datasets/groove">Magenta's Groove MIDI Dataset (GMD)</a>.</p> <p>GMD is a 13.6-hour corpus of expressive drum performances executed by ten drummers on a Roland TD-11 electronic drum kit. It contains 1150 MIDI files along with the corresponding full-kit audio mixtures.</p> <p>As a first step in creating StemGMD, we mapped the 22 different MIDI pitches found in the original files onto nine canonical instruments through the reduction scheme proposed in J. Gillick, A. Roberts, J. Engel, D. Eck, and D. Bamman, "Learning to groove with inverse sequence transformations," in International Conference on Machine Learning (ICML), vol. 97, 2019, pp. 2269–2279.</p> <p>Each of the nine resulting MIDI channels was manually synthesized as a 16-bit/44.1 kHz stereo WAV file using ten realistic-sounding acoustic drum kits sourced from the <a href="https://support.apple.com/en-me/guide/logicpro/lgsi2fb2509e/mac">Logic Pro X sample libraries</a>, i.e., Bluebird, Brooklyn, Detroit Garage, East Bay, Heavy, Motown Revisited, Portland, Retro Rock, Roots, and SoCal.</p> <p>As a result, StemGMD contains 1224 hours of audio, which correspond to more than 136 hours of full-kit mixtures. Moreover, StemGMD also contains single hits for each of the drum pieces at ten different velocities ranging from 30 to 127.</p> <p>To the best of our knowledge, StemGMD is the largest publicly available dataset of drums to date. Moreover, it is the first collection of single-instrument clips from all nine pieces in a canonical drum kit, making it well-suited for training deep drums demixing models.</p> <p> </p> <p><strong>*** THIS IS PART 2 OF 2 ***</strong></p> <p><strong>Download part 1 here:</strong> <a href="../records/7860223">https://zenodo.org/records/7860223 </a></p> <p>After downloading both parts, run <strong>unzip_StemGMD.sh</strong> to build the dataset from the split archive files.<br>Once unzipped, StemGMD will take just over <strong>1.13 TB </strong>of memory.</p> <p> </p> <p>The dataset is made available under a <a href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International (CC BY 4.0) License</a>.<br><br>_____________________________<br><br>We employed the dataset in our paper titled "Toward Deep Drum Source Separation," published in Pattern Recognition Letters. </p> <div> <div> <div> <div> <p>Please, cite this work as: A. I. Mezza, R. Giampiccolo, A. Bernardini, and A. Sarti, "Toward Deep Drum Source Separation," Pattern Recognition Letters, vol. 183, pp. 86-91, 2024, doi: 10.1016/j.patrec.2024.04.026.</p> <pre>@article{mezza2024, title = {Toward deep drum source separation}, author = {Alessandro Ilic Mezza and Riccardo Giampiccolo and Alberto Bernardini and Augusto Sarti}, journal = {Pattern Recognition Letters}, volume = {183}, pages = {86-91}, year = {2024}, issn = {0167-8655}, doi = {https://doi.org/10.1016/j.patrec.2024.04.026} }</pre> </div> </div> </div> </div>
EUV-Induced Hydrogen Desorption As A Step Towards Large-Scale Silicon Quantum Device Patterning
<p><strong>Dataset: </strong>STM, XPS and PEEM raw data, processed data and the codes used for data fitting our <a href="https://doi.org/10.1038/s41467-024-44790-6">published work</a> are all available here.</p> <p><strong>Abstract: </strong>Atomically precise hydrogen desorption lithography using scanning tunnelling microscopy (STM) has enabled the development of single-atom, quantum-electronic devices on a laboratory scale. Scaling up this technology to mass-produce these devices requires bridging the gap between the precision of STM and the processes used in next-generation semiconductor manufacturing. Here, we demonstrate the ability to remove hydrogen from a monohydride Si(001):H surface using extreme ultraviolet (EUV) light. We quantify the desorption characteristics using various techniques, including STM, X-ray photoelectron spectroscopy (XPS), and photoemission electron microscopy (XPEEM). Our results show that desorption is induced by secondary electrons from valence band excitations, consistent with an exactly solvable non-linear differential equation and compatible with the current 13.5 nm (~92 eV) EUV standard for photolithography; the data imply useful exposure times of order minutes for the 300 W sources characteristic of EUV infrastructure. This is an important step towards the EUV patterning of silicon surfaces without traditional resists, by offering the possibility for parallel processing in the fabrication of classical and quantum devices through deterministic doping.</p> <p> </p>
Large-scale annotated dataset for cochlear hair cell detection and classification
<p>Our sense of hearing is mediated by cochlear hair cells, of which there are two types organized in one row of inner hair cells and three rows of outer hair cells. Each cochlea contains 5 - 15 thousand terminally differentiated hair cells, and their survival is essential for hearing as they do not regenerate after insult. It is often desirable in hearing research to quantify the number of hair cells within cochlear samples, in both pathological conditions, and in response to treatment. Machine learning can be used to automate the quantification process but requires a vast and diverse dataset for effective training. In this study, we present a large collection of annotated cochlear hair-cell datasets, labeled with commonly used hair-cell markers and imaged using various fluorescence microscopy techniques. The collection includes samples from mouse, rat, guinea pig, pig, primate, and human cochlear tissue, from normal conditions and following <i>in-vivo</i> and <i>in-vitro</i>ototoxic drug application. The dataset includes over 107,000 hair cells which have been manually identified and annotated as either inner or outer hair cells. This dataset is the result of a collaborative effort from multiple laboratories and has been carefully curated to represent a variety of imaging techniques. With suggested usage parameters and a well-described annotation procedure, this collection can facilitate the development of generalizable cochlear hair-cell detection models or serve as a starting point for fine-tuning models for other analysis tasks. By providing this dataset, we aim to give other hearing research groups the opportunity to develop their own tools with which to analyze cochlear imaging data more fully, accurately, and with greater ease. </p><p>Associated code is provided here: https://github.com/indzhykulianlab/hcat-data</p>
Data for: Characterization of large-scale preferential flow across continental United States
<p>Understanding preferential flow (PF) at large scales is critical for improving land management and groundwater (GW) quality. However, limited knowledge of this process, due to soil surface heterogeneity and observational constraints, hampers progress. In this study, we propose estimating effective PF at remote sensing footprint scale (4 – 9 km) by examining its impact on soil moisture (SM) distribution and shallow GW (SGW) table fluctuations (depth 5 m). Effective PF encompasses macropore, funnel, and finger flow pathways influencing SGW table fluctuations. We compiled daily SGW observations (2019-2021) from 19 continental US (CONUS) sites through USGS. Using inverse modeling in HYDRUS-1D, SGW data, and CHIRPS precipitation data, we inversely estimated soil hydraulic parameters of the dual porosity model (DPM) simulating vertical flow from soil surface to subsurface. Effective PF presence was inferred using three criteria: (1) daily precipitation >= the site-specific average across multiple (calibration) years, (2) daily observed SGW table increase, and (3) daily difference between observed and DPM simulated SGW tables 50% of the site-specific RMSE. Leveraging optimized DPM parameters and associated soil texture, classified PF events, and Soil Moisture Active Passive (SMAP L3E) satellite-based SM, a Random Forest algorithm with 10-fold cross validation predicted large-scale effective PF events. Results indicate seasonal dependence, with spring having the highest occurrence of PF events. The Random Forest model achieved 98% accuracy in predicting large-scale PF events, with SMAP SM and saturated hydraulic conductivity (Ks) among the 4 most impactful variables. Our approach provides a soil hydraulic property, site characteristic, soil texture and remote sensing based generalized tool to analyze large-scale effective PF.</p>
CESM2 land output data for study on hydrological impacts of large-scale forest expansion
<p>This repository contains the land output data from CESM2 which was generated in the study investigating the hydrological impacts of global-scale forestation. The datasets cover the period 2015-2100. The files are labelled according to the experiments they were generated from (base, MF (Max Forest) and No LULCC). Output fields are as follows:</p> <p>discharge_plus_runoff: surface water availability (river discharge plus surface runoff), units m^-3 s^-1</p> <p>EFLX_LH_TOT: total latent heat flux from land to atmosphere, units W m^-2</p> <p>QFLX_EVAP_TOT: total evapotranspiration (canopy evaporation plus canopy transpiration plus soil evaporation), units kg m^-2 s^-1</p> <p>SOILLIQ: soil liquid water content, units kg m^-2</p> <p>SW_surface_albedo: surface albedo, units fraction</p> <p>TSA: 2m air temperature, units K</p> <p>VEGWP: vegetation water potential, units m</p> <p> </p> <p>All data were generated and processed by James A. King.</p>
CESM2 atmosphere output data for study on hydrological impacts of large-scale forest expansion
<p>This repository contains the atmosphere output data from CESM2 which was generated in the study investigating the hydrological impacts of global-scale forestation, for the Max Forest scenario. The datasets cover the period 2015-2100. Output fields are as follows:</p> <p>CCN3: concentration of cloud condensation nuclei at 0.1% supersaturation, units cm^-3</p> <p>CLDLOW: cloud fraction integrated between 1200-700 hPa, units fraction of grid cell</p> <p>CONCLD: convective cloud cover, units fraction of grid cell</p> <p>GCLDLWP: grid cell cloud water path, units kg m^-2</p> <p>LWCF_d1: clean longwave cloud forcing, units W m^-2</p> <p>OMEGA: vertical velocity, units Pa s^-1</p> <p>PRECT: total precipitation, units m s^-1</p> <p>SWCF_d1: clean shortwave cloud forcing, units W m^-2</p> <p>V: meridional wind, units m s^-1</p> <p> </p> <p>All data were generated and processed by James A. King.</p>
CESM2 atmosphere output data for study on hydrological impacts of large-scale forest expansion
<p>This repository contains the atmosphere output data from CESM2 which was generated in the study investigating the hydrological impacts of global-scale forestation, for the base scenario. The datasets cover the period 2015-2100. Output fields are as follows:</p> <p>CCN3: concentration of cloud condensation nuclei at 0.1% supersaturation, units cm^-3</p> <p>CLDLOW: cloud fraction integrated between 1200-700 hPa, units fraction of grid cell</p> <p>CONCLD: convective cloud cover, units fraction of grid cell</p> <p>GCLDLWP: grid cell cloud water path, units kg m^-2</p> <p>LWCF_d1: clean longwave cloud forcing, units W m^-2</p> <p>OMEGA: vertical velocity, units Pa s^-1</p> <p>PRECT: total precipitation, units m s^-1</p> <p>SWCF_d1: clean shortwave cloud forcing, units W m^-2</p> <p>V: meridional wind, units m s^-1</p> <p> </p> <p>All data were generated and processed by James A. King.</p>
CESM2 land use data for study on hydrological impacts of large-scale forest expansion
<p>This repository contains the land use/ land cover input data for CESM2 which was used in the study investigating the hydrological impacts of global-scale forestation. The datasets cover the period 2000-2100. The files correspond to experiments described in the study as follows:</p> <p> </p> <p>Base: landuse.timeseries_0.9x1.25_SSP1-2.6_78pfts_CMIP6_simyr2000-2100_c220715.nc</p> <p>Max Forest: landuse.timeseries_0.9x1.25_hist_78pfts_SSPRFAFRS_SSP1_edit_xarray_4_simyr2000-2100_c221024.nc</p> <p>No LULCC: landuse.timeseries_0.9x1.25_hist_78pfts_SSPNOLULCC_3_simyr2000-2100_c221025.nc</p> <p> </p> <p>Files were created in collaboration by James A. King, James Weber, Peter Lawrence, and Stephanie Roe.</p> <p> </p>
CESM2 atmosphere output data for study on hydrological impacts of large-scale forest expansion
<p>This repository contains the atmosphere output data from CESM2 which was generated in the study investigating the hydrological impacts of global-scale forestation, for the No LULCC scenario. The datasets cover the period 2015-2100. Output fields are as follows:</p> <p>CCN3: concentration of cloud condensation nuclei at 0.1% supersaturation, units cm^-3</p> <p>CLDLOW: cloud fraction integrated between 1200-700 hPa, units fraction of grid cell</p> <p>CONCLD: convective cloud cover, units fraction of grid cell</p> <p>GCLDLWP: grid cell cloud water path, units kg m^-2</p> <p>LWCF_d1: clean longwave cloud forcing, units W m^-2</p> <p>OMEGA: vertical velocity, units Pa s^-1</p> <p>PRECT: total precipitation, units m s^-1</p> <p>SWCF_d1: clean shortwave cloud forcing, units W m^-2</p> <p>V: meridional wind, units m s^-1</p> <p> </p> <p>All data were generated and processed by James A. King.</p>
Increasing sustainability in palaeoproteomics by optimizing digestion times for large-scale archaeological bone analyses
<p>Palaeoproteomic analysis of skeletal proteomes is used to provide taxonomic identifications for an increasing number of archaeological specimens. The success rate depends on a range of taphonomic factors and differences in the extraction protocols employed. By analyzing 12 archaeological bone specimens from two archaeological sites, we demonstrate that reducing digestion duration from 18 to 3 hours has no measurable impact on the obtained taxonomic identifications. Peptide marker recovery, COL1 sequence coverage, or proteome complexity are also not significantly impacted. Although we observe minor differences in sequence coverage and glutamine deamidation, these are not consistent across our dataset. A 6-fold reduction in digestion time reduces electricity consumption, and therefore CO<sub>2</sub> emission intensities. We furthermore demonstrate that working in 96-well plates further reduces electricity consumption by 60%, in comparison to individual microtubes. Reducing digestion time therefore has no impact on the taxonomic identifications, while reducing the environmental impact of palaeoproteomic projects.</p>
Synthetic dataset from - Bar et al., Sifting through the haystack - efficiently finding rare behaviors in large-scale datasets, WACV 2025
<p>This is a synthetic dataset emulating pose estimation data, introduced in the associated paper. </p> <p>Briefly, each sample in the dataset is a sequence of 5-keypoints with 9 timesteps. Movement of each keypoint in time is determined by a sinus with some amplitude A and some frequency f, this is loosely inspired by larval zebrafish swimming movement. <br>There are two types of behaviors - a common behavior (aka forming the majority of the samples in the dataset) where the frequency of the sinus is larger than the amplitude, and a rare one where the amplitude is larger than the frequency. <br>We vary the similarity between the rare and common behaviors by relaxing the standard deviation of the gaussian from which we draw these movement parameters (behavior similarity, sd= [0.5, 1.5, 2.5, 5]). We also test different levels of data imbalance, varying the frequency of the rare behavior (rarity=[1.5%,5%,12%,24%]). </p> <p>Thus we created 16 datasets with all possible combinations.</p> <p>The data generation code will become available in our code repository: https://github.com/shir3bar/SiftingTheHaystack</p> <p>The data was used to create a controlled experimental sandbox in which we could test our pipeline for detecting rare behaviors. Sounds interesting? Read our paper and check out the code :)</p> <p> </p> <p> </p>
Digital repository for: Large-scale forest disturbance and associated management shape bird communities in Central European spruce forests
<p>Repository containing R-script and data to reproduce analysis and main figures on the effect of large-scale forest disturbance and associated pre- and post-disturbance management on bird communities in the Harz Mountains, Germany.</p> <p>R-script includes:</p> <ul> <li>indicator species analysis (R package indicspecies; Cáceres & Legendre, 2009)</li> <li>non-metric multidimensional scaling (R package vegan; Oksanen et al., 2016)</li> <li>rarefaction- and extrapolation of Hill numbers (R package iNEXT; Hsieh et al., 2019)</li> <li>multi-species community distance sampling (R package sp Abundance; Doser et al., 2023)</li> </ul> <p>Attached files:</p> <ul> <li><strong>bird_data_Graser_et_al.csv </strong>(row data of bird species point counts per distance category)</li> <li><strong>bird_data_abundance_100_Graser_et_al.csv </strong>(abundance of species per sampling site, summed up over all four sampling repeats only considering detected individuals up to 100 m around the sampling point)</li> <li><strong>siteCovs_Graser_et_al.csv</strong> (environmental variables for each sampling point)</li> <li><strong>A_species_matrix_100_new_Graser_et_al.csv</strong> (species-site matrix of <strong>bark-beetle disturbance, unlogged </strong>sites for rarefaction and extrapolation, species number summed up over all four sampling repeats only considering detected individuals up to 100 m around the sampling point)</li> <li><strong>B_species_matrix_100_new_Graser_et_al.csv </strong>(species-site matrix of <strong>windthrow disturbance, unlogged </strong>sites for rarefaction and extrapolation, species number summed up over all four sampling repeats only considering detected individuals up to 100 m around the sampling point)</li> <li><strong>C_species_matrix_100_new_Graser_et_al.csv </strong>(species-site matrix of <strong>bark-beetle/windthrow disturbance, underplanted, unlogged </strong>sites for rarefaction and extrapolation, species number summed up over all four sampling repeats only considering detected individuals up to 100 m around the sampling point)</li> <li><strong>D_species_matrix_100_new_Graser_et_al.csv </strong>(species-site matrix of <strong>bark-beetle /windthrow disturbance, salvage-unlogged </strong>sites for rarefaction and extrapolation, species number summed up over all four sampling repeats only considering detected individuals up to 100 m around the sampling point)</li> <li><strong>E_species_matrix_100_new_Graser_et_al.csv </strong>(species-site matrix of <strong>bark-beetle /windthrow disturbance, underplanted, salvage-unlogged </strong>sites for rarefaction and extrapolation, summed up over all four sampling repeats only considering detected individuals up to 100 m around the sampling point)</li> <li><strong> F_species_matrix_100_new_Graser_et_al.cs</strong>v (species-site matrix of <strong>mature spruce plantation </strong>sites for rarefaction and extrapolation, species number summed up over all four sampling repeats only considering detected individuals up to 100 m around the sampling point)</li> <li><strong>msHDS_bird_data_management_model_Graser_et_al.rds</strong> (R-data set for multi-species community distance sampling of the effect of different pre- and post-disturbance management groups)</li> <li><strong>msHDS_bird_data_stand_age_model_Graser_et_al.rds </strong>(R-data set for multi-species community distance sampling of the effect of post-disturbance forest succession)</li> </ul> <p>A more detailed description of the data can be found in the README.txt document.</p> <p><span>References:</span></p> <p><span>Cáceres, M. D., & Legendre, P. (2009). </span><span>Associations between species and groups of sites: Indices and statistical inference. <em>Ecology</em>, <em>90</em>(12), 3566–3574. https://doi.org/10.1890/08-1823.1</span></p> <p><span>Doser, J. W., Finley, A. O., Kéry, M., & Zipkin, E. F. (2023). spAbundance: An R package for single‐species and multi‐species spatially explicit abundance models. <em>Methods in Ecology and Evolution</em>, <em>15</em>(6), 1024–1033. https://doi.org/10.1111/2041-210X.14332</span></p> <p><span>Hsieh, T. C., Ma, K. H., & Chao, A. (2019). <em>iNEXT-package: Interpolation and extrapolation for species diversity</em>. https://cran.r-project.org/web/packages/iNEXT/vignettes/Introduction.html</span></p> <p><span>Oksanen, J., Blanchet, F. G., Kindt, R., Legendre, P., O’hara, R. B., Simpson, G. L., Solymos, P., Stevens, M. H. H., Wagner, H., Minchin, P. R., Gavin, L., & Henry, H. (2016). Vegan: Community ecology package. R package version 1.17-4. <em>Http://CRAN. R-Project. </em></span><em><span>Org/Package=vegan</span></em><span>.</span></p> <p></p> <p></p>
Effectively controlling for sample relatedness in large-scale GWAS: application to 79 EHR-derived longitudinal traits
<p><span>Sample relatedness is a major confounder in genome-wide association studies (GWAS), potentially leading to inflated type I error rates if not appropriately controlled. A common strategy is to incorporate a random effect related to genetic relatedness matrix (GRM) into regression models. However, this approach is challenging for large-scale GWAS of complex traits, such as longitudinal traits. Here we propose a scalable and accurate analysis framework, SPA<sub>GRM</sub>, which controls for sample relatedness via a precise approximation of the joint distribution of genotypes. SPA<sub>GRM</sub> can utilize GRM-free models and thus is applicable to various trait types and statistical methods, including linear mixed models and generalized estimation equations for longitudinal traits. A hybrid strategy incorporating saddlepoint approximation greatly increases the accuracy to analyze low-frequency and rare genetic variants, especially in unbalanced phenotypic distributions. We also introduce SPA<sub>GRM(CCT)</sub> to aggregate the results following different models via Cauchy combination test. Extensive simulations and real data analyses demonstrated that SPA<sub>GRM</sub> maintains well-controlled type I error rates and SPA<sub>GRM(CCT)</sub> can serve as a broadly effective method. Applying SPA<sub>GRM</sub> to 79 longitudinal traits extracted from</span><span> </span><span>UK Biobank primary care data, we identified 7,463 genetic loci, making a pioneering attempt to conduct GWAS for these traits as longitudinal traits.</span></p>
Contributions of Anomalous Large-Scale Circulations to the Absence of Tropical Cyclones over the Western North Pacific in July 2020
<p>The datasets are for the article 'Contributions of Anomalous Large-Scale Circulations to the Absence of Tropical Cyclones over the Western North Pacific in July 2020', including the WRF Initial conditions (horizontal wind at 850 hPa, geopotential height at 850 and 200 hPa) of sensitivity experiments (CTRL, W_WNPSH, W_SAH, W_TUTT, and S_SASM) in July 2020, and the formation records of the simulated TC formation in all experiments.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.