Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,037
datasets available to search
ShareScore release 0.9.0
Dataset results
1,037 results for “large-scale”
MultiSubs: A Large-scale Multimodal and Multilingual Dataset
<p>MultiSubs is a dataset of multilingual subtitles gathered from <a href="https://opus.nlpl.eu/OpenSubtitles.php">the OPUS OpenSubtitles dataset</a>, which in turn was sourced from <a href="http://www.opensubtitles.org/">opensubtitles.org</a>. We have supplemented some text fragments (visually salient nouns in this release) within the subtitles with web images, where the word sense of the fragment has been disambiguated using a cross-lingual approach. </p> <p>Please refer to our paper for a more detailed description of the dataset:</p> <p>Josiah Wang, Pranava Madhyastha, Josiel Figueiredo, Chiraag Lala, Lucia Specia (2021). <a href="https://arxiv.org/abs/2103.01910">MultiSubs: A Large-scale Multimodal and Multilingual Dataset</a>. CoRR, abs/2103.01910. Available at: <a href="https://arxiv.org/abs/2103.01910">https://arxiv.org/abs/2103.01910</a></p>
Single objective light-sheet acquired large-scale 3D dataset
<p>This dataset covers Fig. 5 of the following publication:</p> <p>Title: Tilt-invariant scanned oblique plane illumination microscopy for large-scale volumetric imaging<br> Authors: Manish Kumar and Yevgenia Kozorovitskiy <br> Optics Letters Vol. 44, Issue 7, pp. 1706-1709 (2019)<br> https://doi.org/10.1364/OL.44.001706</p> <p>Briefly: The sample imaged is a Thy1GFP expressing transgenic mice brain slice - fixed and coverslipped. No clearing was performed for these.</p> <p>See "readme.txt" for additional info and details about how to use "shearNscaleObliq" file.</p>
Building Large-Scale Gene-Disease Association Datasets for Biomedical Relation Extraction
<p>This repository contains the GDAb and GDAt datasets. GDAb and GDAt are large-scale, distantly supervised, and manually enhanced datasets for Gene-Disease Association (GDA) extraction. Each dataset consists of three text files, corresponding to train, validation, and test sets, plus an additional JSON file containing the mapping between relation names and IDs. Each record in train, validation, or test files corresponds to a single GDA extracted from a sentence. Records are represented as JSON objects with the following structure:</p> <ul> <li><strong>text:</strong> sentence from which the GDA was extracted.</li> <li><strong>relation:</strong> relation name associated to the given GDA.</li> <li><strong>h: </strong>JSON object representing the gene entity, composed of: <ul> <li><strong>id: </strong>UMLS CUI associated to the gene entity.</li> <li><strong>name:</strong> UMLS preferred name associated to the gene entity.</li> <li><strong>pos: </strong>list consisting of starting position and length of the gene mention within text.</li> </ul> </li> <li><strong>t:</strong> JSON object representing the disease entity, composed of: <ul> <li><strong>id: </strong>UMLS CUI associated to the disease entity.</li> <li><strong>name:</strong> UMLS preferred name associated to the disease entity.</li> <li><strong>pos:</strong> list consisting of starting position and length of the disease mention within text.</li> </ul> </li> </ul> <p>Both datasets contain over 2,500,000 sentences and 500,000 bags.<br> The zip file consists of two folders, GDAb and GDAt, containing the files corresponding to the two datasets, respectively.</p> <p> </p>
Human pancreatic islet microRNAs implicated in diabetes and related traits by large-scale genetic analysis
<p>Genetic studies have identified ≥240 loci associated with risk of type 2 diabetes (T2D), yet most of these loci lie in non-coding regions, masking the underlying molecular mechanisms. Recent studies investigating mRNA expression in human pancreatic islets have yielded important insights into the molecular drivers of normal islet function and T2D pathophysiology. However, similar studies investigating microRNA (miRNA) expression remain limited. Here, we present data from 63 individuals, the largest sequencing-based analysis of miRNA expression in human islets to date. We characterize the genetic regulation of miRNA expression by decomposing the expression of highly heritable miRNAs into <em>cis</em>- and <em>trans</em>-acting genetic components and mapping <em>cis</em>-acting loci associated with miRNA expression (miRNA-eQTLs). We find (i) 84 heritable miRNAs, primarily regulated by <em>trans</em>-acting genetic effects, and (ii) 5 miRNA-eQTLs. We also use several different strategies to identify T2D-associated miRNAs. First, we colocalize miRNA-eQTLs with genetic loci associated with T2D and multiple glycemic traits, identifying one miRNA, miR-1908, that shares genetic signals for blood glucose and glycated hemoglobin (HbA1c). Next, we intersect miRNA seed regions and predicted target sites with credible set SNPs associated with T2D and glycemic traits and find 32 miRNAs that may have altered binding and function due to disrupted seed regions. Finally, we perform differential expression analysis and identify 14 miRNAs associated with T2D status—including miR-187-3p, miR-21-5p, miR-668, and miR-199b-5p—and 4 miRNAs associated with a polygenic score for HbA1c levels—miR-216a, miR-25, miR-30a-3p, and miR-30a-5p.</p>
CrusTome: A transcriptome database resource for large-scale analyses across Crustacea
<p>CrusTome_v0.1.0 Prerelease<br> /ReadMe - this file<br> /crustome_aa_BLAST.tar.gz - CrusTome database of amino acid sequences in BLAST format<br> /crustome_aa_DIAMOND.tar.gz - CrusTome database of amino acid sequences in DIAMOND format<br> /crustome_mrna_BLAST.tar.gz - CrusTome database of mRNA sequences in BLAST format<br> /dict - Dictionary file to translate species IDs. For usage with sed/awk see link to Github site below<br> <br> * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * *<br> * <br> * Please note, most of the data files contained in this DOI are <br> * compressed into GZip files (.gz extension). <br> * Mac and Linux OS's can extract this file type natively. <br> * Windows OS requires software to extract the archive. 7-Zip <br> * (http://www.7-zip.org) is free and open source software that will <br> * allow windows PCs to open and decompress the archive. <br> * <br> * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * *<br> <br> <strong><em>Pérez-Moreno JL, Kozma MT, DeLeo DM, Bracken-Grissom HD, Durica DS, Mykles DL. 2023. CrusTome: A transcriptome database resource for large-scale analyses across Crustacea. G3: Genes, Genomes, Genetics.</em></strong></p> <p><strong>CrusTome: A transcriptome database resource for large-scale analyses across Crustacea</strong></p> <p>Transcriptomes from non-traditional model organisms often harbor a wealth of unexplored data. Examining these datasets can lead <br> to clarity and novel insights in traditional systems, as well as to discoveries across a multitude of fields. Despite significant <br> advances in DNA sequencing technologies and in their adoption, access to genomic and transcriptomic resources for non-traditional <br> model organisms remains limited. Crustaceans, for example, being amongst the most numerous, diverse, and widely distributed taxa on the planet, often serve as excellent systems to address ecological, evolutionary, and organismal questions. While they are <br> ubiquitously present across environments, and of economic and food security importance, they remain severely underrepresented in <br> publicly available sequence databases. Here, we present CrusTome, a multi-species, multi-tissue, transcriptome database of 201 <br> assembled mRNA transcriptomes (189 crustaceans, 30 of which were previously unpublished, and 12 ecdysozoan outgroups) as an evolving, and publicly available resource. This database is suitable for evolutionary, ecological, and functional studies that employ<br> genomic/transcriptomic techniques and datasets. CrusTome is presented in BLAST and DIAMOND formats, providing robust datasets for sequence similarity searches, orthology assignments, phylogenetic inference, etc., and thus allowing for straight-forward incorporation into existing custom pipelines for high-throughput analyses.</p> <p> <br> * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * *<br> <br> For questions regarding released datasets contact:<br> Corresponding Author: Jorge L. Perez-Moreno (Colorado State University)<br> jorgepm@colostate.edu / jpere645@fiu.edu<br> <br> <strong> https://github.com/invertome/crustome</strong></p> <p> </p> <p> <br> <strong>PLEASE CITE:</strong></p> <p>Pérez-Moreno JL, Kozma MT, DeLeo DM, Bracken-Grissom HD, Durica DS, Mykles DL. 2023. CrusTome: A transcriptome database resource for large-scale analyses across Crustacea. G3: Genes, Genomes, Genetics.</p> <p> </p> <p><strong>Funder Information</strong></p> <p>Supported by National Science Foundation grants to DLM (IOS-1922701) and DSD (IOS-1922755). In addition, this work was partially funded by two grants awarded from the National Science Foundation: Doctoral Dissertation Improvement Grant (#1701835) awarded to JPM and HBG and the Division of Environmental Biology Bioluminescence and Vision grant (DEB-1556059) awarded to HBG. Samples in the FICC were collected by grants from The Gulf of Mexico Research Initiative (GOMRI), Florida Institute of Oceanography Shiptime Funding awarded to HBG and DMD; the National Science Foundation Division of Environmental Biology Grant 1556059 awarded to HBG; and the National Oceanic and Atmospheric Administration Ocean Exploration Research (NOAA-OER 2015) grant awarded to HBG.</p>
[Dataset] One year of high-precision operational data including measurement uncertainties from a large-scale solar thermal collector array with flat plate collectors, located in Graz, Austria
<p><strong>Highlights:</strong></p> <ul> <li>High-precision measurement data acquired within a scientific research project, using high-quality measurement equipment and implementing extensive data quality assurance measures.</li> <li>The dataset includes data from one full operational year in a 1-minute sampling rate, covering all seasons.</li> <li>Measured data channels include global, beam and diffuse irradiances in horizontal and collector plane. Heat transfer fluid properties were determined in a dedicated laboratory test.</li> <li>In addition to the measured data channels, calculated data channels, such as thermal power output, mass flow, fluid properties, solar incidence angle and shadowing masks are provided to facilitate further analysis.</li> <li>Uncertainties of data channels are provided based on data sheet specifications and GUM error propagation.</li> <li>The dataset refers to a real-scale application which is representative of typical large-scale solar thermal plant designs (flat plate collectors, common hydraulic layout).</li> <li>Additional information is provided in a "Data in Brief" journal article: <a href="https://doi.org/10.1016/j.dib.2023.109224">https://doi.org/10.1016/j.dib.2023.109224</a></li> </ul> <p> </p> <p><strong>Collector array description: </strong>The data is from a flat plate collector array with a total gross collector area of 516 m<sup>2</sup> (361 kW nominal thermal power). The array consists of four parallel collector rows with a common inlet and outlet manifold. Large-area flat-plate collectors from Arcon-Sunmark A/S are used in the plant. Collectors are all oriented towards the south (180°), have a tilt angle of 30° and a row spacing of 3.1 m. The collector array is part of a large-scale solar thermal plant located at Fernheizwerk Graz, Austria (latitude: 47.047294 N, longitude: 15.436366 E). The plant feeds into the local district heating network and is one of the largest Solar District Heating installations in Central Europe.</p> <p> </p> <p><strong>Data files:</strong></p> <ul> <li><strong>FHW_ArcS__main__2017.csv</strong> – This is the main dataset. It is advised to use this file for further analysis. The file contains the full time series of all measured and all calculated data channels and their (propagated) measurement uncertainty (53 data channels in total). Calculated data channels are derived from measured channels (see script make_data.py below) and have the suffix __calc in their channel names. Uncertainty information is given in terms of standard deviation of a normal distribution (suffix __std); some data channels are assumed to have no uncertainty (e.g., sun azimuth or shadowing).</li> <li><strong>FHW_ArcS__main__2017.parquet</strong> – Same as FHW_ArcS__main__2017.csv, but in parquet file format for smaller file size and improved performance when loading the dataset in software.</li> <li><strong>FHW_ArcS__parameters.json</strong> – Contains various metadata about the dataset, in both human and machine-readable format. Includes plant parameters, data channel descriptions, physical units, etc.</li> <li><strong>FHW_ArcS__raw__2017.csv </strong>– Dataset with time series of all measured data channels and their measurement uncertainty. The main dataset FHW_ArcS__main__2017.csv, which includes all calculated data channels, is a superset of this file.</li> </ul> <p> </p> <p><strong>Scripts: </strong></p> <ul> <li><strong>make_data.py</strong> – This Python script exposes the calculation process of the calculated data channels (suffix __calc), including error propagation. The main calculations are defined as functions in the module utils_data.py.</li> <li><strong>make_plots.py</strong> – This Python script, together with utils_plots.py, generates several figures based on the main dataset.</li> </ul> <p> </p> <p><strong>Data collection and preparation</strong>: AEE — Institute for Sustainable Technologies (AEE INTEC), Feldgasse 19, 8200 Gleisdorf, Austria; and SOLID Solar Energy Systems GmbH (SOLID), Am Pfangberg 117, 8045 Graz, Austria</p> <p> </p> <p><strong>Data owner</strong>: solar.nahwaerme.at Energiecontracting GmbH, Puchstrasse 85, 8020 Graz, Austria</p> <p> </p> <p><strong>Additional information</strong> is provided in a journal article in "Data in Brief", titled <a href="https://doi.org/10.1016/j.dib.2023.109224">"One year of high-precision operational data including measurement uncertainties from a large-scale solar thermal collector array with flat plate collectors in Graz, Austria"</a>.</p> <p> </p> <p><strong>Note: </strong>A Gitlab repository is associated with this dataset, intended as a companion to facilitate maintenance of the Python code that is provided along with the data. If you want to use or contribute to the code, please do so using the Gitlab project: <a href="https://gitlab.com/sunpeek/zenodo-fhw-arconsouth-dataset-2017">https://gitlab.com/sunpeek/zenodo-fhw-arconsouth-dataset-2017</a></p> <p> </p>
Data for "Detection of large-scale cloud microphysical changes within a major shipping corridor after implementation of the IMO 2020 fuel sulfur regulations"
<p>Processed data used for the manuscript "Detection of large-scale cloud microphysical changes within a major shipping corridor after implementation of the IMO 2020 fuel sulfur regulations".</p> <p>Includes input data for kriging algorithm as "SSF1deg_shipkrige_Terra.nc" and output data files as "Data_Terra_[VAR]_[YEAR]_C_M[MONTH].nc" for [VAR] Acld (overcast albedo) or cer (cloud droplet effective radius), [YEAR] the starting year of a three-year period starting with 2002 and ending at 2020 or "clim" for the 2002-2019 climatology, and [MONTH] 1to12 (annual mean) or 9to11 (austral spring).</p> <p>For the output data, "Obs" is the original data, "Est" is the mean counterfactual field obtained via kriging, "lowEst" and "highEst" are the 2.5th and 97.5th percentiles of the kriged fields for each grid box, "krSims" stores the results of the 5,000 simulated kriged fields, "Semivariance" is the binned empirical variogram values, "pVal" is the raw field significance (not adjusted for multiple testing), "nOut" is the number of individually significant grid boxes, "tran" is the transform applied (none for cer, logit for Acld), "iniPhi" and "iniSigma2" are the initial values for the fitted variogram, "Phi" and "Sigma2" are the fitted values using weighted least squares, and "parSel" is the list of selected regressors for the mean function that minimize the Bayesian information criterion.</p>
NPM3D dataset with instance labels used in paper "Toward Accurate Instance Segmentation in Large-scale LiDAR Point Clouds"
<p>NPM3D (https://npm3d.fr/paris-carla-3d) consists of mobile laser scanning (MLS) point clouds collected in four different regions in the French cities of Paris and Lille, where each point has been annotated with two labels: one that assigns it to one out of 10 semantic categories and another one that assigns it to an object instance. When inspecting the data, we found 9 cases where multiple tree instances had not been separated correctly (i.e., they had the same ground truth instance label). These cases were manually corrected using the CloudCompare software (https://www.cloudcompare.org), and 35 individual tree instances were obtained. Our variant of the dataset with 10 semantic categories and enhanced instance labels is publicly available.</p>
Two Large-Scale Meteorological Patterns Are Associated with Short-Duration Dry Spells in the Northeastern United States
<p><strong>Description</strong></p> <p>This dataset contains processed data from the ERA5 dataset for some of the atmospheric fields considered in this study. Original (pre-processed) ERA5 data (Hersbach et al. 2020) is available at <a href="https://cds.climate.copernicus.eu/#!/search?text=ERA5&type=dataset">https://cds.climate.copernicus.eu/#!/search?text=ERA5&type=dataset</a>. For each processed data file (netCDF format), the time steps correspond with the events and numerical order as listed in Table 1 of the main manuscript text. Processed data files are given for some of the 12-day averaged dry periods. Other processed data files are available from the authors upon reasonable request. </p> <p> </p> <p><strong>Abstract</strong></p> <p>Large-scale meteorological pattern (LSMP) – based analysis is used novelly to understand antecedent conditions and characteristics of short-duration dry spell events over the northeastern United States. Dry spell events are identified from histograms of consecutive dry days below a daily precipitation threshold. Events lasting twelve days or longer, which correspond to ~10% of dry spell events, are examined. The 500-hPa stream function anomaly fields for the first twelve days of each event are time-averaged and k-means clustering is applied to isolate the dry spell-related LSMPs. The first cluster has a strong, low-pressure anomaly over the Atlantic Ocean, southeast of the region, and is more common in winter and spring. The second cluster has strong, high-pressure over east-central North America and is most common during autumn. Over the region, both clusters have negative specific humidity anomalies, negative integrated vapor transport from the north, and subsidence associated with a midlatitude jet stream dipole structure that reinforces upper-level convergence. Subsidence is supported by cold air advection in the first cluster and the location on the east side of the lower-level high pressure in the second cluster. Extratropical cyclone storm track density across the Northeast is dramatically reduced during these dry spell events. Individual events lie on a continuum between two distinct clusters. These clusters have similar local, but quite different remote, properties. More (56%) short-duration dry spells occurred during the numerous non-drought months than drought months, however the frequency of dry spells is more than three times greater during drought than non-drought months.</p> <p> </p>
Electron transport measurements in liquid xenon with Xenoscope, a large-scale DARWIN demonstrator
<p>Drift velocity and longitudinal diffusion for the manuscript:</p> <p>Electron transport measurements in liquid xenon with Xenoscope, a large-scale DARWIN demonstrator</p> <p>https://arxiv.org/abs/2303.13963</p> <p> </p>
Audiovisual Moments in Time: A Large-Scale Annotated Dataset of Audiovisual Actions
<p>We present Audiovisual Moments in Time (AVMIT), a large-scale dataset of audiovisual action events. In an extensive annotation task 11 participants labelled a subset of 3-second audiovisual videos from the Moments in Time dataset (MIT). For each trial, participants assessed whether the labelled audiovisual action event was present and whether it was the most prominent feature of the video. The dataset includes the annotation of 57,177 audiovisual videos, each independently evaluated by 3 of 11 trained participants. From this initial collection, we created a curated test set of 16 distinct action classes, with 60 videos each (960 videos). We also offer 2 sets of pre-computed audiovisual feature embeddings, using VGGish/YamNet for audio data and VGG16/EfficientNetB0 for visual data, thereby lowering the barrier to entry for audiovisual DNN research. We further carried out an experiment to explore the utility of the AVMIT annotations and feature embeddings. A series of 6 Recurrent Neural Networks (RNNs) were trained on either AVMIT-filtered audiovisual events or modality-agnostic events from MIT, and then tested on our audiovisual test set. In all RNNs, top 1 accuracy was increased by 2.71-5.94\% by training exclusively on audiovisual events, even outweighing a three-fold increase in training data. We anticipate that the newly annotated AVMIT dataset will serve as a valuable resource for research and comparative experiments involving computational models and human participants, specifically when addressing research questions where audiovisual correspondence is of critical importance.</p>
A scalable, accurate, and universal analysis framework using individual-level allele frequency for large-scale genetic association studies in an admixed population
<p>Inclusion of individuals with diverse or admixed genetic ancestries is crucial to discover novel findings that may be missed by genomics analyses rooted solely in Caucasian population. Here, we present an analysis framework, SPAmix, which is scalable to a large-scale biobank data analysis including hundreds of thousands of admixed individuals and is universally applicable to various types of complex traits including binary trait, quantitative trait, time-to-event trait, longitudinal traits, etc. For each genetic variant, SPAmix uses genotype data and genetic principal components (PCs) to estimate individual-level allele frequency, which is subsequently used to calibrate p values via a retrospective analysis. A hybrid strategy including saddlepoint approximation (SPA) can greatly increase the accuracy to analyze rare genetic variants, especially if the phenotypic distribution is unbalanced or extremely unbalanced. Compared to Tractor, SPAmix does not require local ancestry information and can be straightforwardly applicable to a multi-way admixed population. Meanwhile, SPAmix can also be extended to SPAmix<sub>local</sub> in which the local ancestry can be incorporated if available. In addition, we propose SPAmix<sub>CCT</sub> to combine the p values of SPAmix and SPAmix<sub>local</sub> via Cauchy combination (CCT). SPAmix<sub>local</sub> performs close to Tractor when analyzing quantitative traits and is more accurate when analyzing binary traits with an unbalanced case-control ratio. And SPAmix<sub>CCT </sub>is an optimal unified approach for various cross-ancestry genetic architectures. Extensive simulation studies and real data analyses of 369,314 UK Biobank individuals from multiple ancestries demonstrated that SPAmix is scalable and can discover novel hits while controlling type I error rates well.</p>
Dataset and Software for The Relationships Between Large-scale Variations in Shear Velocity, Density, and Compressional Velocity in the Earth's Mantle
<p><strong>Is there a chemically distinct reservoir in the Earth?</strong><br><strong>Do superplumes overly denser-than-average material?</strong><br><strong>Can we detect these anomalies with seismic data?</strong><br><strong>Can we evaluate statistical significance of the features in tomography?</strong></p><p>This study presents the <strong>strongest evidence</strong> to date (ca. 2015) of <strong>large-scale thermo-chemical heterogeneities in the lowermost mantle</strong> using the full spectrum of seismic data. A large data set of surface-wave phase anomalies, body-wave travel times, normal-mode splitting functions and long-period waveforms is used to investigate the scaling between shear velocity, density and compressional velocity in the Earth's mantle (ϱ=dln ρ/dln vS, ν=dln vS/dln vP). Our preferred joint model consists of denser-than-average anomalies (∼1% peak-to-peak) at the base of the mantle roughly coincident with the low-velocity superplumes. The relative variation of shear velocity, density and compressional velocity in our study disfavors a purely thermal contribution to heterogeneity in the lowermost mantle, with implications for the long-term stability and evolution of superplumes.</p><p><strong>Note on Odd Degree Structure:</strong></p><p>Since the self-coupled normal-mode splitting observations constrain only even-degree density variations, all inversions strongly disfavored even-degree vS-ρ correlation (R2 ~ –0.46 to –0.25) in the lowermost mantle, which also disfavors a purely thermal contribution to heterogeneity in this region. However, the starting assumptions on positive vS-ρ correlation persisted in the remaining regions and for odd degree variations. In viscosity inversions with the geoid, opposing sign of the correlation of the longest wavelength even-versus odd-degree structure maps into a region of reduced viscosity in the lower mantle (Rudolph et al., 2020, doi:10.1029/2020gc009335). While important for such dynamical implications, <strong>odd-degree density variations in the lowermost mantle are poorly constrained in this study and should not be interpreted</strong>. We therefore used even-degree variations up to degree 6 for our inferences on thermo-chemical variations in the lowermost mantle (Figure 14), and provide those values in the files below.</p><p><strong>Feedback/Questions?</strong> Please contact Raj Moulik (<a href="https://rajmoulik.com">rajmoulik.com</a>) at <a href="mailto:moulik@caa.columbia.edu?subject=Query%20from%20Zenodo">moulik@caa.columbia.edu</a> </p><p><strong>Reference:</strong></p><p><i>Please cite the following work if you use this data or software.</i></p><ul><li>Moulik, P. & Ekström, G., 2016. The relationships between large-scale variations in shear velocity, density and compressional velocity in the Earth's mantle, <i>J. Geophys. Res.</i>, <strong>121</strong>, doi: <a href="http://dx.doi.org/10.1002/2015JB012679">10.1002/2015JB012679</a>. <a href="https://rajmoulik.com/Publications/MoulikEkstrom_JGR2016.pdf"><i>pdf</i></a></li></ul><p><i>You can also cite the dataset and software from this Zenodo page (Optional).</i></p><p>Moulik, P. & Ekström, G. (2016). Dataset and Software for The Relationships Between Large-scale Variations in Shear Velocity, Density, and Compressional Velocity in the Earth's Mantle. In J. Geophys. Res. Solid Earth (v1.0, Vol. 121, pp. 2737–2771). Zenodo. doi: <a href="https://doi.org/10.5281/zenodo.8356540">10.5281/zenodo.8356540</a></p><p><strong>Data Products:</strong></p><ul><li><strong>ME16_Figures(</strong><a href="https://zenodo.org/api/files/9aa99409-20ae-495e-b5a3-288fd57ecaeb/ME16_Figures.tar.gz"><strong>.tar.gz</strong></a><strong> or </strong><a href="https://zenodo.org/api/files/9aa99409-20ae-495e-b5a3-288fd57ecaeb/ME16_Figures.pdf"><strong>.pdf</strong></a><strong>)</strong> - contains all figures from the paper in .png format</li><li><a href="https://zenodo.org/api/files/9aa99409-20ae-495e-b5a3-288fd57ecaeb/ME16"><strong>ME16</strong></a><strong> - </strong>Coefficients of the spline basis functions for each parameter. Refer cij in equation 3. This is our preferred global model of anisotropic elastic parameters and density. Density variations are allowed to deviate from a constant scaling with shear-velocity variations in the lowermost mantle, which is required to fit the longest-period normal modes (e.g. 0S2). Radial anisotropy is confined to the uppermost mantle (that is, since the anisotropy is parameterized with only the four uppermost splines, it becomes very small below a depth of 250 km, and vanishes at 410 km). This is an updated version of S362ANI+M (Moulik and Ekström, 2014) which did not solve independently for density and compressional-wave velocity variations and imposed a constant scaling throughout the mantle instead (ϱ=0, ν=1/0.55).</li><li><a href="https://zenodo.org/api/files/9aa99409-20ae-495e-b5a3-288fd57ecaeb/STW105"><strong>STW105</strong></a> - reference model used in ME16. Described in Kustowski et al. (2008)</li><li><a href="https://zenodo.org/api/files/9aa99409-20ae-495e-b5a3-288fd57ecaeb/setup.cfg"><strong>setup.cfg</strong></a><strong> - </strong>Some configuration metadata relevant to this model for reproducibility.</li><li><a href="https://zenodo.org/api/files/9aa99409-20ae-495e-b5a3-288fd57ecaeb/epix.tar.gz"><strong>epix.tar.gz</strong></a> - Perturbations in horizontally (<i>vsh</i>) and vertically polarized shear velocity (<i>vsv</i>), Voigt-average isotropic shear-wave (<i>vs</i>) and compressional-wave velocity (<i>vp</i>), density (<i>rho</i>). anisotropy (<i>as</i>) and topography of the internal boundaries. This is calculated from the spline coefficients at every 1 by 1 degree cell-centered pixel and at every ~25 km depth region from Moho to the core-mantle boundary and stored in extended pixel format (.epix) ASCII files. Even-degree variations up to degree 6 are provided for density (<i>rho_even6)</i> and isotropic shear-wave velocity (<i>vs_even6</i>), which should be used for density inferences on thermochemical structure (See note above).</li><li><a href="https://zenodo.org/api/files/9aa99409-20ae-495e-b5a3-288fd57ecaeb/ME16.BOX25km_PIX1X1.avni.nc4"><strong>ME16.BOX25km_PIX1X1.avni.nc4</strong></a> - The perturbations in a standard AVNI format that utilizes the NETCDF4 container format. This file can be read in Python using either xarray or AVNI libraries. For example, to plot even-degree variations up to degree 6 in Voigt-averaged shear velocity perturbations at the bottom of the mantle (2875-2891 km depth)<ul><li><i>import xarray as xr</i></li><li><i>ds = xr.open_dataset('ME16.BOX25km_PIX1X1.avni.nc4')</i></li><li><i>ds['vs_even6'][-1].plot()</i></li></ul></li><li><a href="https://zenodo.org/api/files/9aa99409-20ae-495e-b5a3-288fd57ecaeb/PROGRAMS.tar.gz"><strong>PROGRAMS.tar.gz</strong></a> - Fortran tools for obtaining model values at specific locations. After creating the executables from source code in the <i>src</i> folder, the <i>readme</i> script generates most of the epix files provided in epix.tar.gz above<strong>.</strong></li><li><a href="https://zenodo.org/api/files/9aa99409-20ae-495e-b5a3-288fd57ecaeb/profilescaling.txt"><strong>profilescaling.txt</strong></a> - contains the median scaling ratios as used in Figure 15(a).</li><li><a href="https://zenodo.org/api/files/9aa99409-20ae-495e-b5a3-288fd57ecaeb/scaling3D_MoulikJGR16.tar.gz"><strong>scaling3D_MoulikJGR16.tar.gz</strong></a> - contains the scaling ratios and poisson ratio calculated from the joint model, as used in Figure 15(b).</li></ul>
Large-scale longitudinal gradients of genetic diversity: a meta-analysis across six phyla in the Mediterranean basins
Predicting patterns of variation in biodiversity across the globe is a fundamental issue in ecology and evolution. Diversity within species, that is, genetic diversity, is of prime importance for understanding past and present evolutionary patterns, and highlighting areas where conservation might be a priority. However, most studies on spatial patterns of genetic diversity have not considered longitude as a potentially important ecological driver of these patterns. Therefore, we carried out a meta-analysis to examine the longitudinal patterns of genetic diversity in the Mediterranean Basin. Using published literature and a systematic review/meta-analysis framework, we collected data on the genetic diversity of species whose populations occur in the Mediterranean basin. We then calculated a coefficient of correlation between within‐population genetic diversity indices and longitude, and estimated the role of biological, ecological, biogeographic, and marker type factors on the strength and magnitude of this correlation in six phylla. The results of this study were published in the paper titled Large‐scale longitudinal gradients of genetic diversity: a meta‐analysis across six phyla in the Mediterranean basin (Conord et al. 2012).
Large-scale woody plant removal outcomes in southern New Mexico, USA, 2007-2024
This data package contains datasets from a large-scale woody plant-specific herbicide experiment in southern New Mexico, USA. The study monitored vegetation cover in 43 pairs of plots representing treated and untreated conditions of the same plant community and environmental setting, including baseline and records at 5, 10, and in some cases 15 years post treatment. We also evaluated environmental factors that may control variation in treatment outcomes. This package supports the manuscript “Large-scale experimental evaluation of woody plant removal outcomes in desert grassland: restoration, novelty, or degradation?” by Bestelmeyer, Burkett, James, Gamon, and Schooley.
Data related to the manuscript "Bayesian Calibration and Validation of a Large-scale and Time-demanding Sediment Transport Model"
<p>1) Riverbed_Elevation_Measurements.txt<br> Description: Measured riverbed geometry of available years<br> Columns: Node ID, Easting [m], Northig [m], Elevation 2002 [m asl], Elevation 2005 [m asl], Elevation 2010 [m asl], Elevation <br> 2013 [m asl]<br> ----------------------------------------------------------------------------------------------------------------------------<br> 2) Hydro_FT_2D_manual.txt<br> Description: Simulation results of the manually calibrated full model<br> Columns: Node ID, Easting [m], Northig [m], Elevation 2005 [m asl], Elevation 2010 [m asl], Elevation 2013 [m asl]</p> <p>3.1) Hydro_FT_2D_CollocationPointBase.txt<br> Description: Parameter combinations of the collocation point base for each of the 20 simulations conducted with the full model to <br> construct the surrogate<br> Rows: Critical Shields parameter, Grain Roughness, Grain Size distribution</p> <p>3.2) Hydro_FT_2D_CollocationResults.txt<br> Description: Simulation results of the 20 simulations conducted with the full model at the collocation points<br> Columns: Node ID, Easting [m], Northig [m], Elevations 2005 [m asl] of simulation 1 through 20, Node ID, Easting [m asl], Northig<br> [m asl], Elevations 2010 [m asl] of simulation 1 through 20, Node ID, Easting [m asl], Northig [m asl], Elevations 2013 [m asl] of<br> simulation 1 through 20<br> ----------------------------------------------------------------------------------------------------------------------------<br> 4.1) aPC_MC_N_Combinations_Weights_prior.txt<br> Description: ID of prior MC runs with tested parameter combinations and corresponding importance weights<br> Rows: ID of MC runs, Critical Shields parameter, Grain Roughness, Grain Size distribution, importance weights<br> 4.2) aPC_MC_2005_prior.txt<br> Description: aPC surrogate results of prior MC runs for 2005<br> Columns: Node ID, Easting [m], Northig [m], Elevations 2005 [m asl] of MC run 1 through 100,000<br> 4.3) aPC_MC_2010_prior.txt<br> Description: aPC surrogate results of prior MC runs for 2010<br> Columns: Node ID, Easting [m], Northig [m], Elevations 2010 [m asl] of MC run 1 through 100,000<br> 4.4) aPC_MC_2013_prior.txt<br> Description: aPC surrogate results of prior MC runs for 2013<br> Columns: Node ID, Easting [m], Northig [m], Elevations 2013 [m asl] of MC run 1 through 100,000<br> <br> 4.5) aPC_MC_N_Combinations_Weights_posterior.txt<br> Description: ID of accepted (posterior) MC runs with tested parameter combinations and corresponding importance weights<br> Rows: ID of accepted MC runs, Critical Shields parameter, Grain Roughness, Grain Size distribution, importance weights<br> 4.6) aPC_MC_2005_posterior.txt<br> Description: aPC surrogate results of posterior MC runs for 2005<br> Columns: Node ID, Easting [m], Northig [m], Elevations 2005 [m asl] of accepted MC run 1 through 857<br> 4.7) aPC_MC_2010_posterior.txt<br> Description: aPC surrogate results of posterior MC runs for 2010<br> Columns: Node ID, Easting [m], Northig [m], Elevations 2010 [m asl] of accepted MC run 1 through 857<br> 4.8) aPC_MC_2013_posterior.txt<br> Description: aPC surrogate results of posterior MC runs for 2013<br> Columns: Node ID, Easting [m], Northig [m], Elevation 2013 [m asl] of accepted MC run 1 through 857<br> ----------------------------------------------------------------------------------------------------------------------------<br> 5) aPC_MAP.txt<br> Description: Simulation results conducted with the stochastically calibrated aPC surrogate model using the MAP parameter <br> combination<br> Columns: Node ID, Easting [m], Northig [m], Elevation 2005 [m asl], Elevation 2010 [m asl], Elevation 2013 [m asl]</p> <p>6) Hydro_FT_2D_MAP.txt<br> Description: Simulation results conducted with the stochastically calibrated full model using the MAP parameter combination<br> Columns: Node ID, Easting [m], Northig [m], Elevation 2005 [m asl], Elevation 2010 [m asl], Elevation 2013 [m asl]<br> ----------------------------------------------------------------------------------------------------------------------------<br> 7) dz.txt<br> Description: Riverbed Evolution for all nodes in the section of interest (n=1138) obtained with differently calibrated models for all <br> considered time periods<br> Columns: Node ID, Easting [m asl], Northig [m asl], aPC_prior 2005 [m], aPC_posterior 2005 [m], aPC_MAP 2005 [m], <br> Hydro_FT-2D_MAP 2005 [m], Hydro_FT-2D_manual 2005 [m], aPC_prior 2010 [m], aPC_posterior 2010 [m], aPC_MAP 2010 [m],<br> Hydro_FT-2D_MAP 2010 [m], Hydro_FT-2D_manual 2010 [m], aPC_prior 2013 [m], aPC_posterior 2013 [m], aPC_MAP 2013 [m],<br> Hydro_FT-2D_MAP 2013 [m], Hydro_FT-2D_manual 2013 [m]</p> <p>8) dz_CalibrationNodes.txt<br> Description: Riverbed Evolution for calibration nodes (n=204) obtained with differently calibrated models for all considered time<br> periods<br> Columns: Node ID, Easting [m asl], Northig [m asl], aPC_prior 2005 [m], aPC_posterior 2005 [m], aPC_MAP 2005 [m], <br> Hydro_FT-2D_MAP 2005 [m], Hydro_FT-2D_manual 2005 [m], aPC_prior 2010 [m], aPC_posterior 2010 [m], aPC_MAP 2010 [m],<br> Hydro_FT-2D_MAP 2010 [m], Hydro_FT-2D_manual 2010 [m], aPC_prior 2013 [m], aPC_posterior 2013 [m], aPC_MAP 2013 [m],<br> Hydro_FT-2D_MAP 2013 [m], Hydro_FT-2D_manual 2013 [m]</p> <p> </p>
A large-scale wide-baseline light field dataset - Part I
<p>This dataset is the Part I of a large-scale, synthetic wide-baseline light field dataset (called WLF), including 345 light fields. </p> <p>Each light field provides 9x9 angular (RGB) images and ground truth disparities. This light field dataset involves the spatial resolution (512x512) images only.</p> <p>The dataset is originally created for training the deep learning-based models for depth estimation. You might use this dataset for other tasks if possible.</p> <p>You might also have a try to play with this dataset using our code in https://github.com/YanWQ/LLF-Net.</p>
ToxicoDB: an integrated database to mine and visualize large-scale toxicogenomic datasets (TGGATEs human dataset)
<p>This data was generated by Igarashi Y, Nakatsu N, Yamashita T, Ono A, Ohno Y, Urushidani T, Yamada H. Open TG-GATEs: a large-scale toxicogenomics database. Nucleic Acids Res [Internet]. 2015 Jan;43(Database issue):D921–7. Available from: http://dx.doi.org/10.1093/nar/gku955 PMCID: PMC4384023. The data have been curated and analyzed using our open-source R package, <em>ToxicoGx</em> (<a href="https://github.com/bhklab/ToxicoGx">https://bioconductor.org/packages/devel/bioc/html/ToxicoGx.html</a>), and are available publicly in the <em>ToxicoDB </em>web application (<a href="http://www.toxicodb.ca">www.toxicodb.ca</a>).</p>
The dataset of figures for "The GPU version of LICOM3 under HIP framework and its large-scale application" (updated)
<p>A high-resolution (1/20°) global ocean general circulation model with Graphics processing units (GPUs) code implementations is developed based on the LASG/IAP Climate system Ocean Model version 3 (LICOM3) under Heterogeneous-compute Interface for Portability (HIP) framework. The dynamic core and physics package of LICOM3 are both ported to the GPU, and 3-dimensional parallelization is applied. The HIP version of the LICOM3 (LICOM3-HIP) is 42 times faster than what the same number of CPU cores dose, when 384 AMD GPUs and CPU cores are used. The LICOM3-HIP has excellent scalability; it can still obtain speedup of more than four on 9216 GPUs comparing to 384 GPUs. In this phase, we successfully performed a test of 1/20° LICOM3-HIP using 6550 nodes and 26200 GPUs, and at the grand scale, the model’s time to solution can still obtain an increasing, about 2.72 simulated years per day (SYPD). The high performance was due to putting almost all of computation processes inside GPUs, and thus greatly reduces the time cost of data transfer between CPUs and GPUs. At the same time, a 14-year spin-up integration following the phase 2 of Ocean Model Intercomparison Project (OMIP-2) protocol of surface forcing has been conducted, and the preliminary results have been evaluated. We found that the model results have little differences from the CPU version. Further comparison with observations and lower-resolution LICOM3 results suggests that the 1/20° LICOM3-HIP can not only reproduce the observations, but also produce much smaller scale activities, such as submesoscale eddies and frontal scales structures.</p>
Large-Scale Gravitational Lens Modeling with Bayesian Neural Networks for Accurate and Precise Inference of the Hubble Constant - Datasets, Trained Models, BNN Samples, and MCMC Chains
<p>We publish the training/validation/test datasets, trained model weights, configuration files, Bayesian neural network samples, and MCMC chains used to produce the figures in the LSST DESC paper, "Large-Scale Gravitational Lens Modeling with Bayesian Neural Networks for Accurate and Precise Inference of the Hubble Constant." They are formatted to be used with the DESC package "H0rton" (<a href="https://github.com/jiwoncpark/h0rton">https://github.com/jiwoncpark/h0rton</a>). Additional descriptions can be found in the README. Please contact Ji Won Park (@jiwoncpark) on GitHub or <a href="https://github.com/jiwoncpark/h0rton/issues">make an issue</a> for any questions.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.