Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,037
datasets available to search
ShareScore release 0.9.0
Dataset results
1,037 results for “large-scale”
Linked collectors and determiners for: Bumble bees collected in a large-scale field experiment in power line clearings, southeast Norway.
Natural history specimen data linked to collectors and determiners held within, "Bumble bees collected in a large-scale field experiment in power line clearings, southeast Norway". Claims or attributions were made on Bionomia by volunteer Scribes, <a href="https://bionomia.net/dataset/78822c79-7646-448d-aac2-35700498c147">https://bionomia.net/dataset/78822c79-7646-448d-aac2-35700498c147</a> using specimen data from the dataset aggregated by the Global Biodiversity Information Facility, <a href="https://gbif.org/dataset/78822c79-7646-448d-aac2-35700498c147">https://gbif.org/dataset/78822c79-7646-448d-aac2-35700498c147</a>. Formatted as a Frictionless Data package.
Code and data of "Uncovering disease-related multicellular pathway modules on large-scale single-cell transcriptomes with scPAFA"
<p>Code and data to reproduce the analyses and figures presented in "Uncovering disease-related multicellular pathway modules on large-scale single-cell transcriptomes with scPAFA"</p>
Making sense of large-scale kinase inhibitor bioactivity data sets: a comparative and integrative analysis
<p>We carried out a systematic evaluation of target selectivity profiles across three recent large-scale biochemical assays of kinase inhibitors and further compared these standardized bioactivity assays with data reported in the widely used databases ChEMBL and STITCH. Our comparative evaluation revealed relative benefits and potential limitations among the bioactivity types, as well as pinpointed biases in the database curation processes. Ignoring such issues in data heterogeneity and representation may lead to biased modeling of drugs' polypharmacological effects as well as to unrealistic evaluation of computational strategies for the prediction of drug-target interaction networks. Toward making use of the complementary information captured by the various bioactivity types, including IC50, K(i), and K(d), we also introduce a model-based integration approach, termed KIBA, and demonstrate here how it can be used to classify kinase inhibitor targets and to pinpoint potential errors in database-reported drug-target interactions. An integrated drug-target bioactivity matrix across 52,498 chemical compounds and 467 kinase targets, including a total of 246,088 KIBA scores, has been made freely available.</p> <p>Please cite: </p> <p>https://pubmed.ncbi.nlm.nih.gov/24521231/ </p> <p>https://pubs.acs.org/doi/10.1021/ci400709d</p>
Research Artefact: What network simulator questions do users ask? a large-scale study of stack overflow posts
<p><strong>Research Artefact: What network simulator questions do users ask? a large-scale study of stack overflow posts</strong></p> <p>This is a research artefact for the paper: <strong>What network simulator questions do users ask? a large-scale study of stack overflow posts</strong>. This artefact is a repository consisting of the collected dataset including 2,322 network-simulator-related Stack Overflow questions. This artefact aims to enable researchers to replicate our dataset of the paper and reuse the dataset for further research.</p>
Data for the manuscript entitled "Factors driving large-scale ungulate carrion production in the Anthropocene"
<p>Data for the article entitled "Factors driving large-scale ungulate carrion production in the Anthropocene". The data includes 5 main datasets, each of those corresponding to 5 main carrion production sources in terrestrial ecosystems in peninsular Spain, namely; 1) Livestock, 2) Big Game Hunting, 3) Roadkills, 4) Predation and 5) Natural mortality. </p> <p>In case of any doubt/s or enquiries regarding this data, please, send an email to the corresponding author; Jon Morant Etxebarria (email: jmorant@aranzadi.eus). </p> <p> </p>
T-Evos: A Large-Scale Longitudinal Study on CI Test Execution and Coverage Evolution
<p><strong>Description</strong></p> <p>T-Evos is a dataset on test results and coverage evolution, covering 6,495 consecutive commits across 12 open-source Java projects.</p> <p> </p> <p><strong>Version 0.1.0</strong></p> <p>We have included the test results and code coverage data. The test results data has been combined into the `test_status.json` for ease of access at the root level of the repository. Each of the node zip files contains the code coverage data. In each of the zip files, the code coverage data folder is structured as followed:</p> <ul> <li>project_name <ul> <li>commit_id <ul> <li>details <ul> <li>test_name.json (individual test coverage data)</li> <li>...</li> </ul> </li> <li>coverage.json (full test coverage data on commit_id)</li> </ul> </li> <li>surefire_test_method_dict.json (available test cases)</li> </ul> </li> <li>project_name.csv (commit build status [success/fail])</li> </ul> <p>Note: the current version of the dataset includes temporary `project_cluster` folders that users might kindly ignore.</p>
FIG. 2. — Souzalopesmyia polleti n in Souzalopesmyia Albuquerque, 1951 (Diptera: Muscidae): new species from South America with an updated phylogeny based on morphological evidence, in Touroult J. (ed.), "Our Planet Reviewed" 2015 large-scale biotic survey in Mitaraka, French Guiana.
FIG. 2. — Souzalopesmyia polleti n. sp.: A-D, ♂: sternite 5, dorsal view (A); epandrium, cercal plate and surstyli, dorsal view (B); epandrium, cercal plate and surstyli, lateral view (C); hypandrium and associated structures, lateral view (D); E-H, ♀: ovipositor, dorsal view (E); ovipositor, ventral view (F); spermatheca (G). Scale bars: 0.5 mm.
Data for - Tracking one-in-a-million: Large-scale benchmark for microbial single-cell tracking with experiment-aware robustness metrics
<p><strong>Large-scale Corynebacterium glutamicum data set with Segmentation and Tracking Annotation</strong></p> <p>We provide five time-lapse sequences with manually corrected segmentation and tracking annotations of growing <strong><em>C. glutamicum</em></strong> cultivations. The dataset contains more than 1.4 million cell observations in 29k cell tracks and 14k cell divisions. We provide videos of the annotations (videos.zip) and the dataset in <a href="http://celltrackingchallenge.net/datasets/">Cell Tracking Challenge</a> format (ctc_format.zip). In the videos, cell contours are rendered in yellow, cell links between frames are colored red and cell divisions, and their links are colored in blue.</p> <p><strong>Data Acquisition</strong></p> <p><strong><em>Corynebacterium glutamicum</em></strong> ATCC 13032 was cultivated in BHI-medium at 30°C in this study. From and overnight preculture, the main culture was inoculated the next day with a starting OD600 of 0.05 and grown at 120 rpm to a OD600 of 0.25. A chip was fabricated, according to <a href="https://doi.org/10.1039/D0LC00711K">(Täuber et al., 2020)</a>, and fixed to the microscope’s holder. The main culture cells were transferred to monolayer growth chambers (height = 720 nm) on the microfluidic chip. Flow through the microfluidic device was mediated by pressure driven pumps with a pressure of 100 mbar on the medium reservoir.</p> <p>The time-lapse phase contrast images of five monolayer growth chambers were taken every minute using an inverted microscope (Nikon Eclipse Ti2) with a 100x oil emersion objective and a DS-QI2 camera (Nikon) at 15 % relative DIA-illumination intensity and 100 ms exposure time. The spatial image resolution is 0.072 μm/px.</p>
WHU-OHS: A benchmark dataset for large-scale Hyperspectral Image classification
<p>The WHU-OHS dataset is made up of 42 OHS satellite images acquired from more than 40 different locations in China. The imagery has a spatial resolution of 10 m (nadir) and a swath width of 60 km (nadir). There are 32 spectral channels ranging from the visible to near-infrared range, with an average spectral resolution of 15 nm. We cropped each image into 512 × 512 pixels with a stride of 32. There are 4822, 513, and 2460 sub-images in the training, validation, and test sets, respectively.</p> <p>For transferability test, we choose eight pairs of OHS images, and each pair contains one source image (S) and one target image (T):</p> <p>S1: Changchun</p> <p>T1: Jilin</p> <p>S2: Wuxi</p> <p>T2: Shanghai</p> <p>S3: Guangzhou</p> <p>T3: Zhongshan</p> <p>S4: Xining</p> <p>T4: Lanzhou</p> <p>S5: Hetian</p> <p>T5: Kelamayi</p> <p>S6: Anyi</p> <p>T6: Nanchang</p> <p>S7: Changde</p> <p>T7: Changsha</p> <p>S8: Tianjin</p> <p>T8: Tangshan</p> <p>The 26 OHS images except for the eight pairs:</p> <p>O1: Baoding</p> <p>O2: Chongqing</p> <p>O3: Fujin</p> <p>O4: Huainan</p> <p>O5: Huhehaote</p> <p>O6: Jinzhong</p> <p>O7: Luliang</p> <p>O8: Manasi_1</p> <p>O9: Manasi_2</p> <p>O10: Nanmulin</p> <p>O11: Neimenggu</p> <p>O12: Qingdao</p> <p>O13: Qinghuangdao</p> <p>O14: Shawan</p> <p>O15: Shenyang</p> <p>O16: Shuozhou</p> <p>O17: Songpan</p> <p>O18: Taian</p> <p>O19: Tongjiang_1</p> <p>O20: Tongjiang_2</p> <p>O21: Wuzhong</p> <p>O22: Xundian</p> <p>O23: Xuzhou</p> <p>O24: Yidu</p> <p>O25: Zangzu</p> <p>O26: Zhongshan</p> <p>The image patches have been normalized and scaled by 10000 to reduce storage cost. Divide the pixel values by 10000 and then the image patches can be used directly.</p>
Input data to model multiple effects of large-scale deployment of grass in crop-rotations at European scale
<p>This is the input dataset to a Python script (<a href="https://github.com/oskeng/MF-bio-grass">https://github.com/oskeng/MF-bio-grass</a>) used to model the effects of widespread deployment of grass in rotations with annual crops to provide biomass while remediating soil organic carbon (SOC) losses and other environmental impacts.</p> <p>For more information about the dataset and the study, see the original article:</p> <p>Englund, O., Mola-Yudego, B., Börjesson, P., Cederberg, C., Dimitriou, I., Scarlat, N., Berndes, G. Large-scale deployment of grass in crop rotations as a multifunctional climate mitigation strategy. GCB Bioenergy</p>
Data for: Zebra finch song ecology: monitoring of breeding, observational transects, focal and year-round acoustic recordings, and a large-scale simultaneous playback experiment
<p class="MsoNormal">Male songbirds sing to establish territories and to attract mates. However, increasing reports of singing in non-reproductive contexts and by females show that song use is more diverse than previously considered. Therefore, alternative functions of song, such as social cohesion and synchronisation of breeding, by and large were overlooked even in such well-studied species as the zebra finch (<em>Taeniopygia guttata</em>). In these social songbirds only the males sing and pairs breed synchronously in loose colonies following aseasonal rain events in their arid habitat. As males are not territorial, and pairs form long-term monogamous bonds early in life, conventional theory predicts that zebra finches should not sing much at all; yet they do and their song is the focus of hundreds of lab-based studies. We hypothesise that zebra finch song functions to maintain social cohesion and to synchronise breeding. Here we test this idea using data from five years of field studies, including observational transects, focal and year-round audio recordings, and a large-scale playback experiment. We show that zebra finches frequently sing while in groups, that breeding status influences song output at the nest and at aggregations, that they sing year-round, and that they predominantly sing when with their partner, suggesting that song remains important after pair formation. Our playback reveals that song actively features in social aggregations as it attracts conspecifics. Together, these results demonstrate that birdsong has important functions beyond territoriality and mate choice, illustrating its importance in coordination and cohesion of social units within larger societies.</p>
Data for: Large-scale long-term passive-acoustic monitoring reveals spatiotemporal activity patterns of boreal bats
<p class="MsoNormal"><span>The distribution ranges and spatio-temporal patterns in the occurrence and activity of boreal bats are yet largely unknown due to their cryptic lifestyle and lack of suitable and efficient study methods. We approached the issue by establishing a permanent passive-acoustic sampling setup spanning the area of Finland to gain an understanding on how latitude affects bat species composition and activity patterns in northern Europe. The recorded bat calls were semi-automatically identified for three target taxa; <em>Myotis</em> spp., <em>Eptesicus nilssonii</em> or <em>Pipistrellus nathusii</em> and the seasonal activity patterns were modeled for each taxa across the seven sampling years (2015–2021). We found an increase in activity since 2015 for <em>E. nilssonii</em> and <em>Myotis </em>spp. For <em>E. nilssonii</em> and <em>Myotis</em> spp. we found significant latitude -dependent seasonal activity patterns, where seasonal variation in patterns appeared stronger in the north. Over the years, activity of <em>P. nathusii</em> increased during activity peak in June and late season but decreased in mid season. We found the passive-acoustic monitoring </span><span>network to be an effective and cost-efficient method for gathering b</span><span>at activity data to analyze spatio-temporal patterns. Long-term data on the composition and dynamics of bat communities facilitates better estimates of abundances and population trend directions for conservation purposes and predicting the effects of cli</span><span>mate change.</span></p>
Large-scale Docking Datasets for Machine Learning
<p><strong>Large-scale virtual screening has become a valuable tool for early-phase drug discovery. Recent expansions of commercial chemical space have made it computationally intractable to evaluate all compounds in the libraries. Machine learning is one of the methods that aim to prioritize specific subsets of these vast libraries. In order to put these methods to the test, access to large-scale datasets is beneficial. To help the community benchmark their work, we share the docking scores of several ultralarge virtual screening campaigns.</strong></p> <p><strong>The datasets we provide contain canonical SMILES, compound identifiers, and docking scores. We docked two different chemical libraries against eight different biological targets with therapeutic relevance. The first dataset contained approximately 15.5 million molecules adhering to the "Rule-of-Four", whereas the second datasets consists of approximately 235 million "lead-like" molecules. The biological targets represent different classes of proteins and binding sites.</strong></p> <p><strong>More details on the datasets and our methods can be found on (<a href="http://github.com/carlssonlab/conformalpredictor">https://github.com/carlssonlab/conformalpredictor</a>) and our pre-print (<a href="https://doi.org/10.26434/chemrxiv-2023-w3x36">https://doi.org/10.26434/chemrxiv-2023-w3x36</a>). </strong></p> <p><strong>Please feel free to download and use these datasets for your own research purposes. We only ask that you cite our pre-print and datasets appropriately if you use it in your work. Thank you for your interest in our research!</strong></p>
Assessment of the acoustic adaptation hypothesis in frogs using large-scale citizen science data
<p>This is the data required to reproduce the results of the manuscript "Assessment of the acoustic adaptation hypothesis in frogs using large-scale citizen science data", including measurements of tree canopy cover extracted from the Global Forest Cover Change dataset (Townshend 2016).</p> <p> </p> <p><strong>Reference</strong></p> <p>Gillard, G. L. & Rowley, J. J. L. (2023). Assessment of the acoustic adaptation hypothesis in frogs using large-scale citizen science data. <em>Journal of Zoology</em>. [In publication].</p> <p> </p> <p><strong>Global Forest Cover Change Dataset</strong></p> <p>Townshend J. 2016. Global Forest Cover Change (GFCC) Tree Cover Multi-Year Global 30 m V003 [Data set]. NASA EOSDIS Land Processes DAAC. Accessed June 22, 2022. doi:10.5067/MEaSUREs/GFCC/GFCC30TC.003.Townshend J. 2016. Global Forest Cover Change (GFCC) Tree Cover Multi-Year Global 30 m V003 [Data set]. NASA EOSDIS Land Processes DAAC. Accessed June 22, 2022. doi:10.5067/MEaSUREs/GFCC/GFCC30TC.003.</p>
PubGraph: A Large-Scale Scientific Knowledge Graph
<p>We present PubGraph, a new resource for studying scientific progress that takes the form of a large-scale knowledge graph (KG) with more than 385M entities, 13B main edges, and 1.5B qualifier edges. PubGraph is comprehensive and unifies data from various sources, including Wikidata, OpenAlex, and Semantic Scholar, using the Wikidata ontology. Beyond the metadata available from these sources, PubGraph includes outputs from auxiliary community detection algorithms and large language models. To further support studies on reasoning over scientific networks, we create several large-scale benchmarks extracted from PubGraph for the core task of knowledge graph completion (KGC). These benchmarks present many challenges for knowledge graph embedding models, including an adversarial community-based KGC evaluation setting, zero-shot inductive learning, and large-scale learning. All of the aforementioned resources are accessible at <a href="https://purl.archive.org/pubgraph">https://purl.archive.org/pubgraph</a> and released under the CC-BY-SA license. We plan to update PubGraph quarterly to accommodate the release of new publications.</p>
Increased precipitation over land due to climate feedback of large-scale bioenergy cultivation
<p>Biophysical effects of different bioenergy crop cultivation scenarios on global water cycles simulated by coupled IPSL-CM model, with ORCHIDEE-MICT-BIOENERGY as the land component and LMDz as the atmosphere component. The spatial resolution of the coupled model was 1.26° latitude × 2.5° longitude (i.e., 143*144 grid cells globally).</p> <p>Datasets includes:</p> <p>1. Source data and plotting code for Figure 1, the global precipitation changes induced by bioenergy cultivation.</p> <p>2. Source data and plotting code for Figure 2, the diagnostic precipitation changes for global land area, in and outside of the bioenergy cultivation area.</p> <p>3. Source data plotting code for Figure 3, the changes in water balance for global land area, in and outside of bioenergy cultivation area, monsoon regions and four different humidity zones.</p>
Figure S2 in Large-scale snake genome analyses provide insights into vertebrate development
Figure S2. Snake genome features, related to Figure 2 (A) Evolution of chromosomes in snakes. In total, 23 proto-chromosomes of Serpentes were reconstructed using four lizards as outgroup. (B) Circos plots showing conserved synteny between the hypothesized Serpentes ancestor: Serpentes (red) and Hong Kong dwarf snake (Csep-blue). (C) Snake body lengths were significantly negatively correlated with the genome evolutionary rate (correlation coefficient = 0.50, p value = 0.011). (D) Length distribution of snake ancestor gain and lost genome segments. Length distribution of snake ancestor unique genome segments (left). Length distribution of snake ancestor lost genome segments (right).
Figure S3 in Large-scale snake genome analyses provide insights into vertebrate development
Figure S3. Snake-specific genome structural variations (SSSVs), snake-diverged conserved non-coding elements (SD-CNEs that were diverged in snakes but were still conserved in the outgroup) (SD-CNEs), and orthologous genes used for evolutionary analysis, related to Figure 3 and STAR Methods (A) SSSVs distribution in different genome regions. (B) Genomic reads coverage of the snake-specific lost gene GHRL across four lizards (green anole [Acar], Anan's rock agama [Lsac], Komodo dragon [Vkom],and viviparous lizard [Zviv]), and 14 newly sequenced snakes. (C) Z-score cut off of SD-CNEs. (D) The enriched MGI terms that related to snake phenotypes for genes with SD-CNE (adjusted p value <0.05). SD-CNEs (dots) are located around the transcription start site of the genes enriched in eye, eyelid, ear, lung, mandible, maxillary, sternum, and tooth development-related terms. (E) Enriched GO terms (adjusted p value <0.05) for the newly evolved coding genes in snakes. Those related to dietary excess, protein digestion, and olfactory receptor activity were colored in red, blue, and brown, respectively. (F) Enriched GO terms (adjusted p value <0.05) for snake-specific coding genes that were lost. Red, blue, orange, and green were used to indicate the terms involved in vision, lens development, appetite, and bile acid biosynthetic process, respectively. (G) SSSV deletion in the potential regulatory region of RP1. (H) Counts of orthologous genes used for evolutionary analysis among 27 selected species.
Figure 3 in Large-scale snake genome analyses provide insights into vertebrate development
Figure 3. Genetic basis of skeletal system evolution and organ adaptation in snakes PSGs, REGs, WGCNA hub genes, lost genes, SSSV-related genes, newly evolved genes, and SD-CNE-associated genes are marked in different colors and are represented by rectangles. (A) Renewal of genomic elements contributing to snake skull development and digestion. Related genes are shown in their corresponding regions. (B) The evolution of genomic elements has potentially facilitated the evolution of the elongated body plan of snakes. (C) The evolution of genes related to lung development and function. The heatmap shows the expression levels of WGCNA hub genes, PSGs, and SSSV-related genes in green anole (Acar), keeled slug snake (Pber), and many-banded krait (Bmul) (The ''_numbers'' represents the different copies of a multi-copy gene). (D) Whole-body X-ray images and relative lengths of the fourth forelimb phalanx and third hindlimb phalanx of wild and PTCH1-mutated mice (eight samples per group, the significance is indicated as *p <0.05). Mean ± SD is shown by error bar. See also Figure S4 and Table S3.
Figure S1 in Large-scale snake genome analyses provide insights into vertebrate development
Figure S1. Phylogenetic and divergent time tree of snakes, related to Figure 1 (A) Phylogenetic tree distributions of selected snake species, topology inferred from a previous study.173 (B) Maximum likelihood (ML) phylogenetic tree inferred from the 31-taxon whole-genome alignments. Ultrafast bootstraps were repeated 3,000 times, with NNI UFBoot tree optimization and SH-like approximate likelihood ratio test (SH-aLRT) performed. SH-aLRT support rate/ultrafast bootstrap support rate is indicated at each node. (C) Inferred ML phylogenetic tree using orthologous genes. 1,980 1:1 orthologous genes were concatenated to a super gene sequence for constructing the ML phylogenetic tree using IQ-TREE. 10,000 ultrafast bootstraps were carried out and the two support rates were marked at each node. (D) ML phylogenetic tree inferred from 4d sites. 4d sites were extracted from the orthologous genes and taken as input for IQ-TREE to infer the ML phylogenetic tree. The two supports were computed by 10,000 ultrafast bootstraps and SH-aLRT. (E) Inferred ML phylogenetic tree of conserved non-coding elements (CNEs). All CNEs were identified and concatenated into a single sequence for the ML phylogenetic tree inferred using 5,000 ultrafast bootstraps and SH-aLRT. (F) Coalescent phylogenetic tree inferred from 51,302 1-kb orthologous genomic segments using ASTRAL-III. (G) Coalescent phylogenetic tree inferred from 1,980 1:1 orthologous genes using ASTRAL-III. The poorly supported nodes are indicated in red. (H) DicoVista gene tree topologies frequency analysis. The frequency of three topologies (t1–t3) is shown, and the red is the main topology. The divergence time (million years ago [mya]) of the species is estimated using r8s and the incongruent clades are in red. (I) Divergence time of the 31 species. r8s estimated divergence time using the whole-genome alignments.The estimated divergence time (million years ago [mya]) is labeled on each node and 6 calibrating nodes used are marked. (J) MCMCTree estimated divergence time of the 31 species using 4d sites. Six calibrating nodes used are marked and the estimated divergence time represented in million years is labeled.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.