Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
358
datasets available to search
ShareScore release 0.9.0
Dataset results
358 results for “dataset generation”
Transcriptome analysis of wildtype elmo1+/+ and homozygous elmo1-/- knockout zebrafish larvae at 120 hpf through next generation RNA sequencing [dataset 1]
GEO Series GSE197824. Danio rerio. 12 samples. Type: Expression profiling by high throughput sequencing.
Next generation sequencing quantitatively analyzes transcriptomes of primary CLL cells tranfected with non-specific control (NSC) or ZAP-70 siRNA (+anti-IgM dataset)
GEO Series GSE149021. Homo sapiens. 12 samples. Type: Expression profiling by high throughput sequencing.
Generation of multi-omic datasets using high-throughput molecular profiling of transcriptomic human data
GEO Series GSE281204. Homo sapiens. 39 samples. Type: Expression profiling by high throughput sequencing.
Combination anti-PD-1 and anti-CTLA-4 therapy generates waves of clonal responses that include progenitor-exhausted CD8 T cells [dataset 4]
GEO Series GSE273718. Homo sapiens. 6 samples. Type: Expression profiling by high throughput sequencing.
Transcriptome analysis of wildtype elmo3+/+ and homozygous elmo3-/- knockout zebrafish larvae at 120 hpf through next generation RNA sequencing [dataset 3]
GEO Series GSE197826. Danio rerio. 12 samples. Type: Expression profiling by high throughput sequencing.
First Generation Tools for the Modeling of Capicua-Family Fusion Oncoprotein-Driven Cancers (May 2024 Dataset)
GEO Series GSE295623. Homo sapiens. 21 samples. Type: Expression profiling by high throughput sequencing.
Generated datasets for Yue et al. (2020, Earth and Space Science): "Combining In-situ and Satellite Observations to Understand the Vertical Structure of Tropical Anvil Cloud Microphysical Properties During the TC4 Experiment"
<p>This archive contains the data sets generated from the research conducted by Yue et al. (2020) titled "Combining In-situ and Satellite Observations to Understand the Vertical Structure of Tropical Anvil Cloud Microphysical Properties During the TC4 Experiment" published in Earth and Space Science. The method to generated the following data sets is described in Yue et al. (2020) and stored as Matlab .mat files.</p> <p>CombiningTC4_Satellite_eof_cov_mat.mat contains the correlation matrix shown in Figure 1a.</p> <p>TC4_processed.mat contains the correlation matrix shown in Figure 1b.</p> <p>RO_processed.mat contains the correlation matrix shown in Figure 2a.</p> <p>RVOD_processed.mat contains the correlation matrix shown in Figure 2b.</p> <p>ICE_processed.mat contains the correlation matrix shown in Figure 2c.</p> <p> </p> <p> </p>
Afro-MNIST: Synthetic generation of MNIST-style datasets for low-resource languages
<p>We present Afro-MNIST, a set of synthetic MNIST-style datasets for four orthographies used in Afro-Asiatic and Niger-Congo languages: Ge`ez (Ethiopic), Vai, Osmanya, and N'Ko.<br> These datasets serve as ``drop-in'' replacements for MNIST. We hope that MNIST-style datasets will be developed for other numeral systems, and that these datasets vitalize machine learning education in underrepresented nations in the research community.</p>
sQTL catalog generated by sqtlseeker2-nf in the GTEx dataset
<p>sQTL catalog generated by <a href="https://github.com/guigolab/sqtlseeker2-nf">sqtlseeker2-nf</a> in the <a href="https://www.gtexportal.org/home/">GTEx</a> dataset, as described in the publication<em> Identification and analysis of splicing quantitative trait loci across multiple tissues in the human genome</em> by Garrido-Martín et al. in Nat. Commun. (<a href="https://doi.org/10.1038/s41467-020-20578-2">https://doi.org/10.1038/s41467-020-20578-2</a>). It contains the results of the sqtlseeker2-nf pipeline, run using different input data: i) RSEM transcript quantifications and genotype data from GTEx V7, ii) LeafCutter intron excision ratios and genotype data from GTEx V7 and iii) RSEM transcript quantifications and genotype data from GTEx V8. For each GTEx tissue, it includes summary statistics corresponding to all the tests performed (both nominal and permutation passes) and the significant sQTLs identified.</p> <p> </p>
Dataset and model weights for paper "Multi-Referenced Training for Dialogue Response Generation"
<p>dataset.txt: JSON file of dataset which has multiple references in each training sample</p> <p>gpt2_medium.floor_rel.seed_42.20200407-135521.model.pt: model weights of the finetuned GPT-2 used as a seq-level teacher model</p> <p>gpt2_small.floor_none.seed_42.20200407-133531.model.pt: model weights of the finetuned GPT-2 used as a token-level teacher model (because medium GPT-2 is too large and too slow for token-level KD)</p> <p>roberta_large.floor_none.seed_42.2020-04-01-12_28_50.semi.supervised_by_overall.model.pt: model weights of Roberta-eval for evaluating</p> <p>mturk_results.json: JSON file of Amazon MTurk human evaluation results</p>
Data from: Angiosperm phylogeny based on 18S/26S rDNA sequence data: constructing a large dataset using next-generation sequence data
The utility of 18S and 26S in broad phylogenetic analyses has been much maligned due in large part to the low signal in both genes. However, few analyses have employed complete 26S rDNA sequences over a broad range of taxa, and most alignments of the two genes are done de novo, without taking into account the secondary structure of the two rRNA genes. Here we mine next-generation sequence data to compile large matrices (429 taxa) of complete 18S + 26S gene sequences, and we compare both de novo alignment methods with curated alignments done by eye that take into account secondary structure and hard-to-align regions (profile alignments). The combined 18S + 26S topology is overall very similar to recently published gene trees for the angiosperms based on three or more genes. Overall support for the backbone or framework of the combined tree is low (bootstrap support below 50%). Few major clades have bootstrap support above 50%. Most well-supported clades are tip clades (families and orders sensu APG III 2009). Importantly, the 18S + 26S rDNA topology is consistent with current estimates of relationships: the basalmost angiosperms are recovered (Amborellaceae, Nymphaeales, Austrobaileyales), as are most major clades, including Mesangiospermae, eudicots (Eudicotyledoneae sensu Cantino et al. 2007), core eudicots (Gunneridae sensu Cantino et al. 2007), rosids (Rosidae sensu Cantino et al. 2007), asterids (Asteridae sensu Cantino et al. 2007), and Caryophyllales. Most clades recognized at the ordinal level (sensu APG III 2009) are also recovered. However, there are also some unusual placements in the 18S + 26S topology, but none of these receives bootstrap support above 50%. The profile and de novo alignments gave very similar topologies. 18S + 26S trees remain useful sources of data in large combined analyses. This is the first time a large data set of complete 26S gene sequences has been employed at this scale; this gene in particular proved to be useful phylogenetically. Targeted sequencing of 18S/26S rDNA is not advocated here, but given that these regions provide useful phylogenetic information and are abundant in next-generation sequencing runs, we suggest that the data be used rather than discarded.
Dataset of point-clouds generated by UAV LiDAR scanning of barley field-plot trial.
<p>Point clouds from eight scanning flights, scanned by UAV LiDAR system VUX-UAV1 (Ricopter GmBh), provided in forms of separated .las files. Dataset represent raw data that were further processed by software ALFA to extract height of canopies in particular field plots. Software serves for monitoring of plant growth dynamics in agricultural or breeding filed-plot trials.</p>
Modelling to Generate Alterantives at Ruhr-University Bochum Backbone dataset
<p>Datasets for the Backbone energy model of the Ruhr-University Bochum:</p> <ul> <li>2019 data for validation, demand time series are not published</li> <li>2030 and 2045 data used for Modelling to generate alternatives, demand time series are aggregated</li> </ul> <p>Backbone version 1.4. was used.</p>
Dataset for "The effect of parallel electron plateau on banded chorus generation: 1-D PIC simulations in mirror geometry"
<p>Dataset for a paper "The effect of parallel electron plateau on banded chorus generation: 1-D PIC simulations in mirror geometry"</p>
Dataset and Software for High-resolution mantle flow models reveal importance of plate boundary geometry and slab pull forces on generating tectonic plate motions
<p>This repository contains the plugin and dataset used to setup models in the manuscript: "High-resolution mantle flow models reveal importance of plate boundary geometry and slab pull forces on generating tectonic plate motions". It also includes the corresponding versions of the geodynamics software ASPECT and Worldbuilder which were used to develop the models in the paper.</p> <p>The material model plugin used in our models is in the "plugins" folder. The "models" folder contains the reference input parameter file described in the paper. All other model configurations shown in the paper can be obtained by modifying this parameter file. The input files used to set up the models are in the respective folder. Additionally, the Jupyter notebook used to compute residuals of our models is provided in the "scripts" folder.</p>
Dataset of "Supersonic: Learning to Generate Source Code Optimizations in C/C++"
<p>This is the dataset of "<a href="https://arxiv.org/abs/2309.14846">Supersonic: Learning to Generate Source Code Optimizations in C/C++</a>".</p>
Dataset and Checkpoints of MolEdit: In-silico 3D Molecular Editing through Physics-Informed and Peference-Aligned Generative Foundation Models
<p>This repository hosts the pre-trained checkpoints and data utilized for the development of MolEdit. The corresponding research paper, titled "<em>In-silico</em> 3D Molecular Editing through Physics-Informed and Peference-Aligned Generative Foundation Models" (an early version is preprinted at <a href="doi.org/10.26434/chemrxiv-2023-j2n6l-v2">doi:10.26434/chemrxiv-2023-j2n6l-v2</a>) and the corresponding GitHub repository of <a href="https://github.com/issacAzazel/MolEdit">MolEdit</a> details the application and validation of these checkpoints and data. For further information, please refer to the README.md file contained within this repository.</p>
Next generation sequencing quantitatively analyzes transcriptomes of primary CLL cells tranfected with non-specific control (NSC) or ZAP-70 siRNA (-anti-IgM dataset)
GEO Series GSE149020. Homo sapiens. 11 samples. Type: Expression profiling by high throughput sequencing.
Data from: Angiosperm phylogeny based on 18S/26S rDNA sequence data: constructing a large dataset using next-generation sequence data
Open the record for dataset details and reuse information.
Dataset to aid educators in using Generative AI in software design
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.