Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
447
datasets available to search
ShareScore release 0.7.1
Dataset results
447 results for “Binning”
September 2002 bin-averaged CTD profiles for the Georgia Coastal Ecosystems Doboy Sound transect
Two hydrographic surveys were performed on September 16, 2002, along an east-to-west transect through Doboy Sound near Sapelo Island, Georgia (Doboy Sound Transect, GCE-DB). Vertical CTD profiles were collected at approximately 2km intervals from -2km to 10km along the transect during low tide and high tide conditions. Conductivity, temperature, pressure and optical backscatter were measured, and depth, salinity and sigma-t were calculated for each profile. Data values collected on the upcast were deleted, and the remaining data were averaged within 0.5m depth bins and interpolated to produce a smooth profile for contouring. This data set was collected as part of the Georgia Coastal Ecosystems LTER quarterly hydrographic monitoring program.
September 2002 bin-averaged CTD profiles for the Georgia Coastal Ecosystems Inner Marsh transect
One hydrographic survey was performed on September 18, 2002, along the Darien River, North River, and north channel of the Altamaha River (Inner Marsh Transect, GCE-IM). Vertical CTD profiles were collected at various nominal stations along the transect during a high tidal regime. Conductivity, temperature, pressure and optical backscatter were measured, and depth, salinity and sigma-t were calculated for each profile. Data values collected on the upcast were deleted, and the remaining data were averaged within 0.5m depth bins and interpolated to produce a smooth profile for contouring. This data set was collected as part of the Georgia Coastal Ecosystems LTER quarterly hydrographic monitoring program.
September 2002 bin-averaged CTD profiles for the Georgia Coastal Ecosystems Intracoastal Waterway transect
Three hydrographic surveys were performed from September 16 to September 18, 2002, along the Intracoastal Waterway between the Altamaha River south of Wolf Island and Sapelo Sound (Intracoastal Waterway Transect, GCE-IC). Vertical CTD profiles were collected at various intervals from 0km to 29km along the transect during various tidal regimes. Conductivity, temperature, pressure and optical backscatter were measured, and depth, salinity and sigma-t were calculated for each profile. Data values collected on the upcast were deleted, and the remaining data were averaged within 0.5m depth bins and interpolated to produce a smooth profile for contouring. This data set was collected as part of the Georgia Coastal Ecosystems LTER quarterly hydrographic monitoring program.
September 2002 bin-averaged CTD profiles for the Georgia Coastal Ecosystems Sapelo River transect
Three hydrographic surveys were performed on September 17, 2002, along a transect from Sapelo Sound up the Sapelo River to Eulonia, Georgia (Sapelo River Transect, GCE-SP). Vertical CTD profiles were collected at approximately 2km intervals from 0km to 36km during low and high tidal regimes. Conductivity, temperature, pressure and optical backscatter were measured, and depth, salinity and sigma-t were calculated for each profile. Data values collected on the upcast were deleted, and the remaining data were averaged within 0.5m depth bins and interpolated to produce a smooth profile for contouring. This data set was collected as part of the Georgia Coastal Ecosystems LTER quarterly hydrographic monitoring program.
September 2002 bin-averaged CTD profiles for the Georgia Coastal Ecosystems Altamaha River transect
Four hydrographic surveys were performed on September 18, 2002, along an east-to-west transect up the Altamaha River in Georgia (Altamaha River Transect, GCE-AL). Vertical CTD profiles were collected at 1-2km intervals from 2km east of the line of demarcation (station -02) to 28km upriver along the transect during low and high tidal regimes. Conductivity, temperature, pressure and optical backscatter were measured, and depth, salinity and sigma-t were calculated for each profile. Data values collected on the upcast were deleted, and the remaining data were averaged within 0.5m depth bins and interpolated to produce a smooth profile for contouring. This data set was collected as part of the Georgia Coastal Ecosystems LTER quarterly hydrographic monitoring program.
A comprehensive evaluation of binning methods to recover human gut microbial species from a non-redundant reference gene catalog - Supporting Data
<p><strong>Description </strong></p> <p>The following files are available : </p> <ul> <li>Simulated non-redundant Gene Catalog (SGC) composed of 128267 genes;</li> <li>Gene abundance profiles across 40 samples: raw read counts, gene length normalized base counts, depth file computed by the jgi_summarize_bam_contig_depth script provided by MetaBAT;</li> <li>Gold Standard (GS) and Gold Standard Single Assignment (GS_SA) binning results;</li> <li>Binning results obtained on the SGC with nine binning methods: MSPminer, MGS-canopy, DAS Tool, MaxBin2, MetaBAT2, SolidBin, CONCOCT, COCACOLA and MyCC.</li> </ul> <p><strong>License</strong></p> <p>These files are licensed under a <a href="http://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</a>.</p>
Mandalay, Myanmar. Shwe in Bin Monastery, south side.
<p>Mandalay, Myanmar. Shwe in Bin Monastery (Burmese: ရွှေအင်ပင်ကျောင်း), south side. Photograph from circa 1910.</p>
Dust mass fractions for dust bins covering the 0.1 to 100 µm size range
<p>This is a file containing the dust mass fractions for dust bins covering the 0.1 to 100 µm size range calculated from the BFT-supercoarse emitted dust PSD parameterization (Meng et al., 2021 GRL submitted). </p>
High-resolution BIOCLIM and ENVIREM grids for Europe in consecutive 100-year bins spanning the last 21,000 years
<p>Here, I provide a dataset of gridded climatic variables at a spatial resolution of 30 arc-seconds for 210 consecutive 100-year bins spanning the period from 21,000 to 0 BP. The dataset includes 19 bioclimatic and 16 ENVIREM variables (described by Title & Bemmels, 2018) commonly used in species distribution modelling. It covers the European continent and adjacent regions within the following boundaries: 32.5°W–70°E and 32.5°N–82.5°N.</p> <p>For each 100-year bin, bioclimatic and ENVIREM variables were calculated based on the downscaled and debiased monthly temperature and precipitation simulations of the Community Climate System Model version 3 (CCSM3; Collins et al., 2006) as provided by the PaleoView software (Fordham et al., 2017). The downscaling procedure was based on the delta-change method (Ramirez Villejas & Jarvis, 2010). As a baseline climatic data, I used monthly temperature and precipitation grids from the CHELSA database for 1940–1989 (Karger et al., 2017).</p> <p>A detailed description of the dataset, including the downscaling method applied, can be found in the Technical specification attached to this dataset.</p>
CAT 4.6 taxonomic binning of the CAMI 2 Mouse Gut Toy data set, gold standard pooled assembly
<p>Taxonomic binning of the gold standard pooled assembly<br> <strong>Software: </strong>CAT<br> <strong>SoftwareVersion: </strong>4.6<br> <strong>DataURL: </strong> https://data.cami-challenge.org/participate<br> <strong>SoftwareURL:</strong> https://github.com/dutilh/CAT<br> <strong>ReferenceDatabase:</strong> prebuilt 2018-12-12<br> <strong>Taxonomy:</strong> NCBI 2018-12-12<br> <strong>ShortReadsUsed:</strong> False<br> <strong>LongReadsUsed:</strong> False<br> <strong>CommandUsed:</strong> CAT contigs -c anonymous_gsa_pooled.fasta -d CAT_prepare_20181212/2018-12-12_CAT_database/ -t CAT_prepare_20181212/2018-12-12_taxonomy/ --tmpdir tmp --nproc 16</p>
PhyloPythiaS+ 1.4 taxonomic binning of the CAMI 2 Mouse Gut Toy data set, gold standard pooled assembly
<p>Taxonomic binning of the gold standard pooled assembly<br> <strong>Software: </strong>PhyloPythiaS+<br> <strong>SoftwareVersion: </strong>1.4<br> <strong>DataURL: </strong> https://data.cami-challenge.org/participate<br> <strong>SoftwareURL:</strong> https://github.com/algbioi/ppsp<br> <strong>DockerImage:</strong> cami/ppsp:1.4<br> <strong>IsBiobox:</strong> False<br> <strong>ReferenceDatabase:</strong> RefSeq 93, SILVA 132<br> <strong>Taxonomy:</strong> NCBI 2018-02-26<br> <strong>ShortReadsUsed:</strong> False<br> <strong>LongReadsUsed:</strong> False<br> <strong>CommandUsed:</strong> run_ppsp.py --pipelineDir ppsp_pipepline --inputFastaFile anonymous_gsa_pooled.fasta --databaseFile ncbi_taxonomy --refSeq refseq93 --s16Database SILVA_132 --mgDatabase reference_NCBI201502/mg5</p>
MetaBAT 2.12.1 genome binning of the CAMI 2 Mouse Gut Toy data set, samples 0-63, gold standard pooled assembly
Genome binning of the gold standard pooled assembly <br><strong>Software: </strong>MetaBAT<br><strong>SoftwareVersion: </strong>2.12.1<br><strong>DataURL: </strong> https://data.cami-challenge.org/participate<br><strong>SoftwareURL:</strong> https://bitbucket.org/berkeleylab/metabat<br><strong>ShortReadsUsed:</strong> True<br><strong>LongReadsUsed:</strong> False<br><strong>CommandUsed:</strong> bowtie2-build anonymous_gsa_pooled.fasta anonymous_gsa_pooled.fasta<br>for i in {0..63}; do bowtie2 -q --threads 30 --fr -x anonymous_gsa_pooled.fasta --interleaved sample_${i}/anonymous_reads.fq -S anonymous_reads_sample_${i}.sam ; done<br>for i in {0..63}; do samtools view -b sample_${i}.sam -o anonymous_reads_sample_${i}.bam & done<br>for i in {0..63}; do samtools sort anonymous_reads_sample_${i}.bam -o anonymous_reads_sample_${i}.sorted.bam ; done<br>for i in {0..63}; do samtools index anonymous_reads_sample_${i}.sorted.bam ; done<br>runMetaBat.sh -l anonymous_gsa_pooled.fasta anonymous_reads_sample_*.sorted.bam
Kraken 2.0.8 beta taxonomic binning of the CAMI 2 Mouse Gut Toy data set, gold standard pooled assembly
<p>Taxonomic binning of the gold standard pooled assembly<br> <strong>Software: </strong>Kraken<br> <strong>SoftwareVersion: </strong>2.0.8 beta<br> <strong>DataURL: </strong> https://data.cami-challenge.org/participate<br> <strong>SoftwareURL:</strong> https://ccb.jhu.edu/software/kraken2/<br> <strong>ReferenceDatabase:</strong> built 2019-05-22<br> <strong>Taxonomy:</strong> NCBI 2019-05-22<br> <strong>ShortReadsUsed:</strong> False<br> <strong>LongReadsUsed:</strong> False<br> <strong>CommandUsed:</strong> kraken2-build --standard --db kraken2db_std --use-ftp<br> kraken2 --db kraken2db_std --threads 16 --output 19122017_mousegut_scaffolds.kraken --report 19122017_mousegut_scaffolds.kreport anonymous_gsa_pooled.fasta<br> cat 19122017_mousegut_scaffolds | awk '{print $2 "\t" $3}' > 19122017_mousegut_scaffolds.cami</p>
Dataset for "Atmospheric oxygen isotopic fractionation in clouds: a bin–resolved microphysics model approach"
<p>This dataset contains the raw model outputs for the paper titled "Atmospheric oxygen isotopic fractionation in clouds: a bin–resolved microphysics model approach" by Thibault Hiron and Andrea Flossmann.</p> <p>The structure of the data and the explaination for the filenames are to be found in the ReadMe.txt file.</p>
Data release: Searching for binary black hole sub-populations in gravitational wave data using binned Gaussian processes
<p>The data required to reproduce the analyses of "Searching for binary black hole sub-populations in gravitational wave data using binned Gaussian processes" (<a href="https://arxiv.org/abs/2404.03166" target="_blank" rel="noopener">arxiv:2404.03166</a>). The main inference code can be found at <a href="https://github.com/AnaryaRay1/gppop/tree/spin-dev" target="_blank" rel="noopener">https://github.com/AnaryaRay1/gppop/tree/spin-dev </a> (commit: <a href="https://github.com/AnaryaRay1/gppop/commit/ee5ffc421e2c96eeed15a0e0d3839da42b982842">ee5ffc</a>). To reproduce the analyses, follow the instructions at <a href="https://github.com/AnaryaRay1/bbh-subpopulations-scripts">https://github.com/AnaryaRay1/bbh-subpopulations-scripts</a> (commit <a href="https://github.com/AnaryaRay1/bbh-subpopulations-scripts/commit/de88f931d8c1a2cb31ad2fa9d6fdf9a5a00a3c3b">de88f93</a>). Frozen versions of these repositories that were used to generate all the results are available as part of this data release, in the files "gppop_spin_dev_ee5ffc421.tar.gz" and "bbh-subpopulations-scripts_de88f931.tar.gz" respectively.</p>
Source data for gamete binning for an autotetraploid potato cultivar Otava
<p>Here we provide the source data for analyzing a highly heterozygous autotetraploid potato cultivar 'Otava'.</p>
"Genome binning of viral entities from bulk metagenomics data" - CAMISIM simulated datasets and genomes
<p><strong>Genome binning of viral entities from bulk metagenomics data</strong></p> <p> </p> <p><strong>Authors</strong></p> <p><strong>Joachim Johansen1,2, Damian R. Plichta2, Jakob Nybo Nissen1,3, Marie Louise Jespersen1,4, Shiraz A. Shah5, Ling Deng6, Jakob Stokholm5,6, Hans Bisgaard5, Dennis Sandris Nielsen6, Søren Sørensen7, Simon Rasmussen1</strong></p> <p> </p> <p><strong>Affiliations</strong></p> <p>1 Novo Nordisk Foundation Center for Protein Research, Faculty of Health and Medical Sciences, University of Copenhagen, Copenhagen N, Denmark</p> <p>2 Infectious Disease and Microbiome Program, Broad Institute of MIT and Harvard, Cambridge, MA, USA</p> <p>3 Statens Serum Institut, Viral & Microbial Special diagnostics, Copenhagen, Denmark</p> <p>4 National Food Institute, Technical University of Denmark, Kongens Lyngby, Denmark</p> <p>5 Copenhagen Prospective Studies on Asthma in Childhood (COPSAC), Herlev and Gentofte Hospital, University of Copenhagen, Copenhagen, Denmark</p> <p>6 Section of Food Microbiology and Fermentation, Department of Food Science, Faculty of Science, University of Copenhagen, Copenhagen, Denmark</p> <p>7 Section of Microbiology, Department of Biology, University of Copenhagen, Copenhagen, Denmark</p> <p><strong>Methods description</strong></p> <p>We compared the viral binning performance of VAMB and MetaBAT2 using the official CAMI consortium method to create assemblies and metagenome profiles. To this end we generated 3 different metagenome compositions with up to 308 reference genomes; one mixed with bacteria, plasmids and viruses to test binning in complex samples i.e. high diversity (1), one with only crass-like viruses to test binning with highly similar viruses i.e. high relatedness (2) and a set of small-viruses (<6,000 bp) including members of the Microviridae family to address the bias of size (3). Bacterial genomes were gathered from NCBIs refseq genome repository 2021, plasmids from the PLSDB database (v. 2021_06_23) and viral genomes from the recent MGV database. </p> <p>Dataset A contained a mixture of bacteria (N=8), plasmids (N=20) and viruses (N=280) to test binning in complex samples, i.e. high diversity. Dataset B contained only crass-like viruses (N=80) to test binning with highly similar viruses i.e. high relatedness. Dataset C contained small-viruses (N=50, <6,000 bp) of the Microviridae family to address the bias of size. Bacterial genomes were sampled from the Refseq genome repository 2021, plasmids from the PLSDB database and viral genomes from the recent MGV database (Nayfach, et al. Nature Microbiology 2021).</p> <p> </p> <p> </p>
DMSP F15/F16/F18 2011-2014 Binned Poynting Flux Data
<p>This dataset, created using the <a href="https://github.com/lkilcommons/esabin">esabin</a> Python library, comprises 9 spacecraft-years of electrodynamics measurements from Defense Meteorology Satellite Program F15, F16 and F18 spacecraft. <strong>Users of this data should prefer files which do not contain 'idm_only' in their filenames, as these do not use both components of DMSP ion drift vector measurements.</strong> They are included here to support reproducibility of a upcoming publication.</p> <p>The HDF5 files herein organize this data in equal area bins in magnetic coordinates. Each Group in the files represents one bin. Each Dataset contains the data for one crossing of that bin by a DMSP spacecraft. The name of each Dataset is the approximate time of the crossing as Julian date. Hourly NASA OMNIWeb solar wind and geomagnetic activity parameters are included as attributes for each Dataset.</p> <p>The electrodynamic parameters herein are magnetic perturbation (dB), electric field (E) and ion drift velocity (V). These quantities are scaled from satellite altitude (~850 km) where they were observed to ionospheric altitude (~110 km) using Modified Magnetic Apex coordinates. In the filenames, 'e' represents magnetic eastward and 'n' represents magnetic northward. Also included are geomagnetic-main-field-aligned Poynting flux also scaled to 110km altitude (files with 'poynting' in the name).</p>
Text-fig. 9. a–d: Nyssidium orientale SAMYLINA, Partizansk, Starosuchan Formation, Barremian, a – spec. BIN 506/3749, general view of four fruits, b1 – spec. BIN 506/3749-1, holotype, b2 – spec. BIN 506/3749-2, c – spec. BIN 506/3749-6, d – spec. BIN 506/3749-5; e–f: Cercidiphyllum sujfunense KRASSILOV, Konstantinovka, Galenki Formation, early-middle Albian, e – spec. IBSS 11-135, fruit, f – spec. IBSS 11-134, leaf, holotype. Scale bar 5 mm in a, e, f and 2 mm in b–d. in Angiosperm Diversification In The Early Cretaceous Of Primorye, Far East Of Russia
Text-fig. 9. a–d: Nyssidium orientale SAMYLINA, Partizansk, Starosuchan Formation, Barremian, a – spec. BIN 506/3749, general view of four fruits, b1 – spec. BIN 506/3749-1, holotype, b2 – spec. BIN 506/3749-2, c – spec. BIN 506/3749-6, d – spec. BIN 506/3749-5; e–f: Cercidiphyllum sujfunense KRASSILOV, Konstantinovka, Galenki Formation, early-middle Albian, e – spec. IBSS 11-135, fruit, f – spec. IBSS 11-134, leaf, holotype. Scale bar 5 mm in a, e, f and 2 mm in b–d.
1H NMR spectra of commercial honey from 400 and 700 MHz spectrometers and tables with data after processing and binning
<p>Datasets contain the 1H NMR original raw spectral data (Bruker format) of commercial honey from 400 MHz and 700 MHz NMR spectrometers and Tables (.xlsx) with data after processing and binning.</p> <p> </p> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.