Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,569

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,569 results for “Reading”

Learn how ShareScore rates datasets ↗
OpenNeuro52/100

Component processes of word reading in adults and children

Open the record for dataset details and reuse information.

openCC0Jan 2021View details →
zenodo52/100

Simulated NGS read datasets for bacterial pathogenic potential prediction

<p>## Predicting pathogenic potentials from NGS reads: novel bacterial species</p> <p>This repository contains simulated Illumina&nbsp;read datasets for bacterial pathogenic potential prediction and associated metadata extracted from the IMG Database (https://img.jgi.doe.gov/). The reads are 250bp long and were simulated with Mason (https://www.seqan.de/apps/mason/) from genomes downloaded from NCBI. The training-validation-test split was done on the species level to ensure &quot;novelty&quot; of validation and test species. The training sets contain 10 million reads per class, validation sets - 1.25 million reads per class, and test sets - 1.25 million paired reads per class. Additional, imbalanced training sets contain 2.5 million &quot;nonpathogenic&quot; and 17.5 million &quot;pathogenic&quot; reads, keeping the mean covarage constant for all species. The temporal benchmark test set contains reads from 3 additional pathogenic species in the Pantoea genus.</p> <p>## Predicting pathogenic potentials from NGS reads: novel strains of known species</p> <p>The BacPaCS datasets contain reads simulated from the dataset compiled by Barash et al. (https://doi.org/10.1093/bioinformatics/bty928). It this case, the training-validation-test split was done on the strain&nbsp;level (so different strains of the same species may be present in all three sets).</p>

opencc-by-4.0Jan 2019View details →
OpenNeuro48/100

Emotion regulation in the Ageing Brain, University of Reading, BBSRC

Open the record for dataset details and reuse information.

openCC0Jan 2019View details →
zenodo48/100

Data for: What's in a game: Video game visual-spatial demand location exhibits a double dissociation with reading speed

<p>The aggregate data in these datasets were used in analyses for &quot;What&rsquo;s in a game: Video game visual-spatial demand location exhibits a double dissociation with reading speed&quot;.</p>

opencc-by-4.0Mar 2022View details →
zenodo48/100

Reading hyperlinks

<p>This is a eye-tracking dataset in the BIDS format (http://bids.neuroimaging.io/).</p> <p>Please find the details on the study here:</p> <p>Gagl B. (2016) Blue hypertext is a good design decision: no perceptual disadvantage in reading and successful highlighting of relevant information. PeerJ 4:e2467 <a href="https://doi.org/10.7717/peerj.2467">https://doi.org/10.7717/peerj.2467</a></p> <p>and all analysis scripts and a previous version of the data here:</p> <p>Gagl, B. (2016, April 17). Reading Hypertext. Retrieved from <a href="https://osf.io/8c57w/">https://osf.io/8c57w/.</a></p> <p>&nbsp;</p>

opencc-by-4.0Apr 2018View details →
zenodo48/100

Genomes plasmids MDR B. fragilis ONT sequence read files in fastq format

<p>Supporting data for the manuscript <em>Complete genome assembly of clinical multidrug resistant Bacteroides fragilis isolates enables comprehensive identification of antimicrobial resistance genes and plasmids.</em></p> <p>Oxford Nanopore reads demultiplexed with <a href="https://www.google.com/url?sa=t&amp;rct=j&amp;q=&amp;esrc=s&amp;source=web&amp;cd=1&amp;cad=rja&amp;uact=8&amp;ved=2ahUKEwjNjqL5tY7iAhUawMQBHZHfDasQFjAAegQIAhAB&amp;url=https%3A%2F%2Fgithub.com%2Frrwick%2FDeepbinner&amp;usg=AOvVaw0wikvIUagLuFV38CwKZtia">Deepbinner</a> v0.2.0 and base-called (with demultiplexing) using Albacore v2.3.3. Barcodes and adapters were removed with <a href="https://github.com/rrwick/Porechop">Porechop</a> v0.2.4 with the --discard_middle option.</p> <p>Data from each isolate was produced from two runs per isolate. Data for the individual runs are included here. They can easily be concatenated eg with cat. Runs are named TVS_01,. TVS_02, TVS_03 and TVS_04.</p> <p>Fast5 (only demultiplexed with deepbinner and basecalled with albacore) as well as illumina reads and genome assemblies can be found via the NCBI bioproject accessions:</p> <p>Isolates, NCBI bioproject accession no:</p> <p>CCUG4856T,&nbsp;<a href="http://www.ncbi.nlm.nih.gov/bioproject/PRJNA525024">PRJNA525024</a></p> <p>BFO17,&nbsp;<a href="http://www.ncbi.nlm.nih.gov/bioproject/PRJNA244943">PRJNA244943</a></p> <p>BFO18,&nbsp;<a href="http://www.ncbi.nlm.nih.gov/bioproject/PRJNA244944">PRJNA244944</a></p> <p>S01,&nbsp;<a href="http://www.ncbi.nlm.nih.gov/bioproject/PRJNA244942">PRJNA244942</a></p> <p>BFO42,&nbsp;<a href="http://www.ncbi.nlm.nih.gov/bioproject/PRJNA253771">PRJNA253771</a></p> <p>BFO67,&nbsp;<a href="http://www.ncbi.nlm.nih.gov/bioproject/PRJNA254401">PRJNA254401</a></p> <p>BFO85,&nbsp;<a href="http://www.ncbi.nlm.nih.gov/bioproject/PRJNA254455">PRJNA254455</a></p> <p>&nbsp;</p> <p><strong>md5sum&#39;s (also found in the file md5.md5):</strong></p> <p>135d0570a1e49e25c8fde59f321cca68&nbsp; BFO17_TVS_03_99377.barcode02_trimmed.fastq.gz<br> 3c6ca800a1f735937c0cffccd263fb82&nbsp; BFO18_TVS_01_97673.barcode03_trimmed.fastq.gz<br> dc37823950f529d3a859b7a7c514af8c&nbsp; BFO18_TVS_03_99377.barcode03_trimmed.fastq.gz<br> abd707404f9ebbc38652e45ed378ca21&nbsp; BFO42_TVS_02.barcode10_trimmed.fastq.gz<br> 9761e9ab082e276624c5778dcbeaffd2&nbsp; BFO42_TVS_04.barcode10_trimmed.fastq.gz<br> ffc27009c0f7fead1af84ce045dbe5f3&nbsp; BFO67_TVS_02.barcode09_trimmed.fastq.gz<br> 06c363a7feeeb88f9d195b3769b37f2b&nbsp; BFO67_TVS_04.barcode09_trimmed.fastq.gz<br> 5553c95cc98f4b9d4cfb38c4f8f8f037&nbsp; BFO85_TVS_02.barcode08_trimmed.fastq.gz<br> b1a8013cba7a079cee6d3bdd6cd97ff2&nbsp; BFO85_TVS_04.barcode08_trimmed.fastq.gz<br> 8169225219a5fb20935d5f0304aa80c5&nbsp; CCUG4856T_TVS_01_97673.barcode01_trimmed.fastq.gz<br> a509b0ae912a798a91e677795198c1c6&nbsp; CCUG5846T_TVS_03_99377.barcode01_trimmed.fastq.gz<br> 8a9d2eb8b626a6e87ed267d31aa220e3&nbsp; S01_TVS_01_97673.barcode04_trimmed.fastq.gz<br> 69e239a17becc25a4f103dfc3bf5886a&nbsp; S01_TVS_03_99377.barcode04_trimmed.fastq.gz</p>

opencc-by-4.0May 2019View details →
zenodo48/100

Metabarcoding data (number of reads per operational taxonomic unit) from a monitoring study of sandy beach meiofauna before and after sand nourishment (Ahrenshoop, Baltic Sea)

<p>We provide metabarcoding data (number of reads per operational taxonomic unit, OTU) determined from sediment samples collected on the sandy-beach water line of Ahrenshoop (Baltic Sea). Five sampling stations lay within the zone impacted by the sand nourishment between the boundary of the nature reserve in the north east and a site just north of the breakwater (AH01&ndash;AH05). An unaffected reference station was located south of Ahrenshoop (close to Niehagen) at the end of the road Pappelallee (PAP). Samples were collected at four dates. The first sampling was carried out before the sand nourishment took place (T0: 14 and 16 September 2021). Three samplings were realised after the impact: T1 (23 March 2022), T2 (27 September 2022), and T3 (28 March 2023). Latitude and longitude of each sampling location per station were recorded at each sampling date using a hand-held GPS application on a mobile phone. At the stations sampling locations varied over time. Prior to the sand nourishment the beach was narrow due to sand erosion in previous years. After the nourishment the additional extent of the beach was approximately 40 m at sampling date T1. Subsequently, progressive sand erosion forced the sampling locations (situated at the water line) further inland at T2 and T3.<br>Samples were taken from the beach-water interface (water line) in the middle of the area between two groynes. Plexiglass cores (inner core diameter 5.4 cm) were inserted vertically into the sediment down to 15 cm depth. Each core was sliced in 5 cm-layers (0&ndash;5, 5&ndash;10 and 10&ndash;15 cm). Sediment horizons were preserved in 96&ndash;99% ethanol. <br>Three cores (2 cores at T0) per sampling date were taken for metabarcoding analyses. The organisms were extracted by decantation over a 32-&mu;m sieve.&nbsp;Genomic DNA was extracted from the filters using the DNeasy PowerSoil pro kit (Qiagen). Realtime-PCR was performed to amplify V1&amp;V2, two hypervariable regions of 18S rDNA gene. The sequencing run was performed using the MiSeq Reagent Nanokit v2 (250 cycles paired end) on an Illumina MiSeq platform at the DZMB Metabarcoding lab in Wilhelmshaven, Germany. High-resolution amplicon sequence variants (ASVs) were obtained and compared to the NCBI database to assign taxonomic information to each ASV. The target meiofauna ASVs were further classified into operational taxonomic units (OTUs) with a 3% cut-off threshold using the statistical software R.</p> <p>Here, we present two Tables (as xlsx and tab-delimited files):<br>(1) the taxonomic description of the 843 OTUs and their assigned ID number;<br>(2) the number of reads per OTU per sample (including metadata for each sample: event; date; latitude; longitude; station, core and sample ID; sediment depth).</p> <p>The metabarcoding data are part of a larger ecological study on the influence of sand nourishment on meiofauna communities, which included grain-size and meiofauna abundances&nbsp;(see &ldquo;related works&rdquo;).</p> <p><strong>Comment: </strong>Our study is related to but not funded by the project ECAS Baltic: Strategies of ecosystem-friendly coastal protection and ecosystem-supporting coastal adaptation for the German Baltic Sea Coast <a href="https://deutsche-kuestenforschung.de/ecas-baltic.html">https://deutsche-kuestenforschung.de/ecas-baltic.html</a></p>

opencc-by-4.0Jul 2024View details →
zenodo48/100

Behavioral and fMRI Data: Nurturing the reading brain: Home literacy practices are associated with children's neural response to printed words through vocabulary skills

<p>This is the behavioral and fMRI dataset described in &quot;Nurturing the reading brain: &nbsp;Home literacy practices are associated with children&rsquo;s neural response to printed words through vocabulary skills&quot;.&nbsp;</p> <p>Because of anonymization concerns within&nbsp;the framework of EU privacy regulations (<a href="https://gdpr-info.eu">GDPR</a>), we cannot provide raw MRI data. Therefore, the fMRI data consists of individual&nbsp;pre-processed volumes, normalized into the MNI&nbsp;template (see paper for details about the preprocessing pipeline). Anonymized behavioral data and first level analyses are also provided for each participant (SPM.mat file as well as beta, con, spmT, RPV and ResMS&nbsp;files). Note that the dataset&nbsp;also include runs and GLM results for a third task (Dots) that was not analyzed in the paper. Finally, the <a href="https://www.psychopy.org">PsychoPy</a> implementation of the tasks is also provided. If you have any questions, please send an email to jerome.prado [at] univ-lyon1.fr.&nbsp;</p> <p><strong>IMPORTANT:</strong></p> <p>In accordance with EU privacy regulations, we ask that you sign and return a Data Use Agreement (DUA) before downloading the data. You can download the DUA&nbsp;<a href="https://zenodo.org/record/4965716/files/DUA.pdf?download=1">here</a>. Please, sign it and send it to jerome.prado [at] univ-lyon1.fr.</p>

opencc-by-4.0Jul 2021View details →
edi48/100

Catalog of GenBank sequence read archive (SRA) entries of 16S and 18S rRNA genes from bacterial and protistan planktonic communities along the Eastern Beaufort Sea coast, North Slope, Alaska, 2011-2013

Microbial communities in the coastal Arctic Ocean experience extreme variability in organic matter and inorganic nutrients driven by seasonal shifts in sea ice extent and freshwater inputs. Lagoons border more than half of the Beaufort Sea coast and provide important habitats for migratory fish and seabirds; yet, little is known about the planktonic food webs supporting these higher trophic levels. To investigate seasonal changes in bacterial and protistan planktonic communities, amplicon sequences of 16S and 18S rRNA genes were generated from samples collected during periods of ice-cover (April), ice break-up (June), and open water (August) from shallow lagoons along the eastern Alaska Beaufort Sea coast from 2011 through 2013. This data package catalogs sequence read archive (SRA) entries available through GenBank BioProject PRJNA530074 at https://www.ncbi.nlm.nih.gov/bioproject/PRJNA530074. This data package is associated with the following publication: Kellogg CTE, McClelland JW, Dunton KH and Crump BC (2019) Strong Seasonality in Arctic Estuarine Microbial Food Webs. Front. Microbiol. 10:2628. doi: 10.3389/fmicb.2019.02628 Environmental variables (physiochemical data from YSI and HOBO data loggers, as well as organic matter analysis and stable isotope data from discrete water samples) associated with this genomic dataset are available from the Arctic Data Center: Kenneth Dunton, Byron Crump, and James McClelland. Physical, chemical, and biological data from lagoons and open coastal waters in the nearshore environment of the eastern Alaska Beaufort Sea, 2011-2013. Arctic Data Center. doi:10.18739/A2DG13. To join the two datasets together, please use the provided site codes (column "site_name" here) and collection dates (column "collection_date" here) in each dataset. Note that the site codes in this package are without hyphens (e.g. JAA) while site codes in the above environmental data package have hyphens (e.g. JA-A). Instead of citing this package which is jus

openCC0Jan 2020View details →
edi48/100

Catalog of GenBank sequence read archive (SRA) entries of metagenomic DNA sequence analyses of bacterial and archaeal water column communities along the Eastern Beaufort Sea coast, North Slope, Alaska, 2012

In contrast to temperate systems, Arctic lagoons that span the Alaska Beaufort Sea coast face extreme seasonality. Nine months of ice cover up to ∼1.7 m thick is followed by a spring thaw that introduces an enormous pulse of freshwater, nutrients, and organic matter into these lagoons over a relatively brief 2–3 week period. Prokaryotic communities link these subsidies to lagoon food webs through nutrient uptake, heterotrophic production, and other biogeochemical processes, but little is known about how the genomic capabilities of these communities respond to seasonal variability. This study characterizes the metabolic capabilities of microbial communities across three seasons in two lagoons and one open coastal site along the eastern Alaska Beaufort Sea coast. We used metagenomic DNA sequence data of bacterial and archaeal water column communities to identify genes of relevant biogeochemical pathways. This data package catalogs sequence read archive (SRA) entries available through GenBank BioProject PRJNA642637 at https://www.ncbi.nlm.nih.gov/bioproject/PRJNA642637. This data package is associated with the following publication: Baker, Kristina D., Colleen T. E. Kellogg, James W. McClelland, Kenneth H. Dunton, and Byron C. Crump. “The Genomic Capabilities of Microbial Communities Track Seasonal Variation in Environmental Conditions of Arctic Lagoons.” Frontiers in Microbiology 12 (2021). https://doi.org/10.3389/fmicb.2021.601901. Environmental variables (physiochemical data from YSI and HOBO data loggers, as well as organic matter analysis and stable isotope data from discrete water samples) associated with this genomic dataset are available from the Arctic Data Center: Kenneth Dunton, Byron Crump, and James McClelland. Physical, chemical, and biological data from lagoons and open coastal waters in the nearshore environment of the eastern Alaska Beaufort Sea, 2011-2013. Arctic Data Center. doi:10.18739/A2DG13. To join the two datasets together, please use the provi

openCC0Apr 2021View details →
zenodo44/100

Raw read counts and phased SNP counts for every single cell in the sequencing datasets of the breast cancer patient S1 from "Characterizing allele- and haplotype-specific copy numbers in single cells with CHISEL"

<p>This dataset contains the raw read counts and phased SNP counts&nbsp;for every single cell in the sequencing datasets of breast cancer patient S1 from &ldquo;Characterizing allele- and haplotype-specific copy numbers in single cells with CHISEL&rdquo; [Zaccaria &amp; Raphael, 2020]. These data enable the full reproduction of all the results in the related manuscript for breast cancer patient S1. Specifically, the data are provided in two files for every dataset <em>DAT</em> of patient S1 with the following format:</p> <ol> <li><em>DAT.raw</em>_<em>read</em>_<em>counts.bed.gz&nbsp;</em>is a multi-cell BED file containing the raw read counts in the following fields: <ul> <li>CHROMOSOME: the name of a human chromosome</li> <li>START: the starting genomic position of a genomic bin in the chromosome</li> <li>END: the ending genomic position of the genomic bin in the chromosome</li> <li>CELL: the cell barcode that uniquely identifies a cell</li> <li>NORMAL: the raw read count&nbsp;for the specified bin from a matched-normal sample</li> <li>COUNT: the raw read count&nbsp;for the specified bin in the specified cell</li> <li>RDR: the estimated read-depth ratio for the specified bin in the specified cell</li> </ul> </li> <li><em>DAT.phased</em>_<em>snps</em>_<em>counts.pos.gz&nbsp;</em>is a multi-cell POS file containing the phased SNP counts in the following fields: <ul> <li>CHROMOSOME: the name of a human chromosome</li> <li>POS: the genomic position in the chromosome of a germline SNP</li> <li>CELL: the cell barcode that uniquely identifies a cell</li> <li>COUNT_HAPLOTYPE_A: the count of reads that cover&nbsp;the SNP and that belong to haplotype A in the specified cell</li> <li>COUNT_HAPLOTYPE_B: the count of reads that cover&nbsp;the SNP and that belong to haplotype B in the specified cell</li> </ul> </li> </ol> <p>All the files have been compressed using standard <em>gzip</em>.</p>

opencc-by-4.0May 2020View details →
zenodo44/100

Raw read counts and phased SNP counts for every single cell in the sequencing datasets of the breast cancer patient S0 from "Characterizing allele- and haplotype-specific copy numbers in single cells with CHISEL"

<p>This dataset contains the raw read counts and phased SNP counts&nbsp;for every single cell in the sequencing datasets of breast cancer patient S0 from &ldquo;Characterizing allele- and haplotype-specific copy numbers in single cells with CHISEL&rdquo; [Zaccaria &amp; Raphael, 2020]. These data enable the full reproduction of all the results in the related manuscript for breast cancer patient S0. Specifically, the data are provided in two files for every dataset <em>DAT</em> of patient S0 with the following format:</p> <ol> <li><em>DAT.raw</em>_<em>read</em>_<em>counts.bed.gz&nbsp;</em>is a multi-cell BED file containing the raw read counts in the following fields: <ul> <li>CHROMOSOME: the name of a human chromosome</li> <li>START: the starting genomic position of a genomic bin in the chromosome</li> <li>END: the ending genomic position of the genomic bin in the chromosome</li> <li>CELL: the cell barcode that uniquely identifies a cell</li> <li>NORMAL: the raw read count&nbsp;for the specified bin from a matched-normal sample</li> <li>COUNT: the raw read count&nbsp;for the specified bin in the specified cell</li> <li>RDR: the estimated read-depth ratio for the specified bin in the specified cell</li> </ul> </li> <li><em>DAT.phased</em>_<em>snps</em>_<em>counts.pos.gz&nbsp;</em>is a multi-cell POS file containing the phased SNP counts in the following fields: <ul> <li>CHROMOSOME: the name of a human chromosome</li> <li>POS: the genomic position in the chromosome of a germline SNP</li> <li>CELL: the cell barcode that uniquely identifies a cell</li> <li>COUNT_HAPLOTYPE_A: the count of reads that cover&nbsp;the SNP and that belong to haplotype A in the specified cell</li> <li>COUNT_HAPLOTYPE_B: the count of reads that cover&nbsp;the SNP and that belong to haplotype B in the specified cell</li> </ul> </li> </ol> <p>All the files have been compressed using standard <em>gzip</em>.</p>

opencc-by-4.0May 2020View details →
zenodo44/100

Body measurements of human adults, with repeated readings

<p>This data file contains body measurements of participants of a course in morphometrics. The data are entirely anonymized and the compliance for publication of the data in the present form was obtained from each person.</p> <p>The participants were asked to measure themselves twice at intervals of one or two days, which resulted in reading a and reading b. The following measurements were taken (in mm):</p> <p>arm.l -&gt; arm length<br> bod.h -&gt; total body height<br> foo.l -&gt; foot length<br> hea.c -&gt; head circumference<br> leg.l -&gt; leg length</p> <p>The dataset may be considered for demonstration of measurement error.</p>

opencc-by-4.0Nov 2020View details →
zenodo44/100

MOSID (Microcontroller On-chip Sensor IDentification): A dataset of readings from the internal monitoring sensors of STM32L152RTXX microcontrollers during the stimulation of their electronic activity

<p>The MOSID (Microcontroller On-chip Sensor IDentification) dataset consists of 5 acquired data subsets (6,72 GB total, compressed into 560 MB), each collected during different experiments and periods using various equipment (HMP4040, DF1731SB &amp; HM305) and acquisition strategies. These subsets contain readings from the temperature and voltage sensors embedded in 20 STM32L-DISCOVERY devices. The data was captured during the execution of 5 different workloads as stimuli, repeated over 20 iterations. The stimuli employed are as follows:</p><ol><li>20x20 Long-type matrix product.</li><li>20x20 Float-type matrix product.</li><li>Algorithm for ascending sorting, Bubble Sort.</li><li>Algorithm for 2D-point clustering, Convex Hull.</li><li>Encryption algorithm AES 128-bit.</li></ol><p>The subsets are structured according to the folder format "X_Y," where X is the manually assigned number to the board, and Y is the corresponding number for the executed algorithm. Within each of these folders, files are present in the format "data_Z.txt," where Z represents the iteration number to which the file belongs. In total, the dataset comprises 9600 files with a final size of approximately 7 GB. The different presented subsets are as follows:</p><ul><li>ACQ1: Derived from the experiment named "Automatic Acquisition 1 (HMP4040)" conducted using a daisy-chain topology (20 out of 20 boards, 2000 files).</li><li>ACQ2: Derived from the experiment named "Automatic Acquisition 2 (HMP4040)" conducted using a daisy-chain topology (20 out of 20 boards, 2000 files).</li><li>ACQ3: Derived from the experiment named "Individual Acquisitions (HMP4040)", performed board by board from idle conditions (20 out of 20 boards, 2000 files).</li><li>ACQ4: Derived from the experiment named "GOLD SOURCE DF1731SB Acquisitions" conducted using a partial daisy-chain setup (2 devices at a time, 18 out of 20 boards excluding boards , 1800 files).</li><li>ACQ5: Derived from the experiment named "HANMATEK HM305 Acquisitions" conducted using a partial daisy-chain setup (2 devices at a time, 18 out of 20 boards, 1800 files).</li></ul><p>In each "data_Z.txt" file, starting from the 5th line, temperature and voltage raw ADC conversions from the sensors are provided, captured during the execution of the stimulus in successive lines. Additionally, a table (Table_UIDS.csv) with metadata for each of the boards used in the experiments is included, which is needed in order to normalize the data in terms of ºC and Volts.</p><ul><li>BOARD_NUM, which contains the manually assigned board number.</li><li>UID, which contains the Unique Identifier of the board assigned by the manufacturer.</li><li>T_CAL_1, which holds the calibration value of the board's temperature sensor at 30ºC.</li><li>T_CAL_2, which holds the calibration value of the board's temperature sensor at 100ºC.</li><li>VREFINT_CAL, which contains the calibration value of the board's voltage sensor.</li></ul>

opencc-by-4.0Oct 2023View details →
zenodo44/100

Simulated query reads used for benchmarks in Metabuli publication.

<p>Simulated query reads used for synthetic&nbsp;benchmarks in Metabuli publication.</p><p><strong>Prokaryote Subspecies Inclusion Test</strong></p><ul><li>Illumina paired-end short reads<ul><li>Simulated using Mason2 based on NovaSeq 6000's error rates.</li><li><strong>prokaryote_inclusion_reads_1.fna.gz</strong></li><li><strong>prokaryote_inclusion_reads_2.fna.gz</strong></li></ul></li><li>PacBio HiFi long reads (Sequel II CCS)<ul><li>Multi-pass reads were simulated using PBSIM3 based on Sequel Il's error profile.</li><li>CCS reads were generated using PacBio's CCS-read tool.</li><li><strong>prokaryote_inclusion_sequel_ccs.fq.gz</strong></li></ul></li><li>PacBio Sequel II long reads<ul><li>Simulated using PBSIM3 based on Sequel Il's error profile.</li><li><strong>prokaryote_inclusion_sequel.fq.gz</strong></li></ul></li><li>ONT long reads<ul><li>Simulated using PBSIM3 based on ONT's error profile.</li><li><strong>prokaryote_inclusion_ont.fq.gz</strong></li></ul></li></ul><p><strong>Prokaryote Subspecies exclusion Test</strong></p><ul><li>Illumina paired-end short reads<ul><li>Simulated using Mason2 based on NovaSeq 6000's error rates.</li><li><strong>prokaryote_ss-exclusion_reads_1.fna.gz</strong></li><li><strong>prokaryote_ss-exclusion_reads_2.fna.gz</strong></li></ul></li><li>PacBio HiFi long reads (Sequel II CCS)<ul><li>Multi-pass reads were simulated using PBSIM3 based on Sequel Il's error profile.</li><li>CCS reads were generated using PacBio's CCS-read tool.</li><li><strong>prokaryote_ss-exclusion_sequel_ccs.fq.gz</strong></li></ul></li><li>PacBio Sequel II long reads<ul><li>Simulated using PBSIM3 based on Sequel Il's error profile.</li><li><strong>prokaryote_ss-exclusion_sequel.fq.gz</strong></li></ul></li><li>ONT long reads<ul><li>Simulated using PBSIM3 based on ONT's error profile.</li><li><strong>prokaryote_ss-exclusion_ont.fq.gz</strong></li></ul></li></ul><p><strong>Prokaryote Species Exclusion Test</strong></p><ul><li>Illumina paired-end short reads<ul><li>Simulated using Mason2 based on NovaSeq 6000's error rates.</li><li><strong>prokaryote_exclusion_reads_1.fna.gz</strong></li><li><strong>prokaryote_exclusion_reads_2.fna.gz</strong></li></ul></li><li>PacBio HiFi long reads (Sequel II CCS)<ul><li>Multi-pass reads were simulated using PBSIM3 based on Sequel Il's error profile.</li><li>CCS reads were generated using PacBio's CCS-read tool.</li><li><strong>prokaryote_sp-exclusion_sequel_ccs.fq.gz</strong></li></ul></li><li>PacBio Sequel II long reads<ul><li>Simulated using PBSIM3 based on Sequel Il's error profile.</li><li><strong>prokaryote_sp-exclusion_sequel.fq.gz</strong></li></ul></li><li>ONT long reads<ul><li>Simulated using PBSIM3 based on ONT's error profile.</li><li><strong>prokaryote_exclusion_ont.fq.gz</strong></li></ul></li></ul>

opencc-by-4.0Jun 2023View details →
zenodo44/100

Explore maps as you read a comic book

<p>This repository contains three documents accompanying the article 'Explore maps as you read a comic book.</p> <p>The first document compiles several types of breaks encountered in various pan-scalar maps. Each type is also assigned a progression note after an analysis of their effects on the user.</p> <p>The next table illustrates an example of good generalisation for each of the hydrographic patterns observed in the article. This generalisation is based on three key concepts of progressivity, promoting a continuity of meaning, coherence, and rhythm.</p> <p>The final document represents a scale master inspired by the methodology of (Brewer and Buttenfield, 2007). In the article, we demonstrate how to adapt it to pan-scalar maps to better visualize sequences, as well as two types of rhythms: map cadence and abstraction cadence.</p>

opencc-by-4.0Dec 2023View details →
zenodo44/100

Metagenome-assembled genomes from Stordalen Mire, Sweden (2019) (MAGs from long-read, short-read, & hybrid assemblies)

<p>METHODS:</p> <p>Soil samples (6 total) were collected at the Stordalen Mire site in 2019 from two depths (1-5 &amp; 20-24 cm below ground) across three habitats (Palsa, Bog, and Fen). DNA was extracted based on the protocol described by&nbsp;<a href="http://dx.doi.org/10.17504/protocols.io.yxmvm244bg3p/v1">Li et al. (2024)</a>. For short reads, libraries were prepared at the Joint Genome Institute (JGI) with the KAPA Hyperprep kit, and sequenced with Illumina NovaSeq 6000. For long reads, libraries were prepared with the SMRTbell Express Template Prep Kit 2.0 (PacBio), then sequenced using PacBio Sequel IIe at JGI. PacBio data was processed at JGI to form filtered CCS (Circular Consensus Sequencing) reads.&nbsp;</p> <p>Assemblies were generated with short-only, long-only, and hybrid read sources: <strong>Short-only</strong> was assembled with <a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC5411777/">metaSPAdes</a>&nbsp;(v3.15.4) using <a href="https://zenodo.org/records/10806928">Aviary</a> (v0.5.3) with default parameters. <strong>Long-only</strong> was assembled with&nbsp;<a href="https://www.nature.com/articles/s41592-020-00971-x">metaFlye</a>&nbsp;(v2.9-b1768) using&nbsp;<a href="https://zenodo.org/records/10806928">Aviary</a> (v0.5.3) with default parameters. <strong>Hybrid</strong> assembly was performed using <a href="https://zenodo.org/records/10806928">Aviary</a> v0.5.3 with default parameters. This involved a step-down procedure with long-read assembly through <a href="https://www.nature.com/articles/s41592-020-00971-x">metaFlye</a> (v2.9-b1768), followed by short-read polishing by <a href="https://genome.cshlp.org/content/27/5/737">Racon</a> (v1.4.3), <a href="https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0112963">Pilon</a> (v1.24) and then Racon again. Next, reads that didn't map to high-quality metaFlye contigs were hybrid assembled with <a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC5411777/">SPAdes (--meta option)</a> and binned out with <a href="https://peerj.com/articles/7359/">MetaBAT2</a> (v2.1.5). For each bin, the reads within the bin were hybrid assembled using <a href="https://journals.plos.org/ploscompbiol/article?id=10.1371/journal.pcbi.1005595">Unicycler</a> (v0.4.8). The high-coverage metaFlye contigs and Unicycler contigs were then combined to form the assembly fasta file. Genome recovery was performed using&nbsp;<a href="https://zenodo.org/records/10806928">Aviary</a> v0.5.3 with samples chosen for differential abundance binning by <a href="https://zenodo.org/records/10939393">Bin Chicken</a> (v0.4.2) using <a href="https://zenodo.org/records/7130825">SingleM metapackage S3.0.5</a>. This involved initial read mapping through <a href="https://zenodo.org/records/10531254">CoverM</a> (v0.6.1)&nbsp;using <a href="https://academic.oup.com/bioinformatics/article/34/18/3094/4994778">minimap2</a> (v2.18)&nbsp;and binning by <a href="https://peerj.com/articles/1165/">MetaBAT</a>, <a href="https://peerj.com/articles/7359/">MetaBAT2</a> (v2.1.5), <a href="https://www.nature.com/articles/s41587-020-00777-4">VAMB</a> (v3.0.2), <a href="http://doi.org/10.1038/s41467-022-29843-y">SemiBin</a> (v1.3.1), <a href="https://zenodo.org/records/10460259">Rosella</a> (v0.4.2), <a href="https://www.nature.com/articles/nmeth.3103">CONCOCT</a> (v1.1.0)&nbsp;and <a href="https://academic.oup.com/bioinformatics/article/32/4/605/1744462">MaxBin2</a> (v2.2.7). Genomes were analyzed using <a href="https://www.nature.com/articles/s41592-023-01940-w">CheckM2</a> (v1.0.2)&nbsp;and clustered at 95% ANI using <a href="https://zenodo.org/records/10526086">Galah</a> (v0.4.0).</p> <p>&nbsp;</p> <p>FILES:</p> <ul> <li><strong>EMERGE_MAGs_2019_long-short-hybrid.tar.gz</strong> - Archive containing the MAG files (.fna).</li> <li><strong>metadata_MAGs_2019_EMERGE.tsv</strong> - Table containing source sample names and accessions, GTDB classifications, CheckM2 quality information, NCBI GenomeBatch- and MIMAG(6.0)-formatted attributes, and other metadata for the MAGs.</li> </ul> <p>&nbsp;</p> <p>FUNDING:</p> <p>This research is a contribution of the EMERGE Biology Integration Institute (<a href="https://emerge-bii.github.io/">https://emerge-bii.github.io/</a>), funded by the National Science Foundation, Biology Integration Institutes Program, Award # 2022070.</p> <p>This study was also funded by the Genomic Science Program of the United States Department of Energy Office of Biological and Environmental Research, grant #s DE-SC0004632. DE-SC0010580. and DE-SC0016440.</p> <p>We thank the Swedish Polar Research Secretariat and SITES for the support of the work done at the Abisko Scientific Research Station. SITES is supported by the Swedish Research Council's grant 4.3-2021-00164.</p> <p>Data from the Joint Genome Institute (JGI) was collected under BER Support Science Proposal 503530 (DOI: <a href="https://doi.org/10.46936/10.25585/60001148">10.46936/10.25585/60001148</a>), conducted by the U.S. Department of Energy Joint Genome Institute (<a href="https://ror.org/04xm1d337">https://ror.org/04xm1d337</a>), a DOE Office of Science User Facility supported by the Office of Science of the U.S. Department of Energy under Contract No. DE-AC02-05CH11231.</p>

opencc-by-4.0Nov 2024View details →
zenodo44/100

GWAS Summary Statistics for Publication: Identifying novel genetic and phenotypic associations to genomic features by leveraging off-target reads in exome sequencing data

<p>This dataset contains summary statistics for genome-wide association studies (GWAS) conducted on genomic features derived from off-target reads in whole-exome sequencing (WES) data. The study utilized tools like Seeing Beyond the Target (SBT) and ImReP to construct novel phenotypic features from unmapped reads in ~50,000 participants in the UK Biobank. Features include mitochondrial DNA (mtDNA) copy number, ribosomal DNA (rDNA) copy number (5S, 18S, 28S), immune repertoire metrics (e.g., T-cell receptor alpha diversity), and microvial genome load (viral and fungal).</p> <p>Summary statistics can be used for replication studies, meta-analyses, or further exploration of these phenotypes.</p>

opencc-by-4.0Nov 2024View details →
zenodo44/100

Long-read sequencing and structural variant characterization in 1,019 samples from the 1000 Genomes Project

SV analysis of the long-read sequencing data of 1,019 samples from the 1000 Genomes Project. The data is hosted at the International Genome Sample Resource (IGSR) in the <a href="https://ftp.1000genomes.ebi.ac.uk/vol1/ftp/data_collections/1KG_ONT_VIENNA/">1KG_ONT_VIENNA</a> directory. Please see the <a href="https://ftp.1000genomes.ebi.ac.uk/vol1/ftp/data_collections/1KG_ONT_VIENNA/README_1KG_ONT_VIENNA.md">README</a> and <a href="https://ftp.1000genomes.ebi.ac.uk/vol1/ftp/data_collections/1KG_ONT_VIENNA/README_1KG_ONT_VIENNA_datareuse_statement.md">data reuse statement</a> for further information about this dataset.

openmit-licenseApr 2024View details →
zenodo44/100

Complex basis of hybrid female sterility and Haldane's rule in Heliconius butterflies: Z-linkage and epistasis - RADseq and RNAseq reads, sterility phenotypes and pedigree

<p>RADseq and RNAseq reads (.fastq files),&nbsp;and sterility phenotypes and pedigree (.xlsx) using for QTL mapping of Heliconius pardalinus sterility crosses in Rosser, N., Edelman, N.B., Queste, L.M., Nelson, M., Seixas, F., Dasmahapatra, K.K. and Mallet, J., 2021. Complex basis of hybrid female sterility and Haldane&rsquo;s rule in Heliconius butterflies: Z-linkage and epistasis, accepted for publication in Molecular Ecology. Queries to Neil Rosser (neil.rosser@york.ac.uk).&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Dec 2020View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record