Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,569
datasets available to search
ShareScore release 0.9.0
Dataset results
1,569 results for “Reading”
Component processes of word reading in adults and children
Open the record for dataset details and reuse information.
Simulated NGS read datasets for bacterial pathogenic potential prediction
<p>## Predicting pathogenic potentials from NGS reads: novel bacterial species</p> <p>This repository contains simulated Illumina read datasets for bacterial pathogenic potential prediction and associated metadata extracted from the IMG Database (https://img.jgi.doe.gov/). The reads are 250bp long and were simulated with Mason (https://www.seqan.de/apps/mason/) from genomes downloaded from NCBI. The training-validation-test split was done on the species level to ensure "novelty" of validation and test species. The training sets contain 10 million reads per class, validation sets - 1.25 million reads per class, and test sets - 1.25 million paired reads per class. Additional, imbalanced training sets contain 2.5 million "nonpathogenic" and 17.5 million "pathogenic" reads, keeping the mean covarage constant for all species. The temporal benchmark test set contains reads from 3 additional pathogenic species in the Pantoea genus.</p> <p>## Predicting pathogenic potentials from NGS reads: novel strains of known species</p> <p>The BacPaCS datasets contain reads simulated from the dataset compiled by Barash et al. (https://doi.org/10.1093/bioinformatics/bty928). It this case, the training-validation-test split was done on the strain level (so different strains of the same species may be present in all three sets).</p>
Emotion regulation in the Ageing Brain, University of Reading, BBSRC
Open the record for dataset details and reuse information.
Data for: What's in a game: Video game visual-spatial demand location exhibits a double dissociation with reading speed
<p>The aggregate data in these datasets were used in analyses for "What’s in a game: Video game visual-spatial demand location exhibits a double dissociation with reading speed".</p>
Reading hyperlinks
<p>This is a eye-tracking dataset in the BIDS format (http://bids.neuroimaging.io/).</p> <p>Please find the details on the study here:</p> <p>Gagl B. (2016) Blue hypertext is a good design decision: no perceptual disadvantage in reading and successful highlighting of relevant information. PeerJ 4:e2467 <a href="https://doi.org/10.7717/peerj.2467">https://doi.org/10.7717/peerj.2467</a></p> <p>and all analysis scripts and a previous version of the data here:</p> <p>Gagl, B. (2016, April 17). Reading Hypertext. Retrieved from <a href="https://osf.io/8c57w/">https://osf.io/8c57w/.</a></p> <p> </p>
Genomes plasmids MDR B. fragilis ONT sequence read files in fastq format
<p>Supporting data for the manuscript <em>Complete genome assembly of clinical multidrug resistant Bacteroides fragilis isolates enables comprehensive identification of antimicrobial resistance genes and plasmids.</em></p> <p>Oxford Nanopore reads demultiplexed with <a href="https://www.google.com/url?sa=t&rct=j&q=&esrc=s&source=web&cd=1&cad=rja&uact=8&ved=2ahUKEwjNjqL5tY7iAhUawMQBHZHfDasQFjAAegQIAhAB&url=https%3A%2F%2Fgithub.com%2Frrwick%2FDeepbinner&usg=AOvVaw0wikvIUagLuFV38CwKZtia">Deepbinner</a> v0.2.0 and base-called (with demultiplexing) using Albacore v2.3.3. Barcodes and adapters were removed with <a href="https://github.com/rrwick/Porechop">Porechop</a> v0.2.4 with the --discard_middle option.</p> <p>Data from each isolate was produced from two runs per isolate. Data for the individual runs are included here. They can easily be concatenated eg with cat. Runs are named TVS_01,. TVS_02, TVS_03 and TVS_04.</p> <p>Fast5 (only demultiplexed with deepbinner and basecalled with albacore) as well as illumina reads and genome assemblies can be found via the NCBI bioproject accessions:</p> <p>Isolates, NCBI bioproject accession no:</p> <p>CCUG4856T, <a href="http://www.ncbi.nlm.nih.gov/bioproject/PRJNA525024">PRJNA525024</a></p> <p>BFO17, <a href="http://www.ncbi.nlm.nih.gov/bioproject/PRJNA244943">PRJNA244943</a></p> <p>BFO18, <a href="http://www.ncbi.nlm.nih.gov/bioproject/PRJNA244944">PRJNA244944</a></p> <p>S01, <a href="http://www.ncbi.nlm.nih.gov/bioproject/PRJNA244942">PRJNA244942</a></p> <p>BFO42, <a href="http://www.ncbi.nlm.nih.gov/bioproject/PRJNA253771">PRJNA253771</a></p> <p>BFO67, <a href="http://www.ncbi.nlm.nih.gov/bioproject/PRJNA254401">PRJNA254401</a></p> <p>BFO85, <a href="http://www.ncbi.nlm.nih.gov/bioproject/PRJNA254455">PRJNA254455</a></p> <p> </p> <p><strong>md5sum's (also found in the file md5.md5):</strong></p> <p>135d0570a1e49e25c8fde59f321cca68 BFO17_TVS_03_99377.barcode02_trimmed.fastq.gz<br> 3c6ca800a1f735937c0cffccd263fb82 BFO18_TVS_01_97673.barcode03_trimmed.fastq.gz<br> dc37823950f529d3a859b7a7c514af8c BFO18_TVS_03_99377.barcode03_trimmed.fastq.gz<br> abd707404f9ebbc38652e45ed378ca21 BFO42_TVS_02.barcode10_trimmed.fastq.gz<br> 9761e9ab082e276624c5778dcbeaffd2 BFO42_TVS_04.barcode10_trimmed.fastq.gz<br> ffc27009c0f7fead1af84ce045dbe5f3 BFO67_TVS_02.barcode09_trimmed.fastq.gz<br> 06c363a7feeeb88f9d195b3769b37f2b BFO67_TVS_04.barcode09_trimmed.fastq.gz<br> 5553c95cc98f4b9d4cfb38c4f8f8f037 BFO85_TVS_02.barcode08_trimmed.fastq.gz<br> b1a8013cba7a079cee6d3bdd6cd97ff2 BFO85_TVS_04.barcode08_trimmed.fastq.gz<br> 8169225219a5fb20935d5f0304aa80c5 CCUG4856T_TVS_01_97673.barcode01_trimmed.fastq.gz<br> a509b0ae912a798a91e677795198c1c6 CCUG5846T_TVS_03_99377.barcode01_trimmed.fastq.gz<br> 8a9d2eb8b626a6e87ed267d31aa220e3 S01_TVS_01_97673.barcode04_trimmed.fastq.gz<br> 69e239a17becc25a4f103dfc3bf5886a S01_TVS_03_99377.barcode04_trimmed.fastq.gz</p>
Metabarcoding data (number of reads per operational taxonomic unit) from a monitoring study of sandy beach meiofauna before and after sand nourishment (Ahrenshoop, Baltic Sea)
<p>We provide metabarcoding data (number of reads per operational taxonomic unit, OTU) determined from sediment samples collected on the sandy-beach water line of Ahrenshoop (Baltic Sea). Five sampling stations lay within the zone impacted by the sand nourishment between the boundary of the nature reserve in the north east and a site just north of the breakwater (AH01–AH05). An unaffected reference station was located south of Ahrenshoop (close to Niehagen) at the end of the road Pappelallee (PAP). Samples were collected at four dates. The first sampling was carried out before the sand nourishment took place (T0: 14 and 16 September 2021). Three samplings were realised after the impact: T1 (23 March 2022), T2 (27 September 2022), and T3 (28 March 2023). Latitude and longitude of each sampling location per station were recorded at each sampling date using a hand-held GPS application on a mobile phone. At the stations sampling locations varied over time. Prior to the sand nourishment the beach was narrow due to sand erosion in previous years. After the nourishment the additional extent of the beach was approximately 40 m at sampling date T1. Subsequently, progressive sand erosion forced the sampling locations (situated at the water line) further inland at T2 and T3.<br>Samples were taken from the beach-water interface (water line) in the middle of the area between two groynes. Plexiglass cores (inner core diameter 5.4 cm) were inserted vertically into the sediment down to 15 cm depth. Each core was sliced in 5 cm-layers (0–5, 5–10 and 10–15 cm). Sediment horizons were preserved in 96–99% ethanol. <br>Three cores (2 cores at T0) per sampling date were taken for metabarcoding analyses. The organisms were extracted by decantation over a 32-μm sieve. Genomic DNA was extracted from the filters using the DNeasy PowerSoil pro kit (Qiagen). Realtime-PCR was performed to amplify V1&V2, two hypervariable regions of 18S rDNA gene. The sequencing run was performed using the MiSeq Reagent Nanokit v2 (250 cycles paired end) on an Illumina MiSeq platform at the DZMB Metabarcoding lab in Wilhelmshaven, Germany. High-resolution amplicon sequence variants (ASVs) were obtained and compared to the NCBI database to assign taxonomic information to each ASV. The target meiofauna ASVs were further classified into operational taxonomic units (OTUs) with a 3% cut-off threshold using the statistical software R.</p> <p>Here, we present two Tables (as xlsx and tab-delimited files):<br>(1) the taxonomic description of the 843 OTUs and their assigned ID number;<br>(2) the number of reads per OTU per sample (including metadata for each sample: event; date; latitude; longitude; station, core and sample ID; sediment depth).</p> <p>The metabarcoding data are part of a larger ecological study on the influence of sand nourishment on meiofauna communities, which included grain-size and meiofauna abundances (see “related works”).</p> <p><strong>Comment: </strong>Our study is related to but not funded by the project ECAS Baltic: Strategies of ecosystem-friendly coastal protection and ecosystem-supporting coastal adaptation for the German Baltic Sea Coast <a href="https://deutsche-kuestenforschung.de/ecas-baltic.html">https://deutsche-kuestenforschung.de/ecas-baltic.html</a></p>
Behavioral and fMRI Data: Nurturing the reading brain: Home literacy practices are associated with children's neural response to printed words through vocabulary skills
<p>This is the behavioral and fMRI dataset described in "Nurturing the reading brain: Home literacy practices are associated with children’s neural response to printed words through vocabulary skills". </p> <p>Because of anonymization concerns within the framework of EU privacy regulations (<a href="https://gdpr-info.eu">GDPR</a>), we cannot provide raw MRI data. Therefore, the fMRI data consists of individual pre-processed volumes, normalized into the MNI template (see paper for details about the preprocessing pipeline). Anonymized behavioral data and first level analyses are also provided for each participant (SPM.mat file as well as beta, con, spmT, RPV and ResMS files). Note that the dataset also include runs and GLM results for a third task (Dots) that was not analyzed in the paper. Finally, the <a href="https://www.psychopy.org">PsychoPy</a> implementation of the tasks is also provided. If you have any questions, please send an email to jerome.prado [at] univ-lyon1.fr. </p> <p><strong>IMPORTANT:</strong></p> <p>In accordance with EU privacy regulations, we ask that you sign and return a Data Use Agreement (DUA) before downloading the data. You can download the DUA <a href="https://zenodo.org/record/4965716/files/DUA.pdf?download=1">here</a>. Please, sign it and send it to jerome.prado [at] univ-lyon1.fr.</p>
Catalog of GenBank sequence read archive (SRA) entries of 16S and 18S rRNA genes from bacterial and protistan planktonic communities along the Eastern Beaufort Sea coast, North Slope, Alaska, 2011-2013
Microbial communities in the coastal Arctic Ocean experience extreme variability in organic matter and inorganic nutrients driven by seasonal shifts in sea ice extent and freshwater inputs. Lagoons border more than half of the Beaufort Sea coast and provide important habitats for migratory fish and seabirds; yet, little is known about the planktonic food webs supporting these higher trophic levels. To investigate seasonal changes in bacterial and protistan planktonic communities, amplicon sequences of 16S and 18S rRNA genes were generated from samples collected during periods of ice-cover (April), ice break-up (June), and open water (August) from shallow lagoons along the eastern Alaska Beaufort Sea coast from 2011 through 2013. This data package catalogs sequence read archive (SRA) entries available through GenBank BioProject PRJNA530074 at https://www.ncbi.nlm.nih.gov/bioproject/PRJNA530074. This data package is associated with the following publication: Kellogg CTE, McClelland JW, Dunton KH and Crump BC (2019) Strong Seasonality in Arctic Estuarine Microbial Food Webs. Front. Microbiol. 10:2628. doi: 10.3389/fmicb.2019.02628 Environmental variables (physiochemical data from YSI and HOBO data loggers, as well as organic matter analysis and stable isotope data from discrete water samples) associated with this genomic dataset are available from the Arctic Data Center: Kenneth Dunton, Byron Crump, and James McClelland. Physical, chemical, and biological data from lagoons and open coastal waters in the nearshore environment of the eastern Alaska Beaufort Sea, 2011-2013. Arctic Data Center. doi:10.18739/A2DG13. To join the two datasets together, please use the provided site codes (column "site_name" here) and collection dates (column "collection_date" here) in each dataset. Note that the site codes in this package are without hyphens (e.g. JAA) while site codes in the above environmental data package have hyphens (e.g. JA-A). Instead of citing this package which is jus
Catalog of GenBank sequence read archive (SRA) entries of metagenomic DNA sequence analyses of bacterial and archaeal water column communities along the Eastern Beaufort Sea coast, North Slope, Alaska, 2012
In contrast to temperate systems, Arctic lagoons that span the Alaska Beaufort Sea coast face extreme seasonality. Nine months of ice cover up to ∼1.7 m thick is followed by a spring thaw that introduces an enormous pulse of freshwater, nutrients, and organic matter into these lagoons over a relatively brief 2–3 week period. Prokaryotic communities link these subsidies to lagoon food webs through nutrient uptake, heterotrophic production, and other biogeochemical processes, but little is known about how the genomic capabilities of these communities respond to seasonal variability. This study characterizes the metabolic capabilities of microbial communities across three seasons in two lagoons and one open coastal site along the eastern Alaska Beaufort Sea coast. We used metagenomic DNA sequence data of bacterial and archaeal water column communities to identify genes of relevant biogeochemical pathways. This data package catalogs sequence read archive (SRA) entries available through GenBank BioProject PRJNA642637 at https://www.ncbi.nlm.nih.gov/bioproject/PRJNA642637. This data package is associated with the following publication: Baker, Kristina D., Colleen T. E. Kellogg, James W. McClelland, Kenneth H. Dunton, and Byron C. Crump. “The Genomic Capabilities of Microbial Communities Track Seasonal Variation in Environmental Conditions of Arctic Lagoons.” Frontiers in Microbiology 12 (2021). https://doi.org/10.3389/fmicb.2021.601901. Environmental variables (physiochemical data from YSI and HOBO data loggers, as well as organic matter analysis and stable isotope data from discrete water samples) associated with this genomic dataset are available from the Arctic Data Center: Kenneth Dunton, Byron Crump, and James McClelland. Physical, chemical, and biological data from lagoons and open coastal waters in the nearshore environment of the eastern Alaska Beaufort Sea, 2011-2013. Arctic Data Center. doi:10.18739/A2DG13. To join the two datasets together, please use the provi
Raw read counts and phased SNP counts for every single cell in the sequencing datasets of the breast cancer patient S1 from "Characterizing allele- and haplotype-specific copy numbers in single cells with CHISEL"
<p>This dataset contains the raw read counts and phased SNP counts for every single cell in the sequencing datasets of breast cancer patient S1 from “Characterizing allele- and haplotype-specific copy numbers in single cells with CHISEL” [Zaccaria & Raphael, 2020]. These data enable the full reproduction of all the results in the related manuscript for breast cancer patient S1. Specifically, the data are provided in two files for every dataset <em>DAT</em> of patient S1 with the following format:</p> <ol> <li><em>DAT.raw</em>_<em>read</em>_<em>counts.bed.gz </em>is a multi-cell BED file containing the raw read counts in the following fields: <ul> <li>CHROMOSOME: the name of a human chromosome</li> <li>START: the starting genomic position of a genomic bin in the chromosome</li> <li>END: the ending genomic position of the genomic bin in the chromosome</li> <li>CELL: the cell barcode that uniquely identifies a cell</li> <li>NORMAL: the raw read count for the specified bin from a matched-normal sample</li> <li>COUNT: the raw read count for the specified bin in the specified cell</li> <li>RDR: the estimated read-depth ratio for the specified bin in the specified cell</li> </ul> </li> <li><em>DAT.phased</em>_<em>snps</em>_<em>counts.pos.gz </em>is a multi-cell POS file containing the phased SNP counts in the following fields: <ul> <li>CHROMOSOME: the name of a human chromosome</li> <li>POS: the genomic position in the chromosome of a germline SNP</li> <li>CELL: the cell barcode that uniquely identifies a cell</li> <li>COUNT_HAPLOTYPE_A: the count of reads that cover the SNP and that belong to haplotype A in the specified cell</li> <li>COUNT_HAPLOTYPE_B: the count of reads that cover the SNP and that belong to haplotype B in the specified cell</li> </ul> </li> </ol> <p>All the files have been compressed using standard <em>gzip</em>.</p>
Raw read counts and phased SNP counts for every single cell in the sequencing datasets of the breast cancer patient S0 from "Characterizing allele- and haplotype-specific copy numbers in single cells with CHISEL"
<p>This dataset contains the raw read counts and phased SNP counts for every single cell in the sequencing datasets of breast cancer patient S0 from “Characterizing allele- and haplotype-specific copy numbers in single cells with CHISEL” [Zaccaria & Raphael, 2020]. These data enable the full reproduction of all the results in the related manuscript for breast cancer patient S0. Specifically, the data are provided in two files for every dataset <em>DAT</em> of patient S0 with the following format:</p> <ol> <li><em>DAT.raw</em>_<em>read</em>_<em>counts.bed.gz </em>is a multi-cell BED file containing the raw read counts in the following fields: <ul> <li>CHROMOSOME: the name of a human chromosome</li> <li>START: the starting genomic position of a genomic bin in the chromosome</li> <li>END: the ending genomic position of the genomic bin in the chromosome</li> <li>CELL: the cell barcode that uniquely identifies a cell</li> <li>NORMAL: the raw read count for the specified bin from a matched-normal sample</li> <li>COUNT: the raw read count for the specified bin in the specified cell</li> <li>RDR: the estimated read-depth ratio for the specified bin in the specified cell</li> </ul> </li> <li><em>DAT.phased</em>_<em>snps</em>_<em>counts.pos.gz </em>is a multi-cell POS file containing the phased SNP counts in the following fields: <ul> <li>CHROMOSOME: the name of a human chromosome</li> <li>POS: the genomic position in the chromosome of a germline SNP</li> <li>CELL: the cell barcode that uniquely identifies a cell</li> <li>COUNT_HAPLOTYPE_A: the count of reads that cover the SNP and that belong to haplotype A in the specified cell</li> <li>COUNT_HAPLOTYPE_B: the count of reads that cover the SNP and that belong to haplotype B in the specified cell</li> </ul> </li> </ol> <p>All the files have been compressed using standard <em>gzip</em>.</p>
Body measurements of human adults, with repeated readings
<p>This data file contains body measurements of participants of a course in morphometrics. The data are entirely anonymized and the compliance for publication of the data in the present form was obtained from each person.</p> <p>The participants were asked to measure themselves twice at intervals of one or two days, which resulted in reading a and reading b. The following measurements were taken (in mm):</p> <p>arm.l -> arm length<br> bod.h -> total body height<br> foo.l -> foot length<br> hea.c -> head circumference<br> leg.l -> leg length</p> <p>The dataset may be considered for demonstration of measurement error.</p>
MOSID (Microcontroller On-chip Sensor IDentification): A dataset of readings from the internal monitoring sensors of STM32L152RTXX microcontrollers during the stimulation of their electronic activity
<p>The MOSID (Microcontroller On-chip Sensor IDentification) dataset consists of 5 acquired data subsets (6,72 GB total, compressed into 560 MB), each collected during different experiments and periods using various equipment (HMP4040, DF1731SB & HM305) and acquisition strategies. These subsets contain readings from the temperature and voltage sensors embedded in 20 STM32L-DISCOVERY devices. The data was captured during the execution of 5 different workloads as stimuli, repeated over 20 iterations. The stimuli employed are as follows:</p><ol><li>20x20 Long-type matrix product.</li><li>20x20 Float-type matrix product.</li><li>Algorithm for ascending sorting, Bubble Sort.</li><li>Algorithm for 2D-point clustering, Convex Hull.</li><li>Encryption algorithm AES 128-bit.</li></ol><p>The subsets are structured according to the folder format "X_Y," where X is the manually assigned number to the board, and Y is the corresponding number for the executed algorithm. Within each of these folders, files are present in the format "data_Z.txt," where Z represents the iteration number to which the file belongs. In total, the dataset comprises 9600 files with a final size of approximately 7 GB. The different presented subsets are as follows:</p><ul><li>ACQ1: Derived from the experiment named "Automatic Acquisition 1 (HMP4040)" conducted using a daisy-chain topology (20 out of 20 boards, 2000 files).</li><li>ACQ2: Derived from the experiment named "Automatic Acquisition 2 (HMP4040)" conducted using a daisy-chain topology (20 out of 20 boards, 2000 files).</li><li>ACQ3: Derived from the experiment named "Individual Acquisitions (HMP4040)", performed board by board from idle conditions (20 out of 20 boards, 2000 files).</li><li>ACQ4: Derived from the experiment named "GOLD SOURCE DF1731SB Acquisitions" conducted using a partial daisy-chain setup (2 devices at a time, 18 out of 20 boards excluding boards , 1800 files).</li><li>ACQ5: Derived from the experiment named "HANMATEK HM305 Acquisitions" conducted using a partial daisy-chain setup (2 devices at a time, 18 out of 20 boards, 1800 files).</li></ul><p>In each "data_Z.txt" file, starting from the 5th line, temperature and voltage raw ADC conversions from the sensors are provided, captured during the execution of the stimulus in successive lines. Additionally, a table (Table_UIDS.csv) with metadata for each of the boards used in the experiments is included, which is needed in order to normalize the data in terms of ºC and Volts.</p><ul><li>BOARD_NUM, which contains the manually assigned board number.</li><li>UID, which contains the Unique Identifier of the board assigned by the manufacturer.</li><li>T_CAL_1, which holds the calibration value of the board's temperature sensor at 30ºC.</li><li>T_CAL_2, which holds the calibration value of the board's temperature sensor at 100ºC.</li><li>VREFINT_CAL, which contains the calibration value of the board's voltage sensor.</li></ul>
Simulated query reads used for benchmarks in Metabuli publication.
<p>Simulated query reads used for synthetic benchmarks in Metabuli publication.</p><p><strong>Prokaryote Subspecies Inclusion Test</strong></p><ul><li>Illumina paired-end short reads<ul><li>Simulated using Mason2 based on NovaSeq 6000's error rates.</li><li><strong>prokaryote_inclusion_reads_1.fna.gz</strong></li><li><strong>prokaryote_inclusion_reads_2.fna.gz</strong></li></ul></li><li>PacBio HiFi long reads (Sequel II CCS)<ul><li>Multi-pass reads were simulated using PBSIM3 based on Sequel Il's error profile.</li><li>CCS reads were generated using PacBio's CCS-read tool.</li><li><strong>prokaryote_inclusion_sequel_ccs.fq.gz</strong></li></ul></li><li>PacBio Sequel II long reads<ul><li>Simulated using PBSIM3 based on Sequel Il's error profile.</li><li><strong>prokaryote_inclusion_sequel.fq.gz</strong></li></ul></li><li>ONT long reads<ul><li>Simulated using PBSIM3 based on ONT's error profile.</li><li><strong>prokaryote_inclusion_ont.fq.gz</strong></li></ul></li></ul><p><strong>Prokaryote Subspecies exclusion Test</strong></p><ul><li>Illumina paired-end short reads<ul><li>Simulated using Mason2 based on NovaSeq 6000's error rates.</li><li><strong>prokaryote_ss-exclusion_reads_1.fna.gz</strong></li><li><strong>prokaryote_ss-exclusion_reads_2.fna.gz</strong></li></ul></li><li>PacBio HiFi long reads (Sequel II CCS)<ul><li>Multi-pass reads were simulated using PBSIM3 based on Sequel Il's error profile.</li><li>CCS reads were generated using PacBio's CCS-read tool.</li><li><strong>prokaryote_ss-exclusion_sequel_ccs.fq.gz</strong></li></ul></li><li>PacBio Sequel II long reads<ul><li>Simulated using PBSIM3 based on Sequel Il's error profile.</li><li><strong>prokaryote_ss-exclusion_sequel.fq.gz</strong></li></ul></li><li>ONT long reads<ul><li>Simulated using PBSIM3 based on ONT's error profile.</li><li><strong>prokaryote_ss-exclusion_ont.fq.gz</strong></li></ul></li></ul><p><strong>Prokaryote Species Exclusion Test</strong></p><ul><li>Illumina paired-end short reads<ul><li>Simulated using Mason2 based on NovaSeq 6000's error rates.</li><li><strong>prokaryote_exclusion_reads_1.fna.gz</strong></li><li><strong>prokaryote_exclusion_reads_2.fna.gz</strong></li></ul></li><li>PacBio HiFi long reads (Sequel II CCS)<ul><li>Multi-pass reads were simulated using PBSIM3 based on Sequel Il's error profile.</li><li>CCS reads were generated using PacBio's CCS-read tool.</li><li><strong>prokaryote_sp-exclusion_sequel_ccs.fq.gz</strong></li></ul></li><li>PacBio Sequel II long reads<ul><li>Simulated using PBSIM3 based on Sequel Il's error profile.</li><li><strong>prokaryote_sp-exclusion_sequel.fq.gz</strong></li></ul></li><li>ONT long reads<ul><li>Simulated using PBSIM3 based on ONT's error profile.</li><li><strong>prokaryote_exclusion_ont.fq.gz</strong></li></ul></li></ul>
Explore maps as you read a comic book
<p>This repository contains three documents accompanying the article 'Explore maps as you read a comic book.</p> <p>The first document compiles several types of breaks encountered in various pan-scalar maps. Each type is also assigned a progression note after an analysis of their effects on the user.</p> <p>The next table illustrates an example of good generalisation for each of the hydrographic patterns observed in the article. This generalisation is based on three key concepts of progressivity, promoting a continuity of meaning, coherence, and rhythm.</p> <p>The final document represents a scale master inspired by the methodology of (Brewer and Buttenfield, 2007). In the article, we demonstrate how to adapt it to pan-scalar maps to better visualize sequences, as well as two types of rhythms: map cadence and abstraction cadence.</p>
Metagenome-assembled genomes from Stordalen Mire, Sweden (2019) (MAGs from long-read, short-read, & hybrid assemblies)
<p>METHODS:</p> <p>Soil samples (6 total) were collected at the Stordalen Mire site in 2019 from two depths (1-5 & 20-24 cm below ground) across three habitats (Palsa, Bog, and Fen). DNA was extracted based on the protocol described by <a href="http://dx.doi.org/10.17504/protocols.io.yxmvm244bg3p/v1">Li et al. (2024)</a>. For short reads, libraries were prepared at the Joint Genome Institute (JGI) with the KAPA Hyperprep kit, and sequenced with Illumina NovaSeq 6000. For long reads, libraries were prepared with the SMRTbell Express Template Prep Kit 2.0 (PacBio), then sequenced using PacBio Sequel IIe at JGI. PacBio data was processed at JGI to form filtered CCS (Circular Consensus Sequencing) reads. </p> <p>Assemblies were generated with short-only, long-only, and hybrid read sources: <strong>Short-only</strong> was assembled with <a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC5411777/">metaSPAdes</a> (v3.15.4) using <a href="https://zenodo.org/records/10806928">Aviary</a> (v0.5.3) with default parameters. <strong>Long-only</strong> was assembled with <a href="https://www.nature.com/articles/s41592-020-00971-x">metaFlye</a> (v2.9-b1768) using <a href="https://zenodo.org/records/10806928">Aviary</a> (v0.5.3) with default parameters. <strong>Hybrid</strong> assembly was performed using <a href="https://zenodo.org/records/10806928">Aviary</a> v0.5.3 with default parameters. This involved a step-down procedure with long-read assembly through <a href="https://www.nature.com/articles/s41592-020-00971-x">metaFlye</a> (v2.9-b1768), followed by short-read polishing by <a href="https://genome.cshlp.org/content/27/5/737">Racon</a> (v1.4.3), <a href="https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0112963">Pilon</a> (v1.24) and then Racon again. Next, reads that didn't map to high-quality metaFlye contigs were hybrid assembled with <a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC5411777/">SPAdes (--meta option)</a> and binned out with <a href="https://peerj.com/articles/7359/">MetaBAT2</a> (v2.1.5). For each bin, the reads within the bin were hybrid assembled using <a href="https://journals.plos.org/ploscompbiol/article?id=10.1371/journal.pcbi.1005595">Unicycler</a> (v0.4.8). The high-coverage metaFlye contigs and Unicycler contigs were then combined to form the assembly fasta file. Genome recovery was performed using <a href="https://zenodo.org/records/10806928">Aviary</a> v0.5.3 with samples chosen for differential abundance binning by <a href="https://zenodo.org/records/10939393">Bin Chicken</a> (v0.4.2) using <a href="https://zenodo.org/records/7130825">SingleM metapackage S3.0.5</a>. This involved initial read mapping through <a href="https://zenodo.org/records/10531254">CoverM</a> (v0.6.1) using <a href="https://academic.oup.com/bioinformatics/article/34/18/3094/4994778">minimap2</a> (v2.18) and binning by <a href="https://peerj.com/articles/1165/">MetaBAT</a>, <a href="https://peerj.com/articles/7359/">MetaBAT2</a> (v2.1.5), <a href="https://www.nature.com/articles/s41587-020-00777-4">VAMB</a> (v3.0.2), <a href="http://doi.org/10.1038/s41467-022-29843-y">SemiBin</a> (v1.3.1), <a href="https://zenodo.org/records/10460259">Rosella</a> (v0.4.2), <a href="https://www.nature.com/articles/nmeth.3103">CONCOCT</a> (v1.1.0) and <a href="https://academic.oup.com/bioinformatics/article/32/4/605/1744462">MaxBin2</a> (v2.2.7). Genomes were analyzed using <a href="https://www.nature.com/articles/s41592-023-01940-w">CheckM2</a> (v1.0.2) and clustered at 95% ANI using <a href="https://zenodo.org/records/10526086">Galah</a> (v0.4.0).</p> <p> </p> <p>FILES:</p> <ul> <li><strong>EMERGE_MAGs_2019_long-short-hybrid.tar.gz</strong> - Archive containing the MAG files (.fna).</li> <li><strong>metadata_MAGs_2019_EMERGE.tsv</strong> - Table containing source sample names and accessions, GTDB classifications, CheckM2 quality information, NCBI GenomeBatch- and MIMAG(6.0)-formatted attributes, and other metadata for the MAGs.</li> </ul> <p> </p> <p>FUNDING:</p> <p>This research is a contribution of the EMERGE Biology Integration Institute (<a href="https://emerge-bii.github.io/">https://emerge-bii.github.io/</a>), funded by the National Science Foundation, Biology Integration Institutes Program, Award # 2022070.</p> <p>This study was also funded by the Genomic Science Program of the United States Department of Energy Office of Biological and Environmental Research, grant #s DE-SC0004632. DE-SC0010580. and DE-SC0016440.</p> <p>We thank the Swedish Polar Research Secretariat and SITES for the support of the work done at the Abisko Scientific Research Station. SITES is supported by the Swedish Research Council's grant 4.3-2021-00164.</p> <p>Data from the Joint Genome Institute (JGI) was collected under BER Support Science Proposal 503530 (DOI: <a href="https://doi.org/10.46936/10.25585/60001148">10.46936/10.25585/60001148</a>), conducted by the U.S. Department of Energy Joint Genome Institute (<a href="https://ror.org/04xm1d337">https://ror.org/04xm1d337</a>), a DOE Office of Science User Facility supported by the Office of Science of the U.S. Department of Energy under Contract No. DE-AC02-05CH11231.</p>
GWAS Summary Statistics for Publication: Identifying novel genetic and phenotypic associations to genomic features by leveraging off-target reads in exome sequencing data
<p>This dataset contains summary statistics for genome-wide association studies (GWAS) conducted on genomic features derived from off-target reads in whole-exome sequencing (WES) data. The study utilized tools like Seeing Beyond the Target (SBT) and ImReP to construct novel phenotypic features from unmapped reads in ~50,000 participants in the UK Biobank. Features include mitochondrial DNA (mtDNA) copy number, ribosomal DNA (rDNA) copy number (5S, 18S, 28S), immune repertoire metrics (e.g., T-cell receptor alpha diversity), and microvial genome load (viral and fungal).</p> <p>Summary statistics can be used for replication studies, meta-analyses, or further exploration of these phenotypes.</p>
Long-read sequencing and structural variant characterization in 1,019 samples from the 1000 Genomes Project
SV analysis of the long-read sequencing data of 1,019 samples from the 1000 Genomes Project. The data is hosted at the International Genome Sample Resource (IGSR) in the <a href="https://ftp.1000genomes.ebi.ac.uk/vol1/ftp/data_collections/1KG_ONT_VIENNA/">1KG_ONT_VIENNA</a> directory. Please see the <a href="https://ftp.1000genomes.ebi.ac.uk/vol1/ftp/data_collections/1KG_ONT_VIENNA/README_1KG_ONT_VIENNA.md">README</a> and <a href="https://ftp.1000genomes.ebi.ac.uk/vol1/ftp/data_collections/1KG_ONT_VIENNA/README_1KG_ONT_VIENNA_datareuse_statement.md">data reuse statement</a> for further information about this dataset.
Complex basis of hybrid female sterility and Haldane's rule in Heliconius butterflies: Z-linkage and epistasis - RADseq and RNAseq reads, sterility phenotypes and pedigree
<p>RADseq and RNAseq reads (.fastq files), and sterility phenotypes and pedigree (.xlsx) using for QTL mapping of Heliconius pardalinus sterility crosses in Rosser, N., Edelman, N.B., Queste, L.M., Nelson, M., Seixas, F., Dasmahapatra, K.K. and Mallet, J., 2021. Complex basis of hybrid female sterility and Haldane’s rule in Heliconius butterflies: Z-linkage and epistasis, accepted for publication in Molecular Ecology. Queries to Neil Rosser (neil.rosser@york.ac.uk). </p> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.