Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
90
datasets available to search
ShareScore release 0.7.1
Dataset results
90 results for “FASTQ”
(Fastq Files) Amplicon sequencing of ama1 and mdr1 to track within-host P. falciparum diversity in Kilifi, KENYA
<p>These data were generated from amplicon sequencing of <em>Plasmodium falciparum</em> <em>ama1 </em>and<em> </em><em>mdr1</em> genes in samples collected from Kilifi, at the coast of Kenya.</p> <p>The two papers that reference these data will soon be included here:</p> <ol> <li> The Journal of Infectious Diseases - https://doi.org/10.1093/infdis/jiac144</li> <li>Wellcome Open Research - https://wellcomeopenresearch.org/articles/7-95</li> </ol> <p>Two objectives were explored:</p> <ol> <li>To determine temporal changes in the genetic diversity of malaria parasites in asymptomatic and febrile infections.</li> <li>To track within-host parasite diversity, throughout treatment in a clinical drug trial.</li> </ol>
Genomes plasmids MDR B. fragilis ONT sequence read files in fastq format
<p>Supporting data for the manuscript <em>Complete genome assembly of clinical multidrug resistant Bacteroides fragilis isolates enables comprehensive identification of antimicrobial resistance genes and plasmids.</em></p> <p>Oxford Nanopore reads demultiplexed with <a href="https://www.google.com/url?sa=t&rct=j&q=&esrc=s&source=web&cd=1&cad=rja&uact=8&ved=2ahUKEwjNjqL5tY7iAhUawMQBHZHfDasQFjAAegQIAhAB&url=https%3A%2F%2Fgithub.com%2Frrwick%2FDeepbinner&usg=AOvVaw0wikvIUagLuFV38CwKZtia">Deepbinner</a> v0.2.0 and base-called (with demultiplexing) using Albacore v2.3.3. Barcodes and adapters were removed with <a href="https://github.com/rrwick/Porechop">Porechop</a> v0.2.4 with the --discard_middle option.</p> <p>Data from each isolate was produced from two runs per isolate. Data for the individual runs are included here. They can easily be concatenated eg with cat. Runs are named TVS_01,. TVS_02, TVS_03 and TVS_04.</p> <p>Fast5 (only demultiplexed with deepbinner and basecalled with albacore) as well as illumina reads and genome assemblies can be found via the NCBI bioproject accessions:</p> <p>Isolates, NCBI bioproject accession no:</p> <p>CCUG4856T, <a href="http://www.ncbi.nlm.nih.gov/bioproject/PRJNA525024">PRJNA525024</a></p> <p>BFO17, <a href="http://www.ncbi.nlm.nih.gov/bioproject/PRJNA244943">PRJNA244943</a></p> <p>BFO18, <a href="http://www.ncbi.nlm.nih.gov/bioproject/PRJNA244944">PRJNA244944</a></p> <p>S01, <a href="http://www.ncbi.nlm.nih.gov/bioproject/PRJNA244942">PRJNA244942</a></p> <p>BFO42, <a href="http://www.ncbi.nlm.nih.gov/bioproject/PRJNA253771">PRJNA253771</a></p> <p>BFO67, <a href="http://www.ncbi.nlm.nih.gov/bioproject/PRJNA254401">PRJNA254401</a></p> <p>BFO85, <a href="http://www.ncbi.nlm.nih.gov/bioproject/PRJNA254455">PRJNA254455</a></p> <p> </p> <p><strong>md5sum's (also found in the file md5.md5):</strong></p> <p>135d0570a1e49e25c8fde59f321cca68 BFO17_TVS_03_99377.barcode02_trimmed.fastq.gz<br> 3c6ca800a1f735937c0cffccd263fb82 BFO18_TVS_01_97673.barcode03_trimmed.fastq.gz<br> dc37823950f529d3a859b7a7c514af8c BFO18_TVS_03_99377.barcode03_trimmed.fastq.gz<br> abd707404f9ebbc38652e45ed378ca21 BFO42_TVS_02.barcode10_trimmed.fastq.gz<br> 9761e9ab082e276624c5778dcbeaffd2 BFO42_TVS_04.barcode10_trimmed.fastq.gz<br> ffc27009c0f7fead1af84ce045dbe5f3 BFO67_TVS_02.barcode09_trimmed.fastq.gz<br> 06c363a7feeeb88f9d195b3769b37f2b BFO67_TVS_04.barcode09_trimmed.fastq.gz<br> 5553c95cc98f4b9d4cfb38c4f8f8f037 BFO85_TVS_02.barcode08_trimmed.fastq.gz<br> b1a8013cba7a079cee6d3bdd6cd97ff2 BFO85_TVS_04.barcode08_trimmed.fastq.gz<br> 8169225219a5fb20935d5f0304aa80c5 CCUG4856T_TVS_01_97673.barcode01_trimmed.fastq.gz<br> a509b0ae912a798a91e677795198c1c6 CCUG5846T_TVS_03_99377.barcode01_trimmed.fastq.gz<br> 8a9d2eb8b626a6e87ed267d31aa220e3 S01_TVS_01_97673.barcode04_trimmed.fastq.gz<br> 69e239a17becc25a4f103dfc3bf5886a S01_TVS_03_99377.barcode04_trimmed.fastq.gz</p>
Fastq data for Vero E6 cell infeceted with SARS CoV 2
<p>Vero cells infected with SARS-CoV 2 and the transcriptome sequenced by dRNAseq on nanopore.</p>
ABRomics genomic paired-end FASTQ data demo files
<p>This dataset contains the demo files for the <em>Genomic paired-end FASTQ</em> template of the ABRomics platform:</p> <ul> <li>Raw data: Paired-end Illumina sequencing files of sample ARDIG49.</li> <li>Metadata: Filled out <em>Genomic paired-end FASTQ</em> template for sample ARDIG49.</li> </ul>
Collection of Schistosoma mansoni ChIP-Seq input fastq files
<p>These are fastq files of ChIP-Seq input files for different life cycle stages of <em>Schistosoma mansoni</em>.</p> <ul> <li>adult female worms</li> <li>pairs of adults</li> <li>female cercariae</li> <li>miracidia</li> <li>primary sporocysts (sp1)</li> </ul> <p>Produced at IHPE (http://ihpe.univ-perp.fr/)</p>
EAW FASTQ files for bioinformatic courses (16S rRNA genes, 2018)
<p>Selection of 6 samples, with 6 Forward (R1) and 6 Reverse (R2) files, including primers. R1 and R2 reads are ca. 300 bp long, and were obtained from Illumina MiSeq technologies, at the FEM facility sequencing platform. The files refer to the16S rRNA gene reads obtained from the analyses carried out on the samples collected and filtered (Sterivex<sup>TM</sup> 0.22 µm) in different areas and depths of Lake Garda on September, 2018 (see EAW_2018_FASTQ_16S_description.docx).</p> <p>Sampling and analyses were carried out in the framework of the project Eco-AlpsWater (ASP569), funded by the Interreg Alpine Space program.</p> <p>A bioinformatic protocol for analyzing these data using DADA2 is available in Zenodo:</p> <pre>https://doi.org/10.5281/zenodo.5232772 </pre>
EAW FASTQ files for bioinformatic courses (18S rRNA genes, 2018)
<p>Selection of 6 samples, with 6 Forward (R1) and 6 Reverse (R2) files, including primers. R1 and R2 reads are ca. 300 bp long, and were obtained from Illumina MiSeq technologies, at the FEM facility sequencing platform. The files refer to the18S rRNA gene reads obtained from the analyses carried out on the samples collected and filtered (Sterivex<sup>TM</sup> 0.22 µm) in different areas and depths of Lake Garda on September, 2018 (see EAW_2018_FASTQ_18S_description.docx).</p> <p>Sampling and analyses were carried out in the framework of the project Eco-AlpsWater (ASP569), funded by the Interreg Alpine Space program.</p> <p>A bioinformatic protocol for analyzing these data using DADA2 is available in Zenodo:</p> <pre>https://doi.org/10.5281/zenodo.5233527</pre> <p>The corresponding 16S rRNA gene reads, obtained from the same eDNA extracts, are saved in <a href="https://doi.org/10.5281/zenodo.5215815">https://doi.org/10.5281/zenodo.5215815</a></p>
Urbanus and Cosgrove et al. Nature Communications (2023) - 12 months scRNAseq fastq files (mouse 1) for Figures 5 and 6
<p>This dataset contains .fastq files for the 12 month timepoint (mouse 1) scRNAseq dataset used in Figures 5 and 6. For any questions about this dataset please contact Leila Perie (leile.perie@curie.fr)</p>
Urbanus and Cosgrove et al. Nature Communications (2023) - 12 months scRNAseq fastq files (mouse 2) for Figures 5 and 6
<p>This dataset contains .fastq files for the 12 month timepoint (mouse 2) scRNAseq dataset used in Figures 5 and 6. For any questions about this dataset please contact Leila Perie (leile.perie@curie.fr)</p>
Heteroplasmy Benchmark Dataset - mitochondrial DNA mixture model - MiSeq - U5-H1-M1-M2-M3-M4-M5 - FASTQ
<p>mtDNA mixture model of 2 mtDNA sequences belonging to haplogroups U5 and H1. Run on Illumina MiSeq with 3 different polymerases (Clontech, Herculase, NEB Taq), and different DNA extraction protocols - Paired-end Fastq files</p> <p>M1 = Mixture 1:2 i.e. 50%</p> <p>M2 = Mixture 1:10 i.e. 10%</p> <p>M3 = Mixture 1:50 i.e. 2%</p> <p>M4 = Mixture 1:100 i.e. 1%</p> <p>M5 = Mixture 1:200 i.e. 0.5%</p>
RNA polymerase evolution FASTQ files (2/2)
<p>RNA polymerase (RNAP) is emblematic of complex biological systems that control multiple traits involving trade-offs such as growth versus maintenance. Laboratory evolution has revealed that mutations in RNAP subunits, including RpoB, are frequently selected. However, we lack a systems view of how mutations alter the RNAP molecular functions to promote adaptation. We, therefore, measured the fitness of thousands of <em>rpoB</em> variants under multiple conditions and genetic backgrounds, to find that adaptive mutations cluster in two separate modules. Mutations in one module favor growth over maintenance through a partial loss of an interaction associated with faster elongation. Mutations in the other favor maintenance over growth through a destabilized RNAP-DNA complex. The two molecular handles capture the versatile RNAP-mediated adaptations. Combining both interaction losses simultaneously improved maintenance and growth, challenging the idea that growth-maintenance tradeoff resorts only from limited resources, and revealing how compensatory evolution operates within RNAP.</p>
Subsampled fastq from GSM7890929 and GSM7890951
<p>This record contains:<br>- 4 fastqs which are subsets of fastqs corresponding to mESC and 168h gastruloids (the whole fastqs are available at SRA).<br>- All command lines which have been used to generate these fastqs:<br> - The main script is pipeline.sh<br> - Other files are temporary files or accessory scripts<br><br>The subsetting is totally biased and promote cells from different clusters so the proportion are absolutely artifical!</p><p>The goal of this subsetting is to be able to use this data in training and workflow testing.</p>
Fastq files for benchmarking somatic variant calling pipelines
<p>The <a href="https://download.imgag.de/public/validation_dataset_somatic/readme.html" target="_blank" rel="noopener">original .bam files</a> were provided by <a href="https://www.medizin.uni-tuebingen.de/de/das-klinikum/mitarbeiter/profil/3377" target="_blank" rel="noopener">Marc Sturm</a> from the <a href="https://www.medizin.uni-tuebingen.de/de/das-klinikum/einrichtungen/institute/medizinische-genetik-und-angewandte-genomik" target="_blank" rel="noopener">Institut für Medizinische Genetik und Angewandte Genomik at the University Clinic Tübingen.</a></p> <p>The files were transformed into .fq.gz files using the <a href="https://github.com/nf-core/bamtofastq" target="_blank" rel="noopener">nf-core/bamtofastq pipeline.</a> All information on the pipeline run can be found in the <a href="../api/records/10805134/draft/files/execution_report_2024-03-11_11-44-17.html/content" target="_blank" rel="noopener noreferrer">execution_report_2024-03-11_11-44-17.html</a>. The FASTQ files for the normal sample were directly uploaded to <a href="https://osf.io/cduyq/files/onedrive">the corresponding osf project</a>.</p> <p> </p> <p> </p>
Sub-sampled Fastq Files for ChIP-seq datasets from Click-Seq Science Paper (Science 2017, 10.1126/science.aal2066)
<p>Data were downloaded from SRA. We then randomly sub-sampled 20% of the reads using seqtk_sample v 1.2 in galaxy (seed 4).</p> <p>This dataset is a support dataset for the <a href="https://www.embl.de/training/events/2019/EPI19-01/index.html">EMBL Course: Chromatin Signatures During Differentiation: Integrated Omics</a></p>
Sub-sampled Fastq Files from Click-Seq Science Paper (Science 2017, 10.1126/science.aal2066)
<p>Data were downloaded from SRA. We then randomly sub-sampled 20% of the reads using seqtk_sample v 1.2 in galaxy (seed 4).</p> <p>This dataset is a support dataset for the <a href="https://www.embl.de/training/events/2019/EPI19-01/index.html">EMBL Course: Chromatin Signatures During Differentiation: Integrated Omics</a> </p>
kChip FASTQ files and associated data
<p>These are the FASTQ for the organisms that were studied in </p> <div> <p><span>Jared Kehe <em>et al.</em></span>, "<span>Positive interactions are common among culturable bacteria". </span><span><em>Sci. Adv. </em></span><span><strong>7</strong>, </span><span>eabi7159 </span><span>(2021). </span><span>DOI:<a href="https://doi.org/10.1126/sciadv.abi7159">10.1126/sciadv.abi7159</a></span></p> <p><span>These sequences were acquired from the original strains in the lab of Jonathan Friedman in Israel and were sequenced in the USA.</span></p> </div>
Extra tables and fastq files for "Virome Sequencing Identifies H5N1 Avian Influenza in Wastewater from Nine Cities."
<p>Tables:</p> <p>"mutation_analysis.xlsx" = detailed notes on variant analysis of H5N1 reads.</p> <p>"TEPHI_samples_H5N1_status1.xlsx" = Table of samples with metadata and H5N1 status</p> <p>"BioSampleObjects.txt" = Table from SRA mapping BioSample IDs to library (sample) IDs</p> <p> </p> <p>H5N1_reads_thru_p1858:</p> <p>Paired-end Illumina read files (.fastq) from all samples. Includes automated H5N1 called reads from iav_serotype tool from samples that were manually validated to have H5N1 specific reads.</p>
ATACSeq fastq files associated with the manuscript entitled 'Interspecies transcriptome analyses identify genes that control the development and evolution of limb skeletal proportion'
<p>This next-generation sequencing dataset is associated with the research manuscript entitled ‘<em>Interspecies transcriptome analyses identify genes that control the development and evolution of limb skeletal proportion</em>’ (https://www.biorxiv.org/content/10.1101/754002v2).</p> <p>The zipped folder ‘<strong>Zenodo_Saxena_etal_2021_ATACSeq_FastqFiles</strong>’ contains raw/unprocessed ATACSeq Fatsq read files for postnatal day 5 (P5) mouse (Mus) and jerboa (Jac) cartilage samples (Metatarsal = MT; Radius/Ulna = RU).</p> <p>> The <strong>Jac_P5</strong> subfolder contains paired-end reads (R1 and R2) for three jerboa metatarsals (MT1-3) and radius/ulna (RU1-3) biological replicates.</p> <p>> The <strong>Mus_P5</strong> subfolder contains paired-end reads (R1 and R2) for two mouse metatarsals (MT1-2) and radius/ulna (RU1-2) biological replicates.</p>
The impact of cefuroxime prophylaxis on human intestinal microbiota in surgical oncological patients - Dataset (FASTQ FILES)
<p>Dataset containing FASTQ files of the sequenced samples, generated by the Illumina MiSeq platform. </p> <p><span>This data is freely available under a CC-BY license; if you use it in your work, please cite our paper, "The impact of cefuroxime prophylaxis on human intestinal microbiota in surgical oncological patients" (DOI 10.3389/frmbi.2022.1092771).</span></p>
Urbanus and Cosgrove Nature Communications (2023) - scRNAseq fastq files GFP positive sample in Figure 2
<p>This dataset contains .fastq files for the GFP negative sample in the scRNAseq dataset used in Figures 2. For any questions about this dataset please contact Leila Perie (leile.perie@curie.fr)</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.