Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

2,439

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

2,439 results for “assembly”

Learn how ShareScore rates datasets ↗
zenodo52/100

Draft genome assembly version 1 of the meadow spittlebug Philaenus spumarius (Linnaeus, 1758) (Hemiptera, Aphrophoridae)

<p>We sequenced the genome of the meadow spittlebug, <em>Philaenus spumarius </em>(Linnaeus, 1758), the main insect vector of <em>Xylella fastidiosa </em>Wells et al. 1987 in Europe (Saponari et al., 2014), using 10x Chromium linked-reads. A single <em>P. spumarius</em> adult female from Portugal (Fontanelas, Sintra; GPS location: 38&deg;50&#39;15.75&quot;N; 9&deg;25&#39;20.77&quot;W), collected in September of 2018, was selected for genome sequencing. This population was initially surveyed for colour polymorphism in 1988 (Quartau &amp; Borges, 1997) and was later included in phylogeographic and population genomic studies of this species (Rodrigues et al., 2014; Seabra et al., unpublished). It is also geographically close to the population from which the individual used for the first partial genome assembly was collected (Rodrigues et al., 2016). The availability of this previous genetic information contributed to the choice of this population as the source of genomic material for whole genome sequencing. A subset of males from the same collection date were analysed for genitalia morphology to confirm species identification, as the best diagnostic characters are the appendages of the aedeagus (Drosopoulos &amp; Quartau, 2002).</p> <p>The genomic DNA of the <em>P. spumarius</em> adult from Sintra was extracted using Illustra Nucleon Phytopure kit according to the manufacturer&rsquo;s instructions (GE Healthcare). We assessed the quality and concentration of the DNA using Femto fragment analyser (Agilent). 10x Chromium library preparation and Illumina genome sequencing (HiSeq X, 150bp paired-end) were performed by Novogene Bioinformatics Technology Co, Beijing, China, in accordance with standard protocols.</p> <p>To create the <em>de novo</em> 10x Chromium assembly we ran Supernova 2.1.1 (Weisenfeld et al., 2017) on the 10x Chromium linked-read data with default parameters, using 1.0 billion reads corresponding to 56X coverage. To improve the initial supernova assembly, we performed iterative scaffolding using all of the 10x raw data (2.3 billion of reads). We ran two rounds of Scaff10x (https://github.com/wtsi-hpag/Scaff10X), followed by mis-assembly detection and correction with Tigmint (Jackman et al., 2018). This was followed by a final round of scaffolding with ARCS (Yeo et al., 2018). The assembly was checked for contamination using the BlobTools pipeline (version 0.9.19; Laetsch and Blaxter 2017;&nbsp;Kumar et al., 2013) and k-mer content was analysed with the KAT comp tool (Mapleson et al., 2017). In order to perform these analyses, it was necessary to remove the 10x linked barcodes from the reads with the script process_10xReads.py (https://github.com/ucdavis-bioinformatics/proc10xG).&nbsp;We assessed the quality of our draft genome assembly by searching for conserved, single copy, arthropod genes (n=1,066) with Benchmarking Universal Single-Copy Orthologs (BUSCO) v3.0 (Waterhouse et al., 2018).</p> <p>With the above assembly procedure, we obtained a final assembly of 2.7 Gb, having a scaffold N50 length of 116 Kb (contig N50 = 18 Kb) and the longest scaffold was 3.7 Mb. The length of the assembly was consistent with the genome size estimated by flow cytometry (Rodrigues et al., 2016). The k-mer distribution indicated high heterozygosity, estimated at 2.3%. BlobTools analyses revealed the presence of contigs assigned to <em>Sodalis </em>spp. (Enterobacteriaceae), a symbiont in members of tribe Philaenini (Koga et al., 2013). These contigs were filtered from the final assembly. Gene completeness assessment shows that 956 (89.6%) among 1,066 BUSCOs were &nbsp;found as complete copies, with only 26 (2.4%) missing. Of the BUSCOs that were detected, 878 (82.4%) were complete and single-copy, 78 (7.3%) were complete and duplicated and 84 (7.9%) were fragmented.</p> <p>In conclusion, due in part to high (2.3%) heterozygosity levels, the <em>P. spumarius</em> version 1 genome assembly is highly fragmented. Nonetheless, the assembly is considered complete and is likely to contain the majority of the gene content of <em>P. spumarius.</em></p>

opencc-by-4.0Jan 2020View details →
zenodo52/100

Pawpaws prevent predictability: A locally-dominant tree alters understory beta-diversity and community assembly

<p>Data used in "Pawpaws Prevent Predictability: A locally-dominant tree alters understory beta-diversity and community assembly" (Wassel and Myers) accepted for publication in Ecosphere.<br><br><strong>Metadata for Zenodo.pdf&nbsp;</strong>contains more information on the following data files including descriptions of the columns.&nbsp;</p> <p>The file <strong>understory_abundance_data2021.csv</strong>&nbsp;contains all species abundances in 1x1m plots. This data was used for analyses in publication. Each row is a plot, each column is a speceis or plot descriptor, values for columns 5 and higher are species abundances. Data was collected July-August 2021 by Anna Wassel in Missouri, USA.&nbsp;</p> <p>The file <strong>understory_species_list2021.csv&nbsp;</strong>contains a list of the species codes used in the first file with their scientific names and their status as herbs or woody. This was used to filter out herbaceous species from the data set for herbaceous-only analyses.&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Nov 2024View details →
zenodo52/100

Orbicella faveolata and O. franksi coral metagenome assemblies from the Lower Florida Keys region of Florida, USA

<div> <p>The enclosed files include mostly <em>Orbicella faveolata</em> and three <em>Orbicella franksi</em> coral metagenome assemblies collected from the Lower Keys in Florida&rsquo;s Coral Reef, USA. Metadata for the files is included in this repository. Apparently healthy coral tissue cores were collected between May 28 and June 21, 2021. The DNA was extracted from the host and associated microorganisms and sequenced in a paired-end 150 bp format on an Illumina NovaSeq. Trimming and quality filtering of DNA sequence reads proceeded, followed by host and photoendosymbiotic dinoflagellate DNA removal. The host-cleaned reads were assembled individually by coral sample into longer contigs using MegaHit v1.1.4. The &ldquo;Assembly_Fastas&rdquo; zipped file contains 41 metagenome assemblies from the individual <em>Orbicella faveolata</em> corals and 3 assemblies from the individual <em>Orbicella franksi&nbsp;</em>colonies for a total of 44 assemblies. In addition, these assemblies were annotated with eggnog-mapper v2.1.6 to generate both predicted gene regions and annotation output files. The &ldquo;Predicted_Gene_Fastas&rdquo; zipped file contains nucleotide fasta files of the predicted gene regions for all 44&nbsp;coral metagenome assemblies. The fasta header of each gene includes the contig ID it originated from in the associated &ldquo;Assembly_Fasta&rdquo;. The &ldquo;Predicted_Gene_Annotations&rdquo; zipped file contains either .csv or .xlsx files with the eggnog-mapper-based annotations. These files contain a &ldquo;query contig&rdquo; that corresponds to the contig ID in the fasta header of the &ldquo;Predicted_Gene_Fasta&rdquo;.&nbsp;</p> <p>In addition to individual assemblies, a co-assembly was generated that included all 41 <em>Orbicella faveolata</em> coral samples. Prior to co-assembly, further removal of eukaryotic DNA proceeded by splitting the indiviudual assemblies into eukaryotic and prokaryotic content with the program EukRep v0.6.7, followed by mapping of the host-clean reads to the eukaryotic DNA to remove them. The eukaryote-clean reads from all 41 corals were input into MegaHit to generate a co-assembly. The co-assembly is included (FLK_OFAV_MG_coassembly_final.contigs.fa). Predicted genes from the co-assembly were generated with Prodigal v2.6.3 and the nucleotide fasta of the output is included in this repository (FLK_OFAV_MG_pred.fna). Like with the indiviudal assemblies, eggnog-mapper was used to generate annotations of the predicted genes from Prodigal (FLK_OFAV_MG.emapper.annotations.xlsx).&nbsp; Additionally, the abundance of each predicted gene was generated using Salmon to map the eukaryote-clean reads to the predicted genes. The number of reads (counts) for each gene across each coral sample were aggregated as integers into one table and included in this repository (FLK_OFAV_MG_pred_NumReads.tsv).&nbsp;&nbsp;</p> </div> <div> <p>These data were processed and generated by Julie Meyer&rsquo;s Lab at the University of Florida, using funding from the Florida Department of Environmental Protection.&nbsp;&nbsp;</p> </div>

opencc-by-4.0Jun 2024View details →
zenodo52/100

Labeled Time Series Data of Force/Torque for Monitoring Assembly Processes with a Delta Robot

<p>This dataset comprises 524 recordings of 6-dimensional time series data, capturing forces in three directions and torques in three directions during the assembly of small car model wheels. The data was collected using an equidistant sampling method with a sampling period of 0.004 seconds. Each time series represents the process of assembling one wheel, specifically the placement of a tire onto a rim, and includes a label indicating whether the assembly was successful (OK). The wheels were assembled in batches of four, and the recordings were obtained over six different days. The labels of recordings from two (days 3 and 4) of the six days are invalid as described in [1].&nbsp; The labels presented in this data set are only binary (they do not describe the reason of the failure). The labels of recordings from days 5 and 6 are created by human while the other labels came from a convolutional neural network based computer vision classifier and can be inaccurate as described in section 5.4 of [1].&nbsp; &nbsp;</p> <h4>Dataset Structure:</h4> <ul> <li><strong>File:</strong> <code>ForceTorqueTimeSeries.csv</code> <ul> <li><strong>Columns:</strong> <ul> <li><code>idx (1-524)</code>: Index of the recording corresponding to the assembly of one wheel.</li> <li><code>label (true/false)</code>: Indicates whether the assembly was successful (TRUE = product is OK).</li> <li><code>meas_id (1-6)</code>: Identifier for the day on which the recording was made (refer to Table 2.1 in [1]).</li> <li><code>force_x</code>: X-component of the force measured by the sensor mounted on the delta robot's end effector.</li> <li><code>force_y</code>: Y-component of the force.</li> <li><code>force_z</code>: Z-component of the force.</li> <li><code>torque_x</code>: X-component of the torque.</li> <li><code>torque_y</code>: Y-component of the torque.</li> <li><code>torque_z</code>: Z-component of the torque.</li> </ul> </li> </ul> </li> </ul> <h4>Additional Files:</h4> <ul> <li><strong><code>IMG_3351.MOV</code>:</strong> A video demonstrating the assembly process for one batch of four wheels.</li> <li><strong><code>F3-BP-2024-Trna-Ales-Ales Trna - 2024 - Anomaly detection in robotic assembly process using force and torque sensors.pdf</code>:</strong> Bachelor thesis [1] detailing the dataset and preliminary experiments on fault detection.</li> <li><strong><code>F3-BP-2024-Hanzlik-Vojtech-Anomaly_Detection_Bachelors_Thesis.pdf</code>:</strong> Bachelor thesis [2] describing the data acquisition process.</li> </ul> <h3>References:</h3> <ol> <li>Trna, A. (2024). <em>Anomaly detection in robotic assembly process using force and torque sensors</em> [Bachelor&rsquo;s thesis, Czech Technical University in Prague].</li> <li>Hanzlik, V. (2024). <em>Edge AI integration for anomaly detection in assembly using Delta robot</em> [Bachelor&rsquo;s thesis, Czech Technical University in Prague].</li> </ol>

opencc-by-4.0Sep 2024View details →
zenodo48/100

DATASET: De novo assembly and functional annotation of the heart + hemolymph transcriptome in the Caribbean spiny lobster Panulirus argus

<p>The spiny lobster <em>Panulirus argus</em> is an ecologically relevant species in shallow water coral reefs and target of the most lucrative fishery in the greater Caribbean region. This study reports, for the first time, the heart + hemolymph transcriptome of the Caribbean spiny lobster<em> Panulirus argus</em> assembled from short Illumina 150&thinsp;bp PE raw reads. A total 80,152,094 raw reads were assembled using the Oyster River Protocol pipeline that aspires to become the standard protocol for <em>de novo</em> transcriptome assembly. The assembly resulted in a total of 254,773 transcripts. Functional gene annotation was conducted using the software package &#39;dammit&#39; that also aspires to become the standard protocol for <em>de novo</em> transcriptome annotation. Lastly, gene enrichment analyses were conducted using the Gene Ontology (GO), KEGG pathway analyses (Kaas), and KOG (WebMGA) databases. This resource will be of utmost importance in future research aiming at exploring the effect of local and regional anthropogenic disturbances as well as global climate change on the molecular physiology of this overexploited species.</p>

opencc-by-4.0Dec 2019View details →
zenodo48/100

Gene Annotations of 49 Bacillariophyta Genome Assemblies

<div>Contact: katharina.hoff@uni-greifswald.de.</div> <div>&nbsp;</div> <div> <h2>Manuscript</h2> <p>The data hosted here is associated with the preprint <a href="https://doi.org/10.48550/arXiv.2410.05467">https://doi.org/10.48550/arXiv.2410.05467</a></p> </div> <h2>Files</h2> <div>The following gff3-files with structural and functional genome annotation are included in the compressed archive Bacillariophyta_annotations.tar.gz:</div> <div>&nbsp;</div> <div>Asterionella_formosa.gff3<br>Asterionellopsis_glacialis.gff3<br>Bacterosira_constricta.gff3<br>Chaetoceros_muellerii.gff3<br>concatenated_output.gff3<br>Conticribra_guillardii.gff3<br>Conticribra_weissflogii.gff3<br>Craspedostauros_australis.gff3<br>Cyclostephanos_invisitatus.gff3<br>Cyclostephanos_tholiformis.gff3<br>Cyclotella_atomus.gff3<br>Cyclotella_baltica.gff3<br>Cyclotella_choctawhatcheeana.gff3<br>Cyclotella_cryptica.gff3<br>Cylindrotheca_fusiformis.gff3<br>Detonula_confervacea.gff3<br>Discostella_pseudostelligera.gff3<br>Discostella_stelligera.gff3<br>Discostella_stelligeroides.gff3<br>Epithemia_pelagica.gff3<br>Fistulifera_pelliculosa.gff3<br>Fistulifera_solaris.gff3<br>Fragilaria_radians.gff3<br>Fragilariopsis_cylindrus.gff3<br>Licmophora_abbreviata.gff3<br>Mediolabrus_comicus.gff3<br>Nitzschia_palea.gff3<br>Nitzschia_putrida.gff3<br>Porosira_glacialis.gff3<br>Psammoneis_japonica.gff3<br>Pseudo-nitzschia_multiseries.gff3<br>Pseudo-nitzschia_pungens.gff3<br>Skeletonema_costatum.gff3<br>Skeletonema_marinoi.gff3<br>Skeletonema_menzelii.gff3<br>Skeletonema_potamos.gff3<br>Skeletonema_tropicum.gff3<br>Stephanocyclus_meneghinianus.gff3<br>Stephanodiscus_minutulus.gff3<br>Stephanodiscus_triporus.gff3<br>Thalassiosira_allenii.gff3<br>Thalassiosira_delicatula.gff3<br>Thalassiosira_exigua.gff3<br>Thalassiosira_gravida.gff3<br>Thalassiosira_livingstoniorum.gff3<br>Thalassiosira_mediterranea.gff3<br>Thalassiosira_oceanica.gff3<br>Thalassiosira_ordinaria.gff3<br>Thalassiosira_pacifica.gff3<br>Thalassiosira_profunda.gff3</div> <div>&nbsp;</div> <div>To extract the dataset, execute the following command:</div> <div>&nbsp;</div> <div><code>tar -xvf Bacillariophyta_annotations.tar.gz</code></div> <h2>Genome Assemblies</h2> <div>The files in this folder attain to genome assemblies are publicly available at NCBI datasets (https://www.ncbi.nlm.nih.gov/datasets/). We used the following versions:</div> <div>&nbsp;</div> <div>Asterionella formosa GCA_002256025.1</div> <div>Asterionellopsis glacialis GCA_014885115.2</div> <div>Bacterosira constricta GCA_037356235.1</div> <div>Chaetoceros muellerii GCA_019693545.1</div> <div>Conticribra guillardii GCA_036939335.1</div> <div>Conticribra weissflogii GCA_036940025.1</div> <div>Craspedostauros australis GCA_026770025.1</div> <div>Cyclostephanos invisitatus GCA_036939675.1</div> <div>Cyclostephanos tholiformis GCA_036939975.1</div> <div>Cyclotella atomus GCA_036939935.1</div> <div>Cyclotella baltica GCA_036939635.1</div> <div>Cyclotella choctawhatcheeana GCA_036939855.1</div> <div>Cyclotella cryptica GCA_013187285.1</div> <div>Cylindrotheca fusiformis GCA_019693525.1</div> <div>Detonula confervacea GCA_036939415.1</div> <div>Discostella pseudostelligera GCA_036940085.1</div> <div>Discostella stelligera GCA_036939735.1</div> <div>Discostella stelligeroides GCA_036939555.1</div> <div>Epithemia pelagica GCA_946965045.2</div> <div>Fistulifera pelliculosa GCA_026008555.1</div> <div>Fistulifera solaris GCA_030295235.1</div> <div>Fragilaria radians GCA_900642245.1</div> <div>Fragilariopsis cylindrus GCA_900095095.1</div> <div>Licmophora abbreviata GCA_900291995.1</div> <div>Mediolabrus comicus GCA_036940125.1</div> <div>Nitzschia palea GCA_019593585.1</div> <div>Nitzschia putrida GCA_016586335.1</div> <div>Porosira glacialis GCA_036939395.1</div> <div>Psammoneis japonica GCA_008632985.1</div> <div>Pseudo-nitzschia multiseries GCA_037355745.1</div> <div>Pseudo-nitzschia pungens GCA_037355855.1</div> <div>Skeletonema costatum GCA_018806925.1</div> <div>Skeletonema marinoi GCA_030544225.1</div> <div>Skeletonema menzelii GCA_036940005.1</div> <div>Skeletonema potamos GCA_036940105.1</div> <div>Skeletonema tropicum GCA_037178625.1</div> <div>Stephanocyclus meneghinianus GCA_036940045.1</div> <div>Stephanodiscus minutulus GCA_036939435.1</div> <div>Stephanodiscus triporus GCA_036939755.1</div> <div>Thalassiosira allenii GCA_036939655.1</div> <div>Thalassiosira delicatula GCA_036939835.1</div> <div>Thalassiosira exigua GCA_036939895.1</div> <div>Thalassiosira gravida GCA_037356215.1</div> <div>Thalassiosira livingstoniorum GCA_036939595.1</div> <div>Thalassiosira mediterranea GCA_036939795.1</div> <div>Thalassiosira oceanica GCA_019693575.1</div> <div>Thalassiosira ordinaria GCA_036939695.1</div> <div>Thalassiosira pacifica GCA_036939875.1</div> <div>Thalassiosira profunda GCA_036939355.1</div> <p>&nbsp;</p> <h2>Converting to Protein FASTA and Coding Sequences FASTA</h2> <p>&nbsp;</p> <div>To save storage place at Zenodo, we did not upload the protein FASTA and coding sequence FASTA files. They can easily be generated from the genome FASTA file in combination with the respective GFF3 file. To do this, you can use the following commands:</div> <p>&nbsp;</p> <div><code># assume that genome.fa ist you respective genome FASTA file downloaded from NCBI datasets</code></div> <div><code>sed '/^&gt;/ s/ .*//' genome.fasta &gt; genome_short_headers.fasta</code></div> <div><code># assume that file.gff is the respective GFF3 file</code></div> <div><code>getAnnoFastaFromJoingenes.py -g genome_short_headers.fasta -3 file.gff -o nameStem</code></div> <div>&nbsp;</div> <div>This will produce the following files: nameStem.aa (protein FASTA file) and nameStem.codingseq (coding sequence FASTA file).</div> <div>&nbsp;</div> <div>The getAnnoFastaFromJoingenes.py script is available at https://raw.githubusercontent.com/Gaius-Augustus/Augustus/master/scripts/getAnnoFastaFromJoingenes.py . It is part of the AUGUSTUS software package.</div> <h2>Release notes</h2> <p>This release contains a gene set where a results of an OrthoFinder run that did not include genes on contigs that are suspected to be contaminants or horizontal gene transfer candidates were used to filter single exon genes. This means the gene and transcript counts changed compared to the previous release.</p> <h2>License</h2> <div>The genome annotation files are licensed under the Creative Commons Attribution 4.0 International License (CC BY 4.0). To view a copy of this license, visit http://creativecommons.org/licenses/by/4.0/ or send a letter to Creative Commons, PO Box 1866, Mountain View, CA 94042, USA.</div>

opencc-by-4.0Sep 2024View details →
zenodo48/100

Introduction to Ancient Metagenomics Textbook (Edition 2025): de novo Genome Assembly

<p>Data and conda software environment file for the chapter &#39;<em>de novo</em> Genome Assembly&#39; of the SPAAM Community&#39;s textbook: Introduction to Ancient Metagenomics (https://www.spaam-community.org/intro-to-ancient-metagenomics-book).</p>

opencc-by-4.0Sep 2024View details →
zenodo48/100

Economical routes to size-specific assembly of self-closing structures

<p>This data contains images related to a publication on the self-assembly of DNA origami particles (<a href="https://www.science.org/doi/10.1126/sciadv.ado5979">https://www.science.org/doi/10.1126/sciadv.ado5979</a>). In this work, we conduct self-assembly experiments with various unique subunit types that target two different diameters of tubule structures.</p> <p>We provide image data of tubules that are associated with the probability distributions reported across several figures in the main text. Images of tubules are in the ZIP archives and show the section of tubules we analyzed to produce the probability distributions in the manuscript. Each folder of images has an associated CSV file that relates an image name to the type of tubule that the image was identified as. Tubule types have "m" and "n" values.</p> <p>We provide full tomogram reconstruction data for the multicomponent tubules that are shown in Figure 2 of the main text. In the ZIP archive, each tubule image has two files associated with it: a REC file that contains the tomogram reconstruction data and an MDOC file that contains imaging metadata. REC files can be opened with the open-source software IMOD.</p> <p>We provide raw image data of pitch- and width-controlled tubules that have been labeled with gold nanoparticles. These accompany the representative images in Figure 4 in the main text. (Pitch Controlled 4-color with GNPs.zip, Width Controlled 4-color with GNPs.zip).</p> <p>We provide raw image data of length-controlled tubules. These images accompany Figure 5 in the main text. (Length Controlled Tubule Images.zip)</p> <p><strong>Associated publication citation:</strong></p> <div> <p><span>Thomas E. Videb&aelig;k&nbsp;<em>et al.,&nbsp;</em></span><span>Economical routes to size-specific assembly of self-closing structures. </span><span><em>Sci. Adv. </em></span><span><strong>10</strong>, </span><span>eado5979 </span><span>(2024). </span><span>DOI:<a href="https://doi.org/10.1126/sciadv.ado5979">10.1126/sciadv.ado5979</a></span></p> </div>

opencc-by-4.0Nov 2023View details →
zenodo48/100

Genome, repeat, and functional annotation associated with the naked mole-rat genome assembly, mHetGlaV3 (GCA_964261345.1)

<p>The naked mole-rat (NMR; Heterocephalus glaber) is a eusocial subterranean rodent with a highly unusual set of physiological traits, such as extreme longevity, that has attracted great interest amongst the scientific community. However, the genetic basis of most of these traits has not been elucidated. To facilitate our understanding of the molecular mechanisms underlying NMR physiology and behaviour, we generated a long-read chromosomal-level genome assembly of the NMR. This genome, mHetGlaV2, was subsequently annotated and incorporated into a &ldquo;91 eutherian mammals&rdquo; multiple whole genome alignment in Ensembl.&nbsp;</p> <p>We identified intra-chromosomal misassemblies within mHetGlaV2. We fixed these misassemblies by comparing syntenic blocks between this assembly and the Canadian Porcupine (EreDor) genome assembly (https://www.ncbi.nlm.nih.gov/datasets/genome/GCA_028451465.1/) and a FISH-Karyotype of the naked mole-rat completed by Romanenko et al., 2023 (PMID: 380307020) to address any misassemblies and place centromeres. Chromosome numbering was identified from a composite karyogram of karyotypes from over 350 cells.&nbsp;This scaffold-corrected assembly is labelled mHetGlaV3 (https://www.ebi.ac.uk/ena/browser/view/GCA_964261345.1).</p> <p>This repository stores the repeat, genome, and epigenome annotations for HetGlaV3.</p> <p>mHetGlaV3.primary.gtf.gz. Gene structures and gene symbols are transferred from ENSEMBL annotations of mHetGlaV2 using liftOff with default parameters. Additional gene symbols were identified using TOGA and manual curation.</p> <p>mHetGlaV3.primary.gtf.gz. Simple repetitive regions and transposable elements were annotated using EarlGrey (https://github.com/TobyBaril/EarlGrey) using "Rodentia" annotations for RepeatMasker.</p> <p>mHetGlaV3.primary.genesymbol_table.txt.txt.gz. A tab-delimited file where rows are gene IDs and columns are gene symbols generated with each method. "Consensus" shows the best matching gene symbol for each gene ID.</p> <p>mHetGlaV3.primary_annotated_blacklist.bed.gz. Provides an assembly "blacklist" for mHetGlaV3. This blacklist is a bed file annotating assembly breakpoints between HetGlaV2 and HetGlaV3. This blacklist contains additional columns (e.g., closest gene, overlapping TE etc.) and should therefore be filtered to the first column before being incorporated into traditional genomic pipelines.</p> <p>mHetGlaV3.primary_hypothalamus_ABC_enhancer.bedpe.gz. Activity-By-Contact enhancers (https://github.com/broadinstitute/ABC-Enhancer-Gene-Prediction) generated in the female subordinate naked mole-rat hypothalamus using Hi-C-seq, ChIP-seq of H3K27Ac data, ATAC-seq, and RNA-seq information.</p> <p>mHetGlaV3.primary_hypothalamus_chromHMM.bed.gz. Chromatin states (using Chromhmm) annotating the female subordinate naked mole-rat hypothalamus using H3K4me3 (promoter), H4K4me2 (promoter-enhancer), H3K27Ac (active enhancer), H3K36me3 (elongated), H3K27me3 (polycomb repressed), H3K9me3 (heterochromatin), and CTCF (whole brain) ChIP-seq data, as well as ATAC-seq and RNA-seq data.</p> <p>mHetGlaV3.primary.fa.gz. Genome assembly fasta file for the naked mole-rat (V3, primary assembly). This assembly matches the primary assembly stored on ENA, however the chromosome names match these files, rather than have chromosome names processed by ENA (e.g. chr 1 instead of "OZ179169.1 Heterocephalus glaber genome assembly, chromosome: 1").</p> <p>&nbsp;</p> <p>UPDATES:</p> <p>* The 1.2 update fixed unscaffolded contig names from those used in-lab to those compatible with ENA.</p> <p>* The 1.3 update added small (50~100kbp) contigs onto mHetGlaV3.primary.fa.gz that were filtered before the ENA submission.</p> <p>* The 1.4 update fixed a small chromosome naming inconsistency spotted in the 1.3 update.</p>

opencc-by-4.0Nov 2024View details →
zenodo48/100

Bolaform Surfactant-Induced Au Nanoparticle Assemblies for Reliable Solution-Based Surface-Enhanced Raman Scattering Detection

<p>Related publication: Garc&iacute;a-Lojo, D; M&eacute;ndez-Merino, D; P&eacute;rez-Juste, I; Acu&ntilde;a, A; Garc&iacute;a-R&iacute;o, L; Rodr&iacute;guez-Pat&oacute;n, A; Pastoriza-Santos, I; P&eacute;rez-Juste, J. Bolaform surfactant-induced Au nanoparticle assemblies for reliable solution-based SERS detection. Adv.Mater. Technol. 2022, 2101726. <a href="https://doi.org/10.1002/admt.202101726">https://doi.org/10.1002/admt.202101726</a></p> <p>&nbsp;</p> <p>&nbsp;</p> <p>Abstract:</p> <p>Solution-based surface-enhanced Raman scattering (SERS) detection typically involves the aggregation of citrate-stabilized Au nanoparticles into colloidal assemblies. Although this sensing methodology offers excellent prospects for sensitivity, portability, and speed, it is still challenging to control the assembly process by a salting-out effect, which affects the reproducibility of the assemblies and, therefore, the reliability of the analysis. This work presents an alternative approach that uses a bolaform surfactant, B<sub>20</sub>, to induce the plasmonic assembly. The decrease of the surface charge and the bridging effect, both promoted by the adsorption of B<sub>20</sub>, are hypothesized as the key points governing the assembly. Furthermore, molecular dynamic simulations supported the bridging effect of the B<sub>20</sub>&nbsp;by showing the preferential bridging of surfactant monomers between two adjacent Au(111) slabs. The colloidal assemblies showed excellent SERS capabilities towards the rapid, on-site detection and quantification of beta-blockers and analgesic drugs in the nanomolar regime, with a portable Raman device. Interestingly, the application of state-of-the-art convolutional neural networks, such as ResNet, allows a 100% accuracy in classifying the concentration of different binary mixtures. Finally, the colloidal approach was successfully implemented in a millifluidic chip allowing the automation of the whole process, as well as improving the performance of the sensor in terms of speed, reliability, and reusability without affecting its sensitivity.</p>

opencc-by-4.0Apr 2022View details →
zenodo48/100

Explosive networking: the role of adaptive host radiations and ecological opportunity in a species-rich host-parasite assembly

<p>Dataset for Cruz-Laufer et al. (2021) Explosive networking: the role of adaptive host radiations and ecological opportunity in a species-rich host-parasite assembly.</p> <p><strong>Abstract: </strong>Many species-rich ecological communities emerge from adaptive radiation events. The effects of this explosive speciation on community assembly remain poorly understood. Here, we explore the well-documented radiations of African cichlid fishes and their interactions with the flatworm gill parasites <em>Cichlidogyrus </em>spp., including 10529 reported infections and 477 different host-parasite combinations collected through a survey of peer-reviewed literature. We assess how evolutionary, ecological, and morphological parameters determine host-parasite meta-communities affected by adaptive radiation events through network metrics, host repertoire measures, and network link prediction. The hosts&rsquo; evolutionary history mostly determined host repertoires of the parasites. Ecological and evolutionary parameters determined host-parasite interactions. Generally, ecological opportunity and fitting have shaped cichlid-<em>Cichlidogyrus</em> meta-communities suggesting an invasive potential for hosts used in aquaculture. Meta-communities affected by adaptive radiations are increasingly specialised with higher environmental stability. These trends should be verified across other systems to infer generalities in the evolution of species-rich host-parasite networks.</p>

opencc-by-4.0Jan 2022View details →
zenodo48/100

Bin-assembled Escherichia coli genomes from a study in Punjab, Pakistan

<h2>Bin-assembled <em>Escherichia coli</em> genomes from Punjab, Pakistan</h2> <p>These assemblies are a part of a cross-sectional study conducted in Punjab, Pakistan aimed at investigating <em>E. coli</em> colonisation diversity in healthy carriage with the use of CLED enrichment plates.</p> <h3><strong>About</strong></h3> <h4><strong>Version history</strong></h4> <p><strong>v0.1.1 (current version)</strong></p> <ul> <li>Added reference to the study.</li> </ul> <p><strong>v0.1.0</strong></p> <ul> <li>Added brief description with a few missing parts.</li> </ul> <h4><strong>Distribution</strong></h4> <p>If you use these assemblies in your study please cite the source as appropriate.&nbsp;These assemblies are made available under a CC-BY 4.0 license.</p> <h4><strong>Citation</strong></h4> <p>Khawaja, T., M&auml;klin, T., Kallonen, T. et al. Deep sequencing of <em>Escherichia coli</em> exposes colonisation diversity and impact of antibiotics in Punjab, Pakistan. Nature Communications 15, 5196 (2024).&nbsp;<a href="https://doi.org/10.1038/s41467-024-49591-5">https://doi.org/10.1038/s41467-024-49591-5</a></p> <h3><strong>Methods briefly</strong></h3> <h4><strong>Species identification</strong></h4> <p>Sequencing data from the ENA project <a href="https://www.ebi.ac.uk/ena/browser/view/PRJEB36642">PRJEB36642</a> was error-corrected with <a href="https://github.com/opengene/fastp">fastp</a> and pseudoaligned with <a href="https://github.com/algbio/themisto">Themisto</a> against a species-level index (available from <a href="https://doi.org/10.5281/zenodo.6656881">https://doi.org/10.5281/zenodo.6656881</a>). Reads were assigned to species using the <a href="https://doi.org/10.1099%2Fmgen.0.000691">mSWEEP/mGEMS pipeline</a> as described in <a href="https://www.nature.com/articles/s41467-022-35178-5">https://www.nature.com/articles/s41467-022-35178-5</a>.</p> <h4><strong>Lineage identification</strong></h4> <p>Read from the species-level bins were again pseudoaligned with Themisto against an <em>E. coli</em> index (will be made available in a later version). Lineage-level assignment was performed using mSWEEP and mGEMS at the level of <a href="https://genome.cshlp.org/content/29/2/304">PopPUNK</a> sequence clusters. The created bins were screened with <a href="https://github.com/tmaklin/coreutils_demix_check">demix_check</a> and bins that received a score of 1 or 2 were kept. Data in the kept bins were assembled with <a href="https://github.com/tseemann/shovill">shovill</a> and the bin-assembled genomes (BAGs) were quality controlled with <a href="https://genome.cshlp.org/content/25/7/1043">checkm</a> for &gt;= 90% completeness and &lt;= 10% contamination. Finally, BAGs shorter than 4 Mb or longer than 6 Mb were removed.</p> <h3><strong>Contact</strong></h3> <p>Tommi M&auml;klin &lt;tommi'at'maklin.fi&gt;.</p>

opencc-by-4.0Jun 2024View details →
zenodo48/100

Experimental data for "Yu-Shiba-Rusinov bands in a self-assembled kagome lattice of magnetic molecules"

<p>Here, we provide all original data used in the manuscript "Yu-Shiba-Rusinov bands in a self-assembled kagome lattice of magnetic molecules"</p> <p>We acknowledge financial support by the Deutsche Forschungsgemeinschaft (DFG, German Research Foundation) through projects 277101999 (CRC 183, project&nbsp;C03) and FR2726/10-1.</p>

opencc-by-4.0Feb 2024View details →
zenodo48/100

Kiwifruit Genome Assembly Red5 Version PS1.68.5

<p>Draft assembly pseudomolecules plus repeat and gene model&nbsp; annotation files for kiwifruit <em>Actinidia chinensis</em> var. <em>chinensis</em> &#39;Red5&#39;.&nbsp;</p>

opencc-by-4.0Jul 2018View details →
zenodo48/100

Data From: The Oyster River Protocol: A multi assembler and kmer approach for de novo transcriptome assembly.

<p>Characterizing transcriptomes in non-model organisms has resulted in a massive increase in our understanding of biological phenomena. This boon, largely made possible via high-throughput sequencing, means that studies of functional, evolutionary and population genomics are now being done by hundreds or even thousands of labs around the world. For many, these studies begin with a <em>de novo</em> transcriptome assembly, which is a technically complicated process involving several discrete steps. The Oyster River Protocol (ORP), described here, implements a standardized and benchmarked set of bioinformatic processes, resulting in an assembly with enhanced qualities over other standard assembly methods. Specifically, ORP produced assemblies have higher Detonate and TransRate scores and mapping rates, which is largely a product of the fact that it leverages a multi-assembler and kmer assembly process, thereby bypassing the shortcomings of any one approach. These improvements are important, as previously unassembled transcripts are included in ORP assemblies, resulting in a significant enhancement of the power of downstream analysis. Further, as part of this study, I show that assembly quality is unrelated with the number of reads generated, above 30 million reads. Code Availability: The version controlled open-source code is available at <a href="https://github.com/macmanes-lab/Oyster_River_Protocol">https://github.com/macmanes-lab/Oyster_River_Protocol</a>. Instructions for software installation and use, and other details are available at <a href="http://oyster-river-protocol.rtfd.org/">http://oyster-river-protocol.rtfd.org/</a>.</p>

opencc-by-4.0Jul 2018View details →
zenodo48/100

Actinidia eriantha Accession EA01_01 low coverage genome assembly

Draft assembly scaffolds of kiwifruit <i>Actinidia eriantha</i> 'EA01_01'. This is a female vine derived from seed collected on Qi-Yuan Mt., Min-Qing, Fukien on 15/11/75 by Li Lai-Yung, Professor of Subtropical Pomology, University of Fukien, Peoples Republic of China and provided to the New Zealand DSIR in 1975

opencc-by-4.0Jul 2018View details →
zenodo48/100

Dataset related to the publication "Lasing in an Assembled Array of Silver Nanocubes"

<p>This archive contains the raw data used to draw the graphs published in the paper "Lasing in an Assembled Array of Silver Nanocubes".&nbsp;</p>

opencc-by-4.0Oct 2024View details →
zenodo48/100

Geometric Frustration Directs the Self-assembly of Nanoparticles with Crystallized Ligand Bundles

<p>This is the supporting dataset of the publication "Geometric Frustration Directs the Self-assembly of Nanoparticles with Crystallized Ligand Bundles".</p> <p><a href="https://doi.org/10.1021/acs.jpcb.4c04562">https://doi.org/10.1021/acs.jpcb.4c04562</a></p> <p>The description of the dataset can be&nbsp; found in the file README.txt</p>

opencc-by-4.0Oct 2024View details →
zenodo48/100

Metagenome-assembled genomes from Stordalen Mire, Sweden (MAGs v2)

<p><strong>This release (MAGs v2) is a major new version of this metagenome-assembled genome (MAG) set.</strong> All previous releases on this page (which only differ in the metadata) are designated "MAGs v1." The current release (MAGs v2) uses<strong>&nbsp;</strong>CheckM2 v1.0.2 filtering (&ge;70% completeness, &le;10% contamination) to expand this dataset to include <strong>36,419 MAGs</strong>, with the following subcategories:</p> <ul> <li>Cronin_v1:&nbsp; Manually-curated subset of the "Field" category from MAGs v1.</li> <li>Cronin_v2:&nbsp; MAGs from raw bin filtering on the same assemblies used to generate Cronin_v1.</li> <li>Woodcroft_v2:&nbsp; MAGs from raw bin filtering on the same assemblies used to generate the MAGs reported in <a href="https://doi.org/10.1038/s41586-018-0338-1">Woodcroft &amp; Singleton et al. (2018)</a>.</li> <li>SIPS:&nbsp; Updated genomes from samples originating from a stable isotope probing (SIP) incubation experiment by Moira Hough et al. ("SIP" in MAGs v1), re-analyzed due to read truncation and sample linkage issues in MAGs v1.</li> <li>JGI:&nbsp; Expanded set of genomes from the Joint Genome Institute's metagenome annotation pipeline.</li> </ul> <p>&nbsp;</p> <p>FILES:</p> <ul> <li><strong>Emerge_MAGs_v2.tar.gz</strong> - Archive containing the MAG files (.fna).</li> <li><strong>metadata_MAGs_v2_EMERGE.tsv</strong>&nbsp;- Table containing source sample names and accessions, GTDB taxonomy information, CheckM2 quality reports, NCBI GenomeBatch- and MIMAG(6.0)-formatted sample attributes and other metadata for the MAGs.&nbsp;</li> </ul> <p>&nbsp;</p> <p>FUNDING:</p> <p>This research is a contribution of the EMERGE Biology Integration Institute (<a href="https://emerge-bii.github.io">https://emerge-bii.github.io/</a>), funded by the National Science Foundation, Biology Integration Institutes Program, Award # 2022070.</p> <p>This study was also funded by the Genomic Science Program of the United States Department of Energy Office of Biological and Environmental Research, grant #s DE-SC0004632. DE-SC0010580. and DE-SC0016440.</p> <p>We thank the Swedish Polar Research Secretariat and SITES for the support of the work done at the Abisko Scientific Research Station. SITES is supported by the Swedish Research Council's grant 4.3-2021-00164.</p> <p>Data collected at the Joint Genome Institute was generated under the following awards:</p> <ul> <li>The majority of sequencing at JGI was supported by BER Support Science Proposal 503530 (DOI: <a href="https://doi.org/10.46936/10.25585/60001148">10.46936/10.25585/60001148</a>), conducted by the U.S. Department of Energy Joint Genome Institute (<a href="https://ror.org/04xm1d337">https://ror.org/04xm1d337</a>), a DOE Office of Science User Facility supported by the Office of Science of the U.S. Department of Energy under Contract No. DE-AC02-05CH11231.</li> <li>Sequencing of SIP samples was performed under the Facilities Integrating Collaborations for User Science (FICUS) initiative (proposal 503547; award DOI:&nbsp;<a href="https://doi.org/10.46936/fics.proj.2017.49950/60006215">10.46936/fics.proj.2017.49950/60006215</a>) and used resources at the DOE Joint Genome Institute (<a href="https://ror.org/04xm1d337">https://ror.org/04xm1d337</a>) and the Environmental Molecular Sciences Laboratory (<a href="https://ror.org/04rc0xn13">https://ror.org/04rc0xn13</a>), which are DOE Office of Science User Facilities. Both facilities are sponsored by the Office of Biological and Environmental Research and operated under Contract Nos. DE-AC02-05CH11231 (JGI) and DE-AC05-76RL01830 (EMSL).</li> </ul>

opencc-by-4.0Oct 2024View details →
zenodo48/100

Metagenome-Assembled Genomes Abundance & Activity Tables. Environmental Parameters Associated with the dataset.

<p>Lake Mendota, WI, USA, is a temperate lake subject to annual temperature and oxygen fluctuations. Each summer, the water column becomes anoxic (no-oxygen). In 2020, we sampled the lake at weekly intervals, at different depths (5, 10, 15, 20 and 23.5m). For each time+depth sample, we collected metagenomes, viromes and metatranscriptomes. Environmental data profiles were collected on-site for each sampling day.&nbsp;</p> <p>Following standard metagenomic binning best practices, we obtained 431 metagenomes-assembled-genomes (MAGs).</p> <p>This record comprises the microbial abundance and expression table for these MAGs, and the environmental profiles collected each day.</p>

opencc-by-4.0Nov 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record