Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
30,265
datasets available to search
ShareScore release 0.7.1
Dataset results
30,265 results for “Functionality”
functional MRI study on the language stress perception in a foreign language
<p>fMRI dataset of 91 participants during a linguistic task about language stress perception in a foreign language.</p> <p>Participants listened to pairs of words in a foreign language (Spanish) and had to indicate if the words were the same or different. The different pairs differed either by the stress pattern, or by the final vowel.</p> <p>This dataset was divided into two groups: 51 participants with French as native language and 40 with Swiss-German as native language. None of the participants had knowledge of Spanish.</p> <p>This repository respects the BIDS standard (<a href="https://bids.neuroimaging.io/">https://bids.neuroimaging.io/</a>), including all the raw data (func, fmap, anat) and metadata in order to reproduce the processing.</p> <p>These data have been used in two papers:</p> <p>S. Schwab, M. Mouthon, L.B. Jost, J. Salvadori, I. Yakoub, E. Ferreira da Silva, N. Giroud, B. Perriard and J.M. Annoni, Neural correlates of lexical stress processing in a foreign free-stress language; Brain and Behavior (2023)</p> <p>L. Rogenmoser, M. Mouthon, F. Etter, J. Kamber, J.M. Annoni and S. Schwab; The processing of stress in a foreign language modulates functional antagonism between default mode and attention network regions, (submitted)</p>
How to measure work functions from aqueous solutions - data
<p>Data set pertaining to the article "How to measure work functions from aqueous solutions", <a href="https://doi.org/10.1039/D3SC01740K" target="_blank" rel="noopener">https://doi.org/10.1039/D3SC01740K</a> (Chemical Science <strong>14</strong>, 9574-9588 (2023)). A new protocol for energy referencing of photoemission data from liquids (<a href="https://doi.org/10.1039/D1SC01908B" target="_blank" rel="noopener">https://doi.org/10.1039/D1SC01908B</a>, Chemical Science <strong>12</strong>, 10558-10582 (2021)) is refined towards determining work functions from liquids.<br><br></p> <p>Files with extension .h5 are hdf5-files structured according to the NeXus v2020.10 standard using the NXmpes user contributed format suggested by the Fairmat consortium, see<br>https://www.nexusformat.org/<br>https://fairmat-experimental.github.io/nexus-fairmat-proposal/50433d9039b3f33299bab338998acb5335cd8951/mpes-structure.html<br>A few extensions specific to liquid jet-experiments were added to the standard, and are explained in the notes-group on the top level of each file.<br>NeXus data files can be opened with any software capable of opening hdf5-structured files. The following viewers are adapted to the specifics of the NeXus data format:<br>* nexpy (distributed with python)<br>* https://h5web.panosc.eu/h5wasm (web-based NeXus viewer maintained by the European Photon and Neutron Open Science Cloud-consortium)</p> <p>In each NeXus file-entry, two types of spectra are included:<br>1. Sweep-averaged spectra, integrated over the non-dispersive coordinate of our detector ('data').<br>2. As-measured data ('raw').</p> <p>Files with extension .txt are tab-separated ascii-files.</p> <p><br>The following files are provided:</p> <p>Photoemission data pertaining to solute measurements and reference measurements using a gold wire:<br>'Figure 3.h5'<br>'Figure 4.h5'<br>'Figure S1.h5'<br>'Figure S2.h5'<br>Kinetic energies are presented as measured. The scale offset of our spectrometer, determined as E_kin(corrected) = E_kin(measured) + 0.224 eV for data sets 'Figure 3.h5', 'Figure 4.h5' ,'Figure S2.h5', has not been taken into account.</p> <p>Numeric representations of the analysis results shown in the article's figures in graphical form:<br>'Figure 5.txt'<br>'Figure 6B.txt'<br>'Figure 7.txt'<br>'Figure S4B.txt'<br>'Figure S5.txt'</p> <p>In case you have any questions regarding this data set please contact: Uwe Hergenhahn, uhe@fhi.mpg.de .</p>
The Biomass and Plant Functional Traits of Leymus chinensis Affected by Genotypic Diversity and Soil Nitrogen Addition through a Two-year Experiment, Tianjin, China, 2021-2023
In order to investigate the effects of soil nitrogen addition on the genotypic diversity of Leymus chinensis, 12 genotypes of Leymus chinensis were used as plant material and a two-factor experimental design was carried out in this study. Factor one was genotypic diversity of L. chinensis, including three levels: mono-genotype (G1), three genotypes (G3), and six genotypes (G6). Factor two was the soil nitrogen addition level, which included four levels: no nitrogen addition (N0), 2.5 g N/(m²·a) nitrogen application (N2.5), 5 g N/(m²·a) nitrogen application (N5), and 10 g N/(m²·a) nitrogen application (N10). Each treatment had 12 combinations as replicates, and 12 genotypes of L. chinensis were used. The frequency of each genotype was standardized across all treatment levels of genotypic diversity × soil nitrogen addition. The experiment commenced in September 2021 and soil nitrogen was applied every 2 months. Plants were cultivated in the experimental field at Nankai University, but were moved to a greenhouse for overwintering from November to February each year. During the experiment, there were no stresses or disturbances such as shading, drought, or insect feeding; weeds were regularly removed.
Block summaries of biomass, carbon, nitrogen, and phosphorus allocation among tissue types, species, and plant functional types from Arctic LTER 1981 Moist Acidic Tussock (MAT81) long-term experiment harvests: 2000 and 2015, Toolik Lake Field Station, Alaska.
A complete accounting of biomass, C, N, and P allocation both among tissue types (leaves, stems, rhizomes, roots) and among species and plant functional types from Arctic LTER 1981 Moist Acidic Tussock (MAT81) long-term experiment’s untreated control plots and plots that were fertilized annually, harvested after 20 and 35 years, near Toolik Lake Field Station, Alaska. Data are gram per meter squared summarized by block.
MCR LTER: Coral Reef: Patterns and implications of spatial covariation in herbivore functions on resilience of coral reefs
These data and code were generated in support of the manuscript: Cook DT, Holbrook SJ, and Schmitt RJ, Scientific Reports. In 2017, we collected biological and physical data from 20 sites along the north shore of Moorea, French Polynesia, to investigate spatial patterns in grazing and browsing functions of herbivorous fishes, environmental correlates, and implications for coral resilience. In addition to the data collected at the 20 north shore sites, we conducted a 10-day field experiment to assess the relationship between browsing intensity and potential of reversing a coral-to-macroalgae shift. This material uses data collected by the U.S. National Science Foundation's (NSF) Moorea Coral Reef Long Term Ecological Research (MCR LTER) site under Grant No. OCE 2224354 (and earlier awards). Additional financial support to the MCR LTER site was provided through a generous gift from the Gordon and Betty Moore Foundation. Research was completed under permits issued by the French Polynesian Government (Délégation à la Recherche) and the Haut-commissariat de la République en Polynésie Francaise (DTRT) (Protocole d'Accueil 2005-2025).
Functional Trait Measurements of Macroalgal Communities in the Santa Barbara Channel
This dataset contains trait and elemental composition data for macroalgal samples collected across depth gradients at multiple sites in the Santa Barbara Channel, California. Each sample represents an individual specimen characterized by morphological measurements (e.g., blade thickness, stipe diameter, total height), biomass of anatomical parts (blade, stipe, holdfast, reproductive tissue), and anchoring strength. In addition, biochemical traits—including carbon (C), nitrogen (N), and hydrogen (H) content—were measured from tissue samples analyzed in the analytical laboratory. These data support a trait-based modeling approach to macroalgal community structure and distribution, contributing to our understanding of functional diversity and ecosystem dynamics in temperate marine systems. Accompanying metadata include collection site, date, time, depth, location coordinates, and substrate type, providing context for environmental variation across samples.
Functional Connectivity of Music-Induced Analgesia in Fibromyalgia
Open the record for dataset details and reuse information.
Human es-fMRI Resource: Concurrent deep-brain stimulation and whole-brain functional MRI
Open the record for dataset details and reuse information.
Effects of Phase Regression on High-Resolution Functional MRI of the Primary Visual Cortex
Open the record for dataset details and reuse information.
DATASET: De novo assembly and functional annotation of the heart + hemolymph transcriptome in the Caribbean spiny lobster Panulirus argus
<p>The spiny lobster <em>Panulirus argus</em> is an ecologically relevant species in shallow water coral reefs and target of the most lucrative fishery in the greater Caribbean region. This study reports, for the first time, the heart + hemolymph transcriptome of the Caribbean spiny lobster<em> Panulirus argus</em> assembled from short Illumina 150 bp PE raw reads. A total 80,152,094 raw reads were assembled using the Oyster River Protocol pipeline that aspires to become the standard protocol for <em>de novo</em> transcriptome assembly. The assembly resulted in a total of 254,773 transcripts. Functional gene annotation was conducted using the software package 'dammit' that also aspires to become the standard protocol for <em>de novo</em> transcriptome annotation. Lastly, gene enrichment analyses were conducted using the Gene Ontology (GO), KEGG pathway analyses (Kaas), and KOG (WebMGA) databases. This resource will be of utmost importance in future research aiming at exploring the effect of local and regional anthropogenic disturbances as well as global climate change on the molecular physiology of this overexploited species.</p>
Data set for Global quantitative synthesis of ecosystem functioning across climatic zones and ecosystem types
<p>Dataset used in the publication: " Global quantitative synthesis of ecosystem functioning across climatic zones and ecosystem types". The dataset gathers estimates of ecosystem standing stocks (biomass, organic carbon, detritus), fluxes (GPP, ER, NEP) and process rates (decomposition and carbon uptake rates) for eight broad ecosystem types (forest, grassland, agroecosystem, desert, stream, lake, pelagic and benthic marine ecosystems) in five broad climatic zones (arctic, boreal, arid, temperate, tropical, arid).</p> <p>The scripts to produce the figures and the statistics of the publication are released along with the txt version of the data, which file is uploaded when running the script.</p>
Density functional theory calculations of coherent bcc Fe-Cu interfacial energy densities
<p>File contains the data required to calculate interfacial energy densities of {100}, {110}, {111}, {210}, {211} and {221} orientated coherence bcc Fe-Cu interfaces.</p> <p>Data produced for the study detailed in: Cu nanoprecipitate morphologies and interfacial energy densities in bcc Fe from density functional theory (DFT) A.M. Garrett and C.P. Race.</p> <p>Submitted to Computational Materials Science.</p> <p>.txt files contain the total energies calculated for relaxed interface-containing and bulk simulation cells at a range of interfacial spacings. This data can be used to calculate the size independent interfacial energy densities for a range of Fe-Cu interface orientations using standard fitting approaches. Columns of the tables in the .txt files are no. atoms, interface-containing simulation cell length, interfacial area, total energy of the relaxed interface-containing simulation cell, total energy of the reference bulk Fe and total energy of the reference bulk Cu. Lengths are in Angstrom and energies are in eV.</p>
Data and R code for Tansley review New Phytologist 2021: "An integrated framework of plant form and function: The belowground perspective"
<p>The files in this archive are related to the paper of Weigelt, Mommer, Andraczek et al. (2021) An integrated framework of plant form and function: The belowground perspective. Tansley Review New Phytologist. The paper developed and tested a new conceptual framework of plant form and function linking above and belowground traits of 2510 species. We found that an integrated, whole-plant trait space required as much as four axes. The two main axes represented the fast-slow ‘conservation’ gradient on which leaf and fine-root traits were well aligned, and the ‘collaboration’ gradient in roots. The two additional axes were separate, orthogonal plant size axes for height and rooting depth.</p> <p>This archives contains four files:</p> <ol> <li><strong>Weigelt et al.2021RCode.DataCleaning.txt</strong> - RCode for the complete data processing starting with the downloaded database files from the Plant Trait Database version 5.0 (TRY, Kattge et al. 2020), the Global Root Trait database (GRooT, Guerrero-Ramirez et al. 2020) and a small number of additional data files listed in Table S2 of the original paper. Additional information was later incorporated using FungalRoot Database (Soudzilovkaia et al. 2020), nodDB Database (Tedersoo et al. 2018) and a compiled dataset on rooting depth (Fan et al. 2017). The code processes, cleans and merges the data and produces a final table for PCA analysis of species specific mean traits. This final table is provided as a second file in this archive (Weigelt_et_al_2021_Main.PCA.Matrix.xlsx). A second part of the RCode.DataCleaning extracts species-specific individual trait data where root and shoot traits were measured on the same plant individual or plot. This data was compiled from 43 studies identified in Table S2 of the original publication. The final table for individual trait data is the third file in this archive (Weigelt_et_al_2021_Individual.PCA.Matrix.xlsx).</li> <li><strong>Weigelt_et_al_2021_Main.PCA.Matrix.xlsx</strong> – Datafile with species-specific global mean trait data for 17 traits of 2510 species with at least one root and one shoot trait available. Meta-data is provided in the data file.</li> <li><strong>Weigelt_et_al_2021_Individual.PCA.Matrix.xlsx</strong> – Datafile with species-specific trait data where root and shoot traits were measured on the same individual or plot for 6 traits of 455 species. Meta-data is provided in the data file.</li> <li><strong>Weigelt et al.2021RCode.Analysis.txt – </strong>RCode for all analyses and figures provided in the paper for both the species mean and individual based dataset. The Code is annotated to help reproducibility of the analysis.</li> </ol>
Temperature-related mortality exposure-response functions for 854 cities in Europe
<p>This repository contains data to reconstruct the exposure-response functions (ERF) of temperature-related mortality by five 5 age groups in 854 cities in Europe.</p><p>These ERFs have been derived in the study by Masselot et al. 2023, <i>Excess mortality attributed to heat and cold: a health impact assessment study in 854 cities in Europe</i>, The Lancet Planetary Health (<a href="https://protect-eu.mimecast.com/s/zqg2Cg204i4ZMYKf3NUKN?domain=doi.org">https://doi.org/10.1016/S2542-5196(23)00023-2</a>). An associated semi-replicable GitHub repository is available at <a href="https://github.com/PierreMasselot/Paper--2023--LancetPH--EUcityTRM">https://github.com/PierreMasselot/Paper--2023--LancetPH--EUcityTRM</a> to reproduce part of the analysis and the full results, as well as to provide technical details on the derivation of these ERFs.</p><p><strong>Note: </strong>This updated version contains revised data after the correction of an error in the code related to the computation of the age-specific baseline mortality rates. Details about the error can be found in the GitHub repository linked above. This correction only affects the figures of excess mortality (found in the `results.zip` archive) while the ERFs are negligibly affected. The originally published results can be found in V1.0.0 of this repository.</p><p><strong>Extraction of the ERFs</strong></p><p>The ERFs are provided as coefficients of B-spline functions that can be used to reconstruct the ERFs, along with variance-covariance matrices and quantiles from location-specific temperature distributions. The parametrisation associated with these coefficients is a quadratic B-spline (degree 2), with knots located at the 10th, 75th and 90th percentiles of the temperature distribution. In R, the associated basis can be constructed using the <i>dlnm</i> package, with a temperature series <i>x</i>, as follows:</p><blockquote><p>library(dlnm) </p><p>basis <- onebasis(x, fun = "bs", degree = 2, knots = quantile(x, c(.1, .75, .9)))</p></blockquote><p>The main files associated with ERFs are the following:</p><p><i>coefs.csv</i>: The B-spline coefficients for each age group and city.</p><p><i>vcov.csv</i>: The variance-covariance matrix of the coefficients in each city and age group. It is provided here as the lower triangular part of the matrix with names indicating the position of each value (v[row][column]). In R, assuming <i>x</i> is a row of this file, the matrix can be reconstructed using <i>xpndMat(x)</i> after loading the <i>mixmeta</i> package.</p><p><i>coef_simu.csv</i>: 1000 simulations from the distribution of each city and age-specific coefficients. Useful to derive empirical confidence intervals for derived measures such as excess deaths or attributable fractions.</p><p><i>tmean_distribution.csv</i>: The city-specific temperature percentiles representing the distribution of the data derived from the ERA5-Land dataset.</p><p><strong>Health impact assessment results</strong></p><p><i>results.zip</i>: A summary of the results from the health impact assessment reported in the analysis. The dataset includes several impact measures provided in files representing different geographical levels, including city, country and regional level. Different files are also provided for age-group specific or all age results.</p><p><strong>Additional data</strong></p><p>We provide additional data that are useful to reproduce or extend the analysis. Please note that due to restrictive data-sharing agreements for the mortality series, only a part of the code is reproducible. See the <a href="https://github.com/PierreMasselot/Paper--2023--LancetPH--EUcityTRM">associated GitHub repository</a> for more details.</p><p><i>metadata.csv</i>: City-specific metadata used to create the ERFs and perform the health impact assessment.</p><p><i>additional_data.zip</i>: contains further data used to replicate the second stage of the analysis and the final health impact assessment. It includes the full city-level daily temperature series (<i>era5series.csv</i>), the detail of extracted metadata for available years (<i>metacityyear.csv</i>), a description of the city-level characteristics (<i>metadesc.csv</i>), and the first-stage ERF coefficients for all available city and age-groups (<i>stage1res.csv</i>). Additionally, the file <i>meta-model.RData</i> contains R object defining the second-stage model that can be used to predict new ERFs. </p>
MarFERReT: an open-source, version-controlled reference library of marine microbial eukaryote functional genes
<p>Metatranscriptomics generates large volumes of sequence data about transcribed genes in natural environments. Taxonomic annotation of these datasets depends on availability of curated reference sequences. For marine microbial eukaryotes, current reference libraries are limited by gaps in sequenced organism diversity and barriers to updating libraries with new sequence data, resulting in taxonomic annotation of only about half of eukaryotic environmental transcripts. Here, we introduce version 1.0 of the Marine Functional EukaRyotic Reference Taxa (MarFERReT), an updated marine microbial eukaryotic sequence library with a version-controlled framework designed for taxonomic annotation of eukaryotic metatranscriptomes. We gathered 902 marine eukaryote genomes and transcriptomes from multiple sources and assessed these candidate entries for sequence quality and cross-contamination issues, selecting 800 validated entries for inclusion in the library. MarFERReT v1 contains reference sequences from 800 marine eukaryotic genomes and transcriptomes, covering 453 species- and strain-level taxa, totaling nearly 28 million protein sequences with associated NCBI and PR2 Taxonomy identifiers and Pfam functional annotations. An accompanying MarFERReT project repository hosts containerized build scripts, documentation on installation and use case examples, and information on new versions of MarFERReT.<br><br>MarFERReT is linked to a code repository hosting containerized build scripts, documentation on installation and use case examples, and information on new versions of MarFERReT here: <a href="https://github.com/armbrustlab/marferret">https://github.com/armbrustlab/marferret</a></p> <p>The raw source data for the 902 candidate entries considered for MarFERReT v1.1.1, including the 800 accepted entries, are available for download from their respective online locations. The source URL for each of the entries is listed here in MarFERReT.v1.1.1.entry_curation.csv, and detailed instructions and code for downloading the raw sequence data from source are available in the MarFERReT code repository (<a href="https://github.com/armbrustlab/marferret/blob/main/docs/process_clean_marmicrodb.log.sh">link</a>). </p> <p>This repository release contains MarFERReT database files from the v1.1.1 MarFERReT release using the following MarFERReT library build scripts: <strong>assemble_marferret.sh</strong>, <strong>pfam_annotate.sh</strong>, and <strong>build_diamond_db.sh</strong><br><br>The following MarFERReT data products are available in this repository:</p> <p><strong>MarFERReT.v1.1.1.metadata.csv</strong><br>This CSV file contains descriptors of each of the 902 database entries, including data source, taxonomy, and sequence descriptors. Data fields are as follows:</p> <ol> <li><strong>entry_id</strong>: Unique MarFERReT sequence entry identifier.</li> <li><strong>accepted: </strong>Acceptance into the final MarFERReT build (Y/N). The Y/N values can be adjusted to customize the final build output according to user-specific needs.</li> <li><strong>marferret_name</strong>: A human and machine friendly string derived from the NCBI Taxonomy organism name; maintaining strain-level designation wherever possible.</li> <li><strong>tax_id</strong>: The NCBI Taxonomy ID (taxID).</li> <li><strong>pr2_accession</strong>: Best-matching PR2 accession ID associated with entry</li> <li><strong>pr2_rank</strong>: The lowest shared rank between the entry and the pr2_accession</li> <li><strong>pr2_taxonomy</strong>: PR2 Taxonomy classification scheme of the pr2_accession</li> <li><strong>data_type</strong>: Type of sequence data; transcriptome shotgun assemblies (TSA), gene models from assembled genomes (genome), and single-cell amplified genomes (SAG) or transcriptomes (SAT).</li> <li><strong>data_source</strong>: Online location of sequence data; the Zenodo data repository (<a href="../">Zenodo</a>), the datadryad.org repository (<a href="http://datadryad.org/">datadryad.org</a>), MMETSP re-assemblies on Zenodo (MMETSP)17, NCBI GenBank (<a href="https://www.ncbi.nlm.nih.gov/genbank/">NCBI</a>), JGI Phycocosm (<a href="https://phycocosm.jgi.doe.gov/phycocosm/home">JGI-Phycocosm</a>), the TARA Oceans portal on Genoscope (<a href="http://www.genoscope.cns.fr/tara/">TARA</a>), or entries from the Roscoff Culture Collection through the METdb database repository (<a href="https://metdb.sb-roscoff.fr/metdb/">METdb</a>).</li> <li><strong>source_link</strong>: URL where the original sequence data and/or metadata was collected.</li> <li><strong>pub_year</strong>: Year of data release or publication of linked reference.</li> <li><strong>ref_link</strong>: Pubmed URL directs to the published reference for entry, if available.</li> <li><strong>ref_doi</strong>: DOI of entry data from source, if available.</li> <li><strong>source_filename</strong>: Name of the original sequence file name from the data source.</li> <li><strong>seq_type</strong>: Entry sequence data retrieved in nucleotide (nt) or amino acid (aa) alphabets.</li> <li><strong>n_seqs_raw</strong>: Number of sequences in the original sequence file.</li> <li><strong>source_name:</strong> Full organism name from entry source</li> <li><strong>original_taxID</strong>: Original NCBI taxID from entry data source metadata, if available</li> <li><strong>alias:</strong> Additional identifiers for the entry, if available</li> </ol> <p><br><strong>MarFERReT.v1.1.1.curation.csv</strong><br>This CSV file contains curation and quality-control information on the 902 candidate entries considered for incorporation into MarFERReT v1, including curated NCBI Taxonomy IDs and entry validation statistics. Data fields are as follows:</p> <ol> <li><strong>entry_id:</strong> Unique MarFERReT sequence entry identifier</li> <li><strong>marferret_name: </strong>Organism name in human and machine friendly format, including additional NCBI taxonomy strain identifiers if available.</li> <li><strong>tax_id</strong>: Verified NCBI taxID used in MarFERReT</li> <li><strong>taxID_status</strong>: Status of the final NCBI taxID (Assigned, Updated, or Unchanged)</li> <li><strong>taxID_notes</strong>: Notes on the original_taxID</li> <li><strong>n_seqs_raw</strong>: Number of sequences in the original sequence file</li> <li><strong>n_pfams</strong>: Number of Pfam domains identified in protein sequences</li> <li><strong>qc_flag</strong>: Early validation quality control flags for the following: LOW_SEQS; less than 1,200 raw sequences; LOW_PFAMS; less than 500 Pfam domain annotations.</li> <li><strong>flag_Lasek</strong>: Flag notes from Lasek-Nesselquist and Johnson (2019); contains the flag 'FLAG_LASEK' indicating ciliate samples reported as contaminated in this study.</li> <li><strong>VV_contam_pct</strong>: Estimated contamination reported for MMETSP entries in Van Vlierberghe et al., (2021).</li> <li><strong>flag_VanVlierberghe: </strong>Flag for a high level of estimated contamination, from 'flag_VanVlierberghe' values over 50%: FLAG_VV.</li> <li><strong>rp63_npfams</strong>: Number of ribosomal protein Pfam domains out of 63 total.</li> <li><strong>rp63_contam_pct</strong>: Percent of total ribosomal protein sequences with an inferred taxonomic identity in any lineage other than the recorded identity, as described in the Technical Validation section from analysis of 63 Pfam ribosomal protein domains.</li> <li><strong>flag_rp63</strong>: Flag for a high level of estimated contamination, from 'rp63_contam_pct' values over 50%: FLAG_RP63.</li> <li><strong>flag_sum: </strong>Count of the number of flag columns (`qc_flag`, `flag_Lasek`, `flag_VanVlierberghe`, and `flag_rp63`). All entries with one or more flag are nominally rejected ('accepted' = N); entries without any flags are validated and accepted ('accepted' = Y).</li> <li><strong>accepted: </strong>Acceptance into the final MarFERReT build (Y or N).</li> </ol> <p> </p> <p><strong>MarFERReT.v1.1.1.proteins.faa.gz</strong><br>This Gzip-compressed FASTA file contains the 27,951,013 final translated and clustered protein sequences for all 800 accepted MarFERReT entries. The sequence defline contains the unique identifier for the sequence and its reference (mftX, where 'X' is a ten-digit integer value). </p> <p> </p> <p><strong>MarFERReT.v1.1.1.taxonomies.tab.gz</strong><br>This Gzip-compressed tab-separated file is formatted for interoperability with the DIAMOND protein alignment tool commonly used for downstream analyses and contains some columns without any data. Each row contains an entry for one of the MarFERReT protein sequences in MarFERReT.v1.proteins.faa.gz. Note that 'accession.version' and 'taxid' are populated columns while 'accession' and 'gi' have NA values; the latter columns are required for back-compatibility as input for the DIAMOND alignment software and LCA analysis. </p> <p>The columns in this file contain the following information:</p> <ol> <li><strong>accession</strong>: (NA)</li> <li><strong>accession.version</strong>: The unique MarFERReT sequence identifier ('mftX').</li> <li><strong>taxid</strong>: The NCBI Taxonomy ID associated with this reference sequence.</li> <li><strong>gi</strong>: (NA).</li> </ol> <p> </p> <p><strong>MarFERReT.v1.1.1.proteins_info.tab.gz</strong><br>This Gzip-compressed tab-separated file contains a row for each final MarFERReT protein sequence with the following columns:</p> <ol> <li><strong>aa_id</strong>: the unique identifier for each MarFERReT protein sequence.</li> <li><strong>entry_id</strong>: The unique numeric identifier for each MarFERReT entry.</li> <li><strong>source_defline</strong>: The original, unformatted sequence identifier</li> </ol> <p> </p> <p><strong>MarFERReT.v1.1.1.best_pfam_annotations.csv.gz<br></strong>This Gzip-compressed CSV file contains the best-scoring Pfam annotation for intra-species clustered protein sequences from the 800 validated MarFERReT entries; derived from the hmmsearch annotations against Pfam 34.0 functional domains. This file contains the following fields:</p> <ol> <li><strong>aa_id</strong>: The unique MarFERReT protein sequence ID ('mftX').</li> <li><strong>pfam_name</strong>: The shorthand Pfam protein family name.</li> <li><strong>pfam_id</strong>: The Pfam identifier.</li> <li><strong>pfam_eval</strong>: hmm profile match e-value score</li> <li><strong>pfam_score:</strong> hmm profile match bitscore</li> </ol> <p><br><strong>MarFERReT.v1.1.1.dmnd</strong><br>This binary file is the indexed database of the MarFERReT protein library with embedded NCBI taxonomic information generated by the DIAMOND makedb tool using the build_diamond_db.sh script from the MarFERReT /scripts/ library. This can be used as the reference DIAMOND database for annotating environment sequences from eukaryotic metatranscriptomes. <br><br></p>
Data for manuscript: Functional Protein Dynamics in a Crystal
<p>The data is provided as a part of the manuscript "<strong>Functional Protein Dynamics in a Crystal</strong>". This repository includes an archive with folders:<br> <br> <strong>md_data</strong></p> <ul> <li>contains various simulation systems (crystal supercell, apo and ligand-bound solution) built from the crystal structure of the PDZ domain (PDB ID: 5E11) and carried out using three force fields: Amber ff14SB, CHARMM36m, Amber ff94. <em>The details of the simulations are provided in the Methods and Supplementary methods sections of the manuscript. </em></li> </ul> <p><strong>fig_data</strong></p> <ul> <li>contains the data sets underlying Figures 1-5 of the manuscript's main text. </li> </ul> <p> </p>
Cherri - Accurate detection of functional RNA-RNA interactions sites
<p><strong>CheRRI</strong> - Pipeline for the Identification of putative RNA-RNA interaction sites.</p> <p> </p> <p>This repository contains all CheRRI's models computed and mentioned in the content.txt, listing data and their descriptions. All models can be used to classify interaction sites in CheRRI's eval mode.</p> <p> </p> <p>The source code for CheRRI is avalbile on <a href="https://github.com/BackofenLab/Cherri#install-cherri-conda-package">GitHub</a> and can be cited using this Software Heritage citation:</p> <ul> <li><span>Müller T, Mautner S, Videm P, Eggenhofer F, Raden M, Backofen R (2024) CheRRI - Accurate classification of the biological relevance of putative RNA-RNA interaction sites (Version 0.8). [Computer software]. Software Heritage, <a href="https://archive.softwareheritage.org/swh:1:snp:ebac091117f9c46fb5f0fedd3ef23ec2905ced6c;origin=https://github.com/BackofenLab/Cherri">https://archive.softwareheritage.org/swh:1:snp:ebac091117f9c46fb5f0fedd3ef23ec2905ced6c;origin=https://github.com/BackofenLab/Cherri</a></span></li> </ul> <div> <div> <div> <p>The pipeline contains Machine Learning segments which were annotated using DOME:</p> </div> </div> </div> <ul> <li><span><a href="https://dome.ds-wizard.org/projects/74d0e01c-6374-41e9-93b8-2889d6a8fe25">https://dome.ds-wizard.org/projects/74d0e01c-6374-41e9-93b8-2889d6a8fe25</a></span></li> </ul>
Divergent evolution of sleep functions - Joyce et al 2024 Dataset - Part 1 of 4
<p>This is the full experimental dataset associated to "Divergent evolution of sleep functions" by Joyce et al 2024, Nature Communications. See https://lab.gilest.ro/papers/divergent-evolution-of-sleep-functions/ for more information.</p> <p>The dataset contains 86 zip files for a total of about 170GB. Once uncompressed, they will explode 622 ethoscope db files for a total of 330Gb, covering behavioural analysis of more than 11.000 animals. These are the RAW data as collected from the ethoscopes. A separate zip archive with all the R/Python scripts and the metadata is also provided. This also contains confocal images for BRP analysis.</p> <ul> <li>Part 1 of 4: <a href="https://doi.org/10.5281/zenodo.10554851" target="_blank" rel="noopener">10.5281/zenodo.10554851 (this page) </a></li> <li>Part 2 of 4: <a href="https://doi.org/10.5281/zenodo.10557238" target="_blank" rel="noopener">10.5281/zenodo.10557238</a></li> <li>Part 3 of 4: <a href="https://doi.org/10.5281/zenodo.10557310" target="_blank" rel="noopener">10.5281/zenodo.10557310</a></li> <li>Part 4 of 4: <a href="https://doi.org/10.5281/zenodo.10966461" target="_blank" rel="noopener">10.5281/zenodo.10966461 </a></li> </ul>
Inter-Chemical Correlation results for the study: HHEARx2017-1729 (Air Pollution, Placenta Function, and Birth Outcomes in Los Angeles)
Title: Air Pollution, Placenta Function, and Birth Outcomes in Los Angeles <br>Species: Homo sapiens <br>Number of samples: 450 <br>Number of named analytes: 14 <br>Datasource url: https://hheardatacenter.mssm.edu/PublicFile/ViewPublicFile?projectid=48 <br>
Inter-Chemical Correlation results for the study: HHEARx2016-1534 (A Nested Case-Control Study of Prenatal Exposure to Phthalates and Psychosocial Stress: Adverse Pregnancy Outcomes and the Mediating Role of Placental Function)
Title: A Nested Case-Control Study of Prenatal Exposure to Phthalates and Psychosocial Stress: Adverse Pregnancy Outcomes and the Mediating Role of Placental Function <br>Species: Homo sapiens <br>Number of samples: 5789 <br>Number of named analytes: 17 <br>Datasource url: https://hheardatacenter.mssm.edu/PublicFile/ViewPublicFile?projectid=14 <br>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.