Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,019
datasets available to search
ShareScore release 0.7.1
Dataset results
1,019 results for “Not assigned”
Genetic assignments for Spring Evolutionary Significant Unit reanalysis, Central Valley Chinook Salmon populations, CA, 2011-2024
Central Valley Chinook Salmon populations differ in their Endangered Species Act listing status. It is difficult to visually distinguish individuals from the different Evolutionarily Significant Units (ESU). As such, many of the salmon monitoring and evaluation efforts in the Central Valley and San Francisco Bay-Delta are hampered by uncertainty about population (stock) identification and proportional effects of management actions (Dekar et al. 2013; IEP 2019). Studies have identified that the current identification method (length-at-date models) of juvenile Chinook salmon (Fisher 1992) captured in the watershed vary in their accuracy, particularly for spring-run (NMFS 2013; Harvey et al. 2014; Merz et al. 2014). The inaccuracy of the size-based methods is likely due to differences in fish distribution during early rearing, habitat-specific growth rates, and inter-annual variability in temperatures and food availability that lead to overlap in size ranges among stocks. The primary objective of this project was the genetic classification (to genetic lineage; Evolutionary Significant Unit) of Chinook Salmon captured from State Water Project and Central Valley Project fish protection facilities and Interagency Ecological Program compliance monitoring programs. The genetic lineage was determined for sampled fish by comparing their genotypes to reference genetic baselines. Genetic methods, having less statistical uncertainty that size-based models for population identification, were intended to directly target (and reduce) one source of uncertainty in the estimation of loss (take) from water diversions (operations) and develop the information necessary for understanding stock-specific distribution, habitat utilization, abundance, and life history variation. This project supports recommendations from the Interagency Ecological Program’s Salmon and Sturgeon Assessment of Indicators by Life Stage and Interagency Ecological Program Science Agenda efforts to improve Central V
Integrated freshwater abundance and connectivity clusters at the Hydrologic Unit 8 scale for the Midwest and Northeast U.S.A. – freshwater metric variables and k-means cluster assignment
This dataset includes integrated freshwater abundance and connectivity cluster output, principal component scores, and lake, wetland, and stream abundance and connectivity metrics measured at the Hydrologic Unit 8 (HU8) scale for 17 U.S. states in the Midwest and Northeast regions (appr. 1,800,000 km2). The intent of the cluster analysis is to characterize the macroscale patterns of the integrated freshwater landscape that includes lakes, wetlands, and streams and their surface connectivity attributes. We define freshwater connectivity as the permanent surface hydrologic connections that link lakes, wetlands, and streams and measure connectivity as the landscape position of systems within stream networks. Geographic data used in the analysis are in LAGOS-NE-GEO database v. 1.03 (Lake multi-scaled geospatial and temporal database), an integrated, multi-thematic geographic database (Soranno et al. 2015). The integrated freshwater clusters were created through a multi-step process as follows: 1) we quantified multiple freshwater connectivity metrics for lakes, streams, and wetlands separately, 2) we performed principal components analysis (PCA) on the connectivity metric values for each freshwater type to reduce collinearity, and 3) we performed k-means cluster analysis to group spatial units with similar freshwater connectivity characteristics. The resulting freshwater clusters are representations of the macroscale patterns of freshwater abundance and connectivity in the landscape.
Freshwater connectivity clusters for lakes, wetlands, and streams at the Hydrologic Unit 12 scale in the Midwest and Northeast U.S.A. – freshwater metric variables and K-means cluster assignment
This dataset includes freshwater connectivity cluster output and principal component scores for lakes, wetlands, and streams measured at the Hydrologic Unit 12 (HU12) scale in 17 U.S. states in the Midwest and Northeast regions (appr. 1,800,000 km2). The intent of the cluster analysis is to characterize the macroscale patterns of freshwater connectivity attributes. We define freshwater connectivity as the permanent surface hydrologic connections that link lakes, wetlands, and streams and measure connectivity as the landscape position of systems within stream networks. Geographic data used in the analysis are in LAGOS-NE-GEO database v. 1.03 (Lake multi-scaled geospatial and temporal database), an integrated, multi-thematic geographic database (Soranno et al. 2015). Freshwater connectivity clusters were created separately for lakes, wetlands, and streams through a multi-step process as follows: 1) we quantified multiple freshwater connectivity metrics, 2) we performed principal components analysis (PCA) on the connectivity metric values for each freshwater type to reduce collinearity, and 3) we performed k-means cluster analysis to group spatial units with similar freshwater connectivity characteristics. The resulting freshwater clusters are representations of the macroscale patterns of lake, wetland, and stream connectivity in the landscape.
Chinook Salmon genetic assignments for the Central Valley Project (CVP) and State Water Projects (SWP), Sacramento and San Joaquin Delta Waters, CA, 2024-25
Central Valley Chinook Salmon populations differ in their Endangered Species Act listing status. It is difficult to visually distinguish individuals from the different Evolutionarily Significant Units. As such, many of the salmon monitoring and evaluation efforts in the Central Valley and San Francisco Bay-Delta are hampered by uncertainty about population (stock) identification and proportional effects of management actions (Dekar et al. 2013; IEP 2019). Studies have identified that the current identification method (length-at-date models) of juvenile Chinook salmon (Fisher 1992) captured in the watershed vary in their accuracy, particularly for spring-run (NMFS 2013; Harvey et al. 2014; Merz et al. 2014). The inaccuracy of the size-based methods is likely due to differences in fish distribution during early rearing, habitat-specific growth rates, and inter-annual variability in temperatures and food availability that lead to overlap in size ranges among stocks. The primary objective of this project was the genetic classification (to genetic lineage; Evolutionary Significant Unit) of Chinook Salmon captured from State Water Project and Central Valley Project fish protection facilities and Interagency Ecological Program compliance monitoring programs. The genetic lineage was determined for sampled fish by comparing their genotypes to reference genetic baselines. Genetic methods, having less statistical uncertainty that size-based models for population identification, were intended to directly target (and reduce) one source of uncertainty in the estimation of loss (take) from water diversions (operations) and develop the information necessary for understanding stock-specific distribution, habitat utilization, abundance, and life history variation. This project supports recommendations from the Interagency Ecological Program’s Salmon and Sturgeon Assessment of Indicators by Life Stage and Interagency Ecological Program Science Agenda efforts to improve Central Valley
Nursing Assignment Help | Best Writing Service By Experts
<p>Students enrolled in nursing degree and diploma programs at Australian universities find writing nursing assignments challenging and seek <a href="https://www.globalassignmenthelp.com.au/nursing-assignment-help">Nursing assignment writing services</a> in Australia.</p>
IV-KAPhE kinase-substrate assignments for the entire human phosphoproteome
<p>This data set includes the full, all-vs-all kinase-substrate assignments by the IV-KAPhE method for the entire human phosphoproteome (union of the PhosphoSitePlus human phosphosite database and the Ochoa et al. 2020 high-confidence human phosphoproteome). This is an unfiltered version of Supplemental Table S1 from Invergo BM (2022) "Accurate, high-coverage assignment of in vivo protein kinases to phosphosites from in vitro phosphoproteomic specificity data".</p> <p>The data set also includes files to facilitate scoring new human phosphosites, particularly the in vitro half of the IV-KAPhE model. "naive-bayes-plus-model.tar.gz" is an archive of HDF5 files comprising the "Naive Bayes+" multi-label, in vitro kinase-substrate assignment model used in the IV-KAPhE model, as described in the manuscript. These files are to be used with the motif-kit software package and can be used to score new sites. "kinase-int-domains-sig.tsv" and "kinase-sub-domains-sig.tsv" contain Pfam domains enriched among each kinase's interacting partners or substrates, respectively. Finally, "human-kinase-interactions.tsv" and "human-kinase-2nd-interactions.tsv" contain physical interactions and indirect ("2 hop") interactions between human protein kinases and other proteins, as described in the manuscript.</p>
PRJNA860062 Assigned Taxonomy and QIIME2 Pipeline
<p><strong>PRJNA860062 Assigned Taxonomy:</strong></p> <p>This upload comprises two datasets with the assigned taxonomy for sequence variants of BioProject PRNJA860062.</p> <ul> <li>PRJNA860062_ASVCounts_NCBItaxonomy.txt</li> <li>PRJNA860062_ASVCounts_SILVAtaxonomy.txt</li> </ul> <p>BioProject PRNJN860062 compares bacterial profiles of zebrafish larvae microbiota resulting from two different microbial colonization methods. The full description and sequence data for this project can be obtained from the Sequence Read Archive (<a href="https://www.ncbi.nlm.nih.gov/bioproject">https://www.ncbi.nlm.nih.gov/bioproject</a>). </p> <p>The dataset with the SILVA taxonomy can directly be obtained using the QIIME2 script included in this upload ('PRJNA860062_QIIME2Script.txt'). As previously noted by Lesack and Birol (2018), SILVA species annotations include nomenclature errors (<a href="https://doi.org/10.1101/441576">DOI: 10.1101/441576</a>). Therefore, the dataset with the NCBI taxonomy comprises a manually corrected taxonomy for BioProject PRNJA860062, based on the family to phylum level nomenclature of the NCBI taxonomy browser (<a href="https://www.ncbi.nlm.nih.gov/taxonomy">https://www.ncbi.nlm.nih.gov/taxonomy</a>).</p> <p>Both files are tab-delimited text files, include the domain to species level taxonomy in the first 7 columns, and include the number of assigned sequence variants (ASVs) per taxon in the final 6 colums, corresponding to BioSample SAMN29820940, SAMN29820941, SAMN29820942, SAMN29820943, SAMN29820944, and SAMN29820945.</p> <p> </p> <p><strong>QIIME 2 Pipeline:</strong></p> <p>The QIIME2 script that was used to obtain the assigned SILVA taxonomy BioProject PRNJA860062 is uploaded as:</p> <ul> <li>PRJNA860062_QIIME2Script.txt</li> </ul> <p>Input files that are required to run this script, including a manifest text file, sample metadata, and the reference sequences and taxonomy from the SILVA 138 small subunit (16S/18S) rRNA database Ref NR 99, are uploaded in the zipped file:</p> <ul> <li>PRJNA860062_InputFiles.zip</li> </ul> <p>FASTQ sequence data for BioSample SAMN29820940, SAMN29820941, SAMN29820942, SAMN29820943, SAMN29820944, and SAMN29820945, can be obtained from the Sequence Read Archive under BioProject PRNJA860062 (<a href="https://www.ncbi.nlm.nih.gov/bioproject">https://www.ncbi.nlm.nih.gov/bioproject</a>).</p> <p>All output files are uploaded in the zipped file:</p> <ul> <li>PRJNA860062_OutputFiles.zip</li> </ul> <p>Data provenance, including the versions of python (3.6.7) and python packages, can be acquired by dragging QIIME2 Visualizations (.qzv output files) into the QIIME2 viewing interface (<a href="http://view.qiime2.org">http://view.qiime2.org</a>).</p>
The Encyclopedia of Domains (TED) structural domains assignments for AlphaFold Database v4
<h3>Dataset description:</h3> <p>The Encyclopedia of Domains (TED) is a joint effort by CATH (Orengo group) and the Jones group at University College London to identify and classify protein domains in AlphaFold2 models from AlphaFold Database version 4, covering over 188 million unique sequences and 365 million domain assignments. </p> <p>In this data release, we will be making available to the community a table of domain boundaries and additional metadata on quality (pLDDT, globularity, number of secondary structures), taxonomy, and putative CATH SuperFamily or Fold assignments, for all 365 million domains (~324 million domains in TED100 and ~40 million domains in TED-redundant).</p> <p>For all chains in the chain-level TED-redundant files, the file contains boundary predictions, consensus level and information on the TED100 representative.</p> <p>For both TED100 and TED-redundant we provide domain boundary predictions outputted by each of the three methods employed in the project (Chainsaw, Merizo, UniDoc). </p> <p>We are making available 7,427 PDB files for potentially novel folds identified during the TED classification process, with an annotation table sorted by novelty, as well as 6,433 highly symmetrical folds representatives.</p> <p>Please use the gunzip command to extract files with a '.gz' extension and "tar -xzvf file.tar.gz" to open .tar.gz files .</p> <p>CATH annotations have been assigned using the Foldseek algorithm applied in various modes, and the Foldclass algorithm, both of which are used to report significant structural similarity to a known CATH domain. </p> <p><br><strong>Note: The TED protocol differs from that of the standard CATH Assignment protocol for superfamily assignment, which also involves HMM-based protocols and manual curation for classification into superfamilies.</strong></p> <h3> </h3> <h3><strong>Changelog Version 5:</strong></h3> <ul> <li><strong>Add</strong>: ted_365m.domain_summary.cath.globularity.taxid.tsv.tar.gz - This table, in the same format as the previous ted_100_324m.domain_summary.cath.globularity.taxid.tsv.tar.gz, contains per-domain annotations for the whole of TED, including metadata on domain quality metrics such as secondary structure elements counts, globularity scores, average pLDDT and taxonomical assignments.</li> <li><strong>Add:</strong> high_symmetry_folds_set.domain_summary.tsv.gz - subset of ted_365m.domain_summary.cath.globularity.taxid.tsv containing information on 6,433 high symmetry folds in TED. The entries are sorted in descending order by Z-score obtained from SymD.</li> <li><strong>Add:</strong> high_symmetry_folds_set_models.tar.gz - TED domain models in PDB format for 6,433 high symmetry folds in TED.</li> <li><strong>Add:</strong> ISP_data.tar.gz - Raw data for Interacting SuperFamily Pairs calculations used in the manuscript. A more detailed description of the ISP data is available below as well as within the tar.gz file. </li> <li><strong>Add:</strong> ted_redundant_40m_domain_id.list.gz - list of TED_domain_ID in TED redundant</li> <li><strong>Add:</strong> ted_100_324m_domain_id.list.gz - list of TED_domain_ID in TED100</li> <li><strong>Fix/Replace</strong>: A domain-level summary of TED, now consolidated into <strong>ted_365m.domain_summary.cath.globularity.taxid.tsv</strong>, is consistent with the protocol used in the manuscript. As Foldclass and Foldseek T-level hits provide all 4 CATH digits, we removed the H portion of the CATH code from each prediction at the T-level. <br>Previously, the following columns <br> 14. cath_label - CATH superfamily code if predicted, either a C.A.T.H. homologous superfamily or C.A.T. fold assignment. i.e. 3.40.50.300<br> 15. cath_assignment_level - H for homologous superfamily assignment, T for fold level assignment.<br> 16. cath_assignment_method - Method used to assign a CATH label, either Foldseek or Foldclass<br>sometimes showed an additional label with a T-level prediction by Foldclass in the case of T-level assignments obtained by Foldseek, e.g.<br>3.40.30,3.40.30 T foldseek,foldclass<br>This has now been corrected to reflect the TED protocol, with Foldclass T-level assignments applied only to domains where a T-level assignment could not be applied using Foldseek, e.g. <br>domain-x 3.40.30 T foldseek<br>domain-y 3.20.20 T foldclass<br><br>Thus, in the current version of the data, CATH assignments label can only be <br>H-level assignment by Foldseek (i.e. 3.40.50.300 H foldseek)<br>T-level assignment by Foldseek (i.e. 3.40.30 T foldseek)<br>T-level assignment by Foldclass (i.e. 3.40.30 T foldclass)<br>or no assignment (- - - )</li> </ul> <h3><br>This dataset contains:</h3> <ul> <li><strong>ted_214m_per_chain_segmentation.tsv</strong><br>The file contains all 214M protein chains in TED with consensus domain boundaries and proteome information in the following columns.<br>1. AFDB_model_ID: chain identifier from AFDB in the format AF-<UniProtID>-F1-model_v4 i.e. AF-A0A1V6M2Y0-F1-model_v4<br>2. md5 hash for chain sequence<br>3. nres - number of residues in chain<br>4. n_high - number of high consensus domains predicted in chain<br>5. n_med - number of medium consensus domains predicted in chain<br>6. n_low - number of low consensus domains predicted in chain<br>7. high_consesnsus - boundaries of high consensus domains predicted in chain. If none, 'na'<br>8. med_consensus - boundaries of medium consensus domains predicted in chain. If none, 'na'<br>9. low_consensus - boundaries of low consensus domains predicted in chain. If none, 'na'<br>10. proteome_id - proteome identifier in the format proteome-tax_id-<taxonID>-<shard>_v4 i.e. proteome-tax_id-67581-0_v4<br><br></li> <li><strong>ted_365m_domain_boundaries_consensus_level.tsv.gz</strong><br>The file contains all domain assignments in TED100 and TED-redundant (365M) in the format:<br>1. TED_ID: TED domain identifier in the format AF-<UniProtID>-F1-model_v4_TED<domain_number_in_chain> i.e. AF-A0A1V6M2Y0-F1-model_v4_TED03<br>2. Boundaries: domain boundaries in the format <start>-<stop> or <start>-<stop>_<start>-<stop> for discontinuous domains.<br>3. Consensus: either high or medium.<br><br></li> <li><strong>ted_100_324m_domain_id.list.gz</strong> - list of ~324 million domain identifiers in TED100, one per line in the format AF-<UniProtID>-F1-model_v4_TED<domain_number_in_chain> i.e. AF-A0A1V6M2Y0-F1-model_v4_TED0<br><br></li> <li><strong>ted_redundant_40m_domain_id.list.gz</strong> - list of ~40 million domain identifiers in TED redundant, one per line in the format AF-<UniProtID>-F1-model_v4_TED<domain_number_in_chain> i.e. AF-A0A1V6M2Y0-F1-model_v4_TED0<br><br></li> <li><strong>ted_365m.domain_summary.cath.globularity.taxid.tsv, novel_folds_set.domain_summary.tsv</strong> and <strong>high_symmetry_folds_set.domain_summary.tsv</strong> are header-less with the following columns separated by tabs (.tsv). novel_folds_set.domain_summary.tsv is sorted by novelty<br><strong>Note: The TED protocol differs from that of the standard CATH Assignment protocol for superfamily assignment, which also involves HMM-based protocols and manual curation for classification into superfamilies.</strong><br><br> 1. ted_id - TED domain identifier in the format AF-<UniProtID>-F1-model_v4_TED<domain_number_in_chain> i.e. AF-A0A1V6M2Y0-F1-model_v4_TED03<br> 2. md5_domain - md5 hash of domain sequence<br> 3. consensus_level - medium (2 methods agreement) or high (3 methods agreement)<br> 4. chopping - domain boundaries in the format <start>-<stop> or <start>-<stop>_<start>-<stop> for discontinuous domains<br> 5. nres_domain - number of residues in domain<br> 6. num_segments - number of individual segments in domain. <br> 7. plddt - average pLDDT for domain (range from 0 to 100)<br> 8. num_helix_strand_turn - number of helix strand turns predicted by STRIDE<br> 9. num_helix - number of helices predicted by STRIDE<br> 10. num_strand - number of strands predicted by STRIDE<br> 11. num_helix_strand - number of helices and strands predicted by STRIDE<br> 12. num_turn - number of turns predicted by STRIDE<br> 13. proteome_id - proteome identifier in the format proteome-tax_id-<taxonID>-<shard>_v4 i.e. proteome-tax_id-67581-0_v4<br> 14. cath_label - CATH superfamily code if predicted, either a C.A.T.H. homologous superfamily or C.A.T. fold assignment. i.e. 3.40.50.300. Otherwise '-'<br> 15. cath_assignment_level - H for homologous superfamily assignment, T for fold level assignment. Otherwise '-'<br> 16. cath_assignment_method - Method used to assign a CATH label, either foldseek or foldclass. Otherwise '-'<br> 17. packing_density - metric used to determine globularity. A domain with packing_density >=10.333 and norm_rg below 0.356 is considered globular<br> 18. norm_rg - normalised radius of gyration. A domain with packing_density >=10.333 AND norm_rg below 0.356 is considered globular. <br> 19. tax_common_name - Common name for organism<br> 20. tax_scientific_name - Scientific name for organism<br> 21. tax_lineage - Full taxonomic lineage.<br><br></li> <li><strong>ted_324m_seq_clustering.cathlabels.tsv.gz</strong> <br>The file contains the results of the domain sequences clustering with MMseqs2. <br>Columns:<br>1. Cluster_representative<br>2. Cluster_member<br>3. CATH code assignment if available i.e. 3.40.50.300 for a domain with a homologous match or 3.20.20 for a domain matching at the fold level in the CATH classification<br>4. CATH assignment type - either Foldseek-T, Foldseek-H or Foldclass<br><br></li> <li><strong>Domain assignments for TED redundant using single-chain and multi-chain consensus in</strong> <strong>ted_redundant_39m.multichain.consensus_domain_summary.taxid.tsv.gz and ted_redundant_39m.singlechain.consensus_domain_summary.taxid.tsv.gz</strong> </li> <li> The file <strong>ted_redundant_39m.multichain.consensus_domain_summary.taxid.tsv.gz</strong> contains a header with the following fields. Each column is tab-separated (.tsv).<br> 1. TED_redundant_id - TED chain identifier in the format AF-<UniProtID>-F1-model_v4 i.e. AF-A0A1V6M2Y0-F1-model_v4<br> 2. md5 - md5 hash for chain sequence<br> 3. nres - number of residues in chain<br> 4. n_high - number of high consensus domains predicted in chain<br> 5. n_med - number of medium consensus domains predicted in chain<br> 6. high_consensus - boundaries of high consensus domains predicted in chain<br> 7. med_consensus - boundaries of medium consensus domains predicted in chain<br> 8. ndom_consensus - number of consensus domains predicted in chain<br> 9. n_targets - number of chains considered for consensus calculation<br> 10. proteome_id - proteome identifier in the format proteome-tax_id-<taxonID>-<shard>_v4 i.e. proteome-tax_id-67581-0_v4<br> 11. TED_redundant_species - Scientific name for organism the chain originally comes from.<br> 12. TED100_chain_rep - TED100 representative for chain <br> 13. TED100_chain_rep_species - Species of TED100 representative for chain.</li> <li>The file <strong>ted_redundant_39m.singlechain.consensus_domain_summary.taxid.tsv</strong> contains a header with the following fields. Each column is tab-separated (.tsv).<br> 1. TED_redundant_id - TED chain identifier in the format AF-<UniProtID>-F1-model_v4 i.e. AF-A0A1V6M2Y0-F1-model_v4<br> 2. md5 - md5 hash for chain sequence<br> 3. nres - number of residues in chain<br> 4. n_high - number of high consensus domains predicted in chain<br> 5. n_med - number of medium consensus domains predicted in chain<br> 6. high_consensus - boundaries of high consensus domains predicted in chain<br> 7. med_consensus - boundaries of medium consensus domains predicted in chain<br> 8. proteome_id - proteome identifier in the format proteome-tax_id-<taxonID>-<shard>_v4 i.e. proteome-tax_id-67581-0_v4 <br> 9. TED_redundant_species - Scientific name for organism the chain originally comes from<br> 10. TED100_chain_rep - TED100 representative for chain <br> 11. TED100_chain_rep_species - Species of TED100 representative for chain.</li> </ul> <p> </p> <ul> <li><strong>novel_folds_set_models.tar.gz</strong> contains PDB files of all novel folds representatives identified in TED100.<br><br></li> <li><strong>high_symmetry_folds_set_models.tar.gz</strong> contains PDB files of all highly symmetrical folds representatives identified in TED100.<br> </li> <li><strong>Per-tool domain boundaries_predictions</strong> - All per-tool domain boundaries predictions for TED100 and TED-redundant are in the same format with the following columns.<br> 1. TED_chainID - TED chain identifier in the format AF-<UniProtID>-F1-model_v4 i.e. AF-A0A1V6M2Y0-F1-model_v4<br> 2. TED_chain_md5 - md5 hash for chain sequence<br> 3. TED_chain_length - number of residues in chain<br> 4. ndoms - number of domains predicted in chains<br> 5. Domain boundaries - domain boundaries in the format <start>-<stop> or <start>-<stop>_<start>-<stop> for discontinuous domains<br> 6. Prediction probability - probability of each per-chain prediction<br>Domain boundaries predictions share the same format, with each segment separated by '_' and segment boundaries (start,stop) separated by '-'<br> <br> i.e.domain prediction by Merizo for AF-A0A000-F1-model_v4<br> AF-A0A000-F1-model_v4 e8872c7a0261b9e88e6ff47eb34e4162 394 2 10-52_289-394,53-288 0.90077<br> <br> Merizo predicts one continuous domain and a discontinuous domain,<br> Domain1 (discontinuous): 10-52_289-394<br> segment1: 10-52<br> segment2: 289-394<br> Domain 2 (continuous):<br> segment 1: 53-288<br><br></li> <li><strong>ISP_data.tar.gz</strong> contains raw data for the Interacting Superfamily Pairs (ISP) calculations featured in the manuscript. The archive contains a README as well as :<br>all_ISP_data_cath.pkl: <br>ISP data for CATH 4.3 in Python pickle format. <br>A Python dictionary with the following contents: <br>Each key is an ISP, e.g. '3.40.640.10-3.90.1150.10' <br>Each value is a dictionary with the following contents:<br>key 'aligned_domain_pairs': value is a Python list of length N, each element is a string specifying the two TED domain IDs in contact, separated by a colon, e.g. "AF-A0A000-F1-model_v4_TED02:AF-A0A000-F1-model_v4_TED01" <br>key 'vectors': value is a numpy.ndarray of shape (N, 3). Each row is a raw unnormalized interaction vector for the corresponding domain pair, after aligning to the reference structure. <br>Any given index in each list or ndarray has the data for a single domain pair; the order is constant in each list/array. <br>---------------------------------- <br>all_ISP_data_afdb.pkl: <br>ISP data for TED100 in Python pickle format. <br>A Python dictionary with the following contents: <br>Each key is an ISP, e.g. '3.40.640.10-3.90.1150.10' <br>Each value is a dictionary with the following contents:<br>key 'aligned_domain_pairs': value is a Python list of length N, each element is a string specifying the two TED domain IDs in contact, separated by a colon, e.g. "AF-A0A000-F1-model_v4_TED02:AF-A0A000-F1-model_v4_TED01"<br>key 'vectors': value is a numpy.ndarray of shape (N, 3). Each row is a raw unnormalized interaction vector for the corresponding domain pair, after aligning to the reference structure. <br>key 'choppings': value is a Python list of length N, each element is a colon-separated string containing the TED chopping strings for the domains in contact, e.g. "54-288:11-41_290-389". The format for each 'chopping' follows that used in the main TED TSV files. <br>key 'pae_score': value is a numpy.ndarray of floats of shape (N,). Each element is the median PAE score between the domains in contact, computed across both relevant parts of the PAE matrix as described in the paper. <br>Any given index in each list or ndarray has the data for a single domain pair; the order is constant in each list/array. <br>NB: each list of domains has not been filtered by PAE score, so the pae_score values include values greater than 4.0, which was the threshold used to filter confident predictions in the manuscript.<br>------------------------------------- <br>isp_data_afdbonly_nopaefilter.csv:<br>A subset of the data in the TED100 .pkl file above, in CSV format. <br>Each row contains the following fields: AFDB ID, e.g. AF-A0A000-F1-model_v4 ISP, e.g. 3.40.640.10-3.90.1150.10<br>Colon-separated domain ID pair, e.g. AF-A0A000-F1-model_v4_TED02:AF-A0A000-F1-model_v4_TED01 PAE score, e.g. '4.0' <br>As the aforementioned files, this data has not been filtered by PAE score values.<br><br></li> <li><strong>ted-tools-main.zip</strong> - copy of the https://github.com/psipred/ted-tools repository, containing tools and software used to generate TED.<br><br></li> <li><strong>cath-alphaflow-main.zip</strong> - copy of CATH-AlphaFlow, used to generate globularity scores for TED domains.<br><br></li> <li><strong>ted-web-master.zip</strong> - copy of TED-web, containing code to generate the web interface of TED (https://ted.cathdb.info)<br><br></li> <li><strong>gofocus_data.tar.bz2</strong> - GOFocus model weights</li> </ul>
A Nu Supersymmetric Anomaly-free Atlas: anomaly-free, flavour-dependent U(1) charge assignments for the Minimally Supersymmetric Standard Model plus three Standard Model-singlet superfields
<p>We present lists of anomaly-free charge assignments up to a maximum magnitude charge Qmax=10 for the chiral fermionic content of the MSSM plus 3 right-handed neutrinos. </p> <p>Due to the large number of solutions, we compress the list into the file MSSMnuRcharges_Qmax10.gz. Please note that the unzipped file is approximately 130GB in size. We additionally include a smaller file, MSSMnuRcharges_Qmax4, containing the subset of anomaly-free charge assignments up to a maximum magnitude charge Qmax=4.</p> <p>The files searchU1MSSM.cpp and searchU1MSSM.h contain C++ files (in the 2014 standard) to produce the solutions. runsearch.sh is a bash script that compiles the programs and then runs it for a sample set of inputs.</p> <p>We provide Mathematica notebooks Analytic_solution_generator.nb and Analytic_Checks.nb which respectively provide the parametrisation of the analytic solution and checks thereof.</p> <p>The files beginning 'filter' contain example programs that read in each line in the solution list, apply a filter and print only the solutions satisfying the conditions of that filter. runfilter.sh is a bash script that compiles the filters and then runs a single filter as an example.</p> <p>These data and programs are based on this paper: https://arxiv.org/abs/2107.07926.</p>
Fig. 18 in New generic assignment to the harvestman Metaphareus punctatus (Opiliones: Stygnidae) and observations about it reproductive behavior
Fig. 18. Geographical distribution of Eutimesius punctatus (Roewer, 1913), comb. nov. and E. albicinctus (Roewer, 1915).
Fig 4 in A mountain of millipedes III: A new genus for three new species from the Udzungwa Mountains and surroundings, Tanzania, as well as several 'orphaned' species previously assigned to Odontopyge Brandt, 1841 (Diplopoda, Spirostreptida, Odontopygidae)
Fig 4. Geotypodon millemanus gen. et sp. nov., paratype from Kiranzi-Kitungulu FR. A–E: Right gonopod. A. Posterior view. B. Anterior view, telomere in red oval. C. (Posterior-)mesal view, solenomere (yellow) and proximal telomeral spine (green) coloured. D. (Anterior-)mesal view. E. Telomere (part of coxa at lower left), basal (dorsal) view. F. Limbus. Abbreviations: atl = anterior distal lobe of telomere; bl = basal lamella of telomere; itl = intermediate distal lamella of telomere; ll = longitudinal lamella; mf = anteriad metaplical flange; mla = metaplical lamella; mp = metaplica; msp = metaplical spine-like process; pn = posttorsal narrowing; pp = proplica; ptl = posterior distal lobe of telomere; pts = proximal telomeral spine; slm = solenomere; tt = torsotope. Scales: A–E = 0.1 mm, F = 0.01 mm.
Fig. 1 in A mountain of millipedes III: A new genus for three new species from the Udzungwa Mountains and surroundings, Tanzania, as well as several 'orphaned' species previously assigned to Odontopyge Brandt, 1841 (Diplopoda, Spirostreptida, Odontopygidae)
Fig. 1. Map of the Udzungwa Mountains, showing the collecting sites for the three new Geotypodon species, as well as names of the Forest Reserves in question and names of individual mountains in West Kilombero FR. Red diamonds = G. millemanus gen. et sp. nov., yellow dot: G. submontanus gen. et sp. nov., blue triangle: G. iringensis gen. et sp. nov. Based on fig. 1 in Marshall et al. (2010) and information in Doody et al. (2001).
Fig. 2 in A mountain of millipedes III: A new genus for three new species from the Udzungwa Mountains and surroundings, Tanzania, as well as several 'orphaned' species previously assigned to Odontopyge Brandt, 1841 (Diplopoda, Spirostreptida, Odontopygidae)
Fig. 2. Geotypodon millemanus gen. et sp. nov., paratype from West Kilombero Scarp FR after nine years in alcohol. Photograph by N. Ioannou.
Fig. 3 in A mountain of millipedes III: A new genus for three new species from the Udzungwa Mountains and surroundings, Tanzania, as well as several 'orphaned' species previously assigned to Odontopyge Brandt, 1841 (Diplopoda, Spirostreptida, Odontopygidae)
Fig. 3. Body size of males of Geotypodon spp. Bold symbols indicate numbers of podous rings and midbody vertical diameter of the new species described here. Small circles and shaded areas indicate published measurements for other Geotypodon species.
Fig. 5 in A mountain of millipedes III: A new genus for three new species from the Udzungwa Mountains and surroundings, Tanzania, as well as several 'orphaned' species previously assigned to Odontopyge Brandt, 1841 (Diplopoda, Spirostreptida, Odontopygidae)
Fig. 5. Geotypodon submontanus gen. et sp. nov., holotype. A–E. Left gonopod. A. Anterior view. B. Posterior view. C. Apical part of coxa and proximal part of telopodite, anterior view. D. (Anterior-) mesal view. E. Posterior distal lobe of telomere; insertion highlights spine-like process. F. Limbus. Abbreviations: atl = anterior distal lobe of telomere; itl = intermediate distal lamella of telomere; ll = longitudinal lamella; mla = metaplical lamella; msp = metaplical spine-like process; ptl = posterior distal
Fig. 6 in A mountain of millipedes III: A new genus for three new species from the Udzungwa Mountains and surroundings, Tanzania, as well as several 'orphaned' species previously assigned to Odontopyge Brandt, 1841 (Diplopoda, Spirostreptida, Odontopygidae)
Fig. 6. Geotypodon iringensis gen. et sp. nov. A–D. Holotype, left gonopod. A. Anterior view. B. Posterior view. C. Mesal-ventral view. D. Telomere and solenomere, basal (dorsal) view. — E. Paratype, limbus. Abbreviations: atl = anterior distal lobe of telomere, bl = basal lamella of telomere, cx = coxa (seen from the basis, with remains of muscles); itl = intermediate distal lamella of telomere; lc = lateral concavity of coxa; lfl = longitudinally folded lamella; mf = anteriad metaplical flange; mp = metaplica; msp = metaplical spine-like process; pn = posttorsal narrowing; pp = proplica; ptl = posterior distal lobe of telomere; pts = proximal telomeral spine; slm = solenomere; tl = terminal lobe of telomere; tt = torsotope. Scales: A–D = 0.2 mm, E = 0.01 mm.
The Experimental Data for the Study "Frequency Fitness Assignment: Making Optimization Algorithms Invariant under Bijective Transformations of the Objective Function Value"
<p>The Experimental Data for the Study "Frequency Fitness Assignment: Making Optimization Algorithms Invariant under Bijective Transformations of the Objective Function Value"</p> <p><strong>1. Introduction</strong></p> <p>Frequency Fitness Assignment (FFA) replaces the objective value in the selection step of an optimization method with its encounter frequency in any selection step so far. It turns static problems into dynamic ones. Here we experimentally investigated this approach in two important contexts: First, we integrated it into a basic (1+1)-EA, obtaining the (1+1)-FEA. We applied both algorithms to several well-known benchmark problems with bit-string based search spaces, including the OneMax, LeadingOnes, TwoMax, Jump, Plateau, and W-Model functions. We also applied them to the Max-3-Sat instances from SATLib. We then also integrated FFA into a Memetic Algorithm for the Job Shop Problem.</p> <p><strong>2. Paper</strong></p> <p>This data is used as the basis for the following article: Thomas Weise, Zhize Wu, Xinlu Li, and Yan Chen. Frequency Fitness Assignment: Making Optimization Algorithms Invariant under Bijective Transformations of the Objective Function Value, originally submitted to <a href="https://arxiv.org/abs/2001.01416">arxiv</a> on 2020-01-06 (under the title Frequency Fitness Assignment: Making Optimization Algorithms Invariant under Bijective Transformations of the Objective Function), updated with the new data in June 2020, and submitted to the IEEE Transactions on Evolutionary Computation.</p> <p><strong>3. Data</strong></p> <p>This data set contains all the results of these experiments, the source codes used in the experiments (i.e., the algorithm implementations), as well as the scripts used for evaluating the results.</p> <p><strong>4. Version History</strong></p> <p>This is the second version of the data set, including extended experiments and more evaluation results. Most importantly, data for larger scales of OneMax and LeadingOnes has been added. The original version is at <a href="http://dx.doi.org/10.5281/zenodo.3598172">10.5281/zenodo.3598172</a>.</p> <p><strong>5. Contact</strong></p> <p>If you have any questions or suggestions, please contact <a href="http://iao.hfuu.edu.cn/team/director">Prof. Dr. Thomas Weise</a> of the <a href="http://iao.hfuu.edu.cn/">Institute of Applied Optimization</a> at <a href="http://www.hfuu.edu.cn">Hefei University</a> in Hefei, Anhui, China via email to <a href="mailto:tweise@hfuu.edu.cn">tweise@hfuu.edu.cn</a> with CC to <a href="mailto:tweise@ustc.edu.cn">tweise@ustc.edu.cn</a>.</p>
Figure 6 in New genera and species of Urothoidae (Amphipoda) from the Brazilian deep sea, with the re-assignment of Pseudurothoe and Urothopsis to Phoxocephalopsidae
Figure 6. Carangolioides hamatus gen. et sp. nov., holotype, female, 5.4 mm, 22°04′32″ S, 39°54′11″ W, 750 m depth, MNRJ21434. Scale bars: 0.5 mm.
Figure 9 in New genera and species of Urothoidae (Amphipoda) from the Brazilian deep sea, with the re-assignment of Pseudurothoe and Urothopsis to Phoxocephalopsidae
Figure 9. Coronaurothoe rotunda gen. et sp. nov., sex unknown, 3.0 mm, 21°52′59″ S, 39°55′32″ W, 750 m depth, MNRJ 21443. Scale bars: 0.5 mm.
Figure 8 in New genera and species of Urothoidae (Amphipoda) from the Brazilian deep sea, with the re-assignment of Pseudurothoe and Urothopsis to Phoxocephalopsidae
Figure 8. Coronaurothoe rotunda gen. et sp. nov., sex unknown, 3.0 mm, 21°52′59″ S, 39°55′32″ W, 750 m depth, MNRJ 21443. Scale bars: 0.1 mm for U1; 0.2 mm for U2°3; 0.5 mm for all others.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.