Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
38
datasets available to search
ShareScore release 0.9.0
Dataset results
38 results for “encyclopedia”
Mars Target Encyclopedia - Labeled LPSC abstracts for four Mars missions
<p>This data set contains annotated text versions of 1635 two-page abstracts published at the Lunar and Planetary Science Conference from 1998 to 2020 of relevance to four Mars missions. The annotations were generated using named entity recognition and relation extraction provided by the MTE processing pipeline (available at https://github.com/wkiri/MTE), followed by manual review. Annotated entities include Element, Mineral, Property, and Target. Annotated relations include <strong>Contains</strong>(Target, Element | Mineral) and <strong>HasProperty</strong>(Target, Property). The extracted information (without full texts) is also available as a database (stored in .csv files) at https://pds-geosciences.wustl.edu/missions/mte/mte.htm . The complete annotated texts are provided here as a resource for further research and experimentation on information extraction methods. For more information about the Mars Target Encyclopedia and these annotations, please see:</p> <ul> <li>"<a href="https://www.hou.usra.edu/meetings/lpsc2022/pdf/1231.pdf">Targets from the Spirit Mars Exploration Rover in the Mars Target Encyclopedia</a>", Kiri L. Wagstaff, Raymond Francis, Matthew Golombek, Steven Lu, Ellen Riloff, Leslie Tamppari, Yuan Zhuang, and Thomas Stein.<br> <em>53rd Lunar and Planetary Science Conference</em>, Abstract #1231, March 2022.</li> <li>"<a href="https://www.hou.usra.edu/meetings/lpsc2021/pdf/1278.pdf">The Mars Target Encyclopedia Now Includes Mars Pathfinder and Mars Phoenix Targets</a>", Kiri L. Wagstaff, Raymond Francis, Matthew Golombek, Steven Lu, Ellen Riloff, Leslie Tamppari, and Thomas C. Stein.<br> <em>52nd Lunar and Planetary Science Conference</em>, Abstract #1278, March 2021.</li> </ul> <p>The original PDF abstracts are available at: </p> <ul> <li>For years prior to 2000: https://www.lpi.usra.edu/meetings/LPSC${two-digit-year}/pdf/${id}.pdf</li> <li>For year 2000: https://www.lpi.usra.edu/meetings/LPSC${four-digit-year}/pdf/${id}.pdf</li> <li>For years 2001-2017 (note lower-case lpsc): https://www.lpi.usra.edu/meetings/lpsc${four-digit-year}/pdf/${id}.pdf</li> <li>For years 2018-2020: https://www.hou.usra.edu/meetings/lpsc${four-digit-year}/pdf/${id}.pdf</li> </ul> <p>where ${id} is a four-digit abstract number, starting with 1001 (if available).</p> <p>The text files provided in this archive were extracted from the PDF files using the Apache Tika PDF parsing tool. They are named as ${four-digit-year}_${id}.txt. The text is provided here so that the annotations can be viewed in context. The text content remains copyright of the original abstract authors.</p> <p>The annotations (entities and relations) are provided in the format used by the brat annotation tool. They are named as ${four-digit-year}_${id}.ann. To view the annotations in a web-based graphical form, install the brat tool (http://brat.nlplab.org/). These annotations were generated using brat v1.3. The annotation files are also human-readable and can be parsed in to be used directly in code. If the .ann file is empty, then there are no relevant annotations for the associated text file.</p> <p><strong>Contents</strong>:</p> <ul> <li>mpf.zip: 591 abstracts relating to the Mars Pathfinder mission (1998-2020)</li> <li>mer-a.zip: 397 abstracts relating to the MER-A (Spirit) rover mission (2004-2020)</li> <li>mer-b.zip: 256 abstracts relating to the MER-B (Opportunity) rover mission (2005-2020)</li> <li>phx.zip: 391 abstracts relating to the Mars Phoenix Lander mission (2009-2020)</li> </ul> <p>Each directory contains a .txt and .ann file for each abstract. The .ann file is in brat standoff format (http://brat.nlplab.org/standoff.html). Additional .conf files are provided to generate color highlighting and keyboard shortcuts. These are used by the brat tool.</p> <p>Note: the same abstract may appear in more than one mission directory, if it discusses targets from more than one mission. It will have a different .ann file for each such appearance. Within each directory, a "Target" annotation is understood to refer to a target of the relevant mission.</p> <p><strong>Attribution</strong>:</p> <p>If you use this data set in your own work, please cite it as follows:</p> <p>Kiri L. Wagstaff, Raymond Francis, Matthew Golombek, Leslie Tamppari, and Steven Lu. (2022). Mars Target Encyclopedia - Labeled LPSC abstracts for four Mars missions (1.0.0.0) [Data set]. Zenodo. DOI: 10.5281/zenodo.7066107</p>
Mars Target Encyclopedia - LPSC abstracts labeled data set
<p>This data set contains annotated text versions of 2-page abstracts published at the Lunar and Planetary Science Conference in 2015 and 2016.</p> <p>The original PDF abstracts are available at:</p> <ul> <li>https://www.hou.usra.edu/meetings/lpsc2015/programAbstracts/view/</li> <li>https://www.hou.usra.edu/meetings/lpsc2016/programAbstracts/view/</li> </ul> <p>The text files in this archive were extracted using the Apache Tika PDF parsing tool. The text is provided here so that the annotations can be viewed. The text content remains copyright of the original abstract authors.</p> <p>The annotations (entities and relations) are provided in the format used by the brat annotation tool. To view the annotations in a web-based graphical form, install the brat tool (http://brat.nlplab.org/). These annotations were generated using brat v1.3. The annotation files are also human-readable and can be parsed in to be used directly in code.</p> <p><strong>Contents</strong>:</p> <ul> <li>lpsc15/: 62 abstracts</li> <li>lpsc16/: 55 abstracts</li> </ul> <p>Each directory contains a .txt and .ann file for each abstract. The .ann file is in brat standoff format (http://brat.nlplab.org/standoff.html).</p> <p>Additional .conf files are provided to generate color highlighting and keyboard shortcuts. These are used by the brat tool.</p> <p><strong>Attribution</strong>:</p> <p>If you use this data set in your own work, please cite this DOI:</p> <p>10.5281/zenodo.1048419</p> <p>Please also cite this paper, which provides additional details about the data set.</p> <p>Kiri L. Wagstaff, Raymond Francis, Thamme Gowda, You Lu, Ellen Riloff, Karanjeet Singh, and Nina Lanza. "Mars Target Encyclopedia: Rock and Soil Composition Extracted from the Literature." <em>Proceedings of the Thirtieth Annual Conference on Innovative Applications of Artificial Intelligence</em>, 2018.</p>
The Encyclopedia of Domains (TED) structural domains assignments for AlphaFold Database v4
<h3>Dataset description:</h3> <p>The Encyclopedia of Domains (TED) is a joint effort by CATH (Orengo group) and the Jones group at University College London to identify and classify protein domains in AlphaFold2 models from AlphaFold Database version 4, covering over 188 million unique sequences and 365 million domain assignments. </p> <p>In this data release, we will be making available to the community a table of domain boundaries and additional metadata on quality (pLDDT, globularity, number of secondary structures), taxonomy, and putative CATH SuperFamily or Fold assignments, for all 365 million domains (~324 million domains in TED100 and ~40 million domains in TED-redundant).</p> <p>For all chains in the chain-level TED-redundant files, the file contains boundary predictions, consensus level and information on the TED100 representative.</p> <p>For both TED100 and TED-redundant we provide domain boundary predictions outputted by each of the three methods employed in the project (Chainsaw, Merizo, UniDoc). </p> <p>We are making available 7,427 PDB files for potentially novel folds identified during the TED classification process, with an annotation table sorted by novelty, as well as 6,433 highly symmetrical folds representatives.</p> <p>Please use the gunzip command to extract files with a '.gz' extension and "tar -xzvf file.tar.gz" to open .tar.gz files .</p> <p>CATH annotations have been assigned using the Foldseek algorithm applied in various modes, and the Foldclass algorithm, both of which are used to report significant structural similarity to a known CATH domain. </p> <p><br><strong>Note: The TED protocol differs from that of the standard CATH Assignment protocol for superfamily assignment, which also involves HMM-based protocols and manual curation for classification into superfamilies.</strong></p> <h3> </h3> <h3><strong>Changelog Version 5:</strong></h3> <ul> <li><strong>Add</strong>: ted_365m.domain_summary.cath.globularity.taxid.tsv.tar.gz - This table, in the same format as the previous ted_100_324m.domain_summary.cath.globularity.taxid.tsv.tar.gz, contains per-domain annotations for the whole of TED, including metadata on domain quality metrics such as secondary structure elements counts, globularity scores, average pLDDT and taxonomical assignments.</li> <li><strong>Add:</strong> high_symmetry_folds_set.domain_summary.tsv.gz - subset of ted_365m.domain_summary.cath.globularity.taxid.tsv containing information on 6,433 high symmetry folds in TED. The entries are sorted in descending order by Z-score obtained from SymD.</li> <li><strong>Add:</strong> high_symmetry_folds_set_models.tar.gz - TED domain models in PDB format for 6,433 high symmetry folds in TED.</li> <li><strong>Add:</strong> ISP_data.tar.gz - Raw data for Interacting SuperFamily Pairs calculations used in the manuscript. A more detailed description of the ISP data is available below as well as within the tar.gz file. </li> <li><strong>Add:</strong> ted_redundant_40m_domain_id.list.gz - list of TED_domain_ID in TED redundant</li> <li><strong>Add:</strong> ted_100_324m_domain_id.list.gz - list of TED_domain_ID in TED100</li> <li><strong>Fix/Replace</strong>: A domain-level summary of TED, now consolidated into <strong>ted_365m.domain_summary.cath.globularity.taxid.tsv</strong>, is consistent with the protocol used in the manuscript. As Foldclass and Foldseek T-level hits provide all 4 CATH digits, we removed the H portion of the CATH code from each prediction at the T-level. <br>Previously, the following columns <br> 14. cath_label - CATH superfamily code if predicted, either a C.A.T.H. homologous superfamily or C.A.T. fold assignment. i.e. 3.40.50.300<br> 15. cath_assignment_level - H for homologous superfamily assignment, T for fold level assignment.<br> 16. cath_assignment_method - Method used to assign a CATH label, either Foldseek or Foldclass<br>sometimes showed an additional label with a T-level prediction by Foldclass in the case of T-level assignments obtained by Foldseek, e.g.<br>3.40.30,3.40.30 T foldseek,foldclass<br>This has now been corrected to reflect the TED protocol, with Foldclass T-level assignments applied only to domains where a T-level assignment could not be applied using Foldseek, e.g. <br>domain-x 3.40.30 T foldseek<br>domain-y 3.20.20 T foldclass<br><br>Thus, in the current version of the data, CATH assignments label can only be <br>H-level assignment by Foldseek (i.e. 3.40.50.300 H foldseek)<br>T-level assignment by Foldseek (i.e. 3.40.30 T foldseek)<br>T-level assignment by Foldclass (i.e. 3.40.30 T foldclass)<br>or no assignment (- - - )</li> </ul> <h3><br>This dataset contains:</h3> <ul> <li><strong>ted_214m_per_chain_segmentation.tsv</strong><br>The file contains all 214M protein chains in TED with consensus domain boundaries and proteome information in the following columns.<br>1. AFDB_model_ID: chain identifier from AFDB in the format AF-<UniProtID>-F1-model_v4 i.e. AF-A0A1V6M2Y0-F1-model_v4<br>2. md5 hash for chain sequence<br>3. nres - number of residues in chain<br>4. n_high - number of high consensus domains predicted in chain<br>5. n_med - number of medium consensus domains predicted in chain<br>6. n_low - number of low consensus domains predicted in chain<br>7. high_consesnsus - boundaries of high consensus domains predicted in chain. If none, 'na'<br>8. med_consensus - boundaries of medium consensus domains predicted in chain. If none, 'na'<br>9. low_consensus - boundaries of low consensus domains predicted in chain. If none, 'na'<br>10. proteome_id - proteome identifier in the format proteome-tax_id-<taxonID>-<shard>_v4 i.e. proteome-tax_id-67581-0_v4<br><br></li> <li><strong>ted_365m_domain_boundaries_consensus_level.tsv.gz</strong><br>The file contains all domain assignments in TED100 and TED-redundant (365M) in the format:<br>1. TED_ID: TED domain identifier in the format AF-<UniProtID>-F1-model_v4_TED<domain_number_in_chain> i.e. AF-A0A1V6M2Y0-F1-model_v4_TED03<br>2. Boundaries: domain boundaries in the format <start>-<stop> or <start>-<stop>_<start>-<stop> for discontinuous domains.<br>3. Consensus: either high or medium.<br><br></li> <li><strong>ted_100_324m_domain_id.list.gz</strong> - list of ~324 million domain identifiers in TED100, one per line in the format AF-<UniProtID>-F1-model_v4_TED<domain_number_in_chain> i.e. AF-A0A1V6M2Y0-F1-model_v4_TED0<br><br></li> <li><strong>ted_redundant_40m_domain_id.list.gz</strong> - list of ~40 million domain identifiers in TED redundant, one per line in the format AF-<UniProtID>-F1-model_v4_TED<domain_number_in_chain> i.e. AF-A0A1V6M2Y0-F1-model_v4_TED0<br><br></li> <li><strong>ted_365m.domain_summary.cath.globularity.taxid.tsv, novel_folds_set.domain_summary.tsv</strong> and <strong>high_symmetry_folds_set.domain_summary.tsv</strong> are header-less with the following columns separated by tabs (.tsv). novel_folds_set.domain_summary.tsv is sorted by novelty<br><strong>Note: The TED protocol differs from that of the standard CATH Assignment protocol for superfamily assignment, which also involves HMM-based protocols and manual curation for classification into superfamilies.</strong><br><br> 1. ted_id - TED domain identifier in the format AF-<UniProtID>-F1-model_v4_TED<domain_number_in_chain> i.e. AF-A0A1V6M2Y0-F1-model_v4_TED03<br> 2. md5_domain - md5 hash of domain sequence<br> 3. consensus_level - medium (2 methods agreement) or high (3 methods agreement)<br> 4. chopping - domain boundaries in the format <start>-<stop> or <start>-<stop>_<start>-<stop> for discontinuous domains<br> 5. nres_domain - number of residues in domain<br> 6. num_segments - number of individual segments in domain. <br> 7. plddt - average pLDDT for domain (range from 0 to 100)<br> 8. num_helix_strand_turn - number of helix strand turns predicted by STRIDE<br> 9. num_helix - number of helices predicted by STRIDE<br> 10. num_strand - number of strands predicted by STRIDE<br> 11. num_helix_strand - number of helices and strands predicted by STRIDE<br> 12. num_turn - number of turns predicted by STRIDE<br> 13. proteome_id - proteome identifier in the format proteome-tax_id-<taxonID>-<shard>_v4 i.e. proteome-tax_id-67581-0_v4<br> 14. cath_label - CATH superfamily code if predicted, either a C.A.T.H. homologous superfamily or C.A.T. fold assignment. i.e. 3.40.50.300. Otherwise '-'<br> 15. cath_assignment_level - H for homologous superfamily assignment, T for fold level assignment. Otherwise '-'<br> 16. cath_assignment_method - Method used to assign a CATH label, either foldseek or foldclass. Otherwise '-'<br> 17. packing_density - metric used to determine globularity. A domain with packing_density >=10.333 and norm_rg below 0.356 is considered globular<br> 18. norm_rg - normalised radius of gyration. A domain with packing_density >=10.333 AND norm_rg below 0.356 is considered globular. <br> 19. tax_common_name - Common name for organism<br> 20. tax_scientific_name - Scientific name for organism<br> 21. tax_lineage - Full taxonomic lineage.<br><br></li> <li><strong>ted_324m_seq_clustering.cathlabels.tsv.gz</strong> <br>The file contains the results of the domain sequences clustering with MMseqs2. <br>Columns:<br>1. Cluster_representative<br>2. Cluster_member<br>3. CATH code assignment if available i.e. 3.40.50.300 for a domain with a homologous match or 3.20.20 for a domain matching at the fold level in the CATH classification<br>4. CATH assignment type - either Foldseek-T, Foldseek-H or Foldclass<br><br></li> <li><strong>Domain assignments for TED redundant using single-chain and multi-chain consensus in</strong> <strong>ted_redundant_39m.multichain.consensus_domain_summary.taxid.tsv.gz and ted_redundant_39m.singlechain.consensus_domain_summary.taxid.tsv.gz</strong> </li> <li> The file <strong>ted_redundant_39m.multichain.consensus_domain_summary.taxid.tsv.gz</strong> contains a header with the following fields. Each column is tab-separated (.tsv).<br> 1. TED_redundant_id - TED chain identifier in the format AF-<UniProtID>-F1-model_v4 i.e. AF-A0A1V6M2Y0-F1-model_v4<br> 2. md5 - md5 hash for chain sequence<br> 3. nres - number of residues in chain<br> 4. n_high - number of high consensus domains predicted in chain<br> 5. n_med - number of medium consensus domains predicted in chain<br> 6. high_consensus - boundaries of high consensus domains predicted in chain<br> 7. med_consensus - boundaries of medium consensus domains predicted in chain<br> 8. ndom_consensus - number of consensus domains predicted in chain<br> 9. n_targets - number of chains considered for consensus calculation<br> 10. proteome_id - proteome identifier in the format proteome-tax_id-<taxonID>-<shard>_v4 i.e. proteome-tax_id-67581-0_v4<br> 11. TED_redundant_species - Scientific name for organism the chain originally comes from.<br> 12. TED100_chain_rep - TED100 representative for chain <br> 13. TED100_chain_rep_species - Species of TED100 representative for chain.</li> <li>The file <strong>ted_redundant_39m.singlechain.consensus_domain_summary.taxid.tsv</strong> contains a header with the following fields. Each column is tab-separated (.tsv).<br> 1. TED_redundant_id - TED chain identifier in the format AF-<UniProtID>-F1-model_v4 i.e. AF-A0A1V6M2Y0-F1-model_v4<br> 2. md5 - md5 hash for chain sequence<br> 3. nres - number of residues in chain<br> 4. n_high - number of high consensus domains predicted in chain<br> 5. n_med - number of medium consensus domains predicted in chain<br> 6. high_consensus - boundaries of high consensus domains predicted in chain<br> 7. med_consensus - boundaries of medium consensus domains predicted in chain<br> 8. proteome_id - proteome identifier in the format proteome-tax_id-<taxonID>-<shard>_v4 i.e. proteome-tax_id-67581-0_v4 <br> 9. TED_redundant_species - Scientific name for organism the chain originally comes from<br> 10. TED100_chain_rep - TED100 representative for chain <br> 11. TED100_chain_rep_species - Species of TED100 representative for chain.</li> </ul> <p> </p> <ul> <li><strong>novel_folds_set_models.tar.gz</strong> contains PDB files of all novel folds representatives identified in TED100.<br><br></li> <li><strong>high_symmetry_folds_set_models.tar.gz</strong> contains PDB files of all highly symmetrical folds representatives identified in TED100.<br> </li> <li><strong>Per-tool domain boundaries_predictions</strong> - All per-tool domain boundaries predictions for TED100 and TED-redundant are in the same format with the following columns.<br> 1. TED_chainID - TED chain identifier in the format AF-<UniProtID>-F1-model_v4 i.e. AF-A0A1V6M2Y0-F1-model_v4<br> 2. TED_chain_md5 - md5 hash for chain sequence<br> 3. TED_chain_length - number of residues in chain<br> 4. ndoms - number of domains predicted in chains<br> 5. Domain boundaries - domain boundaries in the format <start>-<stop> or <start>-<stop>_<start>-<stop> for discontinuous domains<br> 6. Prediction probability - probability of each per-chain prediction<br>Domain boundaries predictions share the same format, with each segment separated by '_' and segment boundaries (start,stop) separated by '-'<br> <br> i.e.domain prediction by Merizo for AF-A0A000-F1-model_v4<br> AF-A0A000-F1-model_v4 e8872c7a0261b9e88e6ff47eb34e4162 394 2 10-52_289-394,53-288 0.90077<br> <br> Merizo predicts one continuous domain and a discontinuous domain,<br> Domain1 (discontinuous): 10-52_289-394<br> segment1: 10-52<br> segment2: 289-394<br> Domain 2 (continuous):<br> segment 1: 53-288<br><br></li> <li><strong>ISP_data.tar.gz</strong> contains raw data for the Interacting Superfamily Pairs (ISP) calculations featured in the manuscript. The archive contains a README as well as :<br>all_ISP_data_cath.pkl: <br>ISP data for CATH 4.3 in Python pickle format. <br>A Python dictionary with the following contents: <br>Each key is an ISP, e.g. '3.40.640.10-3.90.1150.10' <br>Each value is a dictionary with the following contents:<br>key 'aligned_domain_pairs': value is a Python list of length N, each element is a string specifying the two TED domain IDs in contact, separated by a colon, e.g. "AF-A0A000-F1-model_v4_TED02:AF-A0A000-F1-model_v4_TED01" <br>key 'vectors': value is a numpy.ndarray of shape (N, 3). Each row is a raw unnormalized interaction vector for the corresponding domain pair, after aligning to the reference structure. <br>Any given index in each list or ndarray has the data for a single domain pair; the order is constant in each list/array. <br>---------------------------------- <br>all_ISP_data_afdb.pkl: <br>ISP data for TED100 in Python pickle format. <br>A Python dictionary with the following contents: <br>Each key is an ISP, e.g. '3.40.640.10-3.90.1150.10' <br>Each value is a dictionary with the following contents:<br>key 'aligned_domain_pairs': value is a Python list of length N, each element is a string specifying the two TED domain IDs in contact, separated by a colon, e.g. "AF-A0A000-F1-model_v4_TED02:AF-A0A000-F1-model_v4_TED01"<br>key 'vectors': value is a numpy.ndarray of shape (N, 3). Each row is a raw unnormalized interaction vector for the corresponding domain pair, after aligning to the reference structure. <br>key 'choppings': value is a Python list of length N, each element is a colon-separated string containing the TED chopping strings for the domains in contact, e.g. "54-288:11-41_290-389". The format for each 'chopping' follows that used in the main TED TSV files. <br>key 'pae_score': value is a numpy.ndarray of floats of shape (N,). Each element is the median PAE score between the domains in contact, computed across both relevant parts of the PAE matrix as described in the paper. <br>Any given index in each list or ndarray has the data for a single domain pair; the order is constant in each list/array. <br>NB: each list of domains has not been filtered by PAE score, so the pae_score values include values greater than 4.0, which was the threshold used to filter confident predictions in the manuscript.<br>------------------------------------- <br>isp_data_afdbonly_nopaefilter.csv:<br>A subset of the data in the TED100 .pkl file above, in CSV format. <br>Each row contains the following fields: AFDB ID, e.g. AF-A0A000-F1-model_v4 ISP, e.g. 3.40.640.10-3.90.1150.10<br>Colon-separated domain ID pair, e.g. AF-A0A000-F1-model_v4_TED02:AF-A0A000-F1-model_v4_TED01 PAE score, e.g. '4.0' <br>As the aforementioned files, this data has not been filtered by PAE score values.<br><br></li> <li><strong>ted-tools-main.zip</strong> - copy of the https://github.com/psipred/ted-tools repository, containing tools and software used to generate TED.<br><br></li> <li><strong>cath-alphaflow-main.zip</strong> - copy of CATH-AlphaFlow, used to generate globularity scores for TED domains.<br><br></li> <li><strong>ted-web-master.zip</strong> - copy of TED-web, containing code to generate the web interface of TED (https://ted.cathdb.info)<br><br></li> <li><strong>gofocus_data.tar.bz2</strong> - GOFocus model weights</li> </ul>
Encyclopedia of Life v2: Taxon Hierarchies and Associated Taxon Concepts
<p>This archive contains a snapshot of the taxon hierarchies, and associated scientific name strings and image thumbnails, used by the Encyclopedia of Life v2 (Parr et al. 2014, http://eol.org). See https://github.com/jhpoelen/eol-globi-data/issues/274 and https://github.com/EOL/tramea/issues/366 for discussion threads. Taxon hierarchy providers include, but are not limited to, Integrated Taxonomic Information System (ITIS, http://itis.gov) and World Register of Marine Species (WoRMS, http://marinespecies.org).</p>
Repackaged Encyclopedia of Life (EOL) Dynamic Hierarchy hash://sha256/f91877189f3cd14f4066b16b693a5f93105fd23b881b0825d6781be6ac674b88 hash://md5/18ca6625cf1d24093dc104851ab5722b
<p>This publication contains a repackaged publication of the Encyclopedia of Life Dynamic (Taxon) Hierarchy v2.1 as retrieved on 8 July 2024 via https://editors.eol.org/uploaded_resources/00a/db4/7b-57ed-4f6b-8f66-83bfdb5120e8 .</p> <p>File included in this publication:</p> <p>1. dh21.zip - original file retrieved via https://editors.eol.org/uploaded_resources/00a/db4/7b-57ed-4f6b-8f66-83bfdb5120e8</p> <p>2. taxon.tab.gz - extracted, and gzipped, taxon.tab as extracted from dh21.zip</p> <p>3. taxon-first10.tab - first 10 lines of taxon.tab.gz</p> <p>3. meta.xml - as extracted from dh21.zip</p> <p>4. prov.nq - log documenting the provenance of dh21.zip</p>
EncycNet: A Knowledge Graph of Historical German Encyclopedias
<p><strong>EncycNet</strong> is an automatically constructed knowledge graph that uses historical German encyclopedias as a data source. You can read more about the project here: <a href="https://encycnet.github.io/">https://encycnet.github.io/</a></p> <p>The first version of the graph (0.1) contains <em>Meyers Großes Konversations-Lexikon </em>(1905), with 5 more encyclopedias to be added in 2024. Formatted in RDF Turtle.</p> <p>The second version of the graph (0.2) contains 3 encyclopedias in total, meaning triples were changed to quads (see code example at <a href="https://github.com/EncycNet/Encyc-Relations">https://github.com/EncycNet/Encyc-Relations</a>):</p> <ul> <li><em>Meyers Großes Konversations-Lexikon </em>(1905)</li> <li><em>Herders Conversations-Lexikon</em> (1854)</li> <li><em>Brockhaus Bilder-Conversations-Lexikon</em> (1837)</li> </ul> <p>The second version also handled some bug fixes concerning wrongly matched Wikidata entries.</p> <p>The third version (0.3) now additionally contains</p> <ul> <li><em>Brockhaus Conversations-Lexikon oder kurzgefaßtes Handwörterbuch</em> (1809)</li> </ul> <p>It also impoved some namespace issues and introduced <a href="https://www.wikidata.org/wiki/Property:P8371">P8371</a> as encyclopedic refeferences in this case.</p>
GREEN-DB: Genomic Regulatory Elements ENcyclopedia
<p>GREEN-DB is a comprehensive collection of 2.4 million regulatory elements in the human genome collected from previously published databases, high-throughput screenings and functional studies. Regulatory regions are classified as enhancers, promoters, silencers, bivalent and information on the controlled gene(s), tissue(s) and associated phenotype(s) are provided for each element when possible. We also calculated a variation constraint metric (range 0-1) for these regulatory regions and showed that genes controlled by constrained regions are enriched for disease-associated genes and essential genes from mouse knock-out screenings.</p> <p>The database also includes information from ENCODE TFBS and DNase peaks; ultra-conserved non-coding elements (UCNE), super-enhancers (dbSuper) and TAD domains (TAD-KB).</p> <p>This release includes 5 files:</p> <ul> <li>GREEN-DB_v2.5.db.gz: The full database in SQLite format</li> <li>GRCh37_GREEN-DB.bed.gz[.csi]: A indexed BED file using GRCh37 genome coordinates describing the regulatory regions and associated information useful for variant annotations (controlled genes, closest gene/TSS, constraint metric).</li> <li>GRCh38_GREEN-DB.bed.gz[.csi]: A indexed BED file using GRCh38 genome coordinates describing the regulatory regions and associated information useful for variant annotations (controlled genes, closest gene/TSS, constraint metric).</li> </ul> <p>To annotate a VCF file with information from GREEN-DB you can use the bed files and our tool GREEN-VARAN (<a href="https://github.com/edg1983/GREEN-VARAN">https://github.com/edg1983/GREEN-VARAN</a>).</p> <p>For more information on the GREEN-DB please refer to our publication (<a href="https://doi.org/10.1101/2020.09.17.301960">https://doi.org/10.1101/2020.09.17.301960</a>) and to online documentation (<a href="https://green-varan.readthedocs.io/en/latest/">https://green-varan.readthedocs.io/en/latest/</a>)</p> <p>GREEN-DB is free to use for academic users, please refer to the attached LICENSE file.</p> <p> </p> <p><strong>Changes from the previous version:</strong></p> <p>- We fixed an issue with alias symbols conversion that caused a small fraction of region-gene links to point to the wrong gene</p> <p>- Due to the problem above, we removed any region-gene link where the region and the controlled gene were located on different chromosomes</p> <p>- GREEN-DB now includes also TAD domain information from TAD-KB (http://dna.cs.miami.edu/TADKB/) and region-gene interactions are now annotated for occurrence within the same TAD</p> <p>- Better constraint metric model that now takes into account overlap with exonic regions</p> <p>- In addition to the closest gene, an annotation for the closest TSS and its distance is now provided </p>
Encyclopedia [IO Bijapur 453]
<ul> <li><strong>Encyclopedia.</strong></li> <li><strong>This manuscript is now IO Bijapur 453 </strong><strong>in the India Office collections.</strong></li> <li><strong>[metadata:</strong><a href="https://de.wikipedia.org/wiki/Otto_Loth"> <strong>Otto Loth, </strong></a><strong><em><a href="http://doi.org/10.5281/zenodo.3923636">A Catalogue of the Arabic Manuscripts in the Library of the India Office</a></em>, (volume 1), no. 1028 here with further notations and hyperlinks]</strong>.</li> </ul> <p><strong>ENCYCLOPEDIA</strong>.</p> <p> </p> <p><a href="https://archive.org/details/catalogueofarabi01greauoft/page/284/mode/2up"><strong>1028</strong></a>.</p> <p>B453. Size 7<sup>1/2</sup> in. by 5 in.; foll. 12. Twenty-five and twenty-three lines in a page.</p> <p>Foll. 5-12. An encyclopedic treatise, by ḤABÎB ALLAH MÎRZÂ JÂN SHÎRÂZÎ (d. A.H. 994), written for a friend named Muḥammad (سمّی حبیب الله صلعم).</p> <p>It gives specimens of nine sciences, with critical remarks on them; viz., 1.البحث الاول من التفسیر; 2. المعانی; 3. البیان; 4. الاصول; 5. الکلام; 6. المنطق; 7. العلم الطبیعی; 8. الالهی; 9. الهیئة.</p> <p>Begins: جل و علا من تحیر عقول العارفین فی کنه جماله.</p> <p>Written in a good Nasta’lîḳ hand, but without diacritical points. Long notes on the margin. Dated A.H. 1000.</p> <p>It is preceded by-</p> <p>Foll. 1-4. A Commentary on the verse of the Koran, Sû, 2, 256; styled in the conclusion الرسالة الشریفة لحضرت حافظ کویکری (!).</p> <p>Begins: الله لا اله الا هوالله اسم عربی الخ.</p> <p>Legibly written.</p> <p> </p>
Twenty-two Historical Encyclopedias (Original XML Encoding)
<p>This dataset accompanies the publication of "Twenty-two Historical Encyclopedias Encoded in TEI: a New Resource for the Digital Humanities".</p> <p>Licensed as CC-BY.</p>
Database file for Encyclopedia of Finite Graphs: Simple connected graphs, n<=10
<p>This initial release contains all simple connected graphs of order n<=10, and a collection of integer invariants. Up to order n<=6, there is a collection of "special" invariants that are stored in a custom table (see main project for details).</p>
Encyclopedia [IO Islamic 1622] النقایة, fragment of
<ul> <li><strong>Encyclopedia. </strong>النقایة</li> <li><strong>This manuscript is now IO Islamic 1622 </strong><strong>in the India Office collections.</strong></li> <li><strong>[metadata:</strong><a href="https://de.wikipedia.org/wiki/Otto_Loth"> <strong>Otto Loth, </strong></a><strong><em><a href="http://doi.org/10.5281/zenodo.3923636">A Catalogue of the Arabic Manuscripts in the Library of the India Office</a></em>, (volume 1), no. 1029 here with notations and hyperlinks]</strong>.</li> </ul> <p><strong><a href="https://archive.org/details/catalogueofarabi01greauoft/page/284/mode/2up">1029</a></strong>.</p> <p>1622. Size 9 in. by 4<sup>3/4</sup> in.; foll. 50. Eight lines in a page.</p> <p>A fragment of an encyclopedic treatise on the Muḥammadan Sciences, which, from the headings, appears to be <a href="http://worldcat.org/identities/lccn-n80081636/">SUYÛṬÎ</a>’S (d. A.H. 911) النقایة. See regarding this work, Ḥ. Kh. vi. 372; <a href="https://findit.library.yale.edu/catalog/digcoll:2845405">Cat. Mus. Brit</a>. 213 [<strong>ed note</strong>:= CCCCXXXII, p. 213; Add. 7523 Rich]; <a href="http://worldcat.org/identities/lccn-nr88004247/">Flügel</a>, <a href="https://doi.org/10.5281/zenodo.4727583">Hdss. Wien, i</a>. 22.</p> <p>Well written, but damaged and in disorder. Both the beginning and end are wanting. Foll. 1-7 are really the last of this fragment, and fol. 8 begins in what would be the first paragraph of the treatise. The last leaf gives the conclusion of a <em>Persian</em> tract.</p> <p>[<a href="https://doi.org/10.5281/zenodo.4085990">Johnson</a>.]</p> <p> </p>
FAIRmat Tutorial 5: NOMAD Encyclopedia
<p>The NOMAD Encyclopedia is a web-based public infrastructure that provides this materials-oriented view on the NOMAD Repository & Archive. In this tutorial we will discuss how to navigate the materials space using the Encyclopedia GUI as well as the more advanced tasks which are possible using the Encyclopedia API.</p> <p> </p> <p><strong>Disclaimer: </strong>NOMAD is being continuously developed based on input and feedback from the scientific community. Hence the features, services or interface may have changed since the time of recording of this video. For up-to-date information please consult our latest tutorials and the NOMAD documentation <a href="https://nomad-lab.eu/prod/v1/docs/">https://nomad-lab.eu/prod/v1/docs/</a></p>
Encyclopedia of Marine Life of Britain and Ireland Resource: Encyclopedia of Marine Life of Britain and Ireland
Guide to marine life of Britain and Ireland. <p></p>http://www.habitas.org.uk/marinelife/
TaiEOL: Taiwan Encyclopedia of Life - DwCA
<p></p>https://taieol.tw/ Academia Sinica
FIGURE 1 in The Encyclopedia of Life vs. the Brochure of Life: Exploring the relationships between the extinction of species and the inventory of life on Earth
FIGURE 1. Total number of species described after 25 years, expressed as a percentage of the remaining total number of species (y axis), for different rates of discovery of new species (x axis; current values, 1 = 10,000 species/year). Three estimates of the total number of species living in year 2000 are considered: a very low estimate, 3 million (white rectangles), a middle estimate, 10 million (grey rectangles) and a high estimate, 100 million (black rectangles); in all cases, the number of species described up to year 2000 is estimated to be 1.5 million. Only if total biodiversity is very low and rates of discovery are 10 times current values (or higher), can the Encyclopedia of Life be totally accomplished in 25 years (% Described = 100; note the exclamation symbols, '!').
FIGURE 2 in The Encyclopedia of Life vs. the Brochure of Life: Exploring the relationships between the extinction of species and the inventory of life on Earth
FIGURE 2. Relationships between the rate of description of new species ('Discovery', left axis; current values, 1 = 10,000 species/ year; logarithmic scale), the time needed to describe all the species at that rate ('Time', right axis; in years; logarithmic scale), and the number of species that are lost to extinction during this time ('Biodiversity loss', horizontal axis; logarithmic scale). Three estimates of the total number of species living in year 2000 are considered: a very low estimate, 3 million (thin curves); a middle estimate, 10 million (dotted curves), and a high estimate, 100 million (thick curves). For each case, two curves are shown: discovery rates (D curves, like, D3 for 3 million species or D100 for 100 million) and time to complete the Encyclopedia of Life (i.e., time to extinction of the species indicated in the horizontal axis; t curves, like t10 for 10 million). Vertical and horizontal dotted lines are used in the figure as an example and evidence that, if the total number of species existing in year 2000 is estimated in 10 million (dotted curves), ca. 106 species (point a) will be lost during the completion of the Encyclopedia (in a time span above 100 years, point b) if we work at a rate of discovery close to 7 times the current one (point c).
Encyclopedia of Life Taxonomy Patch for Dynamic Hierarchy Version 3.1
<p>A taxonomic patch for the <a href="https://eol.org/docs/eol-dynamic-hierarchy">EOL Dynamic Hierarchy</a>.</p>
Digital Encyclopedia of British Sociability
<p>The DIGIT.EN.S encyclopedia is comprised of 200 entries and an anthology of primary sources on eighteenth- and early nineteenth-century sociability. The entries are divided in 5 categories: people, places, practices, objects and concepts. The entries focus mainly on British sociability but explore also the circulation and transfers of models, forms, practices and values of sociability throughout Europe and in the colonial worlds.</p>
Digital Encyclopedia of British Sociability
Open the record for dataset details and reuse information.
Comparative Validation of the D. melanogaster Encyclopedia of DNA Elements Transcript Models
GEO Series GSE44612. Drosophila mojavensis; Drosophila yakuba; Drosophila ananassae; Drosophila simulans; Drosophila virilis; Drosophila elegans; Drosophila biarmipes; Drosophila melanogaster; Drosophila pseudoobscura; Drosophila eugracilis; Drosophila takahashii; Drosophila ficusphila; Drosophila kikkawai; Drosophila bipectinata; Drosophila rhopaloa. 93 samples. Type: Expression profiling by high throughput sequencing.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.